Loading...
TrafficInternVL: Spatially-Guided Fine-Tuning with Caption Refinement for Fine-Grained Traffic Safety Captioning and Visual Question Answering
Author(s)
Phimsiri, Sasin
Sunpawatr, Sarut
Cherdchusakulchai, Riu
Kiawjak, Pornprom
Tosawadi, Teepakorn
Tungjitnob, Suchat
Trairattanapa, Visarut
Vatathanavaro, Supawit
Utintu, Chaitat
Saetan, Worawit
Kongsawat, Nathamon
Borisuitsawat, Phawat
Mahakijdechachai, Kasisdis
Su-Inn, Nitipan
Thamwiwatthana, Ek
Suttichaya, Vasin
Date Issued
January 1, 2025
Type
Conference Paper
Abstract
Fine-grained traffic understanding requires both detailed visual descriptions and precise answers to safety-critical questions. We present TrafficInternVl, a framework for fine-grained traffic safety description and question answering, developed for AI City Challenge 2025 Track 2. Our approach is based on the InternVL3-38B vision-language model and integrates four key components: (1) spatially guided visual prompting via bounding-box-based cropping and rendering; (2) Adaptive view selection protocols; (3) low-rank adaptation (LoRA) fine-tuning, updating only 1% of model parameters; and (4) caption refinement for intra-scene consistency. Our model achieves a Caption Score of 32.75 (BLEU-4, METEOR, ROUGE-L, CIDEr averaged) and a VQA accuracy of 83.08 %. Code, prompts, and LoRA weights are released at https://github.com/ARV-MLCORE/TrafficInternVL
Citation
Proceedings 2025 IEEE Cvf International Conference on Computer Vision Workshops Iccv W 2025, 5358-5365, 2025
