Unified Multimodal-Multitask Learning for Vehicle Damage Assessment in Insurance Applications

dc.contributor.authorPhuengpanyaloet, Wongsapat
dc.contributor.authorPasupa, Kitsuchart
dc.contributor.authorAngsarawanee, Thanatwit
dc.contributor.authorChetprayoon, Panumate
dc.contributor.authorSakdejayont, Theerat
dc.date.accessioned2026-08-06T10:53:21Z
dc.date.available2026-08-06T10:53:21Z
dc.date.issued2026-01-01
dc.description.abstractAutomated vehicle damage assessment requires both precise localization and clear textual reporting. While existing methods typically treat these as separate tasks, the trade-offs of unified multimodal-multitask learning in this domain remain underexplored. This paper conducts a comparative study between a unified vision-language framework, Generative Region-to-Text Transformer (GRiT), and single-task baselines derived from GRiT by isolating the detection and captioning components. We adapt GRiT to the insurance domain using a dataset enriched with vehicle part annotations and structured damage descriptions. Experimental results demonstrate that the unified model achieves competitive detection performance (F<inf>1</inf>-score: 0.54), slightly outperforming the detection baseline model. Crucially, it significantly surpasses the caption baseline model in description quality (METEOR: 0.75, ROUGE: 0.70, BLEU: 0.46), confirming that object-level visual grounding is essential for accurate reporting. These findings indicate that unified multimodal learning enhances semantic interpretation without compromising localization accuracy, offering a promising direction for automated insurance workflows.
dc.identifier.citationProceedings 23rd International Joint Conference on Computer Science and Software Engineering Jcsse 2026, 249-254, 2026
dc.identifier.doi10.1109/JCSSE68839.2026.11596759
dc.identifier.other2-s2.0-105045258724
dc.identifier.urihttps://dspace.kmitl.ac.th/handle/123456789/17542
dc.sourceProceedings 23rd International Joint Conference on Computer Science and Software Engineering Jcsse 2026
dc.subjectImage Captioning
dc.subjectMultimodal Learning
dc.subjectMultitask Learning
dc.subjectObject Detection
dc.subjectVehicle Damage Assessment
dc.titleUnified Multimodal-Multitask Learning for Vehicle Damage Assessment in Insurance Applications
dc.typeConference Paper

Files

Collections