Unified Multimodal-Multitask Learning for Vehicle Damage Assessment in Insurance Applications

Abstract

Automated vehicle damage assessment requires both precise localization and clear textual reporting. While existing methods typically treat these as separate tasks, the trade-offs of unified multimodal-multitask learning in this domain remain underexplored. This paper conducts a comparative study between a unified vision-language framework, Generative Region-to-Text Transformer (GRiT), and single-task baselines derived from GRiT by isolating the detection and captioning components. We adapt GRiT to the insurance domain using a dataset enriched with vehicle part annotations and structured damage descriptions. Experimental results demonstrate that the unified model achieves competitive detection performance (F1-score: 0.54), slightly outperforming the detection baseline model. Crucially, it significantly surpasses the caption baseline model in description quality (METEOR: 0.75, ROUGE: 0.70, BLEU: 0.46), confirming that object-level visual grounding is essential for accurate reporting. These findings indicate that unified multimodal learning enhances semantic interpretation without compromising localization accuracy, offering a promising direction for automated insurance workflows.

Description

Keywords

Image Captioning, Multimodal Learning, Multitask Learning, Object Detection, Vehicle Damage Assessment

Citation

Proceedings 23rd International Joint Conference on Computer Science and Software Engineering Jcsse 2026, 249-254, 2026

Collections

Endorsement

Review

Supplemented By

Referenced By