Unified Multimodal-Multitask Learning for Vehicle Damage Assessment in Insurance Applications
| dc.contributor.author | Phuengpanyaloet, Wongsapat | |
| dc.contributor.author | Pasupa, Kitsuchart | |
| dc.contributor.author | Angsarawanee, Thanatwit | |
| dc.contributor.author | Chetprayoon, Panumate | |
| dc.contributor.author | Sakdejayont, Theerat | |
| dc.date.accessioned | 2026-08-06T10:53:21Z | |
| dc.date.available | 2026-08-06T10:53:21Z | |
| dc.date.issued | 2026-01-01 | |
| dc.description.abstract | Automated vehicle damage assessment requires both precise localization and clear textual reporting. While existing methods typically treat these as separate tasks, the trade-offs of unified multimodal-multitask learning in this domain remain underexplored. This paper conducts a comparative study between a unified vision-language framework, Generative Region-to-Text Transformer (GRiT), and single-task baselines derived from GRiT by isolating the detection and captioning components. We adapt GRiT to the insurance domain using a dataset enriched with vehicle part annotations and structured damage descriptions. Experimental results demonstrate that the unified model achieves competitive detection performance (F<inf>1</inf>-score: 0.54), slightly outperforming the detection baseline model. Crucially, it significantly surpasses the caption baseline model in description quality (METEOR: 0.75, ROUGE: 0.70, BLEU: 0.46), confirming that object-level visual grounding is essential for accurate reporting. These findings indicate that unified multimodal learning enhances semantic interpretation without compromising localization accuracy, offering a promising direction for automated insurance workflows. | |
| dc.identifier.citation | Proceedings 23rd International Joint Conference on Computer Science and Software Engineering Jcsse 2026, 249-254, 2026 | |
| dc.identifier.doi | 10.1109/JCSSE68839.2026.11596759 | |
| dc.identifier.other | 2-s2.0-105045258724 | |
| dc.identifier.uri | https://dspace.kmitl.ac.th/handle/123456789/17542 | |
| dc.source | Proceedings 23rd International Joint Conference on Computer Science and Software Engineering Jcsse 2026 | |
| dc.subject | Image Captioning | |
| dc.subject | Multimodal Learning | |
| dc.subject | Multitask Learning | |
| dc.subject | Object Detection | |
| dc.subject | Vehicle Damage Assessment | |
| dc.title | Unified Multimodal-Multitask Learning for Vehicle Damage Assessment in Insurance Applications | |
| dc.type | Conference Paper |
