Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kyem, Blessing Agyei, Asamoah, Joshua Kofi, Dontoh, Anthony, Aboah, Armstrong
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910116554473472
author Kyem, Blessing Agyei
Asamoah, Joshua Kofi
Dontoh, Anthony
Aboah, Armstrong
author_facet Kyem, Blessing Agyei
Asamoah, Joshua Kofi
Dontoh, Anthony
Aboah, Armstrong
contents General-purpose vision-language models demonstrate strong performance in everyday domains but struggle with specialized technical fields requiring precise terminology, structured reasoning, and adherence to engineering standards. This work addresses whether domain-specific instruction tuning can enable comprehensive pavement condition assessment through vision-language models. PaveInstruct, a dataset containing 278,889 image-instruction-response pairs spanning 32 task types, was created by unifying annotations from nine heterogeneous pavement datasets. PaveGPT, a pavement foundation model trained on this dataset, was evaluated against state-of-the-art vision-language models across perception, understanding, and reasoning tasks. Instruction tuning transformed model capabilities, achieving improvements exceeding 20% in spatial grounding, reasoning, and generation tasks while producing ASTM D6433-compliant outputs. These results enable transportation agencies to deploy unified conversational assessment tools that replace multiple specialized systems, simplifying workflows and reducing technical expertise requirements. The approach establishes a pathway for developing instruction-driven AI systems across infrastructure domains including bridge inspection, railway maintenance, and building condition assessment.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08212
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment
Kyem, Blessing Agyei
Asamoah, Joshua Kofi
Dontoh, Anthony
Aboah, Armstrong
Computer Vision and Pattern Recognition
General-purpose vision-language models demonstrate strong performance in everyday domains but struggle with specialized technical fields requiring precise terminology, structured reasoning, and adherence to engineering standards. This work addresses whether domain-specific instruction tuning can enable comprehensive pavement condition assessment through vision-language models. PaveInstruct, a dataset containing 278,889 image-instruction-response pairs spanning 32 task types, was created by unifying annotations from nine heterogeneous pavement datasets. PaveGPT, a pavement foundation model trained on this dataset, was evaluated against state-of-the-art vision-language models across perception, understanding, and reasoning tasks. Instruction tuning transformed model capabilities, achieving improvements exceeding 20% in spatial grounding, reasoning, and generation tasks while producing ASTM D6433-compliant outputs. These results enable transportation agencies to deploy unified conversational assessment tools that replace multiple specialized systems, simplifying workflows and reducing technical expertise requirements. The approach establishes a pathway for developing instruction-driven AI systems across infrastructure domains including bridge inspection, railway maintenance, and building condition assessment.
title Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.08212