DOI: 10.1177/03611981261485241 ISSN: 0361-1981

Multi-Modal Deep Learning Approach for Comprehensive Pavement Condition Evaluation Using Vision-Language Models

Neema Jakisa Owor, Yaw Adu-Gyamfi

Timely and accurate assessment of pavement conditions is essential for roadway safety, infrastructure longevity, and cost-effective maintenance planning. Traditional methods rely heavily on manual inspections, which are labor-intensive, subjective, and limited in scalability. Recent advances in vision-language models (VLMs) offer a new opportunity to automate this process by integrating visual analysis with natural language generation. In this study, we propose a novel framework that leverages fine-tuned VLMs to automate pavement condition evaluation by analyzing visual inputs and generating descriptive language outputs. The proposed framework detects key components of pavement distress, including type, severity, and extent, and generates both numerical Pavement Condition Index (PCI) scores and detailed textual descriptions. We fine-tuned three state-of-the-art VLMs—Qwen2.5-VL, InternVL, and LLaVA-Next—using a custom dataset annotated with distress type, severity, and extent, along with polygons annotations, and natural language descriptions. To enhance performance, we implemented a targeted fine-tuning approach incorporating instructional prompt diversification, bounding box grounding, PCI scale normalization, and extended low-rank adaptation across transformer and cross-modal layers. Among the models, Qwen2.5-VL achieved the highest performance, with a METEOR score of 0.8989 and 0.9023 ROUGE-L for descriptive accuracy, and a mean absolute error of 1.19 and root mean square error of 1.89 for PCI prediction. This work demonstrates the effectiveness of VLMs for scalable, interpretable, and accurate pavement assessment, offering a practical alternative to traditional inspection methods and paving the way for intelligent infrastructure management systems.