DOI: 10.3390/su18189656 ISSN: 2071-1050

Auditable and Abstention-Aware Vehicle-Level Evaluation of Vision–Language Model Configurations for Vehicle-Damage Assessment in Intelligent Transportation Maintenance

Yiyang Shao, Lina Mao, Guiliang Zhou

Vehicle-damage assessment in intelligent transportation maintenance often requires one decision from several views of the same vehicle. We developed a paired, auditable protocol for comparing two complete vision–language model configurations at the vehicle level. The protocol uses source_vehicle_id as the sampling unit; isolates Gold and adjudication records during inference; retains structured, non-binary outputs; and combines classification, coverage–risk, paired-bootstrap, and processing-time analyses. The frozen dataset contained 160 vehicles and 355 deduplicated media items, with 119 State 0, 41 State 1, and no State 9 cases. Thus, the primary analysis could not estimate performance for insufficient-reference-evidence cases. Under the frozen configuration-specific inputs; instructions; and preprocessing, software, and output constraints, Models A and B achieved Macro-F1 scores of 0.8645 and 0.1923. The paired difference was 0.6722, with a 95% bootstrap confidence interval of [0.5878, 0.7475] from 10,000 vehicle-level resamples. Binary coverage was 0.9438 for Model A and 0.3938 for Model B, and selective risk was 0.0464 and 0.4444, respectively. End-to-end single-arm processing took 157.75 s for Model A and 654.09 s for Model B under the recorded execution conditions. These findings compare the two complete configurations and do not isolate an architectural effect. Energy use, carbon emissions, repair-material savings, and life-cycle outcomes were not measured. The protocol provides a basis for future matched-task studies of human-in-the-loop maintenance workflows.