DOI: 10.1061/jccee5.cpeng-7383 ISSN: 0887-3801

Comparative Analysis of Multimodal AI Models for Automated Construction Safety Monitoring and Reporting

Trevor Neece, Alessandro Fascetti

Abstract

The construction industry accounts for a significant portion of workplace fatalities in the United States. Health and safety hazards frequently go unrecognized, potentially leading to injuries and death. However, traditional computer vision methods are often limited to specific object detection tasks [e.g., personal protective equipment (PPE) compliance] and their adoption introduces challenges in interpreting complex, contextual hazard scenarios. Furthermore, the scarcity of annotated accident data hinders the development of fully supervised models. Consequently, this manuscript formally characterizes the performance of pretrained multimodal Artificial Intelligence (AI) for automating hazard recognition. Furthermore, the performance of such models is also tested on game engine-based synthetic data, to investigate whether it can be leveraged for data augmentation in order to overcome the scarcity of real-world training examples. To address these questions, this study formally evaluates the effectiveness of AI models over multiple iterations of the models’ architecture for the domain-specific application of automated construction hazard assessment from multimodal inputs. The research presented herein leverages publicly available pretrained vision-language models and evaluates their performance across three core tasks: binary hazard alerts, detailed hazard understanding, and automated citation generation. A dataset of annotated real-world construction images is presented to benchmark the performance of models. The dataset is also made public to provide a standardized means for comparing performance across different models. Results reported herein indicate that leading AI models can achieve high, although not perfect, recognition accuracy without fine-tuning. However, model performance varies across hazard categories, and newer generations show diminishing returns. Consequently, the study introduces and validates high-fidelity, game engine-based synthetic images as a solution. By generating high-fidelity digital images of dangerous activities, synthetic data can help bridge the scarcity of real-world hazard examples. This approach increases opportunities to improve these models with task-specific data, supporting continued progress beyond current limitations.

More from our Archive