DOI: 10.1177/03611981261479733 ISSN: 0361-1981

Comparative Study of Monocular Vision and Vision–Language Models for Vessel Height Estimation in Ship–Bridge Collision Prevention

Mahitha Veeramachaneni, Lu Gao, Yunpeng Zhang, Zhe Han, Jingran Sun

This study presents a vision-based framework for real time ship height estimation from monocular images, which targets the urgent maritime safety challenge of ship–bridge collision prevention. By integrating deep learning, monocular depth estimation, and reference object scaling, we introduce a lightweight, scalable solution that estimates vessel height using a single Red–Green–Blue (RGB) image and a known-size reference object. In addition to the proposed method, vision–language models (VLMs) are employed to provide comparative ship height predictions directly from image inputs. The proposed system combines pixel-to-meter transformation and depth ratio estimation to resolve the scale ambiguity inherent in monocular vision. Deployed via an interactive web-based tool, the proposed method was tested on 60 case studies involving commercial vessels and achieved an average error of 7.65% (3.87 m), comparable with the 17.71 m (35.51%) average error of VLMs. This approach shows potential as a cost-effective, sensor-light alternative to traditional systems. However, practical deployment requires improved robustness under varying lighting, resolution, and image quality conditions, as well as more reliable reference object selection.