AOPQ-Net Acoustic–Optical Proposal Query Network for Underwater Multimodal Object Detection
Yanze Lu, Zhengyan Zhang, Shuoshuo Ding, Haochen Hu, Chih-Yung Wen, Tiedong ZhangOptical cameras and imaging sonars are widely used sensors in autonomous underwater vehicles. However, their different imaging mechanisms introduce substantial cross-modal discrepancies in the acquired data. In addition, underwater optical images are often degraded by low illumination, scattering, and turbidity, whereas sonar images commonly suffer from speckle noise and low spatial resolution. As a result, object detection based on a single optical or acoustic modality is often insufficient in challenging underwater environments. To address this problem, this paper proposes an acoustic–optical fusion network for underwater object detection, termed an Acoustic–Optical Proposal Query Network (AOPQ-Net). First, a Sonar Position Encoding (SPE) module is designed to explicitly encode the geometric priors in sonar images. Second, a Bi-directional Discrepancy-aware Spatial Alignment (BDSA) module is introduced to alleviate spatial misalignment between the two modalities at the feature level. Third, a Proposal Query Transformer (PQT) module performs target-oriented cross-modal interaction at the proposal level. Furthermore, this study constructs a dedicated dataset for underwater acoustic–optical fusion object detection, named Haiqin Underwater Fusion (HUF), and conducts systematic experiments on this dataset. The experimental results show that AOPQ-Net outperforms single-modality baselines and representative multimodal fusion methods in both optical and acoustic image spaces, which demonstrate the effectiveness of the proposed method.