DOI: 10.3390/jmse14191817 ISSN: 2077-1312

RGB-Derived Multi-Representation Texture Enhancement for Underwater Object Detection

Longcan Cheng, Xiaomin Wang

Underwater object detection is hindered by haze, color distortion, and weak object boundaries. This study develops MTUDNet, a YOLOv8n-based framework that combines a dehazed appearance, RGB-derived pseudo-depth, and depth-guided texture cues with frequency-spatial-channel attention and a boundary-adaptive box loss. The contribution lies in their task-oriented integration and in the boundary-discrepancy regression design; the dehazing and monocular depth estimators are adopted from prior work. Relative to YOLOv8n, MTUDNet increases mAP@0.5:0.95 by 6.6, 5.8, and 11.7 percentage points on DUO, UODD, and RUOD, respectively. It obtains 53.4% mAP@0.5:0.95 on UODD with 4.3 million detector parameters. Comparisons with U-DECN and GCC-Net show that this compact detector does not lead every accuracy metric. The 4.3 million parameters characterize the detector only; full-pipeline latency is not reported. Future work will measure complete inference time and assess pseudo-depth reliability under changing underwater conditions.