Physics-Aware Diffusion Synthesis for Robust Underwater Object Detection
Wenxin Xiao, Xiaowei Zhou, Junyu DongUnderwater object detection remains challenging in adverse aquatic environments, where severe image degradation caused by turbidity, light scattering, color attenuation, and low illumination substantially reduces detection reliability. Although real-world underwater datasets are essential, their limited scale and environmental diversity make it difficult to cover the wide range of degraded conditions encountered in practice. To improve detection robustness without collecting additional annotations, we propose physics-aware diffusion synthesis (PADS), a framework that uses a small set of labeled real images to synthesize diverse physically plausible degraded underwater samples. PADS couples a ControlNet-conditioned latent diffusion generator with a physics-based underwater image-formation model inspired by Jaffe–McGlamery and Akkaynak optics. Semantic masks are first employed to preserve object layout during generation. Meanwhile, water-optics parameters are incorporated through cross-attention to guide the degradation process. In addition, the physical model enforces a color and attenuation consistency loss during training and serves as an SDEdit-style latent prior during synthesis. To further improve localization under degradation, especially for small objects, we introduce a training-only scale-aware focaler–NWD (SA-FNWD) bounding-box loss, which emphasizes normalized Wasserstein distance for small boxes while retaining IoU-based regression for larger objects. Experiments on the MOUD dataset demonstrate that detectors trained with PADS-synthesized data achieve substantially stronger robustness under severe degradation. At the harshest turbidity level, PADS retains 59.3% of clean accuracy compared with 14.9% for the copy–paste-based synthesis method and 14.1% for the pix2pix-based synthesis method. SA-FNWD further improves mAP@0.5:0.95 across degradation severities. These results show that physics-grounded diffusion synthesis provides the main robustness gain, while SA-FNWD offers a complementary small-object localization improvement with no inference overhead.