Domain-Conditioned Spatial–Frequency Modulation for Joint Underwater Image Enhancement
Shijian Zheng, Haibo Lin, Yadong Yang, Zeyang Liu, Xiancun Zhou, Guihui Li, Yue TengJoint training across heterogeneous paired underwater image enhancement (UIE) datasets broadens the range of supervised training conditions but requires a shared model to learn how the mapping from an observation to its reference varies across source domains while preserving color fidelity and spatial detail. We study a balanced source-known joint-training setting, in which one shared model is trained across multiple paired datasets and DCSM-Net receives the dataset-of-origin index during both training and evaluation. We propose a Domain-Conditioned Spatial–Frequency Modulation Network (DCSM-Net) for this setting. DCSM-Net integrates three lightweight operations within a hybrid encoder–decoder. A Spatial-Guided (SG) block performs high-resolution local refinement and applies channel-asymmetric latent feature modulation initialized with a wavelength-motivated prior. A Domain Token Module (DTM) combines an image-context descriptor with a learnable source-domain embedding to generate a conditioning token. Conditional Wavelet Prompt Blocks (WPBs) use this token to modulate skip features, applying channel-wise affine transformation to the low-frequency LL component and residual prompt routing to the directional LH, HL, and HH components. Under joint training on UIEB, LSUI, and EUVP, DCSM-Net improves the average PSNR of the same backbone from 25.51 to 26.75 dB with only 0.13 M additional parameters. Under this source-known protocol, it achieves 25.13 dB on UIEB-Test90, 31.69 dB on LSUI-Test150, and 23.42 dB on EUVP-Test100, obtaining the best PSNR, SSIM, and LPIPS among the compared methods. In an additional DCSM-Net analysis, we assess seed sensitivity across seeds 7, 17, and 27 and examine a source-agnostic mean-embedding fallback; the main comparisons otherwise use seed 7. Hierarchical ablations show gains from high-resolution spatial refinement, statistics-conditioned wavelet processing, and source-domain token conditioning.