Residual Conditional Diffusion with Transformer Refinement for Unsupervised Infrared–Visible Image Fusion
Sirui Huang, Lin Tian, Yao ZhangInfrared–visible image fusion aims to integrate thermal target information from infrared images and structural texture information from visible images into a single informative image. Existing deep fusion methods still face challenges in preserving fine textures, maintaining structural consistency, and balancing complementary information under low-light conditions. To address these issues, this paper proposes MRCDFusion, an unsupervised infrared–visible image fusion network based on residual conditional diffusion and Transformer refinement. Specifically, a shared dense encoder is used to extract modality-specific and cross-modal complementary features from infrared and visible images. A Modality-Level Attention Module (MLAM) is then introduced to aggregate strong responses from infrared and visible features and construct modality-aware condition features for guiding the diffusion process. Instead of generating fused features from scratch, the proposed method adopts a base-plus-residual diffusion strategy, in which base features preserve global structures and residual diffusion enhances local details. A deterministic noise strategy is further introduced to improve inference reproducibility. The diffusion-enhanced features are refined by a window Transformer and depthwise separable convolutions, followed by gated feature fusion and progressive image reconstruction. Experiments are conducted primarily on the low-light LLVIP dataset, while FMB, TNO, and RoadScene are used for zero-shot cross-dataset evaluation without additional fine-tuning. The results show that MRCDFusion achieves particularly strong performance in gradient- and edge-related metrics while remaining competitive in visual information fidelity and cross-modal correlation metrics. Ablation studies verify the effectiveness of the main components, and downstream detection and auxiliary segmentation experiments further demonstrate the potential utility of the fused representations for subsequent visual perception tasks.