DOI: 10.3390/photonics13100899 ISSN: 2304-6732

FIRGAN: High-Fidelity Infrared–Visible Image Fusion for UAV Detection Tasks

Huiji Wang, Dandan Huang, Zhi Liu, Zhichao Han, Xingzhao Wang, Jian Fan, Rui Zhang

Infrared–visible image fusion aims to generate informative representations for both visual perception and downstream vision tasks. However, existing methods often prioritize visual quality while overlooking task adaptability, optimization efficiency, and effective spatial-frequency representation learning. To address these issues, we propose FIRGAN, a high-fidelity multi-modal fusion framework for downstream detection applications that jointly optimizes fusion quality and detection-relevant feature representation. Specifically, an SFFM is introduced to exploit MSCN-derived statistical priors for adaptive spectrum enhancement. A Dual-branch Fusion Network is designed to collaboratively model local structural details and global semantic dependencies through complementary CNN and Transformer branches. In addition, an IQA-guided optimization strategy performs quality-aware patch selection to improve training efficiency and robustness. Extensive experiments on public infrared–visible fusion datasets demonstrate that the proposed method achieves competitive visual quality and quantitative performance while providing more discriminative representations for downstream UAV detection tasks. These results verify the effectiveness of the proposed framework for multi-modal fusion for downstream detection applications.