CVIGAN: Contrastive Visible-to-Infrared Generative Adversarial Network with ConvNeXtV2 Multi-Scale Frequency-Aware Attention
Chengyi Wang, Jin Lin, Ting Liu, Xiangyi Lu, Ziheng Wang, Kai ZhengExisting unsupervised visible-infrared image generation methods struggle to produce reasonable target thermal appearances. Consequently, it is challenging to generate infrared images with consistent structural details. To address this limitation, this paper proposes CVIGAN: Contrastive Visible-to-Infrared Generative Adversarial Network with ConvNeXtV2 Multi-Scale Frequency-Aware Attention. Specifically, ConvNeXtV2 is leveraged to construct the encoder of the generator, and a multi-scale frequency-aware attention feature enhancement block (MFEB) is integrated to enhance the feature representation capability. Secondly, multi-scale discriminators are elaborately designed to simultaneously capture local fine-grained details and global structural information of target-domain images. Finally, a composite loss function is formulated by combining multi-layer contrastive loss, unidirectional cycle-consistency loss, identity loss, and least-squares adversarial loss, which enforces cross-domain feature alignment and restricts the quality of generated outputs. The effectiveness of the proposed method is validated through experiments on the unpaired AVIID3, M3FD, and a self-built dataset. The generated infrared images achieve notably enhanced image quality. In addition, the robustness of CVIGAN is validated through cross-scene and cross-resolution generalization experiments, where favorable performance is maintained under unseen scenes and high-resolution input conditions. Thus, the generated infrared images are validated to effectively enhance downstream infrared object detection tasks, offering a more efficient solution for infrared image acquisition and analysis in real-world applications.