DOI: 10.3390/jimaging12080353 ISSN: 2313-433X

Style-Semantic Disentangled Optical-to-Infrared Translation for Infrared Target Recognition

Lizhuo Liu, Jiawei Niu, Lingxia Mu

Infrared target recognition plays an important role in many real-world applications, but its performance is often constrained by the scarcity of annotated infrared data. To alleviate this issue, optical-to-infrared image translation has been widely explored as a data augmentation strategy by leveraging the abundance of optical images. However, existing approaches typically overlook the intrinsically multimodal nature of optical-to-infrared mapping, leading to insufficient diversity in the synthesized infrared images. Moreover, the lack of effective constraints to preserve semantic fidelity further hampers the practical utility of generated samples for recognition tasks. In this paper, we propose a multimodal style translation framework for infrared target recognition. The proposed framework is built upon a style-semantic disentanglement architecture, which decouples domain-general semantic structures from domain-specific style statistics, thereby enabling flexible recombination of optical content with diverse infrared characteristics. Furthermore, we design a multi-level adaptive loss function that explicitly enforces complementary constraints on structural fidelity and semantic consistency during the translation process. Extensive experiments on two public datasets demonstrate the effectiveness of SSD-VI. On RGB-NIR, it achieves an FID of 46.53 and a KID of 0.0331, while increasing classification accuracy by 6.68 percentage points, from 83.37% to 90.05%. On VEDAI, SSD-VI improves mAP@50 by 0.13 for YOLOv8m and 0.14 for RT-DETR, confirming the value of the generated samples for infrared target recognition.

More from our Archive