Foreground Segmentation and Multi-Module CLIP-ReID for Transmission Tower Re-Identification
Changyu Li, Junsheng Lin, Jinchao Guo, Xinlei Zhang, Qianming Wang, Zhenbing ZhaoTransmission tower re-identification (ReID) supports the removal of redundant unmanned aerial vehicle (UAV) inspection images but remains challenging because towers have similar structures, viewpoints vary significantly, and backgrounds are complex. This paper proposes a two-stage framework combining YOLO11n-Seg foreground segmentation with a Contrastive Language–Image Pre-training (CLIP)-based ReID model using a ResNet-50 visual backbone enhanced by foreground-aware cropping (FG-CROP), a depthwise-separable Convolutional Block Attention Module (DS-CBAM), and cross-scale gated fusion (CSGF). YOLO11n-Seg was trained on an independently collected dataset with no image or tower-identity overlap with the re-identification data, and the selected validation checkpoint was used to process all 1772 ReID images. The ReID dataset contains 1437 training images from 464 identities, while the test set contains 106 query images and 229 gallery images from 106 disjoint identities. Across five random seeds, the complete framework achieved 62.34 ± 0.87% mAP and 60.20 ± 1.42% Rank-1, compared with 56.90 ± 0.89% mAP and 54.52 ± 1.40% Rank-1 for the black-background CLIP-ReID baseline with a ResNet-50 visual backbone. The selected YOLO11n-Seg checkpoint achieved a 97.26% mask mAP50 and 71.61% mask mAP50-95 on the independent validation set.