A Transformer-Based Cascade Fusion Network for Vision-Based Power Line Component Detection
Bing Zhang, Ran Zhang, Nana Huang, Lei YangIn the realm of power grid construction, the accurate inspection of critical components within overhead transmission lines holds paramount significance and also serves as a cornerstone for safeguarding the safety and stability of the power system and improving operational efficiency. However, this task faces many challenges, including varied and complex backgrounds, unstructured features, multi-scale issues, etc. To address the above challenges, a Transformer-based cascade fusion network is presented in this paper for automatic detection of key power line components. Specifically, with the aim of acquiring more potent feature representations, an enhanced backbone network integrating convolutional neural network (CNN) and Transformer architectures is built to balance the effective extraction of global context and local detail features. Further, to mitigate detail loss during feature fusion, a cascade neck unit (CNU) is proposed to strengthen the representation of deep and shallow semantic features and adaptively adjust receptive fields for key power line components of varying sizes. Finally, in order to enhance the accuracy of small-target detection, the normalized Wasserstein distance (NWD) loss is introduced to optimize the model’s performance when dealing with small targets that are frequently encountered in the context of power line component detection. To evaluate the performance of the proposed model, a series of quantitative and qualitative experiments have been conducted for a comprehensive performance evaluation. These experimental results, obtained from a dedicated detection dataset specifically focused on key power line components, demonstrate that the proposed model achieves precision, recall, mean average precision (mAP) at an intersection over union (IoU) threshold of 50% (mAP50), and mAP at IoU thresholds ranging from 50% to 95% (mAP50:90) values of 92.8%, 92.8%, 96.4%, and 82.2%, respectively. Taking YOLOv8n as the baseline, our method achieves respective improvements of 1.0%, 1.6%, 3.2%, and 4.2% in the four evaluation metrics mentioned above. Compared with state-of-the-art detection models, the proposed method demonstrates significant superiority in both accuracy and robustness, highlighting its potential for practical applications in the power line inspection domain.