DOI: 10.3390/app16199719 ISSN: 2076-3417

Physics-Guided Deep Learning and Geospatial Aggregation for Photovoltaic Panel Detection and Duplicate-Free Counting in Unmanned Aerial Vehicle Imagery

Sixin Zhu, Kunping Liang, Xu Zhao, Yijie Hu, Hongqian He

Automated inspection of utility-scale photovoltaic plants requires reliable detection and inventory of individual panels from overlapping aerial images. We developed a physics-guided deep learning and geospatial aggregation framework for panel detection and duplicate-free plant-level counting. Scene-targeted augmentation introduces bounded specular-reflection, occlusion, perspective, and boundary-ambiguity perturbations into training images. A You Only Look Once version 11 (YOLO11)-based detector combines adaptive input resolution, reflection-oriented channel attention, and a decoupled feature pyramid network. Frame-level detections are projected into a common geographic coordinate system using unmanned aerial vehicle (UAV) pose and camera geometry, and spatial clustering merges repeated observations without image stitching. On 4200 UAV images containing approximately 110,000 annotated panel instances, the complete detector achieved 74.3% average precision at an intersection-over-union threshold of 0.5 (AP@0.5) and 62.9% AP@0.5:0.95 across five runs. Geospatial aggregation achieved 99.2% plant-level counting accuracy, compared with 75.1% for image stitching. On an independent held-out desert-site dataset containing approximately 45,000 panel instances, the Site-A-trained model was applied without retuning and retained 72.8% AP@0.5 and 98.4% plant-level counting accuracy. Component, attention, regional-counting, layout, cross-site, and runtime experiments further evaluated the framework. Final-pass throughput was 105.6 frames/s; a separate two-pass runtime experiment reported 15.1 ms/image for the complete processing chain, equivalent to 66.2 frames/s. The results support integrated panel detection and inventory under the evaluated imaging and survey conditions.