PA-YOLO: Partial-Channel Grouped Attention and Multi-Scale Prediction for Object Detection in UAV Imagery
Kang Tan, Hu Yao, Zhiyuan Liu, Xingjie Zhang, Qiang Zhang, Xiao Wu, Xinggan Peng, Xingzhou ChenUnmanned aerial vehicle (UAV) imagery presents small-object footprints, dense layouts, occlusion, and weak appearance cues, requiring both local spatial detail and broader context under constrained computation. Existing high-resolution prediction paths and attention mechanisms can strengthen these representations, but their extensive use increases feature-map, attention, and fusion costs. We propose Partial Attention–YOLO (PA-YOLO); its primary methodological contribution is a partial attention module with multi-group query attention, which applies grouped contextual modeling to selected channels while retaining a direct bypass for locally derived convolutional detail. The architecture further integrates established RepVGG-style re-parameterized downsampling with high-resolution four-scale feature fusion and detection. Fixed-recipe component ablation on VisDrone indicates complementary contributions from partial-channel attention and the integrated multi-scale design. On VisDrone, PA-YOLO-n and PA-YOLO-s report AP50/AP50:95 point estimates of 39.4/23.8% and 45.0/27.2%, respectively; under a secondary, non-official 2000-frame UAVDT-derived protocol, the corresponding values are 96.6/69.3% and 97.3/72.0%. These accuracies are single-run point estimates. Controlled forward-only profiling on one RTX 3090 platform gives mean eager-FP32 latencies of 13.795 and 14.197 ms/image, showing that compact parameterization does not necessarily yield lower inference latency than convolution-centric baselines. These findings motivate deployment-oriented optimization and onboard evaluation.