GMS-YOLO11n: A Sheep Detection Model for Challenging Fixed-View Farm Conditions Integrating Spatially Gated Structural Enhancement and Multi-Scale Attention
Wenbo Yu, Ruoya Xie, Yongqi Liu, Zhi Xue, Zhengpeng Yang, Wenqiang HanAccurate sheep detection supports counting, tracking, behavior analysis, and health monitoring in intelligent livestock farming. However, occlusion, scale variation, nighttime low light, and background interference can cause missed detections and inaccurate localization. This study proposes GMS-YOLO11n, an improved YOLO11n-based detector for fixed-view sheep monitoring. A spatially gated bottleneck convolution module was introduced at key backbone downsampling stages to strengthen local structural cues, while a multi-scale attention fusion module combined deep semantic information with shallow-to-intermediate details through feature alignment, cross-scale attention, and channel-wise learnable gating. A dataset of 3531 images of Small-tailed Han sheep, containing 8167 annotated instances, was used for evaluation. Across six independent runs, GMS-YOLO11n achieved precision, recall, F1-score, mAP@0.5, and mAP@0.5:0.95 of 93.62% ± 0.12%, 91.03% ± 1.37%, 92.30% ± 0.75%, 96.53% ± 0.49%, and 76.32% ± 0.57%, respectively. Compared with the COCO-pretrained YOLO11n baseline, the corresponding improvements were 1.05, 4.38, 2.79, 2.76, and 6.92 percentage points. Subset analysis showed mAP@0.5:0.95 gains of 9.83 and 7.62 percentage points under nighttime low-light and obvious-occlusion conditions. These results demonstrate improved detection performance under the challenging visual conditions represented in the evaluated single-farm, fixed-view dataset; broader generalization requires external validation.