DOI: 10.3390/s26196156 ISSN: 1424-8220

DGF–YOLOv12n–LiteP2: Dynamic Gated Local–Context Fusion with a High-Resolution Detection Head for Cattle Body-Part Detection in Complex Barns

Yanhong Liu, Le Yang, Qingqing Li, Jiazi Han, Jingrong Wang, Meng Han, Hua Yang

Reliable cattle body-part localization provides visual inputs for downstream posture, locomotion, and behavior analysis in precision livestock farming. However, complex barn environments introduce scale variation, occlusion, illumination changes, and rail-like background interference. DGF–YOLOv12n–LiteP2 is presented as a YOLOv12n-based detector that combines dynamic fusion of local and contextual features with a lightweight stride-4 prediction branch. The dynamic gated fusion module integrates multi-kernel local feature extraction, progressive receptive-field aggregation, and input-conditioned channel selection. LiteP2 preserves high-resolution spatial information by reusing the backbone P2 feature without constructing a complete high-resolution feature pyramid. Experiments were conducted on a single-site dataset containing 6991 valid images and 132,302 body-part annotations. On the held-out in-domain test set, DGF–YOLOv12n–LiteP2 achieved 93.92% mAP50, 79.27% mAP75, and 69.88% mAP50:95. Compared with YOLOv12n, mAP50:95 increased by 9.61 percentage points. Across three training runs with different random seeds, the mean mAP50:95 reached 70.03±0.30%. The model contained 11.68 M parameters and required 22.89 GFLOPs, indicating a trade-off between detection accuracy and computational cost. These results support the complementary roles of dynamic feature fusion and high-resolution prediction under the evaluated barn conditions. Further validation on additional farms, longer acquisition periods, and agricultural edge devices remains necessary.