DOI: 10.3390/ani16162575 ISSN: 2076-2615

Occlusion-Robust Cattle Pose Estimation for Precision Livestock Monitoring Using Hierarchical Locality Refinement

Yingchao Wang, Na Li, Dan He, Shan Sun, Xinjian Chu, Zixiang Qin, Feng Xue, Jingjun Yi, Hao Wu, Han-Su Zhang, Fan Zhao

Accurate cattle pose estimation is important for precision livestock farming because body landmarks can provide quantitative visual inputs for potential health monitoring, behavior analysis, lameness assessment, and welfare evaluation. This study focuses on keypoint-localization performance and provides a foundation for future task-specific studies of these downstream outcomes. However, real farm images commonly contain inter-cattle occlusion, cluttered backgrounds, small anatomical landmarks, and visually similar animals, which limit the reliability of existing one-stage pose estimators. This study proposes LEC-Pose, a hierarchical locality refinement framework for occlusion-robust cattle pose estimation. Built on YOLOv8-Pose, LEC-Pose first predicts cattle boxes and auxiliary coarse keypoints, then extracts instance-level ROI features to recover local anatomical evidence. A lightweight refinement network directly predicts the final keypoints from heatmaps and offsets on enhanced ROI features. The detected boxes guide the ROI pathway, while the initial keypoints remain auxiliary outputs. During training, an instance-level contrastive loss regularizes a compact global ROI descriptor. During inference, pair construction and contrastive-loss computation are removed, while descriptor fusion remains in the prediction path. Across three runs on CattleEyeView, LEC-Pose reaches 39.47 ± 0.33 mAP@50:95, compared with 33.91 ± 0.28 for YOLOv8-Pose, while running at 111.6 FPS versus 138.2 FPS for YOLOv8-Pose, corresponding to a 19.2% throughput reduction on the evaluated RTX 4090. On NWAFU-Cattle, it obtains 76.08 ± 0.36 and 80.03 ± 0.38 mAP@50:95 under the 50%/50% and 80%/20% protocols, respectively; the former is comparable to FSMC-Pose, whereas the latter is the highest mean among the three repeated principal models. An end-to-end zero-shot evaluation on three external cattle datasets, second-dataset occlusion analysis, detailed module combinations, and descriptor diagnostics further assess the framework within their stated protocols. These results suggest that LEC-Pose can serve as a pose-estimation component for future automated cattle monitoring, and future task-specific studies can connect these keypoints to health, behavior, and welfare outcomes.

More from our Archive