A Standardized Methodological Framework for Model Development in LabGym: A Case Study in Lions
Frej Gammelgård, Silje Marquardsen Lund, Jonas Nielsen, Thea Loumand Faddersbøll, Trine Hammer Jensen, Sussie Pagh, Kirstin Anderson Hansen, Flemming Nielsen, Cino PertoldiBehavioral observation is central to animal welfare assessment in zoological institutions, but manual observation is time-consuming and difficult to apply continuously. We developed and applied a LabGym-based workflow for automated behavioral monitoring of captive lions (Panthera leo). Development footage from three Danish zoological institutions was used to train a lion detector and a seven-class behavior categorizer; video collected on 22–28 December 2025 was temporally held out from model development and used for applied validation. Detector performance was moderate (bounding-box AP 47.1%; segmentation AP 41.6%), while the internal categorizer validation reached 0.62 accuracy and weighted F1. In the held-out validation, 52.6% of manually annotated lion-time received a retained pipeline prediction. Fine-scale end-to-end accuracy was 0.231 and increased to 0.440 when calculated only on covered lion-time. Grouping behaviors as active, inactive/maintenance, and pacing-like locomotion increased these values to 0.317 and 0.602, respectively. Coverage varied substantially among institutions (33.8–59.8%) and was particularly low for cubs and nighttime footage. Although model macro-F1 exceeded simple majority-class baselines, end-to-end accuracy did not. The workflow demonstrates the potential of computer-vision approaches for scalable behavioral quantification and future welfare assessment in zoological settings, while highlighting the importance of accounting for prediction coverage and classification performance when interpreting automated behavioral data.