Cattle-CLIP: A Multimodal Framework for Dairy Cattle Behaviour Recognition from Video
Huimin Liu, Jing Gao, Daria Baran, Axel X. Montout, Neill W. Campbell, Andrew W. DowseyMonitoring cattle behaviour provides valuable insights into animal health, productivity and welfare conditions. However, robust behaviour recognition in real-world farm environments remains challenging due to several data-related limitations, including the scarcity of well-annotated livestock video datasets and the substantial domain gap between large-scale pretraining corpora and agricultural surveillance footage. To address these challenges, Cattle-CLIP, a domain-adaptive vision–language framework based on Contrastive Language-Image Pretraining (CLIP), is proposed, in which cattle behaviour recognition is reformulated as a cross-modal semantic alignment task rather than a purely visual classification problem. Instead of directly fine-tuning visual backbones, Cattle-CLIP incorporates a temporal integration module to extend image-level contrastive pretraining to video-based behaviour understanding, enabling consistent semantic alignment across time. In order to minimise the domain discrepancy between web-scale pretraining data and real-world cattle surveillance scenarios, customised data augmentation and specialised behaviour-oriented prompts are employed. Furthermore, CattleBehaviours6, a curated and behaviour-consistent video dataset comprising 1905 annotated clips across six indoor behaviours was constructed: feeding, drinking, standing-self-grooming, standing-ruminating, lying-self-grooming, lying-ruminating, to support model training and evaluation. Beyond serving as a benchmark for our proposed method, the dataset provides a standardised ethogram definition, offering a practical resource for future research in livestock behaviour analysis. The performance of Cattle-CLIP was assessed under both supervised and few-shot learning paradigms, particularly targeting behaviour recognition in data-constrained environments, an important yet insufficiently studied task in livestock surveillance. Experimental evaluation demonstrated that the proposed model reached 96.1% accuracy across six behaviours in supervised tasks, while achieving both high precision and recall for feeding, drinking and standing-ruminating behaviours. Moreover, the model maintained promising few-shot performance under limited labelled data, highlighting the potential of multimodal learning for data-constrained cow behaviour recognition.