DOI: 10.1002/rse2.70110 ISSN: 2056-3485

Few Annotations, High Accuracy: Transfer Learning and Data Augmentation Improve Passive Acoustic Monitoring of the Vulnerable African Manatee

Lucas Dubus, Auguste Verdier, Nina Giotto, Grace Mbemba, Gabriel Michelin, Baptiste Mulot, Stéphanie Manel, David Mouillot

ABSTRACT

Monitoring elusive and threatened species remains a major challenge in ecology and conservation, particularly in turbid aquatic environments. The African manatee ( Trichechus senegalensis ), listed as Vulnerable by the IUCN, remains difficult to survey using visual methods, limiting reliable assessments of its distribution and population status. Passive acoustic monitoring (PAM) offers a promising alternative monitoring tool, but the large volumes of audio data generated require efficient automated detection of vocalizations. Here, we evaluate and compare three automated approaches for detecting African manatee vocalizations and occurrences across 52 sites over 13 months in the Conkouati‐Douli National Park (Republic of the Congo): (i) a widely used traditional machine learning recognizer implemented in Kaleidoscope, and two deep learning models, (ii) a convolutional neural network (CNN) that learns from local patterns in a visual representation of sounds, and (iii) a transformer model that incorporates attention layers to capture broader acoustic context. We further assess the influence of training dataset size, transfer learning, and data augmentation on model performance under realistic, data‐limited conditions. Both deep learning models substantially outperformed the traditional approach, achieving F1‐scores of 0.89 compared with 0.49 for Kaleidoscope, primarily due to much higher recall while maintaining high precision. Remarkably, deep learning models reach a high performance using only 10% of the annotated training dataset (≈300 vocalizations), demonstrating the effectiveness of transfer learning and data augmentation for bioacoustics applications with limited annotation effort. This ability to reliably detect true occurrences is particularly critical for monitoring threatened and elusive species, for which missed detections can lead to underestimation of habitat occupancy, spatial distribution, and population status. By minimizing false absences, deep learning‐based acoustic recognizers provide more reliable occurrence data for conservation decision‐making and long‐term monitoring even in data‐poor conditions.