Poem meter classification of spoken Arabic poetry: integrating high-resource systems for a low-resource task
Maged S. Al-Shaibani, Zaid Alyafeai, Irfan Ahmad, Abdulkareem Saleh AlzahraniArabic poetry is a cornerstone of Arab cultural and linguistic heritage, yet the computational analysis of spoken Arabic poetry remains critically underexplored. An open question is: how can the meter of a spoken Arabic poem be automatically identified when labeled acoustic data is scarce? Meter identification from audio is more challenging than from text because it must handle both acoustic variability and the complex prosodic rules of Aroud—the classical Arabic science of poetic meter—across all sixteen canonical meters. To the best of our knowledge, no publicly available evaluation covering all sixteen meters from acoustic recordings currently exists, and no standardized benchmark exists for this task. We investigate two methodological approaches: an end-to-end architecture that fine-tunes a pretrained Wav2Vec2 model with a meter classification head, and a novel two-stage integrated framework that chains a high-resource Arabic speech recognition system with a high-resource textual meter classifier, augmented by a domain-specific 4-gram language model trained on Arabic poetry. Both approaches are validated on a controlled baseline set of annotated recordings and a benchmark collected under diverse real-world acoustic conditions. The integrated framework achieves up to 92% accuracy on the benchmark set, demonstrating that combining high-resource systems from adjacent domains can effectively compensate for low-resource constraints. Beyond establishing a strong benchmark result for acoustic Arabic poetry meter classification, this work releases a reusable benchmark that standardizes evaluation for future research in computational Arabic literary heritage.