A Supervised Morphosyntactic Analyzer for Modern and Classical Arabic Text
Majdi SawalhaArabic morphological analysis remains a central challenge for natural language processing (NLP). This challenge is embedded within the complex, rich, and multi-layered morphology of Arabic, where orthographic words encode enclitics, root-and-pattern derivations, and other morphosyntactic features. Another dimension of complexity is the coexistence of different varieties of Arabic, including Modern Standard Arabic (MSA), Classical Arabic (CA), and Dialectal Arabic (DA). In this four-phase empirical investigation of Arabic, a new, standardized CoNLL-U benchmark integrating the Arabella corpus (MSA) and the MASAQ corpus (CA) was evaluated across different settings. This paper aims to build a baseline morphological analyzer by evaluating a wider range of supervised systems, including lexicon-lookup, Conditional Random Fields (CRFs), full fine-tuning of Arabic PLMs, and AraT5 configurations. A parameter-efficient LoRA adaptation of QARiB was implemented, preserving most UPOS performance with minor differences. The results of the comparative evaluation of a variety of models across all settings were reported, and significant results were emphasized. The findings demonstrate severe limitations in current capabilities, particularly that cross-domain transfer for fine-grained language-specific tags (XPOS) remains difficult. Also, models struggle with the asymmetry of cross-register transfer, as moving from Classical to Modern Arabic exposes a lack of coverage for modern named entities and punctuation.