ANERD: A Large-Scale Arabic Corpus for Named Entity Recognition and Disambiguation
Madawi Saqer Alotaibi, Mohamed El Bachir MenaiArabic NER and NED research is limited by the lack of large-scale public datasets that support both tasks with validated entity links and difficulty-aware evaluation. In this paper, we propose ANERD, a large-scale Arabic corpus and benchmark for NER/NED, to overcome these drawbacks. ANERD includes 502,756 sentences from journalistic and encyclopedic texts. The entity mentions are manually annotated with BIO labels and linked to Arabic Wikipedia when a suitable target exists. This corpus is split into Easy, Medium and Hard subsets for diagnostic evaluation based on inter-annotator agreement. To evaluate the corpus, we use three NER baselines: the Arabic-specific Transformer AraBERTv2, multilingual BERT (mBERT), and a BiLSTM–CRF sequence-labeling model. On the held-out ANERD test sets, AraBERTv2 achieved entity-level NER F1 scores of 98.88%, 97.91%, and 96.15% on the Easy, Medium, and Hard partitions, respectively; mBERT achieved 90.45%, 91.79%, and 83.05%; and BiLSTM–CRF achieved 99.40%, 98.31%, and 98.21%. For NED, the no-gold-injection evaluation yielded test Hits@1 scores of 85.34%, 68.30%, and 69.87% and MRR@30 scores of 85.38%, 68.42%, and 73.52% on the Easy, Medium, and Hard partitions, respectively. Before reranking, candidate Recall@30 was 85.45%, 68.54%, and 78.20% on the Easy, Medium, and Hard partitions, respectively. These results indicate that candidate coverage and contextual reranking are separate factors that influence NED performance. Cross-corpus experiments also reveal the effects of domain shift and differences in annotation policies on transfer performance. In summary, ANERD is a large-scale, comprehensive resource for research on Arabic NER, NED, and entity linking.