DOI: 10.1017/nlp.2026.10040 ISSN: 2977-0424

Named Entity Recognition Across Datasets and Domains: Resources and Models for Galician

Albina Sarymsakova, Ettore Mariotti, Helena Pérez Puente, Marcos Garcia

Abstract

Automatic named entity recognition (NER) is essential for many natural language processing applications, particularly in low-resource languages and those with corpora restricted to a single domain, where the lack of diverse data may hinder cross-domain generalisation. In this context, we attempt to contribute to the Galician NER field by extending manually annotated resources and exploring several approaches to improve the performance of automatic classifiers. To this end, we enrich two new datasets with NER annotations in Galician, which we release in this paper. Through experimental evaluation, we investigate the benefits of fine-tuning transformers on NER data by comparing various monolingual and multilingual BERT models and developing a new extra-small DeBERTa model for Galician. Furthermore, we assess the quality of the NER model by expanding the initial training corpus with additional data from related varieties, including Portuguese and Spanish. For evaluation, we use five Galician corpora, including our two newly annotated datasets and three adapted from existing resources, ensuring assessment across multiple domains. Our findings reveal (1) a positive impact of enhanced cross-lingual training data on improved model performance across five Galician test datasets; (2) similar performance of monolingual and multilingual models depending on the training data; and (3) the competitive effectiveness of extra-small BERT models for Galician NER tasks compared to larger models.