DOI: 10.1177/21582440261468849 ISSN: 2158-2440

A Semi-Automatic Approach to Developing a Large-Scale POS Tagged News Corpus for Urdu, a Low-Resource Language

Mushtaq Ali, Muzammil Khan, Huda Alsobhi, Rayed Alakhtar, Fazal Qudus Khan

The development of a large-scale, Part-of-Speech (POS) tagged corpus for Urdu, a morphologically rich and low-resource language, remains a significant challenge due to the scarcity of annotated linguistic data. This paper presents a semi-automatic approach to creating a comprehensive Urdu POS-tagged dataset, addressing the limitations of existing resources. The dataset, named MM-POST (Mushtaq & Muzammil POS Tagged) corpus, comprises 119,276 words/tokens drawn from 2,871 sentences across 75 online BBC Urdu news articles. These articles span seven diverse domains: General, Finance, Entertainment, Politics, Health, Sports, and Science, ensuring broad coverage of linguistic patterns. The tagging process combines the automated tagging capabilities of the Center for Language Engineering (CLE) POS Tagger with rigorous manual annotation, resulting in a high-quality and domain-sensitive corpus with a vocabulary size of 9,395 unique tokens. The study provides a detailed analysis of the corpus construction process, including tagging methodology, tagset design, domain-wise distribution, and corpus characteristics. The MM-POST corpus supports various NLP tasks such as Named Entity Recognition, Information Retrieval, Text Classification, and serves as a benchmark resource for training and evaluating Urdu POS taggers. By preserving syntactic structure, the dataset enables both token- and sentence-level text processing. Additionally, the paper reviews the evolution of Urdu tagsets and highlights prior machine learning-based POS tagging efforts. This work not only contributes a valuable linguistic resource for Urdu but also establishes a scalable framework for developing corpora in other low-resource languages. Future extensions include automated tagger development and Named Entity annotation to enhance the corpus’s applicability across NLP applications.

More from our Archive