DOI: 10.1162/coli.a.650 ISSN: 0891-2017

Using Large Language Models in Formalizing Classical Linguistic Descriptions: A Case Study in Middle Indo-Aryan Sound Change

V. S. D. S. Mahesh Akavarapu, Chinmay Dharurkar, Johannes Dellert, Arnab Bhattacharya, Gerhard Jäger

Abstract

We present a systematic evaluation of Large Language Models (LLMs) in translating classical descriptions of phonological and morphophonological change from Old Indo-Aryan (Sanskrit) to Middle Indo-Aryan (MIA) into the standard notation of modern historical linguistics. Drawing on Vararuci’s Prākṛta Prakāśa (c. 4th century CE) and English translation of Bhāamaha’s commentary (c. 6th century CE), we construct a dataset of 470 phonological and morphophonological rules extracted from these sources using LLMs followed by manual curation. Out of these rules, we compile a benchmark of 216 sound-change rules with example reflexes. Our pipeline integrates Optical Character Recognition (OCR) of non-digitized historical texts, manual gold-standard curation, and LLM-based translation of rule descriptions into contemporary phonological rule notation. Evaluation across several recent LLMs shows accuracies up to 86% on this challenging formalization task. Models incorporating explicit reasoning mechanisms consistently outperform non-reasoning variants, underscoring the importance of reasoning in linguistic formalization. Error analyses reveal systematic weaknesses in modeling complex conditioning environments. We further show that the extracted sound laws generalize well across a broader range of MIA languages. Overall, this work illustrates how contemporary LLMs can engage with millennia-old linguistic scholarship by systematically translating and structuring classical rule descriptions into modern formal representations.

More from our Archive