Corpus linguistics in the age of generative AI: practical and theoretical considerations
Tobias BernaischAbstract
Generative AI tools have implications for corpus linguistics from practical and theoretical angles, both of which the current paper addresses. The first part of the paper establishes whether and how generative AI tools – here ChatGPT o3 – can in practice support a corpus-linguistic workflow comprising data extraction, cleaning, and annotation for the dative alternation. While certain steps in this workflow such as data extraction or object length annotation can be completed reliably in ChatGPT o3, more taxing processes like cleaning extracted data for relevant instances, identifying clause boundaries and clause elements, or annotating objects for animacy are marked by errors and inconsistencies, which even the addition of steps dedicated to facilitating AI support in the corpus-linguistic workflow cannot remedy. The second part of the paper relates to corpus-linguistic theory by discussing whether and how AI-generated texts should be represented in linguistic corpora. It argues that the notions of authenticity and representativeness as key features of linguistic corpora are compatible with the integration of AI-generated texts in linguistic corpora. As said integration marks a departure from corpus-linguistic tradition, various arguments in favour of and against featuring AI-generated texts are provided, yielding a pragmatic practical suggestion for future corpus projects.