DOI: 10.66532/jhai.2026.0023 ISSN: 3092-5533

LLM-ASSISTED MARKUP OF ENTITIES IN TEI: A CASE STUDY

JONAH CAUSIN, MORGAN FUKSA, ARNE KÄFER, JOSEPH (SANG WUK) LEE, CLIFFORD ANDERSON

This paper presents a case study of marking up TEI (Text Encoding Initiative) documents using large language models. We detail our experiments using LLMs for named entity recognition and entity linkage in a corpus of TEI documents. The corpus consists of four journals co-edited by the Swiss-German theologian, Karl Barth (1886–1968). We aimed to enrich the TEI markup by adding XML elements to mark up entities and attributes to link to QIDs on Wikidata. We experimented with using both cloud-based frontier models and local open-source models to enrich a subset of 34 articles from the corpus. After analyzing the outcome using both human reviewers and automated analysis, we indicate where our methodology proved successful and where we encountered problems. We conclude that LLMs can perform named entity recognition and linking in TEI documents, but that the combined financial and labor cost of scaling this procedure to the full corpus would be relatively high without optimizing our current pipeline.