Missing context: a review of natural language processing for archives
Lavinia Dunagan, Jesse Johnston, Dallas CardAbstract
Archives are essential vehicles for both historical research and memory practices like genealogy. The sheer volume of documents produced via digital infrastructures, however, has posed a challenge for traditional archival processes that rely on human judgments. Some archivists have responded to this challenge by engaging with research in natural language processing (NLP), a subfield of AI that offers practical tools for working with large corpora of text. In this paper, we review applications of NLP methods in archival processes in order to surface fundamental tensions between archival values and automation. NLP can be used to classify materials as sensitive, recognize names in documents and facilitate new kinds of search, among other applications. Our discussion of these issues follows the journey of a document through an archive, from acquisition to interpretation. We find that sensitivity to context, transparency, resource limitations and maintenance appear as recurring challenges that are complicated by automation. The benefits that NLP technologies lend to managing large archives, however, suggest that these tools should be considered, not discarded. Turning to a more forward-looking discussion of generative AI and archival infrastructure, we argue that the narrative and dialogic potential of contemporary large language models has an especially contentious relationship with conventional archival desiderata.