DOI: 10.1145/3849382 ISSN: 1046-8188

A Survey of Long-Document Retrieval in the PLM and LLM Era

Minghan Li, Yishuai Zhang, Tianrui Lv, Siqi Zhao, Miyang Luo, Ercong Nie, Guodong Zhou

The rapid growth of long-form documents presents a fundamental challenge to information retrieval (IR). Their length, dispersed evidence, and complex structures demand approaches that go beyond standard passage-level techniques. This survey provides a comprehensive review of long-document retrieval (LDR) and organizes existing work into three major eras of development. It traces progress from classical lexical and early neural models to pre-trained language models (PLMs) and large language models (LLMs). It highlights core paradigms such as passage aggregation, hierarchical encoding, efficient attention, and LLM-driven re-ranking and retrieval. In addition to methods, the survey summarizes domain-specific applications, available evaluation resources, and the benchmarks that shape empirical study. It identifies open challenges related to efficiency trade-offs, evidence localization, multimodal alignment, and interpretability. The principal conclusion is that future progress in LDR will depend on hybrid solutions that combine sub-linear indexing, structure-aware modeling, and LLM-based reasoning (resources are available at https://github.com/lmh0921/Long-Document-Retrieval-SURVEY-PLMs-LLMs ).