DOI: 10.3390/info17090917 ISSN: 2078-2489

Entity Resolution Using Transformer-Based Language Models: A Systematic Scoping Review

Mohammad Beheshti, Maryam Seifaddini, Amir Erfan Zareei Shams Abadi, Karan Karthik, Tarun Mummidi Ramesh Kumar, Suzanne Austin Boren, Iris Zachary

Entity resolution (ER) is fundamental to integrating heterogeneous data, which traditional approaches address through deterministic rule-based methods and probabilistic record linkage. We conducted a systematic scoping review following PRISMA-ScR to characterize the use of transformer-based language models for ER. Five databases were searched, and 155 studies were included for synthesis. The literature expanded sharply after 2023, with 55% of included studies published in 2025 or 2026. General-purpose matching was the most common entity focus, followed by product/e-commerce. Healthcare applications were especially scarce, with only one study applying ER to the patient/healthcare domain. Among studies that performed blocking, embedding-based similarity was most common, followed by string-based approaches. Classification-head approaches remained the most common matching approach, followed by prompt-based approaches. Encoder-only models remained the most widely evaluated architecture, while decoder-only models grew increasingly prominent from 2023 onward. Full-parameter fine-tuning was the predominant learning strategy, followed by zero-shot and few-shot prompting. Among studies classified as general-purpose, nearly one-third were evaluated on only one or two entity types, limiting their generalizability. Reported best F1 scores varied across benchmarks, with no consistent advantage for encoder-only versus decoder-only architectures. These findings support the need for more diverse, end-to-end evaluations that consider efficiency and robustness alongside accuracy.