DOI: 10.1002/widm.70127 ISSN: 1942-4787

Recent Advances in Text Anonymization: A Systematic Review

Marina Litvak, Alípio Jorge

ABSTRACT

Text anonymization has become a critical requirement across healthcare, legal, financial, and online communication domains, where large volumes of sensitive textual data are increasingly used for analytics, information retrieval, and model training. Despite decades of research, anonymization remains challenging due to the diversity of sensitive information types, domain‐specific annotation schemes, and the emergence of complex inference risks that extend beyond explicit identifiers. On the other hand, we have seen a large number of approaches using transformers in recent years. This survey provides the first comprehensive and systematic review of text anonymization methods published between 2020 and 2025, covering 48 primary studies identified through a structured search and rigorous screening procedure. We analyze anonymization approaches across rule‐based, statistical, neural, transformer‐based, and large language model (LLM) paradigms, and compare them across multiple domains, languages, and sensitive‐information taxonomies. Our review also synthesizes available datasets, software resources, and evaluation practices used to assess both privacy protection and text utility. The findings highlight significant fragmentation across domains, persistent limitations in handling quasi‐identifiers and semantic leakage, and substantial inconsistencies in evaluation protocols. We identify key methodological trends, gaps, and emerging challenges, including the integration of LLMs, multilingual settings, and adversarial evaluation. Our analysis reveals a growing shift from identifier‐centric de‐identification toward context‐aware anonymization, while evaluation methodologies have not yet evolved at the same pace. We outline open research directions and propose a roadmap toward more robust, context‐aware, and empirically grounded anonymization systems. This survey aims to establish a unified reference point for researchers and practitioners working on privacy‐preserving NLP and sensitive text processing.

More from our Archive