DOI: 10.3390/electronics15163643 ISSN: 2079-9292

Machine Unlearning Across AI Systems: A Scoping Review of Evidence, Evaluation, and Deployment Contexts

Hyeonwoo Kim, Jihoon Moon, Namkyun Baik

Machine unlearning (MU) is increasingly studied across artificial intelligence systems in which target information may remain in model weights, adapters, retrieval indexes, caches, federated states, and temporally derived representations. However, existing reviews often examine methods, benchmarks, or deployment settings separately. Therefore, this scoping review collected and synthesized 159 sources to compare unlearning evidence across these connected system components. The evidence map separates 42 core MU sources—26 primary empirical or theoretical studies, 12 benchmark or evaluation frameworks, and 4 secondary evidence syntheses—from 117 contextual sources. First, a three-dimensional taxonomy distinguishes guarantee or reference type, update mechanism, and deployment or memory context. This structure prevents partitioning from being treated as a guarantee and influence estimation from being treated as an outcome. Next, the quantitative map identifies 13 centralized or general sources, 16 LLM- or benchmark-focused sources, 3 graph sources, 3 federated sources, 2 temporal sources, 2 quantized-network sources, and 3 cross-domain reviews. These results indicate a stronger reference-based foundation for centralized MU and a benchmark-rich but transformation-sensitive evidence base for large language models. In contrast, graph, federated, and temporal unlearning remain less mature because they are supported by smaller core evidence sets. Moreover, recent studies show that apparent forgetting may fail after 4-bit quantization, probabilistic decoding, alternative reference selection, recovery testing, benign query changes, or overlap between forget and retain knowledge. Accordingly, we present an audit-oriented deletion lifecycle that connects forgetting evidence, retained-utility testing, threat-model-specific attacks, post-transformation verification, and redeployment decisions. Calibration, citation grounding, provenance, and human review are treated as supplementary decision-readiness checks rather than direct proof of unlearning. Finally, the lifecycle is presented as a structured synthesis and reporting framework that still requires prospective validation in real systems and independent practitioner assessment.

More from our Archive