Reinforcement Learning-Based Intelligent Adaptive Protection Relays: A Systematic Comparative Review of Algorithms, Applications, and Performance Metrics
Sipho P. Lafleni, Tlotlollo S. Hlalele, Mbuyu SumbwanyambeThe accelerating integration of inverter-based resources (IBR) and distributed generation (DG)—with IBR penetration now exceeding 40% in many distribution networks and fault-current contribution curtailed to 1.0–1.5 p.u. versus 5–10 p.u. for synchronous generators—has measurably reduced the efficacy of conventional power system protection, which was conceived for unidirectional fault-current topologies. This paper presents a rigorous, PRISMA-guided systematic review of reinforcement learning (RL)-based intelligent adaptive protection relays (APR). A structured database search across IEEE Xplore, Scopus, Web of Science, and ScienceDirect yielded 72 candidate papers published between 1983 and 2025; the application of predefined inclusion/exclusion criteria retained 18 peer-reviewed studies for in-depth analysis. Unlike broad AI survey articles, this review focuses exclusively on the intersection of RL algorithms and power-system protection relaying, providing the following: (i) A detailed algorithmic taxonomy spanning Q-learning, SARSA, DQN, DDPG, PPO, SAC, and REINFORCE. (ii) Two complementary comparative tables: the first presents a standardized summary of algorithm families across seven reinforcement learning categories, highlighting the widespread lack of verified quantitative benchmarks; the second presents a comparative analysis of representative studies, detailing confirmed performance metrics where available and structured qualitative evaluations otherwise. (iii) Structured figures illustrating PRISMA flow, algorithm-performance radar charts, and a technology-readiness-level (TRL) matrix. (iv) A critical synthesis of methodological quality, dataset diversity, and reproducibility. The results confirm that PPO shows the strongest evidence for adaptive coordination performance under high-IBR penetration; SAC performs strongly on adjacent power-system control problems (voltage regulation, energy dispatch), but no protection-specific SAC validation study was identified in the reviewed corpus; and Q-learning remains valuable in constrained, safety-critical environments. The identified research gaps include the absence of standardized fault-scenario benchmarks, limited hardware-in-the-loop (HIL) validation, and under-explored cybersecurity resilience. Future directions encompassing explainable AI, digital twin co-simulation, and safe RL frameworks are systematically outlined.