DOI: 10.1049/rpg2.70373 ISSN: 1752-1416

Ethical Guardrail Bandit Pruning (EGBP): A Fairness‐Constrained Reinforcement Learning System for Equitable Governance of Distributed Bioenergy Grids

Ntebogang Dinah Moroke

ABSTRACT

Distributed bioenergy renewable power networks face an unresolved equity challenge: AI‐governed dispatch optimisers maximise aggregate throughput while systematically under‐serving peripheral nodes. This paper presents ethical guardrail bandit pruning, a fairness‐constrained reinforcement learning system enforcing equity and energy sustainability via a two‐timescale primal‐dual guardrail, coupling a bandit‐gradient importance estimator, an ethical guardrail buffer, cost‐weighted pruning and guardrail‐filtered federated averaging over a constrained Markov decision process. Three theorems establish regret optimality, convergence and a federated exclusion bound. The integrated system was empirically verified across 16 seeds at 60 communication rounds, with no catastrophic instability (maximum Gini below 0.28) and steady‐state Gini 0.188 versus an unconstrained baseline of 0.805, confirmed by ablation and sensitivity analysis. Convergence was checked against logged gradient norms and found shallower than theory predicts. A 5000‐round extended‐horizon run (50,000 local steps) remained within the safety threshold in of rounds; multi‐seed extended‐horizon verification is future work. Results operationalise SDG 7, 10, 12 and 13 as verifiable engineering guardrails.