Reinforcement-Learning-Enabled Supervisory Control of a Full-Scale Anaerobic Digestion Facility Using a Calibrated Digital Twin
Ahmed I. Yunus, Srinivas Jalla, Joe F. Bozeman, Yongsheng ChenAbstract
Anaerobic digestion (AD) is a cornerstone of energy recovery and decarbonization at wastewater resource recovery facilities (WRRFs), yet full-scale digesters are often operated conservatively due to delayed process feedback, nonlinear stability constraints, and limited observability. This study presents a novel reinforcement learning (RL)-enabled supervisory control framework for full-scale AD that integrates a calibrated mechanistic digital twin model (DTM) with offline policy learning and counterfactual uncertainty evaluation. Specifically, the DTM was calibrated to reproduce control-relevant digester behavior. Using 2.7 years of simulated operation, a batch-constrained Q (BCQ)-reinforcement learning policy was trained to maximize biogas production while penalizing instability associated with pH excursions and elevated volatile fatty acids (VFAs)/alkalinity (ALK) ratios. Counterfactual evaluation against historical, greedy, and heuristic controls used 1 year unseen DTM simulation and 1 year of full-scale operational data. The RL policy achieved comparable mean daily biogas production (+0.35%) in simulated testing. Under the 1 year operational data, it achieved plant-equivalent mean production of 9802 m3/day, corresponding to a modest potential cumulative improvement (+2.7%, 93 405 m3/year) subject to substantial uncertainty P(Δ > 0) = 0.524. These results indicate that offline-trained RL, combined with a mechanistic digital twin and counterfactual evaluation, provides a feasible and interpretable framework for developing and assessing supervisory AD control without live-plant experimentation while showing that the resulting benefit must be interpreted under explicit model-plant uncertainty.