Automated Data‐Efficient Symbolic Regression for Interpretable Bioprocess Model Development
Luca Riezzo, Alexander Rogers, Harry Kay, Dongda ZhangABSTRACT
Bioprocessing is central to the sustainable manufacture of pharmaceuticals, food products, and renewable chemicals. Consequently, developing high‐fidelity kinetic models to facilitate accurate process prediction, optimisation, and scale‐up is a top research priority. However, bioprocess model construction remains hindered in practice by incomplete mechanistic understanding and limited data availability. Therefore, to accelerate the development of accurate bioprocess models, this work presents a data‐efficient symbolic regression (SR)‐based framework to simultaneously indentify interpretable model structures and aid knowledge discovery. A generic macroscopic kinetic model backbone was used to capture overall process behaviour, while SR was applied to strategically uncover the structures of critical kinetic terms within the backbone. Two implementation strategies were evaluated using an in‐silico yeast fermentation case study. The first strategy, embedded SR directly into the kinetic model backbone while the second identified time‐varying parameter profiles prior to SR. The results demonstrated that independently identifying individual kinetic terms was crucial for recovering the ground‐truth model, while refining SR‐generated candidates through a novel local iterative structural correction strategy significantly improved convergence to the true kinetic expressions, surpassing model‐based design of experiments in data efficiency. This study therefore enables automated yet interpretable model construction for small‐data bioprocess applications, paving the way towards augmented intelligence driven bioprocess modelling and accelerating digital twin development for process optimisation and control.