Data-Driven Ensemble Machine Learning for Multi-Horizon Microalgal Bioprocess Forecasting
Bartolomeo Cosenza, Riccardo Minardi, Luca Usai, Riccardo Allodi, Alessandro Concas, Daniele Sofia, Giancarlo Cravotto, Giovanni Denaro, Antonio Messineo, Maurizio Volpe, Antonio Picone, Robinson Soto-Ramirez, Catalina Valencia Peroni, Giovanni Antonio LutzuMicroalgal bioprocesses increasingly rely on predictive modelling to support automated control and reduce the cost of laboratory experimentation. Yet, the scarcity of high-quality datasets and the nonlinear nature of microalgal growth severely limit the accuracy and robustness of conventional machine-learning approaches. This study introduces an ensemble-learning framework for forecasting biomass accumulation in Limnospira platensis cultures supplemented with Effective Microorganism (EM) consortia. The method combines temporally consistent, strictly causal feature engineering with tree-based ensemble learning under a prospective validation protocol designed to mirror real deployment. Growth measurements are transformed into temporal descriptors that encode phase transitions, short-term growth dynamics, and treatment effects, enabling tree-based ensemble algorithms to capture nonlinear patterns inaccessible to traditional models. When features are restricted to information available strictly before each prediction, and models are evaluated only on real held-out measurements, one-step nowcasting does not surpass a trivial persistence baseline. For genuine multi-horizon forecasting of a monitored culture, the regime with practical value, a horizon-aware gradient-boosting model trained jointly on two cultivation media (Jordan and Zarrouk), attains R2 ≈ 0.72 on the held-out Jordan observations (0.72 ± 0.02, mean ± SD over 30 seeds; RMSE ≈ 0.20 g L−1, Pearson r ≈ 0.85) and remains stable across forecast horizons of 1–28 days, outperforming persistence roughly fourfold in pooled R2 (0.72 vs. 0.16) at medium-to-long horizons where the naive baseline collapses. Permutation-based feature-importance analysis identifies current biomass and the EM dilution level as the leading predictors, whereas EM treatment identity contributes only modestly. Including an independent second-medium experiment (Zarrouk) in joint training did not materially change held-out Jordan performance, indicating that, under these data-limited conditions, neither additional model complexity nor a second training medium substantially improved forecasting once the causal, horizon-aware framework was in place. Among treatments, EM3 most strongly suppressed biomass accumulation. Overall, the study provides an honest, fully reproducible, two-medium benchmark for microalgal biomass forecasting under data-limited conditions, and identifies the forecast horizon as the regime in which ensemble learning adds genuine value over trivial baselines.