An Interpretable Machine Learning Framework for Forest Biomass Estimation: Stacking Ensemble Architectures and Uncertainty Quantification
Jiecheng Liao, Yin Ren, Shudi Zuo, Xuejing Wu, Birhanie Alemayehu, Xin LiuRegression-based aboveground biomass (AGB) prediction from Earth-observation data often compresses the upper tail of the biomass distribution, yet the relative effectiveness of geospatial residual correction and ensemble learning in fragmented mountain landscapes remains unclear. Thus, we compared the two paths for reducing high-value underestimation: geospatial residual reconstruction using empirical Bayesian kriging regression prediction, and feature-space optimization using Stacking ensemble learning. SHapley Additive exPlanations (SHAP) interpreted feature contributions, and quantile regression forests (QRF) converted high-AGB point estimates into prediction intervals. Results show that geospatial optimization brought limited gain because residual spatial autocorrelation was weak (Moran’s I = 0.10), whereas Stacking improved overall R2 from 0.75 to 0.79 and reduced high-AGB bias (AGB > 80 t/ha) from −13.56 to −5.49 t/ha. This improvement was mainly attributed to complementary heterogeneous learners, with XGBoost capturing the primary non-linear trends, SVR extrapolating to correct high-AGB errors, and RF providing minor marginal calibration. SHAP analysis suggests that LiDAR-derived cubic mean height (Elev_curt_mean_cube) was the dominant feature explaining high-AGB variability, with a threshold response consistent with biomass-height allometry. QRF achieved coverage of 92.8% with a mean interval width of 74.51 t/ha, while coverage in the high-AGB subset was 84.2% with a mean width of 90.50 t/ha. The proposed comparison-and-diagnosis framework provides an interpretable approach for selecting an appropriate correction pathway and supports forest carbon monitoring, carbon accounting, and management decisions in complex mountain ecosystems.