DOI: 10.3390/en19194504 ISSN: 1996-1073

Evaluating Domain-Aware Random Forests for Short-Term Photovoltaic Power Forecasting

Situmbeko Nyirenda, Musa Ndiaye, Esau Zulu, Alan Kabanshi

Data-driven machine learning (ML) models forecast photovoltaic (PV) power accurately, but the conventional Random Forest (RF) samples candidate predictors uniformly and therefore ignores established physical knowledge about how meteorological variables drive PV generation. This paper proposes the Domain-Aware Random Forest (DARF), a mechanism that converts expert structural knowledge into pre-training feature sampling priors. Expert PV knowledge is represented as a directed acyclic graph (DAG), the strength of each edge is estimated from training data using multivariate ridge regression (MVRR), and the edge strengths are propagated through the graph to yield Domain Prior Weights (DPWs) that bias the feature sampling step of RF construction. The method is evaluated on six months of 30 min operational data from a 34 MWp PV plant in Kitwe, Zambia, for a short-term (h < 1 h) power forecasting task, against a conventional RF and an Extreme Gradient Boosting (XGBoost) model trained under identical conditions over 60 seeded runs. Three alternative expert graphs are tested. With the graph that emphasizes direct physical relationships with PV power, DARF reduces MAE by 6.69%, RMSE by 8.85% and MAPE by 1.11% relative to RF and raises R2 from 0.773 to 0.812; the differences are statistically significant (Wilcoxon signed-rank test, p < 0.05) with large effect sizes (rank–biserial r > 0.5) for all metrics except MAPE (r = 0.3). Feature selection frequencies confirm that the mechanism shifts tree construction toward predictors with higher DPWs, at a training-time cost of 6.7% relative to RF and no change in inference time. XGBoost remains more accurate overall, indicating that the contribution of DARF lies in a transparent, knowledge-guided improvement of the RF rather than in state-of-the-art accuracy.