DOI: 10.2174/0126662558474397260915091751 ISSN: 2666-2558

Explainable Machine Learning for Student Outcome Prediction using Boosting and Additive Models

Avishek Barman, Anal Acharya, Soumen Mukherjee

Introduction:

This study proposes an explainability-driven framework for student outcome prediction by comparing a black-box high-performance model with an inherently interpretable model. While machine learning models are increasingly used for academic earlywarning systems, their lack of transparency limits real-world deployment. This work addresses the trade-off between predictive accuracy and explainability in learning analytics.

Materials and Methods:

The proposed framework is evaluated using the Open University Learning Analytics Dataset (OULAD). The modeling pipeline integrates student engagement traces, assessment behavior, and academic workload features to predict pass–fail outcomes. Two models are compared: Gradient Boosting, a black-box predictor supported by post-hoc SHAP explanations, and the Explainable Boosting Machine, a fully interpretable predictive model.

Results:

Gradient Boosting achieved the highest predictive performance with an accuracy of 0.892 and strong ROC–AUC discrimination. The Explainable Boosting Machine achieved comparable accuracy of 0.876 while improving interpretability. Although the Gradient Boosting model required post-hoc SHAP explanations to interpret predictions, the Explainable Boosting Machine provided transparent global shape functions and consistent local additive explanations. The results demonstrate that interpretable models can achieve near–black-box performance while providing clearer insight into the relationships among student engagement, assessment performance, workload, and academic outcomes.

Discussion:

The comparison highlights a critical accuracy–explainability trade-off between black-box and interpretable models. While Gradient Boosting offers marginal performance gains, its dependency on post-hoc explanations limits consistency and auditability. In contrast, the Explainable Boosting Machine demonstrates stronger transparency, stability, and pedagogical compatibility for real-world academic early warning systems.

Conclusion:

The results indicate that inherently interpretable models can achieve predictive performance comparable to high-performing black-box models while providing significantly greater transparency and interpretability to support educational decision-making. The suggested system facilitates the use of learning analytics in teaching and learning institutions by supporting early intervention, informed decision-making, and the ethical deployment of learning analytics. These findings highlight the potential of inherently interpretable machine learning models for trustworthy deployment in academic early-warning systems.