DOI: 10.31681/jetol.1991829 ISSN: 2618-6586

Modelling a dataset-defined skill-retention outcome in generative AI-assisted learning: An exploratory machine-learning and explainability study

Clinton Amponsah, Bernard Kyiewu, Linda Bessa-Simons
Generative artificial intelligence (GenAI) is increasingly used in higher education, but durable learning cannot be inferred from task completion alone. This study presents an exploratory machine-learning benchmark using a public dataset of 50,000 student-like records that is treated here as synthetic/engineered because its source does not document an empirical sampling frame, institution, country, recruitment process, response rate, or primary-study ethics procedures. The target field, Skill Retention Score, is defined in the source schema on a 0-100 scale as representing skills retained and applied after the semester; however, the source provides no assessment instrument, delayed-assessment interval, reliability estimate, or validity evidence. Accordingly, the variable is analysed as a dataset-defined proxy rather than a validated measure of long-term retention. Ridge Regression, Decision Tree Regression, and Extra Trees Regression were compared in the originally reported 80:20 hold-out analysis. On that split, Extra Trees produced R² = .185, RMSE = 11.983, and MAE = 9.696. Its RMSE is approximately 9.8% lower than the full-sample outcome standard deviation of 13.282, indicating modest predictive gain. The explainability output is permutation importance aggregated to the parent-variable level; it is interpreted as a model-internal, non-directional diagnostic, and SHAP results were not available in the submitted analytical record. The results therefore describe the structure of this benchmark dataset and the behaviour of the modelling pipeline; they do not establish effects of GenAI on real students, directional relationships, or instructional recommendations.