DOI: 10.21449/ijate.1802844 ISSN: 2148-7456

Predicting early university dropout in open and distance education: A comparison of data mining models

Selma Tosun, Dilara Bakan Kalaycıoğlu
While open and distance education programs provide access to higher education, they face higher dropout rates than conventional campus-based programs. Early identification of students at risk is a critical step in supporting them. The existing literature on dropout has been quite limited. Despite the potential of Educational Data Mining (EDM) to transform student retention strategies, there is a noticeable lack of studies evaluating algorithm performance on national-level datasets. In particular, the use of data mining techniques in educational systems with big data has not been sufficiently investigated. This study fills the gap by comparing eight different classification algorithms, including various Decision Trees and Artificial Neural Networks (ANN), on a massive dataset of 650,317 students. Focusing on naturally occurring early-stage data, this research aims to determine which models provide the highest prediction validity in the context of big data. The results show that the C5.0 algorithm has the highest classification accuracy, and the ANN provided the most robust discriminative power with an AUC of .867. The findings have the potential to contribute to actions aimed at preventing school dropout in open and distance education.

More from our Archive