DOI: 10.59668/2761.28133 ISSN:

Decision Trees and Random Forests in Educational Data Analysis

Hailey Kuang, Jiawei Xiong, Bowen Wang, Cheng Tang

Decision trees and random forests are supervised learning methods used for both classification and regression tasks (Breiman, 2001; Quinlan, 2014). A decision tree partitions data based on a sequence of decision rules that form a tree-like structure, while a random forest extends this approach by constructing multiple trees and aggregating their predictions to improve accuracy and reduce overfitting. To illustrate, consider an online learning platform where students work through instructional materials and practice activities. While many students remain focused, others may become distracted, such as opening social media, pausing for daydreaming, or rapidly clicking without reading. Detecting these off-task moments is valuable because it allows teachers or the system to offer timely reminders or adjust the learning design. To identify when a student may be off-task, we can examine behaviors such as time spent on each activity, use of hints, and the number of actions completed within short time windows. However, the relationships among those behaviors are often non-linear and vary by context, making their meaning not always straightforward. For example, a tree might identify that a combination of very short response times and frequent rapid actions indicates a high likelihood of off-task behavior. In educational technology research, decision trees and random forests are widely used to analyze complex data, translating meaningful patterns into actionable insights that support instructional design and student learning (Lin et al., 2013; Liu et al., 2022; Matzavela & Alepis, 2021; Rizvi et al., 2019). Tree-based models are particularly well-suited for this domain because they could capture complex non-linear relationships among features while producing transparent decision rules that educators can easily interpret.