Support-Constrained Conservative Reranking for Personalized Learning Activity Plans Under Temporal Distribution Shift
Yuan Ren, Zhanfang Chen, Zeming Du, Xiaoming JiangThis study evaluates whether an offline system can conservatively rerank future learning activity profiles under temporal distribution shift; it does not test whether an intervention improves actual student learning. An activity profile comprises temporally binned activity type proportions, click intensity, and active bin indicators. A multilayer perceptron (MLP) predicts a base profile from the first 30% of a course, seven local residual candidates are transferred from similar training students, unsupported candidates are excluded, and a ridge outcome model fitted with inverse probability weighting (IPW) estimates simulator-defined utility. The proposed support-constrained bootstrap lower quantile rule (SC-LQ) switches only when the empirical 10th percentile of a candidate’s bootstrap gain distribution is positive; this quantity is a ranking statistic, not a calibrated confidence bound. Experiments used five Open University Learning Analytics Dataset (OULAD) courses and frozen semi-synthetic potential utilities. Under natural decision rules, SC-LQ switched for 47.5% of development students and 52.2% of holdout students, whereas IPW argmax switched for 81.8% and 86.9%, respectively. At exactly matched switching coverage, SC-LQ did not significantly improve the mean simulator value over IPW (paired difference 0.00065, 95% confidence interval (CI) [−0.00030, 0.00161], Holm-adjusted p = 0.135) but reduced overall harm by 0.01533 and utility loss above 0.01 by 0.02494. Candidate-level empirical coverage of the q10 (10th-percentile) statistic was only 0.763 with 200 bootstrap models, confirming that SC-LQ provides empirical risk ranking rather than a safety guarantee. The value–risk pattern persisted across temporal resolutions, candidate set sizes, validation-selected anchors, and several outcome and utility models but failed or weakened in deliberately discontinuous and misspecified settings. These findings support conservative fallback as a simulator-tested risk control principle for active OULAD learners, not as evidence of causal learning improvement.