Causal Benefit-Aware Recommendation for Personalized Learning-Path Features: A Targeting-Policy Framework with Provable Guarantees and Randomized Evaluation
Yanfen Huang, Lin Wang, Weihua Bai, Teng Zhou, Xinyang WangEducational platforms increasingly personalize which AI learning-path features (adaptive homework, learner choice) each student receives. The natural correlational baseline ranks students by predicted performance—deliver the feature to those expected to do well—a heuristic that need not identify who actually benefits. We formalize feature recommendation as a causal targeting-policy problem: rank students by the estimated conditional average treatment effect (CATE) of a feature and recommend to the top of the ranking. We prove three results: (i) causal top-CATE targeting maximizes policy value at any budget and weakly dominates predictive (outcome-based) targeting, strictly when the two rankings disagree; (ii) a split-sample doubly robust evaluation of targeting quality is leakage-free (null-exact in finite samples), whereas the naive in-sample version is optimistically biased; and (iii) greedily targeting by CATE traces the optimal cost–benefit (Qini) frontier, with the deployment rule “recommend when τ^>0.” We validate the method on 17 randomized embedded experiments from the ASSISTments platform. Because the 40 held-out splits re-partition the same students, we do not treat them as independent replicates: we calibrate every headline comparison against a within-experiment permutation null and an experiment-clustered bootstrap. Under that calibrated inference, causal targeting outperforms predictive targeting for adaptive homework at every budget (permutation p≤0.005, the resolution floor of 200 replicates; Holm-corrected p≤0.040), while for learner choice the same contrast is directionally consistent but not statistically significant (permutation p=0.23–0.38; clustered p=0.42). The direction is stable in both families: no leave-one-experiment-out refit reverses its sign. Predictive targeting is nonetheless the one rule that is reliably worse than the alternatives, because it recommends the feature to high-performing students who benefit least—realized benefit falls monotonically across predicted-performance deciles (from +0.087 in the lowest to −0.030 in the highest). Against a fuller baseline suite, causal (CATE) targeting does not beat random, a simple risk-based rule (target low performers), or treating everyone. Indeed, the estimated benefit ranking is close to noise—split-half rank agreement is ρ≈0.002–0.008 and its calibration slope is 0.018, far below the ideal of 1—so the gain over predictive targeting comes from avoiding an actively harmful ordering rather than from recovering individual benefit. A fairness analysis shows why this matters: predictive targeting is regressive, concentrating feature access on high-ability students, whereas causal and risk-based targeting reverse that gradient in this corpus; no policy differentiates by neighborhood opportunity zone. Group-conditional policy values, however, are not individually distinguishable from zero once dependence across students and experiments is accounted for; what survives resampling is the allocation itself—predictive targeting directs 0.33 fewer of its recommendations to low-ability than to high-ability students (95% CI [−0.46,−0.01], experiment-clustered)—so we frame the fairness result as improved access, not established equity gains. The actionable finding is therefore narrow and specific: outcome-based targeting systematically mis-allocates learning-path features and should be replaced by some benefit-aware rule; whether that rule needs to be a learned CATE model, rather than a simple risk-based heuristic, is not established by this corpus.