DOI: 10.3390/sym18101632 ISSN: 2073-8994

RECAP-RL: Symmetry-Guided Retention Shaping and Mirrored Credit Assignment for Personalized Instrumental-Practice Recommendation

Yanlu Li, Zhaoen Qu, Zhuodong Liu

Personalized instrumental-practice planning requires selecting and ordering exercises, allocating duration and difficulty, and accounting for prerequisites, fatigue, spacing, and forgetting across multiple sessions. This paper presents RECAP-RL, a symmetry-guided reinforcement-learning framework that post-trains a structured large language model (LLM) policy to generate budget-feasible, teacher-editable practice plans optimized for delayed retention. Symmetry enters at two levels. At the objective level, retention-anchored dense analytic reward shaping (RADAR) re-expresses terminal retention-adjusted learning gain (RALG) as telescoping differences in a retention potential based on the learner’s projected retained mastery; this re-expression preserves core-objective policy ordering and yields an action-independent retention-state baseline whose variance effect is characterized analytically. At the intervention level, matched-intervention replay with residualized order-graph rewards (MIRROR) pairs each exercise with an equal-duration null-practice twin under a common continuation; exchanging the twins reverses the signed retained and prerequisite-unlock contrast, and zero-mean residualization preserves the mean exercise-block advantage when credit is assigned to exercise-token blocks. We instantiate the framework for piano practice in a mechanistic environment with 72 skills and 520 exercises. Against eleven comparison planners on held-out and structurally shifted learner families, RECAP-RL reaches a RALG of 0.399±0.004, exceeding the strongest observable-state model-based planner by 0.034 and terminal-reward group-relative policy optimization by 0.027, while reducing mean cross-family degradation from 26.6% to 18.8% and improving prerequisite repair and spacing alignment. These simulation results support symmetry-guided reward and credit design for delayed educational planning; longitudinal human studies remain necessary to establish educational effectiveness.