Simulated Corrective Subgoal Supervision for Hierarchical Reinforcement Learning in Long-Horizon AntMaze Navigation
Lidong Sun, Ye Wang, Zheheng Fan, Fuchun SunLong-horizon navigation requires a high-level policy to select locally reachable subgoals, yet a scalar task reward provides little information about how an unsuitable proposal should be changed. We introduce Simulated Corrective Subgoal Supervision for Hierarchical Reinforcement Learning (SCS-HRL), a two-level method in which a topology- and clearance-aware programmatic supervisor evaluates each proposed subgoal and returns both a scalar score and a continuous target in the same subgoal space. The score trains the high-level critic, and the target enters a masked regression term for the high-level actor. Primitive actions are always conditioned on the actor’s subgoal; the supervisor is inactive during learned-policy evaluation. In AntMaze, using 6000 training episodes, five seeds, and 100 deterministic evaluation episodes per seed, SCS-HRL attained an 88.4±7.8% final success rate (mean ± sample standard deviation; 95% Student-t confidence interval [78.7%,98.1%]). The matched scalar-only condition and HIRO attained 0% rates. Applying the same route rule directly to the SCS-HRL low-level controllers yielded 82.2±9.9% success; the paired difference favored the learned high-level policy by 6.2 percentage points (95% confidence interval [2.1,10.3], p=0.013). Across three matched seeds, nonzero corrective weights of 0.5, 1.0, and 2.0 remained stable, whereas 0.25 was seed-sensitive. Term-level ablations further show that the continuous target, rather than the exact scalar-shaping formula, was the principal additional signal. Separate fixed-policy tests obtained 0% success rates on two unseen maze layouts. These results indicate that continuous subgoal targets can encode task-specific route information in the source maze, while cross-layout transfer remains unresolved.