Dual‐View Representation Learning for Deployable Humanoid Whole‐Body Control
Haolin Song, Wengang Zhou, Mingxiao Feng, Houqiang Li
Humanoid whole‐body control policies receive limited proprioception at execution time, while simulation exposes richer privileged information during training. Such information can improve policy learning, but using it for deployable control requires transferring training‐only cues without changing the execution‐time sensing interface. While existing methods often rely on teacher distillation or explicit prediction of physical quantities, this work treats the transfer as a representation‐learning problem. To address this,
DVP
is introduced as a Dual‐View Proprioceptive (DVP) framework that turns privileged simulation states into training‐only latent supervision. For each simulated state, DVP pairs the privileged state with a view containing only the inputs available to the deployed actor. A dual‐cosine objective then aligns the privileged and deployable latent features while PPO optimizes the control policy. At execution time, the actor uses only deployment‐available inputs, and the privileged branch is removed. On LimX Oli velocity tracking and whole‐body motion tracking, DVP improves learning progress, control quality, action smoothness, and tracking accuracy over representative representation‐learning and privileged‐learning baselines. Cross‐platform motion‐tracking simulations on Unitree G1, Booster K1, and Noetix E1 test generality across morphology scales, while MuJoCo sim‐to‐sim evaluation and limited LimX Oli hardware demonstrations show execution through the deployable interface. Source code is available at