DOI: 10.3390/fi18090498 ISSN: 1999-5903

A Subsystem-Level Validation and Simulation Framework for a 12-DoF Biped Robot with Deep Reinforcement Learning Locomotion

Michael Felipe Cifuentes-Molano, Kevin David Ortega-Quiñones, Byron Hernandez, Germán Andrés Holguín-Londoño, Mauricio Holguín-Londoño

Simulation-based reinforcement-learning locomotion depends on the physical fidelity of the underlying model. This work presents a subsystem-level modelling, validation, and control framework for a 12-DoF biped robot, combining Denavit–Hartenberg kinematics, Euler–Lagrange dynamics, a Discrete Euler–Lagrange reference integrator, Hunt–Crossley contact, and Soft Actor-Critic training in PyBullet. Validation is scoped. For the fixed-hip leg, numerical damped-least-squares inverse kinematics achieved a round-trip error of 0.017 ± 0.022 mm, while a gravity path-integral test produced a residual of 0.006 J. On a one-DoF reference problem, DEL bounded energy error under a coarse-step stress test, whereas at the 1 ms training step, RK4 was more accurate; no RL-scale DEL advantage was established. Contact realism remained inconclusive because the available prescribed-penetration analysis was not a dynamically consistent whole-body impact test. The same nominal parameters were used in PyBullet for locomotion training, without establishing numerical equivalence between the two simulators. Across three asymmetric-reward runs, forward walking dominated final evaluations, but sustained velocity ranged from 0.62 to 1.24 m/s under unequal training budgets. An exploratory hybrid architecture reached 2.38 m/s in one run without controlled ablation. These results demonstrate subsystem-level diagnostics while identifying full-body validation, contact calibration, equal-budget replication, and architectural ablation as necessary future work.