Compliant Transfer Method for Space Manipulators Based on αPPO-LAG
Lianpeng Li, Di Xin, Mingyang Li, Haibo Zhang, Shuanfeng Xu, Donghao ZhangTo address challenges associated with high-dimensional coordination and strict physical constraints in fixed-base space manipulator tool transfer under zero gravity, this paper proposes a preference-conditioned safe reinforcement learning algorithm, which is called αPPO-LAG (α-Preference Proximal Policy Optimization with Lagrangian). The algorithm introduces a preference factor α to regulate the trade-off between task performance and physical safety through constraint-aware policy optimization. Meanwhile, an adaptive Lagrangian regulation mechanism based on constraint estimation is developed to improve safety satisfaction during training. Experimental results demonstrate that with α=0.5, the proposed method achieves an average reward of 990, outperforming PPO-LAG with 881 and CPO with 915. Furthermore, αPPO-LAG obtains a safety score of 0.935 while reducing the contact-force violation rate to 1.0% and maintaining a low end-effector velocity violation rate of 10.0%. The empirical safety-performance analysis reveals effective operating points under different safety preferences, providing a practical solution for safe and compliant fixed-base tool transfer in zero-gravity environments.