Autonomous Decision‐Making for Multi‐Spacecraft Orbital pursuit–evasion Games via Regularized MADDPG
Lanlin Yu, Congming Peng, Haopeng Wang, Wenjing RenABSTRACT
In this paper, an intelligent real‐time decision‐making framework for multi‐spacecraft orbital pursuit–evasion games is proposed. Based on relative orbital dynamics, the engagement process is formulated as a finite‐horizon nonzero‐sum Markov game. On this basis, the multi‐agent deep deterministic policy gradient (MADDPG) algorithm is employed to obtain the continuous‐thrust policies of the spacecraft, and a policy‐preservation regularization term is introduced into the actor update to constrain excessive policy deviations and improve the stability of adversarial training. Meanwhile, a multi‐scale reward function integrating terminal mission objectives and continuous process guidance is designed to enhance policy learning efficiency. Finally, numerical simulations verify the real‐time performance and effectiveness of the proposed method, providing a feasible approach for continuous‐thrust autonomous decision‐making in multi‐spacecraft cooperative–adversarial pursuit–evasion scenarios.