Risk-Aware Hybrid Decision-Making Combining Reinforcement Learning and Receding Horizon Control for AUV Bistatic Sonar Target Tracking
Weicong Zhan, Yu Tian, Feng Zheng, Jiancheng Yu, Yan HuangReinforcement learning (RL) can guide autonomous underwater vehicle (AUV) maneuvering to improve the relative source–target–receiver geometry for bistatic sonar target tracking. However, learning a reliable RL policy typically requires substantial interactions with the environment. This paper proposes a risk-aware hybrid decision-making framework that combines an RL policy with selective receding horizon control (RHC) to reduce policy training requirements. Specifically, soft actor-critic (SAC) serves as the nominal decision maker, while a tracking-risk detector assesses target existence probability and estimation uncertainty. When a high-risk belief state is identified, RHC temporarily overrides the SAC action and performs finite-horizon planning based on predicted tracking uncertainty and acoustic detectability. Monte Carlo tree search is employed to efficiently solve the resulting planning problem. Numerical simulations show that the proposed framework improves tracking performance and reduces target-loss events under limited SAC training budgets. In particular, the hybrid framework using a SAC policy trained for 160,000 interaction steps achieves tracking performance comparable to that of pure SAC trained for 500,000 steps, while requiring only sparse RHC intervention. These results demonstrate that selective online planning provides an effective trade-off among policy training requirements, online computational cost, and target tracking performance.