DOI: 10.2514/1.i011774 ISSN: 1940-3151

Safety and Performance Considerations in Shielded Autonomous Spacecraft Tasking

Lorenzzo Quevedo Mantovani, Hanspeter Schaub

This paper investigates the impact of shields on the performance of deep reinforcement learning (DRL)-based policies in the context of autonomous agile Earth-observing satellites. The scheduling phase during spacecraft operations determines the sequence of actions that the satellite should take to meet the mission requirements while adhering to the system’s constraints. The increasing number of satellites, the growing demand for their services, and the need for fast replanning due to unexpected events or additional requests push the need for methods that can provide solutions in real time. DRL has shown promise for onboard scheduling, but deploying neural networks in critical systems, such as satellites, raises reliability concerns. Research has proposed using shielded neural networks (SNNs) to provide safety guarantees with DRL. This work provides a comprehensive analysis of the impact of different approaches to training and deploying SNNs on agent performance. Results show that training with action replacement while assigning penalties for shield intervention leads to higher cumulative reward and lower shield interference during testing. Further, shielded agents performed similarly to the best unshielded agents but with fewer safety violations. Although shields provided safety in 30-times-longer-than-training episodes, policies trained with safety concerns outperform those trained without them.

More from our Archive