Validating Online Guidance for Reusability Flight Experiment: Genetic Programming and Reinforcement Learning
Riccardo Cadamuro, Francesco Marchetti, Jose Luis Redondo Gutierrez, David SeelbinderIn this study, two machine learning (ML) techniques—genetic programming (GP) and deep reinforcement learning (DRL)—are leveraged to derive a reentry guidance law and are evaluated using a high-fidelity six-degree-of-freedom simulator developed for validating the Reusability Flight Experiment (ReFEx) vehicle guidance, navigation, and control (GNC) subsystem. Both methods are benchmarked against each other and against the baseline ReFEx optimization-based guidance strategy to assess their applicability to a real mission and to understand their respective strengths and weaknesses. The ReFEx mission focuses on the reentry phase, wherein GP and DRL models are applied to generate real-time corrections to precomputed reference guidance commands, thereby compensating for external disturbances and model inaccuracies. GP is selected for its ability to produce human-readable, continuous, and differentiable models that yield smooth guidance commands, whereas DRL employs a fully connected neural network (NN), delivering superior performance at the expense of a black-box model and nonsmooth guidance signals. The results demonstrate that DRL and GP achieve performance comparable to the mission baseline guidance, hence validating their applicability in mission-grade applications. Moreover, both ML approaches achieve faster online execution times than the baseline guidance method.