A JSBSim Sensor-Interface Protocol for Selecting Learned Fixed-Wing Flight-Dynamics Surrogates
Yihao Feng, Jun Li, Sen Yang, Yulong JiAutonomous flight planners consume sensor-derived estimated states rather than simulator ground truth. Learned dynamics surrogates are queried over a planning horizon through architecture-native multi-step interfaces, so strong short-horizon accuracy can still yield unreliable long-horizon trajectory scores. We address this with a JSBSim sensor-interface evaluation protocol that assesses surrogate predictions from onboard sensing and lightweight estimation—inertial measurement unit (IMU), global navigation satellite system (GNSS)/air-data, barometric, and vertical-speed channels with delay, dropout, asynchronous refresh, and estimator filtering—and guides surrogate selection by a planning task. This paper contributes an evaluation-and-selection methodology for learned flight-dynamics surrogates used by autonomous planners; it is not a new flight-dynamics, guidance, or control method, and it does not propose a new neural architecture. The contribution is a transparent and auditable sensor-interface and estimator-conditioned protocol, instantiated on five surrogate families over a 20 s prediction horizon. We emphasize at the outset that under strict native-unit physical tolerances, all evaluated surrogates diverge on 90–100% of validation windows and none is field-ready; all results are comparative JSBSim stress-test evidence, not field-readiness evidence. Results reveal criterion-dependent ordering: long short-term memory (LSTM) networks are strongest on root-mean-square error (RMSE@1s/RMSE@20s), shared-threshold failure, and absolute fidelity, whereas the Transformer is favored under self-scaled divergence and capped-risk criteria. These aggregate scores combine heterogeneous simulator-coordinate units and are benchmark diagnostics rather than physical safety margins. Because different threshold families select different winners, our central conclusion is not that one surrogate is best, but that the choice of metric and threshold family is itself part of the benchmark assessment. In separately generated JSBSim candidate-screening episodes, Transformer has approximately 8% lower mean cost under abstract sensing corruption, whereas estimator-loop screening shifts toward LSTM; feedback-rich tracking remains exploratory at 40 episodes per scenario and has mixed cost and success endpoints. Cross-aircraft transfer degrades substantially. Surrogates should be assessed using sensing-aware, estimator-conditioned, task-specific long-horizon criteria, not short-horizon accuracy alone.