Beyond Accuracy: A Reliability-Oriented Multi-Dimensional Benchmark of Data-Driven Fault Diagnosis for Liquid Rocket Engines
Mingyang Geng, Gang Zheng, Long He, Siyu Zhao, Xiuwei YuLiquid rocket engines (LREs) are mission-critical propulsion systems whose reliability directly affects launch safety and the repeated operation of reusable launch vehicles. Data-driven LRE fault diagnosis remains difficult because labeled fault samples are scarce, engine-to-engine variations cause distribution shifts, and existing evaluations frequently emphasize classification accuracy without controlling for input representation. This study establishes a reliability-oriented benchmark that evaluates seven representative machine-learning and deep-learning methods from four dimensions: diagnostic capability, sensor informativeness, cross-unit generalization, and data efficiency. The experiments use 1106 fault samples from the XJTU-REF dataset, covering nine fault categories, 19 sensor channels, and three simulated engine units. To ensure an apples-to-apples model comparison, all seven methods are first evaluated using the same 19-dimensional steady-state representation. Under this controlled setting, Random Forest achieves the highest accuracy of 89.15% and Macro-F1 score of 88.74%, while the best deep model, 1D-CNN, reaches 86.42%. An additional representation-ablation experiment shows that concatenating the same steady-state features with learned sequence representations increases 1D-CNN accuracy to 87.58%, reducing its gap from Random Forest to 1.57 percentage points. This result demonstrates that the large difference observed under the original heterogeneous-input setting is partly attributable to representation design rather than model family alone. In cross-unit evaluation, Random Forest obtains an average accuracy of 75.37%, whereas Transformer obtains 61.54%, indicating that engine-to-engine variation remains a major deployment challenge. The top five sensors contribute 47.62% of the total Random Forest importance, and, with five labeled samples per fault category, KNN and Random Forest achieve 71.24% and 70.82% accuracy, respectively. These findings provide quantitative guidance for algorithm selection, sensor prioritization, and data-acquisition planning in reusable launch vehicle health-monitoring systems.