DOI: 10.3390/machines14101124 ISSN: 2075-1702

Leakage-Controlled Grouped Validation and Interval-Safety Calibration for RUL/SOH Prediction Using Engine, Battery, and Bearing Degradation Data

Adel BenAbdennour

In predictive maintenance, low average error is not enough. Serious mistakes can occur near intervention thresholds when a model understates degradation despite acceptable root mean squared error (RMSE) or mean absolute error (MAE). This paper presents a leakage-controlled grouped validation framework for remaining useful life (RUL) and state of health (SOH) prediction across NASA C-MAPSS, NASA Battery, PRONOSTIA/FEMTO, XJTU-SY, and IMS data sets. The gate requires global and urgent/critical coverage of at least 0.90, zero false-safe rate, zero interval-level intervention miss rate, and underwarning rate no higher than 0.05. Across 30 predefined grouped seeds per dataset, the gate-first search resolves all 150 dataset-seed cases; because selection is restricted to eligible candidates, gate satisfaction of those selected rows is selection-conditioned rather than an independent post-selection test. An augmented sharpness/specificity gate requiring mean interval width W¯≤0.50 and nonurgent lower-bound intrusion Owarn≤0.10 resolves 30/30 seeds for IMS, PRONOSTIA/FEMTO, and XJTU-SY, 24/30 for NASA Battery, and 8/30 for the 65-candidate C-MAPSS pool. Point-level intervention misses occurred, especially for NASA Battery and C-MAPSS. The results support empirical interval safety under the primary gate while exposing dataset-dependent sharpness/specificity limitations; the claims remain limited to grouped benchmark evidence and do not imply per-asset coverage guarantees.