DOI: 10.3390/en19194522 ISSN: 1996-1073

Condition-Aware ARIMA Cleaning of Substation Equipment Time Series with Local Residual Scaling and Audit Records

Yin Zhang, Guowei Li, Junbo Wang, Qi Tang, Jintao Yu, Xianglong Gu, Zhipeng Li, Xiaozhen Zhao

Distinguishing data quality defects from genuine operating-condition changes is a central challenge in cleaning substation equipment time series. This study extends the double-loop AO/IO framework of Yan et al. by adding local residual scaling, retain-first state screening, constrained correction, model-conditional prediction intervals, and record-level audit fields; the double-loop structure itself is not claimed as new. Evaluation used six deliberately selected field-derived series (two from each of three data sources) with reproducible synthetic labels. The prior correction cap, which allowed corrections to at most 5% of samples, was disabled. Across 18 representative paired runs and 516 injected points, the proposed method obtained F1 = 0.9647 ± 0.0337, whereas fixed-threshold AO/IO obtained 0.9728 ± 0.0308; the six-series paired comparison did not establish a difference (t-test p = 0.0707; Wilcoxon p = 0.1250). In a constructed heteroscedastic stress test, mean F1 increased from 0.4675 to 0.4941 and low-volatility recall from 0.1458 to 0.2667, but this scenario was designed around variance changes and contains only six independent source series. A second spike-only test retained the naturally varying volatility of measured segments and produced F1 = 0.4835 versus 0.4345, but the series-clustered difference was not significant (p = 0.4489). The evidence therefore supports volatility adaptation and traceable decision output, not universal accuracy superiority or field-proven effectiveness. Prediction intervals are approximate and conditional on the fitted model.