DOI: 10.3390/pr14193049 ISSN: 2227-9717

Predicting Engineering-Defined Failure Risk Levels of Metal-Loss-Related Pipeline Girth Weld Defects Under Incomplete Engineering Data

Ke Wang, Min Zhang, Dan Chen, Weifeng Ma, Shuai Zhang, Wei Wu, Weixin Gao

Predicting engineering-defined failure risk levels for pipeline girth welds …is challenging because engineering records may be incomplete Conventional models often struggle with incomplete engineering inputs, as standard imputation pipelines tend to treat estimated values as equally reliable observations, losing critical missingness patterns. To address this, we propose a missingness-aware tabular deep learning method. Our framework integrates feature values, missingness masks, and feature-confidence scores into a unified representation to distinguish observed and incomplete inputs. Architecturally, the method leverages missingness-aware feature representations through mask and confidence embeddings to learn nonlinear feature interactions directly from incomplete samples. To ensure stability on small, imbalanced datasets, a parameter-efficient ensemble structure with a shared representation layer and multiple lightweight prediction branches is designed. During training, a structured-missingness-robust strategy simulates realistic systematic missing patterns common in engineering records, such as missing material-parameter or defect-size groups. At the output stage, the model provides risk classification probabilities alongside predictive uncertainty (derived from branch disagreement or probability entropy) to flag boundary cases for manual engineering review. Experimental results on a highly imbalanced dataset, in which high-risk samples account for only 0.97% of the observations, demonstrate a more balanced trade-off between overall classification performance and high-risk sample identification compared with representative baselines under the engineering-defined risk classification setting. Under five-fold stratified cross-validation, the model achieved an accuracy of 91.70% ± 1.94%, a Macro-F1 score of 66.54% ± 3.24%, a balanced accuracy of 79.74% ± 11.75%, and a high-risk recall of 52.00% ± 37.76% across folds, corresponding to 51.85% in pooled out-of-fold prediction. Predictive entropy provided supplementary information for prioritizing potentially unreliable predictions, although the effectiveness of different uncertainty indicators varied.