DOI: 10.3390/electronics15153457 ISSN: 2079-9292

The Explanation Cost of Fairness: How Bias Mitigation Affects Explanation Faithfulness in Machine Learning Intrusion Detection

Khalid Alalawi

Machine learning intrusion detectors are accurate but opaque, so explainable AI justifies their alerts and bias mitigation evens out detection across imbalanced categories. These are usually studied independently, and whether making a detector fairer changes how faithfully it can be explained is unknown. In an interventional design, we train an unmitigated baseline and three imbalance mitigations on a one-dimensional CNN, an FT-Transformer, and XGBoost, applying class weighting and oversampling to all three and class-weighted focal loss to the neural models. Faithfulness is measured with deletion-based comprehensiveness and sufficiency, before and after each mitigation, per category, on CICIoT2023 and CICIDS2017. Bias mitigation affects faithfulness unevenly. On the primary dataset, the rare categories most helped gain recall and more faithful explanations, while the largest penalty fell on a category whose detection scarcely changed. XGBoost, explained by exact TreeSHAP, showed no comprehensiveness cost of its own. Among the neural mitigations, focal loss was the most costly, lowering comprehensiveness by up to 0.13 while reducing the recall parity gap by 0.42–0.45, and oversampling was nearly free; this pattern held on both datasets, more weakly on the second, where class weighting and oversampling were not reliably separated. Fairness and interpretability need not be in tension when the mitigation is chosen well, and faithfulness should be evaluated per category whenever a fairness intervention is applied.

More from our Archive