DOI: 10.3390/a19080676 ISSN: 1999-4893

Decision Reliability Profiles for Auditing Targeted Case-Removal Sensitivity in AI Model Comparisons Under Additive Metrics

Khudran M. Alzhrani

Aggregate model-comparison metrics identify which model performs better on average, but not whether that result is broadly supported or depends on a small number of high-impact cases. This study presents Decision Reliability Profiles (DRP) for fixed-prediction comparisons under additive lower-is-better metrics. DRP centers on the Pairwise Reversal Budget (PRB), the minimum number of winner-supporting cases whose removal would produce a tie or make the competing model win. The normalized PRB (nPRB) expresses this budget as a fraction of the evaluation-set size. PRB is a targeted deletion diagnostic conditional on the observed evaluation set and does not estimate sampling uncertainty. The evaluation covered 45 classification and 25 regression OpenML tasks, yielding 9250 pairwise comparisons. Low nPRB values occurred under every metric and were especially common for unique top-vs-runner-up comparisons. Across the probabilistic classification and regression metrics, the proportions of comparisons with nPRB at or below the descriptive 5% threshold ranged from 21.76% for mean absolute error to 43.82% for Brier score. For zero-one error, nPRB equals the absolute empirical error-rate gap; the 72.61% rate therefore represents gaps of at most five percentage points. Low-nPRB comparisons overlapped only partly with paired statistical tests and bootstrap intervals. DRP complements aggregate and statistical reporting by showing how individual evaluation cases support the observed model-comparison result.

More from our Archive