DOI: 10.3390/math14162894 ISSN: 2227-7390

Fairness Evaluation Paradox: How Biased Test Data Masks True Group Fairness Assessment

Sašo Karakatič, Ivona Colakovic, Tjaša Heričko

The EU AI Act makes fairness metrics for high-risk AI systems’ compliance evidence, turning their trustworthiness into a safety and accountability concern. Fairness audits assume that test data reflects the properties of real-world conditions, whereas standard evaluation protocols use test data drawn from the same biased records as the training data. Studies measuring how strongly this bias in test data distorts fairness metrics between the validation phase and real-world deployment are still very rare. We conduct an experiment on synthetic and real data, measuring this discrepancy across five fairness interventions on four unfairness types (1000 repetitions per combination, 20,000 total runs). Synthetic data lets us encode human bias in labels and compare fairness metrics on biased test labels (data available during development) against clean labels (conditions models face in deployment). We find that a systematic evaluation bias is present across all metrics, so the same models on the same test data can support opposite fairness conclusions and mask the mistreatment of the most disadvantaged groups. This pattern of fairness misevaluation is confirmed by a real-world validation on the Adult Census Income dataset. We conclude that trustworthy fairness auditing and regulatory standards should require bias-aware evaluation protocols, in which observed labels are not treated as ground truth.

More from our Archive