A systematic review of standards on bias and data quality from horizontal and vertical perspectives
Anna Schmitz, Maximilian PoretschkinThe recently adopted European AI Act mandates many AI providers to implement data quality and bias mitigation in their systems in order to safeguard fundamental rights, particularly non-discrimination. From a computer science perspective, however, the relevant requirements in the AI Act are not yet clearly linked to specific metrics or methods, highlighting the need for concrete interpretation within real-world applications. The ten harmonized standards which have been requested from the European standardization committee CEN/CLC JTC 21 are of particular importance for operationalizing the AI Act. Notably, the development of these standards is likely to leverage existing results from international standardization. In addition, the upcoming harmonized standards will inevitably interact with vertical regulations, e.g. of medical devices, and thus must be compatible with corresponding standards from the domains.
This paper presents a systematic review of all relevant ISO/IEC and IEEE standards to explore how the requirements regarding fairness and non-discrimination outlined in the AI Act can be operationalized on this basis. Following the Kitchenham method, we analyzed 41 horizontal standards and vertical standards from the medical domain in detail, using a coding scheme to capture the specifications contained in each standard for bias mitigation and data quality. Our analysis of the horizontal standards confirms two prominent trends: i) group-based bias measurement which, in the ISO/IEC standards, additionally focuses on the comparison of model accuracy across groups, ii) a broad array of bias mitigation approaches along the AI lifecycle, surpassing the data focus of the AI Act. Furthermore, the analysis shows the importance of both horizontal and vertical standardization. Interestingly, the bias concept is less formalized and primarily linked to the data in the latter. When it comes to operationalizing the AI Act, what can be learned from vertical standards is the institutionalization of ethical requirements (independent of organizational goals) and the constant reference to intended use, which is explained here with concrete examples and represents a central link to the horizontal standards. However, we also identified several weaknesses such as inconsistencies across different standards as well as overarching gaps, such as a lack of instructions for (pre-trained) generative AI models. In conclusion, by giving a comprehensive overview of the current standardization landscape regarding bias and data quality, pointing out weaknesses therein and possible ways to address these, our review serves as a valuable resource for current standardization efforts in support of the European AI Act.