DOI: 10.1177/03000605261486667 ISSN: 0300-0605

Normalized risk-based evaluation of machine learning–based classification models: A multiclass approach with applications in medical devices

Julius Wiggerthale, Martin Haimerl, Christoph Reich

Objective

Machine learning is increasingly integrated into safety-critical domains such as medical applications. In this context, regulatory frameworks require the assessment and minimization of risks associated with incorrect model predictions. However, classical evaluation methods often focus on quantifying error frequencies, which do not reflect the heterogeneous impact of different types of errors. To address this deficiency, we elaborate an approach for assessing machine learning–based multiclass classification models that conforms to regulatory requirements in the field of medical devices. Additionally, we aim to demonstrate its applicability in a concrete application scenario as an exemplary reference.

Methods

Herein, we provide weighted balanced accuracy W B A n as a risk-based approach for assessing machine learning models for multiclass classification. W B A n is systematically based on regulatory requirements for medical devices. It incorporates core properties such as the representation of risks as a multiplicative combination of likelihood and severity and the relationship between development and real-world scenarios. We demonstrate the application of W B A n using X-ray based detection of lung diseases as a reference scenario.

Results

Our analysis shows that W B A n achieves a risk-based assessment in the lung disease scenario as a reference. W B A n systematically operationalizes the definition of risk from corresponding regulatory requirements and allows the assessment to be adapted according to risk levels that have to be elaborated in the context of the risk-management process.

Conclusion

Although this study uses only a single use case to demonstrate the practical applicability of our approach, it shows the impact of a risk-based approach when assessing the performance of machine learning multiclass classification models in medical applications. Without such an approach, regulatory requirements cannot be implemented in a fully comprehensive way. Consequently, the clinical impact cannot be adequately assessed when applying the model in real-world scenarios.