Garbage in, (mis)diagnosis out: making laboratory measurement bias visible and actionable across clinical algorithm outputs
Niels van Noort, Marith van Schrojenstein Lantman, Mathie P.G. Leers, Remy J.H. Martens, Volkher Scharnhorst, Marc H.M. Thelen, Arjen-Kars BoerAbstract
Objectives
Clinical algorithms increasingly combine multiple laboratory measurands to support diagnosis, prognosis, and treatment decisions. Yet, measurement bias arising from differences in calibration, reagent lots, or analytical platforms can propagate through such algorithms in non-linear ways, shifting outputs and altering clinical classifications. Existing external quality assessment (EQA) programmes evaluate measurands individually and do not capture the cumulative effect of bias on algorithm behaviour. We present a framework that extends EQA to quantify and communicate the impact of laboratory-specific measurement bias on clinical algorithm outputs.
Methods
Bias estimates derived from EQA results are injected into a representative patient cohort and propagated through the algorithm; when clinical outcomes are available, output shifts are translated into changes in clinically interpretable performance characteristics. Measurand-specific what-if analyses estimate the expected benefit of correcting individual bias. The framework is illustrated using the Fibrosis-4 (FIB-4) index and scalability is demonstrated on the ten-input CoLab COVID-19 severity score.
Results
Participant-facing reports were developed that summarise deviation from a zero-bias reference, position each laboratory within the peer distribution, and identify which input measurands offer the greatest potential for performance recovery. Applied to FIB-4, the reports demonstrate how laboratories can locate themselves relative to clinically desired performance thresholds and receive actionable, measurand-specific guidance to prioritise corrective efforts.
Conclusions
By translating bias-propagation analysis into interpretable laboratory reports, this framework bridges the gap between measurand-level analytical performance and clinical algorithm performance. It supports post-deployment monitoring, cross-laboratory comparability and verification, targeted quality improvement, and is applicable to any transparent formula-based clinical algorithm.