DOI: 10.3390/info17080786 ISSN: 2078-2489

H-StreamQ: An Entity-Aware Framework for Data Quality Assessment and Drift Monitoring in Electronic Health Records

Gul Muhammad Soomro, Zaira Hassan Amur, Said Krayem, Bronislav Chramcov, Roman Jasek, Ismail Nooraddin Ismail Allahwerdi

Entity-aware quality assessment may reduce false interpretations of electronic health record (EHR) data, but evidence from small, rule-aligned benchmarks cannot establish operational effectiveness. We revised H-StreamQ as a proof-of-concept framework and evaluated its laboratory component using the complete MIMIC-IV v3.1 labevents file (158,374,764 events; 313,442 patients). Ten thousand patients were sampled across laboratory-activity quintiles and split at patient level into training (6000), threshold-calibration (2000), and test (2000) groups. The independent test set contained 918,651 numeric laboratory events. Without excluding naturally alerted records, 54,788 mutually exclusive defects were introduced using subtle value shifts, unit/scale errors, mapping errors, delayed records, and patient-clustered correlated defects. Rules, a context-aware Isolation Forest, their union (Hybrid), a context-free Isolation Forest, Local Outlier Factor (LOF), and linear and radial-basis-function (RBF) One-Class support vector machines (OCSVMs) were compared at a threshold fixed by a 2.5% calibration alert budget. Patient-cluster bootstrap intervals and event-micro and patient-macro results were reported. Rules alone achieved the highest event-micro F1-score (0.637; 95% confidence interval [CI] 0.547–0.722), followed by Hybrid (0.576; 0.484–0.668) and RBF One-Class SVM (0.559; 0.433–0.670). Hybrid increased recall over rules by only 0.004 (95% CI 0.003–0.006) while reducing F1 by 0.061 and increasing the background-alert rate by 0.015. Context conditioning did not improve aggregate Isolation Forest performance. In six batch-level drift simulations, an exponentially weighted moving average (EWMA) and a fixed-window monitor detected 97–100% and 98–100% of changes, respectively, whereas a custom Hoeffding adaptive-window detector was more conservative and often missed smaller or recurrent changes. These results support H-StreamQ as an explainable research framework, not as a validated clinical or production system. Patient-macro F1, which weights every patient equally, was substantially lower than event-micro F1 for every method (rules 0.395 versus 0.637; Hybrid 0.320 versus 0.576), indicating that event-level performance is weighted towards high-activity patients. Precision and F1 are computed relative to injected synthetic labels and are not clinically adjudicated estimates. The entity-aware architecture spans patients, admissions, diagnoses, transfers, and dictionaries, but the quantitative detection benchmark evaluates the numeric laboratory component only; other entities are used for linkage and contextual attachment and are audited descriptively rather than evaluated against labels.

More from our Archive