DOI: 10.1055/a-2938-1280 ISSN: 0026-1270

Preventing Participant Leakage in Reissued Clinical AI Benchmarks: A Biomedical Informatics Evaluation Protocol and DAIC-WOZ/E-DAIC Audit

Minoru Hattori, Naoko Hasunuma

Background. Biomedical informatics increasingly reuses clinical artificial intelligence (AI) benchmarks, yet successor releases are often treated as independent external corpora. If the same participants reappear under new corpus names, features, or partition files, evaluation can become internal while reported as external. Objectives. To define successor-release participant leakage as a benchmark-integrity threat, operationalize a pre-evaluation audit and leak-free protocol, and demonstrate its impact using the Distress Analysis Interview Corpus, Wizard-of-Oz condition (DAIC-WOZ), and the Extended DAIC (E-DAIC) depression-screening benchmarks. Methods. We audited release lineage and participant identity using persistent identifiers, content hashing, Patient Health Questionnaire-8 (PHQ-8) label reconciliation, and fold-transition cross-tabulation. We compared leaky and participant-disjoint E-DAIC-to-DAIC-WOZ protocols on the same held-out participants using paired bootstrap inference, and derived conservative leak-free reference baselines across standard pipelines. Results. All 189 DAIC-WOZ participants reappeared in E-DAIC with byte-identical recordings and identical PHQ-8 totals. Across official partitions, 104 of 189 changed fold. Training on E-DAIC and evaluating on the DAIC-WOZ test set placed 47/47 test participants in the model-development pool; the reverse direction was clean. The leaky acoustic protocol yielded area under the receiver operating characteristic curve (AUROC) 0.797 versus 0.569 after deduplication, with paired Delta AUROC +0.227 (95% confidence interval +0.106 to +0.372; two-sided p < 0.001). At matched training size, Delta AUROC was +0.249. A generic visual classifier reached AUROC 0.887 on a blended external set but 0.562 on the unseen AI-only portion. Leak-free reference baselines were approximately AUROC 0.60. Conclusions. Successor-release participant leakage is a preventable evaluation failure in clinical AI benchmark reuse. Before pooling, external validation, or model comparison, biomedical informatics studies should document release lineage, verify participant identity at the signal or record level, reconcile labels, deduplicate across releases, and enforce participant-disjoint evaluation.

More from our Archive