Robust heart murmur detection for AI digital stethoscopes: An evaluation protocol using the PhysioNet 2016 open PCG dataset
Marko Milić, Šćepan Sinanović, Dejan Kostić, Nemanja Đenić, Branislav RalićAI-enabled digital stethoscopes can standardize auscultation, but real-world success depends on robustness across devices, sites, and noise. Public, de-identified corpora permit rigorous, ethics-exempt development and transparent benchmarking. To design and describe a reproducible evaluation workflow for robust murmur detection tailored to AI digital stethoscopes, using the open PhysioNet/CinC 2016 heart-sound dataset, with a protocol that prioritizes generalization to unseen sources/devices. We prespecify training/validation splits within the public CinC-2016 training pool and construct an external test set by holding out entire sources/devices (leave-source-out). The proposed modeling pipeline fuses 1D waveforms and 2D Mel-spectrograms in a late-fusion CNN/Transformer, with acoustically targeted augmentation and class-imbalance handling. The primary outcome measure is AUROC on the external test; secondary outcomes include AUPRC, sensitivity/specificity, F1, and calibration (Brier/ECE). We contextualize the planned performance against DOI-indexed CinC-2016 benchmarks and present non-redundant figures (class prevalence, ROC-space summary, multi-metric profiles). All data are de-identified and publicly available; no ethics approval or consent was required. The official CinC-2016 corpus comprises 3,153 training recordings (764 subjects) across six sources and 1,277 hidden-test recordings (308 subjects). Published hidden-test benchmarks range from modified accuracy 0.84 to 0.86 with varying sensitivity/specificity trade-offs. Our protocol operationalizes these constraints into a transparent, device-aware evaluation that can be replicated and extended without access to the hidden test. A prespecified, source-held-out evaluation on a public corpus, paired with calibrated, decision-oriented reporting, provides a strong foundation for robust AI auscultation and prepares the ground for prospective, multi-site validation.