DOI: 10.54569/aair.1998842 ISSN: 2757-7422

Duplicate Clinical Profiles as a Source of Data Leakage in Heart Disease Classification: A Benchmark Audit of Validation, Calibration, and Explainability

Oğuzhan Kilim
Duplicate clinical profiles can place the same observation in both training and test partitions, producing optimistic performance estimates that do not represent generalization to independent cases. This benchmark audit examined a widely used heart disease dataset containing 1025 rows and 13 predictors. Data-integrity analysis identified 302 unique clinical profiles and 723 redundant rows. Logistic Regression, Gaussian Naive Bayes, k-Nearest Neighbors, Random Forest, Extra Trees, and Histogram-Based Gradient Boosting were evaluated under naive row-level, group-safe, and deduplicated repeated five-fold cross-validation. Performance was assessed at the profile level using classification, discrimination, and calibration measures. Under naive validation, k-Nearest Neighbors, Random Forest, and Extra Trees achieved accuracy and ROC-AUC values of 1.000. After deduplication, their accuracy ranged from 0.8245 to 0.8377. For Random Forest, naive validation inflated accuracy by 0.1722 and ROC-AUC by 0.0950. In the leakage-free evaluation, Logistic Regression achieved the highest ROC-AUC of 0.9128 and the lowest Brier score of 0.1142. Sigmoid calibration reduced the Random Forest expected calibration error from 0.0563 to 0.0384. A nested cross-validation sensitivity analysis showed modest and inconsistent changes after hyperparameter optimization, with no statistically significant improvement across the outer folds. Permutation importance identified thal, cp, and ca as the most influential predictors. These findings demonstrate that data integrity and the definition of the independent observation unit should be examined before near-perfect clinical machine-learning performance is interpreted as algorithmic superiority.

More from our Archive