DOI: 10.3390/diagnostics16193106 ISSN: 2075-4418

Diagnostic Accuracy of Machine Learning-Enhanced Methods for Pulmonary Hypertension: A Systematic Review and Meta-Analysis

Faizan Ahmed, Ayesha Zulfiqar, Tushar Kanti Bhadra, Noor Ul Sabah, Taha Alam, Unaiza Iftikhar, Ahmad Zulaid, Hassan Farooq, Qais Bin Abdul Ghaffar, Muhammad Faizan Tahir, Muhammad Haris, Omar Kamel, Haris Bin Tahir, Gabriel Benjamen, Ali Abbas Khaleel, Ali Ghani, Amro Taha, Fawaz Alenezi

Background/Objectives: Right heart catheterization (RHC) is the diagnostic gold standard for pulmonary hypertension (PH), defined as mean pulmonary arterial pressure (mPAP) greater than 20 mmHg at rest. It is invasive and carries procedural risk. Machine learning (ML) and deep learning (DL) applied to non-invasive modalities may provide a non-invasive adjunct for screening, triage, and earlier identification. Methods: Methodological quality was assessed using QUADAS-2. Primary non-comparative pooling was restricted to RHC-confirmed studies using Reitsma hierarchical bivariate random-effects models. Comparative studies were analyzed using a paired hierarchical bivariate model, with between-method differences assessed using paired Wald tests. Sensitivity, specificity, likelihood ratios, and diagnostic odds ratios were estimated; predictive values were evaluated as secondary outcomes. GRADE certainty was assessed separately for non-comparative ML/DL, comparative ML/DL, and traditional methods. Results: Nine studies involving multiple pulmonary hypertension subtypes met the inclusion criteria: five non-comparative (single-arm) studies including approximately 26,530 participants and four comparative (double-arm) studies including 330 paired participants per arm. In the primary synthesis restricted to RHC-confirmed single-arm studies, ML/DL models achieved a pooled sensitivity of 0.720 (95% CI: 0.647–0.783) and specificity of 0.883 (95% CI: 0.796–0.936), with a diagnostic odds ratio of 19.4 (95% CI: 13.6–27.7). Estimates were similar when all five single-arm studies were included (sensitivity 0.716, 95% CI: 0.654–0.771; specificity 0.891, 95% CI: 0.820–0.937). In comparative studies analyzed with a paired hierarchical bivariate model, ML/DL outperformed traditional methods for both sensitivity (0.838, 95% CI: 0.668–0.930 versus 0.707, 95% CI: 0.421–0.889; paired Wald p = 0.019) and specificity (0.849, 95% CI: 0.512–0.968 versus 0.640, 95% CI: 0.376–0.840; paired Wald p = 0.0076). At a 25% pre-test probability, ML/DL was projected to yield 34 additional true positives (95% CI: −30 to 109) and 142 fewer false positives (95% CI: −381 to 129) per 1000 patients tested, although all absolute difference intervals crossed zero. Certainty of evidence was very low across all outcomes. GRADE certainty ranged from low to very low depending on outcome and stratum. Conclusions: ML/DL-based methods showed promising non-invasive diagnostic accuracy for PH, with significantly higher pooled sensitivity and specificity than traditional methods in paired comparisons. However, evidence certainty ranged from low to very low, and projected absolute benefits remained imprecise. These findings are hypothesis-generating; ML/DL tools may serve as adjuncts in the diagnostic pathway but cannot replace RHC at this stage.