From Explainability to Clinical Actionability in Multi-Modal AI for Cardiovascular Prediction: A Systematic Review
Hamza Nouri, Rafae Abderrahim, Mohamed ErritaliCardiovascular diseases (CVDs) remain the leading cause of mortality worldwide, driving the need for reliable tools for early risk prediction. Artificial intelligence (AI) applied to electrocardiograms (ECGs) has shown strong predictive performance for future cardiac events, yet its clinical adoption remains limited by the lack of transparency and trust associated with black-box models. This systematic review examines recent advances in AI-based cardiovascular prediction, focusing on the combined challenges of multi-modal data fusion and clinically actionable explainability. Following PRISMA 2020 guidelines, we analyzed 65 peer-reviewed studies published between 2018 and 2025, identified through a systematic search of PubMed, IEEE Xplore, Web of Science, Scopus, ACM Digital Library, and Google Scholar. The reviewed literature reveals that while most AI-ECG models achieve high predictive accuracy, typically AUC 0.85–0.95, the majority rely on post hoc explainability techniques that offer limited clinical insight, and 61.5% of included studies implement no explainability method at all. External validation remains critically underutilized, performed by only 12.3% of studies, and multi-modal approaches integrating ECG data with electronic health records, biomarkers, or genomics represent only 27.7% of the reviewed literature. While these multi-modal models demonstrate improved contextualization and predictive performance, they remain insufficiently validated and inconsistently interpretable. Among studies employing XAI techniques, attention mechanisms were the most prevalent approach (28% of XAI studies), followed by saliency maps (20%), SHAP (16%), and LIME (8%). Only 9.2% of studies were prospective or clinical trials, underscoring the gap between algorithmic development and real-world clinical deployment. Applying a pre-specified four-level clinical actionability scoring framework (Level 0–3), we found that the majority of studies (61.5%) scored at Level 0 (no actionability), with only 9.2% reaching Level 3 (demonstrated clinical impact), confirming that the clinical translation gap extends beyond trial design to encompass the broader absence of clinically contextualised evaluation of AI-ECG systems. This review highlights a persistent and critical gap between predictive performance and clinical usability, and outlines four key directions for developing AI-ECG systems that can better support trustworthy clinical decision-making: (1) developing inherently interpretable architectures, (2) advancing unified multi-modal fusion and explanation frameworks, (3) establishing standardized benchmarks for explainability evaluation, and (4) conducting robust prospective validation measuring real-world patient outcomes.