DOI: 10.4103/jcls.jcls_235_25 ISSN: 2468-6859

Artificial intelligence for critical decision-making: A systematic review of discriminative performance and clinical implementation in emergency medicine

Yatrik Vasavada, Aniket Patel, Aditya Pundkar

ABSTRACT

Background:

Emergency departments (EDs) face increasing pressure from patient volumes, resource constraints, and time-critical decision-making demands. Artificial intelligence (AI), encompassing machine learning (ML), deep learning (DL), and natural language processing (NLP), has emerged as a potential tool to augment clinical decision support in this environment. However, evidence regarding its clinical effectiveness, safety, and implementation fidelity remains heterogeneous and incompletely synthesized. This systematic review aimed to critically evaluate AI-based clinical decision support systems in emergency medicine, with particular focus on discriminative performance, evidence of clinical impact, safety outcomes, and methodological quality.

Methods:

This review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 guidelines. Databases searched included PubMed/MEDLINE, Elsevier/ScienceDirect, Scopus, and the Cochrane Library through September 2025. Eligible studies included retrospective cohort studies, prospective studies, and AI model validation studies involving adult ED patients and evaluating ML, DL, or NLP algorithms for diagnosis, triage, risk stratification, or outcome prediction. Risk of bias was assessed using the Risk of Bias in Non-randomised Studies of Interventions tool, appropriate for the exclusively retrospective, nonrandomized evidence base. The review was not prospectively registered in PROSPERO, which is acknowledged as a methodological limitation.

Results:

Nine retrospective observational studies from six countries (USA, Singapore, Greece, Taiwan, Israel, and Australia) were included, encompassing over 1.5 million ED encounters. Studies were categorized as model development ( n = 6), external or real-world validation ( n = 2), or implementation/impact studies ( n = 1). AI models consistently demonstrated superior discriminative performance over conventional scoring systems (Emergency Severity Index, Modified Early Warning Score, Sequential Organ Failure Assessment, and Systemic Inflammatory Response Syndrome), with area under the receiver operating characteristic curve (AUROC) values ranging from 0.80 to 0.97 for sepsis prediction, critical care triage, and risk stratification.

Conclusion:

NLP-enhanced models combining structured and unstructured data achieved the highest discrimination. Notably, however, improved AUROC did not uniformly translate into demonstrated clinical benefit. The commercially deployed Epic Sepsis Model exhibited substantially reduced performance in prerecognition settings (AUROC: 0.47), illustrating the risks of overfitting and poor external validity. Only one implementation study (Kotovich et al .) reported measurable reductions in 30-day and 120-day mortality following AI deployment for intracerebral hemorrhage detection, though this was subject to important methodological limitations. Formal safety endpoints were not prospectively measured in any included study. Risk of bias was assessed as moderate to high across the included literature.

Conclusions:

Current evidence supports the discriminative potential of AI-based clinical decision support in emergency medicine, particularly for early sepsis identification and triage prioritization. However, the evidence base is constrained by its exclusively retrospective design, absence of prospective safety measurement, limited external validation, and inadequate attention to calibration, algorithmic fairness, and workflow integration. AUROC improvements should not be conflated with clinical effectiveness. Prospective, multicenter impact trials measuring patient-centered outcomes, alongside rigorous evaluation of implementation fidelity, algorithmic fairness, and governance frameworks, are essential before widespread clinical adoption can be recommended.