Large Language Model Chatbots Cannot Reliably Calculate Clinical Risk Scores—A Comparative Accuracy Study
Philippe Di Cicco, Wesley Bennar, Corentin Volet, Serban G. Puricel, Mario Togni, Stéphane Cook, Dorian GarinBackground: Large language models (LLMs) are increasingly accessible to healthcare providers and patients for clinical decision support, yet their ability to perform precise mathematical calculations required for validated risk scores remains unexplored, and errors could compromise patient safety. The EuroSCORE II requires complex multivariable computation that could reveal fundamental limitations in LLM computational capabilities. Methods: We evaluated four publicly available chatbots (ChatGPT (GPT-4o, OpenAI), Claude (Sonnet 4.5, Anthropic), Gemini (2.5 Flash, Google) and Deepseek (V3)) in calculating EuroSCORE II for 105 patients from the CARDIO-FR registry with gold standard heart team calculations. Each model was tested using two approaches: direct calculation from clinical parameters alone and formula-based calculation with the explicit EuroSCORE II algorithm provided. Performance was assessed through mean absolute error (MAE), correlation coefficients, and clinical agreement within ±2% of gold standard values. Results: Direct LLM calculations demonstrated poor accuracy (MAE range: 3.3–6.4%) with the best performer (Gemini) achieving only 50.5% clinical agreement. Formula provision improved performance in three of four models, with ChatGPT formula achieving the lowest MAE (2.9%) and highest clinical agreement (51.4%), followed by Claude formula (MAE 3.2%, agreement 48.6%). Adding the formula to the prompt significantly improved performance and reduced bias. All methods exhibited significant systematic biases (p < 0.05 for 7/8 strategies). Conclusions: Publicly available LLM chatbots cannot reliably calculate EuroSCORE II for clinical use. Adding the formula to the prompt significantly improved performance, but was not sufficient to reach the clinically required threshold of 90% agreement. Clinicians should rely on validated risk calculation tools rather than LLM chatbots for quantitative clinical risk assessments.