DSAI-12 ACCURACY, SAFETY, AND READABILITY OF PUBLIC-FACING LARGE LANGUAGE MODELS IN CNS METASTASIS
Michael Fiorino, Mei Hainline, Tanay Poddar, Talib Hakim, Austin Isaak, Maria Torrez, Ishaan Patel, Amanda Ennis, Daniel AaronsonAbstract
Public-facing large language models (LLMs) are increasingly used by patients to obtain medical information, yet their accuracy, safety, and readability in the context of central nervous system (CNS) metastases remain poorly characterized. Fifteen simulated patient questions regarding CNS brain metastases were submitted to four LLMs (ChatGPT, Claude, Gemini, and Open Evidence). Responses were evaluated using a 5-point Likert scale for accuracy based on National Comprehensive Cancer Network (NCCN) guidelines. Incorrect responses were defined as scores ≤2. Readability was assessed using Flesch Reading Ease (FRE), Flesch–Kincaid Grade Level (FKGL), and Gunning Fog index. Word count was also analyzed. Mixed-effects models were used to compare performance across models, accounting for repeated measures by question. ChatGPT demonstrated the highest mean accuracy score (4.8), followed by Open Evidence (4.7), Gemini (4.1), and Claude (3.7). Incorrect responses occurred most frequently with Claude (26.7%), followed by Gemini (13.3%), while no incorrect responses were observed for ChatGPT or Open Evidence. In ordinal regression analysis, ChatGPT and Open Evidence demonstrated superior performance compared to Claude and Gemini (p < 0.01), with no significant difference between ChatGPT and Open Evidence. Readability analysis revealed that Gemini produced the most readable responses (FRE 43.8; FKGL 12.3; Gunning Fog 15.0), followed by ChatGPT, while Claude and Open Evidence generated significantly less readable outputs (p < 0.01). Gemini also generated the longest responses (mean 467 words), whereas ChatGPT and Open Evidence produced shorter responses (∼370 words). LLM performance varied substantially across accuracy, safety, and readability. ChatGPT and Open Evidence achieved the highest accuracy with no incorrect responses, whereas other models, despite greater readability, were more likely to generate incorrect information. Although contemporary LLMs increasingly reflect medical consensus for CNS metastases, inconsistent reliability remains a concern, underscoring the need for caution in patient use.