Evaluating Next-Generation Large Language Models for Psychiatric Diagnosis: A Comparative Vignette Study
M. MansoorIntroduction
Diagnostic accuracy in psychiatry is a challenge due to overlapping symptom profiles and the subjective nature of clinical assessment. Large Language Models (LLMs) have emerged as promising decision-support tools, capable of analyzing complex textual data. However, the performance of current models is inconsistent; while proficient with common disorders, they exhibit significant failures in complex cases. This study investigates whether a simulated next-generation LLM can address this critical performance gap.
Objectives
The primary objective was to compare the diagnostic accuracy of three current-generation LLMs (GPT-5, GPT-4, Claude 3 Sonnet). A secondary objective was to evaluate and contrast model performance across a spectrum of psychiatric disorders, specifically comparing accuracy on common conditions versus those known for diagnostic ambiguity, such as early schizophrenia and bipolar disorder with psychotic features.
Methods
A comparative study was conducted using a validated clinical vignette methodology. Ten clinical vignettes, validated by an expert panel to meet DSM-5/ICD-11 criteria, were created for five disorders: Major Depressive Disorder (MDD), Post-Traumatic Stress Disorder (PTSD), Social Anxiety Disorder, Early Schizophrenia, and Bipolar I Disorder (manic episode with psychosis). Three LLMs were tested: GPT-5, GPT-4, and Claude 3 Sonnet, which was conceptualized with enhanced causal reasoning capabilities. Each vignette was presented to each model 20 times (N=600) with a standardized prompt to provide a primary diagnosis. Accuracy was calculated against the expert panel’s gold standard, with statistical significance assessed using χ 2 tests.
Results
As shown in Table 1, all models demonstrated high accuracy (>90%) for MDD, PTSD, and Social Anxiety Disorder, performing comparably to the expert benchmark. However, the accuracy of GPT-4 and Claude 3 Sonnet dropped significantly for the Early Schizophrenia (65% and 60%, respectively) and Bipolar I vignettes (70% and 68%, respectively) (p<.001). In contrast, GPT-5 maintained high accuracy in these complex cases (88% and 91%), performing significantly better than the current-generation models (p<.001) and approaching the expert benchmark.
Diagnostic Accuracy (%) by Model and Disorder Table 1. Long description.