Determinants of Trust and Reliance on Artificial Intelligence in Clinical Settings: A Study on Physicians’ Use of Large Language Model in Psychiatric Symptom Assessment
Ji-Min Kim, Seong Hoon JeongObjective This study investigated physicians’ use of large language model (LLM)-generated assessments in psychiatric symptom evaluation and factors influencing their acceptance of artificial intelligence (AI) suggestions.Methods Twelve psychiatrists evaluated 50 anonymized psychiatric case vignettes, rating the presence and severity of 73 symptoms using a 4-point Likert scale. After initial assessment, they reviewed GPT-4o’s symptom ratings and could revise their evaluations. Accuracy was measured against a gold standard established by expert consensus. We computed initial (I), LLM (L), and revised (R) accuracy scores using percent agreement and weighted kappa. Performance improvement was measured by relative kappa increase. Regression analyses examined the influence of task difficulty and user competence on accuracy improvement.Results LLM assessments outperformed initial physician ratings in both agreement (90.4% vs. 82.9%, p<0.001) and kappa (0.713 vs. 0.559, p<0.001). After revision, physician performance improved (κ=0.651) but remained below LLM levels. The switch rate—cases where physicians revised in response to LLM disagreement—was modest (25.4%), indicating partial reliance on AI. Performance gains were positively associated with LLM accuracy and negatively associated with users’ own baseline competence, suggesting that less confident users rely more on LLM assistance.Conclusion Physicians demonstrated limited but strategic trust in LLM outputs, adjusting their judgments more when the LLM was accurate or when their own competence was lower. Miscalibrated trust—excessive skepticism or overreliance—led to missed gains. Effective human-AI collaboration in psychiatry requires tools and training for accurate self-assessment and AI trust calibration to avoid the pitfalls of algorithm aversion or overreliance.