The Effects of Artificial Intelligence in Quality Assessment of Outpatient Mental Health Services on Human Auditors
Hayden T. McLeod, Joshua L. Cohen, Dena M. Bravata, Roma M. Bajaria, Kelvin Tran, Neha P. Chaudhary, Nicole M. Benson, Jeff GouldAbstract
Background The use of large language model (LLM)-based artificial intelligence (AI) to assess the quality of outpatient mental health documentation represents an understudied application. Critical questions remain about the effects of exposing human auditors to AI-generated recommendations on auditor efficiency, agreement, and patterns of decision-making.
Objectives The objective of this study is to examine how exposure to LLM-generated quality assessments influences auditor performance during reviews of documents of outpatient mental health intake sessions.
Methods We conducted a pre–post observational evaluation of 10 trained human auditors who reviewed outpatient mental health intake notes and assigned quality ratings using a standardized 13-item chart audit rubric before and after implementation of LLM support. The AI system used Anthropic Claude Sonnet 4 with a fixed chain-prompting strategy. We assessed item-level pass frequency for each auditor across the pre-AI and post-AI periods, AI–human agreement (the alignment between auditor pass/fail determinations and AI-generated recommendations), and auditing time. The pre-AI period included 9,711 notes (126,243 item-level observations) and the post-AI period included 2,677 notes (34,801 item-level observations).
Results Following implementation of AI-supported auditing, the mean item-level pass rate increased by approximately 2%. The distribution of auditor-level changes was broad (range: −35 to +35%, standard deviation [SD]: 10%) indicating substantial heterogeneity in individual auditor responses to AI support. High auditor agreement with AI recommendations was associated with larger pass-rate changes for selected rubric items, with some auditors increasing and others decreasing pass rates relative to the pre-AI period. The average time for an auditor to review notes decreased from 8.77 min (SD: 9.64) per review during the pre-AI period to 8.06 min (SD: 10.28) during the post-AI period (p = 0.002), an 8% improvement equivalent to 71 min saved per 100 notes reviewed.
Conclusion LLM-based AI tools may improve the efficiency of clinical quality assurance workflows, but exposure to AI recommendations may meaningfully change human auditor behavior.