DOI: 10.1145/3849709 ISSN: 3068-8590
Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots
Lorenzo de la Loza, Vijay K. Madisetti
Large language models (LLMs) frequently endorse and elaborate on users’ delusional beliefs, a failure mode termed
psychogenicity
in the Psychosis-Bench study of Au Yeung et al., whose framing we adopt. We present the first systematic evaluation of whether anti-sycophancy interventions transfer to psychosis-relevant contexts. Across 1,280 experiments spanning 10 conditions, 8 frontier LLMs, and 16 clinically derived Psychosis-Bench scenarios, a combined anti-sycophancy prompt reduces mean Delusion Confirmation Scores by 73.8% (paired
\(t(127)=11.74\)
,
\(p<10^{-21}\)
, Cohen's
\(d=1.04\)
). Adding a domain-general self-reflection prompt yields a 77.0% reduction and raises Safety Intervention rates by 66.9%, delivered entirely as a system prompt. Classifier-based guardrails (Llama Guard 3) flag only 5 of 3,072 evaluated turns; a reasoning guardrail (o4-mini) flags
\(14\times\)
more. Ablations isolating either mechanism alone plateau at
\(\approx 46\%\)
reduction, establishing anti-sycophancy prompting as a
necessary foundation
that add-on mechanisms augment but cannot replace.