AI-assisted rheumatology triage changes with referral framing
Mahmud Omar, Mohammad E Naffaa, Reem Agbareia, Fadi Hassan, Abdulla Watad, Helana Jeries, Alon Gorenshtein, Yiftach Barash, Olga R Brook, Girish N Nadkarni, Eyal KlangAbstract
Objectives
Two in three US physicians now use healthcare AI, and large language models (LLMs) are entering the triage workflows that determine which patients reach rheumatology and how quickly. We aimed to test whether nine prespecified cues in referral notes, patient descriptions and demographics shift AI-assisted triage decisions when the underlying clinical information is unchanged.
Methods
We conducted a controlled, physician-validated experiment across 30 physician-authored rheumatology vignettes, each independently rephrased three times (90 case variants). We tested five LLMs from three providers - Anthropic, Google, and OpenAI - under nine dimensions spanning demographics, clinical context, and communication framing, with 57 controlled contextual modifications, four system-prompt personas, and five repetitions per cell, yielding more than 200,000 model queries. Sixteen clinical fields were graded against physician-validated ground truth, with excellent inter-rater agreement (Fleiss’ kappa=0.92).
Results
Baseline composite concordance with expert ground truth was high at 0.869. We had expected sociodemographic cues to be the strongest source of distortion. Instead, the largest shifts came from how the case was framed and described. When patients were described as anxious, models attributed symptoms to psychological rather than organic causes nearly three times as often as at baseline (13.1% vs 4.5%; odds ratio 3.2 versus stoic framing), a shift that risks relabeling organic disease as functional. Clinician anchoring in the referral note reduced concordance, consistently across models and rephrasings and significantly for acuity (dismissive anchor, vignette-level p = 0.01), mainly by downgrading urgency. In contrast, race or ethnicity, socioeconomic status, and language barrier produced no detectable effect, including in mixed-effects models that accounted for repeated vignette use.
Conclusion
Although baseline concordance with specialist ground truth was high, it was readily disrupted by how referral notes were worded and how patients described their symptoms, not by patient demographics. Before AI-assisted triage enters rheumatology referral pathways, systems should separate objective clinical evidence from interpretive framing, and urgency and psychological attribution should be treated as auditable safety signals.