DOI: 10.3390/diagnostics16193180 ISSN: 2075-4418

Severity Stratification Changes What Hallucination Rates Mean: An Item-Matched Audit of Five Large Language Models in Orthodontic Decision Support

Ferdi Allaf, Mustafa Ozcan

Background/Objectives: Large language models (LLMs) are consulted for clinical decision support, and their reliability is judged by hallucination prevalence—a measure treating every unsupported element as equivalent, although a fabricated citation and a fabricated protocol differ in what a clinician acting on them would do. We tested whether prevalence tracks clinical risk. Methods: Five LLMs answered a 100-item orthodontic benchmark validated by three external orthodontists (content validity index 0.923). The reference standard was fixed by construction for the 45 items naming a non-existent entity, so any substantive elaboration is unsupported by design; citations were adjudicated against PubMed and CrossRef. Responses were coded with a seven-category taxonomy and assigned to clinical, operational, or epistemic severity tiers in a post hoc exploratory stratification, pre-specified rather than prospectively registered. Because all models answered the same items, comparisons used Cochran’s Q with pairwise McNemar tests and generalised estimating equations clustered on item. Results: Of 500 responses, 449 (89.8%) contained a hallucination but only 65 (13.0%; 95% CI 10.3–16.2) were clinically consequential: prevalence was roughly sevenfold greater than the rate of clinically consequential output as the authors defined it. Between-model differences were large for undifferentiated prevalence (73–100%; Q = 50.51, p < 0.001) and contracted at the clinical tier (10–16%; Q = 10.00, p = 0.040), where no pairwise contrast survived adjustment. Item-level clustering was far stronger for clinically consequential output than for undifferentiated prevalence (intra-class correlation 0.83 versus 0.07). Conclusions: Prevalence and clinically consequential error are not interchangeable and rank models differently. The stratification is exploratory and author-defined, and the benchmark stress-tests susceptibility to fabricated premises rather than surveying natural use.