AI hallucinations in health professions education: A mixed-methods evaluation of citation, factual, and data accuracy
Fuad Farajalla, Momin Rasras, Mousa Farajallah, Ahmad Ayed, Ashraf Jehad Abuejheisheh, Rabia H. HaddadBackground
Generative AI is transforming health professions education, but AI hallucinations pose risks to academic integrity. Existing evidence is largely descriptive, leaving error mechanisms poorly understood.
Objectives
This study evaluated AI hallucination prevalence, severity, and mechanisms across health professions education.
Methods
Using a sequential explanatory mixed-methods design, twenty prompts (citation, factual, data-claim) were submitted to ChatGPT (GPT-5.2, free-tier), yielding 80 responses (four generations per prompt). Data were analyzed collectively. Two PhD-level reviewers with medical and nursing backgrounds assessed outputs using a structured hallucination checklist. Quantitative analyses (chi-square, Kruskal–Wallis) were integrated with qualitative thematic analysis of reviewer justifications.
Results
Hallucinations occurred in 61.3% of responses; 48.9% of these were severe. Prevalence and severity were highest in citation-focused prompts (96.9%), followed by factual (62.5%) and data-claim prompts (12.5%) (χ 2 (2) = 41.16, p < .001; η 2 = 0.63). Qualitative analysis revealed four mechanisms: fabricated/hybrid citations, bibliographic inconsistencies, unsupported numerical extrapolations, and conceptually inaccurate framework descriptions. Hallucinations frequently blended authentic and fabricated elements. Conclusions : AI hallucinations in health professions education are prevalent, domain-sensitive, and often severe. AI cannot function as an autonomous scholarly authority; structured verification and AI literacy training are essential.