DOI: 10.1177/20552076261493048 ISSN: 2055-2076

AI hallucinations in health professions education: A mixed-methods evaluation of citation, factual, and data accuracy

Fuad Farajalla, Momin Rasras, Mousa Farajallah, Ahmad Ayed, Ashraf Jehad Abuejheisheh, Rabia H. Haddad

Background

Generative AI is transforming health professions education, but AI hallucinations pose risks to academic integrity. Existing evidence is largely descriptive, leaving error mechanisms poorly understood.

Objectives

This study evaluated AI hallucination prevalence, severity, and mechanisms across health professions education.

Methods

Using a sequential explanatory mixed-methods design, twenty prompts (citation, factual, data-claim) were submitted to ChatGPT (GPT-5.2, free-tier), yielding 80 responses (four generations per prompt). Data were analyzed collectively. Two PhD-level reviewers with medical and nursing backgrounds assessed outputs using a structured hallucination checklist. Quantitative analyses (chi-square, Kruskal–Wallis) were integrated with qualitative thematic analysis of reviewer justifications.

Results

Hallucinations occurred in 61.3% of responses; 48.9% of these were severe. Prevalence and severity were highest in citation-focused prompts (96.9%), followed by factual (62.5%) and data-claim prompts (12.5%) (χ 2 (2) = 41.16, p < .001; η 2 = 0.63). Qualitative analysis revealed four mechanisms: fabricated/hybrid citations, bibliographic inconsistencies, unsupported numerical extrapolations, and conceptually inaccurate framework descriptions. Hallucinations frequently blended authentic and fabricated elements. Conclusions : AI hallucinations in health professions education are prevalent, domain-sensitive, and often severe. AI cannot function as an autonomous scholarly authority; structured verification and AI literacy training are essential.