DOI: 10.70058/cjm.1962409 ISSN: 3023-7092

Comparative Accuracy of Different ChatGPT Model Configurations in Asthma-Related Questions

Fatma Dindar Çelik, Enes Çelik
Objective: Large language models are increasingly used to obtain medical information; however, their accuracy in asthma-related board-review questions, particularly those requiring guideline-based clinical reasoning, remains uncertain. This study aimed to compare the performance of different ChatGPT model configurations in asthma-related multiple-choice questions.Methods: Forty asthma-related questions obtained from an educational board-review resource were submitted separately to seven ChatGPT model configurations in May 2026: GPT-5.2 Instant, GPT-5.2 Standard Thinking, GPT-5.2 Extended Thinking, OpenAI o3, GPT-5.5 Instant, GPT-5.5 Standard Thinking, and GPT-5.5 Extended Thinking. Each question was entered individually in a new temporary chat session. Responses were compared with reference answers reviewed by adult and pediatric allergist-immunologists according to current guideline recommendations and relevant literature. Accuracy was compared using Cochran’s Q test, McNemar’s test, and the Wilcoxon signed-rank test.Results: Model accuracy ranged from 85.0% to 92.5%. GPT-5.5 Standard Thinking achieved the highest accuracy, with 37 correct responses out of 40, whereas GPT-5.2 Standard Thinking had the lowest accuracy, with 34 correct responses. No statistically significant difference was observed among the seven configurations (p = 0.638). Family-level accuracy was numerically higher for GPT-5.5 than for GPT-5.2, but this difference was not statistically significant (90.8% vs. 86.7%; p = 0.276). Pairwise analyses also showed no statistically significant differences between model comparisons.Conclusion: ChatGPT models showed generally high accuracy in asthma-related questions, although errors occurred in selected guideline- and context-dependent scenarios. ChatGPT may serve as a supportive tool for medical education and clinical information seeking, but its outputs require expert verification.