DOI: 10.16899/jcm.2008162 ISSN: 2667-7180
Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions
Muhammed Nezih Koç, Sait Ramazan Gülbay Aim: In large language model (LLM) examination studies, models answer the same items repeatedly, so observations are clustered and conventional intervals overstate precision. We aimed to measure the accuracy of seven LLMs on a subspecialty-level anaesthesiology question bank and to determine which between-model and subgroup differences remain distinguishable under a cluster-respecting analysis.Materials and Methods: A 120-item, five-option multiple-choice bank in the style of the Anaesthesiology and Reanimation Subspecialty Examination, written by the authors and reviewed against current guidelines and textbooks, was administered to seven LLMs in five runs at temperature 0, with option order re-randomised per run (4,200 responses from 120 unique items). Accuracy is reported with cluster bootstrap 95% confidence intervals (CIs) resampling items (4,000 replicates) and naive Wilson intervals for comparison; subgroup analyses are exploratory.Results: Pooled accuracy was 80.1% (cluster-robust 95% CI 75.8-83.9; naive Wilson 78.8-81.3). Between-model differences were large and robust, spanning 60.8% (54.0-67.5) to 94.0% (90.3-97.0) with non-overlapping extremes. Cluster-robust intervals were wider in 25 of 26 estimates (median factor 2.1, range 0.8-4.0). Most subgroup comparisons were not distinguishable, intervals overlapping substantially: vignette (75.4%, 62.5-86.5) versus non-vignette (80.9%, 76.6-84.9) and guideline-dependent (73.0%, 64.6-80.7) versus other items (80.9%, 76.6-84.9). Only contrasts between weakest and strongest domains persisted. Between-run standard deviation was 1.26-2.80 points.Conclusion: Between-model differences and run-to-run instability are robust; commonly emphasised domain and item-characteristic differences are mostly not distinguishable once clustering is respected. Benchmarks of 120 items can separate models whose accuracies differ widely but cannot reliably localise their weaknesses.
More from our Archive
-
DOI: 10.68381/jca02008 2026
Proximal Smoothness and the Lower-C
2
Property F. H. Clarke, R. J. Stern, P. R. Wolenski
-
DOI: 10.68381/jca13044 2026
Characterizations of Prox-Regular Sets in Uniformly Convex Banach Spaces Frédéric Bernard, Lionel Thibault, Nadia Zlateva
-
DOI: 10.68381/jca15047 2026
Brøndsted-Rockafellar Property and Maximality of Monotone Operators Representable by Convex Functions in Non-Reflexive Banach Spaces Maicon Marques Alves, Benar Fux Svaiter
-
DOI: 10.68381/jca16027 2026
Proximal Smoothness and the Exterior Sphere Condition Chadi Nour, Ron J. Stern, Jean Takche
-
DOI: 10.68381/jca16053 2026
A New Old Class of Maximal Monotone Operators Maicon Marques Alves, Benar Fux Svaiter
-
DOI: 10.68381/jca13045 2026
Maximal Monotonicity via Convex Analysis Jonathan Borwein
-
DOI: 10.68381/jca08009 2026
Variational Inequalities and Regularity Properties of Closed Sets in Hilbert Spaces Giovanni Colombo, Vladimir V. Goncharov
-
DOI: 10.68381/jca17060 2026
Existence and Uniqueness of Solutions for Non-Autonomous Complementarity Dynamical Systems Bernard Brogliato, Lionel Thibault
-
DOI: 10.68381/jca01001 2026
Variational Sum of Monotone Operators H. Attouch, J.-B. Baillon, M. Théra
-
DOI: 10.68381/jca22017 2026
Weak Convexity of Sets and Functions in a Banach Space Grigorii E. Ivanov