Retrospective Assessment of LLM Decision-Making for IOL Power Selection in Cataract Surgery in Comparison with Surgeon Performance
Randle M. Uyeda, Josh Zhe, Mina M. Sitto, Sanjana Molleti, Hanna Pawlowski, Phillip C. Hoopes, Majid MoshirfarObjectives: To evaluate a general-purpose large language model (LLM)’s ability to perform surgeon-level decision-making in choosing an intraocular lens (IOL) power when provided the same preoperative data. Methods: This study included 293 eyes (174 patients) that underwent monofocal or toric cataract surgery performed by a single surgeon. GPT-5.2 analyzed the same biometric printouts, topographic displays, and preoperative data used in surgical planning. Postoperative refraction for GPT-5.2-selected IOLs was simulated from the measured postoperative refraction using a fixed conversion factor. The mean refractive deviation from target, mean absolute error (MAE), and number of eyes within ±0.25 D and ±0.50 D of target refraction were calculated. A subgroup analysis evaluated toric IOL cylinder magnitude and axis. Results: GPT-5.2-selected IOL power differed significantly from that of the surgeon (19.24 ± 4.01 vs. 19.64 ± 4.07 D [mean ± SD], p < 0.001). Using simulated refractive outcomes for GPT-5.2, MAE was higher than for surgeon-selected IOLs (0.45 ± 0.03 vs. 0.35 ± 0.02 D [mean ± SE], p = 0.002), and a greater percentage of eyes fell within ±0.25 D of target refraction for the surgeon (51.9% vs. 44.7%, p < 0.01), with no difference within ±0.50 D (p = 0.06). In toric eyes, the cylinder magnitude did not differ, but simulated residual astigmatism was higher for GPT-5.2 (0.78 ± 0.55 vs. 0.58 ± 0.47 D, p < 0.001), with an axis flip in 47.1% versus 39.2% of eyes. Conclusions: GPT-5.2 synthesized clinical data and approximated surgeon IOL selection, but its simulated refractive outcomes were less accurate than the surgeon’s measured outcomes. These findings suggest that a domain-specific LLM may serve as a support tool in IOL selection and cataract surgery planning.