DOI: 10.3390/jcm15166209 ISSN: 2077-0383

Detection of Myopia from Colour Fundus Photographs Using YOLO-Based Computer-Vision Models: A Comparison with Expert Ophthalmologists

Nicola Rizzieri, Luca Dall’Asta, Maris Ozolinš

Objectives: The purpose of this study was to evaluate the ability of computer vision models to detect myopia from standard colour fundus photographs and to compare their diagnostic performance with that of experienced ophthalmologists. Methods: A previously published dataset of 324 retinal fundus images labelled as myopic or non-myopic based on cycloplegic refraction as used for model training and internal validation. Images were acquired using a non-mydriatic 45° fundus camera. Final model evaluation was performed on an independent test set of 50 images from different patients who were not included in the original dataset. YOLOv8 and YOLOv11 variants were trained for binary classification. Internal validation used patient-level cluster bootstrap confidence intervals, whereas image-level bootstrap confidence intervals were estimated for the independent test set. Pairwise model comparisons were adjusted using the Holm–Bonferroni correction. Five experienced ophthalmologists independently classified the test set, and their consensus was compared with the selected YOLO models using DeLong’s and exact McNemar tests. Results: Internal validation identified YOLOv8-m and YOLOv11-n as the best-performing models according to a predefined composite score used exclusively for model selection. On the independent test set, YOLOv11-n achieved the highest area under the curve (AUC = 0.889), followed by YOLOv8-m (0.806), although the difference was not statistically significant (DeLong test, p > 0.05). The clinical consensus achieved an AUC of 0.832, with no significant difference compared with either model. Exact McNemar testing likewise revealed no statistically significant differences in paired classification outcomes between either AI model and the clinical consensus. Limitations include the small, single-centre, class- and age-imbalanced dataset and the limited number of expert observers. Conclusions: Although neither YOLOv8-m nor YOLOv11-n showed statistically significant differences from the clinical consensus on this independent test set, these findings should be interpreted cautiously given the relatively small, single-centre study population. Larger multicentre studies with independent external validation are warranted to confirm the generalisability, robustness, and potential role of clinician-driven computer vision models as decision support tools for myopia screening.

More from our Archive