DOI: 10.3390/info17080749 ISSN: 2078-2489

A Clinician-in-the-Loop Framework for Validating and Selecting Synthetic Paediatric Dermatology Images

Ali Tariq Nagi, Chiara Bellatreccia, Andrea Borghesi, Arianna Dondi, Luca Pierantoni, Daniele Zama, Iria Neri, Marcello Lanari, Roberta Calegari

Synthetic data are increasingly proposed as a strategy for addressing data scarcity and representation imbalance in medical AI, particularly for paediatric populations and darker skin tones. However, visually plausible synthetic images may still contain clinically implausible features or fairness-relevant inconsistencies that are not adequately captured by automatic image-quality metrics. In this study, we present and empirically evaluate a clinician-guided framework for validating and selecting synthetic paediatric dermatology images. The framework combines a clinician-facing evaluation platform with structured assessments of visual realism, mask quality, diagnostic plausibility, confidence, and skin-tone relevance. Four clinicians with complementary expertise in paediatrics and dermatology completed 282 assessments of 93 real and synthetic images. Synthetic images were often rated as visually realistic but showed lower inter-rater agreement and weaker mask-quality assessments than real images. Clinician realism and confidence ratings were then used to divide 30 synthetic images into 18 approved and 12 non-approved images. To assess downstream utility, we compared a real-only ResNet50 classifier with classifiers augmented using all synthetic images, clinician-approved synthetic images, or non-approved synthetic images. Across three patient-level experimental splits, the clinician-approved condition achieved the strongest overall classification performance and the largest gains for the under-represented Dark-Skin subgroup. Because the Dark-Skin subgroup contained only seven patients and the synthetic subsets differed in size and disease composition, these fairness results should be interpreted as exploratory. The present study therefore provides evidence for clinician-guided validation and data curation rather than for a completed iterative generator-retraining process. Future work will evaluate whether clinician feedback can also support repeated generative-model refinement in larger, multi-centre datasets.

More from our Archive