Effect of Fitzpatrick Skin Type Prompting on Diagnostic Accuracy in Multimodal Large Language Models: A Within-Image Experimental Study
Manoj Bhagwat, Tyler Wittles, Jeffery Tan, Joshua Mijares, Neil K. Jairath, Syril Keena T. QueMultimodal large language models are increasingly used for dermatologic queries, but whether Fitzpatrick skin type (FST) labels affect diagnostic accuracy is unknown. We evaluated 656 biopsy-confirmed photographs from the Diverse Dermatology Images (DDI) dataset under no-FST, DDI-concordant FST, and two DDI-discordant FST conditions using ChatGPT 5.2 Edu and Gemini 3.1 Pro browser configurations (5248 evaluations). Confirmed malignant diagnosis omission from the top three differential diagnoses (“lethal miss”) was a prespecified exploratory outcome. Gemini had higher ordinal accuracy than ChatGPT (odds ratio 2.32; 95% confidence interval 2.01–2.67; false discovery rate-adjusted p < 0.001). No FST prompt condition significantly improved accuracy over no-FST prompting; an equal-image-weighted sensitivity analysis yielded similar results. Among malignant images eligible for extreme discordance, extreme DDI-discordant prompting was associated with more ChatGPT lethal misses than no-FST prompting (paired odds ratio 5.0; p = 0.043); Gemini showed no significant paired shift (p = 1.00). Post hoc binary malignancy detection showed low no-FST specificity for ChatGPT (42.5%) and Gemini (22.9%), indicating frequent overcalling. Gemini showed lower no-FST accuracy for malignant FST V–VI images than for FST I–II and III–IV images. FST prompting provided no measurable diagnostic benefit. Neither configuration supports autonomous dermatologic use.