DOI: 10.12688/f1000research.187443.1 ISSN: 2046-1402

MedGemma Evaluation for Fundamental Radiological Imaging Classification Tasks

Theo Sartoretti, Adrien Jayet, Thomas Saliba, Jonas Richiardi, Chiara Pozzessere, David C. Rotzinger, Guillaume Fahrni
Background This study evaluates the performance of MedGemma-4B, a specialized medical vision-language model (VLM), on six fundamental radiological image classification tasks. We hypothesize that, despite their diagnostic capabilities, such models may underperform on clinically essential perceptual tasks that may be underrepresented in training data. Methods MedGemma-4B was assessed using 600 multicenter radiological images, divided equally into six classification tasks: modality, body part, orientation, contrast, image mode, and organ. Using standardized visual question answering (VQA) prompts, model outputs were categorized by two expert radiologists as correct, incorrect, or imprecise. Performance was evaluated using accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and Matthews correlation coefficient (MCC), with 95% confidence intervals and χ 2 tests. Results Performance varied substantially across tasks. The model demonstrated robust accuracy in modality (97%) and body part identification (80%), but critically underperformed in orientation (41%), contrast phase recognition (32%), and image mode classification (48%). Organ identification achieved moderate accuracy (74%). MCC values showed strong correlation for modality (96.2%) and body part (78.7%), but poor correlation for orientation (33.1%). Statistical analysis confirmed significant performance disparities ( p  < 0.001), with modality and body part tasks outperforming all other tasks ( p  < 0.001). Conclusions MedGemma-4B excels at broad categorizations but struggles with more granular yet fundamental radiological image features.

More from our Archive