Benchmark Readiness of Open-access Musculoskeletal Radiograph Datasets for Foundation Model Segmentation: A Systematic Audit
Sri Nikhil Zallipalli, Kaushik Rao Juvvadi, Gowthamini Asokan, Saichand Linga, Nikitha Devi Thotapalli, Preethii AsokanAbstract
Background and Purpose:
Foundation models including the Segment Anything Model 2 have generated considerable interest as annotation-efficient tools for musculoskeletal (MSK) radiographic image segmentation. Rigorous evaluation requires open-access datasets with well-characterized, high-quality segmentation annotations. No systematic analysis of whether currently available open-access MSK radiograph datasets meet this requirement has been published. This study evaluates the annotation completeness, image quality, and benchmark readiness of three widely cited open-access MSK imaging resources used for radiographic benchmarking: FracAtlas and Musculoskeletal Radiographs (MURA), two native projectional radiograph datasets, and computed tomography (CT)-derived digitally reconstructed radiographs (DRRs) generated from TotalSegmentator, an open-access CT segmentation dataset.
Materials and Methods:
A structured, reproducible audit framework was applied to FracAtlas (4083 images), a 500-study stratified random sample from MURA (seven body regions), and 120 DRRs generated from TotalSegmentator CT volumes. Annotation completeness was assessed across four tiers. Mask quality in FracAtlas was graded independently by two raters: two orthopedic surgical trainees and a radiologist, across four levels, with inter-rater agreement quantified by Cohen’s weighted kappa. Image quality was characterized using five no-reference metrics. Domain gap was quantified using structural similarity (SSIM) and Fréchet Inception Distance (FID). A five-criterion, equally weighted Benchmark Readiness Score (BRS) framework was developed, applied to all three datasets, and evaluated through sensitivity analysis.
Results:
MURA contained no pixel-level segmentation annotations, rendering it unsuitable for direct segmentation benchmarking. FracAtlas provided whole-bone-quality (Grade A) masks for only 175 of 4,083 images (4.3%); 61.4% of masks were fracture-region-only. Inter-rater agreement for mask grading was substantial (Cohen’s κw = 0.76; 95% confidence interval: 0.71–0.81). TotalSegmentator-digitally reconstructed radiograph (DRRs) provided complete synthetic bone masks but demonstrated a substantial domain gap from real radiographs (mean SSIM 0.41 ± 0.09; FID 87.3). BRS scores were: MURA 4/25, FracAtlas 8/25, FracAtlas-Complete subset 12/25, and TotalSegmentator-DRR 16/25. Dataset ranking was stable across sensitivity analysis thresholds of 18/25, 20/25, and 22/25; no dataset reached publication-readiness under any threshold.
Conclusion:
No currently available open-access MSK radiograph dataset fully meets the requirements for rigorous foundation model segmentation benchmarking. The FracAtlas-Complete subset (