DSAI-09 INTERPRETABLE RADIOMICS RISK SIGNATURE FOR DISTINGUISHING TUMOR RECURRENCE FROM RADIATION NECROSIS IN BRAIN METASTASES ON POST-TREATMENT MRI: A COMPARATIVE ANALYSIS WITH ‘BLACK-BOX’ APPROACHES
Dheerendranath Battalapalli, Hyemin Um, Marwa Ismail, Virginia B Hill, Sushant Puri, Jennifer S Yu, Lan Lu, Ameya P Nayate, Anthony Higinbotham, Lisa R Rogers, Prateek Prasanna, Mainak Bardhan, Chengnan Li, Mustafa M Basree, Andrew M Baschnagel, Alan B McMillan, Ankush Bhatia, Manmeet S Ahluwalia, Michael C Veronesi, Pallavi TiwariAbstract
Purpose
Distinguishing tumor recurrence (TuR) from radiation necrosis (RN) after stereotactic radiosurgery remains a persistent clinical challenge on conventional post-treatment MRI, with enhancement patterns and mass effect frequently overlapping despite different underlying biology. This ambiguity often leads to delayed therapy or unnecessary surgical intervention. We therefore evaluated an interpretable, heterogeneity-focused radiomics risk signature (SMART-Risk) to distinguish TuR from RN and compared it with two widely used data-driven “black-box” alternatives: ResNet50 and a fine-tuned pretrained imaging foundation model (MedGemma 1.5), using post-contrast T1-weighted MRI in brain metastases.
Methods
We curated a retrospective multi-institutional cohort (233 studies) with >80% pathology confirmation. Cleveland Clinic (CCF) and University Hospitals Cleveland (UH) constituted the training set, and the University of Wisconsin (UW) served as the external test site. After lesion segmentation, 944 radiomic descriptors per lesion were computed, capturing graph-based spatial organization (GrRAiL), local gradient texture heterogeneity (CoLlAGe), Haralick texture, and morphology. LASSO-selected features were combined to create the SMART-Risk signature and tested within a random-forest classifier. For comparison, ResNet50 was trained from scratch, and MedGemma 1.5 was fine-tuned for binary TuR vs RN classification. We reported cross-validation accuracy (mean±SD) and external test accuracy, F1 score, and AUC. Model explanations were generated with SHAP.
Results
On the UW test set, SMART-Risk achieved 0.79 accuracy, 0.78 F1 score, and 0.86 AUC (CV accuracy 0.78±0.08). SHAP highlighted spatial-organization and heterogeneity features (e.g. average path length, node count, entropy; Mann-Whitney U, p ≤ 0.001), with TuR demonstrating greater spatial complexity than RN. SMART-Risk outperformed ResNet50 (0.60 accuracy, 0.67 F1, 0.63 AUC; CV 0.70±0.09) and MedGemma 1.5 (0.66 accuracy, 0.74 F1, 0.73 AUC; CV 0.70±0.08).
Conclusion
Heterogeneity-driven radiomics (SMART-Risk) may improve TuR versus RN discrimination over black-box ResNet50 and foundation model strategies, potentially improving treatment management and reducing post-treatment diagnostic ambiguity in patients with brain metastases.