Benchmarking Deep Learning for NSCLC PET/CT Segmentation on a Histologically Confirmed Vietnamese Dataset: Validation and Generalization
Quang Tuan Ho, Ngoc Ha Bui, Thuy Duong Tran, Quang Huy Khuat, Ngoc Toan Tran, Xuan Chung Le, Huu Quyet Nguyen, Tat Thang Nguyen, Van Thai Nguyen, Dinh Thuy Mai, Quang Duy To, Dinh Chau Nguyen, Nguyen Huong Giang Trinh, Van Chinh Cao, Tien Hung Bui, Thu Trang Vu, Khac Nam Vo, Hai Quan HoAccurate segmentation of non-small cell lung cancer (NSCLC) on positron emission tomography/computed tomography (PET/CT) is an essential prerequisite for automated metabolic tumor volume (MTV) quantification and staging. Although deep learning models achieve high performance on large-scale datasets, their generalization across different clinical domains is limited by variations in imaging protocols and patient demographics. This study aims to evaluate several deep learning architectures and investigate a transfer learning strategy to mitigate domain shift. Three architectures, including ResNet-backbone 3D U-Net, nnU-Net v2, and Swin UNETR, were benchmarked from scratch and compared with a fine-tuned nnU-Net initialized with AutoPET II weights. Results on the internal dataset showed that the fine-tuned nnU-Net achieved a Dice similarity coefficient (DSC) of 83.4 ± 6.5%, a 95% Hausdorff distance (HD95) of 5.1 ± 3.6 mm, and a precision of 89.6 ± 8.2%. Compared to the nnU-Net v2, the fine-tuned nnU-Net improved the absolute DSC by 6.8% while reducing local training time by 37.5% by bypassing the initial feature-learning phase. The fine-tuned nnU-Net model also demonstrated a high correlation between the MTV and the ground truth (Pearson r = 0.96, p < 0.001), indicating its potential as a reliable automated approach for quantitative MTV extraction and NSCLC prognostic-related analysis.