DOI: 10.3390/life16081309 ISSN: 2075-1729

Attention-Enhanced ResNet–U–Net for Automated Colorectal Tumor Segmentation in CT Scans

Lucian Mihai Florescu, Cosmin Vasile Obleagă, Mădălin Mămuleanu, Ioana Andreea Cîrlig, Alesandra Florescu, Mihai Alexandru Ene, Aurelia Ștefania Domenco, Alexandru Marian Olaru, Alexandra Gabriela Cosmina Țâru, Raluca Elena Nica, Rossy Vlăduț Teică, Ioana Andreea Gheonea

Automated colorectal tumor segmentation on routine computed tomography (CT) remains technically challenging because of limited soft-tissue contrast, heterogeneous tumor morphology, complex bowel anatomy, and marked imbalance between tumor and background pixels. This retrospective technical feasibility study aimed to develop and evaluate an attention-enhanced ResNet50–U-Net architecture for automated segmentation of primary colorectal tumors on contrast-enhanced abdominopelvic CT images. The dataset comprised 493 axial CT image–mask pairs containing visible colorectal tumors, including 70 institutional images obtained from 25 patients and 423 images from publicly available colorectal cancer imaging resources. Reference masks were generated by manual tumor delineation performed by an experienced radiologist and reviewed by a second abdominal radiologist, with uncertain contours resolved by consensus. The proposed model combined an ImageNet-pretrained ResNet50 encoder with an attention-guided U-Net decoder and was trained using a composite Binary Cross-Entropy and Focal Tversky loss function to address foreground–background class imbalance. Performance was assessed using five-fold image-level cross-validation. Predicted probability maps were binarized using a threshold of 0.5, and segmentation metrics were calculated through global micro-averaging within each validation fold. The model achieved a mean Intersection over Union of 0.7406 ± 0.0276, a Dice similarity coefficient of 0.8507 ± 0.0182, a pixel-level sensitivity of 0.9559 ± 0.0266, and a pixel-level background specificity of 0.9956 ± 0.0003. Qualitative assessment demonstrated substantial spatial agreement between predicted masks and expert annotations, although minor boundary discrepancies and small false-positive components were observed. These findings support the technical feasibility of attention-enhanced encoder–decoder architectures for colorectal tumor segmentation on routine CT images. However, because validation was performed at the image level, entirely tumor-negative images were not included (preventing the evaluation of clinical specificity and false-positive rates), and no independent external test cohort was available, further patient-level, multicenter, and DICOM-based validation is required before clinical implementation. While the model demonstrated robust internal performance, future studies must prioritize large-scale external validation and patient-level cross-validation to rigorously assess generalizability and specificity.

More from our Archive