DOI: 10.3390/app16189376 ISSN: 2076-3417

BCSeg-IRUNet: A Hybrid Inception–Residual Encoder–Decoder for Mammographic Mass Segmentation

Fabian Cienfuegos-Caraveo, Abimael Guzman-Pando, Graciela Ramirez-Alonso, Claudia Adriana Holguin-Gomez, Luis Carlos Hinojos-Gallardo

Breast cancer accounted for approximately 2.4 million new cases and 694,000 deaths worldwide in 2024. Deep learning approaches, particularly encoder–decoder architectures, have been investigated for mammographic mass segmentation. However, evaluations are commonly centered on benchmarked datasets, which include only annotated abnormal cases and therefore do not support assessment of false-positive pixel predictions on mammograms without annotated masses. Segmentation performance on the Digital Mammography Imaging Dataset (DMID) remains comparatively underexplored. DMID is particularly relevant for this task because it combines native full-field digital mammography in DICOM format with pixel-level mass annotations, complementary clinical information, and mammograms without annotated masses, enabling evaluation of both mass delineation and false-positive behavior. In this study, BCSeg-IRUNet is presented for binary mammographic mass segmentation on DMID. The model combines Inception-style multi-scale feature extraction, residual pathways, and UNet-style skip connections. Candidate architectures, including attention-based variants, were compared, followed by sequential evaluation of loss functions and batch sizes under stratified five-fold cross-validation. BCSeg-IRUNet achieved mean Dice coefficients of 0.7015 on the Mass Dataset and 0.6762 on the Full Dataset. In controlled fold-paired comparisons, BCSeg-IRUNet achieved higher Dice coefficients than the retrained AttentionUNet and MultiResUNet baselines in all five held-out test folds, with mean paired improvements of 0.0572 and 0.0499, respectively. Stratified analysis showed that smaller masses were more challenging to segment. For cross-study context, the Full Dataset performance exceeded the previously published Full Dataset Dice coefficient of 0.558, although differences in data partitioning and experimental protocols preclude a direct comparison. These findings support further validation on larger, diverse, and independent mammography datasets.