DOI: 10.3390/biomedinformatics6050077 ISSN: 2673-7426

Capsule-Expert Routing UNet: A Hybrid 2.5D Convolution–Attention Architecture with Mixture of Experts for 3D Medical Segmentation

Nand Kumar Yadav, Rodrigue Rizk, William C.W. Chen, KC Santosh

Background: U-Net-style encoder–decoder architectures are widely used for 3D medical image segmentation, but their conventional skip connections usually transfer encoder features to the decoder through static concatenation. Such direct skip fusion may inadequately address the semantic gap between low-level encoder features and high-level decoder representations, especially in heterogeneous volumetric medical images. This study aims to improve skip-connection fusion by making encoder–decoder feature transfer adaptive, spatially expressive, and scale-aligned. Methods: We propose Capsule-Expert Routing UNet (CER-UNet), a skip-enhanced encoder–decoder architecture for 3D medical image segmentation. CER-UNet replaces static skip fusion with a Capsule-based Mixture-of-Experts (CapMoE) routing mechanism that dynamically selects and combines capsule-inspired spatial experts for skip-feature refinement. The routed features are further reorganized using ALIGNER-based multi-scale alignment before being injected into the decoder. The model also incorporates 2.5D Inception-style factorized convolutions and parameter-free Statistical Attention CBAM (SCBAM) to support efficient volumetric representation learning. Experiments were conducted on three public benchmarks: Synapse multi-organ CT, BTCV abdominal CT, and ACDC cardiac MRI. Results: CER-UNet achieved average Dice scores of 86.70% on Synapse, 84.94% on BTCV, and 92.52% on ACDC, using approximately 32M parameters in the proposed 2.5D configuration. These results show competitive segmentation performance compared with CNN-based, Transformer-based, and recent hybrid methods while maintaining a compact parameter budget. Conclusions: The findings suggest that adaptive skip-connection enhancement through capsule-expert routing and scale-aligned feature fusion can improve the effectiveness of encoder–decoder feature transfer for volumetric medical image segmentation.