A Prompt-Preserving SAM ViT-B Adaptation Framework with Historical Branch Fusion and Soft Convolutional Expert Weighting for Medical Image Segmentation
Shengyang Ping, Zhijie Lin, Liliang Lin, Lei Zhao, Lisha Ye, Bangguo Wang, Tao WangPromptable medical segmentation models can be sensitive to small targets, weak boundaries, heterogeneous texture, and inaccurate user prompts. We propose HBF-BCER, a prompt-preserving adaptation of the SAM ViT-B architecture for medical image segmentation. Historical Branch Fusion (HBF) reinjects aligned bottleneck summaries from previous enhanced encoder blocks, while Balanced Convolutional Expert Routing (BCER) applies load-balanced, pixel-wise soft weighting to heterogeneous convolutional experts. The evaluation uses fixed held-out evaluation sets, three training seeds, matched original SAM ViT-B, image-encoder and mask-decoder full fine-tuning, decoder-adapter, LoRA, and DoRA baselines, repeated prompt perturbations, paired statistics, and internal module ablations. HBF-BCER achieved Dice scores of 96.69±0.12% and 92.24±0.18% for REFUGE2 disc and cup, 96.18±0.31% for ISIC2016, and 92.08±0.19% for TNMIX. Image-encoder and mask-decoder full fine-tuning produced slightly higher Dice for the REFUGE2 disc and TNMIX, whereas HBF-BCER produced higher Dice for the REFUGE2 cup and ISIC2016 and the lowest HD95 on all four targets. Among the four paired comparisons with available per-image outputs, HBF-BCER showed significant gains for REFUGE2 cup versus SAM + Adapter and LoRA and for TNMIX versus SAM + Adapter; the ISIC2016 difference versus full fine-tuning was not significant. The method reduced trainable parameters by 81.5% relative to image-encoder and mask-decoder full fine-tuning but increased total parameters, FLOPs, and inference latency. These results support HBF-BCER as a parameter-efficient training strategy rather than a generally lightweight or uniformly superior model.