Improving Cross-Organ Generalization in Histopathology Segmentation via Evidence-Guided Vision–Language Query Decoding
Biwen Meng, Jiahao Wang, Jingxin LiuDomain shift remains a major obstacle to robust histopathology image segmentation, especially when models trained on several source organs are deployed to unseen anatomical sites. This study addresses cross-organ adenocarcinoma segmentation by introducing an evidence-guided vision–language segmentation framework that incorporates pathology-relevant morphological evidence into dense mask prediction. The proposed method uses a pathology vision–language encoder to extract image and text representations, a Semantic Query Booster to form image-aware segmentation queries, and an evidence-guided query recalibration that integrates positive tumor-supporting evidence and negative misleading evidence. Experiments were conducted on cross-organ adenocarcinoma datasets from the COSAS challenge under a source-only domain generalization setting, with colorectum, stomach, and pancreas as source domains and ampullary, gallbladder, and intestine as unseen target domains. The proposed framework achieved the highest pooled performance on the seen, unseen, and overall evaluation sets among the compared segmentation, domain generalization, and foundation model-based systems. These findings support the use of structured pathology evidence for cross-organ tumor segmentation under source-only training.