DOI: 10.3390/jimaging12100476 ISSN: 2313-433X

Multi-Caption Guided Weakly Supervised Disease Localization on Medical Cancer Images

Dawit Shibabaw, Munir Awol, Vukosi Marivate, Tesfa Tegegne

We propose a weakly supervised framework for grounding multiple clinical captions to disease regions in medical images using image level labels without pixel level localization annotations. Unlike approaches that treat all textual information equally, the proposed framework models the hierarchical structure of radiology reports by learning the relative importance of individual captions. Building on the Language meets Vision Transformer (LViT), the framework incorporates three main components: (1) multi-caption encoding using BioClinicalBERT, (2) an attention-based caption importance module trained end-to-end under weak supervision, and (3) a weakly supervised grounding mechanism for disease-region localization. The proposed method achieved a caption importance correlation of 0.9552 with clinician-assigned importance rankings and demonstrated improved attention focus compared with the reported baselines (0.5011 vs. 0.4123, p < 0.01). Qualitative results further suggest that higher-weighted captions contribute more strongly to disease region localization, whereas lower-weighted captions provide supporting clinical context. The framework achieved lesion identification accuracy of 0.9811 under weak supervision. Domain experts evaluated the framework using caption importance correlation, attention focus, localization intersection over union, and lesion identification accuracy. Overall, the findings suggest that incorporating the hierarchical structure of clinical reports may improve the alignment between medical images and their associated textual descriptions and may help clinicians more easily interpret medical images of affected regions. The code developed for this study has been made publicly available on GitHub.