DOI: 10.3390/make8100300 ISSN: 2504-4990

RoboSurg-VQA: A Multimodal Visual Question Answering Benchmark Derived from Surgical Segmentation Data

Chengyi Zhang, Zi Ye, Ziyang Wang

Public surgical segmentation datasets contain information about the operative view that extends beyond their pixel labels. RoboSurg-VQA organises this information as a closed-set visual question answering (VQA) benchmark while retaining the source of each answer. It contains 5632 frames from 29 EndoVis 2017/2018 source sequences and 11 task units, with answers derived from source metadata, mask measurements, or machine-generated candidate labels. Seven visual attributes were examined in a blinded 250-frame audit by two study-team reviewers working independently, with disagreements resolved by a third reviewer. We trained a shared model with frozen BiomedCLIP encoders and a 13-answer vocabulary to examine the contributions of image and question inputs. Against mask-derived and candidate labels on held-out sequences, the model achieved a seven-task mean fixed-label Macro-F1 of 0.540. Answer frequency and task-restricted image-only prediction scored 0.391 and 0.400, respectively. With unseen paraphrases, the score was 0.528. On 40 audited frames from held-out sequences, mean Macro-F1 across bleeding, image quality, and glare was 0.535 against the human reference, compared with 0.450 for answer frequency, a gain of 0.085 driven by bleeding. The audit identified attribute-specific label errors, with the greatest reviewer disagreement for glare.