A Rationale-Conditioned Image-Contrast Audit of OCT Dependence in Vision-Language Models for Anti-VEGF Treatment-Response Prediction
Wonbong Jang, Gwon Yul Jo, Siyun Lee, Shin Jeong Yoon, Gi Young Lee, Hee-Eun Lee, Tae Hyung Kim, Jong Won Baek, Joonhyung KimVision-language models (VLMs) are increasingly evaluated for retinal image interpretation, but end-task discrimination does not establish whether a prediction depends on the optical coherence tomography (OCT) image. We evaluated RetinaVLM, LLaVA-Med, and Qwen3.6-27B under zero-shot and parameter-efficient adaptation using an APTOS-2021 cohort (128 training, 21 validation, and 69 test eyes) and a 100-eye cross-site stress-test cohort. The primary readout compared forced-choice continue/stop scores obtained with a real OCT B-scan and a uniform-grey image while holding the generated rationale fixed; it therefore estimates a rationale-conditioned direct image contrast rather than total image dependence. Across 22 internal cells, 21 confidence intervals included an AUC of 0.5, while one Qwen zero-shot cell was inversely aligned (AUC 0.354, 95% CI 0.224–0.493). No positively aligned cells survived the Benjamini–Hochberg adjustment. Selected backbones produced image-responsive biomarker outputs, including pigment epithelial detachment balanced accuracy up to 0.93, whereas a five-field tabular reference model achieved decision AUC 0.731 (95% CI 0.603–0.846). Similar direct-contrast findings occurred in the second-site stress test. Because the generated rationales were not verified as faithful representations of latent computation, the grey image was out of distribution, and the cohorts were small, null direct contrasts cannot show that the models ignored OCT or establish equivalence to no discrimination.