DOI: 10.1145/3802546 ISSN: 1551-6857
Cross-Modal Comprehensive Multi-Level Granularity Alignment for Vision-Language Pre-Training
Ju Jiang, Darko B. Vukovic, Jie Cao, Haoran Xie, Youquan Wang
Vision-Language Pre-training (VLP) aims to convert vision-language tasks into image-text matching challenges, necessitating robust interaction between visual and textual modalities. However, current research often faces a trade-off: methods pursuing fine-grained alignment typically rely on computationally expensive object detectors and pre-defined annotations, while coarse-grained alignment methods overlook subtle visual details. To alleviate these issues, we present OMEGA, a C
O
mprehensive Multi-l
E
vel
G
ranularity
A
lignment framework. OMEGA achieves sufficient cross-modal alignment without relying on external object detectors or annotations. Specifically, we first design a Text-Aware Patch Selection (TAPS) module, which dynamically constructs patch-level regions for fine-grained excavation based on informative token selection, effectively bypassing the need for bounding boxes. Secondly, to facilitate multi-grained information fusion, we introduce a Cross-Grained Alignment (CGA) module to learn modality-shared features across global and object levels. Experimental results demonstrate that OMEGA surpasses existing approaches in five pivotal vision-language tasks—image-text retrieval, visual question answering, visual reasoning, visual entailment, and image captioning, while maintaining superior inference efficiency compared to detector-based methods. We believe this study offers a promising object-free paradigm for scalable vision-language pre-training.