DOI: 10.3390/universe12100296 ISSN: 2218-1997

AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Model Capabilities in Astronomy

Jinghang Shi, Xiaoyu Tang, Yang Huang, Yuyang Li, Yanxia Zhang, Xiao Kong, Caizhan Yue, A-Li Luo

Astronomical image interpretation requires not only visual perception, but also familiarity with domain-specific representations, observational conventions, and astrophysical reasoning. Although recent multimodal large language models (MLLMs) have shown strong general visual–language capabilities, their performance on specialized astronomical figures remains insufficiently evaluated. To address this gap, we introduce AstroMMBench, an astronomy-specific benchmark for evaluating MLLMs on figure-associated astronomical multiple-choice questions. AstroMMBench contains 592 expert-screened multiple-choice questions covering six major astrophysical subfields: Astrophysics of Galaxies, Cosmology and Extragalactic Astrophysics, Earth and Planetary Astrophysics, High Energy Astrophysical Phenomena, Instrumentation and Methods for Astrophysics, and Solar and Stellar Astrophysics. The questions were generated from astronomical figures and associated textual context through an automated pipeline, followed by multi-stage filtering and expert review to assess scientific correctness, image–question alignment, and answer uniqueness. Using AstroMMBench, we evaluated 25 selected MLLMs, including 22 open-source and 3 closed-source models. The results show substantial performance differences across models and astrophysical subfields. Ovis2-34B achieved the highest point-estimate overall accuracy of 71.3%, with performance comparable to strong closed-source models such as ChatGPT-4o and Doubao-1.5-vision-pro. Subfield-level analysis shows the lowest mean point-estimate performance in cosmology, whereas high-energy astrophysics exhibits substantial between-model variability rather than uniformly low performance. Question-only results on the 592 benchmark questions yielded 27.53% for Qwen2.5-VL-7B and 24.83% for InternVL3-38B, compared with image-present accuracies of 58.11% and 68.24%, respectively. These descriptive comparisons suggest a substantial contribution from visual input in the two models. AstroMMBench provides a structured and extensible resource for comparing model behavior on figure-associated astronomical questions.