MetSEval-1k: A Comprehensive Benchmark for Evaluating Large Language Models in Meteorology
Tingzhao Yu, Kuoyin Wang, Zhimin Li, Muhua Wang, Yu Chen, Hui Chen, Rui Zhao, Yucheng Xu, Yingying Song, Baowen XuThis paper proposes METMAP, a comprehensive 6D evaluation framework designed to assess large language models in meteorological applications. The framework encompasses six critical capabilities including meteorological knowledge comprehension, expert-level meteorological content summarization, multilingual translation of meteorological information, geospatial mapping context understanding, alignment with authoritative meteorological standards, and professional meteorological service communication. To operationalize this framework, the paper further introduces MetSEval-1k, a high-quality benchmark comprising 1083 expert-curated questions spanning operational and public-facing meteorological services. The benchmark integrates both objective multiple-choice items and subjective open-ended tasks to enable holistic model assessment. This paper conducts systematic evaluations of multiple state-of-the-art large language models using MetSEval-1k, revealing substantial performance disparities across the six dimensions. The results highlight the critical necessity for domain-specific adaptation and rigorous validation before deploying large language models in operational meteorological contexts. MetSEval-1k is released as a foundational benchmark to advance research and development of trustworthy, service-oriented artificial intelligence application in meteorology.