DOI: 10.1002/qub2.70050 ISSN: 2095-4689

Benchmarking commercial large language models for gene–disease–phenotype extraction from full‐text human genetics literature

Danqing Yin, Matthew Ka Siu Leung, Darren Wan Ho Pun, Fiona Haixin Chen, Julie Yujin Kwon, Xinyi Lin, Joshua W. K. Ho

Abstract

Manual curation of gene–disease–phenotype relationships from the human genetics literature is a persistent bottleneck for maintaining its bioinformatics databases. Whereas large language models (LLMs) offer a promising alternative, there is currently no systematic benchmark that evaluates whether state‐of‐the‐art commercial LLMs can perform this task reliably on the full‐text articles. To address this gap, we introduce a standardized benchmark comprising 406 full‐text articles covering 180 congenital heart disease‐associated genes, and a multi‐dimensional evaluation framework that incorporates fuzzy matching to account for synonyms and partial matches. We benchmarked seven state‐of‐the‐art LLMs, GPT‐4o, Claude‐Opus‐4, DeepSeek‐R1, Grok‐4, Qwen‐3.5, Gemini‐2.5 (Pro), and GPT‐5 on the extraction of structured gene, disease, and phenotype fields. The top‐performing model, Grok‐4, achieved 97.6% overall accuracy, whereas the lowest‐performing model reached approximately 88%, still surpassing many prior benchmarks employing zero‐shot or n ‐shot prompting in biomedical relation extraction (RE) tasks. Our results provide a rigorous characterization of current LLMs capabilities and limitations. This paper contains two components. First, we conducted a human genetics field benchmark study on LLMs against a curated database. Second we developed the evaluation framework for this task. The benchmark dataset, evaluation framework, and model benchmarking outputs are made available online to support future studies in a reproducible manner.

More from our Archive