DOI: 10.35377/saucis...1780353 ISSN: 2636-8129

Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research

Burcu Baştürk, Aytuğ Onan
In this study, presented a comprehensive evaluation of abstractive and extractive summarization performance across three prominent large language models (LLMs): ChatGPT, DeepSeek, and Gemini. A total of 8,000 cardiovascular-related research abstracts were collected from PubMed and summarized using two distinct prompting strategies: abstractive and extractive. This process yielded a dataset of 48,000 summaries. To assess summarization quality, applied a multi-metric evaluation framework including semantic similarity (SBERT cosine), BLEU, GLEU, ROUGE-1 F1, ROUGE-2 F1, ROUGE-L F1, and METEOR. The results indicate that extractive summaries, particularly those generated by ChatGPT, consistently achieve higher scores across most metrics, suggesting stronger lexical fidelity and sequence retention. While Gemini shows balanced performance between abstraction and extraction, DeepSeek yields lower scores in both approaches. This work highlights critical differences in LLM behavior depending on the summarization method and offers a benchmark dataset and evaluation pipeline for future research on AI-assisted biomedical summarization.