DOI: 10.35377/saucis...1864801 ISSN: 2636-8129

Evaluating Large Language Models with a Unified Prompt-Based Pipeline for Turkish News Summarization and Tag Extraction

Burcu Koçak, Kazım Yıldız
News texts are critically important in facilitating readers quick access to accurate information in today’s substantially increasing information-driven world. The very fast pace of news flow and the abundance of highly informative data across various fields require the analysis of summarization and tag extraction tasks. This study aims to measure the multitasking performance of summarization and tag extraction tasks in Turkish news texts using a unified pipeline with a small datasets and different prompt techniques. Summarization and tag extraction tasks were performed with a single pipeline using GPT-4o and Claude 4.5 Sonnet on the commercial side, and Qwen3-32B and Llama 3.3-70B-Instruction models on the open-source side. Considering the requirements for fair evaluation, the experiments were conducted with the same system prompt, samples, and parameters. Experimenta results show that the Claude 4.5 Sonnet model exhibited the highest performance in the 3-shot scenario with a BERTScore of 0.676 and an LLM-as-a-judge score of 7.67. Furthermore, switching from a zero-shot to a few-shot strategy performed in up to a 35% increase in tagging performance. Additionally, the Llama 3.3 model showed performance close to commercial models in the tagging task with an F1 score of 0.287. These results show that the proposed unified pipeline method exhibits strong performance, validated in terms of accuracy and semantic similarity using BertScore, F1 score,and the LLM-as-a-judge method, which simulates human judgment. The proposed framework is expected to contribute as a reliable automation solution that increases efficiency in Turkish news processing processes.