DOI: 10.1200/cci-26-00002 ISSN: 2473-4276

Data Extraction From Oncology Imaging Reports by Large Language Models: A Comparative Accuracy Study

Lea P. Passweg, Johannes M. Schwenke, Christof M. Schönenberger, Flavio Locher, Julia Picker, Manuel Dieterle, Benjamin Thiele, Dimitri Hasler, Alessia Danelli, Andreas M. Schmitt, Tobias Heye, Thomas Stojanov, Matthias Briel, Benjamin Kasenda

PURPOSE

Manual data extraction from clinical text is resource-intensive. Locally hosted large language models (LLMs) may offer a privacy-preserving solution, but their performance on non-English data remains unclear. We investigated whether the accuracy of locally hosted LLMs is noninferior to human accuracy when determining metastasis status and treatment response from German radiology reports.

METHODS

In this retrospective comparative accuracy study, five locally hosted LLMs (llama3.3:70b, mistral-small:24b, qwq:32b, qwen3:32b, and gpt-oss:120b) were compared against humans. A ground truth was established via duplicate human extraction and adjudication of discrepancies by a senior oncologist. The study was conducted at a tertiary referral hospital in Switzerland. We randomly sampled 400 radiology reports from adult patients with cancer (computed tomography, magnetic resonance imaging, positron emission tomography) generated between January 2023 and May 2025 and split them into a prompt optimization set (n = 100) and test set (n = 300). Primary outcomes were noninferiority (5 percentage points [pp] margin) of LLM classification accuracy compared with human accuracy for metastasis status (presence/absence by anatomic site) and treatment response categories. Secondary outcomes included accuracy for primary tumor diagnosis and radiologic absence of tumor.

RESULTS

The analysis included 400 reports from 317 patients. In the test set (n = 300), the human accuracy for metastasis status was 98.4% (95% CI, 98.0 to 98.8). All LLMs were noninferior; gpt-oss:120b performed best (97.6% accuracy; difference, –0.8 pp [90% CI, –1.3 to –0.3 pp]). For response to treatment, the human accuracy was 86.0% (95% CI, 83.2 to 88.8). All LLMs were inferior; the most accurate model, gpt-oss:120b, achieved 78.3% (difference, –7.7 pp [90% CI, –11.6 to –3.8 pp]).

CONCLUSION

In this study, LLMs were noninferior to human accuracy for classification of metastasis status but were inferior for response to treatment assessment.

More from our Archive