DOI: 10.3390/make8100305 ISSN: 2504-4990

Multidimensional Evaluation Framework for Local LLM-Based Clinical Summarization: A Cross-Lingual Prototype on Multi-Modality PACS-Derived Imaging Reports

Luis Vera, Olga Craveiro, Ricardo Malheiro, Manuel Dias, Ricardo Correia Bezerra

Large language models (LLMs) are increasingly proposed for clinical summarization, yet evaluations relying on textual similarity overlook clinically consequential failure modes such as fabrication, omission, contradiction, and negation inversion. We treat medical summarization as controlled clinical compression, operationalizing a multidimensional framework of seven quality dimensions, an eight-category error taxonomy, and fourteen replicability components. The prototype runs on N=511 de-identified Portuguese-language Picture Archiving and Communication System (PACS)-derived multi-modality imaging reports, generating English impressions via local gemma4:latest under three prompt variants (R03) and temperature control. Factuality was scored claim-by-claim by an LLM judge with a two-pass protocol yielding 100% label coverage on 6190 claims, cross-checked by three external judges and two non-radiologist physicians. Factuality reached S=99.38% (P1, 95% confidence interval (CI) 98.94–99.64), 99.54% (P2, 99.12–99.76), and 92.77% (P3, 91.60–93.79); critical errors were 0.50% (95% CI 0.35–0.71); critical omissions (source-to-summary) are outside the claim-level scheme. Pairwise Cohen’s κ among external judges ranged 0.677–0.908. A controlled Portuguese (PT) → PT re-run (n=50) yielded significantly lower factuality than PT → EN (Δ=−7 to −15 pp per variant, p<0.0001 paired), restricting claims to PT → EN. Human annotation elicited a factuality–completeness criterion gap; BERTScore F1 correlated weakly with factuality (r=0.047); temperature had no significant effect.