Accuracy Evaluation of LLM-Generated Electronic Health Record Interpretations Against Medically Verified Sources
Mia Rovis, Alma Smajić, Marijela Miličević, Sandi Baressi Šegota, Vedran Mrzljak, Antun Gršković, Juraj Ahel, Klara Smolić, Dean Markić, Ivan LorencinThis study evaluates the accuracy of medical-report interpretations generated by large language models in comparison to medically verified sources, with a particular focus on urology. The main goal is to examine the extent to which LLMs can reliably and precisely explain medical findings, with emphasis on expert urological terminology. For the evaluation of interpretation accuracy, real clinical documentation was used, specifically a dataset of 40 urological reports from the Clinical Hospital Center Rijeka. The generated explanations were compared with medically correct definitions using quantitative metrics assessing interpretability, specificity, and clinical accuracy. The results revealed a clear stratification of model performance, with DeepSeek-V3.2, GLM-4.6, MiMo-V2-Flash, and Qwen2.5-7B-Instruct achieving the highest semantic accuracy and consistency, while several models exhibited unstable behavior characterized by occasional catastrophic failures. Overall, the findings indicate that although LLMs can produce clinically coherent explanations of urological findings, their variability and susceptibility to hallucinations necessitate human oversight, supporting their use primarily as decision-support tools rather than autonomous clinical interpreters.