DOI: 10.3390/make8080248 ISSN: 2504-4990

Human-Level Extraction of Modified Rankin Scale Scores from Real-World Neurosurgical Clinical Notes Using a Locally Deployed, Quantized Large Language Model

Alban Bornet, Abiram Sandralegar, Anthony Yazdani, Elias Adam Benguettat, Michael Francis Righini, Feres Ravarelli, Paul Eugène Constanthin, Alexandre Lavé, Julien Haemmerli, Insa Janssen, Jelia Issa, Elham Qaderdan, Ethan Guillaume Godin, Nasser Bouhalassa, Karl Schaller, Philippe Bijlenga, Douglas Teodoro

This study conducts a clinical evaluation of a secure, locally deployed, quantized large language model (LLM) for automating modified Rankin Scale (mRS) score extraction from unstructured neurosurgical notes. We retrospectively selected 103 authentic clinical letters (2007–2025) from aneurysm patients at a tertiary neurosurgical centre. To comply with data privacy constraints, an open-source reasoning LLM (Qwen3-32B) with 4-bit quantization was deployed entirely on-premises. The LLM extracted mRS scores using a zero-shot approach with custom logits processors to enforce strict JSON formatting. Performance was compared to a reference standard (attending neurosurgeons’ consensus) and parallel scoring by medical residents and students. The LLM achieved excellent agreement with the attending consensus (QWK 0.95), matching the reliability of medical students (QWK 0.95) and residents (QWK 0.93). Exact agreement was 75%, and agreement within ±1 mRS point was 96%. Bayesian analysis strongly supported statistical equivalence between the model and human raters. The computationally optimized LLM demonstrated human-level classification reliability without task-specific fine-tuning. This approach successfully addresses key patient data privacy barriers and the formatting inconsistencies typical of open-ended generative models. Securely deploying a general-purpose, quantized LLM provides a scalable pathway for extracting functional outcomes and supports FAIR-aligned data systems.

More from our Archive