DOI: 10.3390/app16157795 ISSN: 2076-3417

Text2FHIRwallet: Automated Generation of FHIR Patient Summaries from Unstructured Cardiology Reports Using Fine-Tuned Portuguese Language Models—Development and Evaluation of a Health Professional Wallet

João C. Ferreira, Isabel Rosa, Ricardo Correia

Background and Objectives: Cardiology departments generate large volumes of unstructured free-text reports that impose substantial manual review burdens on clinicians; at Hospital de Santa Maria—Portugal’s largest public hospital—manual review of 12,651 reports took approximately seven minutes per report, representing over 1475 h of avoidable administrative work. This study presents Text2FHIRwallet, a health professional digital wallet that automates extraction and structuring of clinical entities from unstructured Portuguese cardiology reports using fine-tuned Named Entity Recognition (NER) models and maps the results to Fast Healthcare Interoperability Resources (FHIR) R4 patient summaries. Materials and Methods: Following the Design Science Research Methodology (DSRM) and CRISP-DM, we fine-tuned four transformer-based models—BERTimbau Base, BERTimbau Large, Albertina PT-PT, and MediAlbertina—on 305 manually annotated cardiology reports (77,309 tokens; κ = 0.85 inter-annotator agreement) covering eight clinical entity types, drawn from a corpus of 12,651 anonymised documents. Entities were mapped to FHIR R4 resources and delivered through a secure, role-based mobile wallet (React Native). Evaluation comprised token-level NER benchmarking with bootstrapped confidence intervals and McNemar’s testing, FHIR mapping accuracy assessment on 100 manually reviewed reports, processing-efficiency measurement, and a usability pilot with 10 cardiologists (SUS, NPS). Results: MediAlbertina achieved the highest NER performance (macro F1 = 0.985, 95% CI: 0.979–0.990), significantly outperforming all baseline models (p < 0.01, McNemar’s test) and comparing favourably with—though not directly comparable to, given differing languages and datasets—published benchmarks such as GPT-4 (F1 = 0.962 in ophthalmology NER) and fine-tuned BERT models for lung cancer NER (F1 ≈ 0.85–0.90). FHIR mapping accuracy was 98% on 100 independently reviewed reports. Report processing time was reduced from approximately seven minutes to 15–30 s (93–96% reduction), with peak batch-inference throughput of up to 1000 reports/h under parallelised GPU load (observed end-to-end throughput in pilot deployment was approximately 250 reports/h). The pilot usability evaluation yielded a SUS score of 87 (excellent) and an NPS of 80. Conclusions: Text2FHIRwallet demonstrates that domain-specific fine-tuning of a Portuguese-language pretrained language model achieves near-ceiling clinical NER accuracy, enabling scalable, interoperable, and privacy-compliant patient summary generation from unstructured cardiology text, offering an end-to-end pathway for integrating AI-driven NLP into clinical workflows and FHIR-based health information ecosystems, with implications for administrative efficiency, care coordination, and clinical research in non-English-language settings.

More from our Archive