DOI: 10.68337/cpsm.v1.i1.2026-008 ISSN:

MERGE-Med: Multi-Expert Retrieval-Grounded Engine for Medical Report Summarization Using Hierarchical RAG and Contrastive Learning

D. Sai Tejaswi, P. Raja Sekhar Reddy, K. Shailaja, G. Vishnu Murthy

Electronic health records (EHRs) have created large volumes of clinical data that require automated and accurate summarization that patients can understand. Current large language models (LLMs) are prone to hallucination and show limited task specialization and inadequate clinical grounding. MERGE-Med is a hybrid multi-expert architecture that consists of (i) a mixture-of-experts (MoE) layer with three domain-specialized LLaMA-3 8B models for diagnostic reasoning, treatment protocol generation, and patient-friendly communication; (ii) a hierarchical retrieval-augmented generation (RAG) pipeline that combines FAISS dense retrieval, BM25 sparse retrieval, cross-encoder reranking, and Unified Medical Language System (UMLS) knowledge-graph traversal over 35 million PubMed abstracts, 50,000 clinical guidelines, and 15,000 pharmaceutical records; and (iii) a contrastive fine-tuning framework validated by an ensembled DeBERTa-v3-large fact-checking layer that achieves more than 95% consistency. MERGE-Med achieved accuracies of 92.45%, 82.34%, and 98.76% on MedQA, MedMCQA, and Clinical Knowledge, respectively, improvements of 5.18, 6.44, and 1.78 percentage points (pp) over the baseline LLaMA-3 8B. The hallucination rate decreased by 67% (from 18.3% to 6.1%), ROUGE-L improved by 22.3%, and in an evaluation by board-certified physicians of 100 de-identified reports, 92% of summaries were judged appropriate for patients, with an inference latency of 2.3 s.