Design and Evaluation of a Source-Grounded Medical LLM for Clinical Decision Support and Patient Care in Trustworthy Diagnostic Systems
Muhammad Jamil, Adnan Kavak, Sevinç İlhan Omurca, Hossein FotouhiBackground/Objectives: Large language models (LLMs) are increasingly explored for clinical decision support and digital health applications. However, reliable diagnostic assistance remains challenging for low-resource medical languages such as Turkish due to limited source grounding, transparency, clinical safety, and localized medical knowledge. This study presents TurkishMedLLM, a source-grounded and safety-aware Turkish medical LLM designed for clinician-supervised diagnostic decision support, symptom interpretation, and patient care. Methods: The methodology integrates multi-source Turkish medical data ingestion, schema standardization, duplicate removal, quality filtering, supervised fine-tuning, embedding generation, vector indexing, Qwen3-8B fine-tuning, retrieval-augmented generation (RAG), and multi-layer evaluation. The system was evaluated using retrieval metrics, ROUGE, RAGAS, DeepEval, and clinical safety assessments. As a use case, TurkishMedLLM was integrated into the AI-based Diabetes Care (AIDCare) mHealth platform, which supports patient queries related to lifestyle management, symptoms, diagnosis, and treatment of diabetes. The system generates safety-aware responses with clinician-in-the-loop validation before delivery through the mobile application. Results: After pre-processing, the final dataset comprised 232,926 unique documents, including 210,791 Turkish medical question–answer pairs and 22,135 hospital medical articles. The retrieval module achieved Hit Rate@1, Hit Rate@3, and Hit Rate@5 of 94.67%, 98.67%, and 100.00%, respectively, indicating consistent retrieval of clinically relevant evidence. QLoRA fine-tuning achieved a validation loss of 0.9373 and ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-Lsum scores of 0.8338, 0.5476, 0.7836, and 0.7372, respectively. The fine-tuned RAG system achieved RAGAS faithfulness and response relevancy scores of 0.91 and 0.88, respectively, while DeepEval achieved an answer relevancy score of 0.90. Clinical safety assessment achieved a caution score of 0.93, indicating generally evidence-grounded and clinically cautious responses for symptom interpretation and diagnosis-related patient support. Conclusions: Combining retrieval grounding, parameter-efficient fine-tuning, and multi-layer safety evaluation provides a promising approach for clinician-supervised medical AI in Turkish. TurkishMedLLM demonstrates potential for symptom interpretation, differential diagnostic support, and trustworthy digital health applications.