Evaluation methods for Traditional Chinese Medicine large language models: A clinical practice study
Nanxing Xian, Wentong Zhu, Lei Zhang, Yiqing Liu, Yuping Zhao, Hengyi TianObjective:
To address the critical gap that existing medical large language model evaluation systems are predominantly based on Western medical paradigms and lack specialized assessment standards for the field of traditional Chinese medicine (TCM), this study constructs a comprehensive evaluation framework specifically designed for TCM clinical large language models, providing scientific assessment tools for their development, validation, and clinical application.
Methods:
This framework addresses the critical issue of missing TCM evaluation standards by incorporating TCM-specific elements such as pattern differentiation and treatment determination, as well as formula and herb recommendations, while emphasizing patient safety and data privacy.
Results:
The framework encompasses three core dimensions: (1) clinical scenario coverage and authenticity, including renowned physician case mining, pattern identification assistance, and therapy recommendations; (2) core diagnostic and therapeutic capability support, including symptom terminology recognition, pattern identification reasoning, and medical record generation; and (3) clinical application maturity and safety assurance, including data diversity, output reliability, and privacy protection.
Conclusion:
This evaluation system provides essential tools for the development and clinical deployment of TCM large language models, facilitating the responsible integration of AI technology into traditional medicine practice while preserving the theoretical integrity of TCM.