Multilingual Conversational AI Chatbots for Efficient Healthcare Delivery During Case History-Taking: A Systematic Review
Rajashekhara Bhari Sharanesha, Deepti Virupakshappa, Alwaleed Abushanan, Sara AlghamdiLanguage barriers hinder healthcare, particularly during case history-taking, a key part of diagnosis. While multilingual artificial intelligence (AI) chatbots offer solutions, there is fragmented evidence of their effectiveness and impact. This systematic review followed PRISMA 2020 guidelines, examining studies published between 2015 and 2025 on multilingual AI chatbots in healthcare across four databases (Google Scholar, Scopus, Web of Science, and PubMed), using a two-stage screening process. Data extraction focused on applications, supported languages, underlying technologies, target populations, and clinical outcomes. From 503 records, 49 studies, covering primary care, telemedicine, oncology, mental health, and other areas, met the criteria. Supported languages included English, Spanish, Arabic, Chinese, Hindi, and other underrepresented languages. In individual system evaluations using heterogeneous methodologies and evaluation settings, AI chatbots achieved a diagnostic accuracy ranging from 72–92%. Core technologies included large language models (LLMs), bidirectional encoder representations from transformers (BERT), a generative pre-trained transformer (GPT), retrieval-augmented generation (RAG), speech recognition, and distillation. The findings show that these improve clinical workflow (30–70% time savings) and patient engagement, reduce language barriers, and promote health equity. However, the overall evidence certainty was low to moderate, reflecting the predominance of prototype and proof-of-concept studies. Multilingual AI chatbots demonstrate a boost in healthcare efficiency, a reduction in language barriers, and the promotion of health equity, but exhibit challenges regarding validation, workflow integration, and evaluation standards, along with ethical issues such as privacy and bias. Future research should include real-world studies, diverse populations, standardized outcome measures, and long-term equity assessments.