DOI: 10.3390/healthcare14162497 ISSN: 2227-9032

Google Translate Voice App vs. Qualified Interpreters: An Exploratory Study of Clinical Accuracy in Real-World Speech/Voice Encounters

Iris Feinberg, Heewon Lee-Laminack, Elizabeth L. Tighe, Ifedola Owoeye

Background/Objectives: AI voice interpretation applications like Google Translate in speech/voice mode are increasingly used in clinical settings to address verbal language access challenges, yet evidence comparing their performance to qualified in-person medical interpreters in authentic clinical encounters remains limited. The aim of this study is to compare the linguistic and clinical accuracy of AI-based Google Translate in speech/voice mode and qualified in-person medical interpretation using spoken real-world speech segments from clinical encounters. Methods: Outpatient clinical encounters involving patients with limited English proficiency were audio-recorded. Fourteen physician speech segments (mean length 78.5 words) representing common diagnostic, treatment, and counseling content were extracted, spoken in English into Google Translate speech/voice mode, and translated into six languages. Audio recordings of the same sentence segments were provided to qualified in-person medical interpreters. All non-English translations were back-translated into English by professional interpreters. Researchers and a clinician evaluated the back-translated English speech segments for linguistic accuracy, completeness, and clinical significance. Qualitative analyses examined error patterns and contextual loss; quantitative comparisons assessed error rate differences across languages and interpretation conditions. Results: Google Translate in speech/voice mode exhibited significantly higher linguistic errors (χ2[1] = 19.78, p < 0.001) and clinical accuracy errors (χ2[1] = 45.07, p < 0.001) than qualified medical interpreter translations. Clinically significant error rates were 33.3% for Google Translate in speech/voice mode versus 4.8% for interpreter-generated translations. Error rates were also higher for less commonly spoken languages compared to commonly spoken languages when using Google Translate in speech/voice mode (42.9% vs. 14.3%). Conclusions: Qualified in-person interpreters provided translated speech segments that had fewer errors that were either linguistically or clinically significant, and remain essential for safe, accurate clinical communication.

More from our Archive