Evaluating Artificial Intelligence Translation Tools for Language Equivalence of Oncology-Informed Consent Forms From English to Spanish
Shaurya Khanna, Sean Woo, Sandra P. Susanibar-Adaniya, Gary E. Weissman, Meghan L. Blair, Ximena Jordan Bruno, Carmen E. GuerraPURPOSE
Approximately 8% of the US population speaks primary languages other than English. Limited English proficiency (LEP) contributes to under-representation of Hispanic patients in oncology clinical trials. Although certified translation services exist, they are time-consuming and costly. Artificial intelligence (AI)–generated translations of informed consent forms (ICFs) could provide low-cost alternatives, but data on accuracy and safety remain limited. We evaluated language equivalence of English-to-Spanish translations for three oncology clinical trial ICFs using two general-purpose AI translation tools (DeepL Pro and ChatGPT-4o) and a medically trained AI translation tool (Med_English2Spanish) compared with certified translations.
METHODS
Translational equivalence was assessed using a five-point Likert scale on five domains: Semantic, Idiomatic, Experiential, Conceptual, and Safety. Two native Spanish-speaking bilingual board-certified physicians independently scored each translation. Weighted Cohen's kappa determined inter-rater reliability, and the two-sample t-test compared AI-generated and certified translations.
RESULTS
Weighted Cohen's kappa (0.95, 95% CI 0.85 to 0.97) exhibited high inter-rater agreement. Certified translations exhibited the highest equivalence (mean = 4.99, SD = 0.02). ChatGPT-4o similarly demonstrated high equivalence (mean = 4.89, SD = 0.17). DeepL Pro scored well (mean = 4.43, SD = 0.07) but lower than certified translation (
CONCLUSION
Low-cost AI translations of ICFs exhibited variable language equivalence compared with certified translations across several domains. ChatGPT-4o scored nearly equivalent across domains in translating procedural trial information. While AI-generated translations are currently not suitable for clinical deployment without human review, this exploratory study supports further analysis of AI translation tools for reducing language barriers to LEP population enrollment.