DOI: 10.1200/op-26-00317 ISSN: 2688-1527

Evaluating Artificial Intelligence Translation Tools for Language Equivalence of Oncology-Informed Consent Forms From English to Spanish

Shaurya Khanna, Sean Woo, Sandra P. Susanibar-Adaniya, Gary E. Weissman, Meghan L. Blair, Ximena Jordan Bruno, Carmen E. Guerra

PURPOSE

Approximately 8% of the US population speaks primary languages other than English. Limited English proficiency (LEP) contributes to under-representation of Hispanic patients in oncology clinical trials. Although certified translation services exist, they are time-consuming and costly. Artificial intelligence (AI)–generated translations of informed consent forms (ICFs) could provide low-cost alternatives, but data on accuracy and safety remain limited. We evaluated language equivalence of English-to-Spanish translations for three oncology clinical trial ICFs using two general-purpose AI translation tools (DeepL Pro and ChatGPT-4o) and a medically trained AI translation tool (Med_English2Spanish) compared with certified translations.

METHODS

Translational equivalence was assessed using a five-point Likert scale on five domains: Semantic, Idiomatic, Experiential, Conceptual, and Safety. Two native Spanish-speaking bilingual board-certified physicians independently scored each translation. Weighted Cohen's kappa determined inter-rater reliability, and the two-sample t-test compared AI-generated and certified translations.

RESULTS

Weighted Cohen's kappa (0.95, 95% CI 0.85 to 0.97) exhibited high inter-rater agreement. Certified translations exhibited the highest equivalence (mean = 4.99, SD = 0.02). ChatGPT-4o similarly demonstrated high equivalence (mean = 4.89, SD = 0.17). DeepL Pro scored well (mean = 4.43, SD = 0.07) but lower than certified translation ( P < 0.001). Med_English2Spanish demonstrated the lowest degree of equivalence (mean = 3.32, SD = 0.40) compared with certified translations ( P < 0.001).

CONCLUSION

Low-cost AI translations of ICFs exhibited variable language equivalence compared with certified translations across several domains. ChatGPT-4o scored nearly equivalent across domains in translating procedural trial information. While AI-generated translations are currently not suitable for clinical deployment without human review, this exploratory study supports further analysis of AI translation tools for reducing language barriers to LEP population enrollment.

More from our Archive