DOI: 10.3390/diagnostics16162495 ISSN: 2075-4418

Large Language Model Translation of BI-RADS Breast Imaging Reports into Arabic: A Blinded Expert Evaluation of Diagnostic Communication Safety

Mohammad Alarifi, Jake Luo, Abdulrahman Jabour, Yazeed Alashban, Meaad Almusined, Maram Mobara, Alhanouf Alshedi, Mansour Almanaa

Background/Objectives: Breast imaging reports contain Breast Imaging Reporting and Data System (BI-RADS) assessments and management recommendations that may be difficult for patients to understand across languages. Large language models (LLMs) may support patient-facing communication, but clinically important details must be preserved. This study compared radiologists’ opinions regarding the quality, clinical fidelity, safety of wording, and communication usefulness of patient-friendly Arabic translations of BI-RADS breast imaging reports generated by three LLMs. Methods: Five de-identified reports representing BI-RADS categories 0, 2, 3, 4, and 6 were translated from English into Arabic by DeepSeek, ChatGPT, and Gemini using an identical structured prompt. Fifty radiologists rated the blinded outputs across eight 5-point domains. Model ratings were compared using Friedman tests, Kendall’s W, and multiplicity-adjusted Wilcoxon signed-rank tests. Laterality was also verified against the source reports. Results: Gemini achieved the highest overall mean score (3.73 ± 0.78), followed by DeepSeek (3.54 ± 0.73) and ChatGPT (3.03 ± 0.70). The overall model effect was significant (χ2 = 34.11, df = 2, p < 0.001; Kendall’s W = 0.341). Gemini and DeepSeek each outperformed ChatGPT across all eight domains (adjusted p < 0.001), and Gemini outperformed DeepSeek overall (adjusted p = 0.009). No laterality errors were identified among the 15 translations. Conclusions: Performance remained model-dependent despite the shared prompt. Among the participating radiologists, Gemini received the highest expert ratings, while DeepSeek remained competitive. Because errors involving BI-RADS categories, laterality, measurements, lesion location, or recommendations could change diagnostic understanding, LLM-generated Arabic translations should serve as radiologist-reviewed communication aids rather than autonomous substitutes for clinical explanation.

More from our Archive