Back to all papers

Large Language Model Translation of BI-RADS Breast Imaging Reports into Arabic: A Blinded Expert Evaluation of Diagnostic Communication Safety.

August 7, 2026pubmed logopapers

Authors

Alarifi M,Luo J,Jabour A,Alashban Y,Almusined M,Mobara M,Alshedi A,Almanaa M

Affiliations (4)

  • Radiological Sciences Department, College of Applied Medical Sciences, King Saud University, Riyadh 4545, Saudi Arabia.
  • College of Health Sciences, University of Wisconsin-Milwaukee, Milwaukee, WI 53211, USA.
  • Health Informatics Department, Faculty of Public Health and Tropical Medicine, Jazan University, Jazan 45142, Saudi Arabia.
  • Radiology Department, Prince Sultan Military Medical City, Riyadh 12233, Saudi Arabia.

Abstract

<b>Background/Objectives:</b> Breast imaging reports contain Breast Imaging Reporting and Data System (BI-RADS) assessments and management recommendations that may be difficult for patients to understand across languages. Large language models (LLMs) may support patient-facing communication, but clinically important details must be preserved. This study compared radiologists' opinions regarding the quality, clinical fidelity, safety of wording, and communication usefulness of patient-friendly Arabic translations of BI-RADS breast imaging reports generated by three LLMs. <b>Methods:</b> Five de-identified reports representing BI-RADS categories 0, 2, 3, 4, and 6 were translated from English into Arabic by DeepSeek, ChatGPT, and Gemini using an identical structured prompt. Fifty radiologists rated the blinded outputs across eight 5-point domains. Model ratings were compared using Friedman tests, Kendall's W, and multiplicity-adjusted Wilcoxon signed-rank tests. Laterality was also verified against the source reports. <b>Results:</b> Gemini achieved the highest overall mean score (3.73 ± 0.78), followed by DeepSeek (3.54 ± 0.73) and ChatGPT (3.03 ± 0.70). The overall model effect was significant (χ<sup>2</sup> = 34.11, df = 2, <i>p</i> < 0.001; Kendall's W = 0.341). Gemini and DeepSeek each outperformed ChatGPT across all eight domains (adjusted <i>p</i> < 0.001), and Gemini outperformed DeepSeek overall (adjusted <i>p</i> = 0.009). No laterality errors were identified among the 15 translations. <b>Conclusions:</b> Performance remained model-dependent despite the shared prompt. Among the participating radiologists, Gemini received the highest expert ratings, while DeepSeek remained competitive. Because errors involving BI-RADS categories, laterality, measurements, lesion location, or recommendations could change diagnostic understanding, LLM-generated Arabic translations should serve as radiologist-reviewed communication aids rather than autonomous substitutes for clinical explanation.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.