Back to all papers

Multimodal Large Language Models for Breast Ultrasound Report Auditing: Workflow-Error Detection, False-Positive Burden, and Limits of Key-Image Interpretation.

August 3, 2026pubmed logopapers

Authors

Zou M,Xiao M,Zhu Q,Zhang J,Li J,Lv K

Affiliations (2)

  • Department of Ultrasound, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, No. 1 Shuaifuyuan, Dongcheng District, Beijing, 100730, China.
  • Department of Ultrasound, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, No. 1 Shuaifuyuan, Dongcheng District, Beijing, 100730, China. [email protected].

Abstract

Large language models (LLMs) are increasingly being evaluated for radiology reporting. However, the incremental value of adding report-embedded key images to report text for breast ultrasound report auditing remains unclear. This retrospective study included 818 breast ultrasound examinations with pathology- or follow-up-based reference standards. A workflow-error-enriched 300-report subset was constructed, comprising 240 reports with 329 inserted errors and 60 error-free reports. GPT-5.5 and Gemini 3.1 Pro Preview were evaluated under two input settings: report text alone and multimodal input; the latter additionally included report-embedded key images. A physician reader provided a human benchmark. In the full cohort, GPT-5.5 was evaluated with key-image input versus physician-interpreted findings input for malignancy and BI-RADS risk classification. Multimodal input increased report-level sensitivity from 77.5% to 89.6% for GPT-5.5 and from 86.2% to 96.2% for Gemini. Error-level recall increased from 63.2% to 75.4% and from 72.0% to 83.6%, respectively. Gains were concentrated in errors involving image-displayed cues. Gemini with multimodal input produced false-positive outputs in 18/60 error-free reports and 12/240 error-enriched reports, mainly from body-marker misinterpretation. The physician reader achieved 91.7% report-level sensitivity and 83.0% error-level recall without false positives. In the full cohort, key-image input underperformed physician-interpreted findings input for malignancy classification (AUC, 0.879 vs 0.971) and BI-RADS risk classification (AUC, 0.887 vs 0.984). Multimodal input improved LLM-based breast ultrasound workflow-error detection, with gains concentrated in errors involving image-displayed cues. However, model-specific overcalling and the inferior performance of key-image input compared with physician-interpreted findings warrant caution in clinical implementation.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.