Back to all papers

Diagnostic Performance of Multimodal Large Language Models for Grading Pediatric Vesicoureteral Reflux on MCUG: A Preliminary Study.

May 20, 2026pubmed logopapers

Authors

Saini S,Bhatia A,Saxena AK,Kanojia RP,Sodhi KS

Affiliations (2)

  • Department of Radiodiagnosis and Imaging, Postgraduate Institute of Medical Education and Research, Chandigarh, India.
  • Department of Pediatric Surgery, Postgraduate Institute of Medical Education and Research, Chandigarh, India.

Abstract

To evaluate the diagnostic accuracy of ChatGPT-5 and Gemini 2.5 Pro compared with a pediatric radiologist reference standard in grading vesicoureteral reflux (VUR) on pediatric micturating cystourethrogram (MCUG) studies. This retrospective study included 125 pediatric MCUG cases, comprising 25 cases per VUR Grade (I-V) as determined by radiologists. A balanced distribution of VUR grades was employed to facilitate a comprehensive assessment of model performance across the severity spectrum. For each case, the most representative image was anonymized, and the stored Joint Photographic Experts Group images were randomly uploaded to ChatGPT-5 and Gemini 2.5 Pro with a standardized prompt requesting VUR grading. Grades assigned by the models were compared with the radiologist's consensus standard. VUR was present in 182/250 (72.8%) renal units, more frequently on the left side (80%) than the right side (65.6%). Both models demonstrated the highest accuracy in identifying the absence of reflux (ChatGPT-5 = 92.6% and Gemini 2.5 Pro = 88.2%). Overall diagnostic accuracy was 52% (131/250) for ChatGPT-5 and 46% (115/250) for Gemini 2.5 Pro. ChatGPT-5 performed best for grades III-IV (51% and 49%, respectively), whereas Gemini 2.5 Pro performed best for Grade V (54%). Agreement with radiologists was fair (ChatGPT-5 κ = 0.40; Gemini 2.5 Pro κ = 0.33). ChatGPT-5 and Gemini 2.5 Pro demonstrated fair but suboptimal accuracy in grading VUR. Model performance appears to be influenced by the prompting strategy used. Multimodal large language models show potential as adjunctive tools; however, further optimization of prompting approaches and multimodal integration is required before clinical application.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.