Diagnostic Performance of Multimodal Large Language Models for Grading Pediatric Vesicoureteral Reflux on MCUG: A Preliminary Study.
Authors
Affiliations (2)
Affiliations (2)
- Department of Radiodiagnosis and Imaging, Postgraduate Institute of Medical Education and Research, Chandigarh, India.
- Department of Pediatric Surgery, Postgraduate Institute of Medical Education and Research, Chandigarh, India.
Abstract
To evaluate the diagnostic accuracy of ChatGPT-5 and Gemini 2.5 Pro compared with a pediatric radiologist reference standard in grading vesicoureteral reflux (VUR) on pediatric micturating cystourethrogram (MCUG) studies. This retrospective study included 125 pediatric MCUG cases, comprising 25 cases per VUR Grade (I-V) as determined by radiologists. A balanced distribution of VUR grades was employed to facilitate a comprehensive assessment of model performance across the severity spectrum. For each case, the most representative image was anonymized, and the stored Joint Photographic Experts Group images were randomly uploaded to ChatGPT-5 and Gemini 2.5 Pro with a standardized prompt requesting VUR grading. Grades assigned by the models were compared with the radiologist's consensus standard. VUR was present in 182/250 (72.8%) renal units, more frequently on the left side (80%) than the right side (65.6%). Both models demonstrated the highest accuracy in identifying the absence of reflux (ChatGPT-5 = 92.6% and Gemini 2.5 Pro = 88.2%). Overall diagnostic accuracy was 52% (131/250) for ChatGPT-5 and 46% (115/250) for Gemini 2.5 Pro. ChatGPT-5 performed best for grades III-IV (51% and 49%, respectively), whereas Gemini 2.5 Pro performed best for Grade V (54%). Agreement with radiologists was fair (ChatGPT-5 κ = 0.40; Gemini 2.5 Pro κ = 0.33). ChatGPT-5 and Gemini 2.5 Pro demonstrated fair but suboptimal accuracy in grading VUR. Model performance appears to be influenced by the prompting strategy used. Multimodal large language models show potential as adjunctive tools; however, further optimization of prompting approaches and multimodal integration is required before clinical application.