Back to all papers

Diagnostic performance and short-interval reproducibility of multimodal large language models in differentiating cholesteatoma from chronic otitis media using key-image temporal bone high-resolution computed tomography.

August 11, 2026pubmed logopapers

Authors

Yağcı B,Palaz S,Çetinkaya EA,Şenol U

Affiliations (3)

  • University of Health Sciences Türkiye, Antalya Training and Research Hospital, Clinic of Radiology, Antalya, Türkiye.
  • University of Health Sciences Türkiye, Antalya Training and Research Hospital, Clinic of Otorhinolaryngology, Antalya, Türkiye.
  • Akdeniz University Faculty of Medicine, Department of Radiology, Antalya, Türkiye.

Abstract

This study aimed to assess whether multimodal large language models (LLMs) can distinguish cholesteatoma from non-cholesteatomatous chronic otitis media (COM) on representative key-image temporal bone high-resolution computed tomography (HRCT) and to evaluate the short-interval reproducibility of their outputs. This retrospective, single-center study (2019-2024) included 101 patients (48 with cholesteatoma, 53 with non-cholesteatomatous COM) who underwent surgical treatment. The reference standard was intraoperative diagnosis with histopathological confirmation for cholesteatoma and surgical documentation for COM. For each case, six anonymized representative HRCT images reflecting standard diagnostic criteria were selected by consensus between a 4<sup>th</sup>-year radiology resident and a board-certified head and neck radiologist, both unaware of the diagnosis. Subsequently, the same images were analyzed by the GPT-5 and Gemini 2.5 Pro LLMs through their official web interfaces, utilizing structured prompts and a zero-shot approach. These evaluations were conducted in two distinct sessions (S1 and S2) with a 1-week interval. The primary endpoint was accurate binary classification. Accuracy, sensitivity, specificity, positive predictive value, and negative predictive value were calculated with 95% confidence intervals [(CIs); Wilson method] agreement with the reference standard and between sessions and models was assessed with the Cohen kappa (κ) coefficient; and differences in classification were assessed with the McNemar test. The radiologist achieved an accuracy of 96.0% (95% CI: 90.3-98.4) with almost perfect agreement with the reference standard (κ: 0.921). In S1, GPT-5 and Gemini 2.5 Pro achieved accuracies of 43.6% and 49.5%, and in S2, 46.5% and 47.5%, respectively. Both models combined high sensitivity (83.3%-97.9%) with low specificity (1.9%-13.2%), and balanced accuracy ranged from 0.45 to 0.52. Between-session reproducibility was fair for GPT-5 (κ: 0.360) and moderate for Gemini 2.5 Pro (κ: 0.485), and inter-model agreement was slight at both sessions (κ: 0.035 at S1 and κ: 0.086 at S2). Accuracy did not differ significantly between the two models (<i>P</i> = 0.211). In this single-center study, GPT-5 and Gemini 2.5 Pro, in the versions evaluated, combined high sensitivity with low specificity and showed only fair-to-moderate between-session reproducibility and exhibited slight inter-model agreement on temporal bone key-image HRCT. These findings do not support their use as independent second readers, and broader generalization to other multimodal LLMs would require the evaluation of additional models. The evaluated models lacked the spatial precision and consistency required for the accurate assessment of complex middle ear structures. This finding underscores the necessity for verification by a radiologist and continuous monitoring.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.