Human-AI collaboration using GPT-5 for MR grading of endolymphatic hydrops in Ménière's disease.
Authors
Affiliations (8)
Affiliations (8)
- Department of Otorhinolaryngology, Head, and Neck Surgery, The First Affiliated Hospital of Wenzhou Medical University, Wenzhou, Zhejiang, People's Republic of China.
- Wenzhou Medical University, Wenzhou, Zhejiang, People's Republic of China.
- Department of Radiology, The First Affiliated Hospital of Wenzhou Medical University, Wenzhou, Zhejiang, People's Republic of China.
- Department of Ultrasound, The First Affiliated Hospital of Wenzhou Medical University, Wenzhou, Zhejiang, People's Republic of China.
- Department of Otorhinolaryngology, Head, and Neck Surgery, The First Affiliated Hospital of Wenzhou Medical University, Wenzhou, Zhejiang, People's Republic of China. [email protected].
- Wenzhou Medical University, Wenzhou, Zhejiang, People's Republic of China. [email protected].
- Department of Otorhinolaryngology, Head, and Neck Surgery, The First Affiliated Hospital of Wenzhou Medical University, Wenzhou, Zhejiang, People's Republic of China. [email protected].
- Wenzhou Medical University, Wenzhou, Zhejiang, People's Republic of China. [email protected].
Abstract
To evaluate GPT-5 for grading cochlear and vestibular endolymphatic hydrops (EH) on delayed gadolinium-enhanced 3D-FLAIR MRI in Ménière's disease (MD) and to assess prompting strategies and human-AI collaboration. This retrospective study included 436 patients with MD (872 ears). GPT-5 graded cochlear and vestibular EH under multiple prompting strategies and was compared with junior and senior neuroradiologists in independent and GPT-5-assisted workflows. Accuracy (ACC), AUC, F1 score, agreement, and expert-rated clinical usefulness were assessed. Under basic prompting strategies without structured image descriptions, diagnostic performance remained poor: cochlear hydrops grading ACC reached only 25%, and vestibular hydrops grading ACC reached only 39%, despite the inclusion of clinical history or few-shot examples. In contrast, with enhanced prompting that integrated structured image descriptions, clinical context, and representative examples, GPT-5 showed substantially improved performance, achieving accuracies of 87% for cochlear grading and 80% for vestibular grading, with improved agreement and favorable expert ratings. In human-AI interaction experiments, the Human-first mode improved junior physicians' performance, increasing cochlear grading ACC from 36% to 59% and vestibular grading ACC from 42% to 53%. GPT-5 demonstrated limited capability to independently grade EH on MRI when relying primarily on image input. Acceptable diagnostic performance was achieved only after incorporating structured image descriptions, clinical context and representative examples. Under these enhanced prompting conditions, GPT-5 showed potential as an assistive tool for human-AI collaborative interpretation, particularly for less experienced physicians, rather than as a fully autonomous diagnostic system. Question Accurate grading of endolymphatic hydrops on 3D-FLAIR MRI is critical for managing Ménière's disease, but current interpretation lacks standardization and efficiency. Findings GPT-5 showed poor EH MRI grading accuracy without image descriptions, while optimized prompting and human-AI collaboration markedly improved performance toward expert interpretation. Clinical relevance Acceptable MRI grading of endolymphatic hydrops by GPT-5 requires structured image descriptions and contextual prompting. This supports its role as a human-AI collaborative assistive tool, not an autonomous system, potentially reducing reliance on limited medical resources.