Back to all papers

Bridging the Generalist-Subspecialist Gap with GPT-5-Thinking: Dual-Center Evaluation in Orbital and Head-and-Neck Tumor MRI Reports.

September 15, 2026pubmed logopapers

Authors

Li J,Xiao L,Qu X,Du L,Gong T,Yu Y,Li Y,Xie P,Chu G,Li H,Zhang Y,Gao H,Yuan Q,Han Q,Xian J,Liu J

Affiliations (4)

  • Department of Radiology, The Second Hospital of Jilin University, Changchun, Jilin 130041, China.
  • Department of Ophthalmology, The Second Hospital of Jilin University, Changchun, Jilin, China.
  • Department of Radiology, Beijing Tongren Hospital, Capital Medical University, Beijing, China.
  • Department of Radiology, Weifang People's Hospital, Shandong Second Medical University, Weifang, Shandong, China.

Abstract

Background Identifying tumor histologic subtypes in complex regions is challenging for generalist radiologists. Whether reasoning large language models can bridge this expertise gap by interpreting radiologic descriptions remains underexplored. Purpose To evaluate whether Generative Pretrained Transformer (GPT)-5-Thinking can perform similarly to subspecialists and help bridge the generalist expertise gap in interpretation of MRI scans of orbital and head-and-neck tumors. Materials and Methods This dual-center retrospective study was performed at a generalist practice center and a subspecialist practice center and included convenience series of patients with pathologically confirmed orbital or head-and-neck tumors who had pretreatment MRI reports. GPT-5-Thinking processed narrative MRI reports to output benign-malignant classification and top-3 differential diagnoses. Phase 1 compared model accuracy with routine clinical report diagnoses in generalist and subspecialist practice settings. In Phase 2, seven generalists interpreted selected tumor reports twice, without and with GPT-5-Thinking assistance. McNemar tests and mixed-effects logistic regression models were used for analysis. Results This study included 1000 patients (mean age ± SD, 51 years ± 17.0; 520 men). In phase 1, the top-1 accuracy of GPT-5-Thinking was similar to that of subspecialists in patients with orbital tumors (76.7% [230 of 300] vs 78.3% [235 of 300]; <i>P</i> = .64) and head-and-neck tumors (67.5% [135 of 200] vs 68.0% [136 of 200]; <i>P</i> = .92). GPT-5-Thinking had higher top-1 accuracy than did routine generalists for patients with orbital tumors (74.0% [222 of 300] vs 68.7% [206 of 300]; <i>P</i> = .03) and head-and-neck tumors (64.0% [128 of 200] vs 27.0% [54 of 200]; <i>P</i> < .001). In phase 2, GPT-5-Thinking assistance was associated with an increase in generalist top-1 diagnostic accuracy from 61.4% (430 of 700) to 70.3% (492 of 700) (<i>P</i> < .001) in patients with orbital tumors and from 47.0% (329 of 700) to 61.0% (427 of 700) (<i>P</i> < .001) in patients with head-and-neck tumors. Conclusion Using MRI reports of patients with orbital or head-and-neck tumors, Generative Pretrained Transformer (GPT)-5-Thinking showed diagnostic accuracy similar to that of subspecialists; GPT-5-Thinking assistance was also associated with higher generalist diagnostic accuracy. © RSNA, 2026 <i>Supplemental material is available for this article.</i>

Topics

Head and Neck NeoplasmsMagnetic Resonance ImagingOrbital NeoplasmsImage Interpretation, Computer-AssistedJournal ArticleMulticenter Study

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.