Back to all papers

Commercial large language models for oral cavity cancer staging using descriptive pre-treatment MRI reports: ready for standalone use in clinical practice?

September 21, 2026pubmed logopapers

Authors

Ai QYH,Kwok HM,Lu MY,Lau TTS,Hung KF,Wong LM,So TY,King AD,Leung HS

Affiliations (10)

  • Department of Imaging and Interventional Radiology, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong SAR. [email protected].
  • Department of Imaging and Interventional Radiology, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong SAR.
  • Department of Diagnostic and Interventional Radiology, Princess Margaret Hospital, Hospital Authority, Kowloon, Hong Kong SAR.
  • Department of Oral and Maxillofacial Surgery, Chung Shan Medical University Hospital, Taichung, Taiwan.
  • School of Dentistry, College of Oral Medicine, Chung Shan Medical University, Taichung, Taiwan.
  • Department of Oncology, Princess Margaret Hospital, Hospital Authority, Kowloon, Hong Kong SAR.
  • Oral and Maxillofacial Radiology, Applied Oral Sciences & Community Dental Care, Faculty of Dentistry, The University of Hong Kong, Pokfulam, Hong Kong SAR.
  • Department of Imaging and Interventional Radiology, Prince of Wales Hospital, Hospital Authority, Shatin, Hong Kong SAR.
  • Department of Imaging and Interventional Radiology, Faculty of Medicine, The Chinese University of Hong Kong, Shatin, Hong Kong SAR. [email protected].
  • Department of Imaging and Interventional Radiology, Prince of Wales Hospital, Hospital Authority, Shatin, Hong Kong SAR. [email protected].

Abstract

MRI is used for staging head and neck cancer (HNC), but assigning T- and N-category criteria requires specialised expertise, and so many institutions offer only descriptive MRI reports. This study assessed potentials of commercial large language models (LLMs) to stage oral cavity cancer (OCC) using descriptive MRI reports and compared their accuracy with human experts. 104 eligible MRI reports were processed by five commercial LLMs (ChatGPT5.4, ChatGPT5.0, ChatGPT4.1, Gemini3.1 and DeepSeekV3.2). T- and N-categories, and overall stage were extracted from outputs of the LLMs. MRI reports were also staged by two multidisciplinary team (MDT) members. Accuracies of the LLMs and MDT members for cancer staging were assessed against the reference staging standard and compared using the McNemar test. The LLMs showed accuracy of 54.8-76.0% for T-categorisation, 63.5-88.5% for N-categorisation, and 53.8-80.8% for overall stage. Compared with DeepSeekV3.2, ChatGPT and Gemini3.1 showed significantly higher accuracy (p ≤ 0.001), except for Gemini3.1 for T-categorisation (p = 0.07). No differences in accuracy for staging between ChatGPT versions and between them and Gemini3.1 (p = 0.08 to > 0.99). MDT members showed accuracy of 74.0-76.9% for T-categorisation, 76.9-77.9% for N-categorisation and 68.3-71.2% for overall stage. The tested LLMs did not consistently outperform MDT members for staging. Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. Additionally, Low MDT staging accuracy highlighted the need for structured reports that include clinical cancer staging to facilitate MDT assessment, thus potentially ensuring optimised disease management. Question Can commercial LLMs accurately assign cancer stage based on descriptive MRI reports to overcome the lack of specialised expertise in clinical practice? Findings Tested LLMs demonstrated variable performances, ranging 54.8-76.0% for T-categorisation, 63.5-88.5% for N-categorisation and 53.8-80.8% for overall stage, and failed to consistently outperform MDT members. Critical relevance statement Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. MDT members' low performances highlight that structured MRI reports with clinical cancer staging are needed to improve multidisciplinary assessment.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.