Automating the Management of Extraspinal Findings in Magnetic Resonance Imaging Spine Studies Using a Privacy-Preserving Large Language Model: Retrospective Validation Study.
Authors
Affiliations (7)
Affiliations (7)
- Department of Diagnostic Imaging, National University Hospital, Singapore, Singapore.
- Department of Diagnostic Radiology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore, Singapore.
- Artificial Intelligence Innovation Office, National University Health System, Singapore, Singapore.
- School of Clinical Medicine, University of Cambridge, Cambridge, United Kingdom.
- Agency for Science, Technology and Research (A*STAR), Singapore, Singapore.
- Biostatistics Unit, Yong Loo Lin School of Medicine, National University of Singapore, Singapore, Singapore.
- National University Spine Institute, Department of Orthopaedic Surgery, National University Health System, Singapore, Singapore.
Abstract
Magnetic resonance imaging (MRI) spine studies frequently reveal extraspinal findings (ESFs) that require further evaluation; yet, the current process of manually reviewing radiology reports and navigating electronic medical records (EMRs) is time-consuming, labor-intensive, and prone to human error. To address this challenge, we propose using a privacy-preserving large language model (PP-LLM) to automate the identification, classification, and referral assessment of ESFs. A retrospective analysis of 405 consecutive MRI spine reports from the National University Hospital database, covering February to June 2024, was conducted. Two independent clinicians reviewed the reports and cross-referenced them with EMRs to identify ESFs from the imaging reports. The CT Extracolonic Findings Reporting and Data System (C-RADS) was adapted to determine the clinical significance of ESFs and whether specialty referral was required. The PP-LLM was designed to extract these findings, differentiate between new and preexisting conditions, classify their clinical significance, and generate appropriate referrals. A total of 405 MRI spine reports were initially identified. Five reports were excluded because no relevant EMRs were available, leaving 400 MRI reports from 395 patients for analysis. Among 395 patients (male: 48.1%, n=190; female: 51.9%, n=205; age: mean 54.7, SD 16.7, range 17-89 years), 163 (41.3%) had no ESFs and 232 (58.7%) were reported to have had at least one ESF. A total of 401 ESFs were identified, with the most common findings being renal (n=128, 31.9%), gynecological (n=110, 27.4%), and endocrine-related (n=48, 12%). The PP-LLM correctly detected 99.8% (400/401) of all ESFs and correctly identified all clinically urgent findings (3%, 12/401 of the total, for example, aortic dissection). It misclassified 2% (8/401) of cases into lower C-RADS categories, potentially downgrading clinically significant findings (eg, paranasal sinus mucosal thickening), and 1% (4/401) into higher C-RADS categories, upgrading clinically insignificant findings (eg, dependent changes in the lungs). Additionally, it achieved 98.5% accuracy in distinguishing new from preexisting findings and 96.3% accuracy in assigning the correct specialty referral decision (whether referral is needed, and correct subspecialty referral suggested). The PP-LLM demonstrated almost perfect agreement with the reference standard (Gwet κ=0.959, 95% CI 0.937-0.981), comparable to two human readers (reader 1: κ=0.943; reader 2: κ=0.934), with no statistically significant difference in C-RADS classification accuracy. Notably, it completed the analysis of each report in less than 5 seconds. The PP-LLM demonstrated high accuracy and efficiency in automating the identification and classification of ESFs in MRI spine reports. By integrating this AI-driven automation into clinical workflows, this technology has the potential to enhance efficiency, reduce clinician administrative burden, ensure timely specialist referrals, and improve patient care.