Back to all papers

Automatic extraction of structured information from brain MRI reports using an open-weight large language model.

August 18, 2026pubmed logopapers

Authors

Mouheb K,Pomp A,Manenti A,de Haan R,Faghir F,Martens J,Seelaar H,Mattace-Raso F,Vernooij MW,Wolters FJ,Klein S,Bron EE

Affiliations (7)

  • Department of Radiology & Nuclear Medicine, Erasmus MC, Rotterdam, The Netherlands. [email protected].
  • Department of Epidemiology, Erasmus MC, Rotterdam, The Netherlands.
  • Department of Radiology & Nuclear Medicine, Erasmus MC, Rotterdam, The Netherlands.
  • Department of Electrical and Electronics Engineering, ENSEEIHT, Toulouse, France.
  • Alzheimer Centre Erasmus MC, Erasmus MC, Rotterdam, The Netherlands.
  • Department of Neurology, Erasmus MC, Rotterdam, The Netherlands.
  • Department of Internal Medicine, Erasmus MC, Rotterdam, The Netherlands.

Abstract

Automatic data extraction from free-text radiology reports enables large-scale research, but few studies assessed the performance of large language models (LLMs) on Dutch neuroradiology reports. We analyzed 947 brain MRI reports from a tertiary memory clinic (2016-2021), authored by consultant neuroradiologists. Trained medical students annotated thirty variables; 100 reports were double annotated to assess inter-rater reliability. We evaluated the performance of the open-weight LLM LLaMA 3.1 using different languages (Dutch vs English translation) and few-shot prompting with different example selection strategies. Performance was evaluated using balanced accuracy for categorical variables, accuracy and mean absolute error for counts, and text similarity for free text. Metrics were computed across 10 random splits of the 947 reports. LLaMA 3.1 demonstrated high zero-shot performance for visual rating scores (mean [95% CI]): medial temporal atrophy: 0.90 [0.77-1.00] on the left and 0.96 [0.94-0.99] on the right, Global Cortical Atrophy: 0.87 [0.83-0.91], and Fazekas: 0.94 [0.93-0.96]. Microbleed mentions were detected with 0.93 [0.92-0.95] accuracy, and infarct mentions with 0.82 [0.80-0.84]. Text similarity for lesion location reached 0.95 [0.95-0.96]. Performance was lower for numerical variables: 0.80 [0.78-0.82] for the number of microbleeds and 0.66 [0.63-0.68] for infarcts. English translation yielded comparable results. Few-shot prompting improved performance for numerical variables, achieving 0.92 [0.90-0.93] for microbleeds and 0.81 [0.77-0.85] for infarcts using structural similarity-based selection. LLaMA 3.1 shows strong potential for extracting data from Dutch neuroradiology reports. Few-shot prompting enhances performance for numerical variables, whereas challenges remain for location-specific variables. Question The free-text format of radiology reports limits the accessibility of structured data; can LLM LLaMA 3.1 automatically structure free-text Dutch neuroradiology reports? Findings LLaMA 3.1 accurately extracted 30 variables from Dutch neuroradiology reports, with few-shot prompting significantly improving performance compared to zero-shot prompting, particularly with similarity-based example selection. Clinical relevance By enabling accurate, automated post-hoc structuring of free-text radiology reports, open-weight LLMs facilitate large-scale data reuse, improve consistency in reporting, and support clinical research and decision-making in dementia care without disrupting existing clinical workflows.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.