Back to all papers

Data Extraction From Oncology Imaging Reports by Large Language Models: A Comparative Accuracy Study.

August 10, 2026pubmed logopapers

Authors

Passweg LP,Schwenke JM,Schönenberger CM,Locher F,Picker J,Dieterle M,Thiele B,Hasler D,Danelli A,Schmitt AM,Heye T,Stojanov T,Briel M,Kasenda B

Affiliations (6)

  • Division of Medical Oncology, University and University Hospital Basel, Basel, Switzerland.
  • Division of Clinical Epidemiology, Department of Clinical Research, CLEAR Methods Center, University and University Hospital Basel, Basel, Switzerland.
  • Medical University Clinic, Kantonsspital Aarau (KSA), Aarau, Switzerland.
  • Division of Digitalization & ICT, University Hospital Basel, Basel, Switzerland.
  • Clinic of Radiology and Nuclear Medicine, University and University Hospital Basel, Basel, Switzerland.
  • Surgical Outcome Research Center, University and University Hospital Basel, Basel, Switzerland.

Abstract

Manual data extraction from clinical text is resource-intensive. Locally hosted large language models (LLMs) may offer a privacy-preserving solution, but their performance on non-English data remains unclear. We investigated whether the accuracy of locally hosted LLMs is noninferior to human accuracy when determining metastasis status and treatment response from German radiology reports. In this retrospective comparative accuracy study, five locally hosted LLMs (llama3.3:70b, mistral-small:24b, qwq:32b, qwen3:32b, and gpt-oss:120b) were compared against humans. A ground truth was established via duplicate human extraction and adjudication of discrepancies by a senior oncologist. The study was conducted at a tertiary referral hospital in Switzerland. We randomly sampled 400 radiology reports from adult patients with cancer (computed tomography, magnetic resonance imaging, positron emission tomography) generated between January 2023 and May 2025 and split them into a prompt optimization set (n = 100) and test set (n = 300). Primary outcomes were noninferiority (5 percentage points [pp] margin) of LLM classification accuracy compared with human accuracy for metastasis status (presence/absence by anatomic site) and treatment response categories. Secondary outcomes included accuracy for primary tumor diagnosis and radiologic absence of tumor. The analysis included 400 reports from 317 patients. In the test set (n = 300), the human accuracy for metastasis status was 98.4% (95% CI, 98.0 to 98.8). All LLMs were noninferior; gpt-oss:120b performed best (97.6% accuracy; difference, -0.8 pp [90% CI, -1.3 to -0.3 pp]). For response to treatment, the human accuracy was 86.0% (95% CI, 83.2 to 88.8). All LLMs were inferior; the most accurate model, gpt-oss:120b, achieved 78.3% (difference, -7.7 pp [90% CI, -11.6 to -3.8 pp]). In this study, LLMs were noninferior to human accuracy for classification of metastasis status but were inferior for response to treatment assessment.

Topics

NeoplasmsData MiningMedical OncologyJournal ArticleComparative Study

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.