Back to all papers

Large Language Models for Classifying Usual Interstitial Pneumonia from Radiology Reports: Native Reasoning Versus Structured Prompting.

August 5, 2026pubmed logopapers

Authors

Zhang R,Grist TM,Schiebler M,Wu Y,Sandbo N,Brasier AR,Chen GH

Affiliations (7)

  • Department of Radiology, School of Medicine and Public Health, University of Wisconsin-Madison, 600 Highland Avenue, Madison, WI, 53792, USA. [email protected].
  • Department of Medical Physics, School of Medicine and Public Health, University of Wisconsin-Madison, Madison, WI, USA. [email protected].
  • Department of Radiology, School of Medicine and Public Health, University of Wisconsin-Madison, 600 Highland Avenue, Madison, WI, 53792, USA.
  • Department of Medical Physics, School of Medicine and Public Health, University of Wisconsin-Madison, Madison, WI, USA.
  • Department of Medicine, School of Medicine and Public Health, University of Wisconsin-Madison, Madison, WI, USA.
  • Carbone Cancer Center, University of Wisconsin-Madison, Madison, WI, USA.
  • Institute for Clinical and Translational Research, University of Wisconsin-Madison, Madison, WI, USA.

Abstract

Extracting disease labels from radiology reports is essential for developing deep learning-based diagnostic models and enabling large-scale retrospective clinical research. Classification of usual interstitial pneumonia (UIP) patterns from high-resolution computed tomography (HRCT) reports according to Fleischner Society guidelines is a particularly demanding task, requiring synthesis of spatial distribution, fibrotic features, and exclusion criteria. As open-source large language models (LLMs) are released at an accelerating pace with steadily improving general benchmarks, a practical question arises: Do these improvements translate to better performance on complex, real-world clinical classification, and does the optimal prompting strategy differ across model architectures? While prior studies have evaluated LLMs for radiology report labeling, none have compared how prompting strategies interact with the native reasoning capabilities of newer model architectures. We evaluated 10 open-source LLMs from three architecture families (Llama, Qwen, Gemma) spanning 8 to 405 billion parameters, each tested with three prompting strategies on 270 HRCT reports classified by expert consensus of two senior thoracic radiologists. Four models with native reasoning ("thinking") capability were additionally tested in thinking mode. The best configuration achieved a Cohen's kappa (κ) of 0.70 and 82% four-class accuracy. Structured reasoning prompting improved all Llama models but degraded all models with native reasoning capability (Qwen 3.5 and Gemma 4), revealing an architecture-dependent interaction. Thinking mode hurt performance on criteria-based prompts and never yielded the best configuration. Larger models did not consistently outperform smaller ones: Llama 3.1 405B offered no advantage over Llama 3.3 70B, and the Qwen 397B model underperformed the dense Qwen 27B. These findings demonstrate that newer model generations with improved general benchmarks and larger parameter counts do not guarantee better performance on specialized medical classification tasks and that prompt design must be matched to model architecture.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.