Back to all papers

Benchmarking 54 large language model configurations for CAD-RADS scoring: open-weight models approach human-level agreement.

July 30, 2026pubmed logopapers

Authors

Sandfort V,Vigneault DM,Willemink MJ,Wu J,Hallett RL,Nieman K,Fleischmann D,Mastrodicasa D

Affiliations (6)

  • Department of Radiology, Stanford University School of Medicine, 300 Pasteur Drive, Stanford, 94305, California, USA. [email protected].
  • Veterans Affairs Health Care System, Palo Alto, California, USA. [email protected].
  • Department of Radiology, Stanford University School of Medicine, 300 Pasteur Drive, Stanford, 94305, California, USA.
  • Segmed, Inc., #810-3790 El Camino Real, Palo Alto, 94306, California, USA.
  • Stanford Cardiovascular Institute, Stanford University, 265 Campus Drive, Stanford, 94305, California, USA.
  • Department of Radiology, University of Washington School of Medicine, 1959 NE Pacific St, Seattle, 98195, Washington, USA.

Abstract

The Coronary Artery Disease Reporting and Data System (CAD-RADS) standardizes coronary CT angiography (CCTA) reporting, but not all reports contain CAD-RADS classifications. We benchmarked 54 large language model (LLM) configurations across 50 distinct models, including recent proprietary and open-weight reasoning models, for zero-shot CAD-RADS classification. We retrospectively analyzed 500 anonymized CCTA reports from four hospitals across three U.S. regions. Expert cardiovascular radiologists provided the reference standard (human inter-rater [Formula: see text]). Fifty-four model configurations (50 distinct LLMs; four configurable models tested in both thinking and non-thinking modes) spanning Llama 2 7B through recent thinking models (DeepCogito v2, Gemini 3 Pro) processed reports using identical zero-shot prompts. Performance was measured with unweighted Cohen's κ. Two LLMs met both pre-specified non-inferiority criteria ([Formula: see text] margin and entire 95% CI within the human inter-rater agreement band, [Formula: see text]-0.956): Claude 4.6 Opus ([Formula: see text], 95% CI 0.817-0.889) and the open-weight Gemma 4 31B ([Formula: see text], 95% CI 0.810-0.882), which ranked second overall. Both met the criteria at the pre-specified [Formula: see text] margin; at [Formula: see text] no model qualified. On the reports originally dictated without a CAD-RADS statement (n=343), the top models reached [Formula: see text]. Performance declined with longer thinking chains (proxy for case complexity), but thinking-mode outperformed non-thinking mode on matched difficult reports. Gemma 4 31B remained non-inferior at 3-bit quantization and fits on a 24 GB consumer GPU. Current LLMs extract the CAD-RADS stenosis severity category from unstructured CCTA reports with agreement approaching the human inter-rater band, without task-specific training. Our data shows the remarkable rise of open models. A 31B open-weight model matched top proprietary systems and runs on consumer GPUs, enabling privacy-preserving local deployment for clinical data mining.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.