Benchmarking 54 large language model configurations for CAD-RADS scoring: open-weight models approach human-level agreement.
Authors
Affiliations (6)
Affiliations (6)
- Department of Radiology, Stanford University School of Medicine, 300 Pasteur Drive, Stanford, 94305, California, USA. [email protected].
- Veterans Affairs Health Care System, Palo Alto, California, USA. [email protected].
- Department of Radiology, Stanford University School of Medicine, 300 Pasteur Drive, Stanford, 94305, California, USA.
- Segmed, Inc., #810-3790 El Camino Real, Palo Alto, 94306, California, USA.
- Stanford Cardiovascular Institute, Stanford University, 265 Campus Drive, Stanford, 94305, California, USA.
- Department of Radiology, University of Washington School of Medicine, 1959 NE Pacific St, Seattle, 98195, Washington, USA.
Abstract
The Coronary Artery Disease Reporting and Data System (CAD-RADS) standardizes coronary CT angiography (CCTA) reporting, but not all reports contain CAD-RADS classifications. We benchmarked 54 large language model (LLM) configurations across 50 distinct models, including recent proprietary and open-weight reasoning models, for zero-shot CAD-RADS classification. We retrospectively analyzed 500 anonymized CCTA reports from four hospitals across three U.S. regions. Expert cardiovascular radiologists provided the reference standard (human inter-rater [Formula: see text]). Fifty-four model configurations (50 distinct LLMs; four configurable models tested in both thinking and non-thinking modes) spanning Llama 2 7B through recent thinking models (DeepCogito v2, Gemini 3 Pro) processed reports using identical zero-shot prompts. Performance was measured with unweighted Cohen's κ. Two LLMs met both pre-specified non-inferiority criteria ([Formula: see text] margin and entire 95% CI within the human inter-rater agreement band, [Formula: see text]-0.956): Claude 4.6 Opus ([Formula: see text], 95% CI 0.817-0.889) and the open-weight Gemma 4 31B ([Formula: see text], 95% CI 0.810-0.882), which ranked second overall. Both met the criteria at the pre-specified [Formula: see text] margin; at [Formula: see text] no model qualified. On the reports originally dictated without a CAD-RADS statement (n=343), the top models reached [Formula: see text]. Performance declined with longer thinking chains (proxy for case complexity), but thinking-mode outperformed non-thinking mode on matched difficult reports. Gemma 4 31B remained non-inferior at 3-bit quantization and fits on a 24 GB consumer GPU. Current LLMs extract the CAD-RADS stenosis severity category from unstructured CCTA reports with agreement approaching the human inter-rater band, without task-specific training. Our data shows the remarkable rise of open models. A 31B open-weight model matched top proprietary systems and runs on consumer GPUs, enabling privacy-preserving local deployment for clinical data mining.