At SIIM 2026, LLMs demonstrated high accuracy on radiology-based numerical tasks, particularly in extraction and judgment tests.
Key Details
- 1LLMs evaluated included Llama 3.1 8B, DeepSeek R1-distilled Llama 8B, OpenAI o1-mini, and OpenAI GPT 5-mini.
- 2Tasks tested involved extraction and judgment from DEXA, ultrasound, CT, and PET radiology reports.
- 3Most models, except Llama, achieved over 95% accuracy on extraction tasks; Llama ranged from 86% to 98.7%.
- 4GPT 5-mini achieved highest minimum accuracy (judgment tasks: 91.7%) among tested models.
- 5o1-mini and GPT 5-mini reached perfect accuracy in detecting osteoporosis and made no mathematical errors.
- 6Answer-only output formats reduced accuracy for Llama and DeepSeek, but not OpenAI models.
Why It Matters

Source
AuntMinnie
Related News

Real-World Study: Radiology AI Best in Emergency and Inpatient Settings
A commercial AI tool for intracranial aneurysm detection outperformed in inpatient and emergency settings but yielded limited benefits for outpatients in a major health system study.

New Rubric Enhances Safety of AI-Generated Radiology Summaries
Researchers developed a five-factor rubric to assess the safety and quality of AI-generated, patient-friendly radiology report summaries.

Healthcare Leader Warns AI Will Dominate Diagnostic Radiology
A leading oncologist urges future radiologists to specialize in interventional procedures due to AI advances in image interpretation.