New multimodal large language models (LLMs) like OpenAI o3 and Gemini 2.5 Pro demonstrated significant advancements in answering Japanese radiology board exam questions, particularly with image input.
Key Details
- 1Eight LLMs were tested on the Japan Diagnostic Radiology Board Examination (JDRBE).
- 2OpenAI o3 achieved 67% accuracy (text-only) and 72% with image input.
- 3Gemini 2.5 Pro also showed notable accuracy improvements with image data.
- 4Both OpenAI o3 and Gemini 2.5 Pro received higher legitimacy scores from radiologist raters than some competitors.
- 5The test set included 233 questions and 477 images (184 CT, 159 MRI, 15 x-ray, 90 nuclear medicine).
- 6Image input statistically improved diagnostic accuracy for several models.
Why It Matters

Source
AuntMinnie
Related News

Real-World Study: Radiology AI Best in Emergency and Inpatient Settings
A commercial AI tool for intracranial aneurysm detection outperformed in inpatient and emergency settings but yielded limited benefits for outpatients in a major health system study.

New Rubric Enhances Safety of AI-Generated Radiology Summaries
Researchers developed a five-factor rubric to assess the safety and quality of AI-generated, patient-friendly radiology report summaries.

Healthcare Leader Warns AI Will Dominate Diagnostic Radiology
A leading oncologist urges future radiologists to specialize in interventional procedures due to AI advances in image interpretation.