Generative AI versus physicians in diagnostic radiology: a systematic review and meta-analysis.
Authors
Affiliations (8)
Affiliations (8)
- Department of Diagnostic and Interventional Radiology, Graduate School of Medicine, Osaka Metropolitan University, 1-4-3 Asahi-Machi, Abeno-Ku, Osaka, 545-8585, Japan.
- Department of Artificial Intelligence, Graduate School of Medicine, Osaka Metropolitan University, 1-4-3 Asahi-Machi, Abeno-Ku, Osaka, 545-8585, Japan.
- Center for Mathematical and Data Science, Kobe University, 1-1 Rokkodai-Cho, Nada-Ku, Kobe, Hyogo, 657-8501, Japan.
- Smart Data and Knowledge Services Department, German Research Center for Artificial Intelligence (DFKI GmbH), 67663, Kaiserslautern, Germany.
- Center for Digital Transformation of Health Care, Graduate School of Medicine, Kyoto University, Shogoin-Kawahara-Cho 53, Sakyo-Ku, Kyoto, 606-8507, Japan.
- Department of Diagnostic and Interventional Radiology, Graduate School of Medicine, Osaka Metropolitan University, 1-4-3 Asahi-Machi, Abeno-Ku, Osaka, 545-8585, Japan. [email protected].
- Department of Artificial Intelligence, Graduate School of Medicine, Osaka Metropolitan University, 1-4-3 Asahi-Machi, Abeno-Ku, Osaka, 545-8585, Japan. [email protected].
- Center for Health Science Innovation, Osaka Metropolitan University, 1-4-3 Asahi-Machi, Abeno-Ku, Osaka, 545-8585, Japan. [email protected].
Abstract
Generative artificial intelligence (AI) models are increasingly evaluated for diagnostic tasks in radiology, yet accuracy, study designs, endpoints, and comparators vary widely. The purpose was to synthesize diagnostic accuracy of generative AI for radiology and compare performance with physicians. A systematic review and meta-analysis was prospectively registered in PROSPERO (CRD420251040000) and conducted in accordance with PRISMA-DTA guidance. Searches of Medline, Scopus, Web of Science, Cochrane Central, and medRxiv (June 2018-March 2025) identified studies validating generative AI on diagnostic tasks in radiology. Two reviewers independently screened, extracted data, and assessed risk of bias with PROBAST+AI, with disagreements resolved by a third reviewer. Multilevel random-effects meta-regression was performed to compare AI performance with physician performance and identify sources of heterogeneity, with study-clustered robust inference to account for multiple estimates per study. In total, 48 studies met inclusion criteria. Pooled diagnostic accuracy of generative AI was 42.9% (95% CI, 35.8-50.1%) for free-text tasks and 58.1% (95% CI, 50.0-66.2%) for choice tasks. Generative AI overall showed significantly lower accuracy than expert physicians (difference in accuracy [physicians minus AI], +13.0 percentage points [95% CI, 0.9-25.2]; P = .038). Text-only input was associated with higher accuracy than image-only input (difference in accuracy [image-only minus text-only], -25.9 percentage points [95% CI, -39.1 to -12.8]; P = .001) and text-and-image input (- 10.6 percentage points [95% CI, -20.8 to -0.3]; P = .046). Generative AI remained less accurate than expert physicians on diagnostic tasks in radiology. Text-oriented assistive use warrants further evaluation, but the observed accuracy difference may reflect task difficulty and information content rather than the input modality itself. Standardized, transparently reported, adequately powered evaluations are warranted before clinical deployment.