Back to all papers

How Robust Are Radiology Teaching-Case Benchmarks for Multimodal Large Language Models? An Audit of Exposure Opportunity, Inference Configuration, and Text-Only Solvability Across Eighteen Models.

September 22, 2026pubmed logopapers

Authors

Fan W,Chen X,Lin Y,Song X,Xu Y,Xie S

Affiliations (2)

  • Department of Radiology, China-Japan Friendship Hospital, No. 2 Yinghuayuan East Street, Chaoyang District, Beijing, 100029, China.
  • Department of Radiology, China-Japan Friendship Hospital, No. 2 Yinghuayuan East Street, Chaoyang District, Beijing, 100029, China. [email protected].

Abstract

Multimodal large language models (MLLMs) are benchmarked on radiology multiple-choice questions. We audited how training-data exposure opportunity, image resolution, inference configuration, and text-only solvability affected interpretation of a teaching-case benchmark. Eighteen MLLMs and 2 original reference readers answered 260 questions from 51 AuntMinnie cases; 5 additional readers completed them under a unified interface. Analyses included Holm-adjusted McNemar tests, case-cluster bootstrapping, model-cutoff matching, a case-random-intercept logistic model, cross-cohort configuration re-runs, and text-only ablation. Model accuracy ranged from 61.5% to 80.8%, and the 7 readers ranged from 66.5% to 84.2%. No model exceeded all readers; 17 fell within the reader range and 1 below. Eight models were significantly less accurate than Reader 1 after reader-specific Holm adjustment (6 when all 36 comparisons were adjusted together); none differed significantly from Reader 2. A cohort/publication-period association remained after item adjustment (odds ratio, 4.67; 95% Laplace interval, 1.82-11.98), but 3 within-cohort cutoff comparisons showed no exposure-consistent advantage. Resolution changes were within 3 percentage points in both cohorts, and the configuration-by-version interaction was not significant. Among 5 post hoc selected models, mean image contribution was 6.2 percentage points in the original cohort and 4.0 in the supplementary cohort. The benchmark did not support a single human-model ranking, a training-exposure explanation for the cohort difference, or a configuration-dependent GPT version effect. Public teaching-case benchmarks should report provenance, individual-reader calibration, multiplicity control, inference settings, and image-free performance and should not be interpreted as evidence of clinical diagnostic readiness.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.