Back to all papers

Reporting Quality of Large Language Model Studies: A Cross-Sectional Audit of High-Ranking Radiology and Medical Imaging Journals.

August 4, 2026pubmed logopapers

Authors

Mese I,Gunes ST,Coskun O,Keles A,Banaz T,Kazanbas MF,Kocak B

Affiliations (3)

  • Department of Radiology, Üsküdar State Hospital, Istanbul, Türkiye.
  • Department of Radiology, Başakşehir Çam and Sakura City Hospital, Istanbul, Türkiye.
  • Department of Radiology, Başakşehir Çam and Sakura City Hospital, Istanbul, Türkiye. [email protected].

Abstract

To evaluate adherence to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) in radiology and medical imaging studies involving large language models (LLMs). We conducted a cross-sectional audit of original LLM research studies published between January 1 and December 26, 2025, in Q1 journals within the Web of Science "Radiology, Nuclear Medicine, and Medical Imaging" category. PubMed and Scopus were searched to identify eligible studies. A quota-based subsampling strategy, based on journal publication volume, was used to select approximately 100 studies. All four eligible articles from the <i>Korean Journal of Radiology</i> (<i>KJR</i>) were additionally included as a benchmark. Adherence to the 2025 update of MI-CLEAR-LLM was scored through a two-round, consensus-based process: an initial assessment by one reviewer followed by a critical re-evaluation by secondary reviewers, with consensus adjudication by an additional reviewer when needed. Between-journal differences were analyzed with the Kruskal-Wallis test, followed by Dunn post hoc pairwise comparisons with Holm-adjusted <i>P</i>-values. Of 201 eligible studies identified, 102 were finally analyzed after applying the subsampling strategy. Overall adherence to MI-CLEAR-LLM was moderate (mean, 51.2% ± 14.7%; range, 22.2%-84.2%). Adherence was highest for input data type (100%), test-data independence (80.2%), and adaptation strategy (78.1%), and lowest for prompt execution setup (29.4%) and stochasticity management (33.1%). The least frequently reported items were training-data cutoff date (9.8%) and rationale for prompt wording (15.6%). Adherence varied significantly across journals (<i>P</i> = 0.011), with <i>KJR</i> showing the highest mean adherence (72.8% ± 2.7%). Reporting transparency in radiology and medical imaging LLM studies published in 2025 was inconsistent across reporting items and journals, with substantial deficiencies in some reproducibility-critical elements. Broader adoption of reporting standards is essential to improve the reproducibility and interpretability of future accuracy evaluations.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.