Reporting Quality of Large Language Model Studies: A Cross-Sectional Audit of High-Ranking Radiology and Medical Imaging Journals.
Authors
Affiliations (3)
Affiliations (3)
- Department of Radiology, Üsküdar State Hospital, Istanbul, Türkiye.
- Department of Radiology, Başakşehir Çam and Sakura City Hospital, Istanbul, Türkiye.
- Department of Radiology, Başakşehir Çam and Sakura City Hospital, Istanbul, Türkiye. [email protected].
Abstract
To evaluate adherence to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) in radiology and medical imaging studies involving large language models (LLMs). We conducted a cross-sectional audit of original LLM research studies published between January 1 and December 26, 2025, in Q1 journals within the Web of Science "Radiology, Nuclear Medicine, and Medical Imaging" category. PubMed and Scopus were searched to identify eligible studies. A quota-based subsampling strategy, based on journal publication volume, was used to select approximately 100 studies. All four eligible articles from the <i>Korean Journal of Radiology</i> (<i>KJR</i>) were additionally included as a benchmark. Adherence to the 2025 update of MI-CLEAR-LLM was scored through a two-round, consensus-based process: an initial assessment by one reviewer followed by a critical re-evaluation by secondary reviewers, with consensus adjudication by an additional reviewer when needed. Between-journal differences were analyzed with the Kruskal-Wallis test, followed by Dunn post hoc pairwise comparisons with Holm-adjusted <i>P</i>-values. Of 201 eligible studies identified, 102 were finally analyzed after applying the subsampling strategy. Overall adherence to MI-CLEAR-LLM was moderate (mean, 51.2% ± 14.7%; range, 22.2%-84.2%). Adherence was highest for input data type (100%), test-data independence (80.2%), and adaptation strategy (78.1%), and lowest for prompt execution setup (29.4%) and stochasticity management (33.1%). The least frequently reported items were training-data cutoff date (9.8%) and rationale for prompt wording (15.6%). Adherence varied significantly across journals (<i>P</i> = 0.011), with <i>KJR</i> showing the highest mean adherence (72.8% ± 2.7%). Reporting transparency in radiology and medical imaging LLM studies published in 2025 was inconsistent across reporting items and journals, with substantial deficiencies in some reproducibility-critical elements. Broader adoption of reporting standards is essential to improve the reproducibility and interpretability of future accuracy evaluations.