Back to all papers

A Quality Assessment Rubric for Artificial Intelligence-Generated Patient-Friendly Radiology Reports.

September 9, 2026pubmed logopapers

Authors

Armstrong BA,Koirala A,Na HS,Johnston A,Chang M,Lyu D,Cheuy L,Fang Z,Chaudhari A,Larson DB

Affiliations (4)

  • Department of Radiology, Stanford University, Palo Alto, California 94305, United States.
  • Department of Statistics, Stanford University, Palo Alto, California 94305, United States.
  • Department of Biomedical Data Science, Stanford University, Palo Alto, California 94305, United States.
  • Center for Artificial Intelligence in Medicine and Imaging, Stanford University, Palo Alto, California 94305, United States.

Abstract

<b>Background:</b> Artificial intelligence (AI) tools are being used to translate radiology reports into plain language, but translation errors may compromise comprehension and safety. <b>Objective:</b> To develop and evaluate a rubric for assessing the quality and safety of AI-generated patient-friendly radiology reports. <b>Methods:</b> In this prospective study (conducted from February 2025 to December 2025), survey-workshop cycles, involving lay participants and a multidisciplinary panel, were used to develop a rubric for grading AI-generated patient-friendly report quality across core attributes and determining whether such reports are safe for patient distribution. ChatGPT-4.1 and Claude-4.0 were used to generate patient-friendly reports of varying quality based on radiology report impressions from a public dataset and prespecified quality targets across attributes. Research-team members, additional lay participants and radiologists, and ChatGPT-5 evaluated these patient-friendly reports using the rubric. <b>Results:</b> Development included 19 participants (39±7 years; 11 women, 8 men); evaluation included six research-team members and 111 additional participants (47±18 years; 43 women, 68 men). The final rubric included five core attributes-clarity, content, certainty, tone, verbosity-each graded on a 3-point scale; by the rubric's decision rule, patient-friendly reports assessed as grade 1 (unsafe or unacceptable) in any attribute other than verbosity are unsafe for distribution and warrant withholding. Lay and radiologist research-team members (n=3 participants each; 60 reports) had almost-perfect intergroup agreement (α=0.87) for overall grade assignments. Additional lay (n=19) and radiologist (n=12) participants, each evaluating six reports, had moderate (α=0.51) and substantial (α=0.65) interreader agreement, respectively, for overall grade assignments and 91.2% and 95.8% agreement, respectively, between subjective and rubric rule-based distribution decisions. In wider field testing, 80 lay participants (480 reports) had moderate agreement (κ=0.43) with prespecified reference-standard grades; subjective distribution decisions had 73.5% agreement with rule-based distribution decisions. Across 480 reports, AI had moderate agreement (κ=0.44) with prespecified reference-standard grades; rule-based distribution decisions using AI-assigned grades had 88.1% agreement with rule-based distribution decisions using reference-standard grades. <b>Conclusions:</b> The rubric may provide a standardized safeguard before release of AI-generated patient-friendly reports. <b>Clinical Impact:</b> Although requiring further training and validation, AI rubric application could enable scalable quality assurance and safer clinical integration of AI-generated communications.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.