Artificial Intelligence-enhanced cardiac MRI reporting: expert validation and patient-centered outcomes.
Authors
Affiliations (9)
Affiliations (9)
- Cardiology III Department, EP laboratory, De Gasperis Cardio Center, Great Metropolitan Hospital Niguarda, Piazza dell'Ospedale Maggiore, 3, Milan, 20162, Italy.
- Cardiology III Department, EP laboratory, De Gasperis Cardio Center, Great Metropolitan Hospital Niguarda, Piazza dell'Ospedale Maggiore, 3, Milan, 20162, Italy. [email protected].
- Department of Health Sciences, School of Medicine and Surgery, University of Milano-Bicocca, Milan, Italy. [email protected].
- Department of Biomedical, Surgical and Dental Sciences, University of Milan, Milan, Italy.
- Department of Health Sciences, School of Medicine and Surgery, University of Milano-Bicocca, Milan, Italy.
- Advanced Cardiovascular Imaging and Sports Cardiology Unit, Clinica Villa dei Fiori Acerra, Naples, Italy.
- Cardiology Department, Istituto Clinico Città Studi, Milan, Italy.
- Cardiology IV Department, De Gasperis Cardio Center, Great Metropolitan Hospital Niguarda, Milan, Italy.
- Department of Perioperative Cardiology and Cardiovascular Imaging, Centro Cardiologico Monzino IRCCS, Milan, Italy.
Abstract
As patients increasingly access their own electronic health records, the dense technical language of cardiac magnetic resonance (CMR) reports has become a barrier both to patient comprehension and to decision-making by non-imaging physicians. We evaluated whether ChatGPT-4o can enhance the accessibility of CMR reports and generate clinical recommendations, and we quantified the accuracy, safety, and patient reception of these outputs. We prospectively enrolled 75 consecutive outpatients undergoing CMR at two Italian tertiary centres. Each physician-dictated report was processed with ChatGPT-4o through the web interface to produce a simplified patient-facing explanation and tailored clinical recommendations. Two expert cardiologists rated the correctness and completeness of the simplified reports on 5-point Likert scales, with inter-rater reliability by ICC(2,1). Three additional cardiologists rated the AI-generated recommendations for correctness, completeness, and potential harm after a calibration session using a shared written rubric. Patients completed paired questionnaires comparing the standard and AI-enhanced reports across six domains, analysed with the Wilcoxon signed-rank test. Expert-rated correctness of the simplified reports was 4.63 ± 0.88 and completeness 4.42 ± 1.06, with excellent agreement (ICC 0.95-0.99). Patients rated AI-enhanced reports significantly higher than standard reports across every domain (all p < 0.001), including overall satisfaction (8.87 ± 1.11 vs. 6.56 ± 2.22 on a 10-point scale; +35%). AI-generated recommendations showed moderate correctness (patient-directed 3.73 ± 0.74; physician-directed 3.52 ± 0.74) and low mean harm (2.10 ± 0.67 and 2.16 ± 0.68). High-risk recommendations were infrequent but not negligible, affecting 1 patient (1.3%) for patient-directed and 5 patients (6.7%) for physician-directed content, and were concentrated in clinically complex cases. The AI-simplified report was comparable in length to the original (306 ± 112 vs. 284 ± 62 words) but was accompanied by additional tailored recommendations. ChatGPT-4o improved the accessibility of CMR reports and patients' perceived comprehension and satisfaction while preserving expert-validated report accuracy. Given a small but clinically relevant fraction of high-risk recommendations, large-language-model outputs should be deployed as adjunctive, physician-supervised decision support rather than autonomously.