Performance of ChatGPT (GPT-4o) in automated Lung-Reporting and Data Systems (RADS) classification: a category-stratified evaluation of prompting strategies.
Authors
Affiliations (2)
Affiliations (2)
- Department of Radiology, Trinity Health Mid-Atlantic - Mercy Catholic Medical Center, Darby, PA, USA. Electronic address: [email protected].
- Department of Radiology, Trinity Health Mid-Atlantic - Mercy Catholic Medical Center, Darby, PA, USA.
Abstract
The aim of this study is to explore the capabilities of large language models (LLMs), such as ChatGPT, in assigning Lung-Reporting and Data Systems (RADS) scores when given radiology report findings, and to explore prompt engineering practices for efficient and effective integration of artificial intelligence (AI) into radiology workflows. A total of 327 de-identified low-dose lung screening computed tomography (CT) reports were evaluated using GPT- 4o, with instructions to assign Lung-RADS scores based on the reports provided. Structured and unstructured zero-shot prompting strategies were tested. Accuracy, overcalls, undercalls, and weighted Cohen's kappa (linear weights) were computed for each prompt type compared with the rater. McNemar's test was used to compare accuracy between structured and unstructured prompts. Agreement between GPT outputs and the rater's reference standard was higher for the structured prompt compared with the unstructured prompt. For the structured prompt, overall accuracy was 68.8% with a weighted Cohen's kappa of 0.758, while the unstructured prompt achieved an accuracy of 58.1% and a weighted Cohen's kappa of 0.695. Structured prompting produced fewer overcalls and more undercalls compared with unstructured prompting. LLMs demonstrate promise in automatically assigning scores to screening imaging studies when provided with findings and can potentially be incorporated into reporting workflows to improve reporting efficiency.