Multicenter AI- versus expert-assisted RECIST target lesion measurements in follow-up body CT of patients with cancer.
Authors
Affiliations (12)
Affiliations (12)
- Department of Medical Imaging, Radboud University Medical Center, Nijmegen, The Netherlands.
- Fraunhofer Institute for Digital Medicine MEVIS, Bremen, Germany.
- Department of Diagnostic Imaging, Oncological Radiotherapy and Hematology, Fondazione Policlinico Universitario Agostino Gemelli IRCCS, Rome, Italy.
- Department of Medicine, Surgery and Dentistry, University of Salerno, Baronissi, Italy.
- Department of Oral Medicine and Radiology, SCB Medical College and Hospital, Cuttack, Odisha, India.
- Department of Radiology and Nuclear Medicine, Rijnstate Ziekenhuis, Arnhem, The Netherlands.
- Department of Advanced Biomedical Sciences, University of Naples "Federico II,"Naples, Italy.
- Department of Diagnostic Imaging, Sungroup International Hospital, Hanoi, Vietnam.
- Department of Neuroradiology, University Medical Center of the Johannes Gutenberg University Mainz, Mainz, Germany.
- Department of Radiology, Charité - Universitätsmedizin Berlin, Berlin, Germany.
- Department of Radiology, University Medical Center Groningen, Groningen, The Netherlands.
- Department of Radiology, Jeroen Bosch Hospital, 's-Hertogenbosch, The Netherlands.
Abstract
Manual lesion measurements remain the standard for assessing oncologic treatment response, despite being time-consuming and prone to substantial interreader variability, which may lead to inconsistent response classification and subsequent variability in treatment decisions. To evaluate the impact of an artificial intelligence (AI) system for assisted lesion measurement on reading time and measurement consistency in follow-up CT examinations using the Response Evaluation Criteria in Solid Tumors (RECIST 1.1). In this retrospective reader study, follow-up chest-abdomen-pelvis CT examinations from 212 oncology patients collected at 2 Dutch hospitals were assessed by 23 readers (15 radiologists, 8 residents) recruited from 11 international institutions under 3 conditions: unassisted, AI-assisted, and expert-assisted (using a prior radiologist's unassisted measurement). To prevent bias related to the source of the measurements, readers were informed that all support was AI-derived. Primary outcomes were reading time to completion and interobserver measurement variability. For each outcome, a Bayesian generalized linear mixed model was used to analyze the results. AI assistance significantly reduced per-patient reading time versus unassisted reading (-36.0 s; 95% CI, -53.0 to -22.1). At the lesion level, it was associated with a small increase in variability relative to the expert-derived reference standard (1.32 mm; 95% CI, 0.83-1.91). At the patient level, AI assistance did not meaningfully affect change in the sum of longest diameters (-0.38 mm; 95% CI, -1.57 to 0.72), but increased RECIST outcome agreement by 7.7% (95% CI, 2.8-12.7) compared with unassisted reading. Expert-assisted reading yielded even higher interreader agreement (13.3%; 95% CI, 8.6-18.1). AI assistance reduced reading time and improved patient-level RECIST agreement, despite a small increase in lesion-level measurement variability. These findings suggest that AI-assisted RECIST assessment may improve workflow and response classification consistency, while also providing a benchmark from expert-assisted reading for future AI development.