Accuracy of Retrieval-Augmented Large Language Model-Generated Preconsult Summaries in Breast Surgical Oncology.
Authors
Affiliations (4)
Affiliations (4)
- Department of Surgery, Breast Service, Memorial Sloan Kettering Cancer Center, New York, NY.
- Department of Epidemiology and Biostatistics, Biostatistics Service, Memorial Sloan Kettering Cancer Center, New York, NY.
- Department of Data and Analytics, Artificial Intelligence & Machine Learning (AIML), Memorial Sloan Kettering Cancer Center, New York, NY.
- Department of Radiology, Memorial Sloan Kettering Cancer Center, New York, NY.
Abstract
Radiology and pathology reports are important in breast surgical oncology planning, but are often unstructured and variable, requiring time-intensive previsit review and preparation. Here we evaluate the accuracy, completeness, and safety of retrieval-augmented large language model (LLM)-generated structured preconsult summaries compared with clinical staff-authored summaries. In this single-institution retrospective validation study, 200 randomly sampled new breast surgical oncology consultations were included. LLM summaries were generated and compared with human summaries documented in consultation notes. Field-level accuracy, omission rate, and fabrication rate among extracted radiology and pathology variables were measured. Error analysis included patient level and clinically significant fabrication rates. Artificial intelligence (AI)-generated summaries demonstrated high field-level accuracy, with fabrication rates ≤2%. Accuracy rates were significantly higher than those observed in staff-authored notes across multiple domains, including nodal status (95% <i>v</i> 72%, <i>P</i> < .001) and documentation of invasive tumor component on biopsy (94% <i>v</i> 30%, <i>P</i> < .001). In contrast, staff-authored summaries were more accurate for receptor status (96% <i>v</i> 86%, <i>P</i> = .002). At the patient level, fabrication occurred in 16 AI-generated summaries (8%), including six clinically significant cases (3%). Staff-authored summaries demonstrated fabrication in six cases (3.0%), with one clinically significant instance (0.5%). In this retrospective validation study, structured retrieval-augmented LLM-generated preconsult summaries demonstrated field-level accuracy comparable with or exceeding staff-authored documentation across most radiologic and pathologic variables, with low per-field fabrication rates. Clinically significant errors occurred in a small proportion of summaries and were largely confined to identifiable high-risk variables that are amenable to targeted verification. With appropriate human oversight and safeguards, LLM-based structured information extraction may support documentation standardization and improved workflow efficiency in breast surgical oncology.