Large Language Models for Preoperative Microvascular Invasion Prediction in Hepatocellular Carcinoma: A Multicenter Comparison with Radiologists and Treatment Outcomes.
Authors
Affiliations (3)
Affiliations (3)
- Department of Radiology, 7T Magnetic Resonance Translational Medicine Research Center, Southwest Hospital, Army Medical University (Third Military Medical University), Chongqing, 400038, People's Republic of China.
- College of Mathematics and Statistics, Chongqing University, Chongqing, 400044, People's Republic of China.
- Institute of Pathology and Southwest Cancer Center, Southwest Hospital, Army Medical University (Third Military Medical University), Chongqing, People's Republic of China.
Abstract
Preoperative prediction of microvascular invasion (MVI) in hepatocellular carcinoma (HCC) is critical for prognosis but challenging. This study evaluated the performance of large language models (LLMs) for MVI prediction compared with radiologists with varying levels of experience and explored the association of model-predicted MVI with recurrence and treatment outcomes. In this retrospective multicenter study, 602 patients with pathologically confirmed HCC who underwent preoperative MRI and surgical resection were included. GPT-4o and DeepSeek-R1, prompted in English and Chinese using radiology report narratives and predefined MRI features, were compared with six radiologists using histopathology as the reference standard. An independent cohort of 135 patients who underwent radiofrequency ablation (RFA) was included for treatment-outcome analyses. Prediction performance was assessed using diagnostic accuracy and the area under the curve (AUC). GPT-4o (English input) achieved 75.6% accuracy, outperforming residents (59.3%, P < 0.001) and attendings (66.7%, P < 0.001), and was comparable to senior radiologists (74.2%, P = 1.000). DeepSeek-R1 (Chinese input) achieved the highest accuracy (81.2%), outperforming all radiologists (P < 0.05). Residents and attendings showed high sensitivity but relatively low specificity, whereas senior radiologists demonstrated a more balanced performance; while DeepSeek-R1 showed better specificity (up to 94.2%, P < 0.001). LLMs performance were relatively stable across tumor sizes, whereas accuracy among less-experienced radiologists improved with tumor size. AUCs for LLMs (GPT-4o: 75.4%; DeepSeek-R1 Chinese: 80.1%) were superior or comparable to those of radiologists (61.8%-74.5%). Stacking ensemble-predicted MVI-positive status was associated with shorter time-to-recurrence (TTR) in both small (≤3 cm, P = 0.039) and large (>3 cm, P = 0.001) tumors, and with better outcomes following anatomical rather than non-anatomical resection (≤3 cm, P = 0.021; >3 cm, P < 0.001). After propensity score matching, surgical resection was associated with longer TTR than RFA in the model-predicted MVI-positive subgroup, but not in the MVI-negative subgroup. LLMs, particularly DeepSeek-R1 and GPT-4o, demonstrated diagnostic performance for preoperative MVI prediction comparable to or better than that of senior radiologists, with relatively stable performance across tumor sizes and potential value for recurrence-risk stratification and treatment planning.