Zero-shot segmentation and tumor diameter measurement of soft tissue sarcomas on MRI using a multimodal large language model.
Authors
Affiliations (5)
Affiliations (5)
- Department of Medical Imaging, The Ottawa Hospital, University of Ottawa, 501 Smyth Road, Ottawa, Ontario, K1H 8L6, Canada. [email protected].
- Department of Radiology, Institute of Medical Science, The University of Tokyo, Tokyo, Japan.
- Department of Diagnostic Radiology, McGill University, Montreal, QC, Canada.
- Augmented Intelligence and Precision Health Laboratory (AIPHL), Research Institute of the McGill University Health Centre, Montreal, Canada.
- Diagnostic Radiology and Radiation Oncology, Chiba University Graduate School of Medicine, Chiba, Japan.
Abstract
To evaluate the zero-shot capabilities of a multimodal deep learning model (Gemini 3.0) in segmenting Soft Tissue Sarcomas (STS) on MRI and to compare the accuracy of tumor diameter measurements derived from automated segmentation versus direct model estimation. This retrospective study utilized a public dataset (The Cancer Imaging Archive) consisting of 51 patients (mean age, 54.8 ± 17.0 years; 24 men, 27 women) with histologically confirmed STS. Gemini 3.0 performed zero-shot lesion segmentation and direct diameter estimation on a fat-suppressed T2-weighted or STIR image. Segmentation accuracy was evaluated using the Dice similarity coefficient (DSC) against manual segmentation by a board-certified radiologist. Tumor diameters derived from Gemini segmentation masks and direct estimates were compared with manual measurements using intraclass correlation coefficients (ICC) and Bland-Altman analysis. Bootstrap analysis was performed to compare the performance of the two automated measurement methods. Gemini generated segmentation masks with a mean DSC of 0.935 ± 0.082, comparable to inter-radiologist agreement (0.949 ± 0.037) and showing high reproducibility (DSC, 0.981 ± 0.038). For tumor size assessment, the segmentation-derived measurement yielded an ICC of 0.976, significantly higher than the direct estimation ICC of 0.922 (p = 0.027). Bland-Altman analysis revealed a smaller mean bias for the segmentation-based approach (-0.5 mm) compared with direct estimation (-5.0 mm). Gemini 3.0 achieved high accuracy in zero-shot segmentation of STS on MRI. Deriving quantitative measurements from automated segmentation proved significantly more reliable than direct size estimation by the large language model.