Back to all papers

The Role of Google Gemini in the Interpretation of Ovarian Lesions.

August 13, 2026pubmed logopapers

Authors

Psilopatis I,Zwimpfer TA,Heinzelmann-Schwarz V,Manegold-Brauer G,Reina H

Affiliations (2)

  • University Hospital Basel, Department of Gynecology and Obstetrics, Switzerland, Basel.
  • University Hospital and University of Basel, Department of Biomedicine, Switzerland, Basel.

Abstract

Multimodal large language models (LLMs) can process text and images to support diagnostic processes, yet their ability to apply standardized ovarian ultrasound classification systems remains unexplored. This study evaluates Google Gemini's performance with regard to interpreting ovarian ultrasound using the Ovarian-Adnexal Reporting and Data System (O-RADS) and International Ovarian Tumor Analysis (IOTA) criteria (IOTA simple rules) compared to expert evaluation. This retrospective study included 30 ultrasound examinations from the University Hospital of Basel: 10 normal ovaries, 10 benign lesions, and 10 malignant lesions. Each examination included B-mode and Doppler images. Google Gemini was prompted to assess normal ovarian images for pathology and apply the O-RADS categories and IOTA simple rules to pathologic cases. Two blinded experts independently evaluated responses using the Global Quality Score (GQS, 1-5 scale). The reference standard was histopathology for pathologic cases and expert assessment for normal ovaries. Google Gemini achieved a mean GQS of 3 across all categories. The model mischaracterized 6/10 normal ovaries as pathologic. For benign lesions, 6/10 were incorrectly assigned to high-risk O-RADS categories (4-5) with frequent misidentification of malignant IOTA features. While 6/10 malignant cases were correctly identified (GQS 4-5), 3/10 were misclassified as benign (O-RADS 2, GQS 1). The remaining case of a mucinous borderline tumor was also misclassified by the LLM (O-RADS 3, GQS 1). Google Gemini seems to not be ready for independent clinical application in ovarian ultrasound interpretation. Critical inconsistencies including overdiagnosis and dangerous misclassification of malignancies indicate the need for specialized, domain-specific models with rigorous validation before clinical implementation.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.