Back to all papers

Comparative Performance of GPT-5.1, Claude Sonnet 4.5, Gemini 2.0 Flash, and DeepSeek-R1 on Oral and Maxillofacial Radiology Questions From the Turkish Dental Specialisation Examination.

August 19, 2026pubmed logopapers

Authors

Keser G,Celen B,Namdar Peki̇ner F

Affiliations (2)

  • Department of Oral Medicine and Radiology, Marmara University, Istanbul, TUR.
  • Department of Oral and Maxillofacial Radiology, Marmara University, Istanbul, TUR.

Abstract

Rapid advances in artificial intelligence (AI) and large language models (LLMs) have increased interest in their use in dental education and assessment. This study evaluated the accuracy of GPT-5.1, Claude Sonnet 4.5, Gemini 2.0 Flash, and DeepSeek-R1 on Turkish Dental Specialisation Exam (DUS) questions in oral and maxillofacial radiology. A total of 132 Turkish DUS questions from 2012 to 2021 were obtained from a publicly accessible question bank, including 126 theoretical and six image-based items. Each question had five options and one correct answer. Questions covered radiologic physics, radiation safety, dental anatomy, pathology classification, imaging instrumentation, and visual interpretation of panoramic, periapical, cone beam computed tomography (CBCT), and clinical images. All models received identical prompts, and image-based items were submitted through the image-upload functionality available in the respective web interfaces. Responses were scored against the official answer key by two oral and maxillofacial radiologists. Statistical analyses included Fisher's exact test, Cochran's Q test, and Bonferroni-corrected McNemar tests. An exploratory secondary analysis evaluated session-to-session variation using a model-based performance index. GPT-5.1 answered all 132 questions correctly (100%), compared with Claude Sonnet 4.5 (119/132; 90.2%), Gemini 2.0 Flash (111/132; 84.1%), and DeepSeek-R1 (108/132; 81.8%). Overall performance differed significantly among the models (p < 0.001), and GPT-5.1 significantly outperformed each of the other three models after Bonferroni correction. On the six image-based questions, accuracy was 100% for GPT-5.1, 83.3% for Claude Sonnet 4.5, 50.0% for Gemini 2.0 Flash, and 66.7% for DeepSeek-R1. Because only six image-based items were available, these subgroup findings were considered exploratory and do not permit firm conclusions regarding radiographic interpretation. No significant temporal trend was identified in the exploratory model-based performance index. GPT-5.1 achieved the highest accuracy on this publicly accessible DUS benchmark and significantly outperformed Claude Sonnet 4.5, Gemini 2.0 Flash, and DeepSeek-R1. However, the public availability of the questions and answer keys introduces a risk of benchmark contamination, and the very small image-based subgroup precludes generalisation regarding radiographic interpretation. These findings should therefore be interpreted as comparative performance on a public examination benchmark rather than evidence of clinical diagnostic competence.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.