Classification Performance of General-Purpose Multimodal Large Language Models Across Orthodontic Radiographic Tasks: A Comparative Study of ChatGPT, Gemini, and Claude.
Authors
Affiliations (2)
Affiliations (2)
- Department of Orthodontics, School of Dentistry, Ankara University, 06560 Ankara, Türkiye.
- Independent Researcher, 06560 Ankara, Türkiye.
Abstract
<i>Background and Objectives</i>: General-purpose multimodal large language models (MLLMs) can interpret radiographic images, but their classification performance across orthodontic tasks remains uncertain. This study compared the classification performance of ChatGPT, Gemini, and Claude on lateral cephalometric, hand-wrist, and panoramic radiographs. <i>Materials and Methods</i>: This retrospective diagnostic accuracy study included 250 individuals, each contributing one lateral cephalometric, hand-wrist, and panoramic pretreatment radiograph (750 total). Reference classifications were established by two experienced orthodontists, with disagreements resolved by consensus. Lateral cephalometric radiographs were classified as skeletal Class I, II, or III based on the ANB angle according to Steiner analysis; hand-wrist radiographs as prepubertal, pubertal, or postpubertal; and panoramic radiographs as early mixed, late mixed, or permanent dentition. Each image was evaluated once by each AI platform using identical Turkish prompts in separate chat sessions. Classification accuracy, balanced accuracy, macro-F1, class-specific metrics, and reference agreement were assessed. Generalized estimating equations (GEE) assessed platform, radiograph type, and interaction effects on correct classification. <i>Results</i>: ChatGPT had the highest hand-wrist accuracy (81.6%; 95% CI, 76.3-85.9), whereas Gemini had the highest panoramic accuracy (92.8%; 95% CI, 88.9-95.4). Lateral cephalometric accuracies were 70.0%, 64.4%, and 64.8% for ChatGPT, Gemini, and Claude, respectively, with no significant interplatform difference (<i>p</i> = 0.336). The platform × radiograph type interaction was significant (Wald χ<sup>2</sup> = 42.52; df = 4; <i>p</i> < 0.001). Agreement with the reference standard was highest for ChatGPT on hand-wrist radiographs (κw = 0.758) and Gemini on panoramic radiographs (κw = 0.883). <i>Conclusions</i>: Classification performance was task- and platform-dependent, with no model consistently achieving the highest performance. For the predefined classification tasks, these models should not be used as standalone tools for these classification tasks. Their potential as decision-support tools requires prospective evaluation of AI-assisted clinician performance and external validation.