Information matters, but cues are dangerous: a comparative evaluation of three multimodal AI chatbots in oral and maxillofacial radiology.
Authors
Affiliations (2)
Affiliations (2)
- Department of Oral and Maxillofacial Radiology, School of Dentistry, Institute of Oral Bioscience, Jeonbuk National University, Jeonju, 54896, Republic of Korea.
- Department of Oral and Maxillofacial Radiology, Jeonbuk National University Dental Hospital, Jeonju, 54907, Republic of Korea.
Abstract
To compare the diagnostic performance of three multimodal AI chatbots on oral and maxillofacial radiographic images and to examine how additional information, delivery mode, and user-suggested diagnoses affect accuracy. Three AI chatbots (GPT-5.1, Gemini 3 Flash, Claude Opus 4.7) were tested on 90 cases comprising normal controls, osteomyelitis, and benign jaw lesions. Inputs were combinations of a panoramic image, a cropped panoramic image, an axial CBCT image or a text-based cue. Inputs were delivered all at once or sequentially. Accuracy was scored at category and specific-diagnosis levels using non-parametric tests with false-discovery-rate correction. With the panoramic image alone, accuracy for diseased cases was low (0-60%) but rose to as high as 38-92% in each model's best condition with added information. A cropped image was the most consistently beneficial additional visual input, whereas an axial CBCT image provided less improvement. GPT-5.1 recognised normal cases well but missed lesions, Gemini 3 was sensitive but less specific, and Claude 4.7 defaulted to benign diagnoses. Correct verbal cues increased accuracy, whereas a misleading cue caused decline. Gemini 3 accepted a false benign suggestion in 96% of cases it had initially classified as normal. For benign lesions, specific-diagnosis accuracy was almost half of category-level accuracy. Accuracy was strongly affected by model choice, information type and user-suggested diagnoses, but no model reached a level that would be acceptable for clinical use. This multi-model comparison isolates the effects of information type, delivery mode, and sycophancy in oral and maxillofacial radiology.