Back to all papers

Accuracy of Artificial Intelligence based chatbots in reporting jaw lesions from multimodal radiographic images: A cross-sectional study.

August 4, 2026pubmed logopapers

Authors

Elkersh NM,Atteya AM,Shararah AA

Affiliations (3)

  • Lecturer, Department of Oral medicine, Periodontology, Oral diagnosis and Oral Radiology, Faculty of Dentistry, Alexandria University, Alexandria, Egypt.
  • Associate professor, Department of Maxillofacial and Plastic Surgery, Faculty of Dentistry, Alexandria University, Alexandria, Egypt.
  • Lecturer, Department of Oral & Maxillofacial surgery, Faculty of Dentistry, AAST, Egypt.

Abstract

The current study aimed to quantify the diagnostic accuracy of commonly utilized chatbots including Gemini, Copilot, Claude, and specialized architectures like Manus in the detection and differential diagnosis of various jaw lesions, while concurrently evaluating the clinical safety and fidelity of the information they provide. Cone beam computed tomography (CBCT) dataset from 97 patients presented with jaw lesions were collected and anonymized. Panoramic 2D views were reconstructed from Digital Imaging and Communication in Medicine (DICOM) of all cases using Bluesky Plan software and provided to 4 chatbots (Gemini 2.5 Pro, Copilot, Claude and Manus). Moreover, the DICOM data was provided to Manus followed by prompting. The reports generated were evaluated for accuracy, relevance and feasibility. Statistically significant differences were detected between the chatbots in all measured parameters. In all evaluated parameters Manus CBCT showed the most accurate results (95% of lesions were detected and correctly diagnosed,). Gemini 2.5 pro ranked second where 80% of lesions were detected and 56% were correctly diagnosed. Manus Pan showed less accurate results. The least accurate results were detected in Copilot and Claude. Significant discrepancies exist among artificial intelligence (AI) chatbots regarding their diagnostic accuracy in reporting jaw lesions. Notably, the integration of raw 3-dimensional CBCT data substantially optimizes chatbot performance in lesion detection and diagnosis, as demonstrated by Manus architecture.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.