Back to all papers

Periodontitis Risk Assessment and Prevention Planning: Comparative Study of Multimodal Large Language Models and Periodontist Evaluations.

October 2, 2026pubmed logopapers

Authors

Swinckels L,Alves Rabelo K,Delamare EL,Loos BG,Lahoud P,Bijwaard H,de Keijzer A,Kim J,Bruers J

Affiliations (11)

  • Department of Oral Public Health, Academic Center for Dentistry Amsterdam, Gustav Mahlerplein 3004, Amsterdam, North Holland, 1182 DB, The Netherlands, 31 020 59 80888.
  • Cluster of Health, Sport and Welfare, Inholland University of Applied Sciences, Amsterdam, North Holland, The Netherlands.
  • Sydney Dental School, Faculty of Medicine and Health, The University of Sydney, Sydney, New South Wales, Australia.
  • Department of Periodontology, Academic Center for Dentistry Amsterdam, Amsterdam, North Holland, The Netherlands.
  • Division of Periodontology & Oral Microbiology, Department of Oral Health Sciences, University Hospitals Leuven, KU Leuven, Leuven, Flanders, Belgium.
  • OMFS-IMPATH Research Group, Department of Imaging and Pathology, Universitair Ziekenhuis Leuven, Leuven, Flanders, Belgium.
  • Department of Conservative Dentistry, Periodontology and Digital Dentistry, LMU University Hospital, Ludwig-Maximilians-Universität München, Munich, Bavaria, Germany.
  • National Institute for Public Health and the Environment, Bilthoven, Utrecht, The Netherlands.
  • Applied Responsible Artificial Intelligence Research Group, Avans University of Applied Sciences, Breda, North Brabant, The Netherlands.
  • School of Computer Science, Faculty of Engineering, The University of Sydney, Sydney, New South Wales, Australia.
  • Royal Dutch Society for the Promotion of Dentistry, Utrecht, Utrecht, The Netherlands.

Abstract

Periodontitis is one of the most prevalent yet preventable oral diseases, as indicated by multiple clinical and radiographic factors. As these factors are recorded in electronic health records (EHRs), their reuse offers opportunities for personalized risk assessment and targeted prevention. Predictive AI and traditional machine learning models support fragmented detection tasks but lack the integration of textual and imaging predictors. Emerging multimodal large language models (M-LLMs) show promise in combining these data sources for clinical assessment. Evaluating the capabilities of M-LLMs and comparing them against the current clinical standard are therefore essential to determine their potential as digital assistants. This study aimed to evaluate the ability of M-LLMs to assess periodontitis risk and suggest prevention strategies, based on EHR data and radiographic findings. Each M-LLM was individually evaluated by periodontal experts, benchmarked against other models, and compared with a periodontist as a reference. A vignette study was conducted following TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) guidelines for the evaluation of LLMs. Ten periodontal vignettes were created, each including a panoramic radiograph and textual EHR data. Three LLMs capable of reasoning and handling multimodal data were compared to a periodontist who generated outputs manually, based on the same prompts and input data. Periodontal experts rated all outputs across 6 predefined criteria on a 5-point Likert scale. Statistical analyses evaluated overall performance per model and tested whether performance varied per model, scenario complexity, or rater. GPT o1 Pro and Claude Sonnet 4 showed strong performance, with 86.7% and 85.6% of ratings deemed acceptable-comparable to the periodontist's output (87.8%). Gemini 2.5 Pro was rated significantly lower than both the periodontist and the other models (59.4% acceptable; <i>P</i><.002). Radiographic interpretation consistently received lower scores than other abilities across all models and the periodontist, with Gemini rated below the acceptable threshold. The time required for completion ranged from approximately 10 seconds for Claude to 37 seconds for Gemini; 3 minutes, 22 seconds, for GPT; and 5 minutes, 57 seconds, for the periodontist. M-LLMs demonstrated strong reasoning abilities in periodontal assessment. Across all models, unacceptable elements were consistently related to errors in radiographic interpretation, though refined prompting or newer model versions may improve this. Notably, even when radiographic findings were incorrect and plaque-retentive factors were absent, outputs were still rated well, indicating that EHR data alone provide a substantial basis. For clinical applicability, M-LLMs must at least perform comparably to a periodontist and meet the quality standards set by periodontal experts-a bar that GPT and Claude appear to approach.

Topics

Large Language ModelsPeriodontitisJournal ArticleComparative Study

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.