Back to all papers

Reproducibility of Mandibular Cortical Index Classification Among Dental Examiners and a Single Generative AI Platform: An Observer Agreement Study.

July 31, 2026pubmed logopapers

Authors

Seki K,Kashima M,Akiyama T,Kobayashi A,Dezawa K,Takeuchi Y,Furuchi M,Kamimoto A

Affiliations (2)

  • Department of Comprehensive Dentistry and Clinical Education, Nihon University School of Dentistry, Tokyo 101-8310, Japan.
  • Department of Oral and Maxillofacial Radiology, Nihon University School of Dentistry, Tokyo 101-8310, Japan.

Abstract

<b>Background/Objectives:</b> The mandibular cortical index (MCI) is a valuable screening tool for osteoporosis on dental panoramic radiographs, but its assessment is subject to inter-examiner variability. This study evaluated the reproducibility and inter-rater agreement of MCI classification by a closed-source generative AI tool (NotebookLM, Google) compared with eight dentists of varying clinical experience. <b>Methods:</b> One hundred panoramic radiographs were classified according to the three-category MCI in two sessions held at least two weeks apart. Intra-examiner reliability, inter-examiner agreement, and agreement with a reference radiologist were assessed using linearly weighted kappa coefficients. The study was designed as a descriptive reliability study rather than a formal equivalence trial. <b>Results:</b> The intra-examiner reliability of the AI was exceptionally high (κ = 0.987). However, agreement between the AI and the dentists remained at "slight agreement" or lower (κ < 0.2) for every pairing, with 95% confidence intervals that included zero; no formal global hypothesis test was performed, and these individual interval estimates should not be interpreted as proof of the absence of agreement beyond chance. A "two-level discrepancy," in which the AI interchanged Class 1 (normal) and Class 3 (severe), occurred in 10-18% of cases. The dentists showed a possible learning effect, with inter-examiner agreement improving between sessions. <b>Conclusions:</b> Despite the high reproducibility of the NotebookLM configuration evaluated in this study, agreement with the dentists remained at "slight" or lower (κ < 0.2) in MCI classification. As classifications were not validated against bone mineral density or an adjudicated reference standard, these findings concern reproducibility and agreement rather than diagnostic or screening performance and characterize a single LLM-based platform rather than generative AI in general.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.