State-of-the-art deep learning for dental implant detection: a standardized benchmark for clinical practice.
Authors
Affiliations (5)
Affiliations (5)
- Harvard John A. Paulson School of Engineering and Applied Sciences, Harvard University, Cambridge, MA, USA.
- Department of Periodontology, Daejeon Dental Hospital, Institute of Wonkwang Dental Research, Wonkwang University College of Dentistry, Daejeon, Korea.
- Department of Oral Medicine and Radiology, Faculty of Dental Sciences, King George's Medical University, Lucknow, India.
- Department of Periodontology and Institute of Oral Bioscience, Jeonbuk National University College of Dentistry, Jeonju, Korea.
- Research Institute of Clinical Medicine of Jeonbuk National University-Biomedical Research Institute of Jeonbuk National University Hospital, Jeonju, Korea. [email protected].
Abstract
Although deep-learning approaches show promise for detecting dental implants, the absence of standardized benchmarks has hindered evidence-based model selection. This study systematically compared five contemporary architectures on the same dataset under controlled conditions. A three-seed training-stability analysis was also performed to characterize between-run variability. YOLOv8-Large, YOLOv11-Large, real-time detection transformer (RT-DETR)-Large, lightweight detection transformer (LW-DETR)-Large, and Faster R-CNN ResNet50-FPN were evaluated using 3,996 annotated panoramic radiographs divided into training, validation, and test sets at an 80:10:10 ratio. Each model was trained using three random seeds to assess training stability. Detection accuracy, measured using the F1 score and mean average precision (mAP), end-to-end pipeline speed, measured using frames per second and latency, and graphics processing unit memory use were evaluated. Pairwise statistical comparisons were performed using the independent-samples <i>t</i>-test with Bonferroni correction. To ensure strict comparability among architectures, all mAP values were calculated using a single unified evaluator, pycocotools COCOeval, applied identically to every model. YOLOv11-Large achieved the highest [email protected] (0.9873±0.0004) and the fastest end-to-end processing speed (87.2 frames per second), with an F1 score of 0.9853±0.0010. YOLOv8-Large performed similarly, with an [email protected] of 0.9859±0.0005 and an F1 score of 0.9868±0.0002; the 95% confidence intervals for [email protected] overlapped between the two models. RT-DETR-Large had the third-highest accuracy ([email protected] = 0.9806±0.0055). LW-DETR-Large achieved an [email protected] of 0.9699±0.0025 but had substantially slower inference (7.1 frames per second), attributable to quadratic attention scaling at a resolution of 1,536×1,536 pixels. Faster R-CNN ResNet50-FPN had lower accuracy ([email protected] = 0.9214±0.0051) and processing speed (14.2 frames per second). All four contemporary detectors significantly outperformed Faster R-CNN ResNet50-FPN (Bonferroni-corrected <i>P</i><0.005). Within this single-center benchmark based on panoramic radiographs, YOLOv11-Large provided the best balance between processing speed and mAP on workstation-class hardware. LW-DETR-Large achieved competitive accuracy but was constrained by Vision Transformer attention scaling at the clinical image resolution used in this study. These findings may inform model selection for dental implant detection on panoramic radiographs. External multicenter and cross-modality validation is required before broader claims regarding clinical deployment can be made.