Reliability of the QUADAS-2 checklist in AI-based medical imaging: A comparison of human reviewers and ChatGPT-4o.
Authors
Affiliations (10)
Affiliations (10)
- Dentofacial Deformities Research Center, Research Institute of Dental Sciences, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
- Department of Artificial Intelligence Engineering, Graduate School of Natural and Applied Sciences, Istinye University, Istanbul, Türkiye.
- Department of Restorative Dentistry, UCLA School of Dentistry, Los Angeles, CA, USA.
- Dental Material Research Center, School of Dentistry, Islamic Azad University of Medical Sciences, Tehran, Iran.
- Department of Prosthodontics, School of Dentistry, Islamic Azad University of Medical Sciences, Tehran, Iran.
- Department of Oral and Maxillofacial Medicine, Isfahan University of Medical Sciences, Isfahan, Iran.
- Dental Research Center, Research Institute of Dental Science, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
- Child Growth and Development Research Center, Research Institute for Primordial Prevention of Non-Communicable Disease, Isfahan University of Medical Sciences, Isfahan, Iran.
- Department of Orthodontics and Dentofacial Orthopedics, School of Dentistry, Zanjan University of Medical Sciences, Zanjan, Iran.
- Medical Image and Signal Processing Research Center, Isfahan University of Medical Sciences, Isfahan, Iran.
Abstract
Artificial intelligence is reshaping medical imaging, yet reliable bias assessment in this context remains insufficiently studied. It remains unclear whether the widely used QUADAS-2 checklist provides reproducible evaluations in AI-based imaging studies. Therefore, we compared human and ChatGPT-4o assessments to examine the reproducibility of QUADAS-2 in this setting. Twenty PubMed-indexed AI-based dental imaging studies were analyzed by five human reviewers and five independent ChatGPT-4o runs using identical categorical outputs. Reliability within and between groups was quantified using Fleiss' kappa. Human agreement ranged from -0.07 to 0.22, with only the reference standard domain reaching fair reliability (κ=0.22, <i>P</i><0.001). ChatGPT-4o produced kappa values from -0.17 to 0.56, with the highest consistency observed in the index test domain (κ=0.56, <i>P</i><0.001). Between-group kappa values remained near zero across domains, except in items uniformly classified as low risk. These patterns suggest that QUADAS-2 may yield inconsistent judgments across both groups, despite evaluation of 20 studies and repeated AI classifications. These findings suggest that QUADAS-2 may have limited reproducibility when applied to AI-based imaging research under the present study conditions, although the observed inconsistency may also reflect features of the study design and assessment setting. These findings support consideration of AI-specific tools such as QUADAS-AI for future bias assessment in AI-driven medical imaging.