A Detection, Diagnostic, and Triage AI Buyer's Guide: Questions that Matter.
Authors
Affiliations (7)
Affiliations (7)
- Department of Radiology, Radiology Human Factors Lab, Brown University Health and The Warren Alpert Medical School, Brown University, Providence, RI, USA. [email protected].
- Brown University Health AI Center of Excellence, Brown University Health, Providence, RI, USA. [email protected].
- Department of Radiology, Radiology Human Factors Lab, Brown University Health and The Warren Alpert Medical School, Brown University, Providence, RI, USA.
- Brown University Health AI Center of Excellence, Brown University Health, Providence, RI, USA.
- Radiology Associates of North Texas P.A., Fort Worth, TX, USA.
- Department of Radiology, Mayo Clinic School of Medicine, Rochester, MN, USA.
- Department of Radiology, The Pennsylvania State University College of Medicine, Hershey, PA, USA.
Abstract
The published performance of artificial intelligence (AI) models in radiology is typically based on the reporting of sensitivity, specificity, and receiver operating characteristic area under the curve, both in the peer-reviewed literature and for Food and Drug Administration 510(k) submissions. Interestingly, these metrics cannot inform radiologists, the users of these systems, of the rate or quantity of each type of error to anticipate if they implement the candidate AI product(s) into their own clinical practice. Only the positive predictive value (PPV) and negative predictive value (NPV), or rather, their complements, the false discovery rate (1-PPV, FDR), and the false omission rate (1-NPV, FOR) can provide these error rates. Although some published articles and 510(k) submissions include the PPV and NPV of the concerned AI models, many do not, and the ones that do sometimes test on enriched datasets with artificially high prevalence rates of the target condition, thus inflating PPV and deflating NPV that would be found in clinical practice. This manuscript demonstrates how clinical practices can estimate and evaluate an AI's FDR and FOR for their clinical population using Bayes' Theorem. We also propose a risk-based evaluation matrix (RADDE) which allows radiologists to consider the medical, legal, financial, workflow, psychological, and reputational impact of these estimated AI error rates.