Beyond accuracy: a systematic scoping review of algorithmic bias measurement and mitigation in peer-reviewed publications in medical imaging AI development and evaluation for classification.
Authors
Affiliations (2)
Affiliations (2)
- University of Chicago, Department of Radiology, Chicago, Illinois, United States.
- University of Chicago, Pritzker School of Medicine, Chicago, Illinois, United States.
Abstract
The rapid development of medical imaging artificial intelligence (AI) models has raised concerns over ensuring that the outputs of these models, when deployed in clinical environments, provide unbiased classifications across different patient populations or operational factors. This scoping review aimed to investigate how the current literature evaluates algorithmic bias in medical imaging AI models using quantitative fairness metrics. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR), PubMed and Embase were systematically searched for peer-reviewed literature published between January 2015 and May 2025. Studies that evaluated classification medical imaging AI models for algorithmic bias and reported at least one quantitative measure of fairness were included. Articles that investigated fairness in nonimaging AI models used qualitative bias mitigation methods, examined technical bias (e.g., systematic measurement or preprocessing), or were conference proceedings were excluded. After removing duplicates, 2557 studies were included for title and abstract screening. Following initial screening, 143 articles were retrieved for full-text review. Of these, eight studies were eligible for inclusion. Half of the articles focused on bias detection, whereas the other half investigated mitigation techniques. There was considerable heterogeneity in the bias metrics used across the studies. Nearly all studies ( <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>7</mn> <mo>/</mo> <mn>8</mn></mrow> </math> ) used a chest X-ray radiography classifier for evaluation and examined algorithmic bias with respect to gender, age, and race. The relative paucity of existing literature and variability in fairness evaluation methods underscore the need for more standardized frameworks to evaluate and ensure fairness in imaging AI systems.