Research shows new reasoning language models o3-mini and DeepSeek-R1 still reproduce racial and gender stereotypes in clinical vignettes.
Key Details
- 1Flinders University researchers evaluated o3-mini and DeepSeek-R1 LLMs for bias in generating 36,000 clinical vignettes.
- 2Both models mirrored or exceeded the misrepresentation rates found in older models (up to 89% for race and 67% for gender).
- 3Overrepresentation of Black populations persisted for conditions like sarcoidosis and lupus—44% and 31% median misrepresentation, higher than GPT-4's 15%.
- 4Qualitative review showed models invoked demographic-disease associations without epidemiological justification.
- 5Findings published 28-May-2026, Journal of Medical Internet Research.
Why It Matters

Source
EurekAlert
Related News

TRUECAM AI Framework Boosts Reliability in Cancer Pathology Imaging
PolyU researchers introduce the TRUECAM AI framework to enhance the trustworthiness of pathology AI for cancer diagnosis.

Deep Learning Shows Superior Sensitivity in Sinus Disease CT Changes
Automated deep learning achieves greater sensitivity in detecting treatment-related changes in chronic sinus disease on CT scans than standard visual scoring.

Mayo Clinic's AI Advances Ultrasound Detection of Heart Obstruction
Mayo Clinic researchers have developed an AI model that detects heart obstructions from routine ultrasound, potentially improving early identification of at-risk patients.