Research shows new reasoning language models o3-mini and DeepSeek-R1 still reproduce racial and gender stereotypes in clinical vignettes.
Key Details
- 1Flinders University researchers evaluated o3-mini and DeepSeek-R1 LLMs for bias in generating 36,000 clinical vignettes.
- 2Both models mirrored or exceeded the misrepresentation rates found in older models (up to 89% for race and 67% for gender).
- 3Overrepresentation of Black populations persisted for conditions like sarcoidosis and lupus—44% and 31% median misrepresentation, higher than GPT-4's 15%.
- 4Qualitative review showed models invoked demographic-disease associations without epidemiological justification.
- 5Findings published 28-May-2026, Journal of Medical Internet Research.
Why It Matters

Source
EurekAlert
Related News

AI Model Combines ECG and Blood Tests to Spot Heart Transplant Rejection
An AI model integrating ECG data and blood biomarkers predicts heart transplant rejection with high specificity, reducing the need for invasive biopsies.

NIH Grant Funds AI Tool for Lung Cancer Discovery at Weill Cornell
Dr. Olivier Elemento receives an NIH Pioneer Award to develop an AI-human hybrid platform for cancer research, initially focusing on lung cancer using imaging and tissue data.

Study Shows AI Vision Misses Human-Like Perceptual Illusions
A York University study finds that AI vision systems lack the human-like perceptual shifts seen in biological vision during visual illusions.