Back to all news

Next-Gen AI Models Still Perpetuate Biases in Clinical Content

EurekAlertResearch

Research shows new reasoning language models o3-mini and DeepSeek-R1 still reproduce racial and gender stereotypes in clinical vignettes.

Key Details

  • 1Flinders University researchers evaluated o3-mini and DeepSeek-R1 LLMs for bias in generating 36,000 clinical vignettes.
  • 2Both models mirrored or exceeded the misrepresentation rates found in older models (up to 89% for race and 67% for gender).
  • 3Overrepresentation of Black populations persisted for conditions like sarcoidosis and lupus—44% and 31% median misrepresentation, higher than GPT-4's 15%.
  • 4Qualitative review showed models invoked demographic-disease associations without epidemiological justification.
  • 5Findings published 28-May-2026, Journal of Medical Internet Research.

Why It Matters

Radiology and other specialties increasingly adopt LLMs for report drafting and workflow support. Persistent demographic biases in these systems risk perpetuating health disparities, underscoring the need for active bias monitoring and mitigation as AI tools integrate into clinical practice.

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.