Concerns Raised Over Unverified Datasets in AI Health Prediction Models

A new study finds widely used AI health prediction models are built on datasets with unverifiable origins, raising safety and validity concerns.
Key Details
- 1Researchers at QUT and AusHSI examined two popular health datasets on Kaggle used for stroke and diabetes prediction.
- 2These datasets have been cited in 125 peer-reviewed studies, with little to no information about data provenance.
- 3Three AI models trained on these datasets have been used in clinical practice; one was cited in a medical device patent.
- 4The datasets scored 0/9 on the TRIPOD+AI data provenance criteria, indicating they are unsuitable for clinical use.
- 5Seven articles based on these datasets have already been retracted as unreliable.
- 6Researchers urge journals, funders, and data repositories to enforce stricter data-source disclosure, and recommend removal of the problematic datasets.
Why It Matters

Source
EurekAlert
Related News

AI Pathology Tool SÉMIL Improves Stage II Bowel Cancer Risk Assessment
A La Trobe University-developed AI tool accurately predicts relapse risk in stage II bowel cancer using digital pathology images and descriptions.

AI Tool Predicts Which Rectal Cancer Patients Benefit from Intensive Therapy
UCL researchers developed an AI that analyzes biopsy slides to identify rectal cancer patients who benefit from adding irinotecan to standard chemoradiotherapy.

AI-Guided Handheld Cardiac Ultrasound Reduces Referrals and Costs in Spain
AI-guided handheld cardiac ultrasound enables primary care physicians to detect heart failure, reducing specialist referrals and saving costs.