Back to all papers

The benchmark illusion: on the curious underestimation of AI and the uncertainty of its creators' gold standards.

August 31, 2026pubmed logopapers

Authors

Pohlkamp C,Haferlach T

Affiliations (1)

  • MLL Munich Leukemia Laboratory, Max-Lebsche-Platz 31, 81377, Munich, Germany.

Abstract

Artificial intelligence (AI) systems have quietly crossed an uncomfortable threshold: on a growing range of diagnostic tasks, they often match, and in some well-defined tasks even exceed, human performance. Evidence from large language models such as Med-PaLM and GPT-4, and deep learning systems in pathology, radiology, and hematology, shows high accuracy and reproducibility, challenging traditional human benchmarks, although still being subject to substantial heterogeneity in the studies cited. Yet most evaluation frameworks continue to treat human consensus on curated, retrospectively labeled datasets as ground truth, despite well-documented record of cognitive bias, interobserver variability, and diagnostic errors in clinical practice. This creates a striking asymmetry: AI is scrutinized against a putative gold standard whose gold content has never been formally tested-a benchmark illusion. Acknowledging AI's own distinctive failure modes and generalizability problems, this Viewpoint argues that such limitations do not justify clinging to an untested reference. It calls for explicit modeling of ground truth uncertainty, outcome-based validation, and, where appropriate, using AI as an additional reference layer for human performance. The relevant benchmark is no longer clinicians alone, but the human-AI ensemble evaluated within real diagnostic ecosystems and against patient-relevant outcomes. None.

Topics

Journal ArticleReview

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.