Back to all papers

Evaluation metrics for synthetic medical imaging.

August 12, 2026pubmed logopapers

Authors

de Wilde D,Schärli B,Ackermann K,Bottini M,Zanier O,Da Mutten R,Khan I,Elmi-Terander A,Edström E,Sarwin G,Regli L,Serra C,Staartjes VE

Affiliations (6)

  • Machine Intelligence in Clinical Neuroscience & Microsurgical Neuroanatomy (MICN) Laboratory, Department of Neurosurgery, Clinical Neuroscience Center, University Hospital Zurich, University of Zurich, Zurich, Switzerland.
  • Machine Intelligence in Clinical Neuroscience & Microsurgical Neuroanatomy (MICN) Laboratory, Department of Neurosurgery, Clinical Neuroscience Center, University Hospital Zurich, University of Zurich, Zurich, Switzerland; Department of Oncology, University of Oxford, Oxford, United Kingdom.
  • Machine Intelligence in Clinical Neuroscience & Microsurgical Neuroanatomy (MICN) Laboratory, Department of Neurosurgery, Clinical Neuroscience Center, University Hospital Zurich, University of Zurich, Zurich, Switzerland; Department of Neurosurgery, Cantonal Hospital Lucerne, Lucerne, Switzerland.
  • Department of Clinical Neuroscience, Karolinska Institutet, Stockholm, Sweden.
  • Computer Vision Laboratory, ETH Zurich, Zurich, Switzerland.
  • Machine Intelligence in Clinical Neuroscience & Microsurgical Neuroanatomy (MICN) Laboratory, Department of Neurosurgery, Clinical Neuroscience Center, University Hospital Zurich, University of Zurich, Zurich, Switzerland; Department of Oncology, University of Oxford, Oxford, United Kingdom. Electronic address: [email protected].

Abstract

Advances in generative artificial intelligence (AI) have accelerated the development and application of synthetic medical imaging. Despite this rapid progress, the evaluation of synthetic medical images remains heterogeneous, with numerous metrics proposed to assess fidelity, realism, diversity, and clinical validity. Currently, no standardized framework exists to guide the selection, interpretation, or comparison of these metrics, limiting reproducibility and cross-study comparability. This systematic review aims to comprehensively summarize and categorize existing metrics used to assess these complementary dimensions of synthetic medical images. A systematic review was conducted in accordance with PRISMA guidelines. PubMed/MEDLINE, EMBASE, Scopus, and arXiv were searched for studies published between 2015 and April 30, 2025, supplemented by citation screening of included studies. Eligible studies were full-text articles that applied or proposed metrics to evaluate the fidelity, realism, diversity, and/or clinical validity in synthetic medical images. A total of 47 studies were included. Evaluation practices were highly heterogeneous. Expert evaluation (n = 25, 53%) and reference-based evaluations were most common (n = 25, 53%), followed by no-reference metrics (n = 24, 51%), and task-based evaluations (n = 24, 51%). The most commonly used individual metrics were peak signal-to-noise ratio (PSNR) (n = 16, 34%), structural similarity index (SSIM) (n = 15, 32%), mean absolute error (MAE) (n = 12, 26%), and Fréchet Inception Distance (FID) (n = 12, 26%). Evaluation strategies for synthetic medical imaging showed substantial variability and no single metric captured fidelity, realism, diversity, and clinical validity simultaneously. Metric choice is often dictated by data availability rather than clinical purpose. A task-specific, layered evaluation framework could improve comparability and facilitate clinical adoption.

Topics

Journal ArticleReview

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.