Back to all papers

Real-world assessment of a deep learning neural network algorithm for prostate cancer detection in MRI using true de novo data.

August 3, 2026pubmed logopapers

Authors

Shea SM,Wesolowski M,Hashem A,Goldberg A,Joyce C,Gupta G,Grimm R,von Busch H,Lou B,Kamen A

Affiliations (7)

  • Department of Radiology & Medical Imaging, Stritch School of Medicine, Loyola University Chicago, Maywood, Illinois, USA.
  • Parkinson School of Public Health, Loyola University Chicago, Maywood, Illinois, USA.
  • Clinical Research Office, Stritch School of Medicine, Loyola University Chicago, Maywood, Illinois, USA.
  • Departments of Urology and Urologic Oncology, Stritch School of Medicine, Loyola University Chicago, Maywood, Illinois, USA.
  • Research & Clinical Translation, Magnetic Resonance, Siemens Healthineers AG, Erlangen, Germany.
  • Strategy & Innovation, Digital & Automation, Siemens Healthineers AG, Forchheim, Germany.
  • Digital Technology and Innovation Division, Siemens Healthineers, Princeton, New Jersey, USA.

Abstract

Deep-learning neural network algorithms for detecting prostate cancer in MRI have proliferated in the literature. However, out of 30+ studies published since the PROSTATEx challenge, no studies tested the performance of their algorithm against using true external image data sets (studies came from an outside institution that did not supply any training data to the algorithm) while validating against MR-US fusion biopsy or whole-mount prostatectomy. Using true external data sets paints a much clearer picture of real-world clinical performance of an algorithm. This work will assess the performance of a published deep learning (DL) neural network algorithm to detect prostate cancer using external studies. The main difference from other studies is the combination of using only MR-US fusion biopsy results as a gold standard; using test data from an institution that did not supply any training data for this version of the algorithm (including studies acquired with an endorectal coil, which were not in the original training set); and comparing the performance of algorithm-generated regions-of-interest (ROIs) versus algorithm heat maps. Patients were included in the study if they had a prostate MRI with at least one radiologist-drawn target on MRI and underwent MR-US fusion biopsy where the target was sampled for pathological analysis. Patients were excluded if they had any history of prostate cancer treatment, had previously undergone MR-US fusion biopsy at our institution, were missing MRI acquisitions, had artifacts in image sets, or if the study had been shared for future algorithm development. MR image data was assessed using a DL research prototype (XProstate) from Siemens Healthineers that produced (a) ROIs in suspected cancer areas with a level of suspicion (LoS) score and (b) heat maps with LoS scores across the entire gland. The XProstate prototype had been trained with 2170 studies from eight different academic institutions. Clinical radiologist, XProstate ROI, and XProstate Heat Map scores were assessed with ROC analysis using pathology results from biopsy as a gold standard. 202 unique patients were included for assessment of the XProstate research prototype. The ROC curve for the XProstate Heat Map LoS score generated the highest AUC (0.76, 95% CI: 0.70, 0.82) followed by clinical radiologist PI-RADS score (0.73, 95% CI: 0.68, 0.79) and by XProstate ROI LoS score (0.71, 95% CI: 0.65, 0.77). Neither the XProstate Heat Map (AUC difference = 0.03, 95% CI: -0.04, 0.10, p = 0.38) nor the XProstate ROI (AUC difference = -0.02, 95% CI: -0.09, 0.04, p = 0.43) was significantly different from the radiologist PI-RADS score. The XProstate prototype demonstrated equivalent performance as clinical radiologists when presented with de novo cases that would mirror a real-world clinical deployment. The automatic ROI delineation more closely matched clinical radiologist performance when using a cutoff of PI-RADS 5 for annotating suspicious regions. Overall, the XProstate prototype provided reasonable clinical performance and this study demonstrated the need to assess Deep Learning prototypes with external institutional test data.

Topics

Prostatic NeoplasmsDeep LearningMagnetic Resonance ImagingImage Processing, Computer-AssistedAlgorithmsJournal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.