Back to all papers

Beyond segmentation accuracy: understanding model reliability and failure modes in ultrasound image analysis.

September 23, 2026pubmed logopapers

Authors

Bricout A,Hardman D,Inglis S,Pye S,Moran CM

Affiliations (3)

  • Institute for Neurological and Cardiovascular Research, University of Edinburgh, Edinburgh, United Kingdom.
  • Centre for Medical Informatics, University of Edinburgh, Edinburgh, United Kingdom.
  • Medical Physics, NHS Lothian, Edinburgh, United Kingdom.

Abstract

Segmentation of ultrasound images is complex, and although deep learning has improved automation, trust remains limited, restricting adoption. Interpretability methods may help reveal model behaviours and failure modes, supporting more trustworthy AI systems. A U-Net with a ResNet-18 backbone was trained to segment pipes in images of the Edinburgh Pipe Phantom acquired using ten ultrasound probes (<i>N</i> = 162). Performance was assessed using five-fold cross-validation. Predictive entropy and variance across 20 Monte Carlo dropout forward passes were used as approximations of data-related and model-related uncertainty. Calibration was assessed using the Brier score and expected calibration error (ECE). Input-gradient saliency and Grad-CAM maps from four encoder stages and the final decoder block were evaluated. For exploratory prospective quality control, complete empty-mask failures were assessed directly from model output, while non-empty predictions were classified retrospectively using reference-based performance criteria and reference-free uncertainty and morphological features were evaluated as candidate indicators. The model achieved a Dice score of 0.81 ± 0.03 and an HD95 of 11.98 ± 3.39. Seven segmentation errors were identified: three complete failures in which no foreground mask was predicted and four non-empty predictions in which the incorrect pipe was segmented. Dice score was associated with the Brier score (Spearman <i>ρ</i> = -0.96) and ECE (<i>ρ</i> = -0.72), whereas predictive entropy and Monte Carlo dropout variance showed weak associations with segmentation performance (|<i>ρ</i>| < 0.4). In selected failure examples, uncertainty and activation sometimes overlapped the reference pipe region despite incorrect segmentation. Complete empty-mask failures were deterministically identifiable from the model output, whereas the evaluated reference-free indicators showed limited ability to identify poor-quality non-empty segmentations. Within this phantom dataset and model configuration, calibration metrics were more closely associated with segmentation quality than the evaluated uncertainty approximations. Attention and uncertainty maps provided complementary descriptive information about selected errors, but indicated spatial sensitivity rather than causal model reasoning. Simple output checks enabled deterministic identification of complete segmentation failures; however, the evaluated reference-free indicators were insufficient for validated prospective identification of subtler non-empty errors. These findings support multimodal model evaluation for EPP segmentation while emphasising the need for validation on independent datasets, devices and architectures in future studies.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.