Beyond segmentation accuracy: understanding model reliability and failure modes in ultrasound image analysis.
Authors
Affiliations (3)
Affiliations (3)
- Institute for Neurological and Cardiovascular Research, University of Edinburgh, Edinburgh, United Kingdom.
- Centre for Medical Informatics, University of Edinburgh, Edinburgh, United Kingdom.
- Medical Physics, NHS Lothian, Edinburgh, United Kingdom.
Abstract
Segmentation of ultrasound images is complex, and although deep learning has improved automation, trust remains limited, restricting adoption. Interpretability methods may help reveal model behaviours and failure modes, supporting more trustworthy AI systems. A U-Net with a ResNet-18 backbone was trained to segment pipes in images of the Edinburgh Pipe Phantom acquired using ten ultrasound probes (<i>N</i> = 162). Performance was assessed using five-fold cross-validation. Predictive entropy and variance across 20 Monte Carlo dropout forward passes were used as approximations of data-related and model-related uncertainty. Calibration was assessed using the Brier score and expected calibration error (ECE). Input-gradient saliency and Grad-CAM maps from four encoder stages and the final decoder block were evaluated. For exploratory prospective quality control, complete empty-mask failures were assessed directly from model output, while non-empty predictions were classified retrospectively using reference-based performance criteria and reference-free uncertainty and morphological features were evaluated as candidate indicators. The model achieved a Dice score of 0.81 ± 0.03 and an HD95 of 11.98 ± 3.39. Seven segmentation errors were identified: three complete failures in which no foreground mask was predicted and four non-empty predictions in which the incorrect pipe was segmented. Dice score was associated with the Brier score (Spearman <i>ρ</i> = -0.96) and ECE (<i>ρ</i> = -0.72), whereas predictive entropy and Monte Carlo dropout variance showed weak associations with segmentation performance (|<i>ρ</i>| < 0.4). In selected failure examples, uncertainty and activation sometimes overlapped the reference pipe region despite incorrect segmentation. Complete empty-mask failures were deterministically identifiable from the model output, whereas the evaluated reference-free indicators showed limited ability to identify poor-quality non-empty segmentations. Within this phantom dataset and model configuration, calibration metrics were more closely associated with segmentation quality than the evaluated uncertainty approximations. Attention and uncertainty maps provided complementary descriptive information about selected errors, but indicated spatial sensitivity rather than causal model reasoning. Simple output checks enabled deterministic identification of complete segmentation failures; however, the evaluated reference-free indicators were insufficient for validated prospective identification of subtler non-empty errors. These findings support multimodal model evaluation for EPP segmentation while emphasising the need for validation on independent datasets, devices and architectures in future studies.