Back to all papers

Quantifying User Satisfaction: Weighted Metric Approach for Evaluating Deep Learning-Based Thigh MRI Segmentations.

July 31, 2026pubmed logopapers

Authors

Ensle F,Koska IÖ,Derron N,Koska C,Bach U,Maintz PN,Porta-Vilarό M,Kroschke J,Gerber P,Guggenberger R

Affiliations (7)

  • Diagnostic and Interventional Radiology, University Hospital of Zurich, 8091 Zurich, Switzerland.
  • Department of Pediatric Radiology, Acıbadem Kent Hospital, 35620 Izmir, Türkiye.
  • Biomedical Technologies Department, Engineering Faculty, Dokuz Eylül University, 35390 Izmir, Türkiye.
  • Department of Endocrinology, Diabetology and Clinical Nutrition, University Hospital of Zurich, 8091 Zurich, Switzerland.
  • Department of Electrical Electronical, Engineering Faculty, Yaşar University, 35100 Izmir, Türkiye.
  • Department of Radiology, Hospital Clínic de Barcelona, 08036 Barcelona, Spain.
  • Department of Radiology and Nuclear Medicine, Cantonal Hospital Winterthur, University of Zurich, 8091 Zurich, Switzerland.

Abstract

<b>Background:</b> The aim of this paper is to develop a clinically useful benchmark signature of deep learning (DL)-based MRI segmentations and objectively predict user satisfaction based on a weighted combination of multiple performance metrics. <b>Methods:</b> This post hoc analysis of a prospective study analyzed MRI data from 68 patients (29.4 ± 5.6 years; 45% female) acquired during a randomized clinical trial. Fat fraction maps of axial Dixon MRI were used to segment six classes of the thigh. Two radiologists qualitatively scored DL-based segmentations using a 5-point Likert scale. A hybrid deep learning model (MobileNetV2 and DINOv2) was developed to predict Likert scores directly from MRI slices and segmentation maps. Additionally, a weighted combination of quantitative metrics (Dice score, Hausdorff distance, Jaccard index) was determined to best correlate with the Likert scores. Statistical analyses included Cohen's kappa, Spearman correlation, and regression model performance (MAE, RMSE, ROC_AUC). <b>Results:</b> Interreader agreement for Likert scores was near-perfect (κ = 0.82). The DINOv2 model achieved lower MAE (0.2343) and RMSE (0.1129) than MobileNetV2 (0.3488 and 0.431, respectively) in predicting Likert scores. The weighted metric combination model, using Lasso regression, predicted Likert = 5 with 84.1% accuracy, 100% recall, and ROC_AUC of 0.824. Individual metrics showed weak-to-moderate correlation with Likert scores, with Dice and Jaccard indices for extensor intramuscular fat exhibiting the highest correlation (r = 0.37 and r = 0.34, respectively). <b>Conclusions:</b> The proposed DL-based model accurately predicted radiologist-assigned Likert five scores, while the weighted metric combination model aligned closely with user satisfaction, offering a practical alternative to manual validation for clinical adoption of automated segmentation tools.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.