Clinical scale of artificial intelligence-based nerve segmentation on ultrasound: a pilot validation study.
Authors
Affiliations (4)
Affiliations (4)
- Ramsay Santé, Claude Galien Private Hospital, Quincy-Sous-Sénart, France.
- Department of Anaesthesia, University College London Hospitals NHS Foundation Trust, London, United Kingdom.
- Department of Targeted Intervention, University College London, London, United Kingdom.
- Department of Epidemiology and Data Science, Amsterdam UMC, Vrije Universiteit Amsterdam, Amsterdam, Netherlands.
Abstract
The objective metrics for AI-based nerve segmentation are time-consuming and may not fully reflect their clinical usefulness. We aimed to 1) conduct a pilot validation study of the Clinical Segmentation Evaluation Scale (C-SES), as an estimation for an objective metric, and 2) establish a threshold on those scales, because the thresholds were derived for both the new C-SES scale AND for objective metrics (IoU, DSC) to determine the usefulness of AI predictions. Seven experts rated 74 ultrasound sequences twice using AI-based nerve recognition software (cNerve) in 3 anatomical regions. In Step 1, the C-SES was evaluated for 1) content validity (expert appraisal of relevance, comprehensiveness, and clarity), 2) concurrent criterion validity (by correlating the C-SES with a reference standard, the Intersection over Union-IoU), and 3) reliability (intra and inter-rater) and error measurement. In an additional Step 2, expert-perceived usefulness for non-experts and novice anesthesiologists was evaluated to derive candidate thresholds for Intersection over Union (IoU) and C-SES metrics within the investigated clinical context. In Step 1, content validity was preliminarily supported by expert consensus. Concurrent criterion validity was strong (Pearson's <i>r</i> = 0.861). Reliability was excellent within raters [mean ICC(2,1) = 0.905] and high for aggregated inter-rater ratings [ICC(2,3) = 0.869; ICC(2,k) = 0.940]. Precision increased when several raters were pooled. In Step 2, the Youden thresholds for non-expert usefulness were 0.206 for the IoU and 3.51 for the C-SES; for novice usefulness, the thresholds were 0.299 and 4.689, respectively. In this pilot validation, the C-SES appeared to estimate objective segmentation metrics, although further validation is needed. The proposed thresholds are preliminary and hypothesis-generating rather than universal clinical benchmarks and require confirmation before generalization beyond expert assessors and this spectrum-enriched sample.