Back to all papers

Calibrating artificial intelligence against human expertise using femoral nerve segmentation on ultrasound: a consensus framework.

September 1, 2026pubmed logopapers

Authors

Berggreen J,Johansson A,Möller S,Jansson T,Augustinsson A,Jildenstål P

Affiliations (6)

  • Department of Clinical Sciences Lund/Biomedical Engineering, Lund University, Lund, Sweden.
  • Digitalisering IT/MT, Skåne Regional Council, Lund, Sweden.
  • Department of Health Sciences, Lund University, Lund, Sweden.
  • Department of Anaesthesiology, Surgery and Intensive Care Medicine, Sahlgrenska University Hospital, Gothenburg, Sweden.
  • Faculty of Nursing and Health Sciences, Nord University, Bodø, Norway.
  • Department of Anaesthesiology and Intensive Care, Skåne University Hospital, Lund, Sweden.

Abstract

Meaningful validation of artificial intelligence for medical image interpretation requires comparison against human expert performance, yet multi-rater frameworks establishing such comparisons remain uncommon. We developed and applied a consensus framework using nine clinicians who independently segmented the femoral nerve on 100 ultrasound images, yielding 900 annotations and a combined consensus standard established by majority voting. We then evaluated an academic deep learning model against this consensus and individual human performance. The artificial intelligence model achieved a median (IQR [range]) Dice coefficient of 0.72 (0.56-0.84 [0.00-0.91]) against combined consensus. Sensitivity was 0.94 (0.88-0.97 [0.33-1.00]) and precision 0.60 (0.44-0.76 [0.00-0.89]). Individual human Dice scores ranged from 0.32 to 0.73 (median 0.60). The artificial intelligence model matched median human performance and outperformed five of nine annotators (31%-125% relative improvement), with the greatest benefit for the lowest-performing practitioners. Leave-one-annotator-out analysis confirmed consensus stability (median (IQR [range]) artificial intelligence Dice 0.749 (0.745-0.752 [0.742-0.769])). Inter-rater reliability was moderate overall (Fleiss's κ 0.54, p < 0.001). The sensitivity and precision profile of the artificial intelligence model indicated reliable nerve detection with over-segmentation that remained clinically interpretable. The moderate inter-rater reliability is consistent with the inherent subjectivity of nerve delineation on ultrasound. The circularity inherent in evaluating annotators against a consensus they helped define limits direct comparison of artificial intelligence and human scores. A Dice score of 0.72 represents the upper range of human expert performance rather than moderate accuracy. The framework methodology is independent of the specific artificial intelligence system evaluated and offers a transferable approach for calibrating artificial intelligence performance in clinical imaging where no single correct interpretation exists.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.