Arm and shoulder muscle segmentation in axial MRI with UNet deep learning model.
Authors
Affiliations (7)
Affiliations (7)
- Northwestern University, 633 Clark St, Evanston, IL 60208, USA; Shirley Ryan AbilityLab, 355 E Erie St, Chicago, IL 60611, USA. Electronic address: [email protected].
- North Carolina State University, 915 Partners Way, Raleigh, NC 27606, USA. Electronic address: [email protected].
- Northwestern University, 633 Clark St, Evanston, IL 60208, USA; Shirley Ryan AbilityLab, 355 E Erie St, Chicago, IL 60611, USA. Electronic address: [email protected].
- Northwestern University, 633 Clark St, Evanston, IL 60208, USA; Shirley Ryan AbilityLab, 355 E Erie St, Chicago, IL 60611, USA. Electronic address: [email protected].
- Argonne National Laboratory, 9700 S Cass Ave, Lemont, IL 60539, USA; Northwestern-Argonne Institute of Science and Engineering, 2205 Tech Dr, Evanston, IL 60208, USA. Electronic address: [email protected].
- North Carolina State University, 915 Partners Way, Raleigh, NC 27606, USA. Electronic address: [email protected].
- Northwestern University, 633 Clark St, Evanston, IL 60208, USA; Shirley Ryan AbilityLab, 355 E Erie St, Chicago, IL 60611, USA. Electronic address: [email protected].
Abstract
Quantifying individual upper-limb muscle volumes from MRI provides key insight into muscle-specific strength, deficits, and adaptations. Manual delineation is the gold standard but time‑intensive, and the performance of current deep learning approaches, particularly for small or anatomically complex muscles, remains incompletely characterized. We evaluated a state‑of‑the‑art deep learning framework across the entire upper limb and analyzed factors governing segmentation performance, with attention to the forearm. Three previously published MRI datasets (1.5 T, 3D GRE T1‑weighted; total n = 39) spanning young, middle‑aged, and older adults were curated and quality‑checked, including expert manual segmentations for 31 muscles. Following multiclass mask reconstruction, we trained three 3D nnU‑Net multiclass models matched to the muscle subsets present across datasets, using five‑fold cross‑validation and a composite Dice Similarity Coefficient (DSC) + cross entropy loss. Segmentation accuracy was assessed with DSC. Performance varied across muscles (mean DSC = 0.806 ± 0.098), ranging from 0.920 (Deltoid) to 0.461 (Extensor pollicis brevis). In uncertainty‑weighted regressions, muscle volume was positively associated with DSC (R<sup>2</sup> = 0.36, p < 0.001), whereas training segmentation count and muscle orientation showed negligible associations (R<sup>2</sup> ≤ 0.06). A weighted mixed‑effects model identified volume as the strongest evaluated predictor, explaining 23.9% of variance in DSC; orientation and training count each contributed <1%, leaving 61.5% unexplained. These results indicate that deep learning-based segmentation can accurately quantify muscle volume for many upper‑limb muscles but remains constrained for small, low‑contrast forearm muscles.