Hierarchical Multimodal Fusion of Multi-Sequence MRI and Clinical Metadata for the Classification of Rotator Cuff Tears.
Authors
Affiliations (5)
Affiliations (5)
- Department of Software Engineering, Faculty of Engineering and Architecture, Eskişehir Osmangazi University, Eskişehir 26040, Türkiye.
- Center of Intelligent Systems Applications Research, Eskişehir Osmangazi University, Eskişehir 26040, Türkiye.
- Department of Computer Engineering, Faculty of Engineering and Architecture, Eskişehir Osmangazi University, Eskişehir 26040, Türkiye.
- Department of Orthopedics and Traumatology, Faculty of Medicine, Bilecik Şeyh Edebali University, Bilecik 11230, Türkiye.
- Department of Radiology, Yunus Emre State Hospital, Ministry of Health, Eskişehir 26190, Türkiye.
Abstract
<b>Background/Objectives</b>: Rotator cuff tears are a leading cause of shoulder disability. While multi-sequence MRI is standard, the optimal deep learning integration of heterogeneous image series and clinical metadata remains unresolved. This study evaluated a hierarchical, sequence-aware multimodal framework for patient-level binary rotator cuff tear classification. <b>Methods</b>: A single-center cohort of 199 patients (100 tears, 99 controls) was analyzed across four MRI sequences (T1 coronal, T2 fat-suppressed sagittal, and proton density [PD] fat-suppressed coronal and transverse/axial) and nine demographic features. Under a patient-level stratified three-fold cross-validation scheme preventing data leakage, we evaluated ResNet50 and Vision Transformer baselines (Study 0), full-protocol fusion topologies (Study 1), and systematically mapped sequence-subset combinations with or without metadata (Study 2). <b>Results</b>: In Study 0, the PD coronal ResNet50 model was the top baseline (AUC = 0.9834, F1 = 0.9515). In Study 1, late decision fusion yielded the highest AUC (0.9909), while feature concatenation optimized threshold balance (F1 = 0.9502). In Study 2, a streamlined three-sequence subset with metadata (C14M: T2 + PDc + PDt) achieved peak performance (AUC = 0.9961, 95% CI: 0.9823-0.9987, F1 = 0.9618, MCC = 0.9238), outperforming the full protocol (AUC = 0.9909, F1 = 0.9355). Metadata utility was configuration-dependent, assisting only fluid-sensitive combinations. <b>Conclusions</b>: Rather than indiscriminately aggregating entire clinical protocols, multimodal fusion is optimized by selecting complementary imaging series. For binary classification, excluding non-fat-suppressed T1 images in favor of a streamlined T2 and PD set stabilized by clinical demographics maximized classification performance in this internally validated, single-center cohort.