Deep Learning Method Based on Multimodal Medical Image Fusion for Predicting Acute Vertebral Fracture.
Authors
Affiliations (4)
Affiliations (4)
- Department of Spine Surgery, The Affiliated Jiangning Hospital of Nanjing Medical University, Nanjing, Jiangsu 211100, China (Y.C, Z.W., J.Z., Q.M., C.S.); Division of Spine Surgery, Department of Orthopedic Surgery, Nanjing Drum Tower Hospital, Affiliated Hospital of Medical School, Nanjing University, Nanjing, China (Y.C.).
- School of Natural Resources Science and Technology, Xinjiang University of Technology, Hotan 848000, China (H.H.).
- Department of Spine Surgery, The Affiliated Jiangning Hospital of Nanjing Medical University, Nanjing, Jiangsu 211100, China (Y.C, Z.W., J.Z., Q.M., C.S.).
- Department of Spine Surgery, The Affiliated Jiangning Hospital of Nanjing Medical University, Nanjing, Jiangsu 211100, China (Y.C, Z.W., J.Z., Q.M., C.S.). Electronic address: [email protected].
Abstract
The study aimed to develop and validate a deep learning (DL) model based on X-ray and computed tomography (CT) to diagnose acute vertebral fractures (VFs), and to compare its diagnostic accuracy against spine surgeons. This single-center, retrospective diagnostic accuracy study mainly gathered X-ray, CT, and magnetic resonance imaging (MRI) images of patients diagnosed with vertebral compression fractures who were hospitalized in the Department of Spinal Surgery of the Affiliated Jiangning Hospital of Nanjing Medical University from January 2024 to June 2025. Spine image registration was achieved using the "Landmark Registration" tool in 3D Slicer. The latest You-Only-Look-Once (YOLO) v11 target detection technology was employed to construct the DL model, which was trained using multimodal images. Four configurations (X-ray only, CT only, input-level fusion, and decision-level fusion) were evaluated using metrics such as mean average precision, precision, recall, and F1 score. Finally, the model's performance was compared against that of middle-grade and senior spine surgeons. The input-level fusion model achieved the best performance, with an area under the precision-recall curve of 0.990 and an F1-score of 0.953 at the vertebra level. To minimize missed diagnoses, a threshold of 0.24 was selected, achieving a recall of 1.0 and a precision of 0.855. The model significantly outperformed middle-grade surgeons and demonstrated superior sensitivity (0.962 vs. 0.925) and specificity (0.981 vs. 0.955) compared to senior surgeons. The YOLOv11-based input-level fusion model accurately predicts MRI-defined acute VFs using routine X-ray and CT imaging.