Development and multicenter external validation of a deep learning model for early screening of thoracic ossification of the ligamentum flavum on routine chest radiographs.
Authors
Affiliations (4)
Affiliations (4)
- Department of Orthopedics, Ningbo No.6 Hospital, Ningbo, Zhejiang, China.
- Ningbo Clinical Research Center for Orthopedics, Sports Medicine & Rehabilitation, Ningbo, Zhejiang, China.
- Department of Orthopaedic and Reconstructive Surgery/Pediatric Orthopaedics, South China Hospital Affiliated with Shenzhen University, Shenzhen, China.
- Department of Spine Surgery, Changzheng Hospital Affiliated to the Naval Medical University, Shanghai, China.
Abstract
Thoracic ossification of the ligamentum flavum (TOLF) is frequently underrecognized in its early stage because radiographic abnormalities on routine chest radiographs are often subtle. We aimed to develop and externally validate a deep learning model for opportunistic screening of TOLF using routine chest radiographs. This retrospective multicenter diagnostic study included an internal development cohort from Changzheng Hospital and an independent external validation cohort from South China Hospital. The internal cohort comprised 250 patients with TOLF and 250 control subjects collected between January 2017 and January 2023. The external cohort comprised 150 patients with TOLF and 150 control subjects. TOLF status was established on CT using predefined radiological criteria, whereas frontal and lateral chest radiographs were used only as model inputs. We evaluated multiple backbone architectures, including ResNet101, DenseNet169, Vision Transformer, and Swin Transformer, and additionally explored three dual-view fusion strategies. Model development was performed using 10-fold cross-validation in the internal cohort, and performance was summarized using bootstrap-derived 95% confidence intervals. Human-reader comparison was conducted in the internal cohort. In backbone screening within the internal cohort, ResNet101 emerged as the best-performing architecture. After subsequent input-resolution optimization, the final lateral-view ResNet101 model achieved an accuracy of 97.0%, sensitivity of 94.0%, specificity of 100.0%, and an AUC of 0.970. None of the evaluated dual-view fusion strategies outperformed the best single lateral-view model, and the poorer performance of posterior-fusion models was mainly attributable to reduced sensitivity. Compared with experienced spine surgeons and imaging physicians, the internal ResNet101 model showed significantly higher sensitivity and overall accuracy (both <i>p</i> < 0.001). In the external validation cohort, the locked model maintained robust discrimination, with an AUC of 0.954 for frontal radiographs and 0.995 for lateral radiographs. The corresponding accuracy/sensitivity/specificity values were 90.0%/84.7%/95.3% for frontal radiographs and 93.7%/89.3%/98.0% for lateral radiographs. A deep learning model based on routine chest radiographs may provide accurate and generalizable screening for TOLF across institutions. The lateral-view model showed the most consistent diagnostic performance, supporting its potential role as an opportunistic screening tool to prompt confirmatory CT evaluation.