Opportunistic Osteoporosis Screening from Routine Knee Radiographs Using a Deep Learning-Based Texture-Branch Cross-Attention Network.
Authors
Affiliations (5)
Affiliations (5)
- Faculty of Medicine, Kütahya Health Sciences University, Kütahya, Turkey. [email protected].
- Faculty of Medicine, Kütahya Health Sciences University, Kütahya, Turkey.
- Department of Orthopedics and Traumatology, Kütahya City Hospital, Kütahya, Turkey.
- Faculty of Medicine, Akdeniz University, Antalya, Turkey.
- Department of Computer Engineering, Siirt University, Siirt, Turkey.
Abstract
This study systematically investigates the impact of architectural components on deep learning-based osteoporosis screening using knee radiographs from a single-center archive. A total of 580 posteroanterior and lateral knee radiographs (normal: n = 389; osteoporosis: n = 191) were included. ImageNet-pretrained backbones (ResNet50V2, VGG16, and DenseNet121) were progressively extended with texture channel packages, dual-stream (global + texture) cross-attention fusion, ordinal-aware supervised contrastive loss, and input-conditioned learnable test-time augmentation (L-TTA). All models were evaluated by fivefold stratified cross-validation using the out-of-fold area under the ROC curve (OOF AUC). TBA-Net v1 achieved the highest performance (OOF AUC = 0.8727), exceeding the strongest baseline VGG16 (0.8497) by 0.0230 AUC and ResNet50V2 (0.8101) and DenseNet121 (0.8092) by 0.0626 and 0.0635 AUC. Ablation analysis demonstrated that the texture branch, cross-attention mechanism, and contrastive loss reduced performance when introduced sequentially, but produced a synergistic effect when jointly optimized; a Shapley decomposition showed that this gain arises from the interaction between components rather than from any component alone. L-TTA Net (OOF AUC = 0.8427) generated clinically interpretable attention patterns. At the current dataset scale, performance plateaued at approximately 0.87 OOF AUC, indicating that further improvements are constrained by sampling variability rather than architectural complexity. Architectural components in small medical imaging datasets should be evaluated within a holistic optimization framework rather than in isolation, and learnable test-time augmentation weights provide interpretability beyond predictive performance. The framework is positioned as a triage aid for prioritizing densitometry referral rather than as a substitute for it.