Back to all papers

An explainable ResNet50-BiLSTM-attention framework with spatial token modeling and imbalance-aware learning for multi-class knee osteoarthritis severity grading.

August 31, 2026pubmed logopapers

Authors

Raza A,Abbas A,Hanif F

Affiliations (2)

  • College of Computer Science, King Saud University Electronic Training Platform (KSUx), King Saud University, Riyadh, Saudi Arabia.
  • Department of Computer Science, Applied College, Girls Section, King Khalid University, Mahayil, Saudi Arabia.

Abstract

Knee osteoarthritis (KOA) severity grading from radiographs remains challenging because adjacent Kellgren-Lawrence stages exhibit subtle and overlapping structural characteristics. This study proposes ResBiAtt-KOA-Net, an explainable image-level framework that introduces spatial-sequence reasoning into convolutional KOA classification. A pre-trained ResNet50 extracts a (2, 048 × 7 × 7) feature map, which is transformed into 49 ordered spatial tokens. A two-layer bidirectional long short-term memory network models contextual dependencies among these tokens, while additive attention identifies the most informative anatomical regions. The attention-guided representation is subsequently fused with global average- and maximum-pooled convolutional descriptors. Weighted random sampling, class-weighted cross-entropy, and label smoothing are incorporated to improve learning from underrepresented severity categories. The framework was evaluated on 9,786 knee radiographs using a patient-level split, direct same-protocol baselines, component-wise ablation, bootstrap confidence intervals, McNemar testing, expert explainability assessment, and external validation. On the internal test set, ResBiAtt-KOA-Net achieved an accuracy of 0.9392, balanced accuracy of 0.9283, macro-F1 of 0.9263, and macro ROC-AUC of 0.9478. It improved macro-F1 by 0.0177 over the strongest baseline, with an adjusted McNemar (<i>p</i>)-value of 0.0181. External evaluation on the MOST cohort produced an accuracy of 0.8615 and macro-F1 of 0.8389, indicating reasonable cross-dataset transfer despite domain differences. Expert assessment identified 82.7% of Grad-CAM maps and 78.0% of spatial-token attention maps as clinically relevant. The model required 32.47 million parameters and an average inference time of 7.6 ms per image. These findings demonstrate that combining convolutional representation learning, spatial-token dependency modeling, attention-guided fusion, and imbalance-aware optimization provides an accurate, interpretable, and reproducible approach to multi-class KOA severity grading.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.