Back to all papers

Rhamba: Region-aware hybrid attention-Mamba framework for self-supervised learning in resting-state fMRI.

September 21, 2026pubmed logopapers

Authors

Pandey P,Doodipala RR,Eranki P,Torres-Rojas C,Saikia MJ,Sitaram R

Affiliations (6)

  • Multimodal Functional Brain Imaging Research Lab, St. Jude Children's Research Hospital, Memphis, TN 38105, United States. Electronic address: [email protected].
  • Multimodal Functional Brain Imaging Research Lab, St. Jude Children's Research Hospital, Memphis, TN 38105, United States; Biomedical Sensors & Systems Lab, University of Memphis, Memphis, TN 38152, United States. Electronic address: [email protected].
  • Multimodal Functional Brain Imaging Research Lab, St. Jude Children's Research Hospital, Memphis, TN 38105, United States; Biomedical Sensors & Systems Lab, University of Memphis, Memphis, TN 38152, United States.
  • Multimodal Functional Brain Imaging Research Lab, St. Jude Children's Research Hospital, Memphis, TN 38105, United States.
  • Biomedical Sensors & Systems Lab, University of Memphis, Memphis, TN 38152, United States. Electronic address: [email protected].
  • Multimodal Functional Brain Imaging Research Lab, St. Jude Children's Research Hospital, Memphis, TN 38105, United States. Electronic address: [email protected].

Abstract

Self-supervised pretraining is promising for large-scale neuroimaging representation learning, yet the impact of region-aware masking and hybrid sequence modeling remains underexplored. In this work, we introduce Rhamba, a region-aware pretraining framework that integrates anatomically guided masking with hybrid Attention-Mamba architectures for resting state functional magnetic resonance imaging (fMRI) analysis. Models were pretrained on the ABIDE dataset using region-aligned patch embeddings and three masking strategies with increasing spatial specificity. We evaluated four architectural variants: a Mamba-dominant model, an Alternate architecture with interleaved Mamba and Attention blocks, and two hybrid encoder-decoder configurations (Attention-Mamba (AM) and Mamba-Attention (MA)). The pretrained models were fine-tuned on downstream classification tasks using the COBRE and ADHD-200 datasets for schizophrenia and attention-deficit/hyperactivity disorder discrimination. We employed Integrated Gradients, an explainable AI method, to identify the brain regions contributing to model predictions. Our masking strategy strongly influenced reconstruction behavior, with reconstruction loss following a consistent ordering (Any>Majority>Pure). However, this trend did not directly translate into downstream performance, where differences were modest and dataset-dependent. The hybrid architecture with the MA configuration achieved the highest average AUROC across both datasets. Rhamba demonstrated competitive performance against state-of-the-art methods. Region-wise analysis showed that peak performance depends on the interaction between masking strategy and architecture rather than a single dominant configuration. The hybrid architectures underscore the importance of combining global context modeling with efficient sequence dynamics. Overall, Rhamba offers a flexible framework for balancing interpretability, scalability, and performance in large-scale fMRI representation learning.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.