RV-mixer: Random Fourier and variance-guided global feature learning for medical image segmentation.
Authors
Affiliations (3)
Affiliations (3)
- Department of Hepatobiliary, Pancreas and Spleen Surgery, the People's Hospital of Guangxi Zhuang Autonomous Region (Guangxi Academy of Medical Sciences), Nanning, China; Department of General Surgery, First People's Hospital of Fangchenggang, Guangxi Zhuang Autonomous Region, Fangchenggang, China. Electronic address: [email protected].
- School of Computer Science, Northwestern Polytechnical University, Xi'an, 710129, China; School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University, Xi'an, 710072, China. Electronic address: [email protected].
- Department of Hepatobiliary, Pancreas and Spleen Surgery, the People's Hospital of Guangxi Zhuang Autonomous Region (Guangxi Academy of Medical Sciences), Nanning, China; School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University, Xi'an, 710072, China. Electronic address: [email protected].
Abstract
Hybrid CNN-Transformer architectures have achieved promising performance in medical image segmentation, yet the high computational cost of self-attention limits their application in resource-constrained clinical scenarios. Moreover, medical images exhibit domain-specific challenges, including ultrasound speckle noise, low contrast, and ambiguous lesion boundaries, which require robust feature representation beyond conventional local modeling. To address these issues, we propose RV-Mixer, a lightweight module for efficient global representation learning with medical-domain priors. Instead of explicit attention-based token interaction, RV-Mixer employs Random Fourier Features (RFF) to project local representations into a shared frequency-domain embedding space, providing nonlinear feature enhancement with low computational overhead. Furthermore, Variance-Guided Channel Recalibration (VGCR) introduces global statistical aggregation through channel-wise variance modeling, enabling adaptive emphasis of discriminative pathological textures and lesion boundaries that may be weakened by mean-based feature aggregation. Extensive experiments on five public medical image segmentation datasets, covering ultrasound, endoscopy, dermoscopy, and MRI scenarios, demonstrate that RV-Mixer achieves competitive or superior performance compared with Transformer-, MLP-, and Mamba-based approaches while maintaining fewer parameters and lower computational complexity. These results highlight the effectiveness of domain-informed frequency representation and statistical global aggregation for efficient and robust medical image segmentation.