Back to all papers

DAVLMF-Seg: Vision-language model guided latent frequency-aware diffusion for semi-supervised medical image segmentation.

July 31, 2026pubmed logopapers

Authors

Jiang J,Zhou Q,Chen N,He H,Zhang J,He C

Affiliations (6)

  • Electronic Information School, Wuhan University, Wuhan, 430072, Hubei, China. Electronic address: [email protected].
  • Electronic Information School, Wuhan University, Wuhan, 430072, Hubei, China. Electronic address: [email protected].
  • Electronic Information School, Wuhan University, Wuhan, 430072, Hubei, China. Electronic address: [email protected].
  • Department of Neurology, Zhongnan Hospital of Wuhan University, Wuhan, 430071, Hubei, China. Electronic address: [email protected].
  • Department of Neurology, Zhongnan Hospital of Wuhan University, Wuhan, 430071, Hubei, China. Electronic address: [email protected].
  • Electronic Information School, Wuhan University, Wuhan, 430072, Hubei, China. Electronic address: [email protected].

Abstract

Medical image segmentation is a fundamental task in computer-aided diagnosis and treatment planning. Fully supervised methods achieve strong performance but rely on large-scale annotated datasets. Semi-supervised learning (SSL) alleviates this limitation by leveraging unlabeled data. However, existing SSL methods often lack effective semantic modeling and suffer from domain gaps between natural image pretraining and medical imaging. In this paper, we propose DAVLMF-Seg, a domain-adaptive vision-language model guided frequency-aware SSL framework. The method aligns medical images and textual descriptions in a shared latent space via parameter-efficient adaptation, providing semantic priors for pseudo-label refinement. We further introduce a frequency-domain conditioned diffusion module to progressively enhance feature fusion and reduce decoding ambiguity. An uncertainty-aware regularization strategy is also designed to improve confidence calibration. Extensive experiments on multiple benchmarks demonstrate consistent improvements over state-of-the-art methods. On ACDC, our method achieves 90.25% Dice with 10% labels (+1.2%) and 90.48% Dice with 20% labels (+0.8%), while HD95 is reduced by 2.3 and 1.0, respectively. On M&Ms and MyoPS, it attains 84.57% (+2.0%) and 75.23% (+1.1%) Dice, with HD95 reduced by 1.1 and 4.9, respectively. These results highlight the effectiveness of the proposed framework, especially under limited supervision and cross-domain scenarios.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.