Back to all papers

Self-cascade latent Schrödinger bridge for CT field-of-view extension.

September 10, 2026pubmed logopapers

Authors

Wu X,Liu J,Yu H,Maier A,Huang Y

Affiliations (5)

  • Pattern Recognition Lab, University of Erlangen Nuremberg Department of Computer Science, Martensstr,3, Erlangen, 91058, Germany.
  • School of Physics, Beihang University, No. 37 Xueyuan Road, Haidian District, Beijing, 102206, China.
  • Institute of Medical Technology, Peking University Health Science Center, East Side, 2nd Floor, Pathology Building, 38 Xueyuan Rd, Beijing, 100191, China.
  • University of Erlangen Nuremberg Department of Computer Science, Martensstr. 3, Erlangen, Bayern, 91058, Germany.
  • Institute of Medical Technology, Peking University, Xuyuan Rd. 38, Beijing, 100871, China.

Abstract

Computed tomography (CT) field-of-view (FOV) truncation leaves peripheral anatomy absent, causing errors in radiotherapy dose calculation and body-composition analysis. Existing methods rely on unavailable projection data or suffer over-smoothing and hallucination under severe truncation.

Approach. We propose the self-cascade latent Schrödinger bridge (SCL-SB), combining a latent image-to-image Schrödinger bridge (I2SB) and a self-cascade mechanism in one shared-weight U-Net. I2SB constructs a diffusion bridge from the truncated-image distribution to the full-FOV distribution, enabling high-quality reconstruction in ten denoising steps. Two sequential rounds share identical weights: the first extrapolates, the second refines it using that output. A frequency-aware residual mechanism (FARM) amplifies gradients at high-frequency inter-round residuals to stabilise training and recover faithful detail. The truncation radius is sampled from a stratified distribution spanning mild to severe truncation during training, enabling a single model to cover the full severity spectrum without radius-specific fine-tuning.

Main results. On a head-to-chest CT dataset spanning three truncation severities, SCL-SB, trained as a single model without radius-specific tuning, achieves the highest structural similarity index measure (SSIM) and a high-frequency ratio (HF-Ratio) closest to ideal among compared methods at every severity, running nearly 9× faster than a vision transformer-based baseline. Cascade training alone improves the first round's output without a second inference pass, termed cascade implicit enhancement (CIE). Relative to the same architecture without cascade (I2SB), the model recovers more high-frequency detail (HF-Ratio +45.3%) and a higher SSIM by 0.076 (9.0%) at severe truncation, at no extra cost.

Significance. SCL-SB addresses the posterior-mean regression that blurs high-frequency detail in ill-posed FOV extrapolation, produces continuous Hounsfield unit reconstructions free of Vision-transformer patch-boundary discontinuities, and avoids spectral hallucination characteristic of GAN-based methods. Although clinical validation on larger, multi-centre cohorts is necessary, these properties may improve image guidance in applications such as adaptive radiotherapy, body-composition assessment, and spine surgery.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.