Self-cascade latent Schrödinger bridge for CT field-of-view extension.
Authors
Affiliations (5)
Affiliations (5)
- Pattern Recognition Lab, University of Erlangen Nuremberg Department of Computer Science, Martensstr,3, Erlangen, 91058, Germany.
- School of Physics, Beihang University, No. 37 Xueyuan Road, Haidian District, Beijing, 102206, China.
- Institute of Medical Technology, Peking University Health Science Center, East Side, 2nd Floor, Pathology Building, 38 Xueyuan Rd, Beijing, 100191, China.
- University of Erlangen Nuremberg Department of Computer Science, Martensstr. 3, Erlangen, Bayern, 91058, Germany.
- Institute of Medical Technology, Peking University, Xuyuan Rd. 38, Beijing, 100871, China.
Abstract
Computed tomography (CT) field-of-view (FOV) truncation leaves peripheral anatomy absent, causing errors in radiotherapy dose calculation and body-composition analysis. Existing methods rely on unavailable projection data or suffer over-smoothing and hallucination under severe truncation.

Approach. We propose the self-cascade latent Schrödinger bridge (SCL-SB), combining a latent image-to-image Schrödinger bridge (I2SB) and a self-cascade mechanism in one shared-weight U-Net. I2SB constructs a diffusion bridge from the truncated-image distribution to the full-FOV distribution, enabling high-quality reconstruction in ten denoising steps. Two sequential rounds share identical weights: the first extrapolates, the second refines it using that output. A frequency-aware residual mechanism (FARM) amplifies gradients at high-frequency inter-round residuals to stabilise training and recover faithful detail. The truncation radius is sampled from a stratified distribution spanning mild to severe truncation during training, enabling a single model to cover the full severity spectrum without radius-specific fine-tuning.

Main results. On a head-to-chest CT dataset spanning three truncation severities, SCL-SB, trained as a single model without radius-specific tuning, achieves the highest structural similarity index measure (SSIM) and a high-frequency ratio (HF-Ratio) closest to ideal among compared methods at every severity, running nearly 9× faster than a vision transformer-based baseline. Cascade training alone improves the first round's output without a second inference pass, termed cascade implicit enhancement (CIE). Relative to the same architecture without cascade (I2SB), the model recovers more high-frequency detail (HF-Ratio +45.3%) and a higher SSIM by 0.076 (9.0%) at severe truncation, at no extra cost.

Significance. SCL-SB addresses the posterior-mean regression that blurs high-frequency detail in ill-posed FOV extrapolation, produces continuous Hounsfield unit reconstructions free of Vision-transformer patch-boundary discontinuities, and avoids spectral hallucination characteristic of GAN-based methods. Although clinical validation on larger, multi-centre cohorts is necessary, these properties may improve image guidance in applications such as adaptive radiotherapy, body-composition assessment, and spine surgery.