DiffINR: conditional neural field diffusion for deformation-driven volumetric image estimation from limited measurements.
Authors
Affiliations (3)
Affiliations (3)
- UT Southwestern Department of Radiation Oncology, 2280 inwood road, Dallas, Texas, 75235-6403, United States.
- Department of Radiation Oncology, The University of Texas Southwestern Medical Center, 2280 Inwood Rd, Dallas, Texas, 75390-9096, United States.
- Department of Radiation Oncology, UT Southwestern Medical Center, 2280 inwood road, Dallas, Texas, 75235, United States.
Abstract
We propose DiffINR for rapid deformation-driven estimation of patient-specific target volumes from limited measurements. DiffINR uses a coordinate-based implicit neural representation (INR) to estimate a dense deformation vector field (DVF) that warps a high-quality prior volume (source) to the target anatomy, with patient-specific INR initialization provided by a population-based conditional hyperdiffusion model.
Approach: During offline training, INRs were optimized on fully sampled source-target image volume pairs to represent DVFs. A Transformer diffusion model, hyperdiffusion, learned the distribution of the resulting INR weights. The model was jointly conditioned on anatomical features extracted from the training source and target volumes by a foundation model and on features encoded directly from limited measurements. At inference, hyperdiffusion generated INR initial weights, which were further fine-tuned by patient-specific optimization using measurement-domain fidelity and DVF-smoothness losses. We evaluated DiffINR on two tasks: cone-beam CT (CBCT) estimation from limited-angle X-ray projections and real-time volumetric MRI estimation from ultra-sparse k-space data.
Main results: For orthogonal-view 90° CBCT estimation, DiffINR achieved a structural similarity index measure (SSIM) of 0.962 ± 0.021, relative error of 6.84% ± 1.70%, and target registration error of 3.61 ± 1.27 mm, outperforming the iterative methods, population-based models, and randomly initialized INR baselines. For volumetric MRI estimation with 13 radial spokes per slice, DiffINR achieved an SSIM of 0.985 ± 0.007, Dice coefficient of 0.89 ± 0.04, and center-of-mass error of 1.30 ± 1.09 mm. DiffINR showed high robustness to data distribution shifts at test time, including dataset changes, limited-angle setting variations, and different MRI sampling sparsity levels. Test-time runtime was 23 s for CBCT estimation and 27 s for MRI estimation, compared with 795 s and 928 s, respectively, for INR optimization from scratch.
Significance: By combining hyperdiffusion-based INR initialization with patient-specific data consistency, DiffINR achieves fast and accurate deformation-driven volumetric image estimation.