Conditional Latent Diffusion for Synthetic Brain MRI in Alzheimer's Disease: A Preprocessing-Focused Pipeline.
Authors
Affiliations (1)
Affiliations (1)
- Department of Computer and Information Sciences, Northumbria University London, London E1 7HT, UK.
Abstract
Deep learning for Alzheimer's disease (AD) detection from structural magnetic resonance imaging (MRI) needs large, labelled datasets, yet many cohorts hold only a few hundred participants, for which conventional augmentation adds little anatomical diversity. In a two-stage pipeline, a variational autoencoder compressed 256 × 256 coronal slices to a 32 × 32 × 8 latent space, and a class-conditional latent diffusion model under classifier-free guidance generated AD and cognitively normal (CN) images using 295 participants from the Alzheimer's Disease Neuroimaging Initiative (ADNI). The pipeline reached a Kernel Inception Distance (KID) of 0.030 ± 0.002 and a bias-corrected Fréchet Inception Distance (FID<sub>∞</sub>) of 43.82. A controlled ablation varying preprocessing alone improved KID by 0.0147 and precision by 0.069, both with 95% intervals excluding zero. FID did not separate the configurations. A ResNet-18 trained only on synthetic slices and tested on 44 held-out real participants (18 AD, 26 CN), each scored as the mean probability over twenty slices, reached an area under the curve of 0.779 ± 0.031 against 0.869 ± 0.027 for real data; the difference was not distinguishable at this sample size. No instance memorisation was found among 880 samples, and a size-matched control exposed a 27.7-percentage-point inflation in the standard memorisation metric. Preprocessing, therefore, measurably affects synthesis quality at the small-cohort scale, though not on every measure.