Supporting transformer-based cardiac MRI segmentation with text-to-image controllable diffusion pipelines.
Authors
Affiliations (4)
Affiliations (4)
- Université de Technologie Belfort Montbéliard, UTBM, CIAD, UR 7533, F-90000 Belfort, France; SETIME Laboratory Faculty of Sciences Ibn Tofail University, Kenitra, Morocco. Electronic address: [email protected].
- Université de Technologie Belfort Montbéliard, UTBM, CIAD, UR 7533, F-90000 Belfort, France. Electronic address: [email protected].
- Université de Technologie Belfort Montbéliard, UTBM, CIAD, UR 7533, F-90000 Belfort, France. Electronic address: [email protected].
- SETIME Laboratory Faculty of Sciences Ibn Tofail University, Kenitra, Morocco. Electronic address: [email protected].
Abstract
Cardiac magnetic resonance imaging (MRI) plays a critical role in diagnosing cardiovascular diseases; however, acquiring large, annotated datasets remains a significant challenge due to ethical, economic, and logistical constraints. In this paper, we propose a generative framework based on diffusion models for the controllable synthesis of anatomically consistent cardiac MRI scans with corresponding segmentation labels. Our approach combines Low-Rank Adaptation (LoRA) to generate pathology-aware label maps from textual prompts, and ControlNet to guide image synthesis using both semantic and spatial conditioning. This enables the creation of a fully annotated synthetic dataset aligned with cardiac pathologies and phases, supporting supervised training of segmentation models without additional manual labeling. We evaluate the proposed framework in terms of image realism using Fréchet Inception Distance (FID), Kernel Inception Distance (KID), and Fréchet Radiomic Distance (FRD), as well as downstream segmentation performance using a SegFormer model trained on real, synthetic, and combined datasets. Results show that the proposed method improves segmentation accuracy, particularly in data-limited settings. Beyond in-domain evaluation on the ACDC dataset, cross-dataset experiments on the multi-center M&Ms cohort demonstrate improved generalization and robustness to domain shifts. These findings highlight the potential of text-guided diffusion pipelines to generate high-quality, semantically consistent medical imaging data for robust and generalizable AI training.