Back to all papers

From Routine Imaging to Risk Stratification: Multimodal Vision-Language Survival Modeling for Pancreatic Cancer.

August 12, 2026pubmed logopapers

Authors

Le D,Medero RLC,Tariq A,Joshi V,Zhang Z,Murlidhar F,Dong H,Peng Y,Shih G,Wang Z,Pannala R,Yano M,Banerjee I

Affiliations (4)

  • Mayo Clinic Arizona, Phoenix, AZ, USA.
  • Arizona State University, Tempe, AZ, USA.
  • Weill Cornell Medicine, New York, NY, USA.
  • University of California, San Francisco, San Francisco, CA, USA.

Abstract

Pancreatic ductal adenocarcinoma (PDAC) is frequently diagnosed in advanced stages, substantially limiting opportunities for early intervention. In this work, we present a multimodal survival modeling framework for prediagnostic PDAC risk stratification using routinely acquired clinical data, including abdominal computed tomography (CT) imaging, radiology reports, and structured electronic health record (EHR) variables. Our first contribution is a unified multimodal framework that integrates heterogeneous clinical data sources for prediagnostic risk modeling. Second, to address the sparsity and heterogeneity of EHR data, we introduce text-based encoding of clinical variables, while volumetric variability in CT imaging is mitigated through automated pancreas detection and standardized subvolume selection. Third, we integrate a vision-language foundation model (VLM) with a survival modeling objective based on negative log-likelihood to estimate cancer-free survival. Model performance was evaluated on internal and external validation cohorts using the concordance index (C-index). Across cohorts, multimodal fusion generally outperformed unimodal approaches. Vision-language models demonstrated strong and consistent discriminative performance, while feature-engineered models achieved competitive performance, particularly on external validation. Overall, multimodal integration provided the most robust performance, highlighting the complementary value of combining imaging, text, and structured clinical data. Finally, to address interpretability challenges associated with VLM-based modalities, we conducted systematic ablation studies using modality-specific occlusion and noise perturbation to quantify the contribution of image and text features. These results support the feasibility of opportunistic PDAC risk stratification from routinely collected multimodal clinical data and underscore the potential of multimodal representation learning for early risk identification.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.