Back to all papers

Development and internal validation of a task-conditioned multimodal radiology-pathology model for breast cancer staging and biomarker profiling: a retrospective cohort study.

August 21, 2026pubmed logopapers

Authors

Liu Y,Dong Y,Zhang G,Wu Y,Zhang S,Chen Y,Cao R,Wang R,Yin C,Li Y

Affiliations (3)

  • Department of Medical Oncology, The First Affiliated Hospital of Bengbu Medical University, Bengbu, Anhui, China.
  • Faculty of Computing, Harbin Institute of Technology, Harbin, Heilongjiang, China.
  • Department of Biomedical Engineering, Faculty of Engineering, Universiti Malaya, Kuala Lumpur, Malaysia.

Abstract

To retrospectively develop and evaluate a task-conditioned multimodal artificial intelligence framework, in which a learned task embedding and modality-availability mask generate endpoint-specific weights over predefined radiology-only, pathology-only, and fusion experts, for breast cancer staging and biomarker profiling. We retrospectively assembled a cohort of 923 patients with paired radiology and pathology data. Nine tasks included TNM components [tumor extent (T), nodal involvement (N), and distant metastasis (M)], clinical stage, histological grade, and estrogen receptor, progesterone receptor, HER2, and Ki-67 status. Five convolutional neural network backbones were benchmarked for each modality. Task-specific models were selected by validation AUROC, and a late-fusion transformer was used for paired data. We further evaluated learned MoE gating, task-conditioned MoE routing with task embeddings and modality-availability masks, missing-modality robustness, and paired-bootstrap AUROC comparisons. An independent external cohort assessed feasibility for mammography-based pathological complete response and pathology-based triple-negative breast cancer (TNBC) prediction. Best single-modality models achieved AUROCs ranging from 0.606 to 0.990 across tasks, with an AUROC of 0.990 observed for M staging and strong performance observed for histological grade (AUROC, 0.949; accuracy, 0.908) and HER2 status (AUROC, 0.810; accuracy, 0.762). Fixed late fusion improved the mean AUROC from 0.813 to 0.826. Learned MoE gating and task-conditioned MoE further improved mean AUROC to 0.830 and 0.834, respectively. Task-conditioned MoE improved mean AUROC by 0.009 over fixed late fusion (95% CI, 0.005-0.013; P = 0.011) and by 0.021 over the best single-modality expert. Under 70% random missing-modality simulation, task-conditioned MoE outperformed fixed fusion with zero imputation by 0.067 mean AUROC. In external feasibility analyses, the mammography-based pCR model achieved an AUROC of 0.758, and the pathology-based triple-negative breast cancer model achieved an AUROC of 0.872. OmniBreast demonstrated the feasibility of task-conditioned radiology-only, pathology-only, and fused prediction in the internal nine-task benchmark. The external analyses provided preliminary support for feasibility on two related endpoint-extension tasks but did not establish generalizability across the nine primary tasks; multicenter external validation and prospective clinical evaluation remain necessary.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.