MPDRL: difficulty-aware representation learning with medical priors for vision-language pre-training.
Authors
Affiliations (2)
Affiliations (2)
- School of Low-Altitude Technology and Engineering, Guangdong Polytechnic Normal University, Heyuan, China.
- School of Computer Science, Guangdong Polytechnic Normal University, Heyuan, China.
Abstract
Medical vision-language pre-training (MedVLP) learns transferable cross-modal representations by aligning chest X-ray images with radiology reports. However, most existing CLIP-style methods have two limitations. First, they treat all unpaired image-report samples within a mini-batch as negatives, which may incorrectly penalize medically similar samples as false negatives. Second, they assume that all samples contribute equally to optimization, overlooking the heterogeneous difficulty of medical data arising from factors such as disease rarity, severity, comorbidity, and report complexity. To address these limitations, we propose Medical Prior-Driven Difficulty-Aware Representation Learning (MPDRL), which incorporates sample-level medical difficulty priors into multimodal representation learning. This method first constructs an eight-dimensional medical prior feature vector for each image-report pair to characterize sample heterogeneity. Based on these priors, it then incorporates difficulty-aware feature modulation, adaptive temperature scaling, medical prior-softened contrastive learning, and difficulty consistency regularization to reduce false-negative penalties among medically similar samples and preserve meaningful difficulty structures in multimodal representations. Extensive experiments on five medical imaging datasets across three downstream tasks demonstrate the effectiveness of MPDRL. For zero-shot classification, MPDRL achieves an AUC of 87.7% on RSNA, outperforming the strongest baseline by 0.2 percentage points, and an AUC of 84.1% on CheXpert, which is 0.1 percentage points lower than the best-performing baseline, IMITATE. On CheXpert14, MPDRL achieves the best F1-score and accuracy, reaching 26.9% and 86.5%, respectively, although its AUC of 71.8% remains lower than those of several competing methods. For zero-shot image-text retrieval on CheXpert 8 × 200, MPDRL achieves Precision@1 scores of 39.9% for image-to-text retrieval and 64.7% for text-to-image retrieval. For report generation, MPDRL achieves BLEU-1 scores of 0.508 on IU X-Ray and 0.403 on MIMIC-CXR. Importantly, clinical evaluation using CheXbert and RadGraph further shows improved medical factual consistency. On IU X-Ray, MPDRL achieves CheXbert precision, recall, and F1 scores of 0.462, 0.414, and 0.437, respectively, while obtaining an RG-F1 score of 0.256. On MIMIC-CXR, it achieves CheXbert precision, recall, and F1 scores of 0.405, 0.354, and 0.378, respectively, together with an RG-F1 score of 0.238. Overall, these results indicate that medical prior-driven difficulty modeling can improve the robustness of MedVLP representations across diverse downstream tasks.