Back to all papers

An explainable biomedical foundation model via large-scale concept-enhanced vision-language pretraining.

August 17, 2026pubmed logopapers

Authors

Nie Y,He S,Bie Y,Wang Y,Chen Z,Yang S,Cai Z,Wu L,Wang H,Wang X,Cheng NS,Luo L,Wu M,Jin H,Wu X,Chan RCK,Lau YM,Zhang Z,Xiao S,Yang C,Zhao Y,Duan X,Zhang L,Liang L,Zheng Y,Rajpurkar P,Chen H

Affiliations (20)

  • Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong, China.
  • Department of Biomedical Informatics, Harvard University, Boston, MA, USA.
  • Department of Radiology, Shenzhen People's Hospital, Shenzhen, China.
  • Jarvis Research Center, Tencent YouTu Lab, Shenzhen, China.
  • Department of Anatomical and Cellular Pathology, The Chinese University of Hong Kong, Hong Kong, China.
  • State Key Laboratory of Translational Oncology, The Chinese University of Hong Kong, Hong Kong, China.
  • Division of Dermatology, Department of Medicine, Queen Mary Hospital, Hong Kong, China.
  • Department of Pathology, Nanfang Hospital and School of Basic Medical Sciences, Southern Medical University, Guangzhou, China.
  • Guangdong Provincial Key Laboratory of Molecular Tumor Pathology, Guangzhou, China.
  • Jinfeng Laboratory, Chongqing, China.
  • Department of Ultrasound Medicine, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, China.
  • Department of Mathematics, The Hong Kong University of Science and Technology, Hong Kong, China.
  • Third Affiliated Hospital of Southern Medical University (Academy of Orthopedics, Guangdong Province), Guangzhou, China.
  • Department of Radiology, Sun Yat-Sen Memorial Hospital, Sun Yat-Sen University, Guangzhou, China.
  • Medical Artificial Intelligence Laboratory, Westlake University, Hangzhou, China.
  • Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong, China. [email protected].
  • Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology, Hong Kong, China. [email protected].
  • Division of Life Science, The Hong Kong University of Science and Technology, Hong Kong, China. [email protected].
  • HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, Shenzhen, China. [email protected].
  • State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology, Hong Kong, China. [email protected].

Abstract

Artificial intelligence for medical imaging is required to be accurate and interpretable to clinicians. However, current multimodal biomedical foundation models often prioritize performance over explainability. Here we present ConceptCLIP, an explainable biomedical foundation model that achieves state-of-the-art diagnostic accuracy while delivering human-interpretable explanations across diverse imaging modalities. We curate MedConcept-23M, a large-scale dataset comprising 23 million biomedical image-text-concept triplets. Leveraging this dataset, we pretrain ConceptCLIP via joint image-text and region-concept alignment for precise and interpretable medical image analysis. Across a large-scale benchmark covering 78 datasets in 10 imaging modalities, ConceptCLIP demonstrates superior diagnostic performance while providing human-understandable explanations. In a clinician user study spanning three modalities, the concept-based explanations provided by ConceptCLIP help clinicians verify model predictions and identify potential errors. As an explainable biomedical foundation model, ConceptCLIP represents a critical milestone towards the widespread clinical adoption of AI, thereby advancing trustworthy AI in medicine.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.