Back to all papers

A visual large language foundational model for medical image recognition using clinician-oriented social media

September 7, 2026arxiv logopreprint

Authors

Lingxuan Hou,Yuhua Xie,Yue Hu,Yan Zhuang,Junqi Li,Chengzhi Xia,Binh Phu Nguyen,Abubakar Siddique,Minh Nguyen,Yao Hou,Yanju Bao,Kexin Liu,Ke Chen,Jianjun Sun,Zeqi Li,Trung Nguyen,Jiangli Lin

Abstract

Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in medical settings remains limited by the scarcity of visual question answering (VQA) datasets that capture clinical reasoning and explicit image-text alignment. Here, we leverage de-identified medical images and expert commentaries shared on clinician-oriented social media. By combining an advanced LLM with clinician-in-the-loop verification, we established a rigorous pipeline to construct ThoughtMed-1M, a long-form medical VQA dataset containing over one million VQA pairs and designed to capture structured clinical logic and medical image-text alignment. To demonstrate its utility, we developed a FOundational LLM Trained on ThoughtMed-1M (FOLTMed). FOLTMed achieved state-of-the-art performance across 42 medical VQA benchmark datasets, with a macro accuracy of 85.4%, and generated more clinically coherent responses on the ThoughtMed-1M test set. It outperformed state-of-the-art models by 3--5% across factuality and similarity metrics, highlighting a scalable paradigm for advancing research on clinically grounded multimodal LLMs.

Topics

cs.AI

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.