A novel approach for liver steatosis assessment in ultrasound images using an adapted Vision-Language Foundation Model.
Authors
Affiliations (9)
Affiliations (9)
- Department of Engineering and Architecture, University of Trieste, Trieste, Italy. Electronic address: [email protected].
- Department of Medical, Surgical, and Health Sciences, University of Trieste, Trieste, Italy; Department of Biomedical Informatics & Data Science, Yale School of Medicine, New Haven, CT, USA; Liver Pathology Clinic, Azienda Sanitaria Universitaria Giuliano Isontina (ASUGI), Trieste, Italy.
- Department of Engineering and Architecture, University of Trieste, Trieste, Italy.
- Unit of Surgical Pathology, Azienda Sanitaria Universitaria Giuliana Isontina (ASUGI), Trieste, Italy.
- Department of Medical, Surgical, and Health Sciences, University of Trieste, Trieste, Italy.
- Institute for Maternal and Child Health IRCCS "Burlo Garofolo", Trieste, Italy.
- Surgical Clinic Unit, Azienda Sanitaria Universitaria Giuliana Isontina (ASUGI), Trieste, Italy.
- Department of Medical, Surgical, and Health Sciences, University of Trieste, Trieste, Italy; Surgical Clinic Unit, Azienda Sanitaria Universitaria Giuliana Isontina (ASUGI), Trieste, Italy.
- Department of Medical, Surgical, and Health Sciences, University of Trieste, Trieste, Italy; Liver Pathology Clinic, Azienda Sanitaria Universitaria Giuliano Isontina (ASUGI), Trieste, Italy; Fondazione Italiana Fegato - ONLUS, Liver Cancer Unit, Trieste, Italy.
Abstract
Metabolic dysfunction-associated steatotic liver disease (MASLD) is among the most prevalent chronic liver diseases worldwide. B-mode ultrasound remains the first-line screening modality, yet its qualitative nature and inter-observer variability limit diagnostic consistency. Existing AI approaches largely rely on single-image classification with opaque outputs that do not reflect exam-level clinical reasoning. In this study, we aimed to develop a vision-language model (VLM) framework for automated, exam-level hepatic steatosis assessment from multi-image ultrasound, producing outputs aligned with the Hamaguchi scoring system. We developed a two-stage framework. First, a compact VLM (Qwen3-VL-8B-Instruct) was adapted to abdominal ultrasound using clinically grounded visual question answering on 16,232 images from 1925 patients. Second, the adapted model performed exam-level Hamaguchi scoring by jointly analyzing all ultrasound images from a given examination using a structured clinical prompt, without any additional task-specific training. Performance was evaluated on an independent cohort with liver histology specimen (179 patients, 383 images) and compared with majority-vote scoring by five expert clinicians. At the feature level, the VLM achieved 95.1% accuracy (κ=0.90-0.96), outperforming image-only baselines on clinically relevant tasks. At the exam level, AI-derived Hamaguchi scores showed strong correlation with histological steatosis (ρ=0.94), exceeding clinician performance (ρ=0.83,p<0.001). Discrimination was high across clinical steatosis severity thresholds (AUC=0.92-0.99), with significant improvements over clinicians at lower severity levels. The proposed compact vision-language framework demonstrated the ability to generate exam-level ultrasound assessments of hepatic steatosis that closely align with biopsy-derived histological severity. It supports scalable, auditable assessment for clinical screening and longitudinal monitoring.