Back to all papers

A Clinically Grounded Review of Medical Image Classification: Quantitative Insights into CNNs, Vision Transformers, and Hybrid CNN-ViT Models.

September 3, 2026pubmed logopapers

Authors

Hussaini H,Bano S,Elyan E,Moreno-Garcia CF

Affiliations (2)

  • Computing, Robert Gordon University, Garthdee Road, Aberdeen, AB10 7QB, UK. [email protected].
  • Computing, Robert Gordon University, Garthdee Road, Aberdeen, AB10 7QB, UK.

Abstract

Medical image classification has advanced substantially with convolutional neural networks (CNNs), Vision Transformers (ViTs), and hybrid CNN-ViT architectures, yet clinical translation remains limited by dataset dependency, inconsistent evaluation practices, and insufficient external validation. This review provides a clinically grounded comparative synthesis of these model families across diverse medical imaging modalities. A structured literature review was conducted across PubMed, IEEE Xplore, Scopus, and Web of Science for studies published between 2016 and 2025. Following predefined eligibility criteria, 81 studies were included in the qualitative review, of which 74 contributed to a dataset-aware descriptive quantitative aggregation. The quantitative synthesis was therefore restricted to descriptive aggregation; a formal meta-analysis was not performed because of substantial methodological heterogeneity and insufficient reporting of study-level variance information across the included studies. CNN-based models demonstrated the most consistent performance, achieving the highest weighted accuracy (0.934) and weighted recall (0.901). ViT-based models achieved competitive weighted accuracy (0.906) and recall (0.893), particularly for OCT and X-ray imaging, but appeared more sensitive to dataset scale and quality. Hybrid CNN-ViT models achieved a weighted accuracy of 0.814 and weighted recall of 0.698, with the greatest performance variability. Only 10 of the 81 reviewed studies (12.3%) reported independent external validation, while calibration and other clinically relevant evaluation measures were inconsistently reported. CNNs provide a robust baseline for medical image classification, whereas ViT- and hybrid-based architectures offer complementary strengths under appropriate data and training conditions. However, limited external validation and inconsistent reporting indicate that strong retrospective performance should not be interpreted as evidence of clinical readiness. Future research should prioritise externally validated, interpretable, and clinically deployable AI systems supported by standardised evaluation practices.

Topics

Journal ArticleReview

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.