Back to all papers

Implementation of an artificial intelligence-based system for mammography in the compulsory medical insurance program: results of a 3-year study.

August 10, 2026pubmed logopapers

Authors

Vasilev Y,Rumyantsev D,Vladzymyrskyy A,Omelyanskaya O,Arzamasov K,Bazhin A,Pestrenin L,Rodionova L,Naletov I,Varlamova A,Belotsky V,Starikova E

Affiliations (5)

  • Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Department of Health, Moscow, Russia.
  • National Medical and Surgical Center Named After N.I. Pirogov of the Ministry of Health of the Russian Federation, Moscow, Russia.
  • I.M. Sechenov First Moscow State Medical University (Sechenov University), Moscow, Russia.
  • MIREA - Russian Technological University, Moscow, Russia.
  • Third Opinion Platform, Moscow, Russia.

Abstract

The integration of artificial intelligence (AI) into mammography holds significant potential for addressing the increasing workload and radiologist burnout, yet its widespread clinical adoption is hindered by critical limitations in current validation practices. Existing frameworks frequently fail to account for AI's dynamic evolution through retraining, data heterogeneity, and real-world deployment within healthcare systems such as compulsory medical insurance (CMI). This study aims to ensure continuous quality control of a mammography AI solution during its implementation in the CMI system by applying a novel lifecycle-based testing and monitoring methodology that addresses these specific gaps. The observational study incorporated retrospective functional and calibration testing, alongside prospective technical and clinical monitoring, interspersed with AI system updates. Anonymized digital mammograms from women aged ≥18 years underwent analysis. Prospective monitoring included all mammograms from participating sites (no exclusion criteria), capturing continuous real-world clinical data. The mammography AI system utilized U-Net++ and Mask2Former architectures, trained on ~4,000 mammograms. Key metrics encompassed area under the curve (AUC), accuracy, sensitivity, specificity, technical defect rates, and clinical assessment scores. The test dataset comprised 404,502 mammograms from 206 medical organizations and three mammography equipment manufacturers. A total of 336 radiologists participated. The testing and monitoring period lasted 2 years and 5 months. Over this time, AUC increased by 10.8% (from 0.83 to 0.92), accuracy by 16.9% (from 0.77 to 0.90), sensitivity by 4.8% (from 0.84 to 0.88), and specificity by 30.0% (from 0.70 to 0.91). The average technical defect rate decreased by 26.7% (from 3.0% to 0.8%), and the clinical assessment score rose by 47.8% (from 54.38% to 80.36%). The study culminated in the integration of the AI system into the regional CMI program. A key limitation of this study is the relatively small retrospective calibration testing dataset (100 mammograms) and the lack of external validation on independent datasets from other regions or countries. Iterative testing with prospective real-world monitoring, interleaved developer updates, and radiologist feedback substantially enhanced mammography AI performance. This lifecycle testing methodology demonstrates feasibility for clinical integration and CMI program deployment, balancing rigorous validation with continuous improvement. Future work will focus on expanding the retrospective calibration testing dataset and scaling the approach to a national level within the CMI framework.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.