Unfolding new horizons: Machine learning applications for pediatric intussusception - A systematic review and meta-analysis.
Authors
Affiliations (4)
Affiliations (4)
- Department of Medicine, Universidade de Ribeirão Preto (UNAERP), Ribeirão Preto, São Paulo, Brazil.
- School of medicine, University of Ioannina, Ioannina, Greece. Electronic address: [email protected].
- Department of Pediatrics, Hospital de Clínicas de Porto Alegre (HCPA), Porto Alegre, Rio Grande do Sul, Brazil.
- Department of Pediatric Surgery, McGill University Health Centre, Montreal, Quebec, Canada.
Abstract
Pediatric intussusception is among the most common surgical emergencies in early childhood. Ultrasonography is the diagnostic standard but remains operator-dependent, and no validated machine learning (ML) tools exist for triage or prognostic stratification. We systematically evaluated ML models for diagnosis and prognosis of pediatric intussusception. This PROSPERO-registered systematic review and meta-analysis (CRD420251073590) included studies evaluating ML models for pediatric intussusception diagnosis or prognosis reporting sensitivity, specificity, or AUC. Risk of bias was assessed with PROBAST. A bivariate random-effects model was applied to internally and externally validated datasets, with subgroup analyses by input modality. Eleven studies (37 model variants; 36,863 patients) were included. Internally validated ultrasound (US) models (k=4) achieved pooled sensitivity 0.914 (0.829-0.959) and specificity 0.980 (0.939-0.994). Externally validated US models (k=3) showed sensitivity 0.946 (0.909-0.969) and specificity 0.958 (0.918-0.979). Abdominal radiograph (AXR) triage models performed lower externally (sensitivity 0.800 [0.725-0.859]; specificity 0.748 [0.662-0.819]), consistent with a screening role. Notably, AI assistance improved junior readers' specificity by 15.9 percentage points overall (up to 37.6 points for the most junior readers), and AI-assisted junior sonographers cut ultrasound examination time by 61% without loss of diagnostic accuracy. Possible publication bias was detected (Deeks' p=0.021). Prognostic AUCs ranged from 0.555 to 0.962. US-based ML models demonstrate high, externally robust diagnostic accuracy. AXR models suit screening but need further validation. No study reported clinical endpoint data. Multicenter prospective validation outside Asia and regulatory pathway development are required before implementation. Until then, AXR-based models should be positioned as an adjunct that raises suspicion, never as a gate that determines which child receives an ultrasound. Level III - Systematic Review and Meta-Analysis of retrospective diagnostic accuracy studies.