Machine learning for prediction of large-for-gestational-age neonate with adverse outcome at routine ultrasound examination at 35-37 weeks.
Authors
Affiliations (6)
Affiliations (6)
- Fetal Medicine, St George's University Hospitals NHS Foundation Trust, University of London, London, UK.
- Perinatology Department, Ministry of Health, Bingöl State Hospital, Bingöl, Turkey.
- Tel Aviv University Faculty of Medicine, Tel Aviv, Israel.
- Sheba Medical Center, Tel Hashomer, Tel Aviv, Israel.
- Vascular Biology Research Centre, Molecular and Clinical Sciences Research Institute, City St George's University of London, London, UK.
- Fetal Medicine Unit, Liverpool Women's Hospital, Liverpool, UK.
Abstract
To evaluate the performance of a machine-learning (ML) model compared with traditional logistic regression models for predicting a large-for-gestational-age (LGA) neonate at term and associated adverse perinatal outcomes, using sonographic biometric and Doppler parameters routinely assessed at a late third-trimester scan in combination with maternal demographic characteristics. This cohort study involved a retrospective data analysis of singleton pregnancies that underwent a routine third-trimester ultrasound examination at 35 + 0 to 37 + 6 weeks' gestation and were delivered in a single tertiary referral center between December 2019 and December 2025. Pregnancies with major fetal structural anomalies, known genetic abnormalities or missing outcome data were excluded. We collected maternal demographic and obstetric characteristics, first-trimester biochemical screening results, third-trimester fetal biometry and Doppler parameters, and birth and neonatal outcome data. Both traditional statistical and contemporary ML models were developed to predict a LGA neonate (birth weight > 90<sup>th</sup> centile), as well as LGA complicated by one or more predefined adverse perinatal outcomes. Pregnancies with LGA complicated by adverse maternal outcome and those with adverse neonatal outcome were also analyzed separately to evaluate outcome-specific predictors and improve clinical interpretability. Logistic regression was used to identify predictors of each outcome with model discrimination assessed by the area under the receiver-operating-characteristics curve (AUC). In parallel, ML models (random forest, eXtreme Gradient Boosting (XGBoost), elastic net and support vector machine) were applied to capture potential non-linear relationships and interactions, providing complementary modeling approaches across different data structures. The study included 21 743 singleton pregnancies, with a median gestational age of 36 + 3 (interquartile range (IQR), 36 + 2 to 36 + 5) weeks at third-trimester ultrasound examination and 39 + 4 (IQR, 39 + 0 to 40 + 4) weeks at delivery, and a median birth-weight centile of 58.8 (IQR, 32.2-80.9). There were 3007 (13.8%) LGA neonates and 689 (3.2%) cases of LGA neonate with adverse outcome. The best-performing logistic regression and ML models demonstrated high discriminatory performance for predicting a LGA neonate (AUC, 0.888 (95% CI, 0.874-0.901) for both logistic regression and ML models) as well as LGA neonate with adverse perinatal outcome (AUC, 0.859 (95% CI, 0.831-0.884) for logistic regression and AUC, 0.859 (95% CI, 0.831-0.883) for ML model). Across both traditional logistic regression and ML models, the estimated fetal weight (EFW) centile was the most robust single predictor across all outcomes, with AUCs that ranged from 0.809 to 0.878. Model-specific 95% CIs are provided. Notably, ML models did not improve predictive performance compared to logistic regression. SHapley Additive exPlanations (SHAP) analysis revealed that EFW and abdominal circumference centiles were the principal determinants of risk prediction in our cohort. Routinely collected third-trimester ultrasound data can identify pregnancies at risk of LGA and associated adverse outcomes in an unselected population. EFW was the dominant predictor in both traditional logistic regression and ML models, with minimal incremental value obtained from additional maternal or ultrasound variables. ML did not improve predictive performance over traditional logistic regression, indicating that fetal size alone captured most clinically relevant information for risk stratification. These findings support a simpler, more clinically actionable approach to risk stratification, shifting focus from identifying large fetuses to predicting clinically meaningful complications. © 2026 The Author(s). Ultrasound in Obstetrics & Gynecology published by John Wiley & Sons Ltd on behalf of International Society of Ultrasound in Obstetrics and Gynecology.