Comparison of Machine Learning Models for Predicting Recurrent Lumbar Disc Herniation After Percutaneous Endoscopic Lumbar Discectomy.
Authors
Affiliations (3)
Affiliations (3)
- Department of Anesthesiology, Peking University Third Hospital, Beijing 100191, China.
- Department of Anesthesiology, Tibet Autonomous Region People's Hospital, Lhasa 850002, China.
- Department of Pain Medicine, Peking University Third Hospital, Beijing 100191, China.
Abstract
<b>Background</b>: Recurrent lumbar disc herniation (rLDH) significantly impairs outcomes following percutaneous endoscopic lumbar discectomy (PELD). Accurate individualized risk prediction remains challenging. This study aimed to develop and compare multiple machine learning models for predicting rLDH within two years post-surgery. <b>Methods</b>: A retrospective cohort of 1483 patients undergoing single-level PELD was analyzed. The primary outcome was symptomatic, magnetic resonance imaging-confirmed rLDH requiring reintervention. Candidate predictors included demographic, surgical, and radiographic parameters. The dataset was stratified by outcome and randomly split into training (70%, <i>n</i> = 1038) and validation (30%, <i>n</i> = 445) sets. Feature selection utilized univariate screening (<i>p</i> < 0.2) and least absolute shrinkage and selection operator regression. Six machine learning algorithms were trained and optimized via grid search. Performance was evaluated using the area under the receiver operating characteristic curve (AUC), F1-score, Brier score, calibration and decision curve analysis (DCA). Model interpretability was assessed using SHapley Additive exPlanations (SHAP). <b>Results</b>: The overall recurrence rate was 4.25% (63/1483). Logistic regression achieved the optimal F1-score (0.286), while light Gradient Boosting Machine (LightGBM) demonstrated superior discrimination (AUC = 0.768). DCA indicated clinical utility primarily at low threshold probabilities (<10%). SHAP analysis identified increased sagittal range of motion as the strongest risk factor, followed by reduced facet orientation, advanced age, type II Modic changes, and Michigan State University zone C. <b>Conclusions</b>: This study presents an exploratory, internally validated machine learning framework for rLDH risk stratification. While LightGBM demonstrated moderate discriminative ability, model sensitivity was constrained by the inherent rarity of recurrence events, precluding its use as a definitive standalone screening tool. Notably, clinical utility was restricted to low threshold probabilities (<10%), supporting a focused role in identifying high-risk subgroups for intensified preoperative counseling and postoperative monitoring. Beyond elucidating key radiological and demographic risk factors, our findings underscore that rigorous external prospective validation and probability calibration are indispensable before any future clinical deployment.