Predictive accuracy of a general-purpose artificial intelligence model for cut-out following proximal femoral nailing.
Authors
Affiliations (5)
Affiliations (5)
- Tel Aviv University, Tel Aviv, Israel. [email protected].
- Tel Aviv Sourasky Medical Center, Tel Aviv, Israel. [email protected].
- Tel Aviv University, Tel Aviv, Israel.
- Tel Aviv Sourasky Medical Center, Tel Aviv, Israel.
- Hillel Yaffe Medical Center, Hadera, Israel.
Abstract
Cut-out is the most consequential mechanical complication after proximal femoral nailing and requires prompt recognition. General-purpose artificial intelligence (AI) models with image-interpretation capability are increasingly accessible, yet their diagnostic performance for this task is unknown. We evaluated ChatGPT's accuracy for predicting cut-out following proximal femoral nailing and compared TFNA and Gamma nail subgroups. In this retrospective predictive-accuracy study, 989 patients (683 TFNA, 306 Gamma nail; mean age 83.8 years; cut-out prevalence 2.5%) were analysed. For each case, three baseline radiographs - the injury anteroposterior and lateral views and the immediate postoperative control radiograph, all obtained before any complication could be radiographically present - together with the full clinical data set (excluding complication-related information) were submitted to ChatGPT, which returned a binary prediction (yes/no) and probability estimate (0-100%). The follow-up radiographs on which cut-out becomes apparent were not provided to the model. Ground truth was established on subsequent follow-up radiographs by a departmental review board requiring agreement of two senior surgeons. Accuracy metrics were calculated with Wilson 95% confidence intervals (CI); exact McNemar's and Mann-Whitney U tests assessed directional bias and calibration. Reporting followed STARD guidelines. ChatGPT showed a sensitivity of 68.0% (95% CI 48.4-82.8%), specificity of 62.7% (59.6-65.7%), PPV of 4.5% (2.8-7.1%), NPV of 98.7% (97.4-99.3%), and AUC of 0.694 (0.580-0.790). Performance was higher for TFNA (sensitivity 73.7%, AUC 0.730) than for Gamma nail (sensitivity 50.0%, AUC 0.606). The model over-predicted cut-out, with 360 false positives against 8 false negatives (45:1; McNemar's p < 0.001, Holm-corrected). Calibration was inverted: the median predicted probability was lower for correct (9%) than for incorrect (40%) classifications (Mann-Whitney p < 0.001). ChatGPT demonstrated limited, clinically insufficient accuracy for predicting cut-out following proximal femoral nailing, with a prohibitive false-positive burden and inverted calibration. Subgroup point estimates were lower for Gamma nail cases, but this exploratory difference did not reach statistical significance. General-purpose AI models are not currently suitable as a substitute for, or adjunct to, surgeon-led surveillance of this complication. Task-specific training, external validation, and defined clinical boundaries are required before such tools can be considered for fracture follow-up pathways.