Back to all papers

Predictive accuracy of a general-purpose artificial intelligence model for cut-out following proximal femoral nailing.

August 1, 2026pubmed logopapers

Authors

Ben Arie G,Warschawski Y,Amzallag N,Gan-Or H,Ben-Tov T,Khoury A,Horev I,Graif N,Morag G

Affiliations (5)

  • Tel Aviv University, Tel Aviv, Israel. [email protected].
  • Tel Aviv Sourasky Medical Center, Tel Aviv, Israel. [email protected].
  • Tel Aviv University, Tel Aviv, Israel.
  • Tel Aviv Sourasky Medical Center, Tel Aviv, Israel.
  • Hillel Yaffe Medical Center, Hadera, Israel.

Abstract

Cut-out is the most consequential mechanical complication after proximal femoral nailing and requires prompt recognition. General-purpose artificial intelligence (AI) models with image-interpretation capability are increasingly accessible, yet their diagnostic performance for this task is unknown. We evaluated ChatGPT's accuracy for predicting cut-out following proximal femoral nailing and compared TFNA and Gamma nail subgroups. In this retrospective predictive-accuracy study, 989 patients (683 TFNA, 306 Gamma nail; mean age 83.8 years; cut-out prevalence 2.5%) were analysed. For each case, three baseline radiographs - the injury anteroposterior and lateral views and the immediate postoperative control radiograph, all obtained before any complication could be radiographically present - together with the full clinical data set (excluding complication-related information) were submitted to ChatGPT, which returned a binary prediction (yes/no) and probability estimate (0-100%). The follow-up radiographs on which cut-out becomes apparent were not provided to the model. Ground truth was established on subsequent follow-up radiographs by a departmental review board requiring agreement of two senior surgeons. Accuracy metrics were calculated with Wilson 95% confidence intervals (CI); exact McNemar's and Mann-Whitney U tests assessed directional bias and calibration. Reporting followed STARD guidelines. ChatGPT showed a sensitivity of 68.0% (95% CI 48.4-82.8%), specificity of 62.7% (59.6-65.7%), PPV of 4.5% (2.8-7.1%), NPV of 98.7% (97.4-99.3%), and AUC of 0.694 (0.580-0.790). Performance was higher for TFNA (sensitivity 73.7%, AUC 0.730) than for Gamma nail (sensitivity 50.0%, AUC 0.606). The model over-predicted cut-out, with 360 false positives against 8 false negatives (45:1; McNemar's p < 0.001, Holm-corrected). Calibration was inverted: the median predicted probability was lower for correct (9%) than for incorrect (40%) classifications (Mann-Whitney p < 0.001). ChatGPT demonstrated limited, clinically insufficient accuracy for predicting cut-out following proximal femoral nailing, with a prohibitive false-positive burden and inverted calibration. Subgroup point estimates were lower for Gamma nail cases, but this exploratory difference did not reach statistical significance. General-purpose AI models are not currently suitable as a substitute for, or adjunct to, surgeon-led surveillance of this complication. Task-specific training, external validation, and defined clinical boundaries are required before such tools can be considered for fracture follow-up pathways.

Topics

Bone NailsArtificial IntelligenceFracture Fixation, IntramedullaryPostoperative ComplicationsJournal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.