Back to all papers

Large Language Models for Ankle Fracture Classification and Management Prediction from Routine Clinical Documentation: A Single-Center Exploratory Study.

September 18, 2026pubmed logopapers

Authors

Schwarberg B,Ketzer C,Thiel B,Sauter A,Makowski MR,Spitzl D,Mergen M,Gassert FT

Affiliations (4)

  • Department of Diagnostic and Interventional Radiology, TUM School of Medicine and Health, TUM University Hospital, Technical University of Munich, Rechts Der Isar, Munich, Germany. [email protected].
  • Department of Trauma Surgery, TUM School of Medicine and Health, TUM University Hospital, Technical University of Munich, Rechts Der Isar, Munich, Germany.
  • Department of Diagnostic and Interventional Radiology, TUM School of Medicine and Health, TUM University Hospital, Technical University of Munich, Rechts Der Isar, Munich, Germany.
  • Department of Gastroenterology, TUM School of Medicine and Health, TUM University Hospital, Technical University of Munich, Rechts Der Isar, Munich, Germany.

Abstract

Ankle fractures are among the most common injuries in trauma surgery and require accurate classification, consistent documentation, and individualized management. Large language models (LLMs) offer the potential to translate unstructured clinical text reports into structured, management-related information, yet their role in orthopedic trauma workflows remains insufficiently defined. In this retrospective study, the performance of Llama 3.1 (70B Instruct) was evaluated using routine radiology reports and clinical documentation from 54 patients with acute ankle fractures. Outputs were compared with reference standards derived from routine clinical documentation for Weber fracture classification, operative versus nonoperative management, and surgical procedure category. Four prompting strategies were systematically assessed. Accuracy for Weber classification ranged from 0.759 to 0.815, well above majority-class baseline (0.556), with strong performance for Weber B fractures and most errors occurring between adjacent categories. Macro-averaged F1 across the three Weber classes ranged from 0.744 to 0.822. Operative management prediction reached accuracies of 0.833 to 0.963 (sensitivity 0.872 to 1.000, specificity 0.571 to 0.714). Procedure category prediction demonstrated lower accuracy (0.426 to 0.610 depending on the endpoint), largely at or below a majority-class baseline (0.553), reflecting the limitations of text-only input for operative planning. These findings suggest that LLMs can extract structured information from routine clinical documentation, performing well for standardized classification but less reliably for procedure-level prediction. Prospective and multimodal validation is needed before clinical implementation.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.