Assessing the accuracy of ChatGPT-5.6 in American College of Radiology Thyroid Imaging Reporting and Data System feature scoring and classification from synthetic free-text thyroid ultrasound reports.
Authors
Affiliations (2)
Affiliations (2)
- Department of Health Technology and Informatics, The Hong Kong Polytechnic University, Hong Kong.
- Ultrasound Department, EDAN Instruments, Inc., China.
Abstract
ObjectivesThe American College of Radiology Thyroid Imaging Reporting and Data System (ACR TI-RADS) provides a standardized framework for thyroid nodule risk stratification based on predefined ultrasound features. This study aimed to evaluate the accuracy and robustness of ChatGPT-5.6 Sol in extracting and scoring ACR TI-RADS features and classifying thyroid nodules from synthetic free-text ultrasound reports.MethodsA primary dataset of 100 synthetic thyroid ultrasound reports was constructed, with each report describing 1 nodule and the 5 ACR TI-RADS feature categories. ChatGPT-5.6 Sol was evaluated using a structured prompting strategy for feature identification, feature scoring, total score calculation, and final Thyroid Imaging Reporting and Data System classification. A less-directive prompt was also tested for total score calculation and final classification. Robustness was further assessed using a 30-case synthetic challenge set balanced across TR1 to TR5. Outputs were compared with reference annotations based on the 2017 ACR TI-RADS criteria. Accuracy was reported with Wilson 95% confidence intervals, and final-category agreement was assessed using Cohen's kappa.ResultsIn the primary dataset, accuracy for both feature identification and scoring was 100.0% for each ACR TI-RADS feature category (100/100; 95% confidence interval: 96.3%-100.0%). Total score calculation and final classification were also correct in all cases (100/100; 95% confidence interval: 96.3%-100.0%), with a Cohen's κ of 1.000. In the challenge set, feature identification, feature scoring, total score calculation, and final classification were correct in all 30 cases (30/30; 95% confidence interval: 88.6%-100.0%), with a Cohen's κ of 1.000. The less-directive prompt produced identical total score and classification accuracy in both datasets, with no discordant cases.ConclusionsChatGPT-5.6 Sol demonstrated high accuracy and consistent ACR TI-RADS-based classification across both synthetic datasets. Large language models may have potential for standardized Thyroid Imaging Reporting and Data System interpretation of free-text ultrasound reports; however, further validation in larger and more diverse settings is needed to establish generalizability.