Evaluating Clinical NLP Services for Chest Radiograph Report Labeling: A Comparative Study on an Independent Pediatric Dataset.
Authors
Affiliations (3)
Affiliations (3)
- Cincinnati Children's Hospital Medical Center, Cincinnati, OH, USA. [email protected].
- Texas A&M University, College Station, TX, USA.
- Cincinnati Children's Hospital Medical Center, Cincinnati, OH, USA.
Abstract
General-purpose clinical NLP tools are increasingly used to automatically label clinical reports for research and quality improvement. However, independent evaluations for specific tasks like labeling pediatric CXR reports remain limited. This study compares four commercial clinical NLP models: AWS, AZ, GC, SP, for entity extraction and assertion detection of clinically relevant findings in pediatric CXR reports. Two dedicated CXR report labelers, CheXpert and CheXbert, were also evaluated for comparison. A total of 95,008 pediatric CXR reports from a large academic hospital were analyzed. Entities and their assertion statuses (positive, negative, uncertain) were extracted from findings and impression sections using the four general-purpose models. Entities from impressions were mapped to 12 disease categories plus No Findings using regular expressions. CheXpert and CheXbert processed the same reports for the same 13 labels. Outputs from all six systems were compared across assertion categories using Fleiss' Kappa to quantify inter-model agreement. A manual review was done by a board-certified pediatric radiologist on a subset of 360 exams stratified across the 12 CheXpert labels and across the 3 assertion statuses (positive, negative, and uncertain). Precision, recall, and F1-scores were calculated for each NLP service against the ground truth. Significant differences were observed in the mean number of extracted entities among the NLP models (p < 0.001). SP extracted the most unique entities (49,688), followed by AZ (31,543), AWS (27,216), and GC (16,477). Assertion distributions also differed significantly (p < 0.001). For the CheXpert labels comparison, the mean Fleiss' Kappa across all assertion categories was 0.68 ± 0.17, including all reports, and 0.35 ± 0.14, excluding reports where all six models predicted the disease as absent. For manual validation on the subset, CheXbert and CheXpert achieved the highest macro F1-scores (0.565 and 0.561), followed by Google (0.453), Azure (0.421), Spark NLP (0.420), and AWS (0.266). Overall, positive assertions were best detected (F1 0.782) and uncertain the worst (0.468). Performance varied by label, with the highest F1 for pleural effusion, pneumothorax, and pneumonia and the lowest for lung opacity, enlarged cardiomediastinum, and lung lesion. Marked variability exists among NLP models in entity extraction and uncertainty handling, underscoring the need for further validation before clinical deployment.