Automatic Patient Eligibility for Photon-counting CT Using Discriminative and Generative AI Models in Neuroradiology.
Authors
Affiliations (4)
Affiliations (4)
- MRI unit, Radiology Department, HT Medica, Jaén, Spain (T.M.N., A.L.). Electronic address: [email protected].
- NLP Unit, HT Medica, Jaén, Spain (P.L.U.). Electronic address: [email protected].
- Neurorradiología Diagnostica e Intervencionista, HT Médica, Ávila, Spain (J.E.). Electronic address: [email protected].
- MRI unit, Radiology Department, HT Medica, Jaén, Spain (T.M.N., A.L.). Electronic address: [email protected].
Abstract
Photon -Counting CT (PCCT) provides higher resolution, reduced dose, and valuable spectral data, but generates a large number of images per study, underlining radiologist and Picture Archiving and Communication Systems (PACS) capacity. While no guidelines defining when PCCT should be preferred over conventional Energy Integrating Detector (EID) CT in neuroradiology, and manual decision-making being impractical, Natural Language Processing (NLP)-based large language models (LLMs) may help automate routing of CT requests to the most appropriate CT technology. We conducted a retrospective study using Spanish-language neuroradiology CT requests retrieved from the Radiology Information System (RIS) (January 2012-October 2025). A random sample of 800 requests was independently labeled by two neuroradiologists as Basic Protocol (BP); feasible on EID-CT or basic PCCT) or Advanced Protocol (AP); requiring full-advanced PCCT), considering the gold standard an expert consensus (Cohen's k = 0.74). For discriminative modeling, transformer classifiers (Spanish bidirectional encoder representations from transformers (BERT), biomedical RoBERTa, Llama-3.1-8B, and Mistral-7B) were fine-tuned for binary BP/AP classification and evaluated with stratified five-fold cross-validation. For generative modeling, ChatGPT 5.2 and Gemini 3 were applied in a zero-shot setting using structured prompting and post-processed into binary outputs. Performance was assessed using accuracy, precision, recall, specificity, F1-score, and AUC. Discriminative transformers achieved the best performance for the classification, with RoBERTa leading (accuracy 0.925, precision 0.906, recall 0.967, F1 0.935, AUC 0.920) and the lowest total error, without a significant difference versus Spanish BERT. In contrast, RoBERTa outperformed Llama, Mistral, Gemini, and ChatGPT with statistically significant improvements, while the generative LLMs yielded the poorest overall performance. Discriminative transformer models, particularly RoBERTa, enabled highly accurate automatic identification of CT requests requiring AP to be performed at PCCT, substantially outperforming general-purpose generative LLMs. These findings support this exploratory proof-of-concept study of AI-driven eligibility routing to optimize PCCT utilization and streamline neuroradiology workflows.