Clinical Comparison of Two Versions of a Commercial Artificial Intelligence System for Identifying Candidate Locations of Cerebral Aneurysms: Reduction and Characterization of False-Positive Findings.
Authors
Affiliations (3)
Affiliations (3)
- Department of Neurosurgery, Graduate School of Biomedical and Health Sciences, Hiroshima University, 1-2-3 Kasumi, Minami-ku, Hiroshima 734-8551, Japan.
- Department of Neurosurgery, Shimane Prefectural Central Hospital, Izumo 693-8555, Japan.
- LPIXEL Inc., 1-6-1 Otemachi, Chiyoda-ku, Tokyo 100-0004, Japan.
Abstract
<b>Background/Objectives:</b> Despite high sensitivity in deep learning models for identifying candidate locations of unruptured cerebral aneurysms (UCAs), excessive false-positive (FP) findings and insufficient characterization of their anatomical distribution remain barriers to clinical integration. We evaluated a multi-stage deep learning framework for reducing anatomy-specific FPs in TOF-MRA-based UCA screening. <b>Methods:</b> This retrospective multicenter study analyzed TOF-MRA images from 404 scans across seven institutions. A baseline model (Model-A) was compared with an updated model (Model-B) incorporating a secondary cluster-filtering algorithm. Reference standards were established by expert reviewers. Aneurysm-wise sensitivity was compared using McNemar's test, FPs per case using the Wilcoxon signed-rank test, and FP distributions by vascular location. <b>Results:</b> Aneurysm-wise sensitivity was 94.3% (231/245) for Model-A and 93.9% (230/245) for Model-B, with 1.16 and 0.879 FPs per case, respectively, representing a significant 24% FP reduction (<i>p</i> < 0.001). Per-aneurysm sensitivity did not differ significantly (McNemar's test, <i>p</i> = 1.00). Model-B suppressed mimics in the internal carotid artery (ICA) and middle cerebral artery (MCA). Aneurysms were most frequent in the ICA (29%) and MCA (22%). Setting Model-A FPs as 100%, ICA and MCA accounted for 47.1% and 11.3% in Model-A and decreased to 34.5% and 7.7% in Model-B. In an exploratory regional analysis with Holm adjustment, FP reductions remained significant in the ICA (adjusted <i>p</i> < 0.001) and MCA (adjusted <i>p</i> = 0.003). <b>Conclusions:</b> A secondary cluster-filtering algorithm reduced FP candidate findings by 24% with similar sensitivity point estimates (paired difference -0.4 percentage points, 95% CI -3.1 to +2.1). Region-specific error analysis revealed the anatomical dependency of FP reduction. The impact on reading workflow and patient outcomes requires prospective evaluation.