Towards Explainability in Deep Learning for Detection of Five Major Intracranial Hemorrhage Subtypes on Head CT Using Multi-window DICOM Imaging and Patient-Level Cross-Validation.
Authors
Affiliations (9)
Affiliations (9)
- Sciences and Engineering of Biomedicals, Biophysics and Health Laboratory, Higher Institute of Health Sciences, Hassan First University, 26000, Settat, Morocco. [email protected].
- Higher Institute of Nursing Professions and Health Techniques, Rabat, Morocco. [email protected].
- LPHE-Modeling and Simulation, Faculty of Sciences, Mohammed V University, Rabat, Morocco.
- Sciences and Engineering of Biomedicals, Biophysics and Health Laboratory, Higher Institute of Health Sciences, Hassan First University, 26000, Settat, Morocco.
- Department of Radiotherapy, International Clinic of Settat, Settat, Morocco.
- Laboratory of Electronic Systems, Information Processing, Mechanics and Energetics, Faculty of Sciences, University Ibn Tofail Kenitra, Kenitra, Morocco.
- Department of Radiology, 3GCOM Company, Rabat, Morocco.
- Higher Institute of Nursing Professions and Health Techniques, Rabat, Morocco.
- Mohammed VI University of Sciences and Health, UM6SS, Casablanca, Morocco.
Abstract
Intracranial hemorrhage (ICH) is a time-critical neurologic emergency where delayed diagnosis can worsen outcomes. Non-contrast head CT is the first-line imaging modality, and explainable AI tools may enhance rapid detection and subtype classification. This study presents an explainable deep learning framework for multi-label ICH detection using DICOM-native multi-window CT data with leakage-aware validation. A retrospective secondary analysis was performed using the publicly available training dataset from the RSNA 2019 Intracranial Hemorrhage Detection Challenge, comprising 752,803 axial CT images from 18,938 patients and 21,744 studies. Images were processed directly from DICOM, converted to Hounsfield units, and encoded into three channels (brain, subdural, bone windows). Three independent ResNet34-based multi-label models were trained using weighted binary cross-entropy within a threefold patient- and study-level cross-validation framework. The fold-specific models were evaluated separately, without model merging or ensembling, and performance was summarized across folds. Performance was assessed using ROC AUC and PR AUC, along with threshold-dependent metrics (precision, sensitivity, F1 score, and specificity) and the Brier score after class-specific threshold optimization. Explainability was evaluated using Grad-CAM, Integrated Gradients, occlusion sensitivity, quantitative cross-method agreement, window ablation, and deletion-insertion faithfulness analyses. The model showed strong and consistent discrimination. For any hemorrhage, ROC AUC ranged from 0.973 to 0.974 and PR AUC from 0.894 to 0.899. Macro ROC AUC ranged from 0.968 to 0.971 and micro ROC AUC from 0.976 to 0.978. Mean macro precision, sensitivity, and F1 score were 0.640, 0.670, and 0.646, respectively, with a specificity of 0.984 and a Brier score of 0.054. Best F1 scores were observed for any hemorrhage (0.822), intraventricular (0.788), and intraparenchymal hemorrhage (0.767), while epidural hemorrhage remained low (0.203). Across 72 class-specific explanations from 56 CT images, spatial agreement among attribution methods was moderate, window ablation indicated substantial joint dependence on all three CT windows, and the deletion and insertion areas under the curve were 0.185 ± 0.147 and 0.882 ± 0.110, respectively, supporting a functional association between Grad-CAM-highlighted regions and model confidence. This DICOM-native, explainable approach achieved robust ICH detection, though performance for rare subtypes remains limited due to class imbalance.