INFORMER-Interpretability-Founded Monitoring of Medical Image Deep Learning Models: Application to Chest X-ray Pathologies.
Authors
Affiliations (5)
Affiliations (5)
- ARTORG Center for Biomedical Engineering Research, Bern, Switzerland. [email protected].
- Department of Diagnostic, Interventional, and Pediatric Radiology, Inselspital Bern, University of Bern, Bern, Switzerland.
- Department of Computer Science, Khalifa University, Abu Dhabi, United Arab Emirates.
- ARTORG Center for Biomedical Engineering Research, Bern, Switzerland. [email protected].
- Department of Radiation Oncology, Inselspital, Bern University Hospital and University of Bern, Bern, Switzerland. [email protected].
Abstract
Deep learning has demonstrated strong performance in medical imaging. However, its limited interpretability remains a major barrier to clinical trust and safe deployment. This limitation is particularly relevant in multi-label classification, where quality control methods are still underdeveloped and commonly rely only on model outputs, without incorporating gradient-level information that may better reflect prediction reliability. In this study, we propose a quality control framework for multi-label medical image classification that improves both reliability and interpretability. The framework includes a graph-based class-distinctiveness method that analyzes saliency-derived information to identify unreliable predictions, as well as a retrieval-based extension that provides case-based explanations for flagged outputs. The proposed methods were evaluated on the CheXpert dataset and compared with established output-based quality control approaches. Robustness was assessed using bootstrapped test sets, and differences in ranking across bootstrap samples were analyzed using the Wilcoxon signed-rank test. The proposed framework outperformed baseline methods, achieving a higher mean F1 score (0.574 vs. 0.563), while the best-performing variant showed higher sensitivity (0.752 vs. 0.700). In bootstrapped analyses, it achieved better mean ranks than the baseline approaches. Averaged across input-image noise levels of 0.001-0.005, under IxG-based evaluation, the proposed framework showed improvements in bootstrapped F1 over the baseline, with the retrieval-based variant achieving a 21.92% improvement and the corresponding non-retrieval variant achieving a 13.53% improvement. Clinicians further evaluated the retrieved examples to determine their relevance for interpreting flagged predictions. These findings indicate that gradient-level and graph-based analysis can enhance the effectiveness, transparency, and clinical applicability of quality control in multi-label medical image classification.