Calibration and Reliability of Deep Learning in Medical Imaging Classification: Impact of Image Resolution.
Authors
Affiliations (2)
Affiliations (2)
- Department of Population and Quantitative Health Sciences, School of Medicine, Case Western Reserve University, Cleveland, OH, USA. [email protected].
- Department of Artificial Intelligence, Machine Learning and Data Science, Marwadi University, Rajkot, Gujarat, India.
Abstract
In medical imaging, deep learning models are commonly evaluated based on predictive accuracy, while the reliability of their probabilistic outputs remains less explored. In particular, the role of image resolution in influencing model calibration has not been systematically studied. This work aims to investigate the relationship between image resolution and calibration in deep learning-based medical image classification. We conduct a comprehensive empirical study using multiple MedMNIST datasets across varying image resolutions. Model performance is evaluated using standard classification metrics alongside multiple calibration measures, including expected calibration error (ECE). We further analyze the relationship between resolution changes and predictive uncertainty. Our findings indicate that variations in image resolution can significantly influence model calibration. For one of the three imaging modalities evaluated, reduced resolution was associated with a statistically significant increase in miscalibration even as classification accuracy remained relatively stable, while calibration remained stable for the other two modalities. These results suggest that accuracy alone may not fully capture model reliability under varying input conditions and that resolution sensitivity is task dependent. This study highlights the importance of considering image resolution as a key factor in evaluating the reliability of deep learning models for medical imaging classification. Incorporating calibration-aware evaluation can lead to more trustworthy and robust medical AI systems.