Dual adversarial attacks on Explainable Deep Learning in medical image classification.
Authors
Affiliations (2)
Affiliations (2)
- School of Computer Sciences, USM, 11700, Malaysia.
- School of Computer Sciences, USM, 11700, Malaysia. Electronic address: [email protected].
Abstract
Explainable Deep Learning (XDL) can enhance clinicians' trust in medical image classification by providing interpretable explanations alongside accurate predictions. However, recent studies have shown that XDL models remain vulnerable to subtle adversarial perturbations. Prior work has typically attacked either predictions or explanations in isolation, overlooking the coupled vulnerability that arises when both are targeted simultaneously. To address this gap, we propose dual adversarial attacks that simultaneously target model predictions and explanations, and introduce a novel evaluation metric, Attack Success Rate (ASR), which jointly accounts for misclassification and explanation distortion. We propose a dual adversarial attack framework comprising two strategies: (1) a Location Attack that redistributes attribution mass away from the original high-importance regions while inducing misclassification; and (2) a Top-k Attack that reduces attribution assigned to the most diagnostically significant features in the clean explanation, likewise driving prediction errors. Both attacks are generated via iterative gradient-based optimization under an ℓ<sub>∞</sub> constraint, producing imperceptible perturbations. Effectiveness is quantified by ASR, which measures the proportion of perturbed samples that both mislead the classifier and satisfy the explanation-distortion criterion. Experiments across three benchmark medical imaging datasets (Chest X-ray, Fundoscopy, Dermoscopy), three widely used deep learning models (ResNet50, DenseNet121, EfficientNet-V2), and three attribution methods (Integrated Gradients, GradCAM, ScoreCAM) show that our dual attacks consistently achieve high ASR, demonstrating reliable effectiveness and broad applicability in misclassification and explanation manipulation. These results highlight a coupled vulnerability of prediction and explanation in XDL-based medical imaging systems: small, imperceptible perturbations can simultaneously undermine classification accuracy and explanation reliability. This coupled risk underscores the urgent need for defenses that jointly safeguard both decision accuracy and interpretability, especially in safety-critical clinical applications.