TAD: Tail-Aware Diffusion Model for Long-Tailed Medical Image Classification.
Authors
Affiliations (2)
Affiliations (2)
- Henan Institute of Advanced Technology, Zhengzhou University, 97 Wenhua Road, Zhengzhou, 450003, Henan, China.
- School of Electrical and Information Engineering, Zhengzhou University, 100 Kexue Avenue, Zhengzhou, 450001, Henan, China. [email protected].
Abstract
In recent years, breakthroughs in deep learning have substantially improved the diagnostic performance of computer-aided diagnosis (CAD) systems. However, their reliability remains highly dependent on the quality and distribution of training data. One of the most critical challenges is the long-tailed distribution problem in medical images, which leads to a marked degradation in recognition performance for rare tail categories. Existing studies employ decoupling methods to more fairly attend to uncommon cases and significantly boost tail-class performance. Nevertheless, due to the inherent characteristics of medical datasets, namely, extremely scarce tail samples and a large number of categories, these decoupling approaches still suffer from deficiencies in feature learning, hyperparameter tuning, and classifier calibration. To address these issues, we propose a Tail-Aware Diffusion-Enhanced Long-tailed Medical Diagnosis model (TAD). TAD integrates diffusion-based tail-class augmentation, feature distribution regularization, and tail-aware feature compensation to alleviate the representation bias caused by long-tailed distributions. By combining complementary generative and feature-level strategies, the proposed framework enriches tail-class representations while preserving the discriminative structure of real medical images. Finally, extensive experiments on three public medical datasets demonstrate that our TAD model consistently outperforms state-of-the-art methods. On the Hyper-Kvasir dataset, TAD achieves an F1-score of <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>64.57</mn> <mo>±</mo> <mn>0.21</mn> <mo>%</mo></mrow> </math> and a kappa score of <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>94.23</mn> <mo>±</mo> <mn>0.18</mn> <mo>%</mo></mrow> </math> , outperforming the best competing method by 2.18% and 2.31%, respectively. Further ablation and feature-level analyses confirm the effectiveness of the proposed framework in improving tail-class representation and class discriminability.