TPA-ConvNeXt: Trigonometric Phase Attention for Robust Retinal Disease Classification Across Fundus and OCT.
Authors
Affiliations (3)
Affiliations (3)
- Department of Electrical and Computer Engineering, Faculty of Engineering, King Abdulaziz University, Jeddah 21589, Saudi Arabia.
- Department of Electrical-Electronics Engineering, Faculty of Technology, Firat University, 23100 Elazig, Türkiye.
- Department of Electrical-Electronics Engineering, Faculty of Engineering, Bingol University, 12000 Bingol, Türkiye.
Abstract
This study proposes TPA-ConvNeXt, a ConvNeXt-Small-based deep learning architecture for retinal image classification using a phase-based trigonometric attention mechanism. The proposed Trigonometric Phase Attention (TPA) block recalibrates feature maps through learnable phase and amplitude modulation derived from channel-wise and spatial context. In addition, a stable fine-tuning strategy combining layer-wise learning rate decay (LLRD), linear warmup, and cosine annealing is employed to adapt pretrained backbones to medical image data. The method was evaluated on both fundus and optical coherence tomography (OCT) datasets. On the HYAMD fundus dataset, 5-fold cross-validation yielded a mean Macro F1 of 0.8940 ± 0.0403 and an ROC-AUC of 0.9449. External fundus validation further demonstrated cross-dataset robustness, achieving 0.9312 Macro F1 and 0.9803 ROC-AUC when trained on HYAMD and tested on AMDNet23. In OCT experiments, the model achieved Macro F1 scores of 0.9790 on MAK1_OCT, 0.9387 on OCTDL, and 0.9989 on the cleaned OCT2017 benchmark test set after MD5 duplicate removal. Ablation results showed that warmup contributed most strongly to stable optimization, while the proposed TPA block provided a smaller but consistent performance gain. Additional statistical analysis across the five HYAMD folds showed that TPA-ConvNeXt achieved comparable performance to the ConvNeXt-Small baseline, with a small numerical accuracy difference that did not reach statistical significance. Grad-CAM visualizations indicated that the model focused on clinically relevant retinal regions, and complexity analysis showed that the full model required 51.82 M parameters and 17.41 GFLOPs, adding only 2.36 M parameters and 0.02 GFLOPs over the ConvNeXt-Small baseline. These findings suggest that TPA-ConvNeXt provides a robust and generalizable framework for retinal image classification across both fundus and OCT modalities.