From Volumetrics to 3D Tensors: A Multi-Cohort Evaluation of Machine Learning and Deep Learning for Alzheimer's Classification.
Authors
Affiliations (2)
Affiliations (2)
- Department of Electrical Engineering, Canadian University Dubai, Dubai, UAE. [email protected].
- Department of Electrical Engineering, Canadian University Dubai, Dubai, UAE.
Abstract
Automated Alzheimer's disease (AD) classification from structural MRI typically employs either feature-engineered machine learning (ML) or end-to-end 3D deep learning (DL). Addressing a lack of rigorous statistical comparison and multi-dataset validation in current literature, this study evaluates classical ML algorithms utilizing volumetric and voxel-based morphometry (VBM) features against 3D CNNs processing raw MRI tensors. Using the ADNI and OASIS datasets, models underwent internal 4-fold cross-validation and zero-shot cross-cohort testing to assess predictive performance, domain shift resilience, and clinical sensitivity. Welch's t-tests confirmed that 16 of 19 macro-regional brain aspects differed significantly between AD and cognitively normal subjects ( <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow><msub><mi>p</mi> <mtext>FDR</mtext></msub> <mo><</mo> <mn>0.05</mn></mrow> </math> ), led by the medial-temporal lobe. While DL architectures achieved marginal numerical superiority in peak accuracy on ADNI (DenseNet121: accuracy <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>0.92</mn> <mo>±</mo> <mn>0.02</mn></mrow> </math> , F1 <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>0.85</mn> <mo>±</mo> <mn>0.04</mn></mrow> </math> , ROC-AUC <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>0.96</mn> <mo>±</mo> <mn>0.02</mn></mrow> </math> ), they exhibited higher variance and lacked statistical significance compared to optimized ML baselines (XGBoost-VOL: F1 <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>0.82</mn> <mo>±</mo> <mn>0.03</mn></mrow> </math> ; <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>p</mi> <mo>≥</mo> <mn>0.28</mn></mrow> </math> for all 3D CNNs). On the highly imbalanced OASIS dataset, VBM-enhanced Logistic Regression matched top DL models in overall F1-score ( <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>0.68</mn> <mo>±</mo> <mn>0.05</mn></mrow> </math> ) while delivering statistically superior AD recall ( <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mn>0.67</mn> <mo>±</mo> <mn>0.04</mn></mrow> </math> ; <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>p</mi> <mo>=</mo> <mn>0.0026</mn></mrow> </math> versus baseline), exceeding every 3D CNN. Although both paradigms generalized robustly across cohorts (only a 2-4% zero-shot F1 reduction; ROC-AUC up to 0.9506), the findings highlight a crucial clinical trade-off: 3D CNNs autonomously extract complex spatial features, but simpler, interpretable ML models provide superior inferential stability and diagnostic sensitivity, making them highly viable for real-world deployment.