| ||||
| ||||
![]() Title:End-to-End Multimodal Transformers for Multi-Cohort Alzheimer's Classification Authors:Tanguy Vansnick, Maxime Gloesener, Otmane Amel, Vito Tota, Mathis Delehouzee, Laurence Ris and Saïd Mahmoudi Conference:IEEE CBMS 2026 Tags:Alzheimer’s disease, deep learning, MRI classification and Vision Transformer Abstract: Deep learning for Alzheimer’s disease (AD) classification from MRI has shown promise, yet recent reviews reveal that fewer than 20% of studies employ external validation, with data leakage (e.g., same subjects in train and test sets) inflating reported accuracies to 95-99%. Studies with rigorous subject-level splitting achieve 66-90%. We present an end-to-end multimodal deep learning framework that processes both 3D MRI scans and clinical tabular data. The framework combines a Vision Transformer (ViT) for imaging analysis with a Feature Tokenizer Transformer (FT-Transformer) for clinical features, fusing both modalities through bidirectional cross-attention. This multimodal fusion enables the model to leverage complementary information from structural brain imaging and patient clinical profiles. We validate on 6,065 subjects from three independent cohorts (ADNI, OASIS, NACC-SCAN) using strict subject-level 5-fold cross-validation. Our approach achieves 92.4% accuracy (AUC: 0.96) for distinguishing cognitively normal subjects from AD-trajectory patients (including MCI patients who later developed AD), and 93.3% on established AD cases. External validation across datasets reveals that single-cohort models suffer severe performance degradation (accuracy drops of 25-43 points), while our multi-cohort approach maintains robust generalization. End-to-End Multimodal Transformers for Multi-Cohort Alzheimer's Classification ![]() End-to-End Multimodal Transformers for Multi-Cohort Alzheimer's Classification | ||||
| Copyright © 2002 – 2026 EasyChair |
