PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 12, 2026Neurocomputing0 citationsOpen Access

Foundation models meet multimodal neuroimaging: A generative transformer-based framework for Alzheimer’s disease diagnosis

View Full Paper
LZLuca ZeddaALAndrea LoddoCRCecilia Di Ruberto

Key Points

  • This work aims to improve Alzheimer’s disease diagnosis by addressing challenges posed by incomplete neuroimaging data using a generative transformer framework.
  • Proposed a generative transformer-based framework for multimodal analysis with missing modalities.
  • Investigated multimodal performance using structural MRI, DTI, and PET data, incorporating synthetic image generation for training.
  • Employed vision transformers with Low-Rank Adaptation and integrated clinical variables through a projection module.
  • Achieved an improved multiclass classification F1-score with the transformer-based fusion head and generative augmentation.
  • Noticed that unimodal volumetric PET baselines outperformed in the best-case binary setting.
  • Found that the impact of generative augmentation varied, with some configurations improving performance while others degraded.

Abstract

Incomplete neuroimaging data remains a major challenge in Alzheimer’s disease diagnosis, as many patients undergo only a subset of recommended imaging protocols. This work addresses this limitation by proposing a generative transformer-based framework designed to support multimodal analysis in the presence of missing modalities. We systematically investigate multimodal performance and fairness within a unified foundation model framework for Alzheimer’s disease classification while introducing a generative approach that combines structural MRI, DTI, and PET data and leverages ControlNet-based diffusion models to synthesize anatomically consistent surrogate modalities when data are unavailable. These synthetic images are used exclusively as a training-time augmentation strategy for incomplete-modality settings, rather than as replacements for clinical acquisitions. Vision transformers adapted via Low-Rank Adaptation are employed for efficient feature extraction, while clinical variables are integrated through a dedicated projection module. Experimental results show that a transformer-based fusion head can improve over simple aggregation strategies in some complex multimodal settings, achieving an F1-score of in multiclass classification when combined with generative augmentation and clinical data. However, these benefits are not uniform since strong unimodal volumetric PET baselines remain superior in the best-case binary setting, and the effect of generative augmentation is strongly configuration-dependent, with some settings benefiting and others degrading substantially under non-selective synthetic augmentation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zedda et al. (2026) studied this question.

synapsesocial.com/papers/6a02c2fdce8c8c81e96404b6https://doi.org/10.1016/j.neucom.2026.133916
Ask AI
Helpful
Bookmark
Share
View Full Paper