PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 11, 2026Communications Biology0 citationsOpen Access

Vision transformer autoencoders captures local and non-local features in brain imaging to reveal novel genetic associations

View Full Paper
SISamia R. IslamThe University of Texas Health Science Center at HoustonTXTian XiaThe University of Texas Health Science Center at HoustonWHWei HeThe University of Texas Health Science Center at Houston

Key Points

  • This research aims to link genetic variation to brain structure using advanced imaging analysis techniques.
  • Applied a Vision Transformer-based autoencoder to derive 128-dimensional representations from T1-weighted brain MRI scans.
  • Analyzed data from 6,130 UK Biobank participants to identify significant genetic variants.
  • Conducted genome-wide association studies with a total of 22,867 UK Biobank participants.
  • Identified 63 genetic loci, with 24 loci uniquely detected by the Vision Transformer-based method.
  • The model effectively captured both local and non-local anatomical patterns in brain MRI data.
  • Leveraged attention mechanisms and positional embeddings for feature interpretation.

Abstract

Abstract Linking genetic variation to human brain structure is a key step toward understanding the biological basis of cognition and disease. Progress in this area, however, has been limited by a major challenge: imaging features are often predefined, restricting the discovery of novel associations. Here, we present a framework that applies a Vision Transformer (ViT)-based autoencoder to derive 128-dimensional representations from T1-weighted brain MRI scans of 6,130 UK Biobank participants, which we call unsupervised learning derived image phenotypes from ViT (ViT-UDIP). These ViT-UDIP phenotypes are used in genome-wide association studies (GWAS) of 22,867 UK Biobank participants to identify significant genetic variants, which were further aggregated into genetic loci. The ViT-based approach uncovers a total of 63 loci and out of which 24 were not detected by the CNN-based method. Importantly, feature interpretation reveals that the model captured local as well as non-local anatomical patterns such as left-right hemisphere symmetry within brain MRI data by leveraging its attention mechanism and positional embeddings. This ability of capturing non-local patterns distinguishes the ViT from the previous CNN model. Together, these results demonstrate the value of transformer-based architectures in discovering novel and robust imaging phenotypes for genetic discovery.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Islam et al. (2026) studied this question.

synapsesocial.com/papers/6a2a526080c8f91e7f39e689https://doi.org/10.1038/s42003-026-10430-6
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1SABViT: A Pilot Feasibility Study of a Self-Attention-Based Vision Transformer for Binary Brain Tumor Detection in MRI2025
  2. 2Global context modeling with vision transformers for MRI-based classification of brain tumors.2026
  3. 3ViT-BT: Improving MRI Brain Tumor Classification Using Vision Transformer with Transfer Learning2024 · 4 citations
  4. 4Multi-Class Brain Tumor Diagnosis Using a Vision Transformer with MRI Image Segmentation2025 · 7 citations
  5. 5ADHD prediction from individual-space T1 images using a Vision Transformer with a gross-region grid framework2026