PulseTrendingJournal ClubResearchersJournalsExplore
Instagram
HomeTrendingJournal ClubExplore
Synapse
⌘+K
Synapse
February 12, 2026Applied SciencesOpen Access

CGA-ViT: Channel-Guided Additive Attention for Efficient Vision Recognition

View Full Paper
Ask AI
Bookmark
Share

Authors

YZYayue ZhaoJMJingli MiaoZLZhenping Li

Discussion

Loading...

Member takes

Overview

Introduces CGA-ViT, which enhances vision recognition performance while maintaining computational efficiency in high-resolution tasks.

Key Points

  • The aim is to optimize multi-scale feature extraction and global context modeling in vision transformers while reducing computational complexity.
  • Developed the channel-guided additive attention (CGA) mechanism for long-range semantic interactions.
  • Implemented multi-scale dilated feature embedding (MDFE) for enhanced feature capturing.
  • Adopted a hierarchical structure combining local-global interactions in shallow layers with efficient attention in deep layers.
  • Evaluated performance on ImageNet-1K, comparing with existing models like Swin-T and ConvNeXt-T.
  • CGA-ViT achieved 84.0% Top-1 accuracy with only 4.7 GFLOPs.
  • Outperformed Swin-T (81.3%) and ConvNeXt-T (82.1%) by 2.7 and 1.9 percentage points respectively.
  • MDFE and CGA contributed 65.0% to the performance gains, with additional benefits from token-level supervision.

Cite This Study

Zhao et al. (2026) studied this question.

synapsesocial.com/papers/698d6ebb5be6419ac0d54702https://doi.org/10.3390/app16041740
View Full Paper
Ask AI
Bookmark
Share