PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 19, 2023IEEE Transactions on Pattern Analysis and Machine Intelligence146 citations

Dual Vision Transformer

View Full Paper
TYTing YaoYLYehao LiYPYingwei Pan

Key Points

  • This research aims to develop a new Transformer architecture that efficiently integrates global semantics for self-attention learning.
  • Proposed the Dual Vision Transformer (Dual-ViT) architecture.
  • Integrated a semantic pathway for global semantics and a pixel pathway for local details.
  • Demonstrated improved accuracy over state-of-the-art (SOTA) Transformer architectures with similar training complexity.
  • Dual-ViT achieved superior accuracy compared to SOTA Transformer models.
  • Maintained reduced computational complexity while enhancing self-attention learning.
  • Showed efficient compression of token vectors into global semantics.

Abstract

Recent advances have presented several strategies to mitigate the computations of self-attention mechanism with high-resolution inputs. Many of these works consider decomposing the global self-attention procedure over image patches into regional and local feature extraction procedures that each incurs a smaller computational complexity. Despite good efficiency, these approaches seldom explore the holistic interactions among all patches, and are thus difficult to fully capture the global semantics. In this paper, we propose a novel Transformer architecture that elegantly exploits the global semantics for self-attention learning, namely Dual Vision Transformer (Dual-ViT). The new architecture incorporates a critical semantic pathway that can more efficiently compress token vectors into global semantics with reduced order of complexity. Such compressed global semantics then serve as useful prior information in learning finer local pixel level details, through another constructed pixel pathway. The semantic pathway and pixel pathway are integrated together and are jointly trained, spreading the enhanced self-attention information in parallel through both of the pathways. Dual-ViT is henceforth able to capitalize on global semantics to boost self-attention learning without compromising much computational complexity. We empirically demonstrate that Dual-ViT provides superior accuracy than SOTA Transformer architectures with comparable training complexity.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yao et al. (2023) studied this question.

synapsesocial.com/papers/69deac554838c5c0bab0cb3ahttps://doi.org/10.1109/tpami.2023.3268446
Ask AI
Helpful
Bookmark
Share
View Full Paper