Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
June 1, 2023

BiFormer: Vision Transformer with Bi-Level Routing Attention

View Full Paper
Ask AI
Bookmark
Share

Authors

LZLei ZhuXWXinjiang WangZKZhanghan Ke

Discussion

Loading...

Member takes

Overview

Benchmarking study demonstrates improved computational efficiency in dense visual prediction tasks, highlighting the benefits of query-adaptive sparse attention routing.

Key Points

  • To alleviate the severe computational and memory burdens of pairwise token interactions in vision transformers without relying on rigid, content-agnostic sparsity patterns.
  • Designed a dynamic sparse attention mechanism using bi-level routing to filter out irrelevant key-value pairs across coarse regions before applying fine-grained token-to-token attention.
  • Implemented the routing mechanism using hardware-friendly dense matrix multiplications to maximize operational speed on GPUs.
  • Constructed the BiFormer general vision transformer and evaluated its performance on benchmark tasks including image classification, object detection, and semantic segmentation.
  • Bi-level routing attention successfully directed computational focus to small subsets of task-relevant tokens while disregarding distracting background regions.
  • BiFormer achieved superior performance trade-offs with high computational efficiency and reduced memory consumption, particularly across dense visual prediction benchmarks.

Cite This Study

Zhu et al. (2023) studied this question.

synapsesocial.com/papers/69d72bfe5dca7d66cbbef135https://doi.org/10.1109/cvpr52729.2023.00995
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows2022 · 1,316 citations
  2. 2Mask R-CNN2017 · 29,812 citations
  3. 3Randaugment: Practical automated data augmentation with a reduced search space2020 · 3,051 citations
  4. 4PVT v2: Improved baselines with pyramid vision transformer2022 · 2,368 citations
  5. 5UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning2022 · 108 citations