PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 11, 2025ACM Transactions on Reconfigurable Technology and Systems2 citations

REATA: An Efficient Vision Transformer Accelerator Featuring a Resource-Optimized Attention Design on Versal ACAP

View Full Paper
WZWenbo ZhangYZYan ZhangYLYiqi Liu

Key Points

  • This research aims to enhance the performance of Vision Transformers on edge devices by addressing computational and memory bottlenecks.
  • Proposed a modular and adaptive architecture for Vision Transformers targeting the AMD Versal ACAP platform.
  • Introduced a resource-efficient attention computation module localized within AI Engine core clusters.
  • Developed a resource-aware multi-stage pipeline scheduling strategy for feed-forward networks.
  • Achieved 33.2 TOPS throughput at INT8 precision, outperforming EQ-ViT accelerator by 27%.
  • Maintained competitive efficiency of 510.6 GOPS/W in testing.

Abstract

Deploying Vision Transformers (ViTs) on edge devices poses significant challenges due to their high computational demands and memory access overheads, which severely hinder real-time inference efficiency. This paper proposes a modular and adaptive ViT acceleration architecture targeting the AMD Versal ACAP platform. By leveraging heterogeneous resource collaboration and fine-grained dataflow optimizations, the proposed design addresses performance bottlenecks effectively. We introduce a resource-efficient attention computation module that localizes self-attention operations within AI Engine (AIE) core clusters, thereby reducing inter-module communication and minimizing MAC resource usage. In parallel, a resource-aware multi-stage pipeline scheduling strategy dynamically partitions and parallelizes the computation-intensive feed-forward network (FFN), improving computation reuse and module-level coordination. The architecture integrates parameter tiling and a PLIO-based broadcasting mechanism to construct a decoupled compute-communication dataflow engine, alleviating memory bottlenecks. Experimental results on the Xilinx VCK5000 ACAP platform demonstrate that the proposed design achieves 33.2 TOPS throughput at INT8 precision—outperforming the state-of-the-art EQ-ViT accelerator by 27%—while maintaining a competitive efficiency of 510.6 GOPS/W. Scalability evaluations on ViT-Base and DeiT-Tiny confirm the design’s adaptability in edge scenarios, offering a resource-efficient and reconfigurable hardware paradigm for high-density Transformer inference.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/69401b1e2d562116f28f7750https://doi.org/10.1145/3779444
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design2023 · 118 citations
  2. 2An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT2024
  3. 3Co-optimized Vision Transformer Deployment on Edge Devices: Algorithm-Hardware-Compiler 3D Evolution2025
  4. 4FPGA-Based Vit Inference Accelerator Optimization2024 · 2 citations
  5. 5ME-ViT: A Single-Load Memory-Efficient FPGA Accelerator for Vision Transformers2024