PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 22, 2024122 citations

AttAcc! Unleashing the Power of PIM for Batched Transformer-based Generative Model Inference

View Full Paper
JPJaehyun ParkSeoul National UniversityJCJaewan ChoiChungbuk National UniversityKKKwanhee KyungSeoul National University

Key Points

Key points are not available for this paper at this time.

Abstract

The Transformer-based generative model (TbGM), comprising summarization (Sum) and generation (Gen) stages, has demonstrated unprecedented generative performance across a wide range of applications. However, it also demands immense amounts of compute and memory resources. Especially, the Gen stages, consisting of the attention and fully-connected (FC) layers, dominate the overall execution time. Meanwhile, we reveal that the conventional system with GPUs used for TbGM inference cannot efficiently execute the attention layer, even with batching, due to various constraints. To address this inefficiency, we first propose AttAcc, a processing-in-memory (PIM) architecture for efficient execution of the attention layer. Subsequently, for the end-to-end acceleration of TbGM inference, we propose a novel heterogeneous system architecture and optimizations that strategically use xPU and PIM together. It leverages the high memory bandwidth of AttAcc for the attention layer and the powerful compute capability of the conventional system for the FC layer. Lastly, we demonstrate that our GPU-PIM system outperforms the conventional system with the same memory capacity, improving performance and energy efficiency of running a 175B TbGM by up to 2.81× and 2.67×, respectively.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Park et al. (2024) studied this question.

synapsesocial.com/papers/68e6e0a4b6db64358765cad7https://doi.org/10.1145/3620665.3640422
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Unleashing the Potential of PIM: Accelerating Large Batched Inference of Transformer-Based Generative Models2024 · 3 citations
  2. 2PIM GPT a hybrid process in memory accelerator for autoregressive transformers2024 · 30 citations
  3. 3PIM-GPT: A Hybrid Process-in-Memory Accelerator for Autoregressive Transformers2024
  4. 4NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing2024 · 117 citations
  5. 5NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing2024