PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 2, 20250 citationsOpen Access

BitParticle: Partializing Sparse Dual-Factors to Build Quasi-Synchronizing MAC Arrays for Energy-efficient DNNs

View Full Paper
FQFeilong QiaoyuanJWJihe WangZSZ. Sun

Key Points

  • The proposed MAC unit achieves a 29.2% improvement in area efficiency while maintaining energy performance.
  • By addressing dual-factor sparsity, the design effectively minimizes the impact of partial product explosion.
  • A quasi-synchronous scheme enhances MAC array utilization by adding cycle-level elasticity and reducing pipeline stalls.
  • The approximate version of the MAC unit design further improves energy efficiency by 7.5% over its exact counterpart.

Abstract

Bit-level sparsity in quantized deep neural networks (DNNs) offers significant potential for optimizing Multiply-Accumulate (MAC) operations. However, two key challenges still limit its practical exploitation. First, conventional bit-serial approaches cannot simultaneously leverage the sparsity of both factors, leading to a complete waste of one factor' s sparsity. Methods designed to exploit dual-factor sparsity are still in the early stages of exploration, facing the challenge of partial product explosion. Second, the fluctuation of bit-level sparsity leads to variable cycle counts for MAC operations. Existing synchronous scheduling schemes that are suitable for dual-factor sparsity exhibit poor flexibility and still result in significant underutilization of MAC units. To address the first challenge, this study proposes a MAC unit that leverages dual-factor sparsity through the emerging particlization-based approach. The proposed design addresses the issue of partial product explosion through simple control logic, resulting in a more area- and energy-efficient MAC unit. In addition, by discarding less significant intermediate results, the design allows for further hardware simplification at the cost of minor accuracy loss. To address the second challenge, a quasi-synchronous scheme is introduced that adds cycle-level elasticity to the MAC array, reducing pipeline stalls and thereby improving MAC unit utilization. Evaluation results show that the exact version of the proposed MAC array architecture achieves a 29.2% improvement in area efficiency compared to the state-of-the-art bit-sparsity-driven architecture, while maintaining comparable energy efficiency. The approximate variant further improves energy efficiency by 7.5%, compared to the exact version. Index-Terms: DNN acceleration, Bit-level sparsity, MAC unit

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Qiaoyuan et al. (2025) studied this question.

synapsesocial.com/papers/68de5da783cbc991d0a20cb6https://doi.org/10.48550/arxiv.2507.09780
Ask AI
Helpful
Bookmark
Share
View Full Paper