Compute-in-memory (CIM) is being widely explored to minimize power consumption related to data movement and multiply-and-accumulate (MAC) operations for AI edge devices. Compared to analog based CIMs, digital-based CIMs (DCIM), which include small, distributed SRAM banks and a customized MAC unit, realize massively parallel computation with no accuracy loss and better power-performance-area (PPA) scaling with advanced technologies. However, balancing operating efficiency per area (TOPS/mm 2 ) and bit density (Mb/mm 2 ) is one of the challenges in prior DCIMs because of low operating throughput caused by bit-serial input and a small number of rows. In this paper, we introduce a 3nm SRAM-based DCIM macro, which is based on foundry 6T SRAM bit cells and a parallel-MAC scheme to improve bit cell density and operating throughput. The DCIM macro is implemented with CIM BIST in a test chip, and confirms ultra-low voltage MAC operation down to 360mV, and 1.5GHz operation at 0.9V. The DCIM achieves 32.5TOPS/W (assuming a 25% input toggle rate and a 50% weight = 1 distribution), 55.0TOPS/mm 2 and 3.78 Mb/mm 2 .
No takes yet. Share an insight, caveat, or question.
Fujiwara et al. (2024) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: