PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 2, 2025Electronics2 citationsOpen Access

Systematic HLS Co-Design: Achieving Scalable and Fully-Pipelined NTT Acceleration on FPGAs

View Full Paper
JHJinfa HongBZBohao ZhangGMGaoyu Mao

Key Points

  • An efficient HLS co-design method significantly advances the performance of NTT accelerators on FPGAs.
  • The area-latency product shows performance improvements ranging from 1.93 to 191 times in HLS designs compared to existing methods.
  • Key techniques like memory partitioning and butterfly scheduling enable scalable parallel processing in NTT implementations.
  • Improvements in area-cycle product range from 1.7 to 10.6 times compared to HDL-based designs, highlighting the efficiency gains.

Abstract

Lattice-based cryptography (LBC) is an essential direction in the fields of homomorphic encryption (HE), zero-knowledge proofs (ZK), and post-quantum cryptography (PQC), while number theoretic transformations (NTT) are a performance bottleneck that affects the promotion and deployment of LBC applications. Field-programmable gate arrays (FPGAs) are an ideal platform for accelerating NTT due to their reconfigurability and parallel capabilities. High-level synthesis (HLS) can shorten the FPGA development cycle, but for algorithms such as NTT, the synthesizer struggles to handle the inherent memory dependencies, often resulting in suboptimal synthesis outcomes for direct designs. This paper proposes a systematic HLS co-design to progressively guide the synthesis of NTT accelerators. The approach integrates several key techniques: arithmetic module resource optimization, conflict-free butterfly scheduling, memory partitioning, and template-based automated design fusion. It reveals how to resolve pipeline bottlenecks in HLS-based designs and expand parallel processing, guiding microarchitecture iterations to achieve an efficient design space. Compared to existing HLS-based designs, the area-latency product achieves a performance improvement of 1.93 to 191 times, and compared to existing HDL-based designs, the area-cycle product achieves a performance improvement of 1.7 to 10.6 times.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hong et al. (2025) studied this question.

synapsesocial.com/papers/68de5d9c83cbc991d0a203b2https://doi.org/10.3390/electronics14193922
Ask AI
Helpful
Bookmark
Share
View Full Paper