PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 23, 20251 citationsOpen Access

Survey and Benchmarking of Large Language Models for RTL Code Generation: Techniques and Open Challenges

View Full Paper
ARArun RavindranAPA. C. PatraVBVahid Babaey

Key Points

  • Models achieved up to 89.74% on VerilogEval and 96.08% on RTLLM, surpassing previous pipelines.
  • The review covers 26 techniques including fine-tuning and multi-agent orchestration in RTL generation.
  • Detailed failure analysis identified systematic error modes such as FSM mis-sequencing and handshake drift.
  • Future research focuses on controlled specifications, open benchmarks, and modular assurance frameworks.

Abstract

Large language models (LLMs) are emerging as powerful tools for hardware design, with recent work exploring their ability to generate register-transfer level (RTL) code directly from natural-language specifications. This paper provides a survey and evaluation of LLM-based RTL generation. We review twenty-six published efforts, covering techniques such as fine-tuning, reinforcement learning, retrieval-augmented prompting, and multi-agent orchestration, and we analyze their contributions across eight methodological dimensions including debugging support, post-RTL metrics, and benchmark development. Building on this review, we experimentally evaluate frontier commercial models---GPT-4.1, GPT-4.1-mini, and Claude Sonnet 4---on the VerilogEval and RTLLM benchmarks under both single-shot and lightweight agentic settings. Results show that these models achieve up to 89.74% on VerilogEval and 96.08% on RTLLM, matching or exceeding prior domain-specific pipelines without specialized fine-tuning. Detailed failure analysis reveals systematic error modes, including FSM mis-sequencing, handshake drift, blocking vs. non-blocking misuse, and state-space oversimplification. Finally, we outline a forward-looking research roadmap toward natural-language-to-SoC design, emphasizing controlled specification schemas, open benchmarks and flows, PPA-in-the-loop feedback, and modular assurance frameworks. Together, this work provides both a critical synthesis of recent advances and a baseline evaluation of frontier LLMs, highlighting opportunities and challenges in moving toward AI-native electronic design automation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ravindran et al. (2025) studied this question.

synapsesocial.com/papers/68d4758931b076d99fa6d0b3https://doi.org/10.20944/preprints202509.1681.v1
Ask AI
Helpful
Bookmark
Share
View Full Paper