{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T00:09:27Z","timestamp":1784592567494,"version":"3.55.0"},"reference-count":37,"publisher":"Association for Computing Machinery (ACM)","issue":"6","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>The application of large language models (LLMs) to digital hardware code generation is an emerging field, with most LLMs primarily trained on natural language and software code. Hardware code like Verilog constitutes a small portion of training data, and few hardware benchmarks exist. The open-source VerilogEval benchmark, released in November 2023, provided a consistent evaluation framework for LLMs on code completion tasks. Since then, both commercial and open models have seen significant development.<\/jats:p>\n                  <jats:p>In this work, we evaluate new commercial and open models since VerilogEval\u2019s original release\u2014including GPT-4o, GPT-4 Turbo, Llama3.1 (8B\/70B\/405B), Llama3 70B, Mistral Large, DeepSeek Coder (33B and 6.7B), CodeGemma 7B, and RTL-Coder\u2014against an improved VerilogEval benchmark suite. We find measurable improvements in state-of-the-art models: GPT-4o achieves a 63% pass rate on specification-to-RTL tasks. The recently released and open Llama3.1 405B achieves a 58% pass rate, almost matching GPT-4o, while the smaller domain-specific RTL-Coder 6.7B models achieve an impressive 34% pass rate.<\/jats:p>\n                  <jats:p>Additionally, we enhance VerilogEval\u2019s infrastructure by automatically classifying failures, introducing in-context learning support, and extending the tasks to specification-to-RTL translation. We find that prompt engineering remains crucial for achieving good pass rates and varies widely with model and task. A benchmark infrastructure that allows for prompt engineering and failure analysis is essential for continued model development and deployment.<\/jats:p>","DOI":"10.1145\/3718088","type":"journal-article","created":{"date-parts":[[2025,2,19]],"date-time":"2025-02-19T06:19:52Z","timestamp":1739945992000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":30,"title":["Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code Generation"],"prefix":"10.1145","volume":"30","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6159-8964","authenticated-orcid":false,"given":"Nathaniel","family":"Pinckney","sequence":"first","affiliation":[{"name":"NVIDIA Corp","place":["Austin, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2835-667X","authenticated-orcid":false,"given":"Christopher","family":"Batten","sequence":"additional","affiliation":[{"name":"Cornell University","place":["Ithaca, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3488-9763","authenticated-orcid":false,"given":"Mingjie","family":"Liu","sequence":"additional","affiliation":[{"name":"NVIDIA Corp","place":["Austin, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1028-3860","authenticated-orcid":false,"given":"Haoxing","family":"Ren","sequence":"additional","affiliation":[{"name":"NVIDIA Corp","place":["Austin, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7584-3489","authenticated-orcid":false,"given":"Brucek","family":"Khailany","sequence":"additional","affiliation":[{"name":"NVIDIA Corp","place":["Austin, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,10,21]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"Au Large","author":"AI Mistral","year":"2024","unstructured":"Mistral AI. 2024. Au Large. https:\/\/mistral.ai\/news\/mistral-large\/Section: news."},{"key":"e_1_3_2_3_2","unstructured":"Ekin Aky\u00fcrek Dale Schuurmans Jacob Andreas Tengyu Ma and Denny Zhou. 2023. What Learning Algorithm Is In-context Learning? Investigations with Linear Models. arxiv:2211.15661 [cs.LG]. https:\/\/arxiv.org\/abs\/2211.15661"},{"key":"e_1_3_2_4_2","doi-asserted-by":"crossref","unstructured":"Ahmed Allam and Mohamed Shalan. 2024. RTL-Repo: A Benchmark for Evaluating LLMs on Large-scale RTL Design Projects. arxiv:2405.17378 [cs.LG]","DOI":"10.1109\/LAD62341.2024.10691810"},{"key":"e_1_3_2_5_2","unstructured":"Jacob Austin Augustus Odena Maxwell Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie Cai Michael Terry Quoc Le and Charles Sutton. 2021. Program Synthesis with Large Language Models. arxiv:2108.07732 [cs.PL]"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/mlcad58807.2023.10299874"},{"key":"e_1_3_2_7_2","unstructured":"Tom B. Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al.2020. Language Models Are Few-shot Learners. arxiv:2005.14165 [cs.CL]"},{"key":"e_1_3_2_8_2","article-title":"ChipGPT: How far are we from natural language hardware design","volume":"2305","author":"Chang Kaiyan","year":"2023","unstructured":"Kaiyan Chang, Ying Wang, Haimeng Ren, Mengdi Wang, Shengwen Liang, Yinhe Han, Huawei Li, and Xiaowei Li. 2023. ChipGPT: How far are we from natural language hardware design. Computing Research Repository (CoRR) arXiv:2305.14019 (May2023).","journal-title":"Computing Research Repository (CoRR)"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","unstructured":"Damai Dai Yutao Sun Li Dong Yaru Hao Shuming Ma Zhifang Sui and Furu Wei. 2023. Why Can GPT Learn In-context? Language Models Implicitly Perform Gradient Descent as Meta-optimizers. arxiv:2212.10559 [cs.CL]. https:\/\/arxiv.org\/abs\/2212.10559","DOI":"10.18653\/v1\/2023.findings-acl.247"},{"key":"e_1_3_2_10_2","article-title":"GPT4AIChip: Towards next-generation AI accelerator design automation via large language models","author":"Fu Yonggan","year":"2023","unstructured":"Yonggan Fu, Yongan Zhang, Zhongzhi Yu, Sixu Li, Zhifan Ye, Chaojian Li, Cheng Wan, and Yingyan Celine Lin. 2023. GPT4AIChip: Towards next-generation AI accelerator design automation via large language models. In International Conference on Computer-aided Design (ICCAD \u201923).","journal-title":"International Conference on Computer-aided Design (ICCAD \u201923)"},{"key":"e_1_3_2_11_2","volume-title":"GitHub Copilot  \\(\\cdot\\)  Your AI Pair Programmer","year":"2021","unstructured":"GitHub. 2021. GitHub Copilot \\(\\cdot\\) Your AI Pair Programmer. https:\/\/github.com\/features\/copilot\/"},{"key":"e_1_3_2_12_2","volume-title":"Make\u2014GNU Project\u2014Free Software Foundation","year":"1988","unstructured":"GNU. 1988. Make\u2014GNU Project\u2014Free Software Foundation. https:\/\/www.gnu.org\/software\/make\/"},{"key":"e_1_3_2_13_2","volume-title":"Autoconf\u2014GNU Project\u2014Free Software Foundation","year":"1991","unstructured":"GNU. 1991. Autoconf\u2014GNU Project\u2014Free Software Foundation. https:\/\/www.gnu.org\/software\/autoconf\/"},{"key":"e_1_3_2_14_2","volume-title":"google\/codegemma-7b  \\(\\cdot\\)  Hugging Face","unstructured":"Google. [n.d.]. google\/codegemma-7b \\(\\cdot\\) Hugging Face. https:\/\/huggingface.co\/google\/codegemma-7b"},{"key":"e_1_3_2_15_2","unstructured":"Daya Guo Qihao Zhu Dejian Yang Zhenda Xie Kai Dong Wentao Zhang Guanting Chen Xiao Bi Y. Wu Y. K. Li et\u00a0al. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming\u2014The Rise of Code Intelligence. arxiv:2401.14196 [cs.SE]"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","unstructured":"Roee Hendel Mor Geva and Amir Globerson. 2023. In-context Learning Creates Task Vectors. arxiv:2310.15916 [cs.CL]. https:\/\/arxiv.org\/abs\/2310.15916","DOI":"10.18653\/v1\/2023.findings-emnlp.624"},{"key":"e_1_3_2_17_2","volume-title":"hkust-zhiyao\/RTL-Coder","author":"zhiyao hkust","unstructured":"hkust zhiyao. [n.d.]. hkust-zhiyao\/RTL-Coder. https:\/\/github.com\/hkust-zhiyao\/RTL-Coder.original-date: 2023-11-20T13:01:13Z."},{"key":"e_1_3_2_18_2","article-title":"VerilogCoder: Autonomous Verilog coding agents with graph-based planning and Abstract Syntax Tree (AST)-based waveform tracing tool","author":"Ho Chia-Tung","year":"2024","unstructured":"Chia-Tung Ho, Haoxing Ren, and Brucek Khailany. 2024. VerilogCoder: Autonomous Verilog coding agents with graph-based planning and Abstract Syntax Tree (AST)-based waveform tracing tool. arXiv preprint. arXiv:2408.08927 (2024).","journal-title":"arXiv preprint. arXiv:2408.08927"},{"key":"e_1_3_2_19_2","unstructured":"Hanxian Huang Zhenghan Lin Zixuan Wang Xin Chen Ke Ding and Jishen Zhao. 2024. Towards LLM-powered Verilog RTL Assistant: Self-verification and Self-correction. arxiv:2406.00115 [cs.PL]. https:\/\/arxiv.org\/abs\/2406.00115"},{"key":"e_1_3_2_20_2","unstructured":"Mingjie Liu Nathaniel Pinckney Brucek Khailany and Haoxing Ren. 2023. VerilogEval: Evaluating Large Language Models for Verilog Code Generation. arxiv:2309.07544 [cs.LG]"},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","unstructured":"Shang Liu Wenji Fang Yao Lu Qijun Zhang Hongce Zhang and Zhiyao Xie. 2024. RTLCoder: Outperforming GPT-3.5 in Design RTL Generation with Our Open-source Dataset and Lightweight Solution. arxiv:2312.08617 [cs.PL]","DOI":"10.1109\/LAD62341.2024.10691788"},{"key":"e_1_3_2_22_2","unstructured":"Yao Lu Shang Liu Qijun Zhang and Zhiyao Xie. 2023. RTLLM: An Open-source Benchmark for Design RTL Generation with Large Language Model. arxiv:2308.05345 [cs.LG]"},{"key":"e_1_3_2_23_2","volume-title":"meta-llama\/CodeLlama-70b-Instruct-hf  \\(\\cdot\\)  Hugging Face","unstructured":"Meta. [n.d.]. meta-llama\/CodeLlama-70b-Instruct-hf \\(\\cdot\\) Hugging Face. https:\/\/huggingface.co\/meta-llama\/CodeLlama-70b-Instruct-hf"},{"key":"e_1_3_2_24_2","volume-title":"meta-llama\/llama-models","year":"2024","unstructured":"Meta. 2024. meta-llama\/llama-models. https:\/\/github.com\/meta-llama\/llama-models. Original-date: 2024-06-27T22:14:09Z."},{"key":"e_1_3_2_25_2","unstructured":"Zhendong Mi Renming Zheng Haowen Zhong Yue Sun and Shaoyi Huang. 2024. PromptV: Leveraging LLM-powered Multi-agent Prompting for High-quality Verilog Generation. arxiv:2412.11014 [cs.LG]. https:\/\/arxiv.org\/abs\/2412.11014"},{"key":"e_1_3_2_26_2","unstructured":"OpenAI. 2023. GPT-4 Technical Report. arxiv:2303.08774 [cs.CL]"},{"key":"e_1_3_2_27_2","volume-title":"New Models and Developer Products Announced at DevDay","year":"2023","unstructured":"OpenAI. 2023. New Models and Developer Products Announced at DevDay. https:\/\/openai.com\/index\/new-models-and-developer-products-announced-at-devday\/"},{"key":"e_1_3_2_28_2","volume-title":"Hello GPT-4o","year":"2024","unstructured":"OpenAI. 2024. Hello GPT-4o. https:\/\/openai.com\/index\/hello-gpt-4o\/"},{"key":"e_1_3_2_29_2","article-title":"AutoSVA: Democratizing formal verification of RTL module interactions","author":"Orenes-Vera Marcelo","year":"2021","unstructured":"Marcelo Orenes-Vera, Aninda Manocha, David Wentzlaff, and Margaret Martonosi. 2021. AutoSVA: Democratizing formal verification of RTL module interactions. In Design Automation Conference (DAC \u201921).","journal-title":"Design Automation Conference (DAC \u201921)"},{"key":"e_1_3_2_30_2","article-title":"Benchmarking large language models for automated Verilog RTL code generation","author":"Thakur Shailja","year":"2023","unstructured":"Shailja Thakur, Baleegh Ahmad, Zhenxing Fan, Hammond Pearce, Benjamin Tan, Ramesh Karri, Brendan Dolan-Gavitt, and Siddharth Garg. 2023. Benchmarking large language models for automated Verilog RTL code generation. In Design, Automation, and Test in Europe (DATE \u201923).","journal-title":"Design, Automation, and Test in Europe (DATE \u201923)"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3643681"},{"key":"e_1_3_2_32_2","article-title":"AutoChip: Automating HDL generation using LLM feedback","volume":"2311","author":"Thakur Shailja","year":"2023","unstructured":"Shailja Thakur, Jason Blocklove, Hammond Pearce, Benjamin Tan, Siddharth Garg, and Ramesh Karri. 2023. AutoChip: Automating HDL generation using LLM feedback. Computing Research Repository (CoRR). arXiv:2311.04887 (Nov2023).","journal-title":"Computing Research Repository (CoRR)."},{"key":"e_1_3_2_33_2","article-title":"RTLFixer: Automatically fixing RTL syntax errors with large language models","volume":"2311","author":"Tsai Yun-Da","year":"2023","unstructured":"Yun-Da Tsai, Mingjie Liu, and Haoxing Ren. 2023. RTLFixer: Automatically fixing RTL syntax errors with large language models. Computing Research Repository (CoRR). arXivv:2311.16543 (Nov2023).","journal-title":"Computing Research Repository (CoRR)."},{"key":"e_1_3_2_34_2","unstructured":"Mubashir ul Islam Humza Sami Pierre-Emmanuel Gaillardon and Valerio Tenace. 2024. AIvril: AI-driven RTL Generation with Verification In-the-loop. arxiv:2409.11411 [cs.AI]. https:\/\/arxiv.org\/abs\/2409.11411"},{"key":"e_1_3_2_35_2","unstructured":"Johannes von Oswald Eyvind Niklasson Ettore Randazzo Jo\u00e3o Sacramento Alexander Mordvintsev Andrey Zhmoginov and Max Vladymyrov. 2023. Transformers Learn In-context by Gradient Descent. arxiv:2212.07677 [cs.LG]. https:\/\/arxiv.org\/abs\/2212.07677"},{"key":"e_1_3_2_36_2","unstructured":"Zhiqiang Yuan Junwei Liu Qiancheng Zi Mingwei Liu Xin Peng and Yiling Lou. 2023. Evaluating Instruction-tuned Large Language Models on Code Comprehension and Generation. arxiv:2308.01240 [cs.CL]"},{"key":"e_1_3_2_37_2","article-title":"LLM4DV: Using large language models for hardware test stimuli generation","volume":"2310","author":"Zhang Zixi","year":"2023","unstructured":"Zixi Zhang, Greg Chadwick, Hugo McNally, Yiren Zhao, and Robert Mullins. 2023. LLM4DV: Using large language models for hardware test stimuli generation. Computing Research Repository (CoRR). arXiv:2310.04535 (Oct2023).","journal-title":"Computing Research Repository (CoRR)."},{"key":"e_1_3_2_38_2","unstructured":"Yujie Zhao Hejia Zhang Hanxian Huang Zhongming Yu and Jishen Zhao. 2024. MAGE: A Multi-agent Engine for Automated RTL Code Generation. arxiv:2412.07822 [cs.AR]. https:\/\/arxiv.org\/abs\/2412.07822"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3718088","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,22]],"date-time":"2025-10-22T12:22:20Z","timestamp":1761135740000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3718088"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,21]]},"references-count":37,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3718088"],"URL":"https:\/\/doi.org\/10.1145\/3718088","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,21]]},"assertion":[{"value":"2024-09-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}