{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T09:02:35Z","timestamp":1783069355875,"version":"3.54.6"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"6","funder":[{"name":"Graduate Assistance in Areas of National Need"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>Traditionally, digital hardware designs are written in the Verilog hardware description language (HDL) and debugged manually by engineers. This can be time-consuming and error-prone for complex designs. Large Language Models (LLMs) are emerging as a potential tool to help generate fully functioning HDL code, but most works have focused on generation in the single-shot capacity: i.e., run and evaluate, a process that does not leverage debugging and, as such, does not adequately reflect a realistic development process. In this work, we evaluate the ability of LLMs to leverage feedback from electronic design automation (EDA) tools to fix mistakes in their own generated Verilog. To accomplish this, we present an open-source, highly customizable framework, AutoChip, which combines conversational LLMs with the output from Verilog compilers and simulations to iteratively generate and repair Verilog. To determine the success of these LLMs we leverage the VerilogEval benchmark set. We evaluate four state-of-the-art conversational LLMs, focusing on readily accessible commercial models. EDA tool feedback proved to be consistently more effective than zero-shot prompting only with GPT-4o, the most computationally complex model we evaluated. In the best case, we observed a 5.8% increase in the number of successful designs with a 34.2% decrease in cost over the best zero-shot results. Mixing smaller models with this larger model at the end of the feedback iterations resulted in equally as much success as with GPT-4o using feedback, but incurred 41.9% lower cost (corresponding to an overall decrease in cost over zero-shot by 89.6%).<\/jats:p>","DOI":"10.1145\/3723876","type":"journal-article","created":{"date-parts":[[2025,3,26]],"date-time":"2025-03-26T07:07:47Z","timestamp":1742972867000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["Automatically Improving LLM-based Verilog Generation using EDA Tool Feedback"],"prefix":"10.1145","volume":"30","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-5619-4654","authenticated-orcid":false,"given":"Jason","family":"Blocklove","sequence":"first","affiliation":[{"name":"New York University Tandon School of Engineering","place":["Brooklyn, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9590-5061","authenticated-orcid":false,"given":"Shailja","family":"Thakur","sequence":"additional","affiliation":[{"name":"New York University Tandon School of Engineering","place":["Brooklyn, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7642-3638","authenticated-orcid":false,"given":"Benjamin","family":"Tan","sequence":"additional","affiliation":[{"name":"Electrical and Software Engineering, University of Calgary","place":["Calgary, Canada"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3488-7004","authenticated-orcid":false,"given":"Hammond","family":"Pearce","sequence":"additional","affiliation":[{"name":"University of New South Wales Faculty of Engineering","place":["Sydney, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6158-9512","authenticated-orcid":false,"given":"Siddharth","family":"Garg","sequence":"additional","affiliation":[{"name":"New York University Tandon School of Engineering","place":["New York, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7989-5617","authenticated-orcid":false,"given":"Ramesh","family":"Karri","sequence":"additional","affiliation":[{"name":"New York University Tandon School of Engineering","place":["Brooklyn, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,10,22]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2023. Introducing PaLM 2. (2023). Retrieved from https:\/\/blog.google\/technology\/ai\/google-palm-2-ai-large-language-model\/"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2024.3374558"},{"key":"e_1_3_1_4_2","unstructured":"Anthropic. 2024. Claude 3 Haiku: our fastest model yet. (2024). Retrieved from https:\/\/www.anthropic.com\/news\/claude-3-haiku"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/MLCAD58807.2023.10299874"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","unstructured":"Jason Blocklove Siddharth Garg Ramesh Karri and Hammond Pearce. 2024. Evaluating LLMs for hardware design and test. In 2024 IEEE LLM Aided Design Workshop (LAD). 1\u20136. DOI:10.1109\/LAD62341.2024.10691811","DOI":"10.1109\/LAD62341.2024.10691811"},{"key":"e_1_3_1_7_2","first-page":"1877","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of the Advances in Neural Information Processing Systems. H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33, Curran Associates, Inc., 1877\u20131901. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf"},{"key":"e_1_3_1_8_2","unstructured":"Cadence. 2023. Cadence JedAI Generative AI Solution for Chip System and Product Design. (Sept.2023). https:\/\/www.cadence.com\/en_US\/home\/solutions\/joint-enterprise-data-ai-platform.html"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","unstructured":"Kaiyan Chang Ying Wang Haimeng Ren Mengdi Wang Shengwen Liang Yinhe Han Huawei Li and Xiaowei Li. 2023. ChipGPT: How far are we from natural language hardware design. DOI:10.48550\/arXiv.2305.14019 arXiv:2305.14019 [cs].","DOI":"10.48550\/arXiv.2305.14019"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mohammad Bavarian Clemens Winter Philippe Tillet Felipe Petroski Such Dave Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William Hebgen Guss Alex Nichol Alex Paino Nikolas Tezak Jie Tang Igor Babuschkin Suchir Balaji Shantanu Jain William Saunders Christopher Hesse Andrew N. Carr Jan Leike Josh Achiam Vedant Misra Evan Morikawa Alec Radford Matthew Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating Large Language Models Trained on Code. DOI:10.48550\/arXiv.2107.03374 arXiv:2107.03374 [cs].","DOI":"10.48550\/arXiv.2107.03374"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","unstructured":"Luca Collini Siddharth Garg and Ramesh Karri. 2024. C2HLSC: Can LLMs Bridge the Software-to-Hardware Design Gap? In 2024 IEEE LLM Aided Design Workshop (LAD). 1\u201312. DOI:10.1109\/LAD62341.2024.10691856","DOI":"10.1109\/LAD62341.2024.10691856"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","unstructured":"Matthew DeLorenzo Animesh Basak Chowdhury Vasudev Gohil Shailja Thakur Ramesh Karri Siddharth Garg and Jeyavijayan Rajendran. 2024. Make Every Move Count: LLM-based High-Quality RTL Code Generation Using MCTS. DOI:10.48550\/arXiv.2402.03289 arXiv:2402.03289 [cs].","DOI":"10.48550\/arXiv.2402.03289"},{"key":"e_1_3_1_13_2","first-page":"213","volume-title":"Proceedings of the 28th USENIX Conference on Security Symposium (SEC\u201919)","author":"Dessouky Ghada","year":"2019","unstructured":"Ghada Dessouky, David Gens, Patrick Haney, Garrett Persyn, Arun Kanuparthi, Hareesh Khattri, Jason Fung, Ahmad-Reza Sadeghi, and Jeyavijayan Rajendran. 2019. Hardfails: Insights into software-exploitable hardware bugs. In Proceedings of the 28th USENIX Conference on Security Symposium (SEC\u201919). USENIX Association, Santa Clara, CA, USA, 213\u2013230."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAD57390.2023.10323953"},{"key":"e_1_3_1_15_2","unstructured":"GitHub. 2021. GitHub Copilot \u00b7 Your AI pair programmer. (2021). Retrieved from https:\/\/copilot.github.com\/"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/MLCAD58807.2023.10299852"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","unstructured":"Hanxian Huang Zhenghan Lin Zixuan Wang Xin Chen Ke Ding and Jishen Zhao. 2024. Towards LLM-Powered Verilog RTL Assistant: Self-Verification and Self-Correction. DOI:10.48550\/arXiv.2406.00115 arXiv:2406.00115 [cs].","DOI":"10.48550\/arXiv.2406.00115"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2024.3372809"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","unstructured":"Mingjie Liu Teodor-Dumitru Ene Robert Kirby Chris Cheng Nathaniel Pinckney Rongjian Liang Jonah Alben Himyanshu Anand Sanmitra Banerjee Ismet Bayraktaroglu Bonita Bhaskaran Bryan Catanzaro Arjun Chaudhuri Sharon Clay Bill Dally Laura Dang Parikshit Deshpande Siddhanth Dhodhi Sameer Halepete Eric Hill Jiashang Hu Sumit Jain Brucek Khailany Kishor Kunal Xiaowei Li Hao Liu Stuart Oberman Sujeet Omar Sreedhar Pratty Jonathan Raiman Ambar Sarkar Zhengjiang Shao Hanfei Sun Pratik P. Suthar Varun Tej Kaizhe Xu and Haoxing Ren. 2023. ChipNeMo: Domain-Adapted LLMs for Chip Design. DOI:10.48550\/arXiv.2311.00176 arXiv:2311.00176 [cs].","DOI":"10.48550\/arXiv.2311.00176"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","unstructured":"Mingjie Liu Nathaniel Pinckney Brucek Khailany and Haoxing Ren. 2023. VerilogEval: Evaluating Large Language Models for Verilog Code Generation. In 2023 IEEE\/ACM International Conference on Computer Aided Design (ICCAD). 1\u20138. DOI:10.1109\/ICCAD57390.2023.10323812","DOI":"10.1109\/ICCAD57390.2023.10323812"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","unstructured":"Shang Liu Wenji Fang Yao Lu Qijun Zhang Hongce Zhang and Zhiyao Xie. 2024. RTLCoder: Outperforming GPT-3.5 in Design RTL Generation with Our Open-Source Dataset and Lightweight Solution. In 2024 IEEE LLM Aided DesignWorkshop (LAD). 1\u20135. DOI:10.1109\/LAD62341.2024.10691788","DOI":"10.1109\/LAD62341.2024.10691788"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","unstructured":"Yao Lu Shang Liu Qijun Zhang and Zhiyao Xie. 2024. RTLLM: An open-source benchmark for design RTL Generation with large language model. In Proceedings of the 29th Asia and South Pacific Design Automation Conference (Incheon Republic of Korea) (ASPDAC\u201924). IEEE Press 722\u2013727. DOI:10.1109\/ASP-DAC58780.2024.10473904","DOI":"10.1109\/ASP-DAC58780.2024.10473904"},{"key":"e_1_3_1_23_2","unstructured":"Meta. 2023. Introducing Code Llama an AI Tool for Coding. (Aug.2023). Retrieved from https:\/\/about.fb.com\/news\/2023\/08\/code-llama-ai-for-coding\/"},{"key":"e_1_3_1_24_2","unstructured":"Mistral. 2023. Mistral 7B. (Sept.2023). Retrieved from https:\/\/mistral.ai\/news\/announcing-mistral-7b\/"},{"key":"e_1_3_1_25_2","unstructured":"OpenAI. [n.d.] GPT-4o mini: advancing cost-efficient intelligence. (n.d.). Retrieved from https:\/\/openai.com\/index\/gpt-4o-mini-advancing-cost-efficient-intelligence\/"},{"key":"e_1_3_1_26_2","unstructured":"OpenAI. 2022. Introducing ChatGPT. (Nov.2022). Retrieved from https:\/\/openai.com\/blog\/chatgpt"},{"key":"e_1_3_1_27_2","unstructured":"OpenAI. 2024. GPT-4o. (May2024). Retrieved from https:\/\/openai.com\/index\/hello-gpt-4o\/"},{"key":"e_1_3_1_28_2","first-page":"27730","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"35","author":"Ouyang Long","year":"2022","unstructured":"Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Proceedings of the Advances in Neural Information Processing Systems. S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35, Curran Associates, Inc., 27730\u201327744. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/b1efde53be364a73914f58805a001731-Paper-Conference.pdf"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3380446.3430634"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","unstructured":"Zehua Pei Hui-Ling Zhen Mingxuan Yuan Yu Huang and Bei Yu. 2024. BetterV: Controlled Verilog Generation with Discriminative Guidance. DOI:10.48550\/arXiv.2402.03375 arXiv:2402.03375 [cs].","DOI":"10.48550\/arXiv.2402.03375"},{"key":"e_1_3_1_31_2","unstructured":"Sundar Pichai. 2023. An important next step on our AI journey. (Feb.2023). Retrieved from https:\/\/blog.google\/technology\/ai\/bard-google-ai-search-updates\/"},{"key":"e_1_3_1_32_2","unstructured":"Sundar Pichai and Demis Hassabis. 2023. Introducing Gemini: our largest and most capable AI model. (Dec.2023). Retrieved from https:\/\/blog.google\/technology\/ai\/google-gemini-ai\/"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","unstructured":"Nathaniel Pinckney Christopher Batten Mingjie Liu Haoxing Ren and Brucek Khailany. 2024. Revisiting VerilogEval: Newer LLMs In-Context Learning and Specification-to-RTL Tasks. DOI:10.48550\/arXiv.2408.11053 arXiv:2408.11053 [cs].","DOI":"10.48550\/arXiv.2408.11053"},{"key":"e_1_3_1_34_2","unstructured":"RapidSilicon. 2023. RapidGPT. (2023). Retrieved from https:\/\/rapidsilicon.com\/rapidgpt\/"},{"key":"e_1_3_1_35_2","unstructured":"Synopsys. 2023. Redefining Chip Design with AI-Powered EDA Tools | Synopsys.ai | Synopsys Blog. (March2023). Retrieved from https:\/\/www.synopsys.com\/blogs\/chip-design\/synopsys-ai-eda-tools.html"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE56975.2023.10137086"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","unstructured":"Shailja Thakur Jason Blocklove Hammond Pearce Benjamin Tan Siddharth Garg and Ramesh Karri. 2023. AutoChip: Automating HDL Generation Using LLM Feedback. DOI:10.48550\/arXiv.2311.04887 arXiv:2311.04887 [cs].","DOI":"10.48550\/arXiv.2311.04887"},{"key":"e_1_3_1_38_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems. Curran Associates, Inc.Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-long.319"},{"key":"e_1_3_1_40_2","unstructured":"Stephen Williams. 2023. The ICARUS Verilog Compilation System. (May2023). Retrieved from https:\/\/github.com\/steveicarus\/iverilog"},{"key":"e_1_3_1_41_2","unstructured":"Henry Wong. 2017. HDLBits Problem Sets. (Nov.2017). Retrieved from https:\/\/hdlbits.01xz.net\/wiki\/Problem_sets"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","unstructured":"Yang Zhao Di Huang Chongxiao Li Pengwei Jin Ziyuan Nan Tianyun Ma Lei Qi Yansong Pan Zhenxing Zhang Rui Zhang Xishan Zhang Zidong Du Qi Guo Xing Hu and Yunji Chen. 2024. CodeV: Empowering LLMs for Verilog Generation through Multi-Level Summarization. DOI:10.48550\/arXiv.2407.10424 arXiv:2407.10424 [cs].","DOI":"10.48550\/arXiv.2407.10424"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3723876","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,22]],"date-time":"2025-10-22T12:21:21Z","timestamp":1761135681000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3723876"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,22]]},"references-count":41,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3723876"],"URL":"https:\/\/doi.org\/10.1145\/3723876","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,22]]},"assertion":[{"value":"2024-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-28","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}