{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T18:24:18Z","timestamp":1780597458024,"version":"3.54.1"},"reference-count":24,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,3,19]],"date-time":"2023-03-19T00:00:00Z","timestamp":1679184000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2018YFE0126300"],"award-info":[{"award-number":["2018YFE0126300"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62204111, 62034007, and 62141404"],"award-info":[{"award-number":["62204111, 62034007, and 62141404"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Zhejiang Provincial Key R&D program","award":["2020C01052"],"award-info":[{"award-number":["2020C01052"]}]},{"DOI":"10.13039\/501100020789","name":"Shuangchuang Program of Jiangsu Province","doi-asserted-by":"crossref","award":["JSSCBS20210003"],"award-info":[{"award-number":["JSSCBS20210003"]}],"id":[{"id":"10.13039\/501100020789","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2023,5,31]]},"abstract":"<jats:p>With the popularity of deep learning, the hardware implementation platform of deep learning has received increasing interest. Unlike the general purpose devices, e.g., CPU or GPU, where the deep learning algorithms are executed at the software level, neural network hardware accelerators directly execute the algorithms to achieve higher energy efficiency and performance improvements. However, as the deep learning algorithms evolve frequently, the engineering effort and cost of designing the hardware accelerators are greatly increased. To improve the design quality while saving the cost, design automation for neural network accelerators was proposed, where design space exploration algorithms are used to automatically search the optimized accelerator design within a design space. Nevertheless, the increasing complexity of the neural network accelerators brings the increasing dimensions to the design space. As a result, the previous design space exploration algorithms are no longer effective enough to find an optimized design. In this work, we propose a neural network accelerator design automation framework named GANDSE, where we rethink the problem of design space exploration, and propose a novel approach based on the generative adversarial network (GAN) to support an optimized exploration for high-dimension large design space. The experiments show that GANDSE is able to find the more optimized designs in negligible time compared with approaches including multilayer perceptron and deep reinforcement learning.<\/jats:p>","DOI":"10.1145\/3570926","type":"journal-article","created":{"date-parts":[[2022,11,9]],"date-time":"2022-11-09T11:52:17Z","timestamp":1667994737000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["GANDSE: Generative Adversarial Network-based Design Space Exploration for Neural Network Accelerator Design"],"prefix":"10.1145","volume":"28","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9943-0550","authenticated-orcid":false,"given":"Lang","family":"Feng","sequence":"first","affiliation":[{"name":"Nanjing University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3136-3416","authenticated-orcid":false,"given":"Wenjian","family":"Liu","sequence":"additional","affiliation":[{"name":"Nanjing University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7403-0163","authenticated-orcid":false,"given":"Chuliang","family":"Guo","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6461-3314","authenticated-orcid":false,"given":"Ke","family":"Tang","sequence":"additional","affiliation":[{"name":"Nanjing University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2610-7522","authenticated-orcid":false,"given":"Cheng","family":"Zhuo","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7227-4786","authenticated-orcid":false,"given":"Zhongfeng","family":"Wang","sequence":"additional","affiliation":[{"name":"Nanjing University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,3,19]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"crossref","first-page":"69","DOI":"10.1145\/2967413.2967430","article-title":"A holistic approach for optimizing DSP block utilization of a CNN implementation on FPGA","author":"Abdelouahab Kamel","year":"2016","unstructured":"Kamel Abdelouahab, C\u00e9dric Bourrasset, Maxime Pelcat, Fran\u00e7ois Berry, Jean-Charles Quinton, and Jocelyn Serot. 2016. A holistic approach for optimizing DSP block utilization of a CNN implementation on FPGA. In Proceedings of the ACM International Conference on Distributed Smart Camera. 69\u201375.","journal-title":"Proceedings of the ACM International Conference on Distributed Smart Camera"},{"key":"e_1_3_1_3_2","first-page":"1","article-title":"High Performance convolutional neural networks for document processing","author":"Chellapilla Kumar","year":"2006","unstructured":"Kumar Chellapilla, Sidd Puri, and Patrice Simard. 2006. High Performance convolutional neural networks for document processing. In Proceedings of the International Workshop on Frontiers in Handwriting Recognition. 1\u20137.","journal-title":"Proceedings of the International Workshop on Frontiers in Handwriting Recognition"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2016.2616357"},{"key":"e_1_3_1_5_2","unstructured":"Sharan Chetlur Cliff Woolley Philippe Vandermersch Jonathan Cohen John Tran Bryan Catanzaro and Evan Shelhamer. 2014. cuDNN: Efficient primitives for deep learning. Retrieved from https:\/\/arxiv.org\/abs\/1410.0759."},{"key":"e_1_3_1_6_2","unstructured":"DnnWeaver v2.0. 2016. Retrieved from http:\/\/dnnweaver.org\/."},{"key":"e_1_3_1_7_2","unstructured":"Hasan Genc Ameer Haj-Ali Vighnesh Iyer Alon Amid Howard Mao John Wright Colin Schmidt Jerry Zhao Albert Ou Max Banister Yakun Sophia Shao Borivoje Nikolic Ion Stoica and Krste Asanovic. 2019. Gemmini: An agile systolic array generator enabling systematic evaluations of deep-learning architectures. Retrieved from https:\/\/arxiv.org\/abs\/1911.09925."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3422622"},{"key":"e_1_3_1_9_2","first-page":"675","article-title":"Caffe: Convolutional architecture for fast feature embedding","author":"Jia Yangqing","year":"2014","unstructured":"Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell. 2014. Caffe: Convolutional architecture for fast feature embedding. In Proceedings of the ACM International Conference on Multimedia. 675\u2013678.","journal-title":"Proceedings of the ACM International Conference on Multimedia"},{"key":"e_1_3_1_10_2","first-page":"622","article-title":"ConfuciuX: Autonomous hardware resource assignment for DNN accelerators using reinforcement learning","author":"Kao Sheng-Chun","year":"2020","unstructured":"Sheng-Chun Kao, Geonhwa Jeong, and Tushar Krishna. 2020. ConfuciuX: Autonomous hardware resource assignment for DNN accelerators using reinforcement learning. In Proceedings of the IEEE\/ACM International Symposium on Microarchitecture. 622\u2013636.","journal-title":"Proceedings of the IEEE\/ACM International Symposium on Microarchitecture"},{"key":"e_1_3_1_11_2","first-page":"1051","article-title":"NAAS: Neural accelerator architecture search","author":"Lin Yujun","year":"2021","unstructured":"Yujun Lin, Mengtian Yang, and Song Han. 2021. NAAS: Neural accelerator architecture search. In Proceedings of the ACM\/IEEE Design Automation Conference. 1051\u20131056.","journal-title":"Proceedings of the ACM\/IEEE Design Automation Conference"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/tc.2016.2574353"},{"key":"e_1_3_1_13_2","unstructured":"Mehdi Mirza and Simon Osindero. 2014. Conditional generative adversarial nets. Retrieved from https:\/\/arxiv.org\/abs\/1411.1784."},{"key":"e_1_3_1_14_2","unstructured":"NVDLA. 2018. Retrieved from http:\/\/nvdla.org\/."},{"key":"e_1_3_1_15_2","first-page":"1","article-title":"PyTorch: An imperative style, high-performance deep learning library","volume":"32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An imperative style, high-performance deep learning library. Adv. Neural Info. Process. Syst. 32 (2019), 1\u201312.","journal-title":"Adv. Neural Info. Process. Syst."},{"key":"e_1_3_1_16_2","unstructured":"Ananda Samajdar Jan Moritz Joseph Matthew Denton and Tushar Krishna. 2021. AIRCHITECT: Learning custom architecture design and mapping space. Retrieved from https:\/\/arxiv.org\/abs\/2108.08295."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2019.2943570"},{"key":"e_1_3_1_18_2","first-page":"1","article-title":"From high-level deep neural models to FPGAs","author":"Sharma Hardik","year":"2016","unstructured":"Hardik Sharma, Jongse Park, Divya Mahajan, Emmanuel Amaro, Joon Kyung Kim, Chenkai Shao, Asit Mishra, and Hadi Esmaeilzadeh. 2016. From high-level deep neural models to FPGAs. In Proceedings of the IEEE\/ACM International Symposium on Microarchitecture. 1\u201312.","journal-title":"Proceedings of the IEEE\/ACM International Symposium on Microarchitecture"},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","unstructured":"Atefeh Sohrabizadeh Yunsheng Bai Yizhou Sun and Jason Cong. 2022. Automated accelerator optimization aided by graph neural networks. In Proceedings of the 59th ACM\/IEEE Design Automation Conference (DAC\u201922) . 55\u201360.","DOI":"10.1145\/3489517.3530409"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3494534"},{"key":"e_1_3_1_21_2","first-page":"40","article-title":"AutoDNNchip: An automated DNN chip predictor and builder for both FPGAs and ASICs","author":"Xu Pengfei","year":"2020","unstructured":"Pengfei Xu, Xiaofan Zhang, Cong Hao, Yang Zhao, Yongan Zhang, Yue Wang, Chaojian Li, Zetong Guan, Deming Chen, and Yingyan Lin. 2020. AutoDNNchip: An automated DNN chip predictor and builder for both FPGAs and ASICs. In Proceedings of the ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays. 40\u201350.","journal-title":"Proceedings of the ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays"},{"key":"e_1_3_1_22_2","first-page":"117","article-title":"A framework for generating high throughput CNN implementations on FPGAs","author":"Zeng Hanqing","year":"2018","unstructured":"Hanqing Zeng, Ren Chen, Chi Zhang, and Viktor Prasanna. 2018. A framework for generating high throughput CNN implementations on FPGAs. In Proceedings of the ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays. 117\u2013126.","journal-title":"Proceedings of the ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays"},{"key":"e_1_3_1_23_2","first-page":"1","article-title":"DNNBuilder: An automated tool for building high-performance DNN hardware accelerators for FPGAs","author":"Zhang Xiaofan","year":"2018","unstructured":"Xiaofan Zhang, Junsong Wang, Chao Zhu, Yonghua Lin, Jinjun Xiong, Wen-mei Hwu, and Deming Chen. 2018. DNNBuilder: An automated tool for building high-performance DNN hardware accelerators for FPGAs. In Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design. 1\u20138.","journal-title":"Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design"},{"key":"e_1_3_1_24_2","first-page":"1","article-title":"DNNExplorer: A framework for modeling and exploring a novel paradigm of FPGA-based DNN accelerator","author":"Zhang Xiaofan","year":"2020","unstructured":"Xiaofan Zhang, Hanchen Ye, Junsong Wang, Yonghua Lin, Jinjun Xiong, Wen-mei Hwu, and Deming Chen. 2020. DNNExplorer: A framework for modeling and exploring a novel paradigm of FPGA-based DNN accelerator. In Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design. 1\u20139.","journal-title":"Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design"},{"key":"e_1_3_1_25_2","first-page":"706","article-title":"Cambricon-Q: A hybrid architecture for efficient training","author":"Zhao Yongwei","year":"2021","unstructured":"Yongwei Zhao, Chang Liu, Zidong Du, Qi Guo, Xing Hu, Yimin Zhuang, Zhenxing Zhang, Xinkai Song, Wei Li, Xishan Zhang, Ling Li, Zhiwei Xu, and Tianshi Chen. 2021. Cambricon-Q: A hybrid architecture for efficient training. In Proceedings of the International Symposium on Computer Architecture. 706\u2013719.","journal-title":"Proceedings of the International Symposium on Computer Architecture"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3570926","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3570926","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:12Z","timestamp":1750182552000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3570926"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,19]]},"references-count":24,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,5,31]]}},"alternative-id":["10.1145\/3570926"],"URL":"https:\/\/doi.org\/10.1145\/3570926","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,19]]},"assertion":[{"value":"2022-07-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-10-31","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-19","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}