{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T15:33:45Z","timestamp":1779204825484,"version":"3.51.4"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2022,5,10]],"date-time":"2022-05-10T00:00:00Z","timestamp":1652140800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2022,9,30]]},"abstract":"<jats:p>\n            Recent years have seen an explosion of machine learning applications implemented on\n            <jats:bold>Field-Programmable Gate Arrays<\/jats:bold>\n            <jats:bold>(FPGAs)<\/jats:bold>\n            . FPGA vendors and researchers have responded by updating their fabrics to more efficiently implement machine learning accelerators, including innovations such as enhanced\n            <jats:bold>Digital Signal Processing (DSP)<\/jats:bold>\n            blocks and hardened systolic arrays. Evaluating architectural proposals is difficult, however, due to the lack of publicly available benchmark circuits.\n          <\/jats:p>\n          <jats:p>\n            This paper addresses this problem by presenting an open-source benchmark circuit generator that creates realistic DNN-oriented circuits for use in FPGA architecture studies. Unlike previous generators, which create circuits that are agnostic of the underlying FPGA, our circuits explicitly instantiate embedded blocks, allowing for meaningful comparison of recent architectural proposals without the need for a complete inference\n            <jats:bold>computer-aided design (CAD)<\/jats:bold>\n            flow. Our circuits are compatible with the VTR CAD suite, allowing for architecture studies that investigate routing congestion and other low-level architectural implications.\n          <\/jats:p>\n          <jats:p>In addition to addressing the lack of machine learning benchmark circuits, the architecture exploration flow that we propose allows for a more comprehensive evaluation of FPGA architectures than traditional static benchmark suites. We demonstrate this through three case studies which illustrate how realistic benchmark circuits can be generated to target different heterogeneous FPGAs.<\/jats:p>","DOI":"10.1145\/3503465","type":"journal-article","created":{"date-parts":[[2022,5,10]],"date-time":"2022-05-10T11:45:32Z","timestamp":1652183132000},"page":"1-37","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["FPGA Architecture Exploration for DNN Acceleration"],"prefix":"10.1145","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1905-9577","authenticated-orcid":false,"given":"Esther","family":"Roorda","sequence":"first","affiliation":[{"name":"University of British Columbia, Vancouver, British Columbia, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Seyedramin","family":"Rasoulinezhad","sequence":"additional","affiliation":[{"name":"University of Sydney, Sydney, Australia, NSW"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3923-3499","authenticated-orcid":false,"given":"Philip H. W.","family":"Leong","sequence":"additional","affiliation":[{"name":"University of Sydney, Sydney, Australia, NSW"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Steven J. E.","family":"Wilton","sequence":"additional","affiliation":[{"name":"University of British Columbia, Vancouver, British Columbia, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,5,10]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783725"},{"key":"e_1_3_3_3_2","volume-title":"Proceedings of the 2020 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays","author":"Arora Aman","year":"2020","unstructured":"Aman Arora, Samidh Mehta, Vaughn Betz, and Lizy John. 2020. Tensor slices to the rescue: Supercharging ML acceleration on FPGAs. In Proceedings of the 2020 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays."},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASAP49362.2020.00018"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2019.00030"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.5555\/553523"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289602.3293912"},{"key":"e_1_3_3_8_2","doi-asserted-by":"crossref","unstructured":"Andrew Boutros Eriko Nurvitadhi Rui Ma Sergey Gribok Zhipeng Zhao James Hoe Vaughn Betz and Martin Langhammer. 2020. Beyond peak performance: Comparing the real performance of AI-optimized FPGAs and GPUs(FPT\u201920).","DOI":"10.1109\/ICFPT51103.2020.00011"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2018.00014"},{"key":"e_1_3_3_10_2","unstructured":"Intel Corporation. 2020. Avalon Interface Specifications."},{"key":"e_1_3_3_11_2","doi-asserted-by":"crossref","first-page":"181","DOI":"10.1145\/1950413.1950449","volume-title":"ACM\/SIGDA International Symposium on Field Programmable Gate Arrays","author":"Das J.","year":"2011","unstructured":"J. Das and S. J. E. Wilton. 2011. An analytical model relating FPGA architecture parameters to routability. In ACM\/SIGDA International Symposium on Field Programmable Gate Arrays. 181\u2013184."},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3393668"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/NOCS.2018.8512170"},{"key":"e_1_3_3_14_2","first-page":"656","volume-title":"1998 Design, Automation and Test in Europe","author":"Ghosh Debabrata","year":"1998","unstructured":"Debabrata Ghosh, Nevin Kapur, Franc Brglez, and Justin E. Harlow. 1998. Synthesis of wiring signature-invariant equivalence class circuit mutants and applications to benchmarking. In 1998 Design, Automation and Test in Europe. 656\u2013663."},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358291"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/1391732.1391736"},{"key":"e_1_3_3_17_2","doi-asserted-by":"crossref","unstructured":"Kartik Hegde Rohit Agrawal Yulun Yao and Christopher W. Fletcher. 2018. Morph: Flexible Acceleration for 3D CNN-based Video Understanding. (2018). arxiv:cs.LG\/1810.06807","DOI":"10.1109\/MICRO.2018.00080"},{"key":"e_1_3_3_18_2","unstructured":"Andrew G. Howard Menglong Zhu Bo Chen Dmitry Kalenichenko Weijun Wang Tobias Weyand Marco Andreetto and Hartwig Adam. 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. (2017). arxiv:cs.CV\/1704.04861"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2002.800456"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2020.2997638"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPT.2018.00039"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2004.828132"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3296957.3173176"},{"key":"e_1_3_3_24_2","volume-title":"Proceedings of the 2020 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201920)","author":"Langhammer Martin","year":"2020","unstructured":"Martin Langhammer, Eriko Nurvitadhi, Bogdan Pasca, and Sergey Gribok. 2020. Stratix 10 NX architecture and applications. In Proceedings of the 2020 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201920)."},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.29"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/2331147.2331152"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3388617"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/2629579"},{"key":"e_1_3_3_29_2","unstructured":"Angshuman Parashar Minsoo Rhu Anurag Mukkara Antonio Puglielli Rangharajan Venkatesan Brucek Khailany Joel Emer Stephen W. Keckler and William J. Dally. 2017. SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks. (2017). arxiv:cs.NE\/1708.04485"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/43.892855"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373087.3375303"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2019.00015"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2019.00061"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783720"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783720"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080221"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373087.3375321"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-9260(99)00002-4"},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/2847263.2847276"},{"key":"e_1_3_3_40_2","unstructured":"Vivienne Sze Yu-Hsin Chen Tien-Ju Yang and Joel Emer. 2017. Efficient Processing of Deep Neural Networks: A Tutorial and Survey. (2017). arxiv:cs.CV\/1703.09039"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/1065579.1065770"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2016.22"},{"key":"e_1_3_3_43_2","first-page":"31","volume-title":"International Conference on VLSI","author":"Verplaetse P.","year":"2002","unstructured":"P. Verplaetse, D. Stroobandt, and J. VanCampenhout. 2002. Synthetic benchmark circuits for timing-driven physical design applications. In International Conference on VLSI. 31\u201337."},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/2897937.2898002"},{"key":"e_1_3_3_45_2","first-page":"147","volume-title":"International Symposium on FPGAs","author":"Yan Andy","year":"2002","unstructured":"Andy Yan, Rebecca Cheng, and Steven J. E. Wilton. 2002. On the sensitivity of FPGA architectural conclusions to experimental assumptions, tools, and techniques. In International Symposium on FPGAs. 147\u2013156."},{"key":"e_1_3_3_46_2","first-page":"1","volume-title":"MCNC","author":"Yang S.","year":"1991","unstructured":"S. Yang. 1991. Logic Synthesis and Optimization Benchmarks User Guide 3.0. Technical Report. In MCNC. 1\u20136."},{"key":"e_1_3_3_47_2","article-title":"A systematic approach to blocking convolutional neural networks","volume":"1606","author":"Yang Xuan","year":"2016","unstructured":"Xuan Yang, Jing Pu, Blaine Burton Rister, Nikhil Bhagdikar, Stephen Richardson, Shahar Kvatinsky, Jonathan Ragan-Kelley, Ardavan Pedram, and Mark Horowitz. 2016. A systematic approach to blocking convolutional neural networks. CoRR abs\/1606.04209 (2016). arxiv:1606.04209http:\/\/arxiv.org\/abs\/1606.04209","journal-title":"CoRR"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2016.7577356"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3174243.3174265"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689060"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/2966986.2967011"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240765.3240801"},{"key":"e_1_3_3_53_2","doi-asserted-by":"crossref","unstructured":"Xiaofan Zhang Hanchen Ye Junsong Wang Yonghua Lin Jinjun Xiong Wen Mei Hwu and Deming Chen. 2021. DNNExplorer: A Framework for Modeling and Exploring a Novel Paradigm of FPGA-based DNN Accelerator. (2021). arxiv:cs.AR\/2008.12745","DOI":"10.1145\/3400302.3415609"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503465","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503465","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:02:18Z","timestamp":1750186938000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503465"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,5,10]]},"references-count":52,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,9,30]]}},"alternative-id":["10.1145\/3503465"],"URL":"https:\/\/doi.org\/10.1145\/3503465","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,5,10]]},"assertion":[{"value":"2021-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-05-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}