{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,20]],"date-time":"2026-04-20T18:25:45Z","timestamp":1776709545413,"version":"3.51.2"},"reference-count":37,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2021,11,30]],"date-time":"2021-11-30T00:00:00Z","timestamp":1638230400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2022,3,31]]},"abstract":"<jats:p>\n            The underlying goal of FPGA architecture research is to devise flexible substrates that implement a wide variety of circuits efficiently. Contemporary FPGA architectures have been optimized to support networking, signal processing, and image processing applications through high-precision digital signal processing (DSP) blocks. The recent emergence of machine learning has created a new set of demands characterized by: (1) higher computational density and (2) low precision arithmetic requirements. With the goal of exploring this new design space in a methodical manner, we first propose a problem formulation involving computing nested loops over multiply-accumulate (MAC) operations, which covers many basic linear algebra primitives and standard deep neural network (DNN) kernels. A quantitative methodology for deriving efficient coarse-grained compute block architectures from benchmarks is then proposed together with a family of new embedded blocks, called MLBlocks. An MLBlock instance includes several multiply-accumulate units connected via a flexible routing, where each configuration performs a few parallel dot-products in a systolic array fashion. This architecture is parameterized with support for different data movements, reuse, and precisions, utilizing a columnar arrangement that is compatible with existing FPGA architectures. On synthetic benchmarks, we demonstrate that for 8-bit arithmetic, MLBlocks offer 6\u00d7 improved performance over the commercial Xilinx DSP48E2 architecture with smaller area and delay; and for time-multiplexed 16-bit arithmetic, achieves 2\u00d7 higher performance per area with the same area and frequency. All source codes and data, along with documents to reproduce all the results in this article, are available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"http:\/\/github.com\/raminrasoulinezhad\/MLBlocks\">http:\/\/github.com\/raminrasoulinezhad\/MLBlocks<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3491234","type":"journal-article","created":{"date-parts":[[2021,11,30]],"date-time":"2021-11-30T16:04:28Z","timestamp":1638288268000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Rethinking Embedded Blocks for Machine Learning Applications"],"prefix":"10.1145","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7366-8533","authenticated-orcid":false,"given":"Seyedramin","family":"Rasoulinezhad","sequence":"first","affiliation":[{"name":"The University of Sydney, NSW, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Esther","family":"Roorda","sequence":"additional","affiliation":[{"name":"The University of British Columbia, Vancouver, BC, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Steve","family":"Wilton","sequence":"additional","affiliation":[{"name":"The University of British Columbia, Vancouver, BC, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Philip H. W.","family":"Leong","sequence":"additional","affiliation":[{"name":"The University of Sydney, NSW, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David","family":"Boland","sequence":"additional","affiliation":[{"name":"The University of Sydney, NSW, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,11,30]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3309551"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3020078.3021740"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.5555\/3504035.3504135"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/3122009.3242044"},{"key":"e_1_3_1_6_2","series-title":"Proceedings of the 36th International Conference on Machine Learning","first-page":"7015","author":"Yang Guandao","year":"2019","unstructured":"Guandao Yang, Tianyi Zhang, Polina Kirichenko, Junwen Bai, Andrew Gordon Wilson, and Christopher De Sa. 2019. SWALP: Stochastic weight averaging in low precision training. In Proceedings of the 36th International Conference on Machine Learning(Proceedings of Machine Learning Research Vol. 97), Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, 7015\u20137024. Retrieved from http:\/\/proceedings.mlr.press\/v97\/yang19d.html."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373087.3375352"},{"key":"e_1_3_1_8_2","unstructured":"Xilinx. 2017. Deep Learning with INT8 Optimization on Xilinx Devices - WP486 (v1.0.1). Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/white_papers\/wp486-deep-learning-int8.pdf."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2019.00061"},{"key":"e_1_3_1_10_2","unstructured":"Nvidia. 2020. DS-10184-001: NVIDIA Jetson Xavier NXSystem-on-Module. https:\/\/img.iceasy.com\/product\/product\/files\/202107\/8a8a8a1a7a81d57a017a9eb9bb204157.pdf."},{"key":"e_1_3_1_11_2","unstructured":"Xilinx. 2021. XMP103: Product Selection Guide. (2021). Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/selection-guides\/ultrascale-plus-fpga-product-selection-guide.pdf#VUSP."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378514"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689060"},{"key":"e_1_3_1_14_2","unstructured":"Intel. 2020. SV51001 Stratix V Device Overview. (2020). Retrieved from https:\/\/www.intel.com\/content\/dam\/www\/programmable\/us\/en\/pdfs\/literature\/hb\/stratix-v\/stx5_51001.pdf."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2018.00014"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2019.00015"},{"key":"e_1_3_1_17_2","unstructured":"Intel. 2019. Intel Agilex Variable Precision DSP Blocks User Guide. (2019). Retrieved from https:\/\/www.intel.com\/content\/dam\/altera-www\/global\/en=US\/pdfs\/literature\/hb\/agilex\/ug-ag-dsp.pdf."},{"key":"e_1_3_1_18_2","unstructured":"Xilinx. 2021. Versal ACAP DSP Engine Architecture Manual AM004 (v1.1.2). (2021). Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/architecture-manuals\/am004-versal-dsp-engine.pdf."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373087.3375303"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289602.3293912"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3393668"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289602.3293906"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3174243.3174966"},{"key":"e_1_3_1_24_2","unstructured":"Achronix - Data Acceleration. Speedster7t IP Component Library User Guide (UG086). 2019. Retrieved from https:\/\/www.achronix.com\/sites\/default\/files\/docs\/Speedster7t_IP_Component_Library_User_Guide_UG086.pdf."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICFPT51103.2020.00011"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASAP49362.2020.00018"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3431920.3439282"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3381271.3381300"},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","first-page":"210","DOI":"10.1007\/978-3-030-48513-9_17","volume-title":"Cloud Computing, Smart Grid and Innovative Frontiers in Telecommunications","author":"Sun Wei","year":"2020","unstructured":"Wei Sun, Xijie Zhou, Xiaorui Zhang, and Xiaozheng He. 2020. A lightweight neural network combining dilated convolution and depthwise separable convolution. In Cloud Computing, Smart Grid and Innovative Frontiers in Telecommunications, Xuyun Zhang, Guanfeng Liu, Meikang Qiu, Wei Xiang, and Tao Huang (Eds.). Springer International Publishing, Cham, 210\u2013220."},{"key":"e_1_3_1_30_2","article-title":"Optimizing temporal convolutional network inference on FPGA-based accelerators","volume":"2005","author":"Carreras Marco","year":"2020","unstructured":"Marco Carreras, Gianfranco Deriu, Luigi Raffo, Luca Benini, and Paolo Meloni. 2020. Optimizing temporal convolutional network inference on FPGA-based accelerators. CoRR abs\/2005.03775 (2020).","journal-title":"CoRR"},{"key":"e_1_3_1_31_2","unstructured":"Baidu. 2016. DeepBench. (2016). Retrieved from https:\/\/github.com\/baidu-research\/DeepBench."},{"key":"e_1_3_1_32_2","unstructured":"Xilinx. 2019. DS183 (v1.28) - Virtex-7 T and XT FPGAs Data Sheet:DC and AC Switching Characteristics. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/data_sheets\/ds183_Virtex_7_Data_Sheet.pdf."},{"key":"e_1_3_1_33_2","volume-title":"TRETS FPT Journal Track - under review","author":"Roorda Esther","year":"2021","unstructured":"Esther Roorda, Seyedramin Rasoulinezhad, Philip H. W. Leong, and Steve Wilton. 2021. FPGA architecture exploration for DNN acceleration. In TRETS FPT Journal Track - under review."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689071"},{"key":"e_1_3_1_35_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proceedings of the 3rd International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.). Retrieved from http:\/\/arxiv.org\/abs\/1409.1556."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_37_2","unstructured":"Avnet. 2021. Ultra96-V2 Single Board Computer Hardware User\u2019s Guide Version 1.3. (2021). Retrieved from https:\/\/www.avnet.com\/wps\/wcm\/connect\/onesite\/b85b9556-0b2a-42b3-ad6a-8dcf3eac1ff9\/Ultra96-V2-HW-User-Guide-v1_3.pdf?MOD=AJPERES&CACHEID=ROOTWORKSPACE.Z18_NA5A1I41L0ICD0ABNDMDDG0000-b85b9556-0b2a-42b3-ad6a-8dcf3eac1ff9-nDNP5R3."},{"key":"e_1_3_1_38_2","unstructured":"Xilinx. 2018. ZCU111 Evaluation Board User Guide UG1271 (v1.2). (2018). Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/boards_and_kits\/zcu111\/ug1271-zcu111-eval-bd.pdf."}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3491234","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3491234","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:09:19Z","timestamp":1750183759000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3491234"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,11,30]]},"references-count":37,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,3,31]]}},"alternative-id":["10.1145\/3491234"],"URL":"https:\/\/doi.org\/10.1145\/3491234","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,11,30]]},"assertion":[{"value":"2021-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-11-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}