{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,7]],"date-time":"2026-01-07T23:53:58Z","timestamp":1767830038018,"version":"3.49.0"},"reference-count":25,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2019,8,20]],"date-time":"2019-08-20T00:00:00Z","timestamp":1566259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004359","name":"Vetenskapsr\u00e5det","doi-asserted-by":"crossref","award":["2015-05159"],"award-info":[{"award-number":["2015-05159"]}],"id":[{"id":"10.13039\/501100004359","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2019,9,30]]},"abstract":"<jats:p>Matrix-matrix multiplication is a key computational kernel for numerous applications in science and engineering, with ample parallelism and data locality that lends itself well to high-performance implementations. Many matrix multiplication-dependent applications can use reduced-precision integer or fixed-point representations to increase their performance and energy efficiency while still offering adequate quality of results. However, precision requirements may vary between different application phases or depend on input data, rendering constant-precision solutions ineffective. BISMO, a vectorized bit-serial matrix multiplication overlay for reconfigurable computing, previously utilized the excellent binary-operation performance of FPGAs to offer a matrix multiplication performance that scales with required precision and parallelism. We show how BISMO can be scaled up on Xilinx FPGAs using an arithmetic architecture that better utilizes six-input LUTs. The improved BISMO achieves a peak performance of 15.4 binary TOPS on the Ultra96 board with a Xilinx UltraScale+\u00a0MPSoC.<\/jats:p>","DOI":"10.1145\/3337929","type":"journal-article","created":{"date-parts":[[2019,8,21]],"date-time":"2019-08-21T11:40:27Z","timestamp":1566387627000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["Optimizing Bit-Serial Matrix Multiplication for Reconfigurable Computing"],"prefix":"10.1145","volume":"12","author":[{"given":"Yaman","family":"Umuroglu","sequence":"first","affiliation":[{"name":"Xilinx Research Labs, Dublin, Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5834-0812","authenticated-orcid":false,"given":"Davide","family":"Conficconi","sequence":"additional","affiliation":[{"name":"Xilinx Research Labs, Ireland and Politecnico di Milano, Milano, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lahiru","family":"Rasnayake","sequence":"additional","affiliation":[{"name":"Norwegian University of Science and Technology, Trondheim, Norway"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Thomas B.","family":"Preusser","sequence":"additional","affiliation":[{"name":"Accemic Technologies GmbH, Dresden, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4232-6976","authenticated-orcid":false,"given":"Magnus","family":"Sj\u00e4lander","sequence":"additional","affiliation":[{"name":"Uppsala University, Sweden and Norwegian University of Science and Technology, Trondheim, Norway"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,8,20]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Joseph James Gebis, Parry Husbands, Kurt Keutzer, David A. Patterson, William Lester Plishker, John Shalf, Samuel Webb Williams, and Katherine A. Yelick.","author":"Asanovi\u0107 Krste","year":"2006"},{"key":"e_1_2_1_2_1","unstructured":"AVNET. 2018. ULTRA96. Retrieved from http:\/\/www.ultra96.org\/sites\/default\/files\/product_briefs\/5354-pb-ultra96-v3b.pdf.  AVNET. 2018. ULTRA96. Retrieved from http:\/\/www.ultra96.org\/sites\/default\/files\/product_briefs\/5354-pb-ultra96-v3b.pdf."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3174243.3174258"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"F. Pedersoli","year":"2018"},{"key":"e_1_2_1_5_1","volume-title":"Quantized neural networks: Training neural networks with low precision weights and activations. arXiv preprint arXiv:1609.07061","author":"Hubara Itay","year":"2016"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2228360.2228584"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2018.2795611"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2014.6927468"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.1982.1653825"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the 2013 International Conference on Reconfigurable Computing and FPGAs (ReConFig\u201913)","author":"Matam Kiran Kumar"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2893356"},{"key":"e_1_2_1_12_1","unstructured":"Wojchech Mula. 2018. Scalar version of SSE move mask instruction. Retrieved from http:\/\/0x80.pl\/articles\/scalar-sse-movmask.html.  Wojchech Mula. 2018. Scalar version of SSE move mask instruction. Retrieved from http:\/\/0x80.pl\/articles\/scalar-sse-movmask.html."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/3195638.3195661"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2068716.2068725"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.761"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.23919\/FPL.2017.8056834"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3020078.3021744"},{"key":"e_1_2_1_18_1","volume-title":"Streamlined deployment for quantized neural networks. arXiv preprint arXiv:1709.04060","author":"Umuroglu Yaman","year":"2017"},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the Conference on Field Programmable Logic and Applications.","author":"Umuroglu Y."},{"key":"e_1_2_1_20_1","volume-title":"HAQ: Hardware-aware automated quantization. arXiv preprint arXiv:1811.08886","author":"Wang Kuan","year":"2018"},{"key":"e_1_2_1_21_1","unstructured":"Xilinx. 2017. Vivado Design Suite User Guide\u2014Release Notes Installation and Licensing (UG973 (v2017.4) ed.). Xilinx.  Xilinx. 2017. Vivado Design Suite User Guide\u2014Release Notes Installation and Licensing (UG973 (v2017.4) ed.). Xilinx."},{"key":"e_1_2_1_22_1","unstructured":"Xilinx. 2018. Python Productivity for Zynq (Pynq) Documentation (release 2.2 ed.). Xilinx.  Xilinx. 2018. Python Productivity for Zynq (Pynq) Documentation (release 2.2 ed.). Xilinx."},{"key":"e_1_2_1_23_1","unstructured":"Xilinx. 2018. UltraScale Architecture and Product Data Sheet: Overview. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/data_sheets\/ds890-ultrascale-overview.pdf.  Xilinx. 2018. UltraScale Architecture and Product Data Sheet: Overview. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/data_sheets\/ds890-ultrascale-overview.pdf."},{"key":"e_1_2_1_24_1","unstructured":"Xilinx. 2018. Zynq UltraScale+ MPSoC Data Sheet: Overview. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/data_sheets\/ds891-zynq-ultrascale-plus-overview.pdf.  Xilinx. 2018. Zynq UltraScale+ MPSoC Data Sheet: Overview. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/data_sheets\/ds891-zynq-ultrascale-plus-overview.pdf."},{"key":"e_1_2_1_25_1","volume-title":"Computer Architecture: Single and Parallel Systems","author":"Zargham Mehdi R.","year":"1996"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3337929","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3337929","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:25:42Z","timestamp":1750206342000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3337929"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,8,20]]},"references-count":25,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2019,9,30]]}},"alternative-id":["10.1145\/3337929"],"URL":"https:\/\/doi.org\/10.1145\/3337929","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,8,20]]},"assertion":[{"value":"2018-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-08-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}