{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T16:12:55Z","timestamp":1781194375046,"version":"3.54.1"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,12,17]],"date-time":"2024-12-17T00:00:00Z","timestamp":1734393600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Intel Corporation funding"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>Conventional multiply-accumulate (MAC) operations have long dominated computation time for deep neural networks (DNNs), especially convolutional neural networks (CNNs). Recently, product quantization (PQ) has been applied to these workloads, replacing MACs with memory lookups to pre-computed dot products. To better understand the efficiency tradeoffs of product-quantized DNNs (PQ-DNNs), we create a custom hardware accelerator to parallelize and accelerate nearest-neighbor search and dot-product lookups. Additionally, we perform an empirical study to investigate the efficiency\u2013accuracy tradeoffs of different PQ parameterizations and training methods. We identify PQ configurations that improve performance-per-area for ResNet20 by up to 3.1\u00d7, even when compared to a highly optimized conventional DNN accelerator, with similar improvements on two additional compact DNNs. When comparing to recent PQ solutions, we outperform prior work by 4\u00d7 in terms of performance-per-area with a 0.6% accuracy degradation. Finally, we reduce the bitwidth of PQ operations to investigate the impact on both hardware efficiency and accuracy. With only 2\u20136-bit precision on three compact DNNs, we were able to maintain DNN accuracy eliminating the need for DSPs.<\/jats:p>","DOI":"10.1145\/3656643","type":"journal-article","created":{"date-parts":[[2024,4,18]],"date-time":"2024-04-18T08:22:40Z","timestamp":1713428560000},"page":"1-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration"],"prefix":"10.1145","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6381-2936","authenticated-orcid":false,"given":"Ahmed","family":"Abouelhamayed","sequence":"first","affiliation":[{"name":"Electrical and Computer Engineering, Cornell University, New York, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9997-1701","authenticated-orcid":false,"given":"Angela","family":"Cui","sequence":"additional","affiliation":[{"name":"Cornell University, Ithaca, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3747-6523","authenticated-orcid":false,"given":"Javier","family":"Fernandez-marques","sequence":"additional","affiliation":[{"name":"Flower Labs, Cambridge, United Kingdom of Great Britain and Northern Ireland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2728-8273","authenticated-orcid":false,"given":"Nicholas","family":"Lane","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Technology, University of Cambridge, Cambridge, United Kingdom of Great Britain and Northern Ireland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4568-8932","authenticated-orcid":false,"given":"Mohamed","family":"Abdelfattah","sequence":"additional","affiliation":[{"name":"ECE, Cornell University, New York, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,12,17]]},"reference":[{"key":"e_1_3_4_2_2","first-page":"411","article-title":"DLA: Compiler and FPGA overlay for neural network inference acceleration","author":"Abdelfattah Mohamed S.","year":"2018","unstructured":"Mohamed S. Abdelfattah, David Han, Andrew Bitar, Roberto Dicecco, Shane O.\u2019Connell, Nitika Shanker, Joseph Chu, Ian Prins, Joshua Fender, Andrew C. Ling, and Gordon R. Chiu. 2018. DLA: Compiler and FPGA overlay for neural network inference acceleration. 28th International Conference on Field Programmable Logic and Applications (FPL) (2018), 411\u20134117.","journal-title":"28th International Conference on Field Programmable Logic and Applications (FPL)"},{"key":"e_1_3_4_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00023"},{"key":"e_1_3_4_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3020078.3021738"},{"key":"e_1_3_4_5_2","first-page":"517","volume-title":"Proceedings of Machine Learning and Systems","volume":"3","author":"Banbury Colby","year":"2021","unstructured":"Colby Banbury, Chuteng Zhou, Igor Fedorov, Ramon Matas, Urmish Thakker, Dibakar Gope, Vijay Janapa Reddi, Matthew Mattina, and Paul Whatmough. 2021. MicroNets: Neural network architectures for deploying tinyML applications on commodity microcontrollers. In Proceedings of Machine Learning and Systems, A. Smola, A. Dimakis, and I. Stoica (Eds.), Vol. 3. 517\u2013532."},{"key":"e_1_3_4_6_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2021-1286"},{"key":"e_1_3_4_7_2","first-page":"129","volume-title":"Proceedings of Machine Learning and Systems","volume":"2","author":"Blalock Davis","year":"2020","unstructured":"Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag. 2020. What is the state of neural network pruning?. In Proceedings of Machine Learning and Systems, I. Dhillon, D. Papailiopoulos, and V. Sze (Eds.), Vol. 2. 129\u2013146. https:\/\/proceedings.mlsys.org\/paper\/2020\/file\/d2ddea18f00665ce8623e36bd4e3c7c5-Paper.pdf"},{"key":"e_1_3_4_8_2","series-title":"Proceedings of Machine Learning Research","first-page":"992","volume-title":"Proceedings of the 38th International Conference on Machine Learning","volume":"139","author":"Blalock Davis","year":"2021","unstructured":"Davis Blalock and John Guttag. 2021. Multiplying matrices without multiplying. In Proceedings of the 38th International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 139), Marina Meila and Tong Zhang (Eds.). PMLR, 992\u20131004."},{"key":"e_1_3_4_9_2","doi-asserted-by":"publisher","unstructured":"Andrew Brock Theodore Lim J. M. Ritchie and Nick Weston. 2016. Neural Photo Editing with Introspective Adversarial Networks. 10.48550\/ARXIV.1609.07093","DOI":"10.48550\/ARXIV.1609.07093"},{"key":"e_1_3_4_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00154"},{"key":"e_1_3_4_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2910232"},{"key":"e_1_3_4_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/ijcnn.2017.7966217"},{"key":"e_1_3_4_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1980.1163420"},{"key":"e_1_3_4_14_2","doi-asserted-by":"publisher","unstructured":"Alaaeldin El-Nouby Matthew J. Muckley Karen Ullrich Ivan Laptev Jakob Verbeek and Herv\u00e9 J\u00e9gou. 2022. Image Compression with Product Quantized Masked Image Modeling. 10.48550\/ARXIV.2212.07372","DOI":"10.48550\/ARXIV.2212.07372"},{"key":"e_1_3_4_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00050"},{"key":"e_1_3_4_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.240"},{"key":"e_1_3_4_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/2909437.2909443"},{"key":"e_1_3_4_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373087.3375380"},{"key":"e_1_3_4_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr.2016.90"},{"key":"e_1_3_4_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_4_21_2","unstructured":"Yihui He Xiangyu Zhang and Jian Sun. 2017. Channel Pruning for Accelerating Very Deep Neural Networks. arxiv:1707.06168 [cs.CV]"},{"key":"e_1_3_4_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr.2018.00286"},{"key":"e_1_3_4_23_2","doi-asserted-by":"publisher","unstructured":"Eric Jang Shixiang Gu and Ben Poole. 2016. Categorical Reparameterization with Gumbel-Softmax. 10.48550\/ARXIV.1611.01144","DOI":"10.48550\/ARXIV.1611.01144"},{"key":"e_1_3_4_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00348"},{"key":"e_1_3_4_25_2","doi-asserted-by":"publisher","unstructured":"Pranav Jeevan Kavitha Viswanathan Anandu A. S. and Amit Sethi. 2022. WaveMix: A Resource-efficient Neural Network for Image Analysis. 10.48550\/ARXIV.2205.14375","DOI":"10.48550\/ARXIV.2205.14375"},{"key":"e_1_3_4_26_2","volume-title":"Learning Semantic Image Representations at a Large Scale","author":"Jia Yangqing","year":"2014","unstructured":"Yangqing Jia. 2014. Learning Semantic Image Representations at a Large Scale. Ph. D. Dissertation. EECS Department, University of California, Berkeley."},{"key":"e_1_3_4_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3140659.3080246"},{"key":"e_1_3_4_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.57"},{"key":"e_1_3_4_29_2","doi-asserted-by":"publisher","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Optimization. 10.48550\/ARXIV.1412.6980","DOI":"10.48550\/ARXIV.1412.6980"},{"key":"e_1_3_4_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00518"},{"key":"e_1_3_4_31_2","unstructured":"Alex Krizhevsky. 2009. Learning multiple layers of features from tiny images. University of Toronto."},{"key":"e_1_3_4_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr.2016.435"},{"key":"e_1_3_4_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCAS.2019.8702105"},{"key":"e_1_3_4_34_2","doi-asserted-by":"publisher","unstructured":"Calvin McCarter and Nicholas Dronen. 2022. Look-Ups are Not (yet) all You Need for Deep Learning Inference. 10.48550\/ARXIV.2207.05808","DOI":"10.48550\/ARXIV.2207.05808"},{"key":"e_1_3_4_35_2","volume-title":"International Conference on Learning Representations","author":"Mehta Sachin","year":"2022","unstructured":"Sachin Mehta and Mohammad Rastegari. 2022. MobileViT: Light-weight, general-purpose, and mobile-friendly vision transformer. In International Conference on Learning Representations."},{"key":"e_1_3_4_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20083-0_18"},{"key":"e_1_3_4_37_2","first-page":"8026","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems (NeurIPS). 8026\u20138037."},{"key":"e_1_3_4_38_2","doi-asserted-by":"publisher","unstructured":"Jie Ran Rui Lin Jason Chun Lok Li Jiajun Zhou and Ngai Wong. 2022. PECAN: A Product-Quantized Content Addressable Memory Network. 10.48550\/ARXIV.2208.13571","DOI":"10.48550\/ARXIV.2208.13571"},{"key":"e_1_3_4_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPT.2014.7082748"},{"key":"e_1_3_4_40_2","article-title":"XNOR-Net: ImageNet classification using binary convolutional neural networks","volume":"1603","author":"Rastegari Mohammad","year":"2016","unstructured":"Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. 2016. XNOR-Net: ImageNet classification using binary convolutional neural networks. CoRR abs\/1603.05279 (2016). arxiv:1603.05279","journal-title":"CoRR"},{"key":"e_1_3_4_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr.2018.00474"},{"key":"e_1_3_4_42_2","volume-title":"International Conference on Learning Representations (ICLR)","author":"Stock Pierre","year":"2020","unstructured":"Pierre Stock, Armand Joulin, R\u00e9mi Gribonval, Benjamin Graham, and Herv\u00e9 J\u00e9gou. 2020. And the bit goes down: Revisiting the quantization of neural networks. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_4_43_2","unstructured":"Mingxing Tan and Quoc V. Le. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. arxiv:1905.11946"},{"key":"e_1_3_4_44_2","series-title":"Proceedings of Machine Learning Research","first-page":"4985","volume-title":"Proceedings of the 35th International Conference on Machine Learning","volume":"80","author":"Tschannen Michael","year":"2018","unstructured":"Michael Tschannen, Aran Khanna, and Animashree Anandkumar. 2018. StrassenNets: Deep learning with a multiplication budget. In Proceedings of the 35th International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, Stockholmsm\u00e4ssan, Stockholm Sweden, 4985\u20134994."},{"key":"e_1_3_4_45_2","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc."},{"key":"e_1_3_4_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00881"},{"key":"e_1_3_4_47_2","article-title":"Speech commands: A dataset for limited-vocabulary speech recognition","author":"Warden Pete","year":"2018","unstructured":"Pete Warden. 2018. Speech commands: A dataset for limited-vocabulary speech recognition. arXiv preprint arXiv:1804.03209 (2018).","journal-title":"arXiv preprint arXiv:1804.03209"},{"key":"e_1_3_4_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2019.00030"},{"key":"e_1_3_4_49_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_12"},{"key":"e_1_3_4_50_2","volume-title":"International Conference on Learning Representations","author":"Zhang Jingzhao","year":"2020","unstructured":"Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie. 2020. Why gradient clipping accelerates training: A theoretical justification for adaptivity. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=BJgnXpVYwS"},{"key":"e_1_3_4_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2018.00041"},{"key":"e_1_3_4_52_2","unstructured":"Yundong Zhang Naveen Suda Liangzhen Lai and Vikas Chandra. 2018. Hello Edge: Keyword Spotting on Microcontrollers. arxiv:1711.07128 [cs.SD]"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3656643","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3656643","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T23:57:31Z","timestamp":1750291051000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3656643"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,17]]},"references-count":51,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3656643"],"URL":"https:\/\/doi.org\/10.1145\/3656643","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,17]]},"assertion":[{"value":"2023-12-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-17","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}