{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T22:28:44Z","timestamp":1775514524992,"version":"3.50.1"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2022,1,28]],"date-time":"2022-01-28T00:00:00Z","timestamp":1643328000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"C-BRIC"},{"name":"one of six centers in JUMP"},{"name":"Semiconductor Research Corporation (SRC) program"},{"DOI":"10.13039\/100000185","name":"DARPA","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100000185","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2022,5,31]]},"abstract":"<jats:p>Precision scaling has emerged as a popular technique to optimize the compute and storage requirements of Deep Neural Networks (DNNs). Efforts toward creating ultra-low-precision (sub-8-bit) DNNs for efficient inference suggest that the minimum precision required to achieve a given network-level accuracy varies considerably across networks, and even across layers within a network. This translates to a need to support variable precision computation in DNN hardware. Previous proposals for precision-reconfigurable hardware, such as bit-serial architectures, incur high overheads, significantly diminishing the benefits of lower precision. We propose Ax-BxP, a method for approximate blocked computation wherein each multiply-accumulate operation is performed block-wise (a block is a group of bits), facilitating re-configurability at the granularity of blocks. Further, approximations are introduced by only performing a subset of the required block-wise computations to realize precision re-configurability with high efficiency. We design a DNN accelerator that embodies approximate blocked computation and propose a method to determine a suitable approximation configuration for any given DNN. For the AlexNet, ResNet50, and MobileNetV2 DNNs, Ax-BxP achieves improvement in system energy and performance, respectively, over an 8-bit fixed-point (FxP8) baseline, with minimal loss (&lt;1% on average) in classification accuracy. Further, by varying the approximation configurations at a finer granularity across layers and data-structures within a DNN, we achieve improvement in system energy and performance, respectively.<\/jats:p>","DOI":"10.1145\/3492733","type":"journal-article","created":{"date-parts":[[2022,1,28]],"date-time":"2022-01-28T10:14:07Z","timestamp":1643364847000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Ax-BxP: Approximate Blocked Computation for Precision-reconfigurable Deep Neural Network Acceleration"],"prefix":"10.1145","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9646-5088","authenticated-orcid":false,"given":"Reena","family":"Elangovan","sequence":"first","affiliation":[{"name":"Purdue University, West Lafayette, IN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2291-7712","authenticated-orcid":false,"given":"Shubham","family":"Jain","sequence":"additional","affiliation":[{"name":"IBM T. J. Watson Research Center, Yorktown Heights, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4624-564X","authenticated-orcid":false,"given":"Anand","family":"Raghunathan","sequence":"additional","affiliation":[{"name":"Purdue University, West Lafayette, IN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,1,28]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCSE.2016.74"},{"key":"e_1_3_2_3_2","unstructured":"Dario Amodei T. Brown et al. 2020. Language models are few-shot learners. arXiv:arXiv:2005.14165."},{"key":"e_1_3_2_4_2","first-page":"6105","volume-title":"Proceedings of the 36th International Conference on Machine Learning","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc Le. 2019. EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning. 6105\u20136114."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASPDAC.2016.7428029"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2017.2761740"},{"issue":"2","key":"e_1_3_2_7_2","article-title":"In-datacenter performance analysis of a tensor processing unit","volume":"45","author":"al Doe Hyun, N. P. Jouppi et","year":"2017","unstructured":"Doe Hyun, N. P. Jouppi et al. 2017. In-datacenter performance analysis of a tensor processing unit. SIGARCH Comput. Archit. News 45, 2 (2017). DOI: https:\/\/doi.org\/10.1145\/3140659.3080246","journal-title":"SIGARCH Comput. Archit. News"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00881"},{"key":"e_1_3_2_9_2","unstructured":"Asit Mishra Eriko Nurvitadhi Jeffrey J. Cook and Debbie Marr. 2017. WRPN: Wide reduced-precision networks. arXiv:arXiv:1709.01134."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11390-017-1750-y"},{"key":"e_1_3_2_11_2","unstructured":"Jungwook Choi Zhuo Wang Swagath Venkataramani Pierce I.-Jen Chuang Vijayalakshmi Srinivasan and Kailash Gopalakrishnan. 2018. Pact: Parameterized clipping activation for quantized neural networks. arXiv:arXiv:1805.06085."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3316781.3317783"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/2925426.2926294"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.5555\/3130379.3130725"},{"key":"e_1_3_2_15_2","doi-asserted-by":"crossref","unstructured":"Yaman Umuroglu Lahiru Rasnayake and Magnus Sjalander. 2018. Bismo: A scalable bit-serial matrix multiplication overlay for reconfigurable computing. In Proceedings of the 28th International Conference on Field Programmable Logic and Applications (FPL) . 3007\u20133077. DOI:10.1109\/FPL.2018.00059","DOI":"10.1109\/FPL.2018.00059"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.5555\/3195638.3195661"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3195970.3196072"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00069"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2950386"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/2627369.2627613"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACSSC.2013.6810241"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.5555\/554321"},{"key":"e_1_3_2_23_2","article-title":"DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients","author":"al. S. Zhou et","year":"2016","unstructured":"S. Zhou et al.2016. DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients. arXiv preprint arXiv:1606.06160.","journal-title":"arXiv preprint arXiv:1606.06160"},{"key":"e_1_3_2_24_2","unstructured":"Ananda Samajdar Yuhao Zhu Paul Whatmough Matthew Mattina and Tushar Krishna. 2019. SCALE-Sim: Systolic CNN accelerator. arXiv:arXiv:1811.02883."},{"key":"e_1_3_2_25_2","volume-title":"CACTI 6.0: A Tool to Model Large Caches","author":"Jouppi. Naveen Muralimanohar, Rajeev Balasubramonian, and Norman P.","year":"2009","unstructured":"Naveen Muralimanohar, Rajeev Balasubramonian, and Norman P. Jouppi.2009. CACTI 6.0: A Tool to Model Large Caches. Technical Report. HP Laboratories."},{"key":"e_1_3_2_26_2","unstructured":"Kuan Wang Zhijian Liu Yujun Lin Ji Lin and Song Han. 2019. Retrieved from https:\/\/github.com\/mit-han-lab\/haq."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.5555\/3122009.3242044"},{"key":"e_1_3_2_28_2","unstructured":"Fengfu Li Bo Zhang and Bin Liu. 2016. Ternary weight networks. arXiv:arXiv:1605.04711."},{"key":"e_1_3_2_29_2","unstructured":"Matthieu Courbariaux Itay Hubara Daniel Soudry Ran El-Yaniv and Yoshua Bengio. 2016. BinaryNet: Training deep neural networks with weights and activations constrained to+ 1 or- 1.arXiv:arXiv:1602.02830."},{"key":"e_1_3_2_30_2","volume-title":"Computer Vision \u2013 ECCV","author":"Rastegari Mohammad","year":"2016","unstructured":"Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. 2016. XNOR-Net: ImageNet classification using binary convolutional neural networks. In Computer Vision \u2013 ECCV. Springer."},{"key":"e_1_3_2_31_2","unstructured":"Song Han Huizi Mao and William J. Dally. 2016. Deep compression: Compressing deep neural networks with pruning trained quantization and Huffman coding. In Proceedings of the 4th International Conference on Learning Representations (ICLR) . http:\/\/arxiv.org\/abs\/1510.00149"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3195970.3196012"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2009.2012863"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/2228360.2228504"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TEVC.2014.2336175"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.5555\/2132325.2132474"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/1403375.1403679"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2009.2020591"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.5555\/2016802.2016898"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.5555\/3130379.3130382"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/VLSID.2011.51"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD.2013.6657022"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2011.09.039"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3097264"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2000.862012"},{"key":"e_1_3_2_46_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference of Electron Devices and Solid-State Circuits (EDSSC)","author":"Kyaw Khaing Yin","year":"2010","unstructured":"Khaing Yin Kyaw, Wang Ling Goh, and Kiat Seng Yeo. 2010. Low-power high-speed multiplier for error-tolerant application. In Proceedings of the IEEE International Conference of Electron Devices and Solid-State Circuits (EDSSC). 1\u20134."},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/VLSISP.1993.404467"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/82.486455"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/82.769795"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2014.2333366"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.5555\/2840819.2840878"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2016.2535398"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/92.845894"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3492733","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3492733","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3492733","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:53Z","timestamp":1750188653000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3492733"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,28]]},"references-count":52,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,5,31]]}},"alternative-id":["10.1145\/3492733"],"URL":"https:\/\/doi.org\/10.1145\/3492733","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,28]]},"assertion":[{"value":"2021-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-01-28","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}