{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T17:03:17Z","timestamp":1786554197921,"version":"3.56.0"},"reference-count":36,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2021,6,30]],"date-time":"2021-06-30T00:00:00Z","timestamp":1625011200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2018YFC1002600"],"award-info":[{"award-number":["2018YFC1002600"]}]},{"name":"Science and Technology Planning Project of Guangdong Province, China","award":["2017B090904034, 2017B030314109, 2018B090944002, and 2019B020230003"],"award-info":[{"award-number":["2017B090904034, 2017B030314109, 2018B090944002, and 2019B020230003"]}]},{"name":"Guangdong Peak Project","award":["DFJH201802"],"award-info":[{"award-number":["DFJH201802"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62006050"],"award-info":[{"award-number":["62006050"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["J. Emerg. Technol. Comput. Syst."],"published-print":{"date-parts":[[2021,10,31]]},"abstract":"<jats:p>Deep neural networks have demonstrated their great potential in recent years, exceeding the performance of human experts in a wide range of applications. Due to their large sizes, however, compression techniques such as weight quantization and pruning are usually applied before they can be accommodated on the edge. It is generally believed that quantization leads to performance degradation, and plenty of existing works have explored quantization strategies aiming at minimum accuracy loss. In this paper, we argue that quantization, which essentially imposes regularization on weight representations, can sometimes help to improve accuracy. We conduct comprehensive experiments on three widely used applications: fully connected network for biomedical image segmentation, convolutional neural network for image classification on ImageNet, and recurrent neural network for automatic speech recognition, and experimental results show that quantization can improve the accuracy by 1%, 1.95%, 4.23% on the three applications respectively with 3.5x-6.4x memory reduction.<\/jats:p>","DOI":"10.1145\/3451211","type":"journal-article","created":{"date-parts":[[2021,6,30]],"date-time":"2021-06-30T20:40:58Z","timestamp":1625085658000},"page":"1-11","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":23,"title":["Quantization of Deep Neural Networks for Accurate Edge Computing"],"prefix":"10.1145","volume":"17","author":[{"given":"Wentao","family":"Chen","sequence":"first","affiliation":[{"name":"Guangdong Academy of Medical Sciences, Guangzhou, Guangdong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hailong","family":"Qiu","sequence":"additional","affiliation":[{"name":"Guangdong Academy of Medical Sciences, Guangzhou, Guangdong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jian","family":"Zhuang","sequence":"additional","affiliation":[{"name":"Guangdong Academy of Medical Sciences, Guangzhou, Guangdong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chutong","family":"Zhang","sequence":"additional","affiliation":[{"name":"Easylink Technology Co., Ltd., Wuhan, Hubei"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu","family":"Hu","sequence":"additional","affiliation":[{"name":"Easylink Technology Co., Ltd., Wuhan, Hubei"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qing","family":"Lu","sequence":"additional","affiliation":[{"name":"University of Notre Dame, Notre Dame, IN"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tianchen","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Notre Dame, Notre Dame, IN"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yiyu","family":"Shi","sequence":"additional","affiliation":[{"name":"University of Notre Dame, Notre Dame, IN"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Meiping","family":"Huang","sequence":"additional","affiliation":[{"name":"Guangdong Academy of Medical Sciences, Guangzhou, Guangdong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1046-6379","authenticated-orcid":false,"given":"Xiaowe","family":"Xu","sequence":"additional","affiliation":[{"name":"Guangdong Academy of Medical Sciences, Guangzhou, Guangdong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,6,30]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2019.2902141"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/3015812.3015985"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.273"},{"key":"e_1_2_1_4_1","unstructured":"M. Courbariaux and Y. Bengio. 2016. BinaryNet: Training deep neural networks with weights and activations constrained to+ 1 or -1. arXiv:1602.02830.  M. Courbariaux and Y. Bengio. 2016. BinaryNet: Training deep neural networks with weights and activations constrained to+ 1 or -1. arXiv:1602.02830."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969442.2969588"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_7_1","unstructured":"Yukun Ding Jinglan Liu and Yiyu Shi. 2018. On the universal approximability of quantized ReLU neural networks. arXiv:1802.03646.  Yukun Ding Jinglan Liu and Yiyu Shi. 2018. On the universal approximability of quantized ReLU neural networks. arXiv:1802.03646."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.24251\/HICSS.2017.133"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11427-016-0389-9"},{"key":"e_1_2_1_10_1","volume-title":"TIMIT acoustic-phonetic continuous speech corpus. Web download","author":"Garofolo John S.","unstructured":"John S. Garofolo . 1993. TIMIT acoustic-phonetic continuous speech corpus. Web download . Linguistic Data Consortium , Philadelphia, PA . John S. Garofolo. 1993. TIMIT acoustic-phonetic continuous speech corpus. Web download. Linguistic Data Consortium, Philadelphia, PA."},{"key":"e_1_2_1_11_1","volume-title":"Dally","author":"Han Song","year":"2015","unstructured":"Song Han , Huizi Mao , and William J . Dally . 2015 . Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. arXiv:1510.00149. Song Han, Huizi Mao, and William J. Dally. 2015. Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. arXiv:1510.00149."},{"key":"e_1_2_1_12_1","volume-title":"et\u00a0al","author":"Hannun Awni","year":"2014","unstructured":"Awni Hannun , Carl Case , Jared Casper , Bryan Catanzaro , Greg Diamos , Erich Elsen , Ryan Prenger , et\u00a0al . 2014 . Deep speech: Scaling up end-to-end speech recognition. arXiv:1412.5567. Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, et\u00a0al. 2014. Deep speech: Scaling up end-to-end speech recognition. arXiv:1412.5567."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/3122009.3242044"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00286"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.59"},{"key":"e_1_2_1_16_1","unstructured":"Fengfu Li Bo Zhang and Bin Liu. 2016. Ternary weight networks. arXiv:1605.04711.  Fengfu Li Bo Zhang and Bin Liu. 2016. Ternary weight networks. arXiv:1605.04711."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01297"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_32"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_2_1_20_1","volume-title":"Improving photo search: A step across the semantic gap. Google Research Blog 12","author":"Rosenberg Chuck","year":"2013","unstructured":"Chuck Rosenberg . 2013. Improving photo search: A step across the semantic gap. Google Research Blog 12 ( 2013 ). Chuck Rosenberg. 2013. Improving photo search: A step across the semantic gap. Google Research Blog 12 (2013)."},{"key":"e_1_2_1_21_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556.  Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2016.08.008"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2670313"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00460"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-59725-2_43"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41928-018-0059-3"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3264817"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/3199700.3199821"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00866"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-32245-8_53"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-59719-1_8"},{"key":"e_1_2_1_32_1","volume-title":"Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention. 399\u2013407","author":"Yang Lin","unstructured":"Lin Yang , Yizhe Zhang , Jianxu Chen , Siyuan Zhang , and Danny Z. Chen . 2017. Suggestive annotation: A deep active learning framework for biomedical image segmentation . In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention. 399\u2013407 . Lin Yang, Yizhe Zhang, Jianxu Chen, Siyuan Zhang, and Danny Z. Chen. 2017. Suggestive annotation: A deep active learning framework for biomedical image segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention. 399\u2013407."},{"key":"e_1_2_1_33_1","first-page":"1","article-title":"CarcinoPred-EL: Novel models for predicting the carcinogenicity of chemicals using molecular fingerprints and ensemble learning methods","volume":"7","author":"Zhang Li","year":"2017","unstructured":"Li Zhang , Haixin Ai , Wen Chen , Zimo Yin , Huan Hu , Junfeng Zhu , Jian Zhao , Qi Zhao , and Hongsheng Liu . 2017 . CarcinoPred-EL: Novel models for predicting the carcinogenicity of chemicals using molecular fingerprints and ensemble learning methods . Scientific Reports 7 , 1 (2017), 1 \u2013 14 . Li Zhang, Haixin Ai, Wen Chen, Zimo Yin, Huan Hu, Junfeng Zhu, Jian Zhao, Qi Zhao, and Hongsheng Liu. 2017. CarcinoPred-EL: Novel models for predicting the carcinogenicity of chemicals using molecular fingerprints and ensemble learning methods. Scientific Reports 7, 1 (2017), 1\u201314.","journal-title":"Scientific Reports"},{"key":"e_1_2_1_34_1","unstructured":"Aojun Zhou Anbang Yao Yiwen Guo Lin Xu and Yurong Chen. 2017. Incremental network quantization: Towards lossless CNNs with low-precision weights. arXiv:1702.03044.  Aojun Zhou Anbang Yao Yiwen Guo Lin Xu and Yurong Chen. 2017. Incremental network quantization: Towards lossless CNNs with low-precision weights. arXiv:1702.03044."},{"key":"e_1_2_1_35_1","unstructured":"Shuchang Zhou Yuxin Wu Zekun Ni Xinyu Zhou He Wen and Yuheng Zou. 2016. DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients. arXiv:1606.06160.  Shuchang Zhou Yuxin Wu Zekun Ni Xinyu Zhou He Wen and Yuheng Zou. 2016. DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients. arXiv:1606.06160."},{"key":"e_1_2_1_36_1","volume-title":"Dally","author":"Zhu Chenzhuo","year":"2016","unstructured":"Chenzhuo Zhu , Song Han , Huizi Mao , and William J . Dally . 2016 . Trained ternary quantization. arXiv:1612.01064. Chenzhuo Zhu, Song Han, Huizi Mao, and William J. Dally. 2016. Trained ternary quantization. arXiv:1612.01064."}],"container-title":["ACM Journal on Emerging Technologies in Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3451211","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3451211","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T17:49:25Z","timestamp":1750268965000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3451211"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,6,30]]},"references-count":36,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2021,10,31]]}},"alternative-id":["10.1145\/3451211"],"URL":"https:\/\/doi.org\/10.1145\/3451211","relation":{},"ISSN":["1550-4832","1550-4840"],"issn-type":[{"value":"1550-4832","type":"print"},{"value":"1550-4840","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,6,30]]},"assertion":[{"value":"2020-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-06-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}