{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,17]],"date-time":"2025-09-17T16:27:16Z","timestamp":1758126436704,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":52,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T00:00:00Z","timestamp":1634428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"NSF","award":["1925717,1720256"],"award-info":[{"award-number":["1925717,1720256"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,18]]},"DOI":"10.1145\/3466752.3480090","type":"proceedings-article","created":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T19:12:05Z","timestamp":1634497925000},"page":"1309-1322","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["ENMC: Extreme Near-Memory Classification via Approximate Screening"],"prefix":"10.1145","author":[{"given":"Liu","family":"Liu","sequence":"first","affiliation":[{"name":"University of California, Santa Barbara, United States"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jilan","family":"Lin","sequence":"additional","affiliation":[{"name":"University of California, Santa Barbara, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zheng","family":"Qu","sequence":"additional","affiliation":[{"name":"University of California, Santa Barbara, United States"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yufei","family":"Ding","sequence":"additional","affiliation":[{"name":"University of California, Santa Barbara, United States"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuan","family":"Xie","sequence":"additional","affiliation":[{"name":"University of California, Santa Barbara, United States"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,17]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/375551.375608"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750386"},{"key":"e_1_3_2_1_3_1","volume-title":"Chameleon: Versatile and practical near-DRAM acceleration architecture for large memory systems. In 2016 49th annual IEEE\/ACM international symposium on Microarchitecture (MICRO)","author":"Asghari-Moghaddam Hadi","year":"2016","unstructured":"Hadi Asghari-Moghaddam , Young\u00a0Hoon Son , Jung\u00a0Ho Ahn , and Nam\u00a0Sung Kim . 2016 . Chameleon: Versatile and practical near-DRAM acceleration architecture for large memory systems. In 2016 49th annual IEEE\/ACM international symposium on Microarchitecture (MICRO) . IEEE , 1\u201313. Hadi Asghari-Moghaddam, Young\u00a0Hoon Son, Jung\u00a0Ho Ahn, and Nam\u00a0Sung Kim. 2016. Chameleon: Versatile and practical near-DRAM acceleration architecture for large memory systems. In 2016 49th annual IEEE\/ACM international symposium on Microarchitecture (MICRO). IEEE, 1\u201313."},{"key":"e_1_3_2_1_4_1","unstructured":"K. Bhatia K. Dahiya H. Jain P. Kar A. Mittal Y. Prabhu and M. Varma. 2016. The extreme classification repository: Multi-label datasets and code. http:\/\/manikvarma.org\/downloads\/XC\/XMLRepository.html  K. Bhatia K. Dahiya H. Jain P. Kar A. Mittal Y. Prabhu and M. Varma. 2016. The extreme classification repository: Multi-label datasets and code. http:\/\/manikvarma.org\/downloads\/XC\/XMLRepository.html"},{"key":"e_1_3_2_1_5_1","volume-title":"International Conference on Learning Representations (ICLR). arxiv:1810","author":"Chen Pei\u00a0Hung","year":"2019","unstructured":"Pei\u00a0Hung Chen , Si Si , Sanjiv Kumar , Yang Li , and Cho\u00a0Jui Hsieh . 2019 . Learning to screen for fast softmax inference on large vocabulary neural networks . In International Conference on Learning Representations (ICLR). arxiv:1810 .12406https:\/\/openreview.net\/forum?id=ByeMB3Act7 Pei\u00a0Hung Chen, Si Si, Sanjiv Kumar, Yang Li, and Cho\u00a0Jui Hsieh. 2019. Learning to screen for fast softmax inference on large vocabulary neural networks. In International Conference on Learning Representations (ICLR). arxiv:1810.12406https:\/\/openreview.net\/forum?id=ByeMB3Act7"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2016.2616357"},{"key":"e_1_3_2_1_7_1","volume-title":"Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan.","author":"Choi Jungwook","year":"2018","unstructured":"Jungwook Choi , Zhuo Wang , Swagath Venkataramani , Pierce I-Jen Chuang , Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan. 2018 . Pact : Parameterized clipping activation for quantized neural networks. arXiv preprint arXiv:1805.06085(2018). Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan. 2018. Pact: Parameterized clipping activation for quantized neural networks. arXiv preprint arXiv:1805.06085(2018)."},{"key":"e_1_3_2_1_8_1","unstructured":"Xiaoliang Dai Hongxu Yin and Niraj\u00a0K Jha. 2018. Grow and prune compact fast and accurate LSTMs. arXiv preprint arXiv:1805.11797(2018).  Xiaoliang Dai Hongxu Yin and Niraj\u00a0K Jha. 2018. Grow and prune compact fast and accurate LSTMs. arXiv preprint arXiv:1805.11797(2018)."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2015.7056040"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00071"},{"key":"e_1_3_2_1_11_1","volume-title":"DLUX: a LUT-based Near-Bank Accelerator for Data Center Deep Learning Training Workloads","author":"Gu Peng","year":"2020","unstructured":"Peng Gu , Xinfeng Xie , Shuangchen Li , Dimin Niu , Hongzhong Zheng , Krishna\u00a0 T Malladi , and Yuan Xie . 2020. DLUX: a LUT-based Near-Bank Accelerator for Data Center Deep Learning Training Workloads . IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems ( 2020 ). Peng Gu, Xinfeng Xie, Shuangchen Li, Dimin Niu, Hongzhong Zheng, Krishna\u00a0T Malladi, and Yuan Xie. 2020. DLUX: a LUT-based Near-Bank Accelerator for Data Center Deep Learning Training Workloads. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (2020)."},{"key":"e_1_3_2_1_12_1","unstructured":"Song Han Huizi Mao and William\u00a0J Dally. 2015. Deep compression: Compressing deep neural networks with pruning trained quantization and huffman coding. arXiv preprint arXiv:1510.00149(2015).  Song Han Huizi Mao and William\u00a0J Dally. 2015. Deep compression: Compressing deep neural networks with pruning trained quantization and huffman coding. arXiv preprint arXiv:1510.00149(2015)."},{"key":"e_1_3_2_1_13_1","unstructured":"Song Han Jeff Pool John Tran and William Dally. 2015. Learning both weights and connections for efficient neural network. In Advances in neural information processing systems. 1135\u20131143.  Song Han Jeff Pool John Tran and William Dally. 2015. Learning both weights and connections for efficient neural network. In Advances in neural information processing systems. 1135\u20131143."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_2_1_15_1","volume-title":"RecNMP: Accelerating Personalized Recommendation with Near-Memory Processing. In 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). 790\u2013803","author":"Ke Liu","year":"2020","unstructured":"Liu Ke , Udit Gupta , Benjamin\u00a0Youngjae Cho , David Brooks , Vikas Chandra , Utku Diril , Amin Firoozshahian , Kim Hazelwood , Bill Jia , Hsien-Hsin\u00a0 S. Lee , Meng Li , Bert Maher , Dheevatsa Mudigere , Maxim Naumov , Martin Schatz , Mikhail Smelyanskiy , Xiaodong Wang , Brandon Reagen , Carole-Jean Wu , Mark Hempstead , and Xuan Zhang . 2020 . RecNMP: Accelerating Personalized Recommendation with Near-Memory Processing. In 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). 790\u2013803 . https:\/\/doi.org\/10.1109\/ISCA45697.2020.00070 10.1109\/ISCA45697.2020.00070 Liu Ke, Udit Gupta, Benjamin\u00a0Youngjae Cho, David Brooks, Vikas Chandra, Utku Diril, Amin Firoozshahian, Kim Hazelwood, Bill Jia, Hsien-Hsin\u00a0S. Lee, Meng Li, Bert Maher, Dheevatsa Mudigere, Maxim Naumov, Martin Schatz, Mikhail Smelyanskiy, Xiaodong Wang, Brandon Reagen, Carole-Jean Wu, Mark Hempstead, and Xuan Zhang. 2020. RecNMP: Accelerating Personalized Recommendation with Near-Memory Processing. In 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). 790\u2013803. https:\/\/doi.org\/10.1109\/ISCA45697.2020.00070"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132402.3132426"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126965"},{"key":"e_1_3_2_1_18_1","volume-title":"Ramulator: A fast and extensible DRAM simulator","author":"Kim Yoongu","year":"2015","unstructured":"Yoongu Kim , Weikun Yang , and Onur Mutlu . 2015 . Ramulator: A fast and extensible DRAM simulator . IEEE Computer architecture letters 15, 1 (2015), 45\u201349. Yoongu Kim, Weikun Yang, and Onur Mutlu. 2015. Ramulator: A fast and extensible DRAM simulator. IEEE Computer architecture letters 15, 1 (2015), 45\u201349."},{"key":"e_1_3_2_1_19_1","volume-title":"Gen-Z Chipsetfor Exascale Fabrics. In 2019 IEEE Hot Chips 31 Symposium (HCS). IEEE Computer Society, 1\u201322","author":"Knebel Patrick","year":"2019","unstructured":"Patrick Knebel , Dan Berkram , Al Davis , Darel Emmot , Paolo Faraboschi , and Gary Gostin . 2019 . Gen-Z Chipsetfor Exascale Fabrics. In 2019 IEEE Hot Chips 31 Symposium (HCS). IEEE Computer Society, 1\u201322 . Patrick Knebel, Dan Berkram, Al Davis, Darel Emmot, Paolo Faraboschi, and Gary Gostin. 2019. Gen-Z Chipsetfor Exascale Fabrics. In 2019 IEEE Hot Chips 31 Symposium (HCS). IEEE Computer Society, 1\u201322."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358284"},{"key":"e_1_3_2_1_21_1","volume-title":"Tensor Casting: Co-Designing Algorithm-Architecture for Personalized Recommendation Training. arXiv preprint arXiv:2010.13100(2020).","author":"Kwon Youngeun","year":"2020","unstructured":"Youngeun Kwon , Yunjae Lee , and Minsoo Rhu . 2020 . Tensor Casting: Co-Designing Algorithm-Architecture for Personalized Recommendation Training. arXiv preprint arXiv:2010.13100(2020). Youngeun Kwon, Yunjae Lee, and Minsoo Rhu. 2020. Tensor Casting: Co-Designing Algorithm-Architecture for Personalized Recommendation Training. arXiv preprint arXiv:2010.13100(2020)."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC42613.2021.9365862"},{"key":"e_1_3_2_1_23_1","volume-title":"Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668(2020).","author":"Lepikhin Dmitry","year":"2020","unstructured":"Dmitry Lepikhin , HyoukJoong Lee , Yuanzhong Xu , Dehao Chen , Orhan Firat , Yanping Huang , Maxim Krikun , Noam Shazeer , and Zhifeng Chen . 2020 . Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668(2020). Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668(2020)."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3080834"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS51385.2021.00033"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2507157.2507163"},{"key":"e_1_3_2_1_27_1","unstructured":"Tharun Medini Qixuan Huang Yiqiu Wang Vijai Mohan and Anshumali Shrivastava. 2019. Extreme classification in log memory using count-min sketch: A case study of amazon search with 50m products. In Neural Information Processing Systems (NeurIPS). arxiv:1910.13830  Tharun Medini Qixuan Huang Yiqiu Wang Vijai Mohan and Anshumali Shrivastava. 2019. Extreme classification in log memory using count-min sketch: A case study of amazon search with 50m products. In Neural Information Processing Systems (NeurIPS). arxiv:1910.13830"},{"key":"e_1_3_2_1_28_1","unstructured":"Stephen Merity Nitish\u00a0Shirish Keskar and Richard Socher. 2017. Regularizing and optimizing LSTM language models. arXiv preprint arXiv:1708.02182(2017).  Stephen Merity Nitish\u00a0Shirish Keskar and Richard Socher. 2017. Regularizing and optimizing LSTM language models. arXiv preprint arXiv:1708.02182(2017)."},{"key":"e_1_3_2_1_29_1","unstructured":"Stephen Merity Caiming Xiong James Bradbury and Richard Socher. 2016. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843(2016).  Stephen Merity Caiming Xiong James Bradbury and Richard Socher. 2016. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843(2016)."},{"key":"e_1_3_2_1_30_1","unstructured":"Tomas Mikolov Ilya Sutskever Kai Chen Greg Corrado and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. arXiv preprint arXiv:1310.4546(2013).  Tomas Mikolov Ilya Sutskever Kai Chen Greg Corrado and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. arXiv preprint arXiv:1310.4546(2013)."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2018.00077"},{"key":"e_1_3_2_1_32_1","unstructured":"Sharan Narang Erich Elsen Gregory Diamos and Shubho Sengupta. 2017. Exploring sparsity in recurrent neural networks. arXiv preprint arXiv:1704.05119(2017).  Sharan Narang Erich Elsen Gregory Diamos and Shubho Sengupta. 2017. Exploring sparsity in recurrent neural networks. arXiv preprint arXiv:1704.05119(2017)."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-4009"},{"volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","key":"e_1_3_2_1_34_1","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , Alban Desmaison , Andreas Kopf , Edward Yang , Zachary DeVito , Martin Raison , Alykhan Tejani , Sasank Chilamkurthy , Benoit Steiner , Lu Fang , Junjie Bai , and Soumith Chintala . 2019. PyTorch: An Imperative Style , High-Performance Deep Learning Library . In Advances in Neural Information Processing Systems 32, H.\u00a0Wallach, H.\u00a0Larochelle, A.\u00a0Beygelzimer, F.\u00a0d'Alch\u00e9-Buc, E.\u00a0Fox, and R.\u00a0Garnett (Eds.). Curran Associates, Inc., 8024\u20138035. http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32, H.\u00a0Wallach, H.\u00a0Larochelle, A.\u00a0Beygelzimer, F.\u00a0d'Alch\u00e9-Buc, E.\u00a0Fox, and R.\u00a0Garnett (Eds.). Curran Associates, Inc., 8024\u20138035. http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178876.3185998"},{"key":"e_1_3_2_1_36_1","volume-title":"GPU Technology Conference.","author":"Rossetti Davide","year":"2015","unstructured":"Davide Rossetti and S Team . 2015 . GPUDIRECT: Integrating the GPU with a Network Interface . In GPU Technology Conference. Davide Rossetti and S Team. 2015. GPUDIRECT: Integrating the GPU with a Network Interface. In GPU Technology Conference."},{"key":"e_1_3_2_1_37_1","unstructured":"Kyuhong Shim Minjae Lee Iksoo Choi Yoonho Boo and Wonyong Sung. 2017. SVD-softmax: Fast softmax approximation on large vocabulary neural networks. In Neural Information Processing Systems (NeurIPS). 5464\u20135474.  Kyuhong Shim Minjae Lee Iksoo Choi Yoonho Boo and Wonyong Sung. 2017. SVD-softmax: Fast softmax approximation on large vocabulary neural networks. In Neural Information Processing Systems (NeurIPS). 5464\u20135474."},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2018.2857044"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403342"},{"key":"e_1_3_2_1_40_1","unstructured":"Ilya Sutskever Oriol Vinyals and Quoc\u00a0V Le. 2014. Sequence to sequence learning with neural networks. arXiv preprint arXiv:1409.3215(2014).  Ilya Sutskever Oriol Vinyals and Quoc\u00a0V Le. 2014. Sequence to sequence learning with neural networks. arXiv preprint arXiv:1409.3215(2014)."},{"key":"e_1_3_2_1_41_1","volume-title":"HOTI 2019: Compute Express Link. In 2019 IEEE Symposium on High-Performance Interconnects (HOTI). IEEE Computer Society, 18\u201318","author":"Van\u00a0Doren S","year":"2019","unstructured":"S Van\u00a0Doren . 2019 . HOTI 2019: Compute Express Link. In 2019 IEEE Symposium on High-Performance Interconnects (HOTI). IEEE Computer Society, 18\u201318 . S Van\u00a0Doren. 2019. HOTI 2019: Compute Express Link. In 2019 IEEE Symposium on High-Performance Interconnects (HOTI). IEEE Computer Society, 18\u201318."},{"key":"e_1_3_2_1_42_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan\u00a0N Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. arXiv preprint arXiv:1706.03762(2017).  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan\u00a0N Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. arXiv preprint arXiv:1706.03762(2017)."},{"key":"e_1_3_2_1_43_1","unstructured":"Peiqi Wang Xinfeng Xie Lei Deng Guoqi Li Dongsheng Wang and Yuan Xie. 2018. HitNet: hybrid ternary recurrent neural network. In Advances in Neural Information Processing Systems. 604\u2013614.  Peiqi Wang Xinfeng Xie Lei Deng Guoqi Li Dongsheng Wang and Yuan Xie. 2018. HitNet: hybrid ternary recurrent neural network. In Advances in Neural Information Processing Systems. 604\u2013614."},{"key":"e_1_3_2_1_44_1","unstructured":"Yonghui Wu Mike Schuster Zhifeng Chen Quoc\u00a0V. Le Mohammad Norouzi Wolfgang Macherey Maxim Krikun Yuan Cao Qin Gao Klaus Macherey Jeff Klingner Apurva Shah Melvin Johnson Xiaobing Liu \u0141ukasz Kaiser Stephan Gouws Yoshikiyo Kato Taku Kudo Hideto Kazawa Keith Stevens George Kurian Nishant Patil Wei Wang Cliff Young Jason Smith Jason Riesa Alex Rudnick Oriol Vinyals Greg Corrado Macduff Hughes and Jeffrey Dean. 2016. Google\u2019s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. (2016). arxiv:cs.CL\/1609.08144  Yonghui Wu Mike Schuster Zhifeng Chen Quoc\u00a0V. Le Mohammad Norouzi Wolfgang Macherey Maxim Krikun Yuan Cao Qin Gao Klaus Macherey Jeff Klingner Apurva Shah Melvin Johnson Xiaobing Liu \u0141ukasz Kaiser Stephan Gouws Yoshikiyo Kato Taku Kudo Hideto Kazawa Keith Stevens George Kurian Nishant Patil Wei Wang Cliff Young Jason Smith Jason Riesa Alex Rudnick Oriol Vinyals Greg Corrado Macduff Hughes and Jeffrey Dean. 2016. Google\u2019s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. (2016). arxiv:cs.CL\/1609.08144"},{"key":"e_1_3_2_1_45_1","unstructured":"Chen Xu Jianqiang Yao Zhouchen Lin Wenwu Ou Yuanbin Cao Zhirong Wang and Hongbin Zha. 2018. Alternating multi-bit quantization for recurrent neural networks. arXiv preprint arXiv:1802.00150(2018).  Chen Xu Jianqiang Yao Zhouchen Lin Wenwu Ou Yuanbin Cao Zhirong Wang and Hongbin Zha. 2018. Alternating multi-bit quantization for recurrent neural networks. arXiv preprint arXiv:1802.00150(2018)."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3243176.3243188"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219890"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.5555\/3327345.3327528"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00053"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12337"},{"key":"e_1_3_2_1_51_1","unstructured":"Chenzhuo Zhu Song Han Huizi Mao and William\u00a0J Dally. 2016. Trained ternary quantization. arXiv preprint arXiv:1612.01064(2016).  Chenzhuo Zhu Song Han Huizi Mao and William\u00a0J Dally. 2016. Trained ternary quantization. arXiv preprint arXiv:1612.01064(2016)."},{"key":"e_1_3_2_1_52_1","unstructured":"Michael Zhu and Suyog Gupta. 2017. To prune or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878(2017).  Michael Zhu and Suyog Gupta. 2017. To prune or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878(2017)."}],"event":{"name":"MICRO '21: 54th Annual IEEE\/ACM International Symposium on Microarchitecture","sponsor":["SIGMICRO ACM Special Interest Group on Microarchitectural Research and Processing"],"location":"Virtual Event Greece","acronym":"MICRO '21"},"container-title":["MICRO-54: 54th Annual IEEE\/ACM International Symposium on Microarchitecture"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3466752.3480090","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/abs\/10.1145\/3466752.3480090","content-type":"text\/html","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3466752.3480090","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3466752.3480090","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:18:56Z","timestamp":1750191536000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3466752.3480090"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,17]]},"references-count":52,"alternative-id":["10.1145\/3466752.3480090","10.1145\/3466752"],"URL":"https:\/\/doi.org\/10.1145\/3466752.3480090","relation":{},"subject":[],"published":{"date-parts":[[2021,10,17]]},"assertion":[{"value":"2021-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}