{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T16:55:24Z","timestamp":1778604924002,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":47,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,11,13]],"date-time":"2021-11-13T00:00:00Z","timestamp":1636761600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000001","name":"NSF (National Science Foundation)","doi-asserted-by":"publisher","award":["1925717, 2124039"],"award-info":[{"award-number":["1925717, 2124039"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"name":"U.S. DOE Office of Sci-ence, Office of Advanced Scientific Computing Research","award":["66150"],"award-info":[{"award-number":["66150"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,11,14]]},"DOI":"10.1145\/3458817.3476157","type":"proceedings-article","created":{"date-parts":[[2021,10,21]],"date-time":"2021-10-21T04:49:21Z","timestamp":1634791761000},"page":"1-13","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":38,"title":["APNN-TC"],"prefix":"10.1145","author":[{"given":"Boyuan","family":"Feng","sequence":"first","affiliation":[{"name":"University of California"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuke","family":"Wang","sequence":"additional","affiliation":[{"name":"University of California"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tong","family":"Geng","sequence":"additional","affiliation":[{"name":"Pacific Northwest National Laboratory"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ang","family":"Li","sequence":"additional","affiliation":[{"name":"Pacific Northwest National Laboratory"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yufei","family":"Ding","sequence":"additional","affiliation":[{"name":"University of California"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,11,13]]},"reference":[{"key":"e_1_3_2_2_1_1","unstructured":"AMD. 2013. AMD Accelerated Parallel Processing OpenCL Programming Guide. http:\/\/developer.amd.com\/wordpress\/media\/2013\/07\/AMD_Accelerated_Parallel_Processing_OpenCL_Programming_Guide-rev-2.7.pdf.  AMD. 2013. AMD Accelerated Parallel Processing OpenCL Programming Guide. http:\/\/developer.amd.com\/wordpress\/media\/2013\/07\/AMD_Accelerated_Parallel_Processing_OpenCL_Programming_Guide-rev-2.7.pdf."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.574"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.195"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2021.3061394"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2018.022071134"},{"key":"e_1_3_2_2_6_1","unstructured":"Matthieu Courbariaux Yoshua Bengio and Jean-Pierre David. 2015. BinaryConnect: Training Deep Neural Networks with binary weights during propagations. In NIPS. 3123--3131.  Matthieu Courbariaux Yoshua Bengio and Jean-Pierre David. 2015. BinaryConnect: Training Deep Neural Networks with binary weights during propagations. In NIPS. 3123--3131."},{"key":"e_1_3_2_2_7_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT (1)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT (1) . Association for Computational Linguistics , 4171--4186. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT (1). Association for Computational Linguistics, 4171--4186."},{"key":"e_1_3_2_2_8_1","volume-title":"Alexander Frickenstein, Lukas Frickenstein, and Walter Stechele.","author":"Fasfous Nael","year":"2021","unstructured":"Nael Fasfous , Manoj Rohit Vemparala , Alexander Frickenstein, Lukas Frickenstein, and Walter Stechele. 2021 . BinaryCoP: Binary Neural Network-based COVID-19 Face-Mask Wear and Positioning Predictor on Edge Devices. CoRR abs\/2102.03456 (2021). Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Lukas Frickenstein, and Walter Stechele. 2021. BinaryCoP: Binary Neural Network-based COVID-19 Face-Mask Wear and Positioning Predictor on Edge Devices. CoRR abs\/2102.03456 (2021)."},{"key":"e_1_3_2_2_9_1","volume-title":"XpulpNN: Accelerating Quantized Neural Networks on RISC-V Processors Through ISA Extensions","author":"Garofalo Angelo","unstructured":"Angelo Garofalo , Giuseppe Tagliavini , Francesco Conti , Davide Rossi , and Luca Benini . 2020. XpulpNN: Accelerating Quantized Neural Networks on RISC-V Processors Through ISA Extensions . In DATE. IEEE , 186--191. Angelo Garofalo, Giuseppe Tagliavini, Francesco Conti, Davide Rossi, and Luca Benini. 2020. XpulpNN: Accelerating Quantized Neural Networks on RISC-V Processors Through ISA Extensions. In DATE. IEEE, 186--191."},{"key":"e_1_3_2_2_10_1","unstructured":"Tong Geng Ang Li Tianqi Wang Chunshu Wu Yanfei Li Runbin Shi Wei Wu and Martin Herbordt. [n.d.]. O3BNN-R: An out-of-order architecture for high-performance and regularized BNN inference. TPDS'20 ([n. d.]).  Tong Geng Ang Li Tianqi Wang Chunshu Wu Yanfei Li Runbin Shi Wei Wu and Martin Herbordt. [n.d.]. O3BNN-R: An out-of-order architecture for high-performance and regularized BNN inference. TPDS'20 ([n. d.])."},{"key":"e_1_3_2_2_11_1","volume-title":"Dally","author":"Han Song","year":"2016","unstructured":"Song Han , Huizi Mao , and William J . Dally . 2016 . Deep Compression : Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. In ICLR. Song Han, Huizi Mao, and William J. Dally. 2016. Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. In ICLR."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_2_13_1","volume-title":"Intel Nervana Neural Network Processor-T (NNP-T) Fused Floating Point Many-Term Dot Product. In 2020 IEEE 27th Symposium on Computer Arithmetic (ARITH). IEEE, 133--136","author":"Hickmann Brian","year":"2020","unstructured":"Brian Hickmann , Jieasheng Chen , Michael Rotzin , Andrew Yang , Maciej Urbanski , and Sasikanth Avancha . 2020 . Intel Nervana Neural Network Processor-T (NNP-T) Fused Floating Point Many-Term Dot Product. In 2020 IEEE 27th Symposium on Computer Arithmetic (ARITH). IEEE, 133--136 . Brian Hickmann, Jieasheng Chen, Michael Rotzin, Andrew Yang, Maciej Urbanski, and Sasikanth Avancha. 2020. Intel Nervana Neural Network Processor-T (NNP-T) Fused Floating Point Many-Term Dot Product. In 2020 IEEE 27th Symposium on Computer Arithmetic (ARITH). IEEE, 133--136."},{"key":"e_1_3_2_2_14_1","volume-title":"MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv e-prints (April","author":"Howard Andrew G.","year":"2017","unstructured":"Andrew G. Howard , Menglong Zhu , Bo Chen , Dmitry Kalenichenko , Weijun Wang , Tobias Weyand , Marco Andreetto , and Hartwig Adam . 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv e-prints (April 2017 ). Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv e-prints (April 2017)."},{"key":"e_1_3_2_2_15_1","unstructured":"Intel. 2012. Intel Xeon Phi Coprocessor Instruction Set Architecture Reference Manual. https:\/\/software.intel.com\/content\/dam\/develop\/external\/us\/en\/documents\/327364001en.pdf.  Intel. 2012. Intel Xeon Phi Coprocessor Instruction Set Architecture Reference Manual. https:\/\/software.intel.com\/content\/dam\/develop\/external\/us\/en\/documents\/327364001en.pdf."},{"key":"e_1_3_2_2_16_1","volume-title":"International conference on machine learning. PMLR, 448--456","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy . 2015 . Batch normalization: Accelerating deep network training by reducing internal covariate shift . In International conference on machine learning. PMLR, 448--456 . Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning. PMLR, 448--456."},{"key":"e_1_3_2_2_17_1","volume-title":"Dissecting the NVidia Turing T4 GPU via microbenchmarking. arXiv","author":"Jia Zhe","year":"2019","unstructured":"Zhe Jia , Marco Maggioni , Jeffrey Smith , and Daniele Paolo Scarpazza . 2019. Dissecting the NVidia Turing T4 GPU via microbenchmarking. arXiv ( 2019 ). Zhe Jia, Marco Maggioni, Jeffrey Smith, and Daniele Paolo Scarpazza. 2019. Dissecting the NVidia Turing T4 GPU via microbenchmarking. arXiv (2019)."},{"key":"e_1_3_2_2_18_1","volume-title":"Dissecting the NVIDIA volta GPU architecture via microbenchmarking. arXiv preprint arXiv:1804.06826","author":"Jia Zhe","year":"2018","unstructured":"Zhe Jia , Marco Maggioni , Benjamin Staiger , and Daniele P Scarpazza . 2018. Dissecting the NVIDIA volta GPU architecture via microbenchmarking. arXiv preprint arXiv:1804.06826 ( 2018 ). Zhe Jia, Marco Maggioni, Benjamin Staiger, and Daniele P Scarpazza. 2018. Dissecting the NVIDIA volta GPU architecture via microbenchmarking. arXiv preprint arXiv:1804.06826 (2018)."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_2_2_20_1","volume-title":"Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems 25","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky , Ilya Sutskever , and Geoffrey E Hinton . 2012. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems 25 ( 2012 ), 1097--1105. Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems 25 (2012), 1097--1105."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11265-017-1255-5"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356169"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2016.89"},{"key":"e_1_3_2_2_24_1","volume-title":"Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 1273--1278","author":"Li Ang","year":"2016","unstructured":"Ang Li , Shuaiwen Leon Song , Akash Kumar , Eddy Z Zhang , Daniel Chavarr\u00eda-Miranda , and Henk Corp oraal. 2016 . Critical points based register-concurrency autotuning for GPUs. In 2016 Design , Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 1273--1278 . Ang Li, Shuaiwen Leon Song, Akash Kumar, Eddy Z Zhang, Daniel Chavarr\u00eda-Miranda, and Henk Corporaal. 2016. Critical points based register-concurrency autotuning for GPUs. In 2016 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 1273--1278."},{"key":"e_1_3_2_2_25_1","first-page":"1878","article-title":"Accelerating Binarized Neural Networks via Bit-Tensor-Cores in Turing GPUs","volume":"32","author":"Li Ang","year":"2020","unstructured":"Ang Li and Simon Su . 2020 . Accelerating Binarized Neural Networks via Bit-Tensor-Cores in Turing GPUs . IEEE Transactions on Parallel and Distributed Systems 32 , 7 (2020), 1878 -- 1891 . Ang Li and Simon Su. 2020. Accelerating Binarized Neural Networks via Bit-Tensor-Cores in Turing GPUs. IEEE Transactions on Parallel and Distributed Systems 32, 7 (2020), 1878--1891.","journal-title":"IEEE Transactions on Parallel and Distributed Systems"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749246.2749265"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5924"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5954"},{"key":"e_1_3_2_2_29_1","unstructured":"Bradley McDanel Surat Teerapittayanon and H. T. Kung. 2017. Embedded Binarized Neural Networks. In EWSN. Junction Publishing Canada \/ ACM.  Bradley McDanel Surat Teerapittayanon and H. T. Kung. 2017. Embedded Binarized Neural Networks. In EWSN. Junction Publishing Canada \/ ACM."},{"key":"e_1_3_2_2_30_1","volume-title":"Towards Real-Time DNN Inference on Mobile Platforms with Model Pruning and Compiler Optimization. IJCAI","author":"Niu Wei","year":"2020","unstructured":"Wei Niu , Pu Zhao , Zheng Zhan , Xue Lin , Yanzhi Wang , and Bin Ren . 2020. Towards Real-Time DNN Inference on Mobile Platforms with Model Pruning and Compiler Optimization. IJCAI ( 2020 ). Wei Niu, Pu Zhao, Zheng Zhan, Xue Lin, Yanzhi Wang, and Bin Ren. 2020. Towards Real-Time DNN Inference on Mobile Platforms with Model Pruning and Compiler Optimization. IJCAI (2020)."},{"key":"e_1_3_2_2_31_1","unstructured":"NVIDIA. [n.d.]. CUDA Template Library for Dense Linear Algebra at All Levels and Scales (CUTLASS).  NVIDIA. [n.d.]. CUDA Template Library for Dense Linear Algebra at All Levels and Scales (CUTLASS)."},{"key":"e_1_3_2_2_32_1","unstructured":"Nvidia. [n.d.]. NVIDIA A100 Tensor Core GPU Architecture. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/Data-Center\/nvidia-ampere-architecture-whitepaper.pdf.  Nvidia. [n.d.]. NVIDIA A100 Tensor Core GPU Architecture. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/Data-Center\/nvidia-ampere-architecture-whitepaper.pdf."},{"key":"e_1_3_2_2_33_1","unstructured":"Nvidia. [n.d.]. NVIDIA TESLA V100 GPU ARCHITECTURE. https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf.  Nvidia. [n.d.]. NVIDIA TESLA V100 GPU ARCHITECTURE. https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf."},{"key":"e_1_3_2_2_34_1","unstructured":"NVIDIA. 2021. CUDA Programming Guide: Sub-byte Operations. https:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/#wmma-subbyte.  NVIDIA. 2021. CUDA Programming Guide: Sub-byte Operations. https:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/#wmma-subbyte."},{"key":"e_1_3_2_2_35_1","volume-title":"Energy-Efficient Neural Network Accelerator Based on Outlier-Aware Low-Precision Computation","author":"Park Eunhyeok","unstructured":"Eunhyeok Park , Dongyoung Kim , and Sungjoo Yoo . 2018. Energy-Efficient Neural Network Accelerator Based on Outlier-Aware Low-Precision Computation . In ISCA. IEEE Computer Society , 688--698. Eunhyeok Park, Dongyoung Kim, and Sungjoo Yoo. 2018. Energy-Efficient Neural Network Accelerator Based on Outlier-Aware Low-Precision Computation. In ISCA. IEEE Computer Society, 688--698."},{"key":"e_1_3_2_2_36_1","volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , Alban Desmaison , Andreas Kopf , Edward Yang , Zachary DeVito , Martin Raison , Alykhan Tejani , Sasank Chilamkurthy , Benoit Steiner , Lu Fang , Junjie Bai , and Soumith Chintala . 2019. PyTorch: An Imperative Style , High-Performance Deep Learning Library . In NeurIPS'19, H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alch'e-Buc, E. Fox, and R. Garnett (Eds.). Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In NeurIPS'19, H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alch'e-Buc, E. Fox, and R. Garnett (Eds.)."},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00016"},{"key":"e_1_3_2_2_38_1","unstructured":"Advanced Grid Research. [n.d.]. Sesor Technologies and Data Analytics. https:\/\/www.smartgrid.gov\/files\/Sensor_Technologies_MYPP_12_19_18_final.pdf.  Advanced Grid Research. [n.d.]. Sesor Technologies and Data Analytics. https:\/\/www.smartgrid.gov\/files\/Sensor_Technologies_MYPP_12_19_18_final.pdf."},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"crossref","unstructured":"Mark Sandler Andrew Howard Menglong Zhu Andrey Zhmoginov and Liang-Chieh Chen. 2018. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR.  Mark Sandler Andrew Howard Menglong Zhu Andrey Zhmoginov and Liang-Chieh Chen. 2018. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_2_2_40_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. In NIPS. 5998--6008.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. In NIPS. 5998--6008."},{"key":"e_1_3_2_2_41_1","volume-title":"HAQ: Hardware-Aware Automated Quantization With Mixed Precision","author":"Wang Kuan","year":"2019","unstructured":"Kuan Wang , Zhijian Liu , Yujun Lin , Ji Lin , and Song Han . 2019 . HAQ: Hardware-Aware Automated Quantization With Mixed Precision . In CVPR. Computer Vision Foundation \/ IEEE , 8612--8620. Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han. 2019. HAQ: Hardware-Aware Automated Quantization With Mixed Precision. In CVPR. Computer Vision Foundation \/ IEEE, 8612--8620."},{"key":"e_1_3_2_2_42_1","volume-title":"Quantized Convolutional Neural Networks for Mobile Devices. In CVPR'16","author":"Wu J.","unstructured":"J. Wu , C. Leng , Y. Wang , Q. Hu , and J. Cheng . [n.d.] . Quantized Convolutional Neural Networks for Mobile Devices. In CVPR'16 . J. Wu, C. Leng, Y. Wang, Q. Hu, and J. Cheng. [n.d.]. Quantized Convolutional Neural Networks for Mobile Devices. In CVPR'16."},{"key":"e_1_3_2_2_43_1","volume-title":"Searching for Low-Bit Weights in Quantized Neural Networks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020","author":"Yang Zhaohui","year":"2020","unstructured":"Zhaohui Yang , Yunhe Wang , Kai Han , Chunjing Xu , Chao Xu , Dacheng Tao , and Chang Xu . 2020 . Searching for Low-Bit Weights in Quantized Neural Networks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020 , NeurIPS 2020, December 6--12, 2020, virtual, Hugo Larochelle, Marc'Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.). Zhaohui Yang, Yunhe Wang, Kai Han, Chunjing Xu, Chao Xu, Dacheng Tao, and Chang Xu. 2020. Searching for Low-Bit Weights in Quantized Neural Networks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6--12, 2020, virtual, Hugo Larochelle, Marc'Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)."},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01237-3_23"},{"key":"e_1_3_2_2_45_1","volume-title":"ChrEn: Cherokee-English Machine Translation for Endangered Language Revitalization. In EMNLP'20","author":"Zhang Shiyue","unstructured":"Shiyue Zhang , Benjamin Frey , and Mohit Bansal . [n.d.]. ChrEn: Cherokee-English Machine Translation for Endangered Language Revitalization. In EMNLP'20 . Shiyue Zhang, Benjamin Frey, and Mohit Bansal. [n.d.]. ChrEn: Cherokee-English Machine Translation for Endangered Language Revitalization. In EMNLP'20."},{"key":"e_1_3_2_2_46_1","volume-title":"DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients. CoRR abs\/1606.06160","author":"Zhou Shuchang","year":"2016","unstructured":"Shuchang Zhou , Zekun Ni , Xinyu Zhou , He Wen , Yuxin Wu , and Yuheng Zou . 2016. DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients. CoRR abs\/1606.06160 ( 2016 ). Shuchang Zhou, Zekun Ni, Xinyu Zhou, He Wen, Yuxin Wu, and Yuheng Zou. 2016. DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients. CoRR abs\/1606.06160 (2016)."},{"key":"e_1_3_2_2_47_1","volume-title":"CVPR'20","author":"Zhuang B.","unstructured":"B. Zhuang , L. Liu , M. Tan , C. Shen , and I. Reid . [n.d.]. Training Quantized Neural Networks With a Full-Precision Auxiliary Module . In CVPR'20 . B. Zhuang, L. Liu, M. Tan, C. Shen, and I. Reid. [n.d.]. Training Quantized Neural Networks With a Full-Precision Auxiliary Module. In CVPR'20."}],"event":{"name":"SC '21: The International Conference for High Performance Computing, Networking, Storage and Analysis","location":"St. Louis Missouri","acronym":"SC '21","sponsor":["SIGHPC ACM Special Interest Group on High Performance Computing, Special Interest Group on High Performance Computing","IEEE CS"]},"container-title":["Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3458817.3476157","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3458817.3476157","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3458817.3476157","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T17:49:06Z","timestamp":1750268946000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3458817.3476157"}},"subtitle":["accelerating arbitrary precision neural networks on ampere GPU tensor cores"],"short-title":[],"issued":{"date-parts":[[2021,11,13]]},"references-count":47,"alternative-id":["10.1145\/3458817.3476157","10.1145\/3458817"],"URL":"https:\/\/doi.org\/10.1145\/3458817.3476157","relation":{},"subject":[],"published":{"date-parts":[[2021,11,13]]},"assertion":[{"value":"2021-11-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}