{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,28]],"date-time":"2026-08-28T06:04:30Z","timestamp":1787897070468,"version":"build-2784847793"},"reference-count":145,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2022,1,8]],"date-time":"2022-01-08T00:00:00Z","timestamp":1641600000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2022,8,31]]},"abstract":"<jats:p>\n                    In recent years, the fields of natural language processing (NLP) and information retrieval (IR) have made tremendous progress thanks to deep learning models like Recurrent Neural Networks (RNNs), Gated Recurrent Units (GRUs) and Long Short-Term Memory (LSTMs) networks, and Transformer\u00a0[\n                    <jats:xref ref-type=\"bibr\">121<\/jats:xref>\n                    ] based models like Bidirectional Encoder Representations from Transformers (BERT)\u00a0[\n                    <jats:xref ref-type=\"bibr\">24<\/jats:xref>\n                    ], Generative Pre-training Transformer (GPT-2)\u00a0[\n                    <jats:xref ref-type=\"bibr\">95<\/jats:xref>\n                    ], Multi-task Deep Neural Network (MT-DNN)\u00a0[\n                    <jats:xref ref-type=\"bibr\">74<\/jats:xref>\n                    ], Extra-Long Network (XLNet)\u00a0[\n                    <jats:xref ref-type=\"bibr\">135<\/jats:xref>\n                    ], Text-to-text transfer transformer (T5)\u00a0[\n                    <jats:xref ref-type=\"bibr\">96<\/jats:xref>\n                    ], T-NLG\u00a0[\n                    <jats:xref ref-type=\"bibr\">99<\/jats:xref>\n                    ], and GShard\u00a0[\n                    <jats:xref ref-type=\"bibr\">64<\/jats:xref>\n                    ]. But these models are humongous in size. On the other hand, real-world applications demand small model size, low response times, and low computational power wattage. In this survey, we discuss six different types of methods (Pruning, Quantization, Knowledge Distillation (KD), Parameter Sharing, Tensor Decomposition, and Sub-quadratic Transformer-based methods) for compression of such models to enable their deployment in real industry NLP projects. Given the critical need of building applications with efficient and small models, and the large amount of recently published work in this area, we believe that this survey organizes the plethora of work done by the \u201cdeep learning for NLP\u201d community in the past few years and presents it as a coherent story.\n                  <\/jats:p>","DOI":"10.1145\/3487045","type":"journal-article","created":{"date-parts":[[2022,1,8]],"date-time":"2022-01-08T15:51:00Z","timestamp":1641657060000},"page":"1-55","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":78,"title":["Compression of Deep Learning Models for Text: A Survey"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2843-3110","authenticated-orcid":false,"given":"Manish","family":"Gupta","sequence":"first","affiliation":[{"name":"Microsoft, Gachibowli, Hyderabad, Telangana, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Puneet","family":"Agrawal","sequence":"additional","affiliation":[{"name":"Microsoft, Gachibowli, Hyderabad, Telangana, India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,1,8]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Md. Zahangir Alom Adam T. Moody Naoya Maruyama Brian C. Van Essen and Tarek M. Taha. 2018. Effective quantization approaches for recurrent neural networks. In Proceedings of the 2018 International Joint Conference on Neural Networks . IEEE 1\u20138."},{"key":"e_1_3_2_3_2","article-title":"Large scale distributed neural network training through online distillation","author":"Anil Rohan","year":"2018","unstructured":"Rohan Anil, Gabriel Pereyra, Alexandre Passos, Robert Ormandi, George E. Dahl, and Geoffrey E. Hinton. 2018. Large scale distributed neural network training through online distillation. arXiv:1804.03235. Retrieved from https:\/\/arxiv.org\/abs\/1804.03235.","journal-title":"arXiv:1804.03235"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.5555\/2969033.2969123"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454350"},{"key":"e_1_3_2_6_2","first-page":"016329","article-title":"Hippocampal spine head sizes are highly precise","author":"Bartol Thomas M.","year":"2015","unstructured":"Thomas M. Bartol, Cailey Bromer, Justin Kinney, Michael A. Chirillo, Jennifer N. Bourne, Kristen M. Harris, and Terrence J. Sejnowski. 2015. Hippocampal spine head sizes are highly precise. bioRxiv (2015), 016329.","journal-title":"bioRxiv"},{"key":"e_1_3_2_7_2","article-title":"Longformer: The long-document transformer","author":"Beltagy Iz","year":"2020","unstructured":"Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv:2004.05150. Retrieved from https:\/\/arxiv.org\/abs\/2004.05150.","journal-title":"arXiv:2004.05150"},{"key":"e_1_3_2_8_2","article-title":"Estimating or propagating gradients through stochastic neurons for conditional computation","author":"Bengio Yoshua","year":"2013","unstructured":"Yoshua Bengio, Nicholas L\u00e9onard, and Aaron Courville. 2013. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv:1308.3432. Retrieved from https:\/\/arxiv.org\/abs\/1308.3432.","journal-title":"arXiv:1308.3432"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2877890"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289602.3293898"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF02310791"},{"key":"e_1_3_2_12_2","article-title":"Strategies for training large vocabulary neural language models","author":"Chen Welin","year":"2015","unstructured":"Welin Chen, David Grangier, and Michael Auli. 2015. Strategies for training large vocabulary neural language models. arXiv:1512.04906. Retrieved from https:\/\/arxiv.org\/abs\/1512.04906.","journal-title":"arXiv:1512.04906"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.5555\/3045118.3045361"},{"key":"e_1_3_2_14_2","article-title":"Compressing neural language models by sparse word representations","author":"Chen Yunchuan","year":"2016","unstructured":"Yunchuan Chen, Lili Mou, Yan Xu, Ge Li, and Zhi Jin. 2016. Compressing neural language models by sparse word representations. arXiv:1610.03950. Retrieved from https:\/\/arxiv.org\/abs\/1610.03950.","journal-title":"arXiv:1610.03950"},{"key":"e_1_3_2_15_2","article-title":"A survey of model compression and acceleration for deep neural networks","author":"Cheng Yu","year":"2017","unstructured":"Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang. 2017. A survey of model compression and acceleration for deep neural networks. arXiv:1710.09282. Retrieved from https:\/\/arxiv.org\/abs\/1710.09282.","journal-title":"arXiv:1710.09282"},{"key":"e_1_3_2_16_2","volume-title":"Transformers. Zip: Compressing Transformers with Pruning and Quantization","author":"Cheong Robin","year":"2019","unstructured":"Robin Cheong and Robel Daniel. 2019. Transformers. Zip: Compressing Transformers with Pruning and Quantization. Technical Report. Technical report, Stanford University, Stanford, California, 2019."},{"key":"e_1_3_2_17_2","article-title":"Generating long sequences with sparse transformers","author":"Child Rewon","year":"2019","unstructured":"Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. Generating long sequences with sparse transformers. arXiv:1904.10509. Retrieved from https:\/\/arxiv.org\/abs\/1904.10509.","journal-title":"arXiv:1904.10509"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.5555\/2969442.2969588"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.5555\/3294996.3295182"},{"key":"e_1_3_2_20_2","article-title":"Grow and prune compact, fast, and accurate LSTMs","author":"Dai Xiaoliang","year":"2018","unstructured":"Xiaoliang Dai, Hongxu Yin, and Niraj K. Jha. 2018. Grow and prune compact, fast, and accurate LSTMs. arXiv:1805.11797. Retrieved from https:\/\/arxiv.org\/abs\/1805.11797.","journal-title":"arXiv:1805.11797"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-47426-3_19"},{"key":"e_1_3_2_22_2","article-title":"Universal transformers","author":"Dehghani Mostafa","year":"2018","unstructured":"Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and \u0141ukasz Kaiser. 2018. Universal transformers. arXiv:1807.03819 (2018). Retrieved from https:\/\/arxiv.org\/abs\/1807.03819.","journal-title":"arXiv:1807.03819"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2020.2976475"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.5555\/2999792.2999852"},{"key":"e_1_3_2_25_2","article-title":"Bert: Pre-training of deep bidirectional transformers for language understanding","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805 (2018). Retrieved from https:\/\/arxiv.org\/abs\/1810.04805.","journal-title":"arXiv:1810.04805"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.5555\/3045390.3045604"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1106"},{"key":"e_1_3_2_28_2","article-title":"Reducing transformer depth on demand with structured dropout","author":"Fan Angela","year":"2019","unstructured":"Angela Fan, Edouard Grave, and Armand Joulin. 2019. Reducing transformer depth on demand with structured dropout. arXiv:1909.11556. Retrieved from https:\/\/arxiv.org\/abs\/1909.11556.","journal-title":"arXiv:1909.11556"},{"key":"e_1_3_2_29_2","article-title":"Sparse overcomplete word vector representations","author":"Faruqui Manaal","year":"2015","unstructured":"Manaal Faruqui, Yulia Tsvetkov, Dani Yogatama, Chris Dyer, and Noah Smith. 2015. Sparse overcomplete word vector representations. arXiv:1506.02004. Retrieved from https:\/\/arxiv.org\/abs\/1506.02004.","journal-title":"arXiv:1506.02004"},{"key":"e_1_3_2_30_2","article-title":"Ensemble distillation for neural machine translation","author":"Freitag Markus","year":"2017","unstructured":"Markus Freitag, Yaser Al-Onaizan, and Baskaran Sankaran. 2017. Ensemble distillation for neural machine translation. arXiv:1702.01802. Retrieved from https:\/\/arxiv.org\/abs\/1702.01802.","journal-title":"arXiv:1702.01802"},{"key":"e_1_3_2_31_2","article-title":"Compressing deep convolutional networks using vector quantization","author":"Gong Yunchao","year":"2014","unstructured":"Yunchao Gong, Liu Liu, Ming Yang, and Lubomir Bourdev. 2014. Compressing deep convolutional networks using vector quantization. arXiv:1412.6115. Retrieved from https:\/\/arxiv.org\/abs\/1412.6115.","journal-title":"arXiv:1412.6115"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2019.03.057"},{"key":"e_1_3_2_33_2","article-title":"Reweighted proximal pruning for large-scale language representation","author":"Guo Fu-Ming","year":"2019","unstructured":"Fu-Ming Guo, Sijia Liu, Finlay S Mungall, Xue Lin, and Yanzhi Wang. 2019. Reweighted proximal pruning for large-scale language representation. arXiv:1909.12486. Retrieved from https:\/\/arxiv.org\/abs\/1909.12486.","journal-title":"arXiv:1909.12486"},{"key":"e_1_3_2_34_2","unstructured":"Qipeng Guo Xipeng Qiu Pengfei Liu Yunfan Shao Xiangyang Xue and Zheng Zhang. 2019. Star-Transformer. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 1315\u20131325."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.430"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001163"},{"key":"e_1_3_2_37_2","article-title":"Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding","author":"Han Song","year":"2015","unstructured":"Song Han, Huizi Mao, and William J. Dally. 2015. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv:1510.00149. Retrieved from https:\/\/arxiv.org\/abs\/1510.00149.","journal-title":"arXiv:1510.00149"},{"key":"e_1_3_2_38_2","article-title":"DSD: Dense-sparse-dense training for deep neural networks","author":"Han Song","year":"2016","unstructured":"Song Han, Jeff Pool, Sharan Narang, Huizi Mao, Enhao Gong, Shijian Tang, Erich Elsen, Peter Vajda, Manohar Paluri, John Tran, Bryan Catanzaro, and William J. Dally. 2016. DSD: Dense-sparse-dense training for deep neural networks. arXiv:1607.04381. Retrieved from https:\/\/arxiv.org\/abs\/1607.04381.","journal-title":"arXiv:1607.04381"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.5555\/2969239.2969366"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.5555\/2987061.2987082"},{"key":"e_1_3_2_41_2","article-title":"Effective quantization methods for recurrent neural networks","author":"He Qinyao","year":"2016","unstructured":"Qinyao He, He Wen, Shuchang Zhou, Yuxin Wu, Cong Yao, Xinyu Zhou, and Yuheng Zou. 2016. Effective quantization methods for recurrent neural networks. arXiv:1611.10176. Retrieved from https:\/\/arxiv.org\/abs\/1611.10176.","journal-title":"arXiv:1611.10176"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2014.6853595"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_48"},{"key":"e_1_3_2_44_2","article-title":"Distilling the knowledge in a neural network","author":"Hinton Geoffrey","year":"2015","unstructured":"Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the knowledge in a neural network. arXiv:1503.02531. Retrieved from https:\/\/arxiv.org\/abs\/1503.02531.","journal-title":"arXiv:1503.02531"},{"key":"e_1_3_2_45_2","article-title":"Loss-aware weight quantization of deep networks","author":"Hou Lu","year":"2018","unstructured":"Lu Hou and James T. Kwok. 2018. Loss-aware weight quantization of deep networks. arXiv:1802.08635. Retrieved from https:\/\/arxiv.org\/abs\/1802.08635.","journal-title":"arXiv:1802.08635"},{"key":"e_1_3_2_46_2","article-title":"Loss-aware binarization of deep networks","author":"Hou Lu","year":"2016","unstructured":"Lu Hou, Quanming Yao, and James T. Kwok. 2016. Loss-aware binarization of deep networks. arXiv:1611.01600. Retrieved from https:\/\/arxiv.org\/abs\/1611.01600.","journal-title":"arXiv:1611.01600"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.5555\/3157382.3157557"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.5555\/3122009.3242044"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/SiPS.2014.6986082"},{"key":"e_1_3_2_50_2","article-title":"SqueezeBERT: What can computer vision teach NLP about efficient neural networks?","author":"Iandola Forrest N.","year":"2020","unstructured":"Forrest N. Iandola, Albert E. Shaw, Ravi Krishna, and Kurt W Keutzer. 2020. SqueezeBERT: What can computer vision teach NLP about efficient neural networks? arXiv:2006.11316. Retrieved from https:\/\/arxiv.org\/abs\/2006.11316.","journal-title":"arXiv:2006.11316"},{"key":"e_1_3_2_51_2","article-title":"Tinybert: Distilling bert for natural language understanding","author":"Jiao Xiaoqi","year":"2019","unstructured":"Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019. Tinybert: Distilling bert for natural language understanding. arXiv:1909.10351. Retrieved from https:\/\/arxiv.org\/abs\/1909.10351.","journal-title":"arXiv:1909.10351"},{"key":"e_1_3_2_52_2","article-title":"Exploring the limits of language modeling","author":"Jozefowicz Rafal","year":"2016","unstructured":"Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016. Exploring the limits of language modeling. arXiv:1602.02410. Retrieved from https:\/\/arxiv.org\/abs\/1602.02410.","journal-title":"arXiv:1602.02410"},{"key":"e_1_3_2_53_2","article-title":"Low precision RNNs: Quantizing RNNs without losing accuracy","author":"Kapur Supriya","year":"2017","unstructured":"Supriya Kapur, Asit Mishra, and Debbie Marr. 2017. Low precision RNNs: Quantizing RNNs without losing accuracy. arXiv:1710.07706. Retrieved from https:\/\/arxiv.org\/abs\/1710.07706.","journal-title":"arXiv:1710.07706"},{"key":"e_1_3_2_54_2","article-title":"Transformers are RNNs: Fast autoregressive transformers with linear attention","author":"Katharopoulos Angelos","year":"2020","unstructured":"Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Fran\u00e7ois Fleuret. 2020. Transformers are RNNs: Fast autoregressive transformers with linear attention. arXiv:2006.16236. Retrieved from https:\/\/arxiv.org\/abs\/2006.16236.","journal-title":"arXiv:2006.16236"},{"key":"e_1_3_2_55_2","article-title":"Tensorized embedding layers for efficient model compression","author":"Khrulkov Valentin","year":"2019","unstructured":"Valentin Khrulkov, Oleksii Hrinchuk, Leyla Mirvakhabova, and Ivan Oseledets. 2019. Tensorized embedding layers for efficient model compression. arXiv:1901.10787. Retrieved from https:\/\/arxiv.org\/abs\/1901.10787.","journal-title":"arXiv:1901.10787"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.5555\/3016100.3016285"},{"key":"e_1_3_2_57_2","article-title":"Sequence-level knowledge distillation","author":"Kim Yoon","year":"2016","unstructured":"Yoon Kim and Alexander M. Rush. 2016. Sequence-level knowledge distillation. arXiv:1606.07947. Retrieved from https:\/\/arxiv.org\/abs\/1606.07947.","journal-title":"arXiv:1606.07947"},{"key":"e_1_3_2_58_2","article-title":"Fastformers: Highly efficient transformer models for natural language understanding","author":"Kim Young Jin","year":"2020","unstructured":"Young Jin Kim and Hany Hassan Awadalla. 2020. Fastformers: Highly efficient transformer models for natural language understanding. arXiv:2010.13382. Retrieved from https:\/\/arxiv.org\/abs\/2010.13382.","journal-title":"arXiv:2010.13382"},{"key":"e_1_3_2_59_2","article-title":"Reformer: The efficient transformer","author":"Kitaev Nikita","year":"2020","unstructured":"Nikita Kitaev, \u0141ukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The efficient transformer. arXiv:2001.04451. Retrieved from https:\/\/arxiv.org\/abs\/2001.04451.","journal-title":"arXiv:2001.04451"},{"key":"e_1_3_2_60_2","article-title":"Word2bits-quantized word vectors","author":"Lam Maximilian","year":"2018","unstructured":"Maximilian Lam. 2018. Word2bits-quantized word vectors. arXiv:1803.05651. Retrieved from https:\/\/arxiv.org\/abs\/1803.05651.","journal-title":"arXiv:1803.05651"},{"key":"e_1_3_2_61_2","article-title":"ALBERT: A lite BERT for self-supervised learning of language representations","author":"Lan Zhenzhong","year":"2019","unstructured":"Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. ALBERT: A lite BERT for self-supervised learning of language representations. arXiv:1909.11942. Retrieved from https:\/\/arxiv.org\/abs\/1909.11942.","journal-title":"arXiv:1909.11942"},{"key":"e_1_3_2_62_2","unstructured":"Angeliki Lazaridou Eva Maria Vecchi and Marco Baroni. 2013. Fish transporters and miracle homes: How compositional distributional semantics can help NP parsing. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing . 1908\u20131913."},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.5555\/109230.109298"},{"issue":"3","key":"e_1_3_2_64_2","first-page":"1420","article-title":"Proximal newton-type methods for minimizing composite functions","volume":"24","author":"Lee Jason D.","year":"2014","unstructured":"Jason D. Lee, Yuekai Sun, and Michael A. Saunders. 2014. Proximal newton-type methods for minimizing composite functions. Journal of Optimization 24, 3 (2014), 1420\u20131443.","journal-title":"Journal of Optimization"},{"key":"e_1_3_2_65_2","article-title":"Gshard: Scaling giant models with conditional computation and automatic sharding","author":"Lepikhin Dmitry","year":"2020","unstructured":"Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv:2006.16668. Retrieved from https:\/\/arxiv.org\/abs\/2006.16668.","journal-title":"arXiv:2006.16668"},{"key":"e_1_3_2_66_2","article-title":"Ternary weight networks","author":"Li Fengfu","year":"2016","unstructured":"Fengfu Li, Bo Zhang, and Bin Liu. 2016. Ternary weight networks. arXiv:1605.04711. Retrieved from https:\/\/arxiv.org\/abs\/1605.04711.","journal-title":"arXiv:1605.04711"},{"key":"e_1_3_2_67_2","article-title":"BERT-EMD: Many-to-many layer mapping for BERT compression with earth mover\u2019s distance","author":"Li Jianquan","year":"2020","unstructured":"Jianquan Li, Xiaokang Liu, Honghong Zhao, Ruifeng Xu, Min Yang, and Yaohong Jin. 2020. BERT-EMD: Many-to-many layer mapping for BERT compression with earth mover\u2019s distance. arXiv:2010.06133. Retrieved from https:\/\/arxiv.org\/abs\/2010.06133.","journal-title":"arXiv:2010.06133"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.5555\/3157382.3157588"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.5555\/3504035.3504675"},{"key":"e_1_3_2_70_2","article-title":"Neural networks with few multiplications","author":"Lin Zhouhan","year":"2015","unstructured":"Zhouhan Lin, Matthieu Courbariaux, Roland Memisevic, and Yoshua Bengio. 2015. Neural networks with few multiplications. arXiv:1510.03009. Retrieved from https:\/\/arxiv.org\/abs\/1510.03009.","journal-title":"arXiv:1510.03009"},{"key":"e_1_3_2_71_2","volume-title":"Think Tank: Forty Neuroscientists Explore the Biological Roots of Human Experience","author":"Linden David J.","year":"2018","unstructured":"David J. Linden. 2018. Think Tank: Forty Neuroscientists Explore the Biological Roots of Human Experience. Yale University Press."},{"key":"e_1_3_2_72_2","article-title":"Finding function in form: Compositional character models for open vocabulary word representation","author":"Ling Wang","year":"2015","unstructured":"Wang Ling, Tiago Lu\u00eds, Lu\u00eds Marujo, Ram\u00f3n Fernandez Astudillo, Silvio Amir, Chris Dyer, Alan W. Black, and Isabel Trancoso. 2015. Finding function in form: Compositional character models for open vocabulary word representation. arXiv:1508.02096. Retrieved from https:\/\/arxiv.org\/abs\/1508.02096.","journal-title":"arXiv:1508.02096"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00115"},{"key":"e_1_3_2_74_2","article-title":"Improving multi-task deep neural networks via knowledge distillation for natural language understanding","author":"Liu Xiaodong","year":"2019","unstructured":"Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2019. Improving multi-task deep neural networks via knowledge distillation for natural language understanding. arXiv:1904.09482. Retrieved from https:\/\/arxiv.org\/abs\/1904.09482.","journal-title":"arXiv:1904.09482"},{"key":"e_1_3_2_75_2","article-title":"Multi-task deep neural networks for natural language understanding","author":"Liu Xiaodong","year":"2019","unstructured":"Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2019. Multi-task deep neural networks for natural language understanding. arXiv:1901.11504. Retrieved from https:\/\/arxiv.org\/abs\/1901.11504.","journal-title":"arXiv:1901.11504"},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.1982.1056489"},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7472821"},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454487"},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.3389\/fncom.2016.00131"},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3455544"},{"key":"e_1_3_2_81_2","volume-title":"Proceedings of the 1st International Conference on Learning Representations (ICLR\u201913), Scottsdale, Arizona, USA, May 2-4, 2013","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In Proceedings of the 1st International Conference on Learning Representations (ICLR\u201913), Scottsdale, Arizona, USA, May 2-4, 2013."},{"key":"e_1_3_2_82_2","article-title":"Improved knowledge distillation via teacher assistant","author":"Mirzadeh Seyed-Iman","year":"2019","unstructured":"Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh. 2019. Improved knowledge distillation via teacher assistant. arXiv:1902.03393. Retrieved from https:\/\/arxiv.org\/abs\/1902.03393.","journal-title":"arXiv:1902.03393"},{"key":"e_1_3_2_83_2","article-title":"Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy","author":"Mishra Asit","year":"2017","unstructured":"Asit Mishra and Debbie Marr. 2017. Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy. arXiv:1711.05852. Retrieved from https:\/\/arxiv.org\/abs\/1711.05852.","journal-title":"arXiv:1711.05852"},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.202"},{"key":"e_1_3_2_85_2","article-title":"Rounding methods for neural networks with low resolution synaptic weights","author":"Muller Lorenz K.","year":"2015","unstructured":"Lorenz K. Muller and Giacomo Indiveri. 2015. Rounding methods for neural networks with low resolution synaptic weights. arXiv:1504.05767. Retrieved from https:\/\/arxiv.org\/abs\/1504.05767.","journal-title":"arXiv:1504.05767"},{"key":"e_1_3_2_86_2","article-title":"Auto-sizing neural networks: With applications to n-gram language models","author":"Murray Kenton","year":"2015","unstructured":"Kenton Murray and David Chiang. 2015. Auto-sizing neural networks: With applications to n-gram language models. arXiv:1508.05051. Retrieved from https:\/\/arxiv.org\/abs\/1508.05051.","journal-title":"arXiv:1508.05051"},{"key":"e_1_3_2_87_2","article-title":"Exploring sparsity in recurrent neural networks","author":"Narang Sharan","year":"2017","unstructured":"Sharan Narang, Erich Elsen, Gregory Diamos, and Shubho Sengupta. 2017. Exploring sparsity in recurrent neural networks. arXiv:1704.05119. Retrieved from https:\/\/arxiv.org\/abs\/1704.05119.","journal-title":"arXiv:1704.05119"},{"key":"e_1_3_2_88_2","article-title":"Block-sparse recurrent neural networks","author":"Narang Sharan","year":"2017","unstructured":"Sharan Narang, Eric Undersander, and Gregory Diamos. 2017. Block-sparse recurrent neural networks. arXiv:1711.02782. Retrieved from https:\/\/arxiv.org\/abs\/1711.02782.","journal-title":"arXiv:1711.02782"},{"key":"e_1_3_2_89_2","doi-asserted-by":"publisher","DOI":"10.1137\/090752286"},{"key":"e_1_3_2_90_2","article-title":"Recurrent neural networks with limited numerical precision","author":"Ott Joachim","year":"2016","unstructured":"Joachim Ott, Zhouhan Lin, Ying Zhang, Shih-Chii Liu, and Yoshua Bengio. 2016. Recurrent neural networks with limited numerical precision. arXiv:1608.06902. Retrieved from https:\/\/arxiv.org\/abs\/1608.06902.","journal-title":"arXiv:1608.06902"},{"key":"e_1_3_2_91_2","article-title":"Dropneuron: Simplifying the structure of deep neural networks","author":"Pan Wei","year":"2016","unstructured":"Wei Pan, Hao Dong, and Yike Guo. 2016. Dropneuron: Simplifying the structure of deep neural networks. arXiv:1606.07326. Retrieved from https:\/\/arxiv.org\/abs\/1606.07326.","journal-title":"arXiv:1606.07326"},{"key":"e_1_3_2_92_2","article-title":"Model compression via distillation and quantization","author":"Polino Antonio","year":"2018","unstructured":"Antonio Polino, Razvan Pascanu, and Dan Alistarh. 2018. Model compression via distillation and quantization. arXiv:1802.05668. Retrieved from https:\/\/arxiv.org\/abs\/1802.05668.","journal-title":"arXiv:1802.05668"},{"key":"e_1_3_2_93_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7472823"},{"key":"e_1_3_2_94_2","article-title":"When BERT plays the lottery, all tickets are winning","author":"Prasanna Sai","year":"2020","unstructured":"Sai Prasanna, Anna Rogers, and Anna Rumshisky. 2020. When BERT plays the lottery, all tickets are winning. arXiv:2005.00561. Retrieved from https:\/\/arxiv.org\/abs\/2005.00561.","journal-title":"arXiv:2005.00561"},{"key":"e_1_3_2_95_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.232"},{"issue":"8","key":"e_1_3_2_96_2","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019).","journal-title":"OpenAI Blog"},{"key":"e_1_3_2_97_2","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","author":"Raffel Colin","year":"2019","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv:1910.10683. Retrieved from https:\/\/arxiv.org\/abs\/1910.10683.","journal-title":"arXiv:1910.10683"},{"key":"e_1_3_2_98_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_32"},{"key":"e_1_3_2_99_2","article-title":"Fitnets: Hints for thin deep nets","author":"Romero Adriana","year":"2014","unstructured":"Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2014. Fitnets: Hints for thin deep nets. arXiv:1412.6550. Retrieved from https:\/\/arxiv.org\/abs\/1412.6550.","journal-title":"arXiv:1412.6550"},{"key":"e_1_3_2_100_2","first-page":"2","article-title":"Turing-nlg: A 17-billion-parameter language model by microsoft","volume":"1","author":"Rosset C.","year":"2020","unstructured":"C. Rosset. 2020. Turing-nlg: A 17-billion-parameter language model by microsoft. Microsoft Blog 1 (2020), 2.","journal-title":"Microsoft Blog"},{"key":"e_1_3_2_101_2","article-title":"Long short-term memory based recurrent neural network architectures for large vocabulary speech recognition","author":"Sak Ha\u015fim","year":"2014","unstructured":"Ha\u015fim Sak, Andrew Senior, and Fran\u00e7oise Beaufays. 2014. Long short-term memory based recurrent neural network architectures for large vocabulary speech recognition. arXiv:1402.1128. Retrieved from https:\/\/arxiv.org\/abs\/1402.1128.","journal-title":"arXiv:1402.1128"},{"key":"e_1_3_2_102_2","article-title":"DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter","author":"Sanh Victor","year":"2019","unstructured":"Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv:1910.01108. Retrieved from https:\/\/arxiv.org\/abs\/1910.01108.","journal-title":"arXiv:1910.01108"},{"key":"e_1_3_2_103_2","article-title":"Deep model compression: Distilling knowledge from noisy teachers","author":"Sau Bharat Bhusan","year":"2016","unstructured":"Bharat Bhusan Sau and Vineeth N. Balasubramanian. 2016. Deep model compression: Distilling knowledge from noisy teachers. arXiv:1610.09650. Retrieved from https:\/\/arxiv.org\/abs\/1610.09650.","journal-title":"arXiv:1610.09650"},{"key":"e_1_3_2_104_2","article-title":"Compression of neural machine translation models via pruning","author":"See Abigail","year":"2016","unstructured":"Abigail See, Minh-Thang Luong, and Christopher D. Manning. 2016. Compression of neural machine translation models via pruning. arXiv:1606.09274. Retrieved from https:\/\/arxiv.org\/abs\/1606.09274.","journal-title":"arXiv:1606.09274"},{"key":"e_1_3_2_105_2","article-title":"Q-bert: Hessian based ultra low precision quantization of bert","author":"Shen Sheng","year":"2019","unstructured":"Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2019. Q-bert: Hessian based ultra low precision quantization of bert. arXiv:1909.05840. Retrieved from https:\/\/arxiv.org\/abs\/1909.05840.","journal-title":"arXiv:1909.05840"},{"key":"e_1_3_2_106_2","article-title":"Efficient attention: Attention with linear complexities","author":"Shen Zhuoran","year":"2018","unstructured":"Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li. 2018. Efficient attention: Attention with linear complexities. arXiv:1812.01243. Retrieved from https:\/\/arxiv.org\/abs\/1812.01243.","journal-title":"arXiv:1812.01243"},{"key":"e_1_3_2_107_2","article-title":"Megatron-lm: Training multi-billion parameter language models using gpu model parallelism","author":"Shoeybi Mohammad","year":"2019","unstructured":"Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using gpu model parallelism. arXiv:1909.08053. Retrieved from https:\/\/arxiv.org\/abs\/1909.08053.","journal-title":"arXiv:1909.08053"},{"key":"e_1_3_2_108_2","article-title":"Compressing word embeddings via deep compositional code learning","author":"Shu Raphael","year":"2017","unstructured":"Raphael Shu and Hideki Nakayama. 2017. Compressing word embeddings via deep compositional code learning. arXiv:1711.01068. Retrieved from https:\/\/arxiv.org\/abs\/1711.01068.","journal-title":"arXiv:1711.01068"},{"key":"e_1_3_2_109_2","article-title":"Data-free parameter pruning for deep neural networks","author":"Srinivas Suraj","year":"2015","unstructured":"Suraj Srinivas and R. Venkatesh Babu. 2015. Data-free parameter pruning for deep neural networks. arXiv:1507.06149. Retrieved from https:\/\/arxiv.org\/abs\/1507.06149.","journal-title":"arXiv:1507.06149"},{"key":"e_1_3_2_110_2","article-title":"Energy and policy considerations for deep learning in NLP","author":"Strubell Emma","year":"2019","unstructured":"Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. Energy and policy considerations for deep learning in NLP. arXiv:1906.02243. Retrieved from https:\/\/arxiv.org\/abs\/1906.02243.","journal-title":"arXiv:1906.02243"},{"key":"e_1_3_2_111_2","article-title":"Patient knowledge distillation for bert model compression","author":"Sun Siqi","year":"2019","unstructured":"Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019. Patient knowledge distillation for bert model compression. arXiv:1908.09355. Retrieved from https:\/\/arxiv.org\/abs\/1908.09355.","journal-title":"arXiv:1908.09355"},{"key":"e_1_3_2_112_2","article-title":"Mobilebert: A compact task-agnostic bert for resource-limited devices","author":"Sun Zhiqing","year":"2020","unstructured":"Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020. Mobilebert: A compact task-agnostic bert for resource-limited devices. arXiv:2004.02984. Retrieved from https:\/\/arxiv.org\/abs\/2004.02984.","journal-title":"arXiv:2004.02984"},{"key":"e_1_3_2_113_2","doi-asserted-by":"publisher","DOI":"10.5555\/3060832.3060907"},{"key":"e_1_3_2_114_2","article-title":"Multilingual neural machine translation with knowledge distillation","author":"Tan Xu","year":"2019","unstructured":"Xu Tan, Yi Ren, Di He, Tao Qin, Zhou Zhao, and Tie-Yan Liu. 2019. Multilingual neural machine translation with knowledge distillation. arXiv:1902.10461. Retrieved from https:\/\/arxiv.org\/abs\/1902.10461.","journal-title":"arXiv:1902.10461"},{"key":"e_1_3_2_115_2","article-title":"Distilling task-specific knowledge from BERT into simple neural networks","author":"Tang Raphael","year":"2019","unstructured":"Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019. Distilling task-specific knowledge from BERT into simple neural networks. arXiv:1903.12136. Retrieved from https:\/\/arxiv.org\/abs\/1903.12136.","journal-title":"arXiv:1903.12136"},{"key":"e_1_3_2_116_2","article-title":"Sparse sinkhorn attention","author":"Tay Yi","year":"2020","unstructured":"Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan. 2020. Sparse sinkhorn attention. arXiv:2002.11296. Retrieved from https:\/\/arxiv.org\/abs\/2002.11296.","journal-title":"arXiv:2002.11296"},{"key":"e_1_3_2_117_2","article-title":"Lightweight and efficient neural natural language processing with quaternion networks","author":"Tay Yi","year":"2019","unstructured":"Yi Tay, Aston Zhang, Luu Anh Tuan, Jinfeng Rao, Shuai Zhang, Shuohang Wang, Jie Fu, and Siu Cheung Hui. 2019. Lightweight and efficient neural natural language processing with quaternion networks. arXiv:1906.04393. Retrieved from https:\/\/arxiv.org\/abs\/1906.04393.","journal-title":"arXiv:1906.04393"},{"key":"e_1_3_2_118_2","first-page":"4451","volume-title":"Proceedings of the","author":"Tjandra Andros","year":"2017","unstructured":"Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2017. Compressing recurrent neural network with tensor train. In Proceedings of theInternational Joint Conference on Neural Networks. IEEE, 4451\u20134458."},{"key":"e_1_3_2_119_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF02289464"},{"key":"e_1_3_2_120_2","article-title":"Well-read students learn better: The impact of student initialization on knowledge distillation","author":"Turc Iulia","year":"2019","unstructured":"Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Well-read students learn better: The impact of student initialization on knowledge distillation. arXiv:1908.08962. Retrieved from https:\/\/arxiv.org\/abs\/1908.08962.","journal-title":"arXiv:1908.08962"},{"key":"e_1_3_2_121_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683694"},{"key":"e_1_3_2_122_2","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_2_123_2","article-title":"Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned","author":"Voita Elena","year":"2019","unstructured":"Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv:1905.09418. Retrieved from https:\/\/arxiv.org\/abs\/1905.09418.","journal-title":"arXiv:1905.09418"},{"key":"e_1_3_2_124_2","doi-asserted-by":"publisher","DOI":"10.1038\/502172a"},{"key":"e_1_3_2_125_2","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454581"},{"key":"e_1_3_2_126_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Wang Alex","year":"2019","unstructured":"Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_127_2","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Lin","year":"2020","unstructured":"Lin Wang and Kuk-Jin Yoon. 2020. Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE CVPR."},{"key":"e_1_3_2_128_2","doi-asserted-by":"publisher","DOI":"10.5555\/3326943.3326999"},{"key":"e_1_3_2_129_2","article-title":"Linformer: Self-attention with linear complexity","author":"Wang Sinong","year":"2020","unstructured":"Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. 2020. Linformer: Self-attention with linear complexity. arXiv:2006.04768. Retrieved from https:\/\/arxiv.org\/abs\/2006.04768.","journal-title":"arXiv:2006.04768"},{"key":"e_1_3_2_130_2","article-title":"Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers","author":"Wang Wenhui","year":"2020","unstructured":"Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020. Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. arXiv:2002.10957. Retrieved from https:\/\/arxiv.org\/abs\/2002.10957.","journal-title":"arXiv:2002.10957"},{"key":"e_1_3_2_131_2","article-title":"Structured pruning of large language models","author":"Wang Ziheng","year":"2019","unstructured":"Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2019. Structured pruning of large language models. arXiv:1910.04732. Retrieved from https:\/\/arxiv.org\/abs\/1910.04732.","journal-title":"arXiv:1910.04732"},{"key":"e_1_3_2_132_2","article-title":"Attending to mathematical language with transformers","author":"Wangperawong Artit","year":"2018","unstructured":"Artit Wangperawong. 2018. Attending to mathematical language with transformers. arXiv:1812.02825. Retrieved from https:\/\/arxiv.org\/abs\/1812.02825.","journal-title":"arXiv:1812.02825"},{"key":"e_1_3_2_133_2","article-title":"Alternating multi-bit quantization for recurrent neural networks","author":"Xu Chen","year":"2018","unstructured":"Chen Xu, Jianqiang Yao, Zhouchen Lin, Wenwu Ou, Yuanbin Cao, Zhirong Wang, and Hongbin Zha. 2018. Alternating multi-bit quantization for recurrent neural networks. arXiv:1802.00150. Retrieved from https:\/\/arxiv.org\/abs\/1802.00150.","journal-title":"arXiv:1802.00150"},{"key":"e_1_3_2_134_2","article-title":"Bert-of-theseus: Compressing bert by progressive module replacing","author":"Xu Canwen","year":"2020","unstructured":"Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou. 2020. Bert-of-theseus: Compressing bert by progressive module replacing. arXiv:2002.02925. Retrieved from https:\/\/arxiv.org\/abs\/2002.02925.","journal-title":"arXiv:2002.02925"},{"key":"e_1_3_2_135_2","doi-asserted-by":"crossref","unstructured":"Yafeng Yang Kaihuan Liang Xuefeng Xiao Zecheng Xie Lianwen Jin Jun Sun and Weiying Zhou. 2018. Accelerating and compressing LSTM based model for online handwritten chinese character recognition. In Proceedings of the International Workshop on Frontiers in Handwriting Recognition . IEEE 110\u2013115.","DOI":"10.1109\/ICFHR-2018.2018.00028"},{"key":"e_1_3_2_136_2","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454804"},{"key":"e_1_3_2_137_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00977"},{"key":"e_1_3_2_138_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.754"},{"key":"e_1_3_2_139_2","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098135"},{"key":"e_1_3_2_140_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00711"},{"key":"e_1_3_2_141_2","article-title":"Learning to execute","author":"Zaremba Wojciech","year":"2014","unstructured":"Wojciech Zaremba and Ilya Sutskever. 2014. Learning to execute. arXiv:1410.4615. Retrieved from https:\/\/arxiv.org\/abs\/1410.4615.","journal-title":"arXiv:1410.4615"},{"key":"e_1_3_2_142_2","doi-asserted-by":"crossref","unstructured":"Ying Zhang Tao Xiang Timothy M. Hospedales and Huchuan Lu. 2018. Deep mutual learning. In Proceedings of the Conference on Computer Vision and Pattern Recognition . 4320\u20134328.","DOI":"10.1109\/CVPR.2018.00454"},{"key":"e_1_3_2_143_2","article-title":"Extreme language model compression with optimal subwords and shared projections","author":"Zhao Sanqiang","year":"2019","unstructured":"Sanqiang Zhao, Raghav Gupta, Yang Song, and Denny Zhou. 2019. Extreme language model compression with optimal subwords and shared projections. arXiv:1909.11687. Retrieved from https:\/\/arxiv.org\/abs\/1909.11687.","journal-title":"arXiv:1909.11687"},{"key":"e_1_3_2_144_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11390-017-1750-y"},{"key":"e_1_3_2_145_2","article-title":"Trained ternary quantization","author":"Zhu Chenzhuo","year":"2016","unstructured":"Chenzhuo Zhu, Song Han, Huizi Mao, and William J. Dally. 2016. Trained ternary quantization. arXiv:1612.01064. Retrieved from https:\/\/arxiv.org\/abs\/1612.01064.","journal-title":"arXiv:1612.01064"},{"key":"e_1_3_2_146_2","article-title":"To prune, or not to prune: Exploring the efficacy of pruning for model compression","author":"Zhu Michael","year":"2017","unstructured":"Michael Zhu and Suyog Gupta. 2017. To prune, or not to prune: Exploring the efficacy of pruning for model compression. arXiv:1710.01878. Retrieved from https:\/\/arxiv.org\/abs\/1710.01878.","journal-title":"arXiv:1710.01878"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3487045","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3487045","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:18:47Z","timestamp":1750177127000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3487045"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,8]]},"references-count":145,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,8,31]]}},"alternative-id":["10.1145\/3487045"],"URL":"https:\/\/doi.org\/10.1145\/3487045","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,8]]},"assertion":[{"value":"2020-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-01-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}