{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,7]],"date-time":"2026-08-07T14:33:26Z","timestamp":1786113206038,"version":"3.56.0"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2021,4,15]],"date-time":"2021-04-15T00:00:00Z","timestamp":1618444800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS-1717657, CCF-1937435, CNS-1822085"],"award-info":[{"award-number":["CNS-1717657, CCF-1937435, CNS-1822085"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Cyber-Phys. Syst."],"published-print":{"date-parts":[[2021,7,31]]},"abstract":"<jats:p>The invention of Transformer model structure boosts the performance of Neural Machine Translation (NMT) tasks to an unprecedented level. Many previous works have been done to make the Transformer model more execution-friendly on resource-constrained platforms. These researches can be categorized into three key fields: Model Pruning, Transfer Learning, and Efficient Transformer Variants. The family of model pruning methods are popular for their simplicity in practice and promising compression rate and have achieved great success in the field of convolution neural networks (CNNs) for many vision tasks. Nonetheless, previous Transformer pruning works did not perform a thorough model analysis and evaluation on each Transformer component on off-the-shelf mobile devices. In this work, we analyze and prune transformer models at the line-wise granularity and also implement our pruning method on real mobile platforms. We explore the properties of all Transformer components as well as their sparsity features, which are leveraged to guide Transformer model pruning. We name our whole Transformer analysis and pruning pipeline as TPrune. In TPrune, we first propose Block-wise Structured Sparsity Learning (BSSL) to analyze Transformer model property. Then, based on the characters derived from BSSL, we apply Structured Hoyer Square (SHS) to derive the final pruned models. Comparing with the state-of-the-art Transformer pruning methods, TPrune is able to achieve a higher model compression rate with less performance degradation. Experimental results show that our pruned models achieve 1.16\u00d7\u20131.92\u00d7 speedup on mobile devices with 0%\u20138% BLEU score degradation compared with the original Transformer model.<\/jats:p>","DOI":"10.1145\/3446640","type":"journal-article","created":{"date-parts":[[2021,4,15]],"date-time":"2021-04-15T10:58:46Z","timestamp":1618484326000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":41,"title":["TPrune"],"prefix":"10.1145","volume":"5","author":[{"given":"Jiachen","family":"Mao","sequence":"first","affiliation":[{"name":"Duke University, Durham, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huanrui","family":"Yang","sequence":"additional","affiliation":[{"name":"Duke University, Durham, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ang","family":"Li","sequence":"additional","affiliation":[{"name":"Duke University, Durham, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hai","family":"Li","sequence":"additional","affiliation":[{"name":"Duke University, Durham, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yiran","family":"Chen","sequence":"additional","affiliation":[{"name":"Duke University, Durham, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,4,15]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems","author":"Vaswani Ashish","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N. Gomez , \u0141ukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need . In Proceedings of the International Conference on Advances in Neural Information Processing Systems , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.). Curran Associates, Inc. , 5998--6008. Retrieved from http:\/\/papers.nips.cc\/paper\/7181-attention-is-all-you-need.pdf. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the International Conference on Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.). Curran Associates, Inc., 5998--6008. Retrieved from http:\/\/papers.nips.cc\/paper\/7181-attention-is-all-you-need.pdf."},{"key":"e_1_2_1_2_1","volume-title":"BERT: Pre-training of deep bidirectional transformers for language understanding. CoRR abs\/1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . BERT: Pre-training of deep bidirectional transformers for language understanding. CoRR abs\/1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. CoRR abs\/1810.04805 (2018)."},{"key":"e_1_2_1_3_1","volume-title":"Le","author":"Yang Zhilin","year":"2019","unstructured":"Zhilin Yang , Zihang Dai , Yiming Yang , Jaime Carbonell , Russ R. Salakhutdinov , and Quoc V . Le . 2019 . XLNet: Generalized autoregressive pretraining for language understanding. In Proceedings of the International Conference on Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch\u00e9-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., 5754--5764. Retrieved from http:\/\/papers.nips.cc\/paper\/8812-xlnet-generalized-autoregressive-pretraining-for-language-understanding.pdf. Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R. Salakhutdinov, and Quoc V. Le. 2019. XLNet: Generalized autoregressive pretraining for language understanding. In Proceedings of the International Conference on Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch\u00e9-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., 5754--5764. Retrieved from http:\/\/papers.nips.cc\/paper\/8812-xlnet-generalized-autoregressive-pretraining-for-language-understanding.pdf."},{"key":"e_1_2_1_4_1","volume-title":"RoBERTa: A robustly optimized BERT pretraining approach. CoRR abs\/1907.11692","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu , Myle Ott , Naman Goyal , Jingfei Du , Mandar Joshi , Danqi Chen , Omer Levy , Mike Lewis , Luke Zettlemoyer , and Veselin Stoyanov . 2019. RoBERTa: A robustly optimized BERT pretraining approach. CoRR abs\/1907.11692 ( 2019 ). Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. CoRR abs\/1907.11692 (2019)."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3316781.3317865"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.23919\/DATE.2017.7927211"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240765.3240851"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3316781.3317936"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design (ICCAD\u201917)","author":"Mao J.","year":"2017","unstructured":"J. Mao , Z. Qin , Z. Xu , K. W. Nixon , X. Chen , H. Li , and Y. Chen . 2017. AdaLearner: An adaptive distributed mobile learning system for neural networks . In Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design (ICCAD\u201917) . 291--296. DOI:http:\/\/dx.doi.org\/10.1109\/ICCAD. 2017 .8203791 J. Mao, Z. Qin, Z. Xu, K. W. Nixon, X. Chen, H. Li, and Y. Chen. 2017. AdaLearner: An adaptive distributed mobile learning system for neural networks. In Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design (ICCAD\u201917). 291--296. DOI:http:\/\/dx.doi.org\/10.1109\/ICCAD.2017.8203791"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASPDAC.2018.8297378"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2019.00027"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3287624.3287642"},{"key":"e_1_2_1_13_1","volume-title":"Dally","author":"Han Song","year":"2015","unstructured":"Song Han , Huizi Mao , and William J . Dally . 2015 . Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. arXiv preprint arXiv:1510.00149 (2015). Song Han, Huizi Mao, and William J. Dally. 2015. Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. arXiv preprint arXiv:1510.00149 (2015)."},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems, D. D. Lee, M. Sugiyama, U. V. Luxburg, I. Guyon, and R. Garnett (Eds.). Curran Associates, Inc.","author":"Wen Wei","year":"2016","unstructured":"Wei Wen , Chunpeng Wu , Yandan Wang , Yiran Chen , and Hai Li . 2016 . Learning structured sparsity in deep neural networks . In Proceedings of the International Conference on Advances in Neural Information Processing Systems, D. D. Lee, M. Sugiyama, U. V. Luxburg, I. Guyon, and R. Garnett (Eds.). Curran Associates, Inc. , 2074--2082. Retrieved from http:\/\/papers.nips.cc\/paper\/6504-learning-structured-sparsity-in-deep-neural-networks.pdf. Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2016. Learning structured sparsity in deep neural networks. In Proceedings of the International Conference on Advances in Neural Information Processing Systems, D. D. Lee, M. Sugiyama, U. V. Luxburg, I. Guyon, and R. Garnett (Eds.). Curran Associates, Inc., 2074--2082. Retrieved from http:\/\/papers.nips.cc\/paper\/6504-learning-structured-sparsity-in-deep-neural-networks.pdf."},{"key":"e_1_2_1_15_1","volume-title":"Learning intrinsic sparse structures within long short-term memory. arXiv preprint arXiv:1709.05027","author":"Wen Wei","year":"2017","unstructured":"Wei Wen , Yuxiong He , Samyam Rajbhandari , Minjia Zhang , Wenhan Wang , Fang Liu , Bin Hu , Yiran Chen , and Hai Li. 2017. Learning intrinsic sparse structures within long short-term memory. arXiv preprint arXiv:1709.05027 ( 2017 ). Wei Wen, Yuxiong He, Samyam Rajbhandari, Minjia Zhang, Wenhan Wang, Fang Liu, Bin Hu, Yiran Chen, and Hai Li. 2017. Learning intrinsic sparse structures within long short-term memory. arXiv preprint arXiv:1709.05027 (2017)."},{"key":"e_1_2_1_16_1","volume-title":"DASNet: Dynamic activation sparsity for neural network efficiency improvement. arXiv preprint arXiv:1909.06964","author":"Yang Qing","year":"2019","unstructured":"Qing Yang , Jiachen Mao , Zuoguan Wang , and Hai Li. 2019. DASNet: Dynamic activation sparsity for neural network efficiency improvement. arXiv preprint arXiv:1909.06964 ( 2019 ). Qing Yang, Jiachen Mao, Zuoguan Wang, and Hai Li. 2019. DASNet: Dynamic activation sparsity for neural network efficiency improvement. arXiv preprint arXiv:1909.06964 (2019)."},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"\u00a0al Mengye Ren","year":"2018","unstructured":"Mengye Ren et \u00a0al . 2018 . SBNet: Sparse blocks network for fast inference . In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR\u201918) . 8711--8720. Mengye Ren et\u00a0al. 2018. SBNet: Sparse blocks network for fast inference. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR\u201918). 8711--8720."},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems. 947--955","author":"\u00a0al Mikhail Figurnov","year":"2016","unstructured":"Mikhail Figurnov et \u00a0al . 2016 . PerforatedCNNs: Acceleration through elimination of redundant convolutions . In Proceedings of the International Conference on Advances in Neural Information Processing Systems. 947--955 . Mikhail Figurnov et\u00a0al. 2016. PerforatedCNNs: Acceleration through elimination of redundant convolutions. In Proceedings of the International Conference on Advances in Neural Information Processing Systems. 947--955."},{"key":"e_1_2_1_19_1","volume-title":"Reducing transformer depth on demand with structured dropout. arXiv preprint arXiv:1909.11556","author":"Fan Angela","year":"2019","unstructured":"Angela Fan , Edouard Grave , and Armand Joulin . 2019. Reducing transformer depth on demand with structured dropout. arXiv preprint arXiv:1909.11556 ( 2019 ). Angela Fan, Edouard Grave, and Armand Joulin. 2019. Reducing transformer depth on demand with structured dropout. arXiv preprint arXiv:1909.11556 (2019)."},{"key":"e_1_2_1_20_1","volume-title":"Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv preprint arXiv:1905.09418","author":"Voita Elena","year":"2019","unstructured":"Elena Voita , David Talbot , Fedor Moiseev , Rico Sennrich , and Ivan Titov . 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv preprint arXiv:1905.09418 ( 2019 ). Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv preprint arXiv:1905.09418 (2019)."},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems. 14014--14024","author":"Michel Paul","year":"2019","unstructured":"Paul Michel , Omer Levy , and Graham Neubig . 2019 . Are sixteen heads really better than one? In Proceedings of the International Conference on Advances in Neural Information Processing Systems. 14014--14024 . Paul Michel, Omer Levy, and Graham Neubig. 2019. Are sixteen heads really better than one? In Proceedings of the International Conference on Advances in Neural Information Processing Systems. 14014--14024."},{"key":"e_1_2_1_22_1","volume-title":"Auto-sizing the transformer network: Improving speed, efficiency, and performance for low-resource machine translation. arXiv preprint arXiv:1910.06717","author":"Murray Kenton","year":"2019","unstructured":"Kenton Murray , Jeffery Kinnison , Toan Q. Nguyen , Walter Scheirer , and David Chiang . 2019. Auto-sizing the transformer network: Improving speed, efficiency, and performance for low-resource machine translation. arXiv preprint arXiv:1910.06717 ( 2019 ). Kenton Murray, Jeffery Kinnison, Toan Q. Nguyen, Walter Scheirer, and David Chiang. 2019. Auto-sizing the transformer network: Improving speed, efficiency, and performance for low-resource machine translation. arXiv preprint arXiv:1910.06717 (2019)."},{"key":"e_1_2_1_24_1","volume-title":"a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108","author":"Sanh Victor","year":"2019","unstructured":"Victor Sanh , Lysandre Debut , Julien Chaumond , and Thomas Wolf . 2019. DistilBERT , a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108 ( 2019 ). Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108 (2019)."},{"key":"e_1_2_1_25_1","volume-title":"Mobilebert: A compact task-agnostic bert for resource-limited devices. arXiv preprint arXiv:2004.02984.","author":"Sun Zhiqing","year":"2020","unstructured":"Zhiqing Sun , Hongkun Yu , Xiaodan Song , Renjie Liu , Yiming Yang , and Denny Zhou . 2020 . Mobilebert: A compact task-agnostic bert for resource-limited devices. arXiv preprint arXiv:2004.02984. Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020. Mobilebert: A compact task-agnostic bert for resource-limited devices. arXiv preprint arXiv:2004.02984."},{"key":"e_1_2_1_26_1","volume-title":"Accelerating neural transformer via an average attention network. arXiv preprint arXiv:1805.00631","author":"Zhang Biao","year":"2018","unstructured":"Biao Zhang , Deyi Xiong , and Jinsong Su. 2018. Accelerating neural transformer via an average attention network. arXiv preprint arXiv:1805.00631 ( 2018 ). Biao Zhang, Deyi Xiong, and Jinsong Su. 2018. Accelerating neural transformer via an average attention network. arXiv preprint arXiv:1805.00631 (2018)."},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the International Conference on Learning Representitive (ICLR\u201920)","author":"Wu Zhanghao","year":"2020","unstructured":"Zhanghao Wu , Zhijian Liu , Ji Lin , Yujun Lin , and Song Han . 2020 . Efficient transformer for mobile applicatoins . In Proceedings of the International Conference on Learning Representitive (ICLR\u201920) . Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han. 2020. Efficient transformer for mobile applicatoins. In Proceedings of the International Conference on Learning Representitive (ICLR\u201920)."},{"key":"e_1_2_1_28_1","volume-title":"Klaus Macherey et\u00a0al","author":"Wu Yonghui","year":"2016","unstructured":"Yonghui Wu , Mike Schuster , Zhifeng Chen , Quoc V. Le , Mohammad Norouzi , Wolfgang Macherey , Maxim Krikun , Yuan Cao , Qin Gao , Klaus Macherey et\u00a0al . 2016 . Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144 (2016). Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey et\u00a0al. 2016. Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144 (2016)."},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the 31st AAAI Conference on Artificial Intelligence.","author":"Szegedy Christian","unstructured":"Christian Szegedy , Sergey Ioffe , Vincent Vanhoucke , and Alexander A. Alemi . 2017. Inception-v4, Inception-ResNet and the impact of residual connections on learning . In Proceedings of the 31st AAAI Conference on Artificial Intelligence. Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A. Alemi. 2017. Inception-v4, Inception-ResNet and the impact of residual connections on learning. In Proceedings of the 31st AAAI Conference on Artificial Intelligence."},{"key":"e_1_2_1_30_1","volume-title":"Geoffrey Hinton et\u00a0al","author":"Krizhevsky Alex","year":"2009","unstructured":"Alex Krizhevsky , Geoffrey Hinton et\u00a0al . 2009 . Learning multiple layers of features from tiny images. Citeseer . Alex Krizhevsky, Geoffrey Hinton et\u00a0al. 2009. Learning multiple layers of features from tiny images. Citeseer."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAD.2017.8203852"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_2_1_33_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_2_1_34_1","volume-title":"Multi-granularity self-attention for neural machine translation. arXiv preprint arXiv:1909.02222","author":"Hao Jie","year":"2019","unstructured":"Jie Hao , Xing Wang , Shuming Shi , Jinfeng Zhang , and Zhaopeng Tu. 2019. Multi-granularity self-attention for neural machine translation. arXiv preprint arXiv:1909.02222 ( 2019 ). Jie Hao, Xing Wang, Shuming Shi, Jinfeng Zhang, and Zhaopeng Tu. 2019. Multi-granularity self-attention for neural machine translation. arXiv preprint arXiv:1909.02222 (2019)."},{"key":"e_1_2_1_35_1","volume-title":"Structured pruning of large language models. arXiv preprint arXiv:1910.04732","author":"Wang Ziheng","year":"2019","unstructured":"Ziheng Wang , Jeremy Wohlwend , and Tao Lei . 2019. Structured pruning of large language models. arXiv preprint arXiv:1910.04732 ( 2019 ). Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2019. Structured pruning of large language models. arXiv preprint arXiv:1910.04732 (2019)."},{"key":"e_1_2_1_36_1","volume-title":"Pruning a BERT-based question answering model. arXiv preprint arXiv:1910.06360","author":"McCarley J. S.","year":"2019","unstructured":"J. S. McCarley . 2019. Pruning a BERT-based question answering model. arXiv preprint arXiv:1910.06360 ( 2019 ). J. S. McCarley. 2019. Pruning a BERT-based question answering model. arXiv preprint arXiv:1910.06360 (2019)."},{"key":"e_1_2_1_37_1","volume-title":"Reweighted proximal pruning for large-scale language representation. arXiv preprint arXiv:1909.12486","author":"Guo Fu-Ming","year":"2019","unstructured":"Fu-Ming Guo , Sijia Liu , Finlay S. Mungall , Xue Lin , and Yanzhi Wang . 2019. Reweighted proximal pruning for large-scale language representation. arXiv preprint arXiv:1909.12486 ( 2019 ). Fu-Ming Guo, Sijia Liu, Finlay S. Mungall, Xue Lin, and Yanzhi Wang. 2019. Reweighted proximal pruning for large-scale language representation. arXiv preprint arXiv:1909.12486 (2019)."},{"key":"e_1_2_1_38_1","volume-title":"Q-BERT: Hessian based ultra low precision quantization of BERT. arXiv preprint arXiv:1909.05840","author":"Shen Sheng","year":"2019","unstructured":"Sheng Shen , Zhen Dong , Jiayu Ye , Linjian Ma , Zhewei Yao , Amir Gholami , Michael W. Mahoney , and Kurt Keutzer . 2019. Q-BERT: Hessian based ultra low precision quantization of BERT. arXiv preprint arXiv:1909.05840 ( 2019 ). Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. 2019. Q-BERT: Hessian based ultra low precision quantization of BERT. arXiv preprint arXiv:1909.05840 (2019)."},{"key":"e_1_2_1_39_1","volume-title":"Q8BERT: Quantized 8bit BERT. arXiv preprint arXiv:1910.06188","author":"Zafrir Ofir","year":"2019","unstructured":"Ofir Zafrir , Guy Boudoukh , Peter Izsak , and Moshe Wasserblat . 2019. Q8BERT: Quantized 8bit BERT. arXiv preprint arXiv:1910.06188 ( 2019 ). Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019. Q8BERT: Quantized 8bit BERT. arXiv preprint arXiv:1910.06188 (2019)."},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems. 1135--1143","author":"Han Song","year":"2015","unstructured":"Song Han , Jeff Pool , John Tran , and William Dally . 2015 . Learning both weights and connections for efficient neural network . In Proceedings of the International Conference on Advances in Neural Information Processing Systems. 1135--1143 . Song Han, Jeff Pool, John Tran, and William Dally. 2015. Learning both weights and connections for efficient neural network. In Proceedings of the International Conference on Advances in Neural Information Processing Systems. 1135--1143."},{"key":"e_1_2_1_41_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 806--814","author":"Liu Baoyuan","year":"2015","unstructured":"Baoyuan Liu , Min Wang , Hassan Foroosh , Marshall Tappen , and Marianna Pensky . 2015 . Sparse convolutional neural networks . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 806--814 . Baoyuan Liu, Min Wang, Hassan Foroosh, Marshall Tappen, and Marianna Pensky. 2015. Sparse convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 806--814."},{"key":"e_1_2_1_42_1","volume-title":"DeepHoyer: Learning sparser neural network with differentiable scale-invariant sparsity measures. arXiv preprint arXiv:1908.09979","author":"Yang Huanrui","year":"2019","unstructured":"Huanrui Yang , Wei Wen , and Hai Li. 2019. DeepHoyer: Learning sparser neural network with differentiable scale-invariant sparsity measures. arXiv preprint arXiv:1908.09979 ( 2019 ). Huanrui Yang, Wei Wen, and Hai Li. 2019. DeepHoyer: Learning sparser neural network with differentiable scale-invariant sparsity measures. arXiv preprint arXiv:1908.09979 (2019)."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1167"},{"key":"e_1_2_1_44_1","volume-title":"Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201916)","author":"\u00a0al Mart\u00edn Abadi","year":"2016","unstructured":"Mart\u00edn Abadi et \u00a0al . 2016 . Tensorflow: A system for large-scale machine learning . In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201916) . 265--283. Mart\u00edn Abadi et\u00a0al. 2016. Tensorflow: A system for large-scale machine learning. In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201916). 265--283."},{"key":"e_1_2_1_45_1","volume-title":"Tensor2Tensor for neural machine translation. CoRR abs\/1803.07416","author":"Vaswani Ashish","year":"2018","unstructured":"Ashish Vaswani , Samy Bengio , Eugene Brevdo , Francois Chollet , Aidan N. Gomez , Stephan Gouws , Llion Jones , \u0141ukasz Kaiser , Nal Kalchbrenner , Niki Parmar , Ryan Sepassi , Noam Shazeer , and Jakob Uszkoreit . 2018. Tensor2Tensor for neural machine translation. CoRR abs\/1803.07416 ( 2018 ). Ashish Vaswani, Samy Bengio, Eugene Brevdo, Francois Chollet, Aidan N. Gomez, Stephan Gouws, Llion Jones, \u0141ukasz Kaiser, Nal Kalchbrenner, Niki Parmar, Ryan Sepassi, Noam Shazeer, and Jakob Uszkoreit. 2018. Tensor2Tensor for neural machine translation. CoRR abs\/1803.07416 (2018)."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.2478\/pralin-2018-0002"}],"container-title":["ACM Transactions on Cyber-Physical Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3446640","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3446640","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3446640","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:47:31Z","timestamp":1750193251000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3446640"}},"subtitle":["Efficient Transformer Pruning for Mobile Devices"],"short-title":[],"issued":{"date-parts":[[2021,4,15]]},"references-count":45,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2021,7,31]]}},"alternative-id":["10.1145\/3446640"],"URL":"https:\/\/doi.org\/10.1145\/3446640","relation":{},"ISSN":["2378-962X","2378-9638"],"issn-type":[{"value":"2378-962X","type":"print"},{"value":"2378-9638","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,4,15]]},"assertion":[{"value":"2020-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-04-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}