{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T18:36:43Z","timestamp":1772303803514,"version":"3.50.1"},"reference-count":88,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2022,6,6]],"date-time":"2022-06-06T00:00:00Z","timestamp":1654473600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2022,9,30]]},"abstract":"<jats:p>\n            Weight pruning is an effective model compression technique to tackle the challenges of achieving real-time deep neural network (DNN) inference on mobile devices. However, prior pruning schemes have limited application scenarios due to accuracy degradation, difficulty in leveraging hardware acceleration, and\/or restriction on certain types of DNN layers. In this article, we propose a general, fine-grained structured pruning scheme and corresponding compiler optimizations that are applicable to any type of DNN layer while achieving high accuracy and hardware inference performance. With the flexibility of applying different pruning schemes to different layers enabled by our compiler optimizations, we further probe into the new problem of determining the best-suited pruning scheme considering the different acceleration and accuracy performance of various pruning schemes. Two pruning scheme mapping methods\u2014one -search based and the other is rule based\u2014are proposed to automatically derive the best-suited pruning regularity and block size for each layer of any given DNN. Experimental results demonstrate that our pruning scheme mapping methods, together with the general fine-grained structured pruning scheme, outperform the state-of-the-art DNN optimization framework with up to 2.48\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\( \\times \\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            and 1.73\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\( \\times \\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            DNN inference acceleration on CIFAR-10 and ImageNet datasets without accuracy loss.\n          <\/jats:p>","DOI":"10.1145\/3495532","type":"journal-article","created":{"date-parts":[[2022,2,24]],"date-time":"2022-02-24T17:13:41Z","timestamp":1645722821000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Automatic Mapping of the Best-Suited DNN Pruning Schemes for Real-Time Mobile Acceleration"],"prefix":"10.1145","volume":"27","author":[{"given":"Yifan","family":"Gong","sequence":"first","affiliation":[{"name":"Northeastern University, Boston, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9844-992X","authenticated-orcid":false,"given":"Geng","family":"Yuan","sequence":"additional","affiliation":[{"name":"Northeastern University, Boston, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zheng","family":"Zhan","sequence":"additional","affiliation":[{"name":"Northeastern University, Boston, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Niu","sequence":"additional","affiliation":[{"name":"College of William and Mary, Williamsburg, VA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhengang","family":"Li","sequence":"additional","affiliation":[{"name":"Northeastern University, Boston, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pu","family":"Zhao","sequence":"additional","affiliation":[{"name":"Northeastern University, Boston, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuxuan","family":"Cai","sequence":"additional","affiliation":[{"name":"Northeastern University, Boston, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sijia","family":"Liu","sequence":"additional","affiliation":[{"name":"Michigan State University, East Lansing, MI"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bin","family":"Ren","sequence":"additional","affiliation":[{"name":"College of William and Mary, Williamsburg, VA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xue","family":"Lin","sequence":"additional","affiliation":[{"name":"Northeastern University, Boston, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xulong","family":"Tang","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, Pittsburgh, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yanzhi","family":"Wang","sequence":"additional","affiliation":[{"name":"Northeastern University, Boston, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,6,6]]},"reference":[{"key":"e_1_3_4_2_2","unstructured":"TensorFlow. n.d. TensorFlow Lite. Retrieved March 2 2022 from https:\/\/github.com\/tensorflow\/tflite-support."},{"key":"e_1_3_4_3_2","unstructured":"GitHub. n.d. alibaba\/MNN. Retrieved March 2 2022 from https:\/\/github.com\/alibaba\/MNN."},{"key":"e_1_3_4_4_2","unstructured":"PyTorch. n.d. PyTorch Mobile. Retrieved March 2 2022 from https:\/\/pytorch.org\/mobile\/home."},{"key":"e_1_3_4_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/2858788.2688521"},{"key":"e_1_3_4_6_2","doi-asserted-by":"publisher","DOI":"10.1137\/141000671"},{"key":"e_1_3_4_7_2","article-title":"Yolov4: Optimal speed and accuracy of object detection","author":"Bochkovskiy Alexey","year":"2020","unstructured":"Alexey Bochkovskiy, Chien-Yao Wang, and Hong-Yuan Mark Liao. 2020. Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934 (2020).","journal-title":"arXiv preprint arXiv:2004.10934"},{"key":"e_1_3_4_8_2","article-title":"On optimizing operator fusion plans for large-scale machine learning in SystemML","author":"Boehm Matthias","year":"2018","unstructured":"Matthias Boehm, Berthold Reinwald, Dylan Hutchison, Alexandre V. Evfimievski, and Prithviraj Sen. 2018. On optimizing operator fusion plans for large-scale machine learning in SystemML. arXiv preprint arXiv:1801.00829 (2018).","journal-title":"arXiv preprint arXiv:1801.00829"},{"key":"e_1_3_4_9_2","article-title":"ProxylessNAS: Direct neural architecture search on target task and hardware","author":"Cai Han","year":"2018","unstructured":"Han Cai, Ligeng Zhu, and Song Han. 2018. ProxylessNAS: Direct neural architecture search on target task and hardware. arXiv preprint arXiv:1812.00332 (2018).","journal-title":"arXiv preprint arXiv:1812.00332"},{"key":"e_1_3_4_10_2","first-page":"955","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Cai Yuxuan","year":"2021","unstructured":"Yuxuan Cai, Hongjia Li, Geng Yuan, Wei Niu, Yanyu Li, Xulong Tang, Bin Ren, and Yanzhi Wang. 2021. YOLObile: Real-time object detection on mobile devices via compression-compilation co-design. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 955\u2013963."},{"key":"e_1_3_4_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00041-008-9045-x"},{"key":"e_1_3_4_12_2","volume-title":"Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201918)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, et\u00a0al. 2018. TVM: An automated end-to-end optimizing compiler for deep learning. In Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201918)."},{"key":"e_1_3_4_13_2","article-title":"Bayesian optimization in Alphago","author":"Chen Yutian","year":"2018","unstructured":"Yutian Chen, Aja Huang, Ziyu Wang, Ioannis Antonoglou, Julian Schrittwieser, David Silver, and Nando de Freitas. 2018. Bayesian optimization in Alphago. arXiv preprint arXiv:1812.06855 (2018).","journal-title":"arXiv preprint arXiv:1812.06855"},{"key":"e_1_3_4_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.195"},{"key":"e_1_3_4_15_2"},{"key":"e_1_3_4_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2019.2954495"},{"key":"e_1_3_4_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2019.2914438"},{"key":"e_1_3_4_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_4_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18072.2020.9218499"},{"key":"e_1_3_4_20_2","first-page":"759","volume-title":"Proceedings of the 2019 Conference on Neural Information Processing Systems (NeurIPS\u201919)","author":"Dong Xuanyi","year":"2019","unstructured":"Xuanyi Dong and Yi Yang. 2019. Network pruning via transformable architecture search. In Proceedings of the 2019 Conference on Neural Information Processing Systems (NeurIPS\u201919). 759\u2013770."},{"key":"e_1_3_4_21_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201918)","author":"Frankle Jonathan","year":"2018","unstructured":"Jonathan Frankle and Michael Carbin. 2018. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In Proceedings of the International Conference on Learning Representations (ICLR\u201918)."},{"key":"e_1_3_4_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3386263.3407650"},{"key":"e_1_3_4_23_2","volume-title":"Proceedings of the 2016 Conference on Neural Information Processing Systems (NeurIPS\u201916)","author":"Guo Yiwen","year":"2016","unstructured":"Yiwen Guo, Anbang Yao, and Yurong Chen. 2016. Dynamic network surgery for efficient DNNs. In Proceedings of the 2016 Conference on Neural Information Processing Systems (NeurIPS\u201916)."},{"key":"e_1_3_4_24_2","volume-title":"Proceedings of the 2015 Conference on Neural Information Processing Systems (NeurIPS\u201915)","author":"Han Song","year":"2015","unstructured":"Song Han, Jeff Pool, John Tran, and William Dally. 2015. Learning both weights and connections for efficient neural network. In Proceedings of the 2015 Conference on Neural Information Processing Systems (NeurIPS\u201915)."},{"key":"e_1_3_4_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/2906388.2906396"},{"key":"e_1_3_4_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_4_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_48"},{"key":"e_1_3_4_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00447"},{"key":"e_1_3_4_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.155"},{"key":"e_1_3_4_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/2968455.2968511"},{"key":"e_1_3_4_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3081333.3081360"},{"key":"e_1_3_4_32_2","article-title":"Radio frequency fingerprinting on the edge","author":"Jian Tong","year":"2022","unstructured":"Tong Jian, Yifan Gong, Zheng Zhan, Runbin Shi, Nasim Soltani, Zifeng Wang, Jennifer G. Dy, Kaushik Roy Chowdhury, Yanzhi Wang, and Stratis Ioannidis. 2022. Radio frequency fingerprinting on the edge. IEEE Transactions on Mobile Computing. Early access, March 8, 2022.","journal-title":"IEEE Transactions on Mobile Computing."},{"key":"e_1_3_4_33_2","first-page":"528","volume-title":"Proceedings of the 20th International Conference on Artificial Intelligence and Statistics","author":"Klein Aaron","year":"2017","unstructured":"Aaron Klein, Stefan Falkner, Simon Bartels, Philipp Hennig, and Frank Hutter. 2017. Fast Bayesian optimization of machine learning hyperparameters on large datasets. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics. 528\u2013536."},{"key":"e_1_3_4_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPSN.2016.7460664"},{"key":"e_1_3_4_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/MPRV.2017.2940968"},{"key":"e_1_3_4_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/2750858.2804262"},{"key":"e_1_3_4_37_2","article-title":"Pruning filters for efficient ConvNets","author":"Li Hao","year":"2017","unstructured":"Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. 2017. Pruning filters for efficient ConvNets. In Proceedings of the International Conference on Learning Representations (ICLR\u201917).","journal-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201917)."},{"key":"e_1_3_4_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394885.3431627"},{"key":"e_1_3_4_39_2","article-title":"Learning to optimize","author":"Li Ke","year":"2016","unstructured":"Ke Li and Jitendra Malik. 2016. Learning to optimize. arXiv preprint arXiv:1606.01885 (2016).","journal-title":"arXiv preprint arXiv:1606.01885"},{"key":"e_1_3_4_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00410"},{"key":"e_1_3_4_41_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201920)","author":"Li Yuhang","year":"2020","unstructured":"Yuhang Li, Xin Dong, and Wei Wang. 2020. Additive powers-of-two quantization: An efficient non-uniform discretization for neural networks. In Proceedings of the International Conference on Learning Representations (ICLR\u201920)."},{"key":"e_1_3_4_42_2","article-title":"SS-Auto: A single-shot, automatic structured weight pruning framework of DNNs with ultra-high efficiency","author":"Li Zhengang","year":"2020","unstructured":"Zhengang Li, Yifan Gong, Xiaolong Ma, Sijia Liu, Mengshu Sun, Zheng Zhan, Zhenglun Kong, Geng Yuan, and Yanzhi Wang. 2020. SS-Auto: A single-shot, automatic structured weight pruning framework of DNNs with ultra-high efficiency. arXiv preprint arXiv:2001.08839 (2020).","journal-title":"arXiv preprint arXiv:2001.08839"},{"key":"e_1_3_4_43_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_4_44_2","article-title":"AutoCompress: An automatic DNN structured pruning framework for ultra-high compression rates","author":"Liu Ning","year":"2019","unstructured":"Ning Liu, Xiaolong Ma, Zhiyuan Xu, Yanzhi Wang, Jian Tang, and Jieping Ye. 2019. AutoCompress: An automatic DNN structured pruning framework for ultra-high compression rates. arXiv preprint arXiv:1907.03141 (2019).","journal-title":"arXiv preprint arXiv:1907.03141"},{"key":"e_1_3_4_45_2","article-title":"AutoSlim: An automatic DNN structured pruning framework for ultra-high compression rates","author":"Liu Ning","year":"2019","unstructured":"Ning Liu, Xiaolong Ma, Zhiyuan Xu, Yanzhi Wang, Jian Tang, and Jieping Ye. 2019. AutoSlim: An automatic DNN structured pruning framework for ultra-high compression rates. arXiv preprint arXiv:1907.03141 (2019).","journal-title":"arXiv preprint arXiv:1907.03141"},{"key":"e_1_3_4_46_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5924"},{"key":"e_1_3_4_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.298"},{"key":"e_1_3_4_48_2"},{"key":"e_1_3_4_49_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201918)","author":"Liu Zhuang","year":"2018","unstructured":"Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell. 2018. Rethinking the value of network pruning. In Proceedings of the International Conference on Learning Representations (ICLR\u201918)."},{"key":"e_1_3_4_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.541"},{"key":"e_1_3_4_51_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5954"},{"key":"e_1_3_4_52_2","article-title":"BLK-REW: A unified block-based DNN pruning framework using reweighted regularization method","author":"Ma Xiaolong","year":"2020","unstructured":"Xiaolong Ma, Zhengang Li, Yifan Gong, Tianyun Zhang, Wei Niu, Zheng Zhan, Pu Zhao, et\u00a0al. 2020. BLK-REW: A unified block-based DNN pruning framework using reweighted regularization method. arXiv preprint arXiv:2001.08357 (2020).","journal-title":"arXiv preprint arXiv:2001.08357"},{"key":"e_1_3_4_53_2","unstructured":"Xiaolong Ma Sheng Lin Shaokai Ye Zhezhi He Linfeng Zhang Geng Yuan Sia Huat Tan et\u00a0al. 2019. Non-structured DNN weight pruning\u2014Is it beneficial in any platform? arxiv:1907.02124 [cs.LG] (2019)."},{"key":"e_1_3_4_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58601-0_37"},{"key":"e_1_3_4_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASP-DAC47756.2020.9045658"},{"key":"e_1_3_4_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASP-DAC47756.2020.9045658"},{"key":"e_1_3_4_57_2","first-page":"1","volume-title":"Proceedings of the 2019 IEEE\/ACM International Symposium on Nanoscale Architectures (NANOARCH\u201919)","author":"Ma Xiaolong","year":"2019","unstructured":"Xiaolong Ma, Geng Yuan, Sheng Lin, Zhengang Li, Hao Sun, and Yanzhi Wang. 2019. ResNet can be pruned 60 \\( \\times \\) : Introducing network purification and unused path removal (P-RM) after weight pruning. In Proceedings of the 2019 IEEE\/ACM International Symposium on Nanoscale Architectures (NANOARCH\u201919). IEEE, Los Alamitos, CA, 1\u20132."},{"key":"e_1_3_4_58_2","article-title":"2PFPCE: Two-phase filter pruning based on conditional entropy","author":"Min Chuhan","year":"2018","unstructured":"Chuhan Min, Aosen Wang, Yiran Chen, Wenyao Xu, and Xin Chen. 2018. 2PFPCE: Two-phase filter pruning based on conditional entropy. arXiv preprint arXiv:1809.02220 (2018).","journal-title":"arXiv preprint arXiv:1809.02220"},{"key":"e_1_3_4_59_2","article-title":"Achieving real-time execution of transformer-based large-scale models on mobile with compiler-aware neural architecture optimization","author":"Niu Wei","year":"2020","unstructured":"Wei Niu, Zhenglun Kong, Geng Yuan, Weiwen Jiang, Jiexiong Guan, Caiwen Ding, Pu Zhao, Sijia Liu, Bin Ren, and Yanzhi Wang. 2020. Achieving real-time execution of transformer-based large-scale models on mobile with compiler-aware neural architecture optimization. arXiv preprint arXiv:2009.06823 (2020).","journal-title":"arXiv preprint arXiv:2009.06823"},{"key":"e_1_3_4_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378534"},{"issue":"3","key":"e_1_3_4_61_2","first-page":"1","article-title":"Deep learning for mobile multimedia: A survey","volume":"13","author":"Ota Kaoru","year":"2017","unstructured":"Kaoru Ota, Minh Son Dao, Vasileios Mezaris, and Francesco G. B. De Natale. 2017. Deep learning for mobile multimedia: A survey. ACM Transactions on Multimedia Computing, Communications, and Applications 13, 3s (2017), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_4_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3297858.3304076"},{"key":"e_1_3_4_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3297858.3304076"},{"key":"e_1_3_4_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_4_65_2","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556 (2014).","journal-title":"arXiv:1409.1556"},{"key":"e_1_3_4_66_2","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton Richard S.","year":"2018","unstructured":"Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. MIT Press, Cambridge, MA."},{"key":"e_1_3_4_67_2","doi-asserted-by":"publisher","DOI":"10.5555\/3009657.3009806"},{"key":"e_1_3_4_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00293"},{"key":"e_1_3_4_69_2","article-title":"Non-structured DNN weight pruning considered harmful","author":"Wang Yanzhi","year":"2019","unstructured":"Yanzhi Wang, Shaokai Ye, Zhezhi He, Xiaolong Ma, Linfeng Zhang, Sheng Lin, Geng Yuan, et\u00a0al. 2019. Non-structured DNN weight pruning considered harmful. arXiv:1907.02124 (2019).","journal-title":"arXiv:1907.02124"},{"key":"e_1_3_4_70_2","volume-title":"Proceedings of the 2016 Conference on Neural Information Processing Systems (NeurIPS\u201916)","author":"Wen Wei","year":"2016","unstructured":"Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2016. Learning structured sparsity in deep neural networks. In Proceedings of the 2016 Conference on Neural Information Processing Systems (NeurIPS\u201916)."},{"key":"e_1_3_4_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01099"},{"key":"e_1_3_4_72_2","doi-asserted-by":"publisher","DOI":"10.1145\/3241539.3241563"},{"key":"e_1_3_4_73_2"},{"key":"e_1_3_4_74_2","doi-asserted-by":"publisher","DOI":"10.1145\/3038912.3052577"},{"key":"e_1_3_4_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00958"},{"key":"e_1_3_4_76_2"},{"key":"e_1_3_4_77_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISLPED.2019.8824944"},{"key":"e_1_3_4_78_2","article-title":"A SOT-MRAM-based processing-in-memory engine for highly compressed DNN implementation","author":"Yuan Geng","year":"2019","unstructured":"Geng Yuan, Xiaolong Ma, Sheng Lin, Zhengang Li, and Caiwen Ding. 2019. A SOT-MRAM-based processing-in-memory engine for highly compressed DNN implementation. arXiv preprint arXiv:1912.05416 (2019).","journal-title":"arXiv preprint arXiv:1912.05416"},{"key":"e_1_3_4_79_2","unstructured":"Geng Yuan Xiaolong Ma Wei Niu Zhengang Li Zhenglun Kong Ning Liu Yifan Gong et\u00a0al. 2021. MEST: Accurate and fast memory-economic sparse training framework on the edge. arxiv:2110.14032 [cs.LG] (2021)."},{"key":"e_1_3_4_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00478"},{"key":"e_1_3_4_81_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMST.2019.2904897"},{"key":"e_1_3_4_82_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01237-3_12"},{"key":"e_1_3_4_83_2","article-title":"Adam-ADMM: A unified, systematic framework of structured weight pruning for DNNs","author":"Zhang Tianyun","year":"2018","unstructured":"Tianyun Zhang, Kaiqi Zhang, Shaokai Ye, Jiayu Li, Jian Tang, Wujie Wen, Xue Lin, Makan Fardad, and Yanzhi Wang. 2018. Adam-ADMM: A unified, systematic framework of structured weight pruning for DNNs. arXiv:1807.11091 (2018).","journal-title":"arXiv:1807.11091"},{"key":"e_1_3_4_84_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00289"},{"key":"e_1_3_4_85_2","article-title":"Achieving real-time LiDAR 3D object detection on a mobile device","author":"Zhao Pu","year":"2020","unstructured":"Pu Zhao, Wei Niu, Geng Yuan, Yuxuan Cai, Hsin-Hsuan Sung, Wujie Wen, Sijia Liu, et\u00a0al. 2020. Achieving real-time LiDAR 3D object detection on a mobile device. arXiv preprint arXiv:2012.13801 (2020).","journal-title":"arXiv preprint arXiv:2012.13801"},{"key":"e_1_3_4_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00257"},{"key":"e_1_3_4_87_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2018\/453"},{"key":"e_1_3_4_88_2","volume-title":"Proceedings of the 2018 Conference on Neural Information Processing Systems (NeurIPS\u201918)","author":"Zhuang Zhuangwei","year":"2018","unstructured":"Zhuangwei Zhuang, Mingkui Tan, Bohan Zhuang, Jing Liu, Yong Guo, Qingyao Wu, Junzhou Huang, and Jinhui Zhu. 2018. Discrimination-aware channel pruning for deep neural networks. In Proceedings of the 2018 Conference on Neural Information Processing Systems (NeurIPS\u201918)."},{"key":"e_1_3_4_89_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201917)","author":"Zoph Barret","year":"2017","unstructured":"Barret Zoph and Quoc V. Le. 2017. Neural architecture search with reinforcement learning. In Proceedings of the International Conference on Learning Representations (ICLR\u201917)."}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3495532","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3495532","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:23Z","timestamp":1750182563000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3495532"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,6]]},"references-count":88,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2022,9,30]]}},"alternative-id":["10.1145\/3495532"],"URL":"https:\/\/doi.org\/10.1145\/3495532","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,6]]},"assertion":[{"value":"2021-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-06-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}