{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,26]],"date-time":"2025-12-26T07:08:57Z","timestamp":1766732937024,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":41,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,6,3]],"date-time":"2021-06-03T00:00:00Z","timestamp":1622678400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"DARPA","award":["FA8650-20-2-7007"],"award-info":[{"award-number":["FA8650-20-2-7007"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,6,3]]},"DOI":"10.1145\/3447818.3460378","type":"proceedings-article","created":{"date-parts":[[2021,6,4]],"date-time":"2021-06-04T15:09:36Z","timestamp":1622819376000},"page":"291-303","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Accelerating DNNs inference with predictive layer fusion"],"prefix":"10.1145","author":[{"given":"MohammadHossein","family":"Olyaiy","sequence":"first","affiliation":[{"name":"The University of British Columbia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christopher","family":"Ng","sequence":"additional","affiliation":[{"name":"The University of British Columbia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mieszko","family":"Lis","sequence":"additional","affiliation":[{"name":"The University of British Columbia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,6,4]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafa\u0142 J\u00f3zefowicz \u0141ukasz Kaiser Manjunath Kudlur Josh Levenberg Dan Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. http:\/\/tensorflow.org\/ Software available from tensorflow.org.  Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafa\u0142 J\u00f3zefowicz \u0141ukasz Kaiser Manjunath Kudlur Josh Levenberg Dan Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. http:\/\/tensorflow.org\/ Software available from tensorflow.org."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00061"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001138"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783725"},{"key":"e_1_3_2_1_5_1","volume-title":"Proc. Int. Conf. High Perform. Comput., Netw., Storage, Anal.(SC).","author":"Anderson Michael","year":"2018","unstructured":"Michael Anderson , Evangelos Georganas , Sasikanth Avancha , and Alexander Heinecke . 2018 . Tensorfolding: Improving convolutional neural network performance with fused microkernels . In Proc. Int. Conf. High Perform. Comput., Netw., Storage, Anal.(SC). Michael Anderson, Evangelos Georganas, Sasikanth Avancha, and Alexander Heinecke. 2018. Tensorfolding: Improving convolutional neural network performance with fused microkernels. In Proc. Int. Conf. High Perform. Comput., Netw., Storage, Anal.(SC)."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1137\/060676489"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICTC.2018.8539530"},{"key":"e_1_3_2_1_8_1","unstructured":"Tianqi Chen Thierry Moreau Ziheng Jiang Lianmin Zheng Eddie Yan Haichen Shen Meghan Cowan Leyuan Wang Yuwei Hu Luis Ceze etal 2018. {TVM}: An automated end-to-end optimizing compiler for deep learning. In 13th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 18). 578--594.  Tianqi Chen Thierry Moreau Ziheng Jiang Lianmin Zheng Eddie Yan Haichen Shen Meghan Cowan Leyuan Wang Yuwei Hu Luis Ceze et al. 2018. {TVM}: An automated end-to-end optimizing compiler for deep learning. In 13th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 18) . 578--594."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2016.2616357"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750389"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3297858.3304014"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358291"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.30"},{"key":"e_1_3_2_1_14_1","volume-title":"Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. CoRR abs\/1510.00149","author":"Han Song","year":"2016","unstructured":"Song Han , Huizi Mao , and W. Dally . 2016 b. Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. CoRR abs\/1510.00149 (2016). Song Han, Huizi Mao, and W. Dally. 2016b. Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. CoRR abs\/1510.00149 (2016)."},{"key":"e_1_3_2_1_15_1","volume-title":"Proceedings of the 28th International Conference on Neural Information Processing Systems-Volume 1. 1135--1143","author":"Han Song","year":"2015","unstructured":"Song Han , Jeff Pool , John Tran , and William J Dally . 2015 . Learning both weights and connections for efficient neural networks . In Proceedings of the 28th International Conference on Neural Information Processing Systems-Volume 1. 1135--1143 . Song Han, Jeff Pool, John Tran, and William J Dally. 2015. Learning both weights and connections for efficient neural networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems-Volume 1. 1135--1143."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_3_2_1_18_1","volume-title":"International conference on machine learning. PMLR, 448--456","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy . 2015 . Batch normalization: Accelerating deep network training by reducing internal covariate shift . In International conference on machine learning. PMLR, 448--456 . Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning. PMLR, 448--456."},{"volume-title":"2017 ACM\/IEEE 44th Annual International Symposium on Computer Architecture (ISCA). 1--12","author":"Jouppi N. P.","key":"e_1_3_2_1_19_1","unstructured":"N. P. Jouppi , C. Young , N. Patil , D. Patterson , G. Agrawal , R. Bajwa , S. Bates , S. Bhatia , N. Boden , A. Borchers , R. Boyle , P. Cantin , C. Chao , C. Clark , J. Coriell , M. Daley , M. Dau , J. Dean , B. Gelb , T. V. Ghaemmaghami , R. Gottipati , W. Gulland , R. Hagmann , C. R. Ho , D. Hogberg , J. Hu , R. Hundt , D. Hurt , J. Ibarz , A. Jaffey , A. Jaworski , A. Kaplan , H. Khaitan , D. Killebrew , A. Koch , N. Kumar , S. Lacy , J. Laudon , J. Law , D. Le , C. Leary , Z. Liu , K. Lucke , A. Lundin , G. MacKean , A. Maggiore , M. Mahony , K. Miller , R. Nagarajan , R. Narayanaswami , R. Ni , K. Nix , T. Norrie , M. Omernick , N. Penukonda , A. Phelps , J. Ross , M. Ross , A. Salek , E. Samadiani , C. Severn , G. Sizikov , M. Snelham , J. Souter , D. Steinberg , A. Swing , M. Tan , G. Thorson , B. Tian , H. Toma , E. Tuttle , V. Vasudevan , R. Walter , W. Wang , E. Wilcox , and D. H. Yoon . 2017. In-datacenter performance analysis of a tensor processing unit . In 2017 ACM\/IEEE 44th Annual International Symposium on Computer Architecture (ISCA). 1--12 . N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon. 2017. In-datacenter performance analysis of a tensor processing unit. In 2017 ACM\/IEEE 44th Annual International Symposium on Computer Architecture (ISCA). 1--12."},{"key":"e_1_3_2_1_20_1","unstructured":"Wonkyung Jung Daejin Jung Byeongho Kim Sunjung Lee Wonjong Rhee and Jung Ho Ahn. 2019. Restructuring Batch Normalization to Accelerate CNN Training. [arxiv]1807.01702 [cs.CV]  Wonkyung Jung Daejin Jung Byeongho Kim Sunjung Lee Wonjong Rhee and Jung Ho Ahn. 2019. Restructuring Batch Normalization to Accelerate CNN Training. [arxiv]1807.01702 [cs.CV]"},{"key":"e_1_3_2_1_22_1","unstructured":"Ching-En Lee Yakun Sophia Shao Jie-Fang Zhang Angshuman Parashar Joel Emer Stephen W Keckler and Zhengya Zhang. [n.d.]. Stitch-x: An accelerator architecture for exploiting unstructured sparsity in deep neural networks.  Ching-En Lee Yakun Sophia Shao Jie-Fang Zhang Angshuman Parashar Joel Emer Stephen W Keckler and Zhengya Zhang. [n.d.]. Stitch-x: An accelerator architecture for exploiting unstructured sparsity in deep neural networks."},{"key":"e_1_3_2_1_23_1","volume-title":"Ternary Weight Networks. CoRR abs\/1605.04711","author":"Li Fengfu","year":"2016","unstructured":"Fengfu Li and Bin Liu . 2016. Ternary Weight Networks. CoRR abs\/1605.04711 ( 2016 ). Fengfu Li and Bin Liu. 2016. Ternary Weight Networks. CoRR abs\/1605.04711 (2016)."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00290"},{"key":"e_1_3_2_1_25_1","volume-title":"DUET: Boosting Deep Neural Network Efficiency on Dual-Module Architecture. In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 738--750","author":"Liu Liu","year":"2020","unstructured":"Liu Liu , Zheng Qu , Lei Deng , Fengbin Tu , Shuangchen Li , Xing Hu , Zhenyu Gu , Yufei Ding , and Yuan Xie . 2020 . DUET: Boosting Deep Neural Network Efficiency on Dual-Module Architecture. In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 738--750 . Liu Liu, Zheng Qu, Lei Deng, Fengbin Tu, Shuangchen Li, Xing Hu, Zhenyu Gu, Yufei Ding, and Yuan Xie. 2020. DUET: Boosting Deep Neural Network Efficiency on Dual-Module Architecture. In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 738--750."},{"key":"e_1_3_2_1_26_1","volume-title":"SGDR: Stochastic Gradient Descent with Warm Restarts. arXiv: Learning","author":"Loshchilov Ilya","year":"2017","unstructured":"Ilya Loshchilov and Frank Hutter . 2017 . SGDR: Stochastic Gradient Descent with Warm Restarts. arXiv: Learning (2017). Ilya Loshchilov and Frank Hutter. 2017. SGDR: Stochastic Gradient Descent with Warm Restarts. arXiv: Learning (2017)."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356156"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3330345.3330385"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/3104322.3104425"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00042"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080254"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2014.6853196"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58607-2_14"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/MWSCAS48704.2020.9184599"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00068"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00293"},{"key":"e_1_3_2_1_39_1","volume-title":"Accelergy: An Architecture-Level Energy Estimation Methodology for Accelerator Designs. In 2019 IEEE\/ACM International Conference on Computer-Aided Design (ICCAD). 1--8.","author":"Wu Yannan N.","year":"2019","unstructured":"Yannan N. Wu , Joel S. Emer , and Vivienne Sze . 2019 . Accelergy: An Architecture-Level Energy Estimation Methodology for Accelerator Designs. In 2019 IEEE\/ACM International Conference on Computer-Aided Design (ICCAD). 1--8. Yannan N. Wu, Joel S. Emer, and Vivienne Sze. 2019. Accelergy: An Architecture-Level Energy Estimation Methodology for Accelerator Designs. In 2019 IEEE\/ACM International Conference on Computer-Aided Design (ICCAD). 1--8."},{"key":"e_1_3_2_1_40_1","unstructured":"Shunzhi Yang Zheng Gong Kai Ye Yungen Wei Zheng Huang and Zhenhua Huang. 2019. EdgeCNN: Convolutional Neural Network Classification Model with small inputs for Edge Computing. [arxiv]1909.13522 [cs.CV]  Shunzhi Yang Zheng Gong Kai Ye Yungen Wei Zheng Huang and Zhenhua Huang. 2019. EdgeCNN: Convolutional Neural Network Classification Model with small inputs for Edge Computing. [arxiv]1909.13522 [cs.CV]"},{"key":"e_1_3_2_1_41_1","volume-title":"Gate decorator: Global filter pruning method for accelerating deep convolutional neural networks. arXiv preprint arXiv:1909.08174","author":"You Zhonghui","year":"2019","unstructured":"Zhonghui You , Kun Yan , Jinmian Ye , Meng Ma , and Ping Wang . 2019. Gate decorator: Global filter pruning method for accelerating deep convolutional neural networks. arXiv preprint arXiv:1909.08174 ( 2019 ). Zhonghui You, Kun Yan, Jinmian Ye, Meng Ma, and Ping Wang. 2019. Gate decorator: Global filter pruning method for accelerating deep convolutional neural networks. arXiv preprint arXiv:1909.08174 (2019)."},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00716"},{"key":"e_1_3_2_1_43_1","unstructured":"Michael Zhu and Suyog Gupta. 2017. To prune or not to prune: exploring the efficacy of pruning for model compression. [arxiv]1710.01878 [stat.ML]  Michael Zhu and Suyog Gupta. 2017. To prune or not to prune: exploring the efficacy of pruning for model compression. [arxiv]1710.01878 [stat.ML]"}],"event":{"name":"ICS '21: 2021 International Conference on Supercomputing","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture"],"location":"Virtual Event USA","acronym":"ICS '21"},"container-title":["Proceedings of the ACM International Conference on Supercomputing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447818.3460378","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3447818.3460378","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3447818.3460378","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T17:49:27Z","timestamp":1750268967000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447818.3460378"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,6,3]]},"references-count":41,"alternative-id":["10.1145\/3447818.3460378","10.1145\/3447818"],"URL":"https:\/\/doi.org\/10.1145\/3447818.3460378","relation":{},"subject":[],"published":{"date-parts":[[2021,6,3]]},"assertion":[{"value":"2021-06-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}