{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T05:02:49Z","timestamp":1783141369484,"version":"3.54.6"},"publisher-location":"New York, NY, USA","reference-count":32,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,4,25]],"date-time":"2022-04-25T00:00:00Z","timestamp":1650844800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"NSF (National Science Foundation)","doi-asserted-by":"publisher","award":["2011146"],"award-info":[{"award-number":["2011146"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,4,25]]},"DOI":"10.1145\/3487553.3524856","type":"proceedings-article","created":{"date-parts":[[2022,8,16]],"date-time":"2022-08-16T22:41:30Z","timestamp":1660689690000},"page":"548-554","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Optimizing Data Layout for Training Deep Neural Networks"],"prefix":"10.1145","author":[{"given":"Bingyao","family":"Li","sequence":"first","affiliation":[{"name":"University of Pittsburgh, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qi","family":"Xue","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Geng","family":"Yuan","sequence":"additional","affiliation":[{"name":"Northeastern University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sheng","family":"Li","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaolong","family":"Ma","sequence":"additional","affiliation":[{"name":"Northeastern University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanzhi","family":"Wang","sequence":"additional","affiliation":[{"name":"Northeastern University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xulong","family":"Tang","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,8,16]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Tensorflow: A system for large-scale machine learning. In OSDI.","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi , Paul Barham , Jianmin Chen , Zhifeng Chen , Andy Davis , Jeffrey Dean , Matthieu Devin , Sanjay Ghemawat , Geoffrey Irving , Michael Isard , 2016 . Tensorflow: A system for large-scale machine learning. In OSDI. Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, 2016. Tensorflow: A system for large-scale machine learning. In OSDI."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2910232"},{"key":"e_1_3_2_1_3_1","unstructured":"Yoojin Choi Mostafa El-Khamy and Jungwon Lee. 2016. Towards the limit of network quantization. arXiv preprint arXiv:1612.01543(2016).  Yoojin Choi Mostafa El-Khamy and Jungwon Lee. 2016. Towards the limit of network quantization. arXiv preprint arXiv:1612.01543(2016)."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2019.2954495"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2019.2914438"},{"key":"e_1_3_2_1_6_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805(2018).","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805(2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805(2018)."},{"key":"e_1_3_2_1_7_1","volume-title":"The International Conference on Learning Representations (ICLR).","author":"Frankle Jonathan","year":"2018","unstructured":"Jonathan Frankle and Michael Carbin . 2018 . The lottery ticket hypothesis: Finding sparse, trainable neural networks . In The International Conference on Learning Representations (ICLR). Jonathan Frankle and Michael Carbin. 2018. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In The International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_1_8_1","unstructured":"Priya Goyal Piotr Doll\u00e1r Ross Girshick Pieter Noordhuis Lukasz Wesolowski Aapo Kyrola Andrew Tulloch Yangqing Jia and Kaiming He. 2017. Accurate large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677(2017).  Priya Goyal Piotr Doll\u00e1r Ross Girshick Pieter Noordhuis Lukasz Wesolowski Aapo Kyrola Andrew Tulloch Yangqing Jia and Kaiming He. 2017. Accurate large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677(2017)."},{"key":"e_1_3_2_1_9_1","unstructured":"Yiwen Guo Anbang Yao and Yurong Chen. 2016. Dynamic network surgery for efficient dnns. In Advances in Neural Information Processing Systems (NeurIPS).  Yiwen Guo Anbang Yao and Yurong Chen. 2016. Dynamic network surgery for efficient dnns. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"crossref","unstructured":"Song Han Xingyu Liu Huizi Mao Jing Pu Ardavan Pedram Mark\u00a0A Horowitz and William\u00a0J Dally. 2016. EIE: efficient inference engine on compressed deep neural network. In ISCA.  Song Han Xingyu Liu Huizi Mao Jing Pu Ardavan Pedram Mark\u00a0A Horowitz and William\u00a0J Dally. 2016. EIE: efficient inference engine on compressed deep neural network. In ISCA.","DOI":"10.1109\/HOTCHIPS.2016.7936226"},{"key":"e_1_3_2_1_11_1","unstructured":"Yang He Ping Liu Ziwei Wang Zhilan Hu and Yi Yang. 2019. Filter pruning via geometric median for deep convolutional neural networks acceleration. In CVPR.  Yang He Ping Liu Ziwei Wang Zhilan Hu and Yi Yang. 2019. Filter pruning via geometric median for deep convolutional neural networks acceleration. In CVPR."},{"key":"e_1_3_2_1_12_1","unstructured":"Yihui He Xiangyu Zhang and Jian Sun. 2017. Channel pruning for accelerating very deep neural networks. In ICCV.  Yihui He Xiangyu Zhang and Jian Sun. 2017. Channel pruning for accelerating very deep neural networks. In ICCV."},{"key":"e_1_3_2_1_13_1","unstructured":"Zhihao Jia Matei Zaharia and Alex Aiken. 2019. Beyond data and model parallelism for deep neural networks. In SysML.  Zhihao Jia Matei Zaharia and Alex Aiken. 2019. Beyond data and model parallelism for deep neural networks. In SysML."},{"key":"e_1_3_2_1_14_1","unstructured":"Alex Krizhevsky. 2014. One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv:1404.5997(2014).  Alex Krizhevsky. 2014. One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv:1404.5997(2014)."},{"key":"e_1_3_2_1_15_1","unstructured":"Alex Krizhevsky Geoffrey Hinton 2009. Learning multiple layers of features from tiny images. (2009).  Alex Krizhevsky Geoffrey Hinton 2009. Learning multiple layers of features from tiny images. (2009)."},{"key":"e_1_3_2_1_16_1","unstructured":"Sameer Kumar James Bradbury Cliff Young Yu\u00a0Emma Wang Anselm Levskaya Blake Hechtman Dehao Chen HyoukJoong Lee Mehmet Deveci Naveen Kumar 2020. Exploring the limits of Concurrency in ML Training on Google TPUs. arXiv preprint arXiv:2011.03641(2020).  Sameer Kumar James Bradbury Cliff Young Yu\u00a0Emma Wang Anselm Levskaya Blake Hechtman Dehao Chen HyoukJoong Lee Mehmet Deveci Naveen Kumar 2020. Exploring the limits of Concurrency in ML Training on Google TPUs. arXiv preprint arXiv:2011.03641(2020)."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Jinho Lee Inseok Hwang Soham Shah and Minsik Cho. 2020. FlexReduce: Flexible All-reduce for Distributed Deep Learning on Asymmetric Network Topology. In DAC.  Jinho Lee Inseok Hwang Soham Shah and Minsik Cho. 2020. FlexReduce: Flexible All-reduce for Distributed Deep Learning on Asymmetric Network Topology. In DAC.","DOI":"10.1109\/DAC18072.2020.9218538"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2741948.2741965"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Mu Li David\u00a0G Andersen Jun\u00a0Woo Park Alexander\u00a0J Smola Amr Ahmed Vanja Josifovski James Long Eugene\u00a0J Shekita and Bor-Yiing Su. 2014. Scaling distributed machine learning with the parameter server. In OSDI.  Mu Li David\u00a0G Andersen Jun\u00a0Woo Park Alexander\u00a0J Smola Amr Ahmed Vanja Josifovski James Long Eugene\u00a0J Shekita and Bor-Yiing Su. 2014. Scaling distributed machine learning with the parameter server. In OSDI.","DOI":"10.1145\/2640087.2644155"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"crossref","unstructured":"Ning Liu Xiaolong Ma Zhiyuan Xu Yanzhi Wang Jian Tang and Jieping Ye. 2020. AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates. In AAAI.  Ning Liu Xiaolong Ma Zhiyuan Xu Yanzhi Wang Jian Tang and Jieping Ye. 2020. AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates. In AAAI.","DOI":"10.1609\/aaai.v34i04.5924"},{"key":"e_1_3_2_1_21_1","unstructured":"Shaoshan Liu Bin Ren Xipeng Shen and Yanzhi Wang. 2020. CocoPIE: Making Mobile AI Sweet As PIE\u2013Compression-Compilation Co-Design Goes a Long Way. arXiv preprint arXiv:2003.06700(2020).  Shaoshan Liu Bin Ren Xipeng Shen and Yanzhi Wang. 2020. CocoPIE: Making Mobile AI Sweet As PIE\u2013Compression-Compilation Co-Design Goes a Long Way. arXiv preprint arXiv:2003.06700(2020)."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5954"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359646"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378534"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2019.2935967"},{"key":"e_1_3_2_1_26_1","volume-title":"Zero: Memory optimizations toward training trillion parameter models. In SC20.","author":"Rajbhandari Samyam","year":"2020","unstructured":"Samyam Rajbhandari , Jeff Rasley , Olatunji Ruwase , and Yuxiong He . 2020 . Zero: Memory optimizations toward training trillion parameter models. In SC20. Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. Zero: Memory optimizations toward training trillion parameter models. In SC20."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3297858.3304076"},{"key":"e_1_3_2_1_28_1","unstructured":"Wei Wen Chunpeng Wu Yandan Wang Yiran Chen and Hai Li. 2016. Learning structured sparsity in deep neural networks. In NeurIPS.  Wei Wen Chunpeng Wu Yandan Wang Yiran Chen and Hai Li. 2016. Learning structured sparsity in deep neural networks. In NeurIPS."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Yang You and James Demmel. 2017. Runtime data layout scheduling for machine learning dataset. In ICPP.  Yang You and James Demmel. 2017. Runtime data layout scheduling for machine learning dataset. In ICPP.","DOI":"10.1109\/ICPP.2017.54"},{"key":"e_1_3_2_1_30_1","volume-title":"Nisp: Pruning networks using neuron importance score propagation. In CVPR.","author":"Yu Ruichi","year":"2018","unstructured":"Ruichi Yu , Ang Li , Chun-Fu Chen , Jui-Hsin Lai , Vlad\u00a0 I Morariu , Xintong Han , Mingfei Gao , Ching-Yung Lin , and Larry\u00a0 S Davis . 2018 . Nisp: Pruning networks using neuron importance score propagation. In CVPR. Ruichi Yu, Ang Li, Chun-Fu Chen, Jui-Hsin Lai, Vlad\u00a0I Morariu, Xintong Han, Mingfei Gao, Ching-Yung Lin, and Larry\u00a0S Davis. 2018. Nisp: Pruning networks using neuron importance score propagation. In CVPR."},{"key":"e_1_3_2_1_31_1","unstructured":"Shuai Zheng Haibin Lin Sheng Zha and Mu Li. 2020. Accelerated Large Batch Optimization of BERT Pretraining in 54 minutes. arxiv:2006.13484\u00a0[cs.LG]  Shuai Zheng Haibin Lin Sheng Zha and Mu Li. 2020. Accelerated Large Batch Optimization of BERT Pretraining in 54 minutes. arxiv:2006.13484\u00a0[cs.LG]"},{"key":"e_1_3_2_1_32_1","unstructured":"Zhuangwei Zhuang Mingkui Tan Bohan Zhuang Jing Liu Yong Guo Qingyao Wu Junzhou Huang and Jinhui Zhu. 2018. Discrimination-aware channel pruning for deep neural networks. In NeurIPS.  Zhuangwei Zhuang Mingkui Tan Bohan Zhuang Jing Liu Yong Guo Qingyao Wu Junzhou Huang and Jinhui Zhu. 2018. Discrimination-aware channel pruning for deep neural networks. In NeurIPS."}],"event":{"name":"WWW '22: The ACM Web Conference 2022","location":"Virtual Event, Lyon France","acronym":"WWW '22","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web"]},"container-title":["Companion Proceedings of the Web Conference 2022"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3487553.3524856","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3487553.3524856","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3487553.3524856","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:23Z","timestamp":1750188623000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3487553.3524856"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,25]]},"references-count":32,"alternative-id":["10.1145\/3487553.3524856","10.1145\/3487553"],"URL":"https:\/\/doi.org\/10.1145\/3487553.3524856","relation":{},"subject":[],"published":{"date-parts":[[2022,4,25]]},"assertion":[{"value":"2022-08-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}