{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T16:53:54Z","timestamp":1783788834328,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":38,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,8,14]],"date-time":"2021-08-14T00:00:00Z","timestamp":1628899200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"NSF (National Science Foundation)","doi-asserted-by":"publisher","award":["CCF-1704967"],"award-info":[{"award-number":["CCF-1704967"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,8,14]]},"DOI":"10.1145\/3447548.3467080","type":"proceedings-article","created":{"date-parts":[[2021,8,13]],"date-time":"2021-08-13T18:21:39Z","timestamp":1628878899000},"page":"2928-2936","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":25,"title":["Training Recommender Systems at Scale: Communication-Efficient Model and Data Parallelism"],"prefix":"10.1145","author":[{"given":"Vipul","family":"Gupta","sequence":"first","affiliation":[{"name":"University of California, Berkeley, Berkeley, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dhruv","family":"Choudhary","sequence":"additional","affiliation":[{"name":"Facebook Inc., Menlo Park, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peter","family":"Tang","sequence":"additional","affiliation":[{"name":"Facebook Inc., Menlo Park, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaohan","family":"Wei","sequence":"additional","affiliation":[{"name":"Facebook Inc., Menlo Park, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xing","family":"Wang","sequence":"additional","affiliation":[{"name":"Facebook Inc., Menlo Park, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuzhen","family":"Huang","sequence":"additional","affiliation":[{"name":"Facebook Inc., Menlo Park, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Arun","family":"Kejariwal","sequence":"additional","affiliation":[{"name":"Facebook Inc., Menlo Park, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kannan","family":"Ramchandran","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Michael W.","family":"Mahoney","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,8,14]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1045"},{"key":"e_1_3_2_2_2_1","volume-title":"QSGD: Communication-efficient SGD via gradient quantization and encoding. In Advances in Neural Information Processing Systems. 1709--1720.","author":"Alistarh Dan","year":"2017","unstructured":"Dan Alistarh , Demjan Grubic , Jerry Li , Ryota Tomioka , and Milan Vojnovic . 2017 . QSGD: Communication-efficient SGD via gradient quantization and encoding. In Advances in Neural Information Processing Systems. 1709--1720. Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic. 2017. QSGD: Communication-efficient SGD via gradient quantization and encoding. In Advances in Neural Information Processing Systems. 1709--1720."},{"key":"e_1_3_2_2_3_1","unstructured":"Dan Alistarh Torsten Hoefler Mikael Johansson Nikola Konstantinov Sarit Khirirat and C\u00e9dric Renggli. 2018. The convergence of sparsified gradient methods. In Advances in Neural Information Processing Systems. 5973--5983.  Dan Alistarh Torsten Hoefler Mikael Johansson Nikola Konstantinov Sarit Khirirat and C\u00e9dric Renggli. 2018. The convergence of sparsified gradient methods. In Advances in Neural Information Processing Systems. 5973--5983."},{"key":"e_1_3_2_2_4_1","unstructured":"Jimmy Ba and Brendan Frey. 2013. Adaptive dropout for training deep neural networks. In Advances in neural information processing systems. 3084--3092.  Jimmy Ba and Brendan Frey. 2013. Adaptive dropout for training deep neural networks. In Advances in neural information processing systems. 3084--3092."},{"key":"e_1_3_2_2_5_1","volume-title":"Conference of european association for machine translation. 261--268","author":"Cettolo Mauro","year":"2012","unstructured":"Mauro Cettolo , Christian Girardi , and Marcello Federico . 2012 . Wit3: Web inventory of transcribed and translated talks . In Conference of european association for machine translation. 261--268 . Mauro Cettolo, Christian Girardi, and Marcello Federico. 2012. Wit3: Web inventory of transcribed and translated talks. In Conference of european association for machine translation. 261--268."},{"key":"e_1_3_2_2_6_1","volume-title":"Slide: In defense of smart algorithms over hardware acceleration for large-scale deep learning systems. preprint arXiv:1903.03129","author":"Chen Beidi","year":"2019","unstructured":"Beidi Chen , Tharun Medini , James Farwell , Sameh Gobriel , Charlie Tai , and Anshumali Shrivastava . 2019 . Slide: In defense of smart algorithms over hardware acceleration for large-scale deep learning systems. preprint arXiv:1903.03129 (2019). Beidi Chen, Tharun Medini, James Farwell, Sameh Gobriel, Charlie Tai, and Anshumali Shrivastava. 2019. Slide: In defense of smart algorithms over hardware acceleration for large-scale deep learning systems. preprint arXiv:1903.03129 (2019)."},{"key":"e_1_3_2_2_7_1","volume-title":"Revisiting distributed synchronous SGD. preprint arXiv:1604.00981","author":"Chen Jianmin","year":"2016","unstructured":"Jianmin Chen , Xinghao Pan , Rajat Monga , Samy Bengio , and Rafal Jozefowicz . 2016. Revisiting distributed synchronous SGD. preprint arXiv:1604.00981 ( 2016 ). Jianmin Chen, Xinghao Pan, Rajat Monga, Samy Bengio, and Rafal Jozefowicz. 2016. Revisiting distributed synchronous SGD. preprint arXiv:1604.00981 (2016)."},{"key":"e_1_3_2_2_8_1","unstructured":"Jeffrey Dean Greg Corrado Rajat Monga Kai Chen Matthieu Devin Mark Mao Andrew Senior Paul Tucker Ke Yang Quoc V Le etal 2012. Large scale distributed deep networks. In Advances in neural information processing systems. 1223--1231.  Jeffrey Dean Greg Corrado Rajat Monga Kai Chen Matthieu Devin Mark Mao Andrew Senior Paul Tucker Ke Yang Quoc V Le et al. 2012. Large scale distributed deep networks. In Advances in neural information processing systems. 1223--1231."},{"key":"e_1_3_2_2_9_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2000064.2000108"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2016.2637349"},{"key":"e_1_3_2_2_12_1","volume-title":"Proceedings of the fourteenth international conference on artificial intelligence and statistics. 315--323","author":"Glorot Xavier","year":"2011","unstructured":"Xavier Glorot , Antoine Bordes , and Yoshua Bengio . 2011 . Deep sparse rectifier neural networks . In Proceedings of the fourteenth international conference on artificial intelligence and statistics. 315--323 . Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011. Deep sparse rectifier neural networks. In Proceedings of the fourteenth international conference on artificial intelligence and statistics. 315--323."},{"key":"e_1_3_2_2_13_1","volume-title":"Deep learning","author":"Goodfellow Ian","unstructured":"Ian Goodfellow , Yoshua Bengio , and Aaron Courville . 2016. Deep learning . Vol. 1 . Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep learning. Vol. 1."},{"key":"e_1_3_2_2_14_1","volume-title":"large minibatch sgd: Training imagenet in 1 hour. preprint arXiv:1706.02677","author":"Goyal Priya","year":"2017","unstructured":"Priya Goyal , Piotr Doll\u00e1r , Ross Girshick , Pieter Noordhuis , Lukasz Wesolowski , Aapo Kyrola , Andrew Tulloch , Yangqing Jia , and Kaiming He. 2017. Accurate , large minibatch sgd: Training imagenet in 1 hour. preprint arXiv:1706.02677 ( 2017 ). Priya Goyal, Piotr Doll\u00e1r, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017. Accurate, large minibatch sgd: Training imagenet in 1 hour. preprint arXiv:1706.02677 (2017)."},{"key":"e_1_3_2_2_15_1","volume-title":"Stochastic Weight Averaging in Parallel: Large-Batch Training That Generalizes Well. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=rygFWAEFwS","author":"Gupta Vipul","year":"2020","unstructured":"Vipul Gupta , Santiago Akle Serrano , and Dennis DeCoste . 2020 . Stochastic Weight Averaging in Parallel: Large-Batch Training That Generalizes Well. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=rygFWAEFwS Vipul Gupta, Santiago Akle Serrano, and Dennis DeCoste. 2020. Stochastic Weight Averaging in Parallel: Large-Batch Training That Generalizes Well. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=rygFWAEFwS"},{"key":"e_1_3_2_2_16_1","unstructured":"Elad Hoffer Itay Hubara and Daniel Soudry. 2017. Train longer generalize better: closing the generalization gap in large batch training of neural networks. In NIPS .  Elad Hoffer Itay Hubara and Daniel Soudry. 2017. Train longer generalize better: closing the generalization gap in large batch training of neural networks. In NIPS ."},{"key":"e_1_3_2_2_17_1","unstructured":"Yanping Huang Youlong Cheng Ankur Bapna Orhan Firat Dehao Chen Mia Chen HyoukJoong Lee Jiquan Ngiam Quoc V Le Yonghui Wu and zhifeng Chen. 2019. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism. In Advances in Neural Information Processing Systems 32. 103--112.  Yanping Huang Youlong Cheng Ankur Bapna Orhan Firat Dehao Chen Mia Chen HyoukJoong Lee Jiquan Ngiam Quoc V Le Yonghui Wu and zhifeng Chen. 2019. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism. In Advances in Neural Information Processing Systems 32. 103--112."},{"key":"e_1_3_2_2_18_1","unstructured":"Nikita Ivkin Daniel Rothchild Enayat Ullah Vladimir Braverman Ion Stoica and Raman Arora. 2019. Communication-efficient distributed sgd with sketching. In Advances in Neural Information Processing Systems. 13144--13154.  Nikita Ivkin Daniel Rothchild Enayat Ullah Vladimir Braverman Ion Stoica and Raman Arora. 2019. Communication-efficient distributed sgd with sketching. In Advances in Neural Information Processing Systems. 13144--13154."},{"key":"e_1_3_2_2_19_1","unstructured":"Zhihao Jia Matei Zaharia and Alex Aiken. 2019. Beyond data and model parallelism for deep neural networks. (2019).  Zhihao Jia Matei Zaharia and Alex Aiken. 2019. Beyond data and model parallelism for deep neural networks. (2019)."},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196894"},{"key":"e_1_3_2_2_21_1","volume-title":"International Conference on Machine Learning. 3252--3261","author":"Karimireddy Sai Praneeth","year":"2019","unstructured":"Sai Praneeth Karimireddy , Quentin Rebjock , Sebastian Stich , and Martin Jaggi . 2019 . Error Feedback Fixes SignSGD and other Gradient Compression Schemes . In International Conference on Machine Learning. 3252--3261 . Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi. 2019. Error Feedback Fixes SignSGD and other Gradient Compression Schemes. In International Conference on Machine Learning. 3252--3261."},{"key":"e_1_3_2_2_22_1","volume-title":"On large-batch training for deep learning: Generalization gap and sharp minima. arXiv preprint arXiv:1609.04836","author":"Keskar Nitish Shirish","year":"2016","unstructured":"Nitish Shirish Keskar , Dheevatsa Mudigere , Jorge Nocedal , Mikhail Smelyanskiy , and Ping Tak Peter Tang . 2016. On large-batch training for deep learning: Generalization gap and sharp minima. arXiv preprint arXiv:1609.04836 ( 2016 ). Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2016. On large-batch training for deep learning: Generalization gap and sharp minima. arXiv preprint arXiv:1609.04836 (2016)."},{"key":"e_1_3_2_2_23_1","volume-title":"Xun Zheng, Qirong Ho, Garth A Gibson, and Eric P Xing.","author":"Lee Seunghak","year":"2014","unstructured":"Seunghak Lee , Jin Kyu Kim , Xun Zheng, Qirong Ho, Garth A Gibson, and Eric P Xing. 2014 . On model parallelization and scheduling strategies for distributed machine learning. In Advances in neural inf. processing systems. 2834--2842. Seunghak Lee, Jin Kyu Kim, Xun Zheng, Qirong Ho, Garth A Gibson, and Eric P Xing. 2014. On model parallelization and scheduling strategies for distributed machine learning. In Advances in neural inf. processing systems. 2834--2842."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213556.2213579"},{"key":"e_1_3_2_2_25_1","volume-title":"K-sparse autoencoders. arXiv preprint arXiv:1312.5663","author":"Makhzani Alireza","year":"2013","unstructured":"Alireza Makhzani and Brendan Frey . 2013. K-sparse autoencoders. arXiv preprint arXiv:1312.5663 ( 2013 ). Alireza Makhzani and Brendan Frey. 2013. K-sparse autoencoders. arXiv preprint arXiv:1312.5663 (2013)."},{"key":"e_1_3_2_2_26_1","unstructured":"Alireza Makhzani and Brendan J Frey. 2015. Winner-take-all autoencoders. In Advances in neural information processing systems. 2791--2799.  Alireza Makhzani and Brendan J Frey. 2015. Winner-take-all autoencoders. In Advances in neural information processing systems. 2791--2799."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359646"},{"key":"e_1_3_2_2_28_1","volume-title":"Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, et al.","author":"Naumov Maxim","year":"2019","unstructured":"Maxim Naumov , Dheevatsa Mudigere , Hao-Jun Michael Shi , Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, et al. 2019 . Deep learning recommendation model for personalization and recommendation systems. arXiv preprint arXiv:1906.00091 (2019). Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, et al. 2019. Deep learning recommendation model for personalization and recommendation systems. arXiv preprint arXiv:1906.00091 (2019)."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-4009"},{"key":"e_1_3_2_2_30_1","volume-title":"Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Advances in neural information processing systems. 693--701.","author":"Recht Benjamin","year":"2011","unstructured":"Benjamin Recht , Christopher Re , Stephen Wright , and Feng Niu . 2011 . Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Advances in neural information processing systems. 693--701. Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu. 2011. Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Advances in neural information processing systems. 693--701."},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401113"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994522"},{"key":"e_1_3_2_2_33_1","volume-title":"Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research","author":"Srivastava Nitish","year":"2014","unstructured":"Nitish Srivastava , Geoffrey Hinton , Alex Krizhevsky , Ilya Sutskever , and Ruslan Salakhutdinov . 2014. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research , Vol. 15 , 1 ( 2014 ), 1929--1958. Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, Vol. 15, 1 (2014), 1929--1958."},{"key":"e_1_3_2_2_34_1","unstructured":"Sebastian U Stich Jean-Baptiste Cordonnier and Martin Jaggi. 2018. Sparsified SGD with memory. In Advances in Neural Inf. Processing Systems. 4447--4458.  Sebastian U Stich Jean-Baptiste Cordonnier and Martin Jaggi. 2018. Sparsified SGD with memory. In Advances in Neural Inf. Processing Systems. 4447--4458."},{"key":"e_1_3_2_2_35_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008."},{"key":"e_1_3_2_2_36_1","volume-title":"Sai Praneeth Karimireddy, and Martin Jaggi","author":"Vogels Thijs","year":"2019","unstructured":"Thijs Vogels , Sai Praneeth Karimireddy, and Martin Jaggi . 2019 . PowerSGD: Practical low-rank gradient compression for distributed optimization. In Advances in Neural Information Processing Systems . 14259--14268. Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi. 2019. PowerSGD: Practical low-rank gradient compression for distributed optimization. In Advances in Neural Information Processing Systems. 14259--14268."},{"key":"e_1_3_2_2_38_1","volume-title":"PipeMare: Asynchronous Pipeline Parallel DNN Training. arXiv preprint arXiv:1910.05124","author":"Yang Bowen","year":"2019","unstructured":"Bowen Yang , Jian Zhang , Jonathan Li , Christopher R\u00e9 , Christopher R Aberger , and Christopher De Sa. 2019. PipeMare: Asynchronous Pipeline Parallel DNN Training. arXiv preprint arXiv:1910.05124 ( 2019 ). Bowen Yang, Jian Zhang, Jonathan Li, Christopher R\u00e9, Christopher R Aberger, and Christopher De Sa. 2019. PipeMare: Asynchronous Pipeline Parallel DNN Training. arXiv preprint arXiv:1910.05124 (2019)."},{"key":"e_1_3_2_2_39_1","volume-title":"Scaling SGD batch size to 32k for ImageNet training. arXiv preprint arXiv:1708.03888","author":"You Yang","year":"2017","unstructured":"Yang You , Igor Gitman , and Boris Ginsburg . 2017. Scaling SGD batch size to 32k for ImageNet training. arXiv preprint arXiv:1708.03888 ( 2017 ). Yang You, Igor Gitman, and Boris Ginsburg. 2017. Scaling SGD batch size to 32k for ImageNet training. arXiv preprint arXiv:1708.03888 (2017)."}],"event":{"name":"KDD '21: The 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","location":"Virtual Event Singapore","acronym":"KDD '21","sponsor":["SIGMOD ACM Special Interest Group on Management of Data","SIGKDD ACM Special Interest Group on Knowledge Discovery in Data"]},"container-title":["Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery &amp; Data Mining"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447548.3467080","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3447548.3467080","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3447548.3467080","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:25:11Z","timestamp":1750195511000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447548.3467080"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,14]]},"references-count":38,"alternative-id":["10.1145\/3447548.3467080","10.1145\/3447548"],"URL":"https:\/\/doi.org\/10.1145\/3447548.3467080","relation":{},"subject":[],"published":{"date-parts":[[2021,8,14]]},"assertion":[{"value":"2021-08-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}