{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,27]],"date-time":"2025-08-27T15:52:55Z","timestamp":1756309975048,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":51,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,8,14]],"date-time":"2021-08-14T00:00:00Z","timestamp":1628899200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,8,14]]},"DOI":"10.1145\/3447548.3467084","type":"proceedings-article","created":{"date-parts":[[2021,8,13]],"date-time":"2021-08-13T18:21:39Z","timestamp":1628878899000},"page":"3050-3058","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Hierarchical Training: Scaling Deep Recommendation Models on Large CPU Clusters"],"prefix":"10.1145","author":[{"given":"Yuzhen","family":"Huang","sequence":"first","affiliation":[{"name":"Facebook, Menlo Park, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaohan","family":"Wei","sequence":"additional","affiliation":[{"name":"Facebook, Menlo Park, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xing","family":"Wang","sequence":"additional","affiliation":[{"name":"Facebook, Menlo Park, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiyan","family":"Yang","sequence":"additional","affiliation":[{"name":"Facebook, Menlo Park, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bor-Yiing","family":"Su","sequence":"additional","affiliation":[{"name":"Facebook, Menlo Park, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shivam","family":"Bharuka","sequence":"additional","affiliation":[{"name":"Facebook, Menlo Park, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dhruv","family":"Choudhary","sequence":"additional","affiliation":[{"name":"Facebook, Menlo Park, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zewei","family":"Jiang","sequence":"additional","affiliation":[{"name":"Facebook, Menlo Park, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hai","family":"Zheng","sequence":"additional","affiliation":[{"name":"Facebook, Menlo Park, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jack","family":"Langman","sequence":"additional","affiliation":[{"name":"Facebook, Menlo Park, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,8,14]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"12th USENIX symposium on operating systems design and implementation (OSDI 16)","author":"Abadi Martin","year":"2016","unstructured":"Martin Abadi , Paul Barham , Jianmin Chen , Zhifeng Chen , Andy Davis , Jeffrey Dean , Matthieu Devin , Sanjay Ghemawat , Geoffrey Irving , Michael Isard , 2016 . Tensorflow: A system for large-scale machine learning . In 12th USENIX symposium on operating systems design and implementation (OSDI 16) . 265--283. Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. 2016. Tensorflow: A system for large-scale machine learning. In 12th USENIX symposium on operating systems design and implementation (OSDI 16). 265--283."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"crossref","unstructured":"Bilge Acun Matthew Murphy Xiaodong Wang Jade Nie Carole-Jean Wu and Kim Hazelwood. 2020. Understanding Training Efficiency of Deep Learning Recommendation Models at Scale. arXiv preprint arXiv:2011.05497(2020).  Bilge Acun Matthew Murphy Xiaodong Wang Jade Nie Carole-Jean Wu and Kim Hazelwood. 2020. Understanding Training Efficiency of Deep Learning Recommendation Models at Scale. arXiv preprint arXiv:2011.05497(2020).","DOI":"10.1109\/HPCA51647.2021.00072"},{"key":"e_1_3_2_1_3_1","unstructured":"Caffe2 Operators Catalog. 2021. https:\/\/caffe2.ai\/docs\/operators-catalogue.html.(2021).  Caffe2 Operators Catalog. 2021. https:\/\/caffe2.ai\/docs\/operators-catalogue.html.(2021)."},{"key":"e_1_3_2_1_4_1","volume-title":"TensorOpt: Exploring the Tradeoffs in Distributed DNN Training with Auto-Parallelism. CoRRabs\/2004.10856","author":"Cai Zhenkun","year":"2020","unstructured":"Zhenkun Cai , Kaihao Ma , Xiao Yan , Yidi Wu , Yuzhen Huang , James Cheng , Teng Su , and Fan Yu. 2020. TensorOpt: Exploring the Tradeoffs in Distributed DNN Training with Auto-Parallelism. CoRRabs\/2004.10856 ( 2020 ). arXiv:2004.10856 https:\/\/arxiv.org\/abs\/2004.10856 Zhenkun Cai, Kaihao Ma, Xiao Yan, Yidi Wu, Yuzhen Huang, James Cheng, Teng Su, and Fan Yu. 2020. TensorOpt: Exploring the Tradeoffs in Distributed DNN Training with Auto-Parallelism. CoRRabs\/2004.10856 (2020). arXiv:2004.10856 https:\/\/arxiv.org\/abs\/2004.10856"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Kai Chen and Qiang Huo. 2016. Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering. In2016 ieee international conference on acoustics speech and signal processing (icassp). IEEE 5880--5884.  Kai Chen and Qiang Huo. 2016. Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering. In2016 ieee international conference on acoustics speech and signal processing (icassp). IEEE 5880--5884.","DOI":"10.1109\/ICASSP.2016.7472805"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2988450.2988454"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2959100.2959190"},{"key":"e_1_3_2_1_8_1","unstructured":"Jeffrey Dean Greg Corrado Rajat Monga Kai Chen Matthieu Devin Mark Mao Marc'aurelio Ranzato Andrew Senior Paul Tucker Ke Yang etal 2012. Large scale distributed deep networks. Advances in neural information processing systems 25 (2012) 1223--1231.  Jeffrey Dean Greg Corrado Rajat Monga Kai Chen Matthieu Devin Mark Mao Marc'aurelio Ranzato Andrew Senior Paul Tucker Ke Yang et al. 2012. Large scale distributed deep networks. Advances in neural information processing systems 25 (2012) 1223--1231."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2736277.2741667"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2843948"},{"key":"e_1_3_2_1_11_1","unstructured":"Priya Goyal Piotr Doll\u00e1r Ross Girshick Pieter Noordhuis Lukasz Wesolowski Aapo Kyrola Andrew Tulloch Yangqing Jia and Kaiming He. 2017. Accurate large minibatch sgd: Training imagenet in 1 hour. arXiv:1706.02677(2017).  Priya Goyal Piotr Doll\u00e1r Ross Girshick Pieter Noordhuis Lukasz Wesolowski Aapo Kyrola Andrew Tulloch Yangqing Jia and Kaiming He. 2017. Accurate large minibatch sgd: Training imagenet in 1 hour. arXiv:1706.02677(2017)."},{"key":"e_1_3_2_1_12_1","unstructured":"Hui Guan Andrey Malevich Jiyan Yang Jongsoo Park and Hector Yuen. 2019. Post-training 4-bit quantization on embedding tables. arXiv preprintarXiv:1911.02079(2019).  Hui Guan Andrey Malevich Jiyan Yang Jongsoo Park and Hector Yuen. 2019. Post-training 4-bit quantization on embedding tables. arXiv preprintarXiv:1911.02079(2019)."},{"key":"e_1_3_2_1_13_1","unstructured":"Huifeng Guo Ruiming Tang Yunming Ye Zhenguo Li and Xiuqiang He. 2017. DeepFM: a factorization-machine based neural network for CTR prediction. arXivpreprint arXiv:1703.04247(2017).  Huifeng Guo Ruiming Tang Yunming Ye Zhenguo Li and Xiuqiang He. 2017. DeepFM: a factorization-machine based neural network for CTR prediction. arXivpreprint arXiv:1703.04247(2017)."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00084"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00047"},{"key":"e_1_3_2_1_16_1","volume-title":"Xiaohan Wei, Xing Wang, Yuzhen Huang, Arun Kejariwal, Kannan Ramchandran, and Michael W. Mahoney.","author":"Gupta Vipul","year":"2020","unstructured":"Vipul Gupta , Dhruv Choudhary , Ping Tak Peter Tang , Xiaohan Wei, Xing Wang, Yuzhen Huang, Arun Kejariwal, Kannan Ramchandran, and Michael W. Mahoney. 2020 . Fast Distributed Training of Deep Neural Networks: Dynamic Communication Thresholding for Model and Data Parallelism. CoRRabs\/ 2010.08899 (2020). arXiv:2010.08899 https:\/\/arxiv.org\/abs\/2010.08899 Vipul Gupta, Dhruv Choudhary, Ping Tak Peter Tang, Xiaohan Wei, Xing Wang, Yuzhen Huang, Arun Kejariwal, Kannan Ramchandran, and Michael W. Mahoney. 2020. Fast Distributed Training of Deep Neural Networks: Dynamic Communication Thresholding for Model and Data Parallelism. CoRRabs\/2010.08899 (2020). arXiv:2010.08899 https:\/\/arxiv.org\/abs\/2010.08899"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00059"},{"volume-title":"Proceedings of the Eighth International Workshop on Data Mining for Online Advertising. 1--9.","author":"He Xinran","key":"e_1_3_2_1_18_1","unstructured":"Xinran He , Junfeng Pan , Ou Jin , Tianbing Xu , Bo Liu , Tao Xu , Yanxin Shi , Antoine Atallah , Ralf Herbrich , Stuart Bowers , Practical lessons from predicting clicks on ads at facebook . In Proceedings of the Eighth International Workshop on Data Mining for Online Advertising. 1--9. Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, et al.2014. Practical lessons from predicting clicks on ads at facebook. In Proceedings of the Eighth International Workshop on Data Mining for Online Advertising. 1--9."},{"volume-title":"Proceedings of the Eighth International Workshop on Data Mining for Online Advertising (ADKDD'14)","author":"He Xinran","key":"e_1_3_2_1_19_1","unstructured":"Xinran He , Junfeng Pan , Ou Jin , Tianbing Xu , Bo Liu , Tao Xu , Yanxin Shi , Antoine Atallah , Ralf Herbrich , Stuart Bowers , and et al. 2014. Practical Lessons from Predicting Clicks on Ads at Facebook . In Proceedings of the Eighth International Workshop on Data Mining for Online Advertising (ADKDD'14) . Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/2648584.2648589 10.1145\/2648584.2648589 Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, and et al. 2014. Practical Lessons from Predicting Clicks on Ads at Facebook. In Proceedings of the Eighth International Workshop on Data Mining for Online Advertising (ADKDD'14). Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/2648584.2648589"},{"key":"e_1_3_2_1_20_1","volume-title":"Dehao Chen, Hyouk Joong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al.","author":"Huang Yanping","year":"2018","unstructured":"Yanping Huang , Youlong Cheng , Ankur Bapna , Orhan Firat , Mia Xu Chen , Dehao Chen, Hyouk Joong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al. 2018 . Gpipe : Efficient training of giant neural networks using pipeline parallelism. arXiv preprint arXiv:1811.06965(2018). Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, Hyouk Joong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al. 2018. Gpipe: Efficient training of giant neural networks using pipeline parallelism. arXiv preprint arXiv:1811.06965(2018)."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3187009.3177734"},{"key":"e_1_3_2_1_22_1","volume-title":"2019 USENIX Annual Technical Conference (USENIX ATC 19)","author":"Huang Yuzhen","year":"2019","unstructured":"Yuzhen Huang , Xiao Yan , Guanxian Jiang , Tatiana Jin , James Cheng , An Xu , Zhanhao Liu , and Shuo Tu . 2019 . Tangram: bridging immutable and mutable abstractions for distributed data analytics . In 2019 USENIX Annual Technical Conference (USENIX ATC 19) . 191--206. Yuzhen Huang, Xiao Yan, Guanxian Jiang, Tatiana Jin, James Cheng, An Xu,Zhanhao Liu, and Shuo Tu. 2019. Tangram: bridging immutable and mutable abstractions for distributed data analytics. In 2019 USENIX Annual Technical Conference (USENIX ATC 19). 191--206."},{"key":"e_1_3_2_1_23_1","unstructured":"Zhihao Jia Matei Zaharia and Alex Aiken. 2018. Beyond data and model parallelism for deep neural networks. arXiv preprint arXiv:1807.05358(2018).  Zhihao Jia Matei Zaharia and Alex Aiken. 2018. Beyond data and model parallelism for deep neural networks. arXiv preprint arXiv:1807.05358(2018)."},{"volume-title":"International Conference on Machine Learning. 3252--3261","author":"Karimireddy Sai Praneeth","key":"e_1_3_2_1_24_1","unstructured":"Sai Praneeth Karimireddy , Quentin Rebjock , Sebastian Stich , and Martin Jaggi .2019. Error Feedback Fixes Sign SGD and other Gradient Compression Schemes . In International Conference on Machine Learning. 3252--3261 . Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi.2019. Error Feedback Fixes Sign SGD and other Gradient Compression Schemes. In International Conference on Machine Learning. 3252--3261."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2901318.2901331"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358284"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2640087.2644155"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3220023"},{"key":"e_1_3_2_1_29_1","volume-title":"International Conference on Machine Learning. PMLR, 3043--3052","author":"Lian Xiangru","year":"2018","unstructured":"Xiangru Lian , Wei Zhang , Ce Zhang , and Ji Liu . 2018 . Asynchronous decentralized parallel stochastic gradient descent . In International Conference on Machine Learning. PMLR, 3043--3052 . Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu. 2018. Asynchronous decentralized parallel stochastic gradient descent. In International Conference on Machine Learning. PMLR, 3043--3052."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359646"},{"key":"e_1_3_2_1_31_1","unstructured":"Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G. Azzolini Dmytro Dzhulgakov Andrey Mallevich Ilia Cherniavskii Yinghai Lu Raghuraman Krishnamoorthi Ansha Yu Volodymyr Kondratenko Stephanie Pereira Xianjie Chen Wenlin Chen Vijay Rao Bill Jia LiangXiong and Misha Smelyanskiy. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems.arXiv preprint arXiv:1906.00091(2019).  Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G. Azzolini Dmytro Dzhulgakov Andrey Mallevich Ilia Cherniavskii Yinghai Lu Raghuraman Krishnamoorthi Ansha Yu Volodymyr Kondratenko Stephanie Pereira Xianjie Chen Wenlin Chen Vijay Rao Bill Jia LiangXiong and Misha Smelyanskiy. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems.arXiv preprint arXiv:1906.00091(2019)."},{"key":"e_1_3_2_1_32_1","volume-title":"Pytorch: An imperative style, high-performance deep learning library. In Advances in neural information processing systems. 8026--8037.","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , 2019 . Pytorch: An imperative style, high-performance deep learning library. In Advances in neural information processing systems. 8026--8037. Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. In Advances in neural information processing systems. 8026--8037."},{"key":"e_1_3_2_1_33_1","volume-title":"A lock-free approach to parallelizing stochastic gradient descent. Advances in neural information processing systems 24","author":"Recht Benjamin","year":"2011","unstructured":"Benjamin Recht , Christopher Re , Stephen Wright , and Feng Niu . 2011. Hogwild! : A lock-free approach to parallelizing stochastic gradient descent. Advances in neural information processing systems 24 ( 2011 ), 693--701. Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu. 2011. Hogwild!: A lock-free approach to parallelizing stochastic gradient descent. Advances in neural information processing systems 24 (2011), 693--701."},{"key":"e_1_3_2_1_34_1","volume-title":"Factorization Machines. In 2010 IEEE International Conference on Data Mining. 995--1000","author":"Rendle S.","year":"2010","unstructured":"S. Rendle . 2010 . Factorization Machines. In 2010 IEEE International Conference on Data Mining. 995--1000 . https:\/\/doi.org\/10.1109\/ICDM.2010.127 10.1109\/ICDM.2010.127 S. Rendle. 2010. Factorization Machines. In 2010 IEEE International Conference on Data Mining. 995--1000. https:\/\/doi.org\/10.1109\/ICDM.2010.127"},{"key":"e_1_3_2_1_35_1","unstructured":"Noam Shazeer Youlong Cheng Niki Parmar Dustin Tran Ashish Vaswani Pen-porn Koanantakool Peter Hawkins Hyouk Joong Lee Mingsheng Hong CliffYoung et al.2018. Mesh-tensor flow: Deep learning for supercomputers. arXivpreprint arXiv:1811.02084(2018).  Noam Shazeer Youlong Cheng Niki Parmar Dustin Tran Ashish Vaswani Pen-porn Koanantakool Peter Hawkins Hyouk Joong Lee Mingsheng Hong CliffYoung et al.2018. Mesh-tensor flow: Deep learning for supercomputers. arXivpreprint arXiv:1811.02084(2018)."},{"key":"e_1_3_2_1_36_1","volume-title":"Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 165--175","author":"Michael Shi Hao-Jun","year":"2020","unstructured":"Hao-Jun Michael Shi , Dheevatsa Mudigere , Maxim Naumov , and Jiyan Yang . 2020 . Compositional embeddings using complementary partitions for memory-efficient recommendation systems . In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 165--175 . Hao-Jun Michael Shi, Dheevatsa Mudigere, Maxim Naumov, and Jiyan Yang. 2020. Compositional embeddings using complementary partitions for memory-efficient recommendation systems. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 165--175."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403137"},{"key":"e_1_3_2_1_38_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. arXiv preprint arXiv:1706.03762(2017).  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. arXiv preprint arXiv:1706.03762(2017)."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3302424.3303953"},{"key":"e_1_3_2_1_40_1","volume-title":"Elastic Deep Learning in Multi-Tenant GPU Clusters","author":"Wu Yidi","year":"2021","unstructured":"Yidi Wu , Kaihao Ma , Xiao Yan , Zhi Liu , Zhenkun Cai , Yuzhen Huang , James Cheng , Han Yuan , and Fan Yu. 2021. Elastic Deep Learning in Multi-Tenant GPU Clusters . IEEE Transactions on Parallel and Distributed Systems( 2021 ). Yidi Wu, Kaihao Ma, Xiao Yan, Zhi Liu, Zhenkun Cai, Yuzhen Huang, James Cheng, Han Yuan, and Fan Yu. 2021. Elastic Deep Learning in Multi-Tenant GPU Clusters. IEEE Transactions on Parallel and Distributed Systems(2021)."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.634"},{"key":"e_1_3_2_1_42_1","volume-title":"Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, and Yaoliang Yu.","author":"Xing Eric P","year":"2015","unstructured":"Eric P Xing , Qirong Ho , Wei Dai , Jin Kyu Kim , Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, and Yaoliang Yu. 2015 . Petuum : A new platform for distributed machine learning on big data.IEEE Transactions on Big Data 1, 2 (2015), 49--67. Eric P Xing, Qirong Ho, Wei Dai, Jin Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, and Yaoliang Yu. 2015. Petuum: A new platform for distributed machine learning on big data.IEEE Transactions on Big Data 1, 2 (2015), 49--67."},{"key":"e_1_3_2_1_43_1","unstructured":"Xpress Optimization. 2021. https:\/\/www.fico.com\/en\/products\/fico-xpress-optimization. (2021).  Xpress Optimization. 2021. https:\/\/www.fico.com\/en\/products\/fico-xpress-optimization. (2021)."},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.14778\/3067421.3067424"},{"key":"e_1_3_2_1_45_1","volume-title":"Nomad: Non-locking, stochastic multi-machine algorithm for asynchronous and decentralized matrix completion. arXiv preprint arXiv:1312.0193(2013).","author":"Yun Hyokun","year":"2013","unstructured":"Hyokun Yun , Hsiang-Fu Yu , Cho-Jui Hsieh , SVN Vishwanathan , and Inderjit Dhillon . 2013 . Nomad: Non-locking, stochastic multi-machine algorithm for asynchronous and decentralized matrix completion. arXiv preprint arXiv:1312.0193(2013). Hyokun Yun, Hsiang-Fu Yu, Cho-Jui Hsieh, SVN Vishwanathan, and Inderjit Dhillon. 2013. Nomad: Non-locking, stochastic multi-machine algorithm for asynchronous and decentralized matrix completion. arXiv preprint arXiv:1312.0193(2013)."},{"key":"e_1_3_2_1_46_1","unstructured":"Sixin Zhang Anna E Choromanska and Yann LeCun. 2015. Deep learning with elastic averaging SGD. Advances in neural information processing systems 28(2015) 685--693.  Sixin Zhang Anna E Choromanska and Yann LeCun. 2015. Deep learning with elastic averaging SGD. Advances in neural information processing systems 28(2015) 685--693."},{"key":"e_1_3_2_1_47_1","unstructured":"Weijie Zhao Deping Xie Ronglai Jia Yulei Qian Ruiquan Ding Mingming Sun and Ping Li. 2020. Distributed Hierarchical GPU Parameter Server for Massive Scale Deep Learning Ads Systems. arXiv preprint arXiv:2003.05622(2020).  Weijie Zhao Deping Xie Ronglai Jia Yulei Qian Ruiquan Ding Mingming Sun and Ping Li. 2020. Distributed Hierarchical GPU Parameter Server for Massive Scale Deep Learning Ads Systems. arXiv preprint arXiv:2003.05622(2020)."},{"key":"e_1_3_2_1_48_1","volume-title":"Shadow Sync: Performing Synchronization in the Background for Highly Scalable Distributed Training. arXiv preprint arXiv:2003.03477(2020).","author":"Zheng Qinqing","year":"2020","unstructured":"Qinqing Zheng , Bor-Yiing Su , Jiyan Yang , Alisson Azzolini , Qiang Wu , Ou Jin , Shri Karandikar , Hagay Lupesko , Liang Xiong , and Eric Zhou . 2020 . Shadow Sync: Performing Synchronization in the Background for Highly Scalable Distributed Training. arXiv preprint arXiv:2003.03477(2020). Qinqing Zheng, Bor-Yiing Su, Jiyan Yang, Alisson Azzolini, Qiang Wu, Ou Jin,Shri Karandikar, Hagay Lupesko, Liang Xiong, and Eric Zhou. 2020. Shadow Sync: Performing Synchronization in the Background for Highly Scalable Distributed Training. arXiv preprint arXiv:2003.03477(2020)."},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33015941"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219823"},{"key":"e_1_3_2_1_51_1","unstructured":"Martin Zinkevich Markus Weimer Lihong Li and Alex J Smola. 2010. Parallelized stochastic gradient descent. In Advances in neural information processing systems. 2595--2603.  Martin Zinkevich Markus Weimer Lihong Li and Alex J Smola. 2010. Parallelized stochastic gradient descent. In Advances in neural information processing systems. 2595--2603."}],"event":{"name":"KDD '21: The 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","sponsor":["SIGMOD ACM Special Interest Group on Management of Data","SIGKDD ACM Special Interest Group on Knowledge Discovery in Data"],"location":"Virtual Event Singapore","acronym":"KDD '21"},"container-title":["Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery &amp; Data Mining"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447548.3467084","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3447548.3467084","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:25:11Z","timestamp":1750195511000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447548.3467084"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,14]]},"references-count":51,"alternative-id":["10.1145\/3447548.3467084","10.1145\/3447548"],"URL":"https:\/\/doi.org\/10.1145\/3447548.3467084","relation":{},"subject":[],"published":{"date-parts":[[2021,8,14]]},"assertion":[{"value":"2021-08-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}