{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,25]],"date-time":"2026-07-25T16:14:25Z","timestamp":1784996065817,"version":"3.55.0"},"reference-count":171,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2020,3,20]],"date-time":"2020-03-20T00:00:00Z","timestamp":1584662400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2021,3,31]]},"abstract":"<jats:p>The demand for artificial intelligence has grown significantly over the past decade, and this growth has been fueled by advances in machine learning techniques and the ability to leverage hardware acceleration. However, to increase the quality of predictions and render machine learning solutions feasible for more complex applications, a substantial amount of training data is required. Although small machine learning models can be trained with modest amounts of data, the input for training larger models such as neural networks grows exponentially with the number of parameters. Since the demand for processing training data has outpaced the increase in computation power of computing machinery, there is a need for distributing the machine learning workload across multiple machines, and turning the centralized into a distributed system. These distributed systems present new challenges: first and foremost, the efficient parallelization of the training process and the creation of a coherent model. This article provides an extensive overview of the current state-of-the-art in the field by outlining the challenges and opportunities of distributed machine learning over conventional (centralized) machine learning, discussing the techniques used for distributed machine learning, and providing an overview of the systems that are available.<\/jats:p>","DOI":"10.1145\/3377454","type":"journal-article","created":{"date-parts":[[2020,3,20]],"date-time":"2020-03-20T21:04:17Z","timestamp":1584738257000},"page":"1-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":736,"title":["A Survey on Distributed Machine Learning"],"prefix":"10.1145","volume":"53","author":[{"given":"Joost","family":"Verbraeken","sequence":"first","affiliation":[{"name":"Delft University of Technology, Delft, Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Matthijs","family":"Wolting","sequence":"additional","affiliation":[{"name":"Delft University of Technology, Delft, Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jonathan","family":"Katzy","sequence":"additional","affiliation":[{"name":"Delft University of Technology, Delft, Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jeroen","family":"Kloppenburg","sequence":"additional","affiliation":[{"name":"Delft University of Technology, Delft, Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tim","family":"Verbelen","sequence":"additional","affiliation":[{"name":"imec - Ghent University, Ghent, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3791-7114","authenticated-orcid":false,"given":"Jan S.","family":"Rellermeyer","sequence":"additional","affiliation":[{"name":"Delft University of Technology, Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,3,20]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. Retrieved from https:\/\/www.tensorflow.org\/.  Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. Retrieved from https:\/\/www.tensorflow.org\/."},{"key":"e_1_2_1_2_1","volume-title":"Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201916)","author":"Abadi Martin","year":"2016"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2976749.2978318"},{"key":"e_1_2_1_4_1","unstructured":"Adapteva Inc. 2017. E64G401 Epiphany 64-core Microprocessor Datasheet. Retrieved from http:\/\/www.adapteva.com\/docs\/e64g401_datasheet.pdf.  Adapteva Inc. 2017. E64G401 Epiphany 64-core Microprocessor Datasheet. Retrieved from http:\/\/www.adapteva.com\/docs\/e64g401_datasheet.pdf."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2638571"},{"key":"e_1_2_1_6_1","volume-title":"What does fault tolerant deep learning need from MPI? CoRR abs\/1709.03316","author":"Amatya Vinay","year":"2017"},{"key":"e_1_2_1_7_1","unstructured":"Amazon Web Services. 2018. Amazon SageMaker. Retrieved from https:\/\/aws.amazon.com\/sagemaker\/developer-resources\/.  Amazon Web Services. 2018. Amazon SageMaker. Retrieved from https:\/\/aws.amazon.com\/sagemaker\/developer-resources\/."},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the 33rd International Conference on Machine Learning, Maria Florina Balcan and Kilian Q. Weinberger (Eds.)","volume":"48","author":"Amodei Dario","year":"2016"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the ACM\/IEEE Conference on Supercomputing. IEEE Computer Society Press, 2--11","author":"Anderson Edward","year":"1990"},{"key":"e_1_2_1_10_1","unstructured":"Apple. 2017. Core ML Model Format Specification. Retrieved from https:\/\/apple.github.io\/coremltools\/coremlspecification\/.  Apple. 2017. Core ML Model Format Specification. Retrieved from https:\/\/apple.github.io\/coremltools\/coremlspecification\/."},{"key":"e_1_2_1_11_1","unstructured":"Apple. 2018. A12 Bionic. Retrieved from https:\/\/www.apple.com\/iphone-xs\/a12-bionic\/.  Apple. 2018. A12 Bionic. Retrieved from https:\/\/www.apple.com\/iphone-xs\/a12-bionic\/."},{"key":"e_1_2_1_12_1","volume-title":"How to backdoor federated learning. arXiv preprint arXiv:1807.00459","author":"Bagdasaryan Eugene","year":"2018"},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems. 91--98","author":"Bagnell Drew"},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the Conference on Learning Theory. 26--1.","author":"Balcan Maria-Florina","year":"2012"},{"key":"e_1_2_1_15_1","doi-asserted-by":"crossref","volume-title":"On Distributed Communication Networks","author":"Baran Paul","DOI":"10.7249\/P2626"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2003.1196112"},{"key":"e_1_2_1_17_1","volume-title":"Random search for hyper-parameter optimization. J. Mach. Learn. Res. 13 (Feb","author":"Bergstra James","year":"2012"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.25080\/Majora-92bf1922-003"},{"key":"e_1_2_1_19_1","volume-title":"Bernstein and Eric Newcomer","author":"Philip","year":"2009"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/567806.567807"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2133806.2133826"},{"key":"e_1_2_1_22_1","first-page":"993","article-title":"Latent Dirichlet allocation","author":"Blei David M.","year":"2003","journal-title":"J. Mach. Learn. Res. 3"},{"key":"e_1_2_1_23_1","volume-title":"Davide Del Testa","author":"Bojarski Mariusz","year":"2016"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-7908-2604-3_16"},{"key":"e_1_2_1_25_1","volume-title":"Random forests. Mach. Learn. 45, 1 (1","author":"Breiman Leo","year":"2001"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1111\/1467-9884.00117"},{"key":"e_1_2_1_27_1","unstructured":"Rajkumar Buyya et al. 1999. High Performance Cluster Computing: Architectures and Systems. Prentice Hall Upper SaddleRiver NJ 999.  Rajkumar Buyya et al. 1999. High Performance Cluster Computing: Architectures and Systems. Prentice Hall Upper SaddleRiver NJ 999."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1137\/140954362"},{"key":"e_1_2_1_29_1","first-page":"113","article-title":"Sibyl: A system for large scale supervised machine learning","volume":"1","author":"Canini K.","year":"2012","journal-title":"Tech. Talk"},{"key":"e_1_2_1_30_1","volume-title":"Revisiting distributed synchronous SGD. CoRR abs\/1604.00981","author":"Chen Jianmin","year":"2016"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7472805"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2644865.2541967"},{"key":"e_1_2_1_33_1","volume-title":"MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems. CoRR abs\/1512.01274","author":"Chen Tianqi","year":"2015"},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201914)","author":"Chilimbi Trishul","year":"2014"},{"key":"e_1_2_1_35_1","unstructured":"Fran\u00e7ois Chollet et al. 2015. Keras. Retrieved from https:\/\/keras.io\/.  Fran\u00e7ois Chollet et al. 2015. Keras. Retrieved from https:\/\/keras.io\/."},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems. 281--288","author":"Chu Cheng-Tao"},{"key":"e_1_2_1_37_1","volume-title":"Buchanan","author":"Clearwater Scott H.","year":"1989"},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the International Conference on Machine Learning. 1337--1345","author":"Coates Adam","year":"2013"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2018.03.032"},{"key":"e_1_2_1_40_1","volume-title":"Distributed Systems: Concepts and Design. Pearson Education.","author":"Coulouris George F.","year":"2005"},{"key":"e_1_2_1_41_1","volume-title":"Proceedings of the USENIX Annual Technical Conference. 37--48","author":"Cui Henggang","year":"2014"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1137\/S1052623497318992"},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the 25th International Conference on Neural Information Processing Systems","volume":"1","author":"Dean Jeffrey"},{"key":"e_1_2_1_44_1","volume-title":"Proceedings of the 6th Conference on Operating Systems Design 8 Implementation","volume":"6","author":"Dean Jeffrey","year":"2004"},{"key":"e_1_2_1_45_1","volume-title":"Adaptive subgradient methods for online learning and stochastic optimization. J. Mach. Learn. Res. 12 (July","author":"Duchi John","year":"2011"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/568522.568525"},{"key":"e_1_2_1_47_1","unstructured":"Facebook. 2017. Gloo. Retrieved from https:\/\/github.com\/facebookincubator\/gloo.  Facebook. 2017. Gloo. Retrieved from https:\/\/github.com\/facebookincubator\/gloo."},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2011.5981829"},{"key":"e_1_2_1_49_1","volume-title":"Anastasia Ailamaki, and Babak Falsafi.","author":"Ferdman Michael","year":"2012"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2697065"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.1972.5009071"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-45172-3_11"},{"key":"e_1_2_1_53_1","unstructured":"Ermias Gebremeskel. 2018. Analysis and comparison of distributed training techniques for deep neural networks in a dynamic environment. (2018).  Ermias Gebremeskel. 2018. Analysis and comparison of distributed training techniques for deep neural networks in a dynamic environment. (2018)."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1992.4.1.1"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/945445.945450"},{"key":"e_1_2_1_56_1","unstructured":"Andrew Gibiansky. 2017. Bringing HPC Techniques to Deep Learning. Retrieved from http:\/\/research.baidu.com\/bringing-hpc-techniques-deep-learning\/.  Andrew Gibiansky. 2017. Bringing HPC Techniques to Deep Learning. Retrieved from http:\/\/research.baidu.com\/bringing-hpc-techniques-deep-learning\/."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2015.04.061"},{"key":"e_1_2_1_58_1","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems","volume":"27","author":"Goodfellow Ian","year":"2014"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-25591-5_39"},{"key":"e_1_2_1_60_1","unstructured":"Google. 2017. Google Cloud TPU. Retrieved from https:\/\/cloud.google.com\/tpu.  Google. 2017. Google Cloud TPU. Retrieved from https:\/\/cloud.google.com\/tpu."},{"key":"e_1_2_1_61_1","volume-title":"large minibatch SGD: Training ImageNet in 1 hour. CoRR abs\/1706.02677","author":"Goyal Priya","year":"2017"},{"key":"e_1_2_1_62_1","volume-title":"Using MPI: Portable Parallel Programming with the Message-passing Interface","author":"Gropp William D."},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/bts360"},{"key":"e_1_2_1_64_1","volume-title":"Breeze: Numerical Processing Library for Scala.","author":"Hall D.","year":"2009"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.14778\/2777598.2777604"},{"key":"e_1_2_1_66_1","volume-title":"Proceedings of the 4th Workshop on General Purpose Processing on Graphics Processing Units. 3 (Mar. 2011)","author":"Han Tianyi David"},{"key":"e_1_2_1_67_1","unstructured":"Elmar Hau\u00dfmann. 2018. Accelerating I\/O Bound Deep Learning on Shared Storage. Retrieved from https:\/\/blog.riseml.com\/accelerating-io-bound-deep-learning-e0e3f095fd0.  Elmar Hau\u00dfmann. 2018. Accelerating I\/O Bound Deep Learning on Shared Storage. Retrieved from https:\/\/blog.riseml.com\/accelerating-io-bound-deep-learning-e0e3f095fd0."},{"key":"e_1_2_1_68_1","volume-title":"Deep residual learning for image recognition. CoRR abs\/1512.03385","author":"He Kaiming","year":"2015"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.6028\/jres.049.044"},{"key":"e_1_2_1_70_1","volume-title":"Salakhutdinov","author":"Hinton Geoffrey E.","year":"2012"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133956.3134012"},{"key":"e_1_2_1_72_1","volume-title":"Proceedings of the 15th Conference on Uncertainty in Artificial Intelligence. Morgan Kaufmann Publishers Inc., 289--296","author":"Hofmann Thomas","year":"1999"},{"key":"e_1_2_1_73_1","volume-title":"Proceedings of the 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201917)","author":"Hsieh Kevin","year":"2017"},{"key":"e_1_2_1_74_1","unstructured":"IBM Cloud. 2018. IBM Watson Machine Learning. Retrieved from https:\/\/www.ibm.com\/cloud\/machine-learning.  IBM Cloud. 2018. IBM Watson Machine Learning. Retrieved from https:\/\/www.ibm.com\/cloud\/machine-learning."},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/1272996.1273005"},{"key":"e_1_2_1_76_1","unstructured":"Sylvain Jeaugey. 2017. NCCL 2.0. Retrieved from http:\/\/on-demand.gputechconf.com\/gtc\/2017\/presentation\/s7155-jeaugey-nccl.pdf.  Sylvain Jeaugey. 2017. NCCL 2.0. Retrieved from http:\/\/on-demand.gputechconf.com\/gtc\/2017\/presentation\/s7155-jeaugey-nccl.pdf."},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-77018-3_32"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654889"},{"key":"e_1_2_1_79_1","volume-title":"Mitchell","author":"Jordan Michael I.","year":"2015"},{"key":"e_1_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.5555\/1622737.1622748"},{"key":"e_1_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbankfin.2010.06.001"},{"key":"e_1_2_1_83_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2009.263"},{"key":"e_1_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126916"},{"key":"e_1_2_1_85_1","volume-title":"Kim","author":"Kwon Donghwoon","year":"2017"},{"key":"e_1_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.5555\/1324616.1324684"},{"key":"e_1_2_1_87_1","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-015-0032-1"},{"key":"e_1_2_1_88_1","volume-title":"Proceedings of the International Conference on Machine Learning","volume":"227","author":"Larochelle Hugo"},{"key":"e_1_2_1_89_1","doi-asserted-by":"publisher","DOI":"10.1145\/355841.355847"},{"key":"e_1_2_1_90_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2013.6639343"},{"key":"e_1_2_1_91_1","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. ACM, 8.","author":"Li Guanpeng"},{"key":"e_1_2_1_92_1","volume-title":"Proceedings of the 27th International Conference on Neural Information Processing Systems (NIPS\u201914)","volume":"1","author":"Li Mu","year":"2014"},{"key":"e_1_2_1_93_1","doi-asserted-by":"publisher","DOI":"10.1145\/2640087.2644155"},{"key":"e_1_2_1_94_1","volume-title":"Big Learning NIPS Workshop","volume":"6","author":"Li Mu","year":"2013"},{"key":"e_1_2_1_95_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF01589116"},{"key":"e_1_2_1_96_1","unstructured":"A. R. Mamidala G. Kollias C. Ward and F. Artico. 2018. MXNET-MPI: Embedding MPI parallelism in parameter server task model for scaling deep learning. ArXiv e-prints (Jan. 2018). arxiv:cs.DC\/1801.03855.  A. R. Mamidala G. Kollias C. Ward and F. Artico. 2018. MXNET-MPI: Embedding MPI parallelism in parameter server task model for scaling deep learning. ArXiv e-prints (Jan. 2018). arxiv:cs.DC\/1801.03855."},{"key":"e_1_2_1_97_1","unstructured":"H. Brendan McMahan Eider Moore Daniel Ramage and Blaise Ag\u00fcera y Arcas. 2016. Federated learning of deep networks using model averaging. (2016).  H. Brendan McMahan Eider Moore Daniel Ramage and Blaise Ag\u00fcera y Arcas. 2016. Federated learning of deep networks using model averaging. (2016)."},{"key":"e_1_2_1_98_1","doi-asserted-by":"publisher","DOI":"10.5555\/2946645.2946679"},{"key":"e_1_2_1_99_1","unstructured":"Cade Metz. 2018. Big bets on AI open a new frontier for chip start-ups too. The New York Times 14 Jan. (2018). Retrieved from https:\/\/www.nytimes.com\/2018\/01\/14\/technology\/artificial-intelligence-chip-start-ups.html.  Cade Metz. 2018. Big bets on AI open a new frontier for chip start-ups too. The New York Times 14 Jan. (2018). Retrieved from https:\/\/www.nytimes.com\/2018\/01\/14\/technology\/artificial-intelligence-chip-start-ups.html."},{"key":"e_1_2_1_100_1","doi-asserted-by":"publisher","DOI":"10.1109\/MPOT.2007.906096"},{"key":"e_1_2_1_101_1","unstructured":"Microsoft. 2018. Microsoft Azure Machine Learning. Retrieved from https:\/\/azure.microsoft.com\/en-us\/overview\/machine-learning\/.  Microsoft. 2018. Microsoft Azure Machine Learning. Retrieved from https:\/\/azure.microsoft.com\/en-us\/overview\/machine-learning\/."},{"key":"e_1_2_1_102_1","unstructured":"Microsoft Inc. 2015. Distributed Machine Learning Toolkit (DMTK). Retrieved from http:\/\/www.dmtk.io.  Microsoft Inc. 2015. Distributed Machine Learning Toolkit (DMTK). Retrieved from http:\/\/www.dmtk.io."},{"key":"e_1_2_1_103_1","doi-asserted-by":"publisher","DOI":"10.1109\/IRI.2013.6642536"},{"key":"e_1_2_1_104_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00500-014-1511-6"},{"key":"e_1_2_1_105_1","first-page":"1801","article-title":"Distributed algorithms for topic models","author":"Newman David","year":"2009","journal-title":"J. Mach. Learn. Res. 10"},{"key":"e_1_2_1_106_1","unstructured":"NVIDIA Corporation. 2015. NVIDIA Collective Communications Library (NCCL). Retrieved from https:\/\/developer.nvidia.com\/nccl.  NVIDIA Corporation. 2015. NVIDIA Collective Communications Library (NCCL). Retrieved from https:\/\/developer.nvidia.com\/nccl."},{"key":"e_1_2_1_107_1","unstructured":"NVIDIA Corporation. 2017. Nvidia Tesla V100. Retrieved from https:\/\/www.nvidia.com\/en-us\/data-center\/tesla-v100\/.  NVIDIA Corporation. 2017. Nvidia Tesla V100. Retrieved from https:\/\/www.nvidia.com\/en-us\/data-center\/tesla-v100\/."},{"key":"e_1_2_1_108_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2004.01.013"},{"key":"e_1_2_1_109_1","unstructured":"Andreas Olofsson. 2016. Epiphany-V: A 1024 processor 64-bit RISC system-on-chip. (2016).  Andreas Olofsson. 2016. Epiphany-V: A 1024 processor 64-bit RISC system-on-chip. (2016)."},{"key":"e_1_2_1_110_1","volume-title":"Kickstarting high-performance energy-efficient manycore architectures with epiphany. arXiv preprint arXiv:1412.5538","author":"Olofsson Andreas","year":"2014"},{"key":"e_1_2_1_111_1","volume-title":"Opitz and Richard Maclin","author":"David","year":"1999"},{"key":"e_1_2_1_112_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.2858"},{"key":"e_1_2_1_113_1","doi-asserted-by":"publisher","DOI":"10.1109\/TEVC.2008.925144"},{"key":"e_1_2_1_114_1","volume-title":"Chung","author":"Ovtcharov Kalin","year":"2015"},{"key":"e_1_2_1_115_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2012.04.026"},{"key":"e_1_2_1_116_1","volume-title":"Genomic big data hitting the storage bottleneck. EMBnet. J. 24","author":"Papageorgiou Louis","year":"2018"},{"key":"e_1_2_1_117_1","unstructured":"Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in PyTorch. (2017).  Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in PyTorch. (2017)."},{"key":"e_1_2_1_118_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2008.09.002"},{"key":"e_1_2_1_119_1","doi-asserted-by":"publisher","DOI":"10.1007\/s13748-012-0035-5"},{"key":"e_1_2_1_120_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2005.06.076"},{"key":"e_1_2_1_121_1","volume-title":"Machine learning and cloud computing: Survey of distributed and SaaS solutions. arXiv preprint arXiv:1603.08767","author":"Pop Daniel","year":"2016"},{"key":"e_1_2_1_122_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1009876119989"},{"key":"e_1_2_1_123_1","doi-asserted-by":"publisher","DOI":"10.1186\/s13634-016-0355-x"},{"key":"e_1_2_1_124_1","volume-title":"Proceedings of the TeraGrid Conference. 12--15","author":"Raicu Ioan","year":"2006"},{"key":"e_1_2_1_125_1","volume-title":"Proceedings of the 26th International Conference on Machine Learning. ACM, 873--880","author":"Raina Rajat"},{"key":"e_1_2_1_126_1","volume-title":"AVX-512 instructions","author":"Reinders James","year":"2013"},{"key":"e_1_2_1_127_1","doi-asserted-by":"publisher","DOI":"10.5555\/2946645.3007028"},{"key":"e_1_2_1_128_1","volume-title":"An in-depth look at Google\u2019s first Tensor Processing Unit (TPU). Google Cloud Big Data Mach. Learn. Blog 12","author":"Sato Kaz","year":"2017"},{"key":"e_1_2_1_129_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2945397"},{"key":"e_1_2_1_130_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2014-274"},{"key":"e_1_2_1_131_1","volume-title":"Horovod: Fast and easy distributed deep learning in TensorFlow.","author":"Sergeev Alexander","year":"2018"},{"key":"e_1_2_1_132_1","volume-title":"Proceedings of the 21st International Conference on Pattern Recognition (ICPR\u201912)","author":"Sermanet Pierre","year":"2012"},{"key":"e_1_2_1_133_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2011.6033589"},{"key":"e_1_2_1_134_1","unstructured":"Amazon Web Services. 2016. Introducing Amazon EC2 P2 Instances the Largest GPU-Powered Virtual Machine in the Cloud. Retrieved from https:\/\/aws.amazon.com\/about-aws\/whats-new\/2016\/09\/introducing-amazon-ec2-p2-instances-the-largest-gpu-powered-virtual-machine-in-the-cloud\/.  Amazon Web Services. 2016. Introducing Amazon EC2 P2 Instances the Largest GPU-Powered Virtual Machine in the Cloud. Retrieved from https:\/\/aws.amazon.com\/about-aws\/whats-new\/2016\/09\/introducing-amazon-ec2-p2-instances-the-largest-gpu-powered-virtual-machine-in-the-cloud\/."},{"key":"e_1_2_1_135_1","unstructured":"Amazon Web Services. 2017. Amazon EC2 F1 Instances. Retrieved from https:\/\/aws.amazon.com\/ec2\/instance-types\/f1\/.  Amazon Web Services. 2017. Amazon EC2 F1 Instances. Retrieved from https:\/\/aws.amazon.com\/ec2\/instance-types\/f1\/."},{"key":"e_1_2_1_136_1","unstructured":"Shai Shalev-Shwartz and Tong Zhang. 2013. Stochastic dual coordinate ascent methods for regularized loss minimization. (2013).  Shai Shalev-Shwartz and Tong Zhang. 2013. Stochastic dual coordinate ascent methods for regularized loss minimization. (2013)."},{"key":"e_1_2_1_137_1","volume-title":"Proceedings of the 21st ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, 2323--2324","author":"James"},{"key":"e_1_2_1_138_1","volume-title":"Performance modeling and evaluation of distributed deep learning frameworks on GPUs. CoRR abs\/1711.05979","author":"Shi Shaohuai","year":"2017"},{"key":"e_1_2_1_139_1","doi-asserted-by":"publisher","DOI":"10.1109\/CCBD.2016.029"},{"key":"e_1_2_1_140_1","volume-title":"Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security. ACM, 1310--1321","author":"Shokri Reza","year":"2015"},{"key":"e_1_2_1_141_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSST.2010.5496972"},{"key":"e_1_2_1_142_1","volume-title":"Proceedings of the IEEE 27th International Workshop on Machine Learning for Signal Processing (MLSP\u201917)","author":"Simm J.","year":"2017"},{"key":"e_1_2_1_143_1","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-014-0008-6"},{"key":"e_1_2_1_144_1","volume-title":"Application-specific Integrated Circuits","author":"Sebastian Smith Michael John"},{"key":"e_1_2_1_145_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920931"},{"key":"e_1_2_1_146_1","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems","volume":"25","author":"Snoek Jasper"},{"key":"e_1_2_1_147_1","volume-title":"Going deeper with convolutions. CoRR abs\/1409.4842","author":"Szegedy Christian","year":"2014"},{"key":"e_1_2_1_148_1","volume-title":"Rethinking the inception architecture for computer vision. CoRR abs\/1512.00567","author":"Szegedy Christian","year":"2015"},{"key":"e_1_2_1_149_1","volume-title":"Proceedings of the International Conference on Machine Learning (ICML\u201913)","author":"Tak\u00e1c Martin","year":"2013"},{"key":"e_1_2_1_150_1","unstructured":"The Khronos Group. 2018. Neural Network Exchange Format (NNEF). Retrieved from https:\/\/www.khronos.org\/registry\/NNEF\/specs\/1.0\/nnef-1.0.pdf.  The Khronos Group. 2018. Neural Network Exchange Format (NNEF). Retrieved from https:\/\/www.khronos.org\/registry\/NNEF\/specs\/1.0\/nnef-1.0.pdf."},{"key":"e_1_2_1_151_1","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems.","author":"Tsianos K. I."},{"key":"e_1_2_1_152_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2012.11.001"},{"key":"e_1_2_1_153_1","volume-title":"No peek: A survey of private distributed deep learning. arXiv preprint arXiv:1812.03288","author":"Vepakomma Praneeth","year":"2018"},{"key":"e_1_2_1_154_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2012.06.011"},{"key":"e_1_2_1_155_1","volume-title":"Xing","author":"Wei Jinliang","year":"2015"},{"key":"e_1_2_1_156_1","doi-asserted-by":"publisher","DOI":"10.5555\/1622620.1622633"},{"key":"e_1_2_1_157_1","doi-asserted-by":"publisher","DOI":"10.1162\/evco.1995.3.2.149"},{"key":"e_1_2_1_158_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-015-0892-3"},{"key":"e_1_2_1_159_1","unstructured":"P. Xie J. K. Kim Y. Zhou Q. Ho A. Kumar Y. Yu and E. Xing. 2015. Distributed machine learning via sufficient factor broadcasting. (2015).  P. Xie J. K. Kim Y. Zhou Q. Ho A. Kumar Y. Yu and E. Xing. 2015. Distributed machine learning via sufficient factor broadcasting. (2015)."},{"key":"e_1_2_1_160_1","volume-title":"Petuum: A new platform for distributed machine learning on big data. ArXiv e-prints (Dec.","author":"Xing E. P.","year":"2013"},{"key":"e_1_2_1_161_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.ENG.2016.02.008"},{"key":"e_1_2_1_162_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2016.25"},{"key":"e_1_2_1_163_1","doi-asserted-by":"publisher","DOI":"10.1145\/2987550.2987576"},{"key":"e_1_2_1_164_1","volume-title":"Learning to compose words into sentences with reinforcement learning. CoRR abs\/1611.09100","author":"Yogatama Dani","year":"2016"},{"key":"e_1_2_1_165_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126912"},{"key":"e_1_2_1_166_1","doi-asserted-by":"publisher","DOI":"10.1145\/566340.566342"},{"key":"e_1_2_1_167_1","volume-title":"Proceedings of the 9th USENIX Conference on Networked Systems Design and Implementation. USENIX Association, 2--2.","author":"Zaharia Matei","year":"2012"},{"key":"e_1_2_1_168_1","volume-title":"Proceedings of the 2nd USENIX Conference on Hot Topics in Cloud Computing (HotCloud\u201910)","author":"Zaharia Matei","year":"2010"},{"key":"e_1_2_1_169_1","doi-asserted-by":"publisher","DOI":"10.1145\/2934664"},{"key":"e_1_2_1_170_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10799-012-0124-y"},{"key":"e_1_2_1_171_1","volume-title":"Xing","author":"Zhang Hao","year":"2017"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3377454","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3377454","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:33:18Z","timestamp":1750199598000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3377454"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,20]]},"references-count":171,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,3,31]]}},"alternative-id":["10.1145\/3377454"],"URL":"https:\/\/doi.org\/10.1145\/3377454","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,3,20]]},"assertion":[{"value":"2018-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-03-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}