{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T15:43:15Z","timestamp":1782834195333,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":38,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,3,25]],"date-time":"2019-03-25T00:00:00Z","timestamp":1553472000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,3,25]]},"DOI":"10.1145\/3302424.3303975","type":"proceedings-article","created":{"date-parts":[[2019,3,22]],"date-time":"2019-03-22T13:10:03Z","timestamp":1553260203000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":51,"title":["Fast Distributed Deep Learning over RDMA"],"prefix":"10.1145","author":[{"given":"Jilong","family":"Xue","sequence":"first","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Youshan","family":"Miao","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cheng","family":"Chen","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ming","family":"Wu","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lintao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lidong","family":"Zhou","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,3,25]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Retrieved in 2017. CUDA Driver API. http:\/\/docs.nvidia.com\/cuda\/cuda-driver-api.  Retrieved in 2017. CUDA Driver API. http:\/\/docs.nvidia.com\/cuda\/cuda-driver-api."},{"key":"e_1_3_2_1_2_1","unstructured":"Retrieved in 2017. gRPC - An RPC library and framework. https:\/\/github.com\/grpc\/grpc.  Retrieved in 2017. gRPC - An RPC library and framework. https:\/\/github.com\/grpc\/grpc."},{"key":"e_1_3_2_1_3_1","unstructured":"Retrieved in 2017. TensorFlow. https:\/\/github.com\/tensorflow\/tensorflow\/tree\/r1.2.  Retrieved in 2017. TensorFlow. https:\/\/github.com\/tensorflow\/tensorflow\/tree\/r1.2."},{"key":"e_1_3_2_1_4_1","unstructured":"Retrieved in 2017. The Apache Thrift. http:\/\/thrift.apache.org.  Retrieved in 2017. The Apache Thrift. http:\/\/thrift.apache.org."},{"key":"e_1_3_2_1_5_1","volume-title":"WTM'15 Machine Translation Dataset. http:\/\/www.statmt.org\/wmt15","author":"Retrieved","year":"2017","unstructured":"Retrieved in 2017 . WTM'15 Machine Translation Dataset. http:\/\/www.statmt.org\/wmt15 . Retrieved in 2017. WTM'15 Machine Translation Dataset. http:\/\/www.statmt.org\/wmt15."},{"key":"e_1_3_2_1_6_1","unstructured":"Retrieved in 2017. ZeroMQ. http:\/\/zeromq.org\/.  Retrieved in 2017. ZeroMQ. http:\/\/zeromq.org\/."},{"key":"e_1_3_2_1_7_1","volume-title":"TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16)","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi , Paul Barham , Jianmin Chen , Zhifeng Chen , Andy Davis , Jeffrey Dean , Matthieu Devin , Sanjay Ghemawat , Geoffrey Irving , Michael Isard , Manjunath Kudlur , Josh Levenberg , Rajat Monga , Sherry Moore , Derek G. Murray , Benoit Steiner , Paul Tucker , Vijay Vasudevan , Pete Warden , Martin Wicke , Yuan Yu , and Xiaoqiang Zheng . 2016 . TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16) . USENIX Association, GA, 265--283. https:\/\/www.usenix.org\/conference\/osdi16\/technical-sessions\/presentation\/abadi Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). USENIX Association, GA, 265--283. https:\/\/www.usenix.org\/conference\/osdi16\/technical-sessions\/presentation\/abadi"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2080.357392"},{"key":"e_1_3_2_1_9_1","volume-title":"MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. In NIPS Workshop on Machine Learning Systems (LearningSys).","author":"Chen Tianqi","year":"2016","unstructured":"Tianqi Chen , Mu Li , Yutian Li , Min Lin , Naiyan Wang , Minjie Wang , Tianjun Xiao , Bing Xu , Chiyuan Zhang , and Zheng Zhang . 2016 . MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. In NIPS Workshop on Machine Learning Systems (LearningSys). Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. 2016. MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. In NIPS Workshop on Machine Learning Systems (LearningSys)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2901318.2901349"},{"key":"e_1_3_2_1_11_1","volume-title":"arXiv:1606.07792","author":"Cheng Heng-Tze","year":"2016","unstructured":"Heng-Tze Cheng , Levent Koc , Jeremiah Harmsen , Tal Shaked , Tushar Chandra , Hrishi Aradhye , Glen Anderson , Greg Corrado , Wei Chai , Mustafa Ispir , Rohan Anil , Zakaria Haque , Lichan Hong , Vihan Jain , Xiaobing Liu , and Hemal Shah . 2016. Wide & Deep Learning for Recommender Systems . arXiv:1606.07792 ( 2016 ). http:\/\/arxiv.org\/abs\/1606.07792 Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, Rohan Anil, Zakaria Haque, Lichan Hong, Vihan Jain, Xiaobing Liu, and Hemal Shah. 2016. Wide & Deep Learning for Recommender Systems. arXiv:1606.07792 (2016). http:\/\/arxiv.org\/abs\/1606.07792"},{"key":"e_1_3_2_1_12_1","volume-title":"11th USENIX Symposium on Operating Systems Design and Implementation (OSDI'14)","author":"Chilimbi Trishul","year":"2014","unstructured":"Trishul Chilimbi , Yutaka Suzue , Johnson Apacible , and Karthik Kalyanaraman . 2014 . Project Adam: Building an Efficient and Scalable Deep Learning Training System . In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI'14) . USENIX. https:\/\/www.usenix.org\/conference\/osdi14\/technical-sessions\/presentation\/chilimbi Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman. 2014. Project Adam: Building an Efficient and Scalable Deep Learning Training System. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI'14). USENIX. https:\/\/www.usenix.org\/conference\/osdi14\/technical-sessions\/presentation\/chilimbi"},{"key":"e_1_3_2_1_13_1","volume-title":"Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. CoRR abs\/1406.1078","author":"Cho Kyunghyun","year":"2014","unstructured":"Kyunghyun Cho , Bart van Merrienboer , \u00c7aglar G\u00fcl\u00e7ehre , Fethi Bougares , Holger Schwenk , and Yoshua Bengio . 2014. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. CoRR abs\/1406.1078 ( 2014 ). http:\/\/arxiv.org\/abs\/1406.1078 Kyunghyun Cho, Bart van Merrienboer, \u00c7aglar G\u00fcl\u00e7ehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. CoRR abs\/1406.1078 (2014). http:\/\/arxiv.org\/abs\/1406.1078"},{"key":"e_1_3_2_1_14_1","volume-title":"NIPS Workshop.","author":"Collobert Ronan","year":"2011","unstructured":"Ronan Collobert , Koray Kavukcuoglu , and Cl\u00e9ment Farabet . 2011 . Torch7: A Matlab-like Environment for Machine Learning. In BigLearn , NIPS Workshop. Ronan Collobert, Koray Kavukcuoglu, and Cl\u00e9ment Farabet. 2011. Torch7: A Matlab-like Environment for Machine Learning. In BigLearn, NIPS Workshop."},{"key":"e_1_3_2_1_15_1","volume-title":"Ng","author":"Dean Jeffrey","year":"2012","unstructured":"Jeffrey Dean , Greg Corrado , Rajat Monga , Kai Chen , Matthieu Devin , Mark Mao , Marc'aurelio Ranzato , Andrew Senior , Paul Tucker , Ke Yang , Quoc V. Le , and Andrew Y . Ng . 2012 . Large Scale Distributed Deep Networks. In Advances in Neural Information Processing Systems 25. Curran Associates, Inc . http:\/\/papers.nips.cc\/paper\/4687-large-scale-distributed-deep-networks.pdf Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc'aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, Quoc V. Le, and Andrew Y. Ng. 2012. Large Scale Distributed Deep Networks. In Advances in Neural Information Processing Systems 25. Curran Associates, Inc. http:\/\/papers.nips.cc\/paper\/4687-large-scale-distributed-deep-networks.pdf"},{"key":"e_1_3_2_1_16_1","volume-title":"FaRM: Fast Remote Memory. In 11th USENIX Symposium on Networked Systems Design and Implementation (NSDI'14)","author":"Dragojevi\u0107 Aleksandar","year":"2014","unstructured":"Aleksandar Dragojevi\u0107 , Dushyanth Narayanan , Miguel Castro , and Orion Hodson . 2014 . FaRM: Fast Remote Memory. In 11th USENIX Symposium on Networked Systems Design and Implementation (NSDI'14) . USENIX. https:\/\/www.usenix.org\/conference\/nsdi14\/technical-sessions\/dragojevi{\u0107} Aleksandar Dragojevi\u0107, Dushyanth Narayanan, Miguel Castro, and Orion Hodson. 2014. FaRM: Fast Remote Memory. In 11th USENIX Symposium on Networked Systems Design and Implementation (NSDI'14). USENIX. https:\/\/www.usenix.org\/conference\/nsdi14\/technical-sessions\/dragojevi{\u0107}"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2815400.2815425"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2619239.2626299"},{"key":"e_1_3_2_1_20_1","volume-title":"Design Guidelines for High Performance RDMA Systems. In 2016 USENIX Annual Technical Conference (USENIX ATC 16)","author":"Kalia Anuj","unstructured":"Anuj Kalia , Michael Kaminsky , and David G. Andersen . 2016 . Design Guidelines for High Performance RDMA Systems. In 2016 USENIX Annual Technical Conference (USENIX ATC 16) . USENIX Association, Denver, CO, 437--450. https:\/\/www.usenix.org\/conference\/atc16\/technical-sessions\/presentation\/kalia Anuj Kalia, Michael Kaminsky, and David G. Andersen. 2016. Design Guidelines for High Performance RDMA Systems. In 2016 USENIX Annual Technical Conference (USENIX ATC 16). USENIX Association, Denver, CO, 437--450. https:\/\/www.usenix.org\/conference\/atc16\/technical-sessions\/presentation\/kalia"},{"key":"e_1_3_2_1_21_1","volume-title":"Scalable and Simple Distributed Transactions with Two-Sided (RDMA) Datagram RPCs. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16)","author":"Kalia Anuj","unstructured":"Anuj Kalia , Michael Kaminsky , and David G. Andersen . 2016. FaSST: Fast , Scalable and Simple Distributed Transactions with Two-Sided (RDMA) Datagram RPCs. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16) . USENIX Association, GA, 185--201. https:\/\/www.usenix.org\/conference\/osdi16\/technical-sessions\/presentation\/kalia Anuj Kalia, Michael Kaminsky, and David G. Andersen. 2016. FaSST: Fast, Scalable and Simple Distributed Transactions with Two-Sided (RDMA) Datagram RPCs. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). USENIX Association, GA, 185--201. https:\/\/www.usenix.org\/conference\/osdi16\/technical-sessions\/presentation\/kalia"},{"key":"e_1_3_2_1_22_1","volume-title":"GPUnet: Networking Abstractions for GPU Programs. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14)","author":"Kim Sangman","year":"2014","unstructured":"Sangman Kim , Seonggu Huh , Xinya Zhang , Yige Hu , Amir Wated , Emmett Witchel , and Mark Silberstein . 2014 . GPUnet: Networking Abstractions for GPU Programs. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14) . USENIX Association, Broomfield, CO, 201--216. https:\/\/www.usenix.org\/conference\/osdi14\/technical-sessions\/presentation\/kim Sangman Kim, Seonggu Huh, Xinya Zhang, Yige Hu, Amir Wated, Emmett Witchel, and Mark Silberstein. 2014. GPUnet: Networking Abstractions for GPU Programs. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14). USENIX Association, Broomfield, CO, 201--216. https:\/\/www.usenix.org\/conference\/osdi14\/technical-sessions\/presentation\/kim"},{"key":"e_1_3_2_1_24_1","volume-title":"Advances in Neural Information Processing Systems 25","author":"Krizhevsky Alex","unstructured":"Alex Krizhevsky , Ilya Sutskever , and Geoffrey E Hinton . 2012. ImageNet Classification with Deep Convolutional Neural Networks . In Advances in Neural Information Processing Systems 25 , F. Pereira, C.J. C. Burges, L. Bottou, and K. Q. Weinberger (Eds.). Curran Associates, Inc. , 1097--1105. http:\/\/papers.nips.cc\/paper\/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In Advances in Neural Information Processing Systems 25, F. Pereira, C.J. C. Burges, L. Bottou, and K. Q. Weinberger (Eds.). Curran Associates, Inc., 1097--1105. http:\/\/papers.nips.cc\/paper\/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf"},{"key":"e_1_3_2_1_25_1","volume-title":"Scaling Distributed Machine Learning with the Parameter Server. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI'14)","author":"Li Mu","year":"2014","unstructured":"Mu Li , David G. Andersen , Jun Woo Park , Alexander J. Smola , Amr Ahmed , Vanja Josifovski , James Long , Eugene J. Shekita , and Bor-Yiing Su . 2014 . Scaling Distributed Machine Learning with the Parameter Server. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI'14) . USENIX. https:\/\/www.usenix.org\/conference\/osdi14\/technical-sessions\/presentation\/limu Mu Li, David G. Andersen, Jun Woo Park, Alexander J. Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J. Shekita, and Bor-Yiing Su. 2014. Scaling Distributed Machine Learning with the Parameter Server. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI'14). USENIX. https:\/\/www.usenix.org\/conference\/osdi14\/technical-sessions\/presentation\/limu"},{"key":"e_1_3_2_1_26_1","volume-title":"Proceedings of the 2013 USENIX Conference on Annual Technical Conference (USENIX ATC'13). USENIX, 12","author":"Mitchell Christopher","year":"2013","unstructured":"Christopher Mitchell , Yifeng Geng , and Jinyang Li . 2013 . Using One-sided RDMA Reads to Build a Fast, CPU-efficient Key-value Store . In Proceedings of the 2013 USENIX Conference on Annual Technical Conference (USENIX ATC'13). USENIX, 12 . http:\/\/dl.acm.org\/citation.cfm?id=2535461.2535475 Christopher Mitchell, Yifeng Geng, and Jinyang Li. 2013. Using One-sided RDMA Reads to Build a Fast, CPU-efficient Key-value Store. In Proceedings of the 2013 USENIX Conference on Annual Technical Conference (USENIX ATC'13). USENIX, 12. http:\/\/dl.acm.org\/citation.cfm?id=2535461.2535475"},{"key":"e_1_3_2_1_27_1","volume-title":"Latency-Tolerant Software Distributed Shared Memory. In 2015 USENIX Annual Technical Conference (USENIX ATC'15)","author":"Nelson Jacob","year":"2015","unstructured":"Jacob Nelson , Brandon Holt , Brandon Myers , Preston Briggs , Luis Ceze , Simon Kahan , and Mark Oskin . 2015 . Latency-Tolerant Software Distributed Shared Memory. In 2015 USENIX Annual Technical Conference (USENIX ATC'15) . USENIX. https:\/\/www.usenix.org\/conference\/atc15\/technical-session\/presentation\/nelson Jacob Nelson, Brandon Holt, Brandon Myers, Preston Briggs, Luis Ceze, Simon Kahan, and Mark Oskin. 2015. Latency-Tolerant Software Distributed Shared Memory. In 2015 USENIX Annual Technical Conference (USENIX ATC'15). USENIX. https:\/\/www.usenix.org\/conference\/atc15\/technical-session\/presentation\/nelson"},{"key":"e_1_3_2_1_28_1","volume-title":"On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. In 5th International Conference on Learning Representations (ICLR 17)","author":"Mikhail Smelyanskiy Nitish Shirish Jorge Nocedal","unstructured":"Jorge Nocedal Mikhail Smelyanskiy Nitish Shirish Keskar, Dheevatsa Mudigere and Ping Tak Peter Tang . {n. d.}. On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. In 5th International Conference on Learning Representations (ICLR 17) . Jorge Nocedal Mikhail Smelyanskiy Nitish Shirish Keskar, Dheevatsa Mudigere and Ping Tak Peter Tang. {n. d.}. On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. In 5th International Conference on Learning Representations (ICLR 17)."},{"key":"e_1_3_2_1_29_1","unstructured":"K. Simonyan and A. Zisserman. 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. CoRR abs\/1409.1556 (2014).  K. Simonyan and A. Zisserman. 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. CoRR abs\/1409.1556 (2014)."},{"key":"e_1_3_2_1_30_1","volume-title":"Proceedings of the 27th International Conference on Neural Information Processing Systems (NIPS'14)","author":"Sutskever Ilya","unstructured":"Ilya Sutskever , Oriol Vinyals , and Quoc V. Le . 2014. Sequence to Sequence Learning with Neural Networks . In Proceedings of the 27th International Conference on Neural Information Processing Systems (NIPS'14) . MIT Press, Cambridge, MA, USA, 3104--3112. http:\/\/dl.acm.org\/citation.cfm?id=2969033.2969173 Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to Sequence Learning with Neural Networks. In Proceedings of the 27th International Conference on Neural Information Processing Systems (NIPS'14). MIT Press, Cambridge, MA, USA, 3104--3112. http:\/\/dl.acm.org\/citation.cfm?id=2969033.2969173"},{"key":"e_1_3_2_1_31_1","volume-title":"Rethinking the Inception Architecture for Computer Vision. CoRR abs\/1512.00567","author":"Szegedy Christian","year":"2015","unstructured":"Christian Szegedy , Vincent Vanhoucke , Sergey Ioffe , Jonathon Shlens , and Zbigniew Wojna . 2015. Rethinking the Inception Architecture for Computer Vision. CoRR abs\/1512.00567 ( 2015 ). http:\/\/arxiv.org\/abs\/1512.00567 Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2015. Rethinking the Inception Architecture for Computer Vision. CoRR abs\/1512.00567 (2015). http:\/\/arxiv.org\/abs\/1512.00567"},{"key":"e_1_3_2_1_32_1","volume-title":"Nessie: A Decoupled, Client-Driven, Key-Value Store using RDMA. Technical Report CS-2015-09","author":"Szepesi Tyler","year":"2015","unstructured":"Tyler Szepesi , Benjamin Cassell , Bernard Wong , Tim Brecht , and Xiaoyi Liu . 2015 . Nessie: A Decoupled, Client-Driven, Key-Value Store using RDMA. Technical Report CS-2015-09 , University of Waterloo , David R. Cheriton School of Computer Science. (2015). Tyler Szepesi, Benjamin Cassell, Bernard Wong, Tim Brecht, and Xiaoyi Liu. 2015. Nessie: A Decoupled, Client-Driven, Key-Value Store using RDMA. Technical Report CS-2015-09, University of Waterloo, David R. Cheriton School of Computer Science. (2015)."},{"key":"e_1_3_2_1_33_1","volume-title":"Theano: A Python framework for fast computation of mathematical expressions. arXiv e-prints abs\/1605.02688 (May","author":"Team Theano Development","year":"2016","unstructured":"Theano Development Team . 2016 . Theano: A Python framework for fast computation of mathematical expressions. arXiv e-prints abs\/1605.02688 (May 2016). http:\/\/arxiv.org\/abs\/1605.02688 Theano Development Team. 2016. Theano: A Python framework for fast computation of mathematical expressions. arXiv e-prints abs\/1605.02688 (May 2016). http:\/\/arxiv.org\/abs\/1605.02688"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2588555.2595641"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2815400.2815419"},{"key":"e_1_3_2_1_36_1","volume-title":"Hadoop: The Definitive Guide","author":"White Tom","year":"2009","unstructured":"Tom White . 2009 . Hadoop: The Definitive Guide ( 1 st ed.). O'Reilly Media, Inc. Tom White. 2009. Hadoop: The Definitive Guide (1st ed.). O'Reilly Media, Inc.","edition":"1"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2806777.2806849"},{"key":"e_1_3_2_1_38_1","volume-title":"14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17)","author":"Xiao Wencong","year":"2017","unstructured":"Wencong Xiao , Jilong Xue , Youshan Miao , Zhen Li , Cheng Chen , Ming Wu , Wei Li , and Lidong Zhou . 2017 . TuX2: Distributed Graph Computation for Machine Learning . In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) . USENIX Association, Boston, MA, 669--682. https:\/\/www.usenix.org\/conference\/nsdi17\/technical-sessions\/presentation\/xiao Wencong Xiao, Jilong Xue, Youshan Miao, Zhen Li, Cheng Chen, Ming Wu, Wei Li, and Lidong Zhou. 2017. TuX2: Distributed Graph Computation for Machine Learning. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). USENIX Association, Boston, MA, 669--682. https:\/\/www.usenix.org\/conference\/nsdi17\/technical-sessions\/presentation\/xiao"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2783258.2783323"}],"event":{"name":"EuroSys '19: Fourteenth EuroSys Conference 2019","location":"Dresden Germany","acronym":"EuroSys '19","sponsor":["SIGOPS ACM Special Interest Group on Operating Systems"]},"container-title":["Proceedings of the Fourteenth EuroSys Conference 2019"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3302424.3303975","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3302424.3303975","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:01:49Z","timestamp":1750208509000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3302424.3303975"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,3,25]]},"references-count":38,"alternative-id":["10.1145\/3302424.3303975","10.1145\/3302424"],"URL":"https:\/\/doi.org\/10.1145\/3302424.3303975","relation":{},"subject":[],"published":{"date-parts":[[2019,3,25]]},"assertion":[{"value":"2019-03-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}