{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T02:24:18Z","timestamp":1783650258486,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":32,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,8,29]],"date-time":"2022-08-29T00:00:00Z","timestamp":1661731200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Natural Science Foundation of Hunan Province, China","award":["2021JJ30867"],"award-info":[{"award-number":["2021JJ30867"]}]},{"name":"Key Research and Development Program of Hunan","award":["2022WK2005"],"award-info":[{"award-number":["2022WK2005"]}]},{"name":"National Natural Science Foundation of China","award":["62132022, 61872387"],"award-info":[{"award-number":["62132022, 61872387"]}]},{"name":"Postgraduate Scientific Research Innovation Project of Hunan Province","award":["QL20210059"],"award-info":[{"award-number":["QL20210059"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,8,29]]},"DOI":"10.1145\/3545008.3545024","type":"proceedings-article","created":{"date-parts":[[2023,1,15]],"date-time":"2023-01-15T01:04:08Z","timestamp":1673744648000},"page":"1-11","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["HSP: Hybrid Synchronous Parallelism for Fast Distributed Deep Learning"],"prefix":"10.1145","author":[{"given":"Yijun","family":"Li","sequence":"first","affiliation":[{"name":"Central South University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiawei","family":"Huang","sequence":"additional","affiliation":[{"name":"Central South University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhaoyi","family":"Li","sequence":"additional","affiliation":[{"name":"Central South University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shengwen","family":"Zhou","sequence":"additional","affiliation":[{"name":"Central South University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wanchun","family":"Jiang","sequence":"additional","affiliation":[{"name":"Central South University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianxin","family":"Wang","sequence":"additional","affiliation":[{"name":"Central South University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,1,13]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"e_1_3_2_1_2_1","volume-title":"Natural language processing (almost) from scratch. Journal of machine learning research, 12:2493\u20132537","author":"Collobert Ronan","year":"2011","unstructured":"[2] Ronan Collobert, Jason Weston, L\u00e9on Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. Natural language processing (almost) from scratch. Journal of machine learning research, 12:2493\u20132537, 2011."},{"key":"e_1_3_2_1_3_1","volume-title":"et\u00a0al. Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups","author":"Hinton Geoffrey","year":"2012","unstructured":"[3] Geoffrey Hinton, Li\u00a0Deng, Dong Yu, George\u00a0E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara\u00a0N Sainath, et\u00a0al. Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal processing magazine, 29(6):82\u201397, 2012."},{"key":"e_1_3_2_1_4_1","first-page":"1223","article-title":"Large scale distributed deep networks","volume":"25","author":"Dean Jeffrey","year":"2012","unstructured":"[4] Jeffrey Dean, Gregory\u00a0S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc\u00a0V. Le, Mark\u00a0Z. Mao, Marc\u2019Aurelio Ranzato, Andrew\u00a0W. Senior, Paul\u00a0A. Tucker, Ke\u00a0Yang, and A.\u00a0Ng. Large scale distributed deep networks. Advances in Neural Information Processing Systems, 25:1223\u20131231, 2012.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_5_1","volume-title":"Proc. USENIX OSDI","author":"Aurick Qiao","year":"2021","unstructured":"[5] Qiao Aurick, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, and Eric P. Xing. Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning. In Proc. USENIX OSDI, 2021."},{"key":"e_1_3_2_1_6_1","first-page":"598","volume-title":"Proc. USENIX OSDI","author":"Andersen G","year":"2014","unstructured":"[6] Mu\u00a0Li, David\u00a0G Andersen, Jun\u00a0Woo Park, Alexander\u00a0J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene\u00a0J Shekita, and Bor-Yiing Su. Scaling distributed machine learning with the parameter server. In Proc. USENIX OSDI, pages 583\u2013598, 2014."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1327452.1327492"},{"key":"e_1_3_2_1_8_1","first-page":"15","volume-title":"Proc. ACM SOSP","author":"Narayanan Deepak","year":"2019","unstructured":"[8] Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil\u00a0R Devanur, Gregory\u00a0R Ganger, Phillip\u00a0B Gibbons, and Matei Zaharia. Pipedream: generalized pipeline parallelism for dnn training. In Proc. ACM SOSP, pages 1\u201315, 2019."},{"key":"e_1_3_2_1_9_1","first-page":"193","volume-title":"Proc. USENIX ATC","author":"Zhang Hao","year":"2017","unstructured":"[9] Hao Zhang, Zeyu Zheng, Shizhen Xu, Wei Dai, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, and Eric\u00a0P Xing. Poseidon: An efficient communication architecture for distributed deep learning on gpu clusters. In Proc. USENIX ATC, pages 181\u2013193, 2017."},{"key":"e_1_3_2_1_10_1","first-page":"582","volume-title":"Proc. USENIX OSDI","author":"Chilimbi Trishul","year":"2014","unstructured":"[10] Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman. Project adam: Building an efficient and scalable deep learning training system. In Proc. USENIX OSDI, pages 571\u2013582, 2014."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3035933"},{"key":"e_1_3_2_1_12_1","first-page":"1231","volume-title":"Proc. NIPS","author":"Ho Qirong","year":"2013","unstructured":"[12] Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin\u00a0Kyu Kim, Phillip\u00a0B Gibbons, Garth\u00a0A Gibson, Greg Ganger, and Eric\u00a0P Xing. More effective distributed ml via a stale synchronous parallel parameter server. In Proc. NIPS, pages 1223\u20131231, 2013."},{"key":"e_1_3_2_1_13_1","first-page":"538","volume-title":"Proc. IEEE ICDCS","author":"Li Shijian","year":"2021","unstructured":"[13] Shijian Li, Oren Mangoubi, Lijie Xu, and Tian Guo. Sync-switch: Hybrid parameter synchronization for distributed deep learning. In Proc. IEEE ICDCS, pages 528\u2013538, 2021."},{"key":"e_1_3_2_1_14_1","first-page":"540","volume-title":"Proc. IEEE INFOCOM","author":"Chen Chen","year":"2019","unstructured":"[14] Chen Chen, Wei Wang, and Bo\u00a0Li. Round-robin synchronization: Mitigating communication bottlenecks in parameter servers. In Proc. IEEE INFOCOM, pages 532\u2013540, 2019."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1137\/16M1080173"},{"key":"e_1_3_2_1_16_1","volume-title":"Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems. arXiv preprint arXiv:1512.01274","author":"Chen Tianqi","year":"2015","unstructured":"[16] Tianqi Chen, Mu\u00a0Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems. arXiv preprint arXiv:1512.01274, 2015."},{"key":"e_1_3_2_1_17_1","first-page":"19","article-title":"Communication efficient distributed machine learning with the parameter server","volume":"27","author":"Andersen G","year":"2014","unstructured":"[17] Mu\u00a0Li, David\u00a0G Andersen, Alexander\u00a0J Smola, and Kai Yu. Communication efficient distributed machine learning with the parameter server. Advances in Neural Information Processing Systems, 27:19\u201327, 2014.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_18_1","first-page":"235","volume-title":"Proc. ACM SIGCOMM","author":"Montazeri Behnam","year":"2018","unstructured":"[18] Behnam Montazeri, Yilong Li, Mohammad Alizadeh, and John Ousterhout. Homa: A receiver-driven low-latency transport protocol using network priorities. In Proc. ACM SIGCOMM, pages 221\u2013235, 2018."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_20_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"[20] Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014."},{"key":"e_1_3_2_1_21_1","volume-title":"Learning multiple layers of features from tiny images","author":"Krizhevsky Alex","year":"2009","unstructured":"[21] Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. 2009."},{"key":"e_1_3_2_1_22_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"[22] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1264"},{"key":"e_1_3_2_1_24_1","first-page":"109","volume-title":"Proc. IEEE ICDCS","author":"Zhang Chengliang","year":"2018","unstructured":"[24] Chengliang Zhang, Huangshi Tian, Wei Wang, and Feng Yan. Stay fresh: Speculative synchronization for fast distributed machine learning. In Proc. IEEE ICDCS, pages 99\u2013109, 2018."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS.2019.00150"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2020.3040601"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM42981.2021.9488815"},{"key":"e_1_3_2_1_28_1","first-page":"430","volume-title":"Proc. SysML","author":"Hashemi Sayed\u00a0Hadi","year":"2019","unstructured":"[28] Sayed\u00a0Hadi Hashemi, Sangeetha\u00a0Abdu Jyothi, and Roy\u00a0H Campbell. Tictac: Accelerating distributed deep learning with communication scheduling. In Proc. SysML, pages 418\u2013430, 2019."},{"key":"e_1_3_2_1_29_1","first-page":"145","volume-title":"Proc. MLSys","author":"Jayarajan Anand","year":"2019","unstructured":"[29] Anand Jayarajan, Jinliang Wei, Garth Gibson, Alexandra Fedorova, and Gennady Pekhimenko. Priority-based parameter propagation for distributed dnn training. In Proc. MLSys, pages 132\u2013145, 2019."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359642"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386367.3431307"},{"key":"e_1_3_2_1_32_1","first-page":"1687","volume-title":"Proc. IEEE INFOCOM","author":"Wang Shuai","unstructured":"[32] Shuai Wang, Dan Li, and Jinkun Geng. Geryon: Accelerating distributed cnn training by network-level flow scheduling. In Proc. IEEE INFOCOM, pages 1678\u20131687. IEEE, 2020."}],"event":{"name":"ICPP '22: 51st International Conference on Parallel Processing","location":"Bordeaux France","acronym":"ICPP '22"},"container-title":["Proceedings of the 51st International Conference on Parallel Processing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3545008.3545024","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3545008.3545024","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:02:43Z","timestamp":1750186963000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3545008.3545024"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,8,29]]},"references-count":32,"alternative-id":["10.1145\/3545008.3545024","10.1145\/3545008"],"URL":"https:\/\/doi.org\/10.1145\/3545008.3545024","relation":{},"subject":[],"published":{"date-parts":[[2022,8,29]]},"assertion":[{"value":"2023-01-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}