{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T15:57:40Z","timestamp":1782835060985,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":76,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,3,25]],"date-time":"2023-03-25T00:00:00Z","timestamp":1679702400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,3,25]]},"DOI":"10.1145\/3582016.3582037","type":"proceedings-article","created":{"date-parts":[[2023,3,20]],"date-time":"2023-03-20T16:59:03Z","timestamp":1679331543000},"page":"376-391","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":38,"title":["In-Network Aggregation with Transport Transparency for Distributed Training"],"prefix":"10.1145","author":[{"given":"Shuo","family":"Liu","sequence":"first","affiliation":[{"name":"Huawei Technologies, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qiaoling","family":"Wang","sequence":"additional","affiliation":[{"name":"Huawei Technologies, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Junyi","family":"Zhang","sequence":"additional","affiliation":[{"name":"Huawei Technologies, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenfei","family":"Wu","sequence":"additional","affiliation":[{"name":"Peking University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qinliang","family":"Lin","sequence":"additional","affiliation":[{"name":"Huawei Technologies, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yao","family":"Liu","sequence":"additional","affiliation":[{"name":"Sun Yat-sen University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Meng","family":"Xu","sequence":"additional","affiliation":[{"name":"Huawei Technologies, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marco","family":"Canini","sequence":"additional","affiliation":[{"name":"King Abdullah University of Science and Technology, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ray C. C.","family":"Cheung","sequence":"additional","affiliation":[{"name":"City University of Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianfei","family":"He","sequence":"additional","affiliation":[{"name":"City University of Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,3,25]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"TOFINO: World\u2019s fastest P4-programmable Ethernet switch ASICs.  https:\/\/barefootnetworks.com\/products\/brief-tofino\/","year":"2019","unstructured":"Barefoot. 2019 . TOFINO: World\u2019s fastest P4-programmable Ethernet switch ASICs. https:\/\/barefootnetworks.com\/products\/brief-tofino\/ Barefoot. 2019. TOFINO: World\u2019s fastest P4-programmable Ethernet switch ASICs. https:\/\/barefootnetworks.com\/products\/brief-tofino\/"},{"key":"e_1_3_2_1_2_1","volume-title":"Proceedings of IEEE Scalable High Performance Computing Conference. 357\u2013364","author":"Barnett Mike","year":"1994","unstructured":"Mike Barnett , Lance Shuler , Robert van De Geijn , Satya Gupta , David G Payne , and Jerrell Watts . 1994 . Interprocessor collective communication library (InterCom) . In Proceedings of IEEE Scalable High Performance Computing Conference. 357\u2013364 . https:\/\/ieeexplore.ieee.org\/abstract\/document\/296665 Mike Barnett, Lance Shuler, Robert van De Geijn, Satya Gupta, David G Payne, and Jerrell Watts. 1994. Interprocessor collective communication library (InterCom). In Proceedings of IEEE Scalable High Performance Computing Conference. 357\u2013364. https:\/\/ieeexplore.ieee.org\/abstract\/document\/296665"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3317550.3321436"},{"key":"e_1_3_2_1_4_1","unstructured":"Li Chen Ge Chen Justinas Lingys and Kai Chen. 2018. Programmable switch as a parallel computing device. arXiv preprint arXiv:1803.01491. \t\t\t\t  Li Chen Ge Chen Justinas Lingys and Kai Chen. 2018. Programmable switch as a parallel computing device. arXiv preprint arXiv:1803.01491."},{"key":"e_1_3_2_1_5_1","volume-title":"LightNF: Simplifying Network Function Offloading in Programmable Networks. In 2021 IEEE\/ACM 29th International Symposium on Quality of Service (IWQOS). 1\u201310","author":"Chen Xiang","year":"2021","unstructured":"Xiang Chen , Qun Huang , Peiqiao Wang , Zili Meng , Hongyan Liu , Yuxin Chen , Dong Zhang , Haifeng Zhou , Boyang Zhou , and Chunming Wu . 2021 . LightNF: Simplifying Network Function Offloading in Programmable Networks. In 2021 IEEE\/ACM 29th International Symposium on Quality of Service (IWQOS). 1\u201310 . Xiang Chen, Qun Huang, Peiqiao Wang, Zili Meng, Hongyan Liu, Yuxin Chen, Dong Zhang, Haifeng Zhou, Boyang Zhou, and Chunming Wu. 2021. LightNF: Simplifying Network Function Offloading in Programmable Networks. In 2021 IEEE\/ACM 29th International Symposium on Quality of Service (IWQOS). 1\u201310."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3106989.3106998"},{"key":"e_1_3_2_1_7_1","volume-title":"Gossipgrad: Scalable deep learning using gossip communication based asynchronous gradient descent. arXiv preprint arXiv:1803.05880,  arxiv:1803.05880","author":"Daily Jeff","year":"2018","unstructured":"Jeff Daily , Abhinav Vishnu , Charles Siegel , Thomas Warfel , and Vinay Amatya . 2018 . Gossipgrad: Scalable deep learning using gossip communication based asynchronous gradient descent. arXiv preprint arXiv:1803.05880, arxiv:1803.05880 Jeff Daily, Abhinav Vishnu, Charles Siegel, Thomas Warfel, and Vinay Amatya. 2018. Gossipgrad: Scalable deep learning using gossip communication based asynchronous gradient descent. arXiv preprint arXiv:1803.05880, arxiv:1803.05880"},{"key":"e_1_3_2_1_8_1","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1\u201316","author":"Sensi Daniele De","year":"2021","unstructured":"Daniele De Sensi , Salvatore Di Girolamo , Saleh Ashkboos , Shigang Li , and Torsten Hoefler . 2021 . Flare: flexible in-network allreduce . In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1\u201316 . Daniele De Sensi, Salvatore Di Girolamo, Saleh Ashkboos, Shigang Li, and Torsten Hoefler. 2021. Flare: flexible in-network allreduce. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1\u201316."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_1_10_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805,  arxiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, arxiv:1810.04805 Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, arxiv:1810.04805"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2012.01.020"},{"key":"e_1_3_2_1_12_1","volume-title":"Workshop on I\/O Virtualization. 2.","author":"Dong Yaozu","year":"2008","unstructured":"Yaozu Dong , Zhao Yu , and Greg Rose . 2008 . SR-IOV Networking in Xen: Architecture, Design and Implementation .. In Workshop on I\/O Virtualization. 2. Yaozu Dong, Zhao Yu, and Greg Rose. 2008. SR-IOV Networking in Xen: Architecture, Design and Implementation.. In Workshop on I\/O Virtualization. 2."},{"key":"e_1_3_2_1_13_1","volume-title":"Proceedings of SIGCOMM.","author":"Fei Jiawei","year":"2021","unstructured":"Jiawei Fei , Chen-Yu Ho , Atal Narayan Sahu , Marco Canini , and Amedeo Sapio . 2021 . Efficient Sparse Collective Communication and its application to Accelerate Distributed Deep Learning . In Proceedings of SIGCOMM. Jiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini, and Amedeo Sapio. 2021. Efficient Sparse Collective Communication and its application to Accelerate Distributed Deep Learning. In Proceedings of SIGCOMM."},{"key":"e_1_3_2_1_14_1","first-page":"829","article-title":"In-network Aggregation for Shared Machine Learning Clusters","volume":"3","author":"Gebara Nadeen","year":"2021","unstructured":"Nadeen Gebara , Manya Ghobadi , and Paolo Costa . 2021 . In-network Aggregation for Shared Machine Learning Clusters . Proceedings of Machine Learning and Systems , 3 (2021), 829 \u2013 844 . Nadeen Gebara, Manya Ghobadi, and Paolo Costa. 2021. In-network Aggregation for Shared Machine Learning Clusters. Proceedings of Machine Learning and Systems, 3 (2021), 829\u2013844.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_2_1_15_1","volume-title":"2019 IEEE 35th International Conference on Data Engineering (ICDE). 100\u2013111","author":"Geng Jinkun","year":"2019","unstructured":"Jinkun Geng , Dan Li , and Shuai Wang . 2019 . Rima: an RDMA-accelerated model-parallelized solution to large-scale matrix factorization . In 2019 IEEE 35th International Conference on Data Engineering (ICDE). 100\u2013111 . Jinkun Geng, Dan Li, and Shuai Wang. 2019. Rima: an RDMA-accelerated model-parallelized solution to large-scale matrix factorization. In 2019 IEEE 35th International Conference on Data Engineering (ICDE). 100\u2013111."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/COMHPC.2016.006"},{"key":"e_1_3_2_1_17_1","volume-title":"Scalable Hierarchical Aggregation and Reduction Protocol (SHARP) Streaming-Aggregation Hardware Design and Evaluation. In International Conference on High Performance Computing. 41\u201359","author":"Graham Richard L","year":"2020","unstructured":"Richard L Graham , Lion Levi , Devendar Burredy , Gil Bloch , Gilad Shainer , David Cho , George Elias , Daniel Klein , Joshua Ladd , and Ophir Maor . 2020 . Scalable Hierarchical Aggregation and Reduction Protocol (SHARP) Streaming-Aggregation Hardware Design and Evaluation. In International Conference on High Performance Computing. 41\u201359 . https:\/\/link.springer.com\/chapter\/10.1007\/978-3-030-50743-5_3 Richard L Graham, Lion Levi, Devendar Burredy, Gil Bloch, Gilad Shainer, David Cho, George Elias, Daniel Klein, Joshua Ladd, and Ophir Maor. 2020. Scalable Hierarchical Aggregation and Reduction Protocol (SHARP) Streaming-Aggregation Hardware Design and Evaluation. In International Conference on High Performance Computing. 41\u201359. https:\/\/link.springer.com\/chapter\/10.1007\/978-3-030-50743-5_3"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2934872.2934908"},{"key":"e_1_3_2_1_19_1","volume-title":"SoftNIC: A software NIC to augment hardware. EECS Department","author":"Han Sangjin","year":"2015","unstructured":"Sangjin Han , Keon Jang , Aurojit Panda , Shoumik Palkar , Dongsu Han , and Sylvia Ratnasamy . 2015. SoftNIC: A software NIC to augment hardware. EECS Department , University of California , Berkeley, Tech . Rep. UCB\/EECS- 2015 -155. Sangjin Han, Keon Jang, Aurojit Panda, Shoumik Palkar, Dongsu Han, and Sylvia Ratnasamy. 2015. SoftNIC: A software NIC to augment hardware. EECS Department, University of California, Berkeley, Tech. Rep. UCB\/EECS-2015-155."},{"key":"e_1_3_2_1_20_1","volume-title":"Sangeetha Abdu Jyothi, and Roy H Campbell","author":"Hashemi Sayed Hadi","year":"2018","unstructured":"Sayed Hadi Hashemi , Sangeetha Abdu Jyothi, and Roy H Campbell . 2018 . Tictac : Accelerating distributed deep learning with communication scheduling. arXiv preprint arXiv:1803.03288. Sayed Hadi Hashemi, Sangeetha Abdu Jyothi, and Roy H Campbell. 2018. Tictac: Accelerating distributed deep learning with communication scheduling. arXiv preprint arXiv:1803.03288."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3387514.3405849"},{"key":"e_1_3_2_1_23_1","unstructured":"Huggingface. 2020. Transformers:State-of-the-art Natural Language Processing for PyTorch and TensorFlow 2.0.  https:\/\/github.com\/huggingface\/transformers \t\t\t\t  Huggingface. 2020. Transformers:State-of-the-art Natural Language Processing for PyTorch and TensorFlow 2.0.  https:\/\/github.com\/huggingface\/transformers"},{"key":"e_1_3_2_1_24_1","unstructured":"Sylvain Jeaugey. 2017. NCCL 2.0.  http:\/\/on-demand.gputechconf .com\/gtc\/2017\/ presentation\/s7155-jeaugey-nccl.pdf \t\t\t\t  Sylvain Jeaugey. 2017. NCCL 2.0.  http:\/\/on-demand.gputechconf .com\/gtc\/2017\/ presentation\/s7155-jeaugey-nccl.pdf"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10766-017-0520-3"},{"key":"e_1_3_2_1_26_1","unstructured":"Xianyan Jia Shutao Song Wei He Yangzihao Wang Haidong Rong Feihu Zhou Liqiang Xie Zhenyu Guo Yuanzhou Yang and Liwei Yu. 2018. Highly scalable deep learning training system with mixed-precision: Training Imagenet in four minutes. arXiv preprint arXiv:1807.11205 arxiv:1807.11205 \t\t\t\t  Xianyan Jia Shutao Song Wei He Yangzihao Wang Haidong Rong Feihu Zhou Liqiang Xie Zhenyu Guo Yuanzhou Yang and Liwei Yu. 2018. Highly scalable deep learning training system with mixed-precision: Training Imagenet in four minutes. arXiv preprint arXiv:1807.11205 arxiv:1807.11205"},{"key":"e_1_3_2_1_27_1","volume-title":"NetChain: Scale-Free Sub-RTT Coordination. In 15th USENIX Symposium on Networked Systems Design and Implementation (NSDI 18)","author":"Jin Xin","year":"2018","unstructured":"Xin Jin , Xiaozhou Li , Haoyu Zhang , Nate Foster , Jeongkeun Lee , Robert Soul\u00e9 , Changhoon Kim , and Ion Stoica . 2018 . NetChain: Scale-Free Sub-RTT Coordination. In 15th USENIX Symposium on Networked Systems Design and Implementation (NSDI 18) . USENIX Association, Renton, WA. 35\u201349. isbn:978-1-939133-01-4 https:\/\/www.usenix.org\/conference\/nsdi18\/presentation\/jin Xin Jin, Xiaozhou Li, Haoyu Zhang, Nate Foster, Jeongkeun Lee, Robert Soul\u00e9, Changhoon Kim, and Ion Stoica. 2018. NetChain: Scale-Free Sub-RTT Coordination. In 15th USENIX Symposium on Networked Systems Design and Implementation (NSDI 18). USENIX Association, Renton, WA. 35\u201349. isbn:978-1-939133-01-4 https:\/\/www.usenix.org\/conference\/nsdi18\/presentation\/jin"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132747.3132764"},{"key":"e_1_3_2_1_29_1","volume-title":"Proceedings of the 2021 ACM SIGCOMM 2021 Conference. 657\u2013675","author":"Khani Mehrdad","year":"2021","unstructured":"Mehrdad Khani , Manya Ghobadi , Mohammad Alizadeh , Ziyi Zhu , Madeleine Glick , Keren Bergman , Amin Vahdat , Benjamin Klenk , and Eiman Ebrahimi . 2021 . SiP-ML: high-bandwidth optical network interconnects for machine learning training . In Proceedings of the 2021 ACM SIGCOMM 2021 Conference. 657\u2013675 . Mehrdad Khani, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu, Madeleine Glick, Keren Bergman, Amin Vahdat, Benjamin Klenk, and Eiman Ebrahimi. 2021. SiP-ML: high-bandwidth optical network interconnects for machine learning training. In Proceedings of the 2021 ACM SIGCOMM 2021 Conference. 657\u2013675."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3230543.3230572"},{"key":"e_1_3_2_1_31_1","volume-title":"FreeFlow: Software-based Virtual RDMA Networking for Containerized Clouds. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19)","author":"Kim Daehyeok","year":"2019","unstructured":"Daehyeok Kim , Tianlong Yu , Hongqiang Harry Liu , Yibo Zhu , Jitu Padhye , Shachar Raindel , Chuanxiong Guo , Vyas Sekar , and Srinivasan Seshan . 2019 . FreeFlow: Software-based Virtual RDMA Networking for Containerized Clouds. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19) . USENIX Association, Boston, MA. 113\u2013126. isbn:978-1-93 1971-49-2 https:\/\/www.usenix.org\/conference\/nsdi19\/presentation\/kim Daehyeok Kim, Tianlong Yu, Hongqiang Harry Liu, Yibo Zhu, Jitu Padhye, Shachar Raindel, Chuanxiong Guo, Vyas Sekar, and Srinivasan Seshan. 2019. FreeFlow: Software-based Virtual RDMA Networking for Containerized Clouds. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). USENIX Association, Boston, MA. 113\u2013126. isbn:978-1-931971-49-2 https:\/\/www.usenix.org\/conference\/nsdi19\/presentation\/kim"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00085"},{"key":"e_1_3_2_1_33_1","unstructured":"Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097\u20131105.  http:\/\/papers.nips.cc\/paper\/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf \t\t\t\t  Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097\u20131105.  http:\/\/papers.nips.cc\/paper\/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf"},{"key":"e_1_3_2_1_34_1","volume-title":"Proceedings of the ACM Special Interest Group on Data Communication. 351\u2013366","author":"Kumar Praveen","year":"2019","unstructured":"Praveen Kumar , Nandita Dukkipati , Nathan Lewis , Yi Cui , Yaogong Wang , Chonggang Li , Valas Valancius , Jake Adriaens , Steve Gribble , and Nate Foster . 2019 . PicNIC: predictable virtualized NIC . In Proceedings of the ACM Special Interest Group on Data Communication. 351\u2013366 . Praveen Kumar, Nandita Dukkipati, Nathan Lewis, Yi Cui, Yaogong Wang, Chonggang Li, Valas Valancius, Jake Adriaens, Steve Gribble, and Nate Foster. 2019. PicNIC: predictable virtualized NIC. In Proceedings of the ACM Special Interest Group on Data Communication. 351\u2013366."},{"key":"e_1_3_2_1_35_1","volume-title":"ATP: In-network Aggregation for Multi-tenant Learning. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21)","author":"Lao ChonLam","year":"2021","unstructured":"ChonLam Lao , Yanfang Le , Kshiteej Mahajan , Yixi Chen , Wenfei Wu , Aditya Akella , and Michael Swift . 2021 . ATP: In-network Aggregation for Multi-tenant Learning. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21) . USENIX Association, 741\u2013761. isbn:978-1-939133-21-2 https:\/\/www.usenix.org\/conference\/nsdi21\/presentation\/lao ChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen, Wenfei Wu, Aditya Akella, and Michael Swift. 2021. ATP: In-network Aggregation for Multi-tenant Learning. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21). USENIX Association, 741\u2013761. isbn:978-1-939133-21-2 https:\/\/www.usenix.org\/conference\/nsdi21\/presentation\/lao"},{"key":"e_1_3_2_1_36_1","unstructured":"Alberto Lerner Rana Hussein Philippe Cudre-Mauroux and U eXascale Infolab. 2019. The Case for Network Accelerated Query Processing.. In CIDR. \t\t\t\t  Alberto Lerner Rana Hussein Philippe Cudre-Mauroux and U eXascale Infolab. 2019. The Case for Network Accelerated Query Processing.. In CIDR."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10766-018-00623-w"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322259"},{"key":"e_1_3_2_1_39_1","volume-title":"DistCache: Provable Load Balancing for Large-Scale Storage Systems with Distributed Caching. In 17th USENIX Conference on File and Storage Technologies (FAST 19)","author":"Liu Zaoxing","year":"2019","unstructured":"Zaoxing Liu , Zhihao Bai , Zhenming Liu , Xiaozhou Li , Changhoon Kim , Vladimir Braverman , Xin Jin , and Ion Stoica . 2019 . DistCache: Provable Load Balancing for Large-Scale Storage Systems with Distributed Caching. In 17th USENIX Conference on File and Storage Technologies (FAST 19) . USENIX Association, Boston, MA. 143\u2013157. isbn:978-1-939133-09-0 https:\/\/www.usenix.org\/conference\/fast19\/presentation\/liu Zaoxing Liu, Zhihao Bai, Zhenming Liu, Xiaozhou Li, Changhoon Kim, Vladimir Braverman, Xin Jin, and Ion Stoica. 2019. DistCache: Provable Load Balancing for Large-Scale Storage Systems with Distributed Caching. In 17th USENIX Conference on File and Storage Technologies (FAST 19). USENIX Association, Boston, MA. 143\u2013157. isbn:978-1-939133-09-0 https:\/\/www.usenix.org\/conference\/fast19\/presentation\/liu"},{"key":"e_1_3_2_1_40_1","volume-title":"Proc. of MLSys.","author":"Luo Liang","year":"2020","unstructured":"Liang Luo , Peter West , Arvind Krishnamurthy , Luis Ceze , and Jacob Nelson . 2020 . PLink: Discovering and Exploiting Datacenter Network Locality for Efficient Cloud-based Distributed Training . Proc. of MLSys. Liang Luo, Peter West, Arvind Krishnamurthy, Luis Ceze, and Jacob Nelson. 2020. PLink: Discovering and Exploiting Datacenter Network Locality for Efficient Cloud-based Distributed Training. Proc. of MLSys."},{"key":"e_1_3_2_1_41_1","volume-title":"Themis: Fair and efficient $GPU$ cluster scheduling. In 17th $USENIX$ Symposium on Networked Systems Design and Implementation ($NSDI$ 20). 289\u2013304.","author":"Mahajan Kshiteej","year":"2020","unstructured":"Kshiteej Mahajan , Arjun Balasubramanian , Arjun Singhvi , Shivaram Venkataraman , Aditya Akella , Amar Phanishayee , and Shuchi Chawla . 2020 . Themis: Fair and efficient $GPU$ cluster scheduling. In 17th $USENIX$ Symposium on Networked Systems Design and Implementation ($NSDI$ 20). 289\u2013304. Kshiteej Mahajan, Arjun Balasubramanian, Arjun Singhvi, Shivaram Venkataraman, Aditya Akella, Amar Phanishayee, and Shuchi Chawla. 2020. Themis: Fair and efficient $GPU$ cluster scheduling. In 17th $USENIX$ Symposium on Networked Systems Design and Implementation ($NSDI$ 20). 289\u2013304."},{"key":"e_1_3_2_1_42_1","unstructured":"Mellanox. 2022. ConnectX-5 EN Single\/Dual-Port Adapter Supporting 100Gb\/s Ethernet.  https:\/\/www.mellanox.com\/products\/ethernet-adapters\/connectx-5-en \t\t\t\t  Mellanox. 2022. ConnectX-5 EN Single\/Dual-Port Adapter Supporting 100Gb\/s Ethernet.  https:\/\/www.mellanox.com\/products\/ethernet-adapters\/connectx-5-en"},{"key":"e_1_3_2_1_43_1","unstructured":"Mellanox. 2022. InfiniBand Switch Silicon: Mellanox Quantum.  https:\/\/www.mellanox.com\/products\/infiniband-switches-ic\/quantum \t\t\t\t  Mellanox. 2022. InfiniBand Switch Silicon: Mellanox Quantum.  https:\/\/www.mellanox.com\/products\/infiniband-switches-ic\/quantum"},{"key":"e_1_3_2_1_44_1","unstructured":"Jeffrey C Mogul. 2003. TCP Offload Is a Dumb Idea Whose Time Has Come.. In HotOS. 25\u201330. \t\t\t\t  Jeffrey C Mogul. 2003. TCP Offload Is a Dumb Idea Whose Time Has Come.. In HotOS. 25\u201330."},{"key":"e_1_3_2_1_45_1","volume-title":"Jumpgate: In-network processing as a service for data analytics. In 11th $USENIX$ Workshop on Hot Topics in Cloud Computing (HotCloud 19).","author":"Mustard Craig","year":"2019","unstructured":"Craig Mustard , Fabian Ruffy , Anny Gakhokidze , Ivan Beschastnikh , and Alexandra Fedorova . 2019 . Jumpgate: In-network processing as a service for data analytics. In 11th $USENIX$ Workshop on Hot Topics in Cloud Computing (HotCloud 19). Craig Mustard, Fabian Ruffy, Anny Gakhokidze, Ivan Beschastnikh, and Alexandra Fedorova. 2019. Jumpgate: In-network processing as a service for data analytics. In 11th $USENIX$ Workshop on Hot Topics in Cloud Computing (HotCloud 19)."},{"key":"e_1_3_2_1_46_1","volume-title":"14th $USENIX$ Symposium on Operating Systems Design and Implementation ($OSDI$ 20). 481\u2013498.","author":"Narayanan Deepak","unstructured":"Deepak Narayanan , Keshav Santhanam , Fiodar Kazhamiaka , Amar Phanishayee , and Matei Zaharia . 2020. Heterogeneity-aware cluster scheduling policies for deep learning workloads . In 14th $USENIX$ Symposium on Operating Systems Design and Implementation ($OSDI$ 20). 481\u2013498. Deepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee, and Matei Zaharia. 2020. Heterogeneity-aware cluster scheduling policies for deep learning workloads. In 14th $USENIX$ Symposium on Operating Systems Design and Implementation ($OSDI$ 20). 481\u2013498."},{"key":"e_1_3_2_1_47_1","unstructured":"NVIDIA. 2017. NVIDIA DGX-1 with Tesla V100 System Architecture.  https:\/\/www.nvidia.com\/en-us\/data-center\/resources\/dgx-1-system-architecture-whitepaper\/ \t\t\t\t  NVIDIA. 2017. NVIDIA DGX-1 with Tesla V100 System Architecture.  https:\/\/www.nvidia.com\/en-us\/data-center\/resources\/dgx-1-system-architecture-whitepaper\/"},{"key":"e_1_3_2_1_48_1","volume-title":"NCCL: Optimized primitives for collective multi-GPU communication.  https:\/\/github.com\/NVIDIA\/nccl","author":"NVIDIA.","year":"2019","unstructured":"NVIDIA. 2019 . NCCL: Optimized primitives for collective multi-GPU communication. https:\/\/github.com\/NVIDIA\/nccl NVIDIA. 2019. NCCL: Optimized primitives for collective multi-GPU communication. https:\/\/github.com\/NVIDIA\/nccl"},{"key":"e_1_3_2_1_49_1","unstructured":"NVIDIA. 2019. NVIDIA NVLink Fabric.  https:\/\/www.nvidia.com\/en-sg\/data-center\/nvlink\/ \t\t\t\t  NVIDIA. 2019. NVIDIA NVLink Fabric.  https:\/\/www.nvidia.com\/en-sg\/data-center\/nvlink\/"},{"key":"e_1_3_2_1_50_1","unstructured":"NVIDIA. 2020. NVIDIA V100: The First Tensor Core GPU.  https:\/\/www.nvidia.com\/en-sg\/data-center\/v100\/ \t\t\t\t  NVIDIA. 2020. NVIDIA V100: The First Tensor Core GPU.  https:\/\/www.nvidia.com\/en-sg\/data-center\/v100\/"},{"key":"e_1_3_2_1_51_1","unstructured":"NVIDIA. 2021. NVIDIA Collective Communication Library (NCCL).  https:\/\/developer.nvidia.com\/nccl \t\t\t\t  NVIDIA. 2021. NVIDIA Collective Communication Library (NCCL).  https:\/\/developer.nvidia.com\/nccl"},{"key":"e_1_3_2_1_52_1","volume-title":"GeForce RTX 2080","author":"NVIDIA.","year":"2023","unstructured":"NVIDIA. 2023 . GeForce RTX 2080 . https:\/\/www.nvidia.com\/en-us\/geforce\/graphics-cards\/rtx-2080\/ NVIDIA. 2023. GeForce RTX 2080. https:\/\/www.nvidia.com\/en-us\/geforce\/graphics-cards\/rtx-2080\/"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359642"},{"key":"e_1_3_2_1_54_1","unstructured":"Rolf Rabenseifner. 1997. A new optimized MPI reduce algorithm.  https:\/\/fs.hlrs.de\/projects\/par\/mpi\/\/myreduce.html \t\t\t\t  Rolf Rabenseifner. 1997. A new optimized MPI reduce algorithm.  https:\/\/fs.hlrs.de\/projects\/par\/mpi\/\/myreduce.html"},{"key":"e_1_3_2_1_55_1","unstructured":"Alec Radford Jeffrey Wu Dario Amodei Daniela Amodei Jack Clark Miles Brundage and Ilya Sutskever. 2019. Better language models and their implications. OpenAI Blog https:\/\/openai. com\/blog\/better-language-models \t\t\t\t  Alec Radford Jeffrey Wu Dario Amodei Daniela Amodei Jack Clark Miles Brundage and Ilya Sutskever. 2019. Better language models and their implications. OpenAI Blog https:\/\/openai. com\/blog\/better-language-models"},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"crossref","unstructured":"Pranav Rajpurkar Jian Zhang Konstantin Lopyrev and Percy Liang. 2016. SQUAD: 100 000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250 arxiv:1606.05250 \t\t\t\t  Pranav Rajpurkar Jian Zhang Konstantin Lopyrev and Percy Liang. 2016. SQUAD: 100 000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250 arxiv:1606.05250","DOI":"10.18653\/v1\/D16-1264"},{"key":"e_1_3_2_1_57_1","volume-title":"irdma: Efficient use of rdma in distributed deep learning systems. In 2017 IEEE 19th International Conference on High Performance Computing and Communications","author":"Ren Yufei","unstructured":"Yufei Ren , Xingbo Wu , Li Zhang , Yandong Wang , Wei Zhang , Zijun Wang , Michel Hack , and Song Jiang . 2017. irdma: Efficient use of rdma in distributed deep learning systems. In 2017 IEEE 19th International Conference on High Performance Computing and Communications ; IEEE 15th International Conference on Smart City; IEEE 3rd International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS) . 231\u2013238. Yufei Ren, Xingbo Wu, Li Zhang, Yandong Wang, Wei Zhang, Zijun Wang, Michel Hack, and Song Jiang. 2017. irdma: Efficient use of rdma in distributed deep learning systems. In 2017 IEEE 19th International Conference on High Performance Computing and Communications; IEEE 15th International Conference on Smart City; IEEE 3rd International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS). 231\u2013238."},{"key":"e_1_3_2_1_58_1","volume-title":"Scaling Distributed Machine Learning with In-Network Aggregation. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21)","author":"Sapio Amedeo","year":"2021","unstructured":"Amedeo Sapio , Marco Canini , Chen-Yu Ho , Jacob Nelson , Panos Kalnis , Changhoon Kim , Arvind Krishnamurthy , Masoud Moshref , Dan Ports , and Peter Richtarik . 2021 . Scaling Distributed Machine Learning with In-Network Aggregation. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21) . USENIX Association, 785\u2013808. isbn:978-1-939133-21-2 https:\/\/www.usenix.org\/conference\/nsdi21\/presentation\/sapio Amedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson, Panos Kalnis, Changhoon Kim, Arvind Krishnamurthy, Masoud Moshref, Dan Ports, and Peter Richtarik. 2021. Scaling Distributed Machine Learning with In-Network Aggregation. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21). USENIX Association, 785\u2013808. isbn:978-1-939133-21-2 https:\/\/www.usenix.org\/conference\/nsdi21\/presentation\/sapio"},{"key":"e_1_3_2_1_59_1","unstructured":"Alexander Sergeev and Mike Del Balso. 2018. Horovod: fast and easy distributed deep learning in TensorFlow. arXiv preprint arXiv:1802.05799 arxiv:1802.05799 \t\t\t\t  Alexander Sergeev and Mike Del Balso. 2018. Horovod: fast and easy distributed deep learning in TensorFlow. arXiv preprint arXiv:1802.05799 arxiv:1802.05799"},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM.2019.8737367"},{"key":"e_1_3_2_1_61_1","unstructured":"Jinwoo Shin and KyoungSoo Park. 2021. Elastic Resource Sharing for Distributed Deep Learning. \t\t\t\t  Jinwoo Shin and KyoungSoo Park. 2021. Elastic Resource Sharing for Distributed Deep Learning."},{"key":"e_1_3_2_1_62_1","volume-title":"Megatron-lm: Training multi-billion parameter language models using gpu model parallelism. arXiv preprint arXiv:1909.08053,  arxiv:1909.08053","author":"Shoeybi Mohammad","year":"2019","unstructured":"Mohammad Shoeybi , Mostofa Patwary , Raul Puri , Patrick LeGresley , Jared Casper , and Bryan Catanzaro . 2019 . Megatron-lm: Training multi-billion parameter language models using gpu model parallelism. arXiv preprint arXiv:1909.08053, arxiv:1909.08053 Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using gpu model parallelism. arXiv preprint arXiv:1909.08053, arxiv:1909.08053"},{"key":"e_1_3_2_1_63_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 arxiv:1409.1556 \t\t\t\t  Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 arxiv:1409.1556"},{"key":"e_1_3_2_1_64_1","volume-title":"Proceedings of the Twentieth ACM Workshop on Hot Topics in Networks. 61\u201368","author":"Stephens Brent E","year":"2021","unstructured":"Brent E Stephens , Darius Grassi , Hamidreza Almasi , Tao Ji , Balajee Vamanan , and Aditya Akella . 2021 . TCP is Harmful to In-Network Computing: Designing a Message Transport Protocol (MTP) . In Proceedings of the Twentieth ACM Workshop on Hot Topics in Networks. 61\u201368 . Brent E Stephens, Darius Grassi, Hamidreza Almasi, Tao Ji, Balajee Vamanan, and Aditya Akella. 2021. TCP is Harmful to In-Network Computing: Designing a Message Transport Protocol (MTP). In Proceedings of the Twentieth ACM Workshop on Hot Topics in Networks. 61\u201368."},{"key":"e_1_3_2_1_65_1","unstructured":"PyTorch Team. 2023. PyTorch.  https:\/\/github.com\/pytorch\/pytorch \t\t\t\t  PyTorch Team. 2023. PyTorch.  https:\/\/github.com\/pytorch\/pytorch"},{"key":"e_1_3_2_1_66_1","unstructured":"TensorFlow. 2019. A benchmark framework for Tensorflow.  https:\/\/github.com\/tensorflow\/benchmarks \t\t\t\t  TensorFlow. 2019. A benchmark framework for Tensorflow.  https:\/\/github.com\/tensorflow\/benchmarks"},{"key":"e_1_3_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389698"},{"key":"e_1_3_2_1_68_1","volume-title":"Proceedings of the 11th ACM Symposium on Cloud Computing. 447\u2013461","author":"Viswanathan Raajay","year":"2020","unstructured":"Raajay Viswanathan , Arjun Balasubramanian , and Aditya Akella . 2020 . Network-accelerated distributed machine learning for multi-tenant settings . In Proceedings of the 11th ACM Symposium on Cloud Computing. 447\u2013461 . Raajay Viswanathan, Arjun Balasubramanian, and Aditya Akella. 2020. Network-accelerated distributed machine learning for multi-tenant settings. In Proceedings of the 11th ACM Symposium on Cloud Computing. 447\u2013461."},{"key":"e_1_3_2_1_69_1","volume-title":"GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461,  arxiv:1804.07461","author":"Wang Alex","year":"2018","unstructured":"Alex Wang , Amanpreet Singh , Julian Michael , Felix Hill , Omer Levy , and Samuel R Bowman . 2018 . GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, arxiv:1804.07461 Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, arxiv:1804.07461"},{"key":"e_1_3_2_1_70_1","volume-title":"Blink: Fast and generic collectives for distributed ml. arXiv preprint arXiv:1910.04940.","author":"Wang Guanhua","year":"2019","unstructured":"Guanhua Wang , Shivaram Venkataraman , Amar Phanishayee , Jorgen Thelin , Nikhil Devanur , and Ion Stoica . 2019 . Blink: Fast and generic collectives for distributed ml. arXiv preprint arXiv:1910.04940. Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, Jorgen Thelin, Nikhil Devanur, and Ion Stoica. 2019. Blink: Fast and generic collectives for distributed ml. arXiv preprint arXiv:1910.04940."},{"key":"e_1_3_2_1_71_1","unstructured":"Xilinx. 2023. Virtex UltraScale - Xilinx.  https:\/\/www.xilinx.com\/products\/silicon-devices\/fpga\/virtex-ultrascale.html#productAdvantages \t\t\t\t  Xilinx. 2023. Virtex UltraScale - Xilinx.  https:\/\/www.xilinx.com\/products\/silicon-devices\/fpga\/virtex-ultrascale.html#productAdvantages"},{"key":"e_1_3_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/3302424.3303975"},{"key":"e_1_3_2_1_73_1","volume-title":"Traffic Management for Distributed Machine Learning in RDMA-enabled Data Center Networks. In ICC 2021-IEEE International Conference on Communications. 1\u20136.","author":"Yang Weihong","year":"2021","unstructured":"Weihong Yang , Yang Qin , Zukai Jiang , and Xiaowen Chu . 2021 . Traffic Management for Distributed Machine Learning in RDMA-enabled Data Center Networks. In ICC 2021-IEEE International Conference on Communications. 1\u20136. Weihong Yang, Yang Qin, Zukai Jiang, and Xiaowen Chu. 2021. Traffic Management for Distributed Machine Learning in RDMA-enabled Data Center Networks. In ICC 2021-IEEE International Conference on Communications. 1\u20136."},{"key":"e_1_3_2_1_74_1","volume-title":"Unlocking the Power of Inline Floating-Point Operations on Programmable Switches. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22)","author":"Yuan Yifan","year":"2022","unstructured":"Yifan Yuan , Omar Alama , Jiawei Fei , Jacob Nelson , Dan R. K. Ports , Amedeo Sapio , Marco Canini , and Nam Sung Kim . 2022 . Unlocking the Power of Inline Floating-Point Operations on Programmable Switches. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22) . Yifan Yuan, Omar Alama, Jiawei Fei, Jacob Nelson, Dan R. K. Ports, Amedeo Sapio, Marco Canini, and Nam Sung Kim. 2022. Unlocking the Power of Inline Floating-Point Operations on Programmable Switches. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22)."},{"key":"e_1_3_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.14778\/3368289.3368301"},{"key":"e_1_3_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/2785956.2787484"}],"event":{"name":"ASPLOS '23: 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3","location":"Vancouver BC Canada","acronym":"ASPLOS '23","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture","SIGOPS ACM Special Interest Group on Operating Systems","SIGPLAN ACM Special Interest Group on Programming Languages","SIGBED ACM Special Interest Group on Embedded Systems"]},"container-title":["Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3582016.3582037","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:45Z","timestamp":1750178805000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3582016.3582037"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,25]]},"references-count":76,"alternative-id":["10.1145\/3582016.3582037","10.1145\/3582016"],"URL":"https:\/\/doi.org\/10.1145\/3582016.3582037","relation":{},"subject":[],"published":{"date-parts":[[2023,3,25]]},"assertion":[{"value":"2023-03-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}