{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T03:28:18Z","timestamp":1784258898415,"version":"3.55.0"},"reference-count":25,"publisher":"Association for Computing Machinery (ACM)","issue":"12","license":[{"start":{"date-parts":[[2024,11,22]],"date-time":"2024-11-22T00:00:00Z","timestamp":1732233600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"ETH Zurich"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Commun. ACM"],"published-print":{"date-parts":[[2024,12]]},"abstract":"<jats:p>Numerous microarchitectural optimizations unlocked tremendous processing power for deep neural networks that in turn fueled the ongoing AI revolution. With the exhaustion of such optimizations, the growth of modern AI is now gated by the performance of training systems, especially their data movement. Instead of focusing on single accelerators, we investigate data-movement characteristics of large-scale training at full system scale. Based on our workload analysis, we design HammingMesh, a novel network topology that provides high bandwidth at low cost with high job scheduling flexibility. Specifically, HammingMesh can support full bandwidth and isolation to deep learning training jobs with two dimensions of parallelism. Furthermore, it also supports high global bandwidth for generic traffic. Thus, HammingMesh will power future large-scale deep learning systems with extreme bandwidth requirements.<\/jats:p>","DOI":"10.1145\/3623490","type":"journal-article","created":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T15:54:37Z","timestamp":1732204477000},"page":"97-105","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["HammingMesh: A Network Topology for Large-Scale Deep Learning"],"prefix":"10.1145","volume":"67","author":[{"given":"Torsten","family":"Hoefler","sequence":"first","affiliation":[{"name":"ETH Zurich, Zurich, Switzerland"},{"name":"Microsoft Corp., Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tommaso","family":"Bonato","sequence":"additional","affiliation":[{"name":"ETH Zurich, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniele","family":"De Sensi","sequence":"additional","affiliation":[{"name":"ETH Zurich, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Salvatore","family":"Di Girolamo","sequence":"additional","affiliation":[{"name":"ETH Zurich, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shigang","family":"Li","sequence":"additional","affiliation":[{"name":"ETH Zurich, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marco","family":"Heddes","sequence":"additional","affiliation":[{"name":"Microsoft Corp, Redmond, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Deepak","family":"Goel","sequence":"additional","affiliation":[{"name":"Microsoft Corp, Sunnyvale, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Miguel","family":"Castro","sequence":"additional","affiliation":[{"name":"Microsoft Corp, Cambridge, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Steve","family":"Scott","sequence":"additional","affiliation":[{"name":"Microsoft Corp, Redmond, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,11,22]]},"reference":[{"issue":"2","key":"e_1_3_1_2_2","first-page":"5773","article-title":"A simulator for large-scale parallel computer architectures","volume":"1","author":"Adalsteinsson H.","year":"2010","unstructured":"Adalsteinsson, H. et al. A simulator for large-scale parallel computer architectures. Int. J. Distrib. Syst. Technol. 1, 2 (apr 2010), 5773.","journal-title":"Int. J. Distrib. Syst. Technol."},{"key":"e_1_3_1_3_2","volume-title":"Advances in Neural Information Processing Systems 31","author":"Alistarh D.","year":"2018","unstructured":"Alistarh, D. et al. The convergence of sparsified gradient methods. Advances in Neural Information Processing Systems 31. Curran Associates, Inc., Dec. 2018."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1155\/S0161171204307325"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1006\/jpdc.1995.1018"},{"issue":"4","key":"e_1_3_1_6_2","first-page":"65:1","article-title":"Demystifying parallel and distributed deep learning: An in-depth concurrency analysis","volume":"52","author":"Ben-Nun T.","year":"2019","unstructured":"Ben-Nun, T. and Hoefler, T. Demystifying parallel and distributed deep learning: An in-depth concurrency analysis. ACM Comput. Surv. 52, 4 (Aug. 2019), 65:1\u201365:43.","journal-title":"ACM Comput. Surv."},{"key":"e_1_3_1_7_2","doi-asserted-by":"crossref","unstructured":"Besta M. and Hoefler T. Slim fly: A cost effective low-diameter network topology. In Proceedings of the Intern. Conf. On High Performance Computing Networking Storage and Analysis (SC14) Nov. 2014.","DOI":"10.1109\/SC.2014.34"},{"key":"e_1_3_1_8_2","unstructured":"Brown T.B. et al. Language Models Are Few-Shot Learners 2020."},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","unstructured":"De Sensi D. et al. An in-depth analysis of the slingshot interconnect. In Proceedings of the Intern. Conf. For High Performance Computing Networking Storage and Analysis (SC20) Nov. 2020.","DOI":"10.1109\/SC41405.2020.00039"},{"issue":"241","key":"e_1_3_1_10_2","first-page":"1","article-title":"Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks","volume":"22","author":"Hoefler T.","year":"2021","unstructured":"Hoefler, T. et al. Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. J. of Machine Learning Research 22, 241 (Sep. 2021), 1\u2013124.","journal-title":"J. of Machine Learning Research"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41404.2022.00016"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","unstructured":"Hoefler T. Heddes M.C. and Belk J.R. Distributed processing architecture. US Patent Us11076210b1 Jul. 2021.","DOI":"10.1109\/SC41404.2022.00016"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Hoefler T. Heddes M.C. Goel D. and Belk J.R. Distributed processing architecture. US Patent Us20210209460a1 Jul. 2021.","DOI":"10.1109\/SC41404.2022.00016"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/1362622.1362692"},{"key":"e_1_3_1_15_2","volume-title":"Proceedings of Machine Learning and Systems 3 (Mlsys 2021)","author":"Ivanov A.","unstructured":"Ivanov, A. et al. Data movement is all you need: A case study on optimizing transformers. In Proceedings of Machine Learning and Systems 3 (Mlsys 2021), Apr. 2021."},{"key":"e_1_3_1_16_2","unstructured":"Kaplan J. et al. Scaling Laws for Neural Language Models 2020."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/2807591.2807652"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","unstructured":"Kim J. Dally W.J. Scott S. and Abts D. Technology-driven highly-scalable dragony topology. In Proceedings of 2008 Intern. Symp. On Computer Architecture 2008 77\u201388.","DOI":"10.1109\/ISCA.2008.19"},{"key":"e_1_3_1_19_2","unstructured":"Lepikhin D. et al. Gshard: Scaling Giant Models with Conditional Computation and Automatic Sharding 2020."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476145"},{"key":"e_1_3_1_21_2","unstructured":"Naumov M. et al. Deep learning recommendation model for personalization and recommendation systems. Arxiv Preprint Arxiv:1906.00091 2019."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/2464996.2465434"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"Renggli C. Alistarh D. Aghagolzadeh M. and Hoefler T. Sparcml: High-performance sparse communication for machine learning. In Proceedings of the Intern. Conf. For High Performance Computing Networking Storage and Analysis (SC19) Nov. 2019.","DOI":"10.1145\/3295500.3356222"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/2686882"},{"key":"e_1_3_1_25_2","unstructured":"Shoeybi M. et al. Megatron-Lm: Training Multi-Billion Parameter Language Models Using Model Parallelism 2020."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1177\/1094342005051521"}],"container-title":["Communications of the ACM"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3623490","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3623490","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:51:01Z","timestamp":1750287061000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3623490"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,22]]},"references-count":25,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2024,12]]}},"alternative-id":["10.1145\/3623490"],"URL":"https:\/\/doi.org\/10.1145\/3623490","relation":{},"ISSN":["0001-0782","1557-7317"],"issn-type":[{"value":"0001-0782","type":"print"},{"value":"1557-7317","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,22]]},"assertion":[{"value":"2024-11-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}