{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T16:06:22Z","timestamp":1784822782631,"version":"3.55.0"},"reference-count":47,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2025,4,11]],"date-time":"2025-04-11T00:00:00Z","timestamp":1744329600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"US National Science Foundation","doi-asserted-by":"crossref","award":["CCF-2413597"],"award-info":[{"award-number":["CCF-2413597"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000015","name":"Department of Energy","doi-asserted-by":"crossref","award":["DE-SC0024271"],"award-info":[{"award-number":["DE-SC0024271"]}],"id":[{"id":"10.13039\/100000015","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Model. Comput. Simul."],"published-print":{"date-parts":[[2025,4,30]]},"abstract":"<jats:p>Dragonfly is an indispensable interconnect topology for exascale high-performance computing (HPC) systems. To link tens of thousands of compute nodes at a reasonable cost, Dragonfly shares network resources with the entire system such that network bandwidth is not exclusive to any single application. Since HPC systems are usually shared among multiple co-running applications at the same time, network competition between co-existing workloads is inevitable. This network contention manifests as workload interference, in which a job\u2019s network communication can be severely delayed by other jobs. This study presents a comprehensive examination of leveraging intelligent routing and flexible job placement to mitigate workload interference on Dragonfly systems. Specifically, we leverage the parallel discrete event simulation toolkit, the Structural Simulation Toolkit (SST), to investigate workload interference on Dragonfly with three contributions. We first present Q-adaptive routing, a multi-agent reinforcement learning routing scheme, and a flexible job placement strategy that, together, can mitigate workload interference based on workload communication characteristics. Next, we enhance SST with Q-adaptive routing and develop an automatic module that serves as the bridge between the SST and HPC job scheduler for automatic simulation configuration and automated simulation launching. Finally, we extensively examine workload interference under various job placement and routing configurations.<\/jats:p>","DOI":"10.1145\/3706104","type":"journal-article","created":{"date-parts":[[2024,12,2]],"date-time":"2024-12-02T10:59:52Z","timestamp":1733137192000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Preventing Workload Interference with Intelligent Routing and Flexible Job Placement Strategy on Dragonfly System"],"prefix":"10.1145","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3692-2483","authenticated-orcid":false,"given":"Xin","family":"Wang","sequence":"first","affiliation":[{"name":"Computer Science, University of Illinois Chicago, Chicago, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1026-0557","authenticated-orcid":false,"given":"Yao","family":"Kang","sequence":"additional","affiliation":[{"name":"NVIDIA Corp, Santa Clara, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1047-8724","authenticated-orcid":false,"given":"Zhiling","family":"Lan","sequence":"additional","affiliation":[{"name":"Computer Science, University of Illinois Chicago, Chicago, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,4,11]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"Aurora: Argonne Leadership Computing Facility","year":"2022","unstructured":"ALCF. 2022. Aurora: Argonne Leadership Computing Facility. Retrieved from https:\/\/www.alcf.anl.gov\/aurora"},{"key":"e_1_3_2_3_2","unstructured":"ALCF. 2022. Polaris User Guide. Retrieved from https:\/\/www.alcf.anl.gov\/support\/user-guides\/polaris\/queueing-and-running-jobs\/job-and-queue-scheduling\/index.html#resource-selection-and-job-placement"},{"key":"e_1_3_2_4_2","article-title":"Cray XC series network","author":"Alverson Bob","year":"2012","unstructured":"Bob Alverson, Edwin Froese, Larry Kaplan, and Duncan Roweth. 2012. Cray XC series network. Cray Inc., White Paper WP-Aries01-1112 (2012).","journal-title":"Cray Inc., White Paper WP-Aries01-1112"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2010.40"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/2063384.2063486"},{"key":"e_1_3_2_7_2","first-page":"671","article-title":"Packet routing in dynamically changing networks: A reinforcement learning approach","volume":"6","author":"Boyan Justin","year":"1993","unstructured":"Justin Boyan and Michael Littman. 1993. Packet routing in dynamically changing networks: A reinforcement learning approach. Advances in Neural Information Processing Systems 6 (1993), 671\u2013678.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-78713-4_8"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126926"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","unstructured":"Dally and Seitz. 1987. Deadlock-free message routing in multiprocessor interconnection networks. IEEE Transactions on Computers C\u201336 5 (1987) 547\u2013553. DOI:10.1109\/TC.1987.1676939","DOI":"10.1109\/TC.1987.1676939"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356196"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2012.39"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/MLHPC54614.2021.00009"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-92040-5_15"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.33"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/1555754.1555783"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.5555\/AAI29997000"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3431379.3460650"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41404.2022.00025"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3573900.3591119"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3316480.3325517"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2008.19"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevB.47.558"},{"key":"e_1_3_2_24_2","unstructured":"Martin Lauer and Martin A. Riedmiller. 2000. An algorithm for distributed reinforcement learning in cooperative multi-agent systems. In Proceedings of the Seventeenth International Conference on Machine Learning (ICML\u201900) Morgan Kaufmann Publishers Inc. San Francisco CA USA 535\u2013542."},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2018.00068"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2007.4399095"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/PMBS54543.2021.00010"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-20656-7_1"},{"key":"e_1_3_2_30_2","volume-title":"Frontier","year":"2022","unstructured":"ORNL. 2022. Frontier. Retrieved from https:\/\/www.olcf.ornl.gov\/frontier\/"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2002.10019"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/1964218.1964225"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00039"},{"key":"e_1_3_2_34_2","article-title":"Horovod: Fast and easy distributed deep learning in TensorFlow","author":"Sergeev Alexander","year":"2018","unstructured":"Alexander Sergeev and Mike Del Balso. 2018. Horovod: Fast and easy distributed deep learning in TensorFlow. arXiv preprint arXiv:1802.05799 (2018).","journal-title":"arXiv preprint arXiv:1802.05799"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/HiPINEB.2017.11"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2005.1526010"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2018.00030"},{"key":"e_1_3_2_38_2","volume-title":"Aurora: Argonne\u2019s Next-generation Exascale Supercomputer","author":"Stevens Rick","year":"2019","unstructured":"Rick Stevens, Jini Ramprakash, Paul Messina, Michael Papka, and Katherine Riley. 2019. Aurora: Argonne\u2019s Next-generation Exascale Supercomputer. Technical Report. Argonne National Lab (ANL), Argonne, IL."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1063\/1.874055"},{"key":"e_1_3_2_40_2","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton Richard S.","year":"2018","unstructured":"Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. MIT Press."},{"key":"e_1_3_2_41_2","volume-title":"Top500 List","year":"2022","unstructured":"top500.org. 2022. Top500 List. Retrieved from https:\/\/www.top500.org\/lists\/top500\/2022\/11\/"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/1995896.1995932"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS47924.2020.00089"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2018.00120"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER49012.2020.00021"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2015.7056051"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2016.63"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1007\/10968987_3"}],"container-title":["ACM Transactions on Modeling and Computer Simulation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3706104","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3706104","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3706104","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:03Z","timestamp":1750295883000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3706104"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,11]]},"references-count":47,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,4,30]]}},"alternative-id":["10.1145\/3706104"],"URL":"https:\/\/doi.org\/10.1145\/3706104","relation":{},"ISSN":["1049-3301","1558-1195"],"issn-type":[{"value":"1049-3301","type":"print"},{"value":"1558-1195","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,11]]},"assertion":[{"value":"2024-01-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-07","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}