{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T22:26:17Z","timestamp":1782944777606,"version":"3.54.5"},"reference-count":88,"publisher":"Springer Science and Business Media LLC","issue":"10","license":[{"start":{"date-parts":[[2024,3,28]],"date-time":"2024-03-28T00:00:00Z","timestamp":1711584000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,3,28]],"date-time":"2024-03-28T00:00:00Z","timestamp":1711584000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100007065","name":"Universit\u00e0 degli Studi di Salerno","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100007065","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2024,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The development of new exascale supercomputers has dramatically increased the need for fast, high-performance networking technology. Efficient network topologies, such as Dragonfly+, have been introduced to meet the demands of data-intensive applications and to match the massive computing power of GPUs and accelerators. However, these supercomputers still face performance variability mainly caused by the network that affects system and application performance. This study comprehensively analyzes performance variability on a large-scale HPC system with Dragonfly+ network topology, focusing on factors such as communication patterns, message size, job placement locality, MPI collective algorithms, and overall system workload. The study also proposes an easy-to-measure metric for estimating network background traffic generated by other users, which can be used to estimate the performance of our job accurately. The insights gained from this study contribute to improving performance predictability, enhancing job placement policies and MPI algorithm selection, and optimizing resource management strategies in supercomputers.<\/jats:p>","DOI":"10.1007\/s11227-024-06040-w","type":"journal-article","created":{"date-parts":[[2024,3,28]],"date-time":"2024-03-28T06:02:31Z","timestamp":1711605751000},"page":"14978-15005","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Analysis and prediction of performance variability in large-scale computing systems"],"prefix":"10.1007","volume":"80","author":[{"given":"Majid","family":"Salimi Beni","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sascha","family":"Hunold","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Biagio","family":"Cosenza","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,3,28]]},"reference":[{"key":"6040_CR1","doi-asserted-by":"crossref","unstructured":"Thoman P, Salzmann P, Cosenza B, Fahringer T (2019) Celerity: high-level C++ for accelerator clusters. In: Euro-Par 2019: Parallel Processing: 25th International Conference on Parallel and Distributed Computing, G\u00f6ttingen, Germany, August 26\u201330, 2019, Proceedings 25. Springer, pp 291\u2013303","DOI":"10.1007\/978-3-030-29400-7_21"},{"key":"6040_CR2","doi-asserted-by":"publisher","first-page":"3165","DOI":"10.1007\/s11227-020-03390-z","volume":"77","author":"AH Sojoodi","year":"2021","unstructured":"Sojoodi AH, Salimi Beni M, Khunjush F (2021) Ignite-gpu: a gpu-enabled in-memory computing architecture on clusters. J Supercomput 77:3165\u20133192","journal-title":"J Supercomput"},{"issue":"9","key":"6040_CR3","doi-asserted-by":"publisher","DOI":"10.1063\/5.0065859","volume":"28","author":"A Bhattacharjee","year":"2021","unstructured":"Bhattacharjee A, Wells J (2021) Preface to special topic: bilding the bridge to the exascale-applications and opportunities for plasma physics. Phys Plasmas 28(9):090401","journal-title":"Phys Plasmas"},{"key":"6040_CR4","doi-asserted-by":"crossref","unstructured":"Tr\u00e4ff JL, L\u00fcbbe FD, Rougier A, Hunold S (2015) Isomorphic, sparse MPI-like collective communication operations for parallel stencil computations. In: Proceedings of the 22nd European MPI Users\u2019 Group Meeting, pp 1\u201310","DOI":"10.1145\/2802658.2802663"},{"key":"6040_CR5","doi-asserted-by":"crossref","unstructured":"Salzmann P, Knorr F, Thoman P, Cosenza B (2022) Celerity: how (well) does the sycl api translate to distributed clusters? In: International workshop on OpenCL, pp 1\u20132","DOI":"10.1145\/3529538.3530004"},{"issue":"2","key":"6040_CR6","doi-asserted-by":"publisher","first-page":"68","DOI":"10.1109\/MM.2022.3148670","volume":"42","author":"YH Temu\u00e7in","year":"2022","unstructured":"Temu\u00e7in YH, Sojoodi AH, Alizadeh P, Kitor B, Afsahi A (2022) Accelerating deep learning using interconnect-aware UCX communication for MPI collectives. IEEE Micro 42(2):68\u201376","journal-title":"IEEE Micro"},{"key":"6040_CR7","doi-asserted-by":"crossref","unstructured":"Jain N, Bhatele A, White S, Gamblin T, Kale LV (2016) Evaluating HPC networks via simulation of parallel workloads. In: SC\u201916: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, pp 154\u2013165","DOI":"10.1109\/SC.2016.13"},{"key":"6040_CR8","doi-asserted-by":"crossref","unstructured":"Temu\u00e7in YH, Sojoodi A, Alizadeh P, Afsahi A (2021) Efficient multi-path NVLink\/PCIe-aware UCX based collective communication for deep learning. In: 2021 IEEE Symposium on High-Performance Interconnects (HOTI). IEEE, pp 25\u201334","DOI":"10.1109\/HOTI52880.2021.00018"},{"key":"6040_CR9","doi-asserted-by":"crossref","unstructured":"Alizadeh P, Sojoodi A, Hassan\u00a0Temucin Y, Afsahi A (2022) Efficient process arrival pattern aware collective communication for deep learning. In: Proceedings of the 29th European MPI Users\u2019 Group Meeting, pp 68\u201378","DOI":"10.1145\/3555819.3555857"},{"key":"6040_CR10","unstructured":"NVLink and NVSwitch. https:\/\/www.nvidia.com\/en-us\/data-center\/nvlink\/. Accessed 2023-06-30"},{"key":"6040_CR11","unstructured":"Pentakalos OI (2002) An introduction to the Infini-Band architecture. In: International CMG Conference 2002, Reno, USA, pp 425\u2013432"},{"key":"6040_CR12","doi-asserted-by":"crossref","unstructured":"Kim J, Dally WJ, Scott S, Abts D (2008) Technology-driven, highly-scalable Dragonfly topology. In: 2008 International Symposium on Computer Architecture. IEEE, pp 77\u201388","DOI":"10.1109\/ISCA.2008.19"},{"issue":"12","key":"6040_CR13","doi-asserted-by":"publisher","first-page":"1765","DOI":"10.1109\/TPDS.2010.30","volume":"21","author":"JM Camara","year":"2010","unstructured":"Camara JM, Moreto M, Vallejo E, Beivide R, Miguel-Alonso J, Mart\u00ednez C, Navaridas J (2010) Twisted torus topologies for enhanced interconnection networks. IEEE Trans Parallel Distrib Syst 21(12):1765\u20131778","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"6040_CR14","doi-asserted-by":"crossref","unstructured":"Jain N, Bhatele A, Howell LH, B\u00f6hme D, Karlin I, Le\u00f3n EA, Mubarak M, Wolfe N, Gamblin T, Leininger ML (2017) Predicting the performance impact of different Fat-Tree configurations. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pp 1\u201313","DOI":"10.1145\/3126908.3126967"},{"key":"6040_CR15","doi-asserted-by":"crossref","unstructured":"Chunduri S, Harms K, Parker S, Morozov V, Oshin S, Cherukuri N, Kumaran K (2017) Run-to-run variability on Xeon Phi based Cray XC systems. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pp 1\u201313","DOI":"10.1145\/3126908.3126926"},{"key":"6040_CR16","doi-asserted-by":"crossref","unstructured":"Yu H, Chung I-H, Moreira J (2006) Topology mapping for Blue Gene\/L supercomputer. In: Proceedings of the 2006 ACM\/IEEE Conference on Supercomputing, p 116","DOI":"10.1145\/1188455.1188576"},{"key":"6040_CR17","doi-asserted-by":"crossref","unstructured":"Jyothi SA, Singla A, Godfrey PB, Kolla A (2016) Measuring and understanding throughput of network topologies. In: SC\u201916: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, pp 761\u2013772","DOI":"10.1109\/SC.2016.64"},{"key":"6040_CR18","unstructured":"Top500, MARCONI-100. https:\/\/www.top500.org\/system\/179845\/. Accessed 2023-07-01"},{"key":"6040_CR19","doi-asserted-by":"crossref","unstructured":"Shpiner A, Haramaty Z, Eliad S, Zdornov V, Gafni B, Zahavi E (2017) Dragonfly+: low cost topology for scaling datacenters. In: 2017 IEEE 3rd International Workshop on High-Performance Interconnection Networks in the Exascale and Big-Data Era (HiPINEB). IEEE, pp 1\u20138","DOI":"10.1109\/HiPINEB.2017.11"},{"key":"6040_CR20","doi-asserted-by":"crossref","unstructured":"Zhou Z, Yang X, Lan Z, Rich P, Tang W, Morozov V, Desai N (2015) Improving batch scheduling on blue Gene\/Q by relaxing 5d torus network allocation constraints. In: 2015 IEEE International Parallel and Distributed Processing Symposium. IEEE, pp 439\u2013448","DOI":"10.1109\/IPDPS.2015.110"},{"key":"6040_CR21","doi-asserted-by":"crossref","unstructured":"Tang W, Desai N, Buettner D, Lan Z (2010) Analyzing and adjusting user runtime estimates to improve job scheduling on the Blue Gene\/P. In: 2010 IEEE International Symposium on Parallel & Distributed Processing (IPDPS). IEEE, pp 1\u201311","DOI":"10.1109\/IPDPS.2010.5470474"},{"key":"6040_CR22","doi-asserted-by":"crossref","unstructured":"Skinner D, Kramer W (2005) Understanding the causes of performance variability in HPC workloads. In: IEEE International. 2005 Proceedings of the IEEE Workload Characterization Symposium, 2005. IEEE, pp 137\u2013149","DOI":"10.1109\/IISWC.2005.1526010"},{"issue":"02","key":"6040_CR23","doi-asserted-by":"publisher","first-page":"623","DOI":"10.1109\/TPDS.2022.3221085","volume":"34","author":"A Afzal","year":"2023","unstructured":"Afzal A, Hager G, Wellein G (2023) The role of idle waves, desynchronization, and bottleneck evasion in the performance of parallel programs. IEEE Transa Parallel Distrib Syst 34(02):623\u2013638","journal-title":"IEEE Transa Parallel Distrib Syst"},{"key":"6040_CR24","doi-asserted-by":"crossref","unstructured":"Bhatele A, Thiagarajan JJ, Groves T, Anirudh R, Smith SA, Cook B, Lowenthal DK (2020) The case of performance variability on Dragonfly-based systems. In: 2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, pp 896\u2013905","DOI":"10.1109\/IPDPS47924.2020.00096"},{"key":"6040_CR25","doi-asserted-by":"crossref","unstructured":"Chester D, Groves T, Hammond SD, Law TR, Wright SA, Smedley-Stevenson R, Fahmy SA, Mudalige GR, Jarvis S (2021) Stressbench: a configurable full system network and I\/O benchmark framework. In: IEEE High Performance Extreme Computing Conference. York","DOI":"10.1109\/HPEC49654.2021.9774494"},{"key":"6040_CR26","unstructured":"Propagation and Decay of Injected One-Off Delays on Clusters: A Case Study | IEEE Conference Publication|IEEE Xplore. https:\/\/ieeexplore.ieee.org\/document\/8890995. Accessed 02\/04\/2024"},{"key":"6040_CR27","doi-asserted-by":"crossref","unstructured":"Salimi\u00a0Beni M, Crisci L, Cosenza B (2023) EMPI: enhanced message passing interface in modern c++. In: 2023 23rd IEEE International Symposium on Cluster, Cloud and Internet Computing (CCGrid). IEEE","DOI":"10.1109\/CCGrid57682.2023.00023"},{"key":"6040_CR28","doi-asserted-by":"crossref","unstructured":"Wang X, Mubarak M, Yang X, Ross RB, Lan Z (2018) Trade-off study of localizing communication and balancing network traffic on a Dragonfly system. In: 2018 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, pp 1113\u20131122","DOI":"10.1109\/IPDPS.2018.00120"},{"key":"6040_CR29","doi-asserted-by":"crossref","unstructured":"De\u00a0Sensi D, Di\u00a0Girolamo S, Hoefler T (2019) Mitigating network noise on Dragonfly networks through application-aware routing. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pp 1\u201332","DOI":"10.1145\/3295500.3356196"},{"key":"6040_CR30","doi-asserted-by":"crossref","unstructured":"Liu Y, Liu Z, Kettimuthu R, Rao N, Chen Z, Foster I (2019) Data transfer between scientific facilities - bottleneck analysis, insights and optimizations. In: 2019 19th IEEE\/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID), pp 122\u2013131","DOI":"10.1109\/CCGRID.2019.00023"},{"key":"6040_CR31","doi-asserted-by":"crossref","unstructured":"Kousha P, Sankarapandian Dayala Ganesh\u00a0Ram KR, Kedia M, Subramoni H, Jain A, Shafi A, Panda D, Dockendorf T, Na H, Tomko K (2021) Inam: cross-stack profiling and analysis of communication in MPI-based applications. In: Practice and Experience in Advanced Research Computing, pp 1\u201311","DOI":"10.1145\/3437359.3465582"},{"key":"6040_CR32","doi-asserted-by":"crossref","unstructured":"Brown KA, McGlohon N, Chunduri S, Borch E, Ross RB, Carothers CD, Harms K (2021) A tunable implementation of Quality-of-Service classes for HPC networks. In: International Conference on High Performance Computing. Springer, pp 137\u2013156","DOI":"10.1007\/978-3-030-78713-4_8"},{"key":"6040_CR33","doi-asserted-by":"crossref","unstructured":"Suresh KK, Ramesh B, Ghazimirsaeed SM, Bayatpour M, Hashmi J, Panda DK (2020) Performance characterization of network mechanisms for non-contiguous data transfers in MPI. In: 2020 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). IEEE, pp 896\u2013905","DOI":"10.1109\/IPDPSW50202.2020.00150"},{"key":"6040_CR34","doi-asserted-by":"publisher","DOI":"10.2172\/1592888","volume-title":"Evaluating trade-offs in potential exascale interconnect technologies","author":"KS Hemmert","year":"2020","unstructured":"Hemmert KS, Bair R, Bhatale A, Groves T, Jain N, Lewis C, Mubarak M, Pakin SD, Ross R, Wilke JJ (2020) Evaluating trade-offs in potential exascale interconnect technologies. Lawrence Livermore National Lab. (LLNL), Livermore"},{"key":"6040_CR35","doi-asserted-by":"crossref","unstructured":"Cheng Q, Huang Y, Bahadori M, Glick M, Rumley S, Bergman K (2018) Advanced routing strategy with highly-efficient fabric-wide characterization for optical integrated switches. In: 2018 20th International Conference on Transparent Optical Networks (ICTON). IEEE, pp 1\u20134","DOI":"10.1109\/ICTON.2018.8473671"},{"key":"6040_CR36","doi-asserted-by":"crossref","unstructured":"Zacarias FV, Nishtala R, Carpenter P (2020) Contention-aware application performance prediction for disaggregated memory systems. In: Proceedings of the 17th ACM International Conference on Computing Frontiers, pp 49\u201359","DOI":"10.1145\/3387902.3392625"},{"key":"6040_CR37","doi-asserted-by":"crossref","unstructured":"Ponce M, Zon R, Northrup S, Gruner D, Chen J, Ertinaz F, Fedoseev A, Groer L, Mao F, Mundim BC et al (2019) Deploying a top-100 supercomputer for large parallel workloads: the Niagara supercomputer. In: Proceedings of the Practice and Experience in Advanced Research Computing on Rise of the Machines (learning), pp 1\u20138","DOI":"10.1145\/3332186.3332195"},{"key":"6040_CR38","unstructured":"Marconi100. The new accelerated system. https:\/\/www.hpc.cineca.it\/hardware\/marconi100. Accessed 2023-07-01"},{"key":"6040_CR39","doi-asserted-by":"crossref","unstructured":"Kang Y, Wang X, McGlohon N, Mubarak M, Chunduri S, Lan Z (2019) Modeling and analysis of application interference on Dragonfly+. In: Proceedings of the 2019 ACM SIGSIM Conference on Principles of Advanced Discrete Simulation, pp 161\u2013172","DOI":"10.1145\/3316480.3325517"},{"key":"6040_CR40","doi-asserted-by":"crossref","unstructured":"Wang X, Mubarak M, Kang Y, Ross RB, Lan, Z (2020) Union: an automatic workload manager for accelerating network simulation. In: 2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS), pp 821\u2013830","DOI":"10.1109\/IPDPS47924.2020.00089"},{"key":"6040_CR41","doi-asserted-by":"crossref","unstructured":"Salimi\u00a0Beni M, Cosenza B (2023) An analysis of long-tailed network latency distribution and background traffic on dragonfly+. In: International Symposium on Benchmarking, Measuring and Optimization. Springer, pp 123\u2013142","DOI":"10.1007\/978-3-031-31180-2_8"},{"key":"6040_CR42","doi-asserted-by":"crossref","unstructured":"Beni MS, Cosenza B (2022) An analysis of performance variability on Dragonfly+ topology. In: 2022 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, pp 500\u2013501","DOI":"10.1109\/CLUSTER51413.2022.00061"},{"key":"6040_CR43","doi-asserted-by":"crossref","unstructured":"Navaridas J, Lant J, Pascual JA, Lujan M, Goodacre J (2019) Design exploration of multi-tier interconnection networks for exascale systems. In: Proceedings of the 48th International Conference on Parallel Processing, pp 1\u201310","DOI":"10.1145\/3337821.3337903"},{"key":"6040_CR44","doi-asserted-by":"crossref","unstructured":"Hashmi JM, Xu S, Ramesh B, Bayatpour M, Subramoni H, Panda DKD (2020) Machine-agnostic and communication-aware designs for MPI on emerging architectures. In: 2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, pp 32\u201341","DOI":"10.1109\/IPDPS47924.2020.00014"},{"key":"6040_CR45","doi-asserted-by":"crossref","unstructured":"Subramoni H, Lu X, Panda DK (2017) A scalable network-based performance analysis tool for MPI on large-scale HPC systems. In: 2017 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, pp 354\u2013358","DOI":"10.1109\/CLUSTER.2017.78"},{"key":"6040_CR46","doi-asserted-by":"crossref","unstructured":"Teh MY, Wilke JJ, Bergman K, Rumley S (2017) Design space exploration of the Dragonfly topology. In: International Conference on High Performance Computing. Springer, pp 57\u201374","DOI":"10.1007\/978-3-319-67630-2_5"},{"key":"6040_CR47","doi-asserted-by":"crossref","unstructured":"Zahn F, Fr\u00f6ning H (2020) On network locality in MPI-based HPC applications. In: 49th International Conference on Parallel Processing-ICPP, pp 1\u201310","DOI":"10.1145\/3404397.3404436"},{"key":"6040_CR48","doi-asserted-by":"crossref","unstructured":"Hoefler T, Schneider T, Lumsdaine A (2010)Characterizing the influence of system noise on large-scale applications by simulation. In: SC\u201910: Proceedings of the 2010 ACM\/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, pp 1\u201311","DOI":"10.1109\/SC.2010.12"},{"key":"6040_CR49","unstructured":"Maricq A, Duplyakin D, Jimenez I, Maltzahn C, Stutsman R, Ricci R (2018) Taming performance variability. In: 13th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 18), pp 409\u2013425"},{"key":"6040_CR50","unstructured":"Vetter J, Chambreau C (2005) MPIP: lightweight, scalable MPI profiling"},{"key":"6040_CR51","doi-asserted-by":"crossref","unstructured":"Arnold DC, Ahn DH, De\u00a0Supinski B, Lee G, Miller B, Schulz M (2007) Stack trace analysis for large scale applications. In: 21st IEEE International Parallel & Distributed Processing Symposium (IPDPS\u201907), Long Beach, CA","DOI":"10.1109\/IPDPS.2007.370254"},{"key":"6040_CR52","doi-asserted-by":"crossref","unstructured":"Petrini F, Kerbyson DJ, Pakin S (2003) The case of the missing supercomputer performance: achieving optimal performance on the 8,192 processors of asci q. In: Proceedings of the 2003 ACM\/IEEE Conference on Supercomputing, p 55","DOI":"10.1145\/1048935.1050204"},{"key":"6040_CR53","doi-asserted-by":"crossref","unstructured":"Sato K, Ahn DH, Laguna I, Lee GL, Schulz M, Chambreau CM (2017) Noise injection techniques to expose subtle and unintended message races. In: Proceedings of the 22Nd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, pp 89\u2013101","DOI":"10.1145\/3018743.3018767"},{"key":"6040_CR54","doi-asserted-by":"crossref","unstructured":"Smith SA, Cromey CE, Lowenthal DK, Domke J, Jain N, Thiagarajan JJ, Bhatele A(2018) Mitigating inter-job interference using adaptive flow-aware routing. In: SC18: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, pp 346\u2013360","DOI":"10.1109\/SC.2018.00030"},{"key":"6040_CR55","doi-asserted-by":"crossref","unstructured":"McGlohon N, Carothers CD, Hemmert KS, Levenhagen M, Brown KA, Chunduri S, Ross RB (2021) Exploration of congestion control techniques on Dragonfly-class HPC networks through simulation. In: 2021 International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS). IEEE, pp 40\u201350","DOI":"10.1109\/PMBS54543.2021.00010"},{"key":"6040_CR56","doi-asserted-by":"crossref","unstructured":"Shah A, M\u00fcller M, Wolf F (2018) Estimating the impact of external interference on application performance. In: Euro-Par 2018: Parallel Processing: 24th International Conference on Parallel and Distributed Computing, Turin, Italy, August 27-31, 2018, Proceedings 24. Springer, pp 46\u201358","DOI":"10.1007\/978-3-319-96983-1_4"},{"key":"6040_CR57","doi-asserted-by":"crossref","unstructured":"Zhang Y, Groves T, Cook B, Wright NJ, Coskun AK (2020)Quantifying the impact of network congestion on application performance and network metrics. In: 2020 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, pp 162\u2013168","DOI":"10.1109\/CLUSTER49012.2020.00026"},{"key":"6040_CR58","doi-asserted-by":"crossref","unstructured":"Brown KA, Jain N, Matsuoka S, Schulz M, Bhatele A (2018) Interference between I\/O and MPI traffic on Fat-Tree networks. In: Proceedings of the 47th International Conference on Parallel Processing, pp 1\u201310","DOI":"10.1145\/3225058.3225144"},{"key":"6040_CR59","doi-asserted-by":"crossref","unstructured":"Tang X, Zhai J, Qian X, He B, Xue W, Chen W (2018) Vsensor: leveraging fixed-workload snippets of programs for performance variance detection. In: Proceedings of the 23rd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, pp 124\u2013136","DOI":"10.1145\/3200691.3178497"},{"key":"6040_CR60","doi-asserted-by":"crossref","unstructured":"Zheng L, Zhai J, Tang X, Wang H, Yu T, Jin Y, Song SL, Chen W (2022) Vapro: performance variance detection and diagnosis for production-run parallel applications. In: Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, pp 150\u2013162","DOI":"10.1145\/3503221.3508411"},{"key":"6040_CR61","doi-asserted-by":"crossref","unstructured":"Besta M, Schneider M, Konieczny M, Cynk K, Henriksson E, Di\u00a0Girolamo S, Singla A, Hoefler T (2020) Fatpaths: routing in supercomputers and data centers when shortest paths fall short. In: SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, pp 1\u201318","DOI":"10.1109\/SC41405.2020.00031"},{"key":"6040_CR62","doi-asserted-by":"crossref","unstructured":"Kang Y, Wang X, Lan Z (2020) Q-adaptive: a multi-agent reinforcement learning based routing on Dragonfly network. In: Proceedings of the 30th International Symposium on High-Performance Parallel and Distributed Computing, pp 189\u2013200","DOI":"10.1145\/3431379.3460650"},{"key":"6040_CR63","doi-asserted-by":"crossref","unstructured":"Newaz MN, Mollah MA, Faizian P, Tong Z (2021) Improving adaptive routing performance on large scale Megafly topology. In: 2021 IEEE\/ACM 21st International Symposium on Cluster, Cloud and Internet Computing (CCGrid). IEEE, pp 406\u2013416","DOI":"10.1109\/CCGrid51090.2021.00050"},{"issue":"2","key":"6040_CR64","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3349620","volume":"6","author":"MA Mollah","year":"2019","unstructured":"Mollah MA, Wang W, Faizian P, Rahman MS, Yuan X, Pakin S, Lang M (2019) Modeling universal globally adaptive load-balanced routing. ACM Trans Parallel Comput 6(2):1\u201323","journal-title":"ACM Trans Parallel Comput"},{"issue":"4","key":"6040_CR65","doi-asserted-by":"publisher","first-page":"931","DOI":"10.1109\/TMSCS.2018.2877264","volume":"4","author":"P Faizian","year":"2018","unstructured":"Faizian P, Alfaro JF, Rahman MS, Mollah MA, Yuan X, Pakin S, Lang M (2018) Tpr: traffic pattern-based adaptive routing for Dragonfly networks. IEEE Trans Multi-Scale Comput Syst 4(4):931\u2013943","journal-title":"IEEE Trans Multi-Scale Comput Syst"},{"key":"6040_CR66","doi-asserted-by":"crossref","unstructured":"De\u00a0Sensi D, Di\u00a0Girolamo S, McMahon KH, Roweth D, Hoefler T (2020) An in-depth analysis of the slingshot interconnect. In: SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, pp 1\u201314","DOI":"10.1109\/SC41405.2020.00039"},{"key":"6040_CR67","doi-asserted-by":"crossref","unstructured":"Wen K, Samadi P, Rumley S, Chen CP, Shen Y, Bahadori M, Bergman K, Wilke J (2016)Flexfly: enabling a reconfigurable Dragonfly through silicon photonics. In: SC\u201916: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, pp 166\u2013177","DOI":"10.1109\/SC.2016.14"},{"key":"6040_CR68","doi-asserted-by":"crossref","unstructured":"Rahman MS, Bhowmik S, Ryasnianskiy Y, Yuan X, Lang M (2019) Topology-custom ugal routing on Dragonfly. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. SC \u201919. Association for Computing Machinery, New York, NY, USA","DOI":"10.1145\/3295500.3356208"},{"key":"6040_CR69","doi-asserted-by":"crossref","unstructured":"Rocher-Gonzalez J, Escudero-Sahuquillo J, Garcia PJ, Quiles FJ, Mora G (2019) Efficient congestion management for high-speed interconnects using adaptive routing. In: 2019 19th IEEE\/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID). IEEE, pp 221\u2013230","DOI":"10.1109\/CCGRID.2019.00036"},{"key":"6040_CR70","doi-asserted-by":"crossref","unstructured":"Kaplan F, Tuncer O, Leung VJ, Hemmert SK, Coskun AK (2017) Unveiling the interplay between global link arrangements and network management algorithms on Dragonfly networks. In: 2017 17th IEEE\/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID). IEEE, pp 325\u2013334","DOI":"10.1109\/CCGRID.2017.93"},{"key":"6040_CR71","doi-asserted-by":"crossref","unstructured":"Michelogiannakis G, Ibrahim KZ, Shalf J, Wilke JJ, Knight S, Kenny JP (2017) Aphid: hierarchical task placement to enable a tapered Fat Tree topology for lower power and cost in HPC networks. In: 2017 17th IEEE\/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID). IEEE, pp 228\u2013237","DOI":"10.1109\/CCGRID.2017.33"},{"key":"6040_CR72","doi-asserted-by":"crossref","unstructured":"Zhang Y, Tuncer O, Kaplan F, Olcoz K, Leung VJ, Coskun AK (2018) Level-spread: a new job allocation policy for Dragonfly networks. In: 2018 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, pp 1123\u20131132","DOI":"10.1109\/IPDPS.2018.00121"},{"key":"6040_CR73","doi-asserted-by":"crossref","unstructured":"Wang X, Yang X, Mubarak M, Ross RB, Lan Z (2017) A preliminary study of intra-application interference on Dragonfly network. In: 2017 IEEE International Conference on Cluster Computing (CLUSTER). IEEE, pp 643\u2013644","DOI":"10.1109\/CLUSTER.2017.95"},{"key":"6040_CR74","doi-asserted-by":"publisher","first-page":"e6508","DOI":"10.1002\/cpe.6508","volume":"35","author":"SA Aseeri","year":"2021","unstructured":"Aseeri SA, Gopal Chatterjee A, Verma MK, Keyes DE (2021) A scheduling policy to save 10% of communication time in parallel fast Fourier transform. Concurr Comput Pract Exp 35:e6508","journal-title":"Concurr Comput Pract Exp"},{"issue":"2","key":"6040_CR75","doi-asserted-by":"publisher","first-page":"278","DOI":"10.1145\/146628.140384","volume":"20","author":"CJ Glass","year":"1992","unstructured":"Glass CJ, Ni LM (1992) The turn model for adaptive routing. ACM SIGARCH Comput Architect News 20(2):278\u2013287","journal-title":"ACM SIGARCH Comput Architect News"},{"key":"6040_CR76","unstructured":"OSU Micro-Benchmarks 5.8, The Ohio State University.https:\/\/mvapich.cse.ohio-state.edu\/benchmarks\/. Accessed 2023-07-01"},{"issue":"1","key":"6040_CR77","doi-asserted-by":"publisher","first-page":"16","DOI":"10.3847\/1538-4365\/ab4da1","volume":"245","author":"K Heitmann","year":"2019","unstructured":"Heitmann K, Finkel H, Pope A, Morozov V, Frontiere N, Habib S, Rangel E, Uram T, Korytov D, Child H et al (2019) The outer rim simulation: a path to many-core supercomputers. Astrophys J Suppl Ser 245(1):16","journal-title":"Astrophys J Suppl Ser"},{"key":"6040_CR78","unstructured":"Heroux MA, Doerfler DW, Crozier PS, Willenbring JM, Edwards HC, Williams A, Rajan M, Keiter ER, Thornquist HK, Numrich RW (2009) Improving performance via mini-applications. Sandia National Laboratories, Technical Report SAND2009-5574, vol 3"},{"issue":"12","key":"6040_CR79","doi-asserted-by":"publisher","first-page":"3617","DOI":"10.1109\/TPDS.2016.2539167","volume":"27","author":"S Hunold","year":"2016","unstructured":"Hunold S, Carpen-Amarie A (2016) Reproducible MPI benchmarking is still not as easy as you think. IEEE Trans Parallel Distrib Syst 27(12):3617\u20133630","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"6040_CR80","unstructured":"Slurm, Slurm\u2019s job allocation policy for Dragonfly network. https:\/\/github.com\/SchedMD\/slurm\/blob\/master\/src\/plugins\/select\/linear\/select_linear.c. Accessed 2023-07-01"},{"key":"6040_CR81","unstructured":"GitHub - cea-hpc\/hp2p: Heavy Peer To Peer: a MPI based benchmark for network diagnostic. https:\/\/github.com\/cea-hpc\/hp2p. Accessed 15 May 2023"},{"key":"6040_CR82","doi-asserted-by":"publisher","unstructured":"(2008) Pearson\u2019s correlation coefficient. In: Kirch W (eds) Encyclopedia of public health. Springer, Dordrecht. https:\/\/doi.org\/10.1007\/978-1-4020-5614-7_256","DOI":"10.1007\/978-1-4020-5614-7_256"},{"key":"6040_CR83","doi-asserted-by":"publisher","unstructured":"Zar JH (2005) Spearman rank correlation. In: Encyclopedia of biostatistics, vol 7. https:\/\/doi.org\/10.1002\/0470011815.b2a15150","DOI":"10.1002\/0470011815.b2a15150"},{"key":"6040_CR84","doi-asserted-by":"crossref","unstructured":"Hunold S, Carpen-Amarie A (2018) Autotuning MPI collectives using performance guidelines. In: Proceedings of the International Conference on High Performance Computing in Asia-Pacific Region, pp 64\u201374","DOI":"10.1145\/3149457.3149461"},{"key":"6040_CR85","doi-asserted-by":"crossref","unstructured":"Hunold S, Carpen-Amarie A (2018) Algorithm selection of MPI collectives using machine learning techniques. In: 2018 IEEE\/ACM Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS). IEEE, pp 45\u201350","DOI":"10.1109\/PMBS.2018.8641622"},{"key":"6040_CR86","doi-asserted-by":"crossref","unstructured":"Hunold S, Steiner S (2022) OMPICollTune: Autotuning MPI collectives by incremental online learning. In: 2022 IEEE\/ACM International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS). IEEE, pp 123\u2013128","DOI":"10.1109\/PMBS56514.2022.00016"},{"key":"6040_CR87","doi-asserted-by":"crossref","unstructured":"Salimi\u00a0Beni M, Hunold S, Cosenza B (2023) Algorithm selection of MPI collectives considering system utilization. In: Euro-Par 2023: Parallel Processing Workshops. Springer","DOI":"10.1007\/978-3-031-48803-0_37"},{"issue":"2","key":"6040_CR88","doi-asserted-by":"publisher","first-page":"74","DOI":"10.1145\/2408776.2408794","volume":"56","author":"J Dean","year":"2013","unstructured":"Dean J, Barroso LA (2013) The tail at scale. Commun ACM 56(2):74\u201380. https:\/\/doi.org\/10.1145\/2408776.2408794","journal-title":"Commun ACM"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-024-06040-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-024-06040-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-024-06040-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,10]],"date-time":"2024-06-10T11:13:00Z","timestamp":1718017980000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-024-06040-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,28]]},"references-count":88,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2024,7]]}},"alternative-id":["6040"],"URL":"https:\/\/doi.org\/10.1007\/s11227-024-06040-w","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,28]]},"assertion":[{"value":"3 March 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 March 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}