{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T10:46:10Z","timestamp":1781606770611,"version":"3.54.5"},"reference-count":72,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2025,4,12]],"date-time":"2025-04-12T00:00:00Z","timestamp":1744416000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"European Union\u2019s Horizon 2020 research and innovation programme","award":["671553"],"award-info":[{"award-number":["671553"]}]},{"name":"EuroEXA","award":["754337"],"award-info":[{"award-number":["754337"]}]},{"name":"RED-SEA EuroHPC","award":["955776x"],"award-info":[{"award-number":["955776x"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2025,6,30]]},"abstract":"<jats:p>We present and evaluate the ExaNeSt prototype, which compactly packages 128 Xilinx ZU9EG MPSoCs, two TBytes of DRAM, and eight TBytes of SSD into a liquid-cooled rack, using a custom interconnection hardware based on 10 GB\/s links. We developed this testbed in 2016\u20132019 in order to leverage the flexibility of FPGAs for experimenting with efficient hardware support for HPC communication among tens of thousands of processors and accelerators in the quest toward Exascale systems and beyond. In the years since then, we carefully studied this system, and we present our key design choices and insights resulting from our measurement and analysis.<\/jats:p>\n          <jats:p>We developed this testbed, from architecture to the PCBs and the run-time software, within the ExaNeSt project. It is fully operational in configurations with up to 8 \u00d7 4 \u00d7 4 MPSoC nodes. It achieves high density through tight board design, while also leveraging state-of-the-art liquid cooling technology. In this article, we present a thorough architectural analysis, along with important aspects of our infrastructure development. Our custom interconnect includes a low-cost low-latency network interface, offering user-level, zero-copy RDMA, which we coupled with the ARMv8 processors in the MPSoCs. We further developed the corresponding runtimes that allow us to test real MPI applications on the large-scale testbed.<\/jats:p>\n          <jats:p>\n            We evaluated our platform through MPI microbenchmarks, mini application, and full MPI applications. Single-hop, one-way latency is 1.3 \u03bcs; approximately 0.47 \u03bcs out of these are attributed to network interface and the user-space library that exposes its functionality to the runtime. Latency over longer paths increases as expected, reaching 2.55 \u03bcs for a five-hop path. Bandwidth tests show that, for single-hop, link utilization reaches\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(82\\%\\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            of the theoretical capacity. Microbenchmarks based on MPI collectives reveal that broadcast latency scales as expected when the number of participating ranks increases. We also implemented a custom MPI_Allreduce accelerator in the network interface, which reduces the latency of such collectives by up to\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(88\\%\\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            . We assess performance scaling through weak and strong scaling tests for HPCG, LAMMPS, and the miniFE mini application; for all these tests, parallelization efficiency is at least\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(69\\%\\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            , or better.\n          <\/jats:p>","DOI":"10.1145\/3715152","type":"journal-article","created":{"date-parts":[[2025,2,4]],"date-time":"2025-02-04T15:08:22Z","timestamp":1738681702000},"page":"1-34","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["The ExaNeSt Prototype: Evaluation of Efficient HPC Communication Hardware in an ARM-based Multi-FPGA Rack"],"prefix":"10.1145","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2173-062X","authenticated-orcid":false,"given":"Manolis","family":"Ploumidis","sequence":"first","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-9438-6436","authenticated-orcid":false,"given":"Fabien","family":"Chaix","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-8497-6985","authenticated-orcid":false,"given":"Nikolaos","family":"Chrysos","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-2660-6308","authenticated-orcid":false,"given":"Marios","family":"Assiminakis","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0331-1475","authenticated-orcid":false,"given":"Nikolaos","family":"Kallimanis","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8559-0557","authenticated-orcid":false,"given":"Nikolaos","family":"Kossifidis","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-6932-9650","authenticated-orcid":false,"given":"Michael","family":"Nikoloudakis","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5788-0892","authenticated-orcid":false,"given":"Nikolaos","family":"Dimou","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-8678-770X","authenticated-orcid":false,"given":"Michalis","family":"Gianioudis","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9776-985X","authenticated-orcid":false,"given":"George","family":"Ieronymakis","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6805-3107","authenticated-orcid":false,"given":"Aggelos","family":"Ioannou","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6032-1140","authenticated-orcid":false,"given":"George","family":"Kalokerinos","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-3091-920X","authenticated-orcid":false,"given":"Pantelis","family":"Xirouchakis","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-3944-5156","authenticated-orcid":false,"given":"Astrinos","family":"Damianakis","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-9362-4671","authenticated-orcid":false,"given":"Michael","family":"Ligerakis","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece and Exapsys, PLC, Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-1912-3797","authenticated-orcid":false,"given":"Theocharis","family":"Vavouris","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-5437-4709","authenticated-orcid":false,"given":"Manolis","family":"Katevenis","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece, and Department of Computer Science, University of Crete, Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5443-6470","authenticated-orcid":false,"given":"Vassilis","family":"Papaefstathiou","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece, and Department of Computer Science, University of Crete, Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4768-3289","authenticated-orcid":false,"given":"Manolis","family":"Marazakis","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2665-5203","authenticated-orcid":false,"given":"Iakovos","family":"Mavroidis","sequence":"additional","affiliation":[{"name":"Foundation for Research and Technology Hellas (FORTH), Heraklion, Greece and Exapsys, PLC, Heraklion, Greece."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,4,12]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"AMD. 2016. Integrated Logic Analyzer v6.2. Retrieved from https:\/\/docs.amd.com\/v\/u\/en-US\/pg172-ila"},{"key":"e_1_3_3_3_2","unstructured":"ExaNeSt. 2019. European Exascale System Interconnect and Storage - ExaNeSt. Retrieved from https:\/\/www.exanest.eu"},{"key":"e_1_3_3_4_2","unstructured":"HPCG. 2022. HPCG Benchmark. Retrieved from https:\/\/hpcg-benchmark.org\/index.html"},{"key":"e_1_3_3_5_2","unstructured":"LAMMPS. 2024. LAMMPS User Guide. Retrieved from https:\/\/docs.lammps.org\/Speed_bench.html"},{"key":"e_1_3_3_6_2","unstructured":"OSU. 2024. OSU Micro-Benchmarks. Retrieved from http:\/\/mvapich.cse.ohio-state.edu\/benchmarks\/"},{"key":"e_1_3_3_7_2","unstructured":"Advanced Micro Devices Inc. 2022. AMD Alveo Adaptable Accelerator Cards. Retrieved from https:\/\/www.amd.com\/en\/products\/accelerators\/alveo.html"},{"key":"e_1_3_3_8_2","first-page":"280","volume-title":"High Performance Computing","author":"Agarwal Saurabh","year":"2005","unstructured":"Saurabh Agarwal, Rahul Garg, and Nisheeth K. Vishnoi. 2005. The impact of noise on the scaling of collectives: A theoretical approach. In D. A. Bader, M. Parashar, V. Sridhar, and V. K. Prasanna (Eds.), High Performance Computing. Springer, Berlin, 280\u2013289."},{"key":"e_1_3_3_9_2","first-page":"481","volume-title":"Cluster","author":"Ammendola Roberto","year":"2004","unstructured":"Roberto Ammendola, M. Guagnelli, G. Mazza, Filippo Palombi, Roberto Petronzio, Davide Rossetti, Andrea Salamon, and Piero Vicini. 2004. APENet: A high speed, low latency 3D interconnect network. In Cluster. Citeseer, 481."},{"key":"e_1_3_3_10_2","doi-asserted-by":"crossref","first-page":"287","DOI":"10.1109\/AHS.2007.71","volume-title":"Proceedings of the 2nd NASA\/ESA Conference on Adaptive Hardware and Systems (AHS \u201907)","author":"Baxter Rob","year":"2007","unstructured":"Rob Baxter, Stephen Booth, Mark Bull, Geoff Cawood, James Perry, Mark Parsons, Alan Simpson, Arthur Trew, Andrew McCormick, Graham Smart, et al. 2007. Maxwell - A 64 FPGA supercomputer. In Proceedings of the 2nd NASA\/ESA Conference on Adaptive Hardware and Systems (AHS \u201907), 287\u2013294. DOI: 10.1109\/AHS.2007.71"},{"key":"e_1_3_3_11_2","volume-title":"In Proceedings of the 25th Euromicro Conference on Digital System Design (DSD) and SEAA (Software Engineering and Advanced Applications) Conference","author":"Biagioni Andrea","year":"2022","unstructured":"Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, and Matteo Turisini, et al. 2022. RED-SEA: Network solution for exascale architectures. In Proceedings of the 25th Euromicro Conference on Digital System Design (DSD) and SEAA (Software Engineering and Advanced Applications) Conference."},{"key":"e_1_3_3_12_2","unstructured":"BittWare. 2019. BittWare FPGA Acceleration. Retrieved from https:\/\/www.bittware.com\/"},{"key":"e_1_3_3_13_2","first-page":"130","volume-title":"Proceedings of the International Conference on High Performance Computing Simulation (HPCS)","author":"Blott M.","year":"2016","unstructured":"M. Blott. 2016. Reconfigurable future for HPC. In Proceedings of the International Conference on High Performance Computing Simulation (HPCS), 130\u2013131."},{"key":"e_1_3_3_14_2","first-page":"60","volume-title":"Proceedings of the ACM\/IEEE Conference on Supercomputing (SC)","author":"Adiga N. R.","year":"2002","unstructured":"N. R. Adiga, G. Almasi, G. S. Almasi, Y. Aridor, R. Barik, D. Beece, R. Bellofatto, G. Bhanot, R. Bickford, M. Blumrich, et al. 2002. An overview of the BlueGene\/L supercomputer. In Proceedings of the ACM\/IEEE Conference on Supercomputing (SC), 60 pages."},{"key":"e_1_3_3_15_2","unstructured":"B. Brech J. Rubio and M. Hollinger. 2015. Data Engine for NoSQL - IBM Power Systems Edition. White Paper."},{"key":"e_1_3_3_16_2","first-page":"1","volume-title":"Proceedings of the 2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO)","author":"Caulfield Adrian M.","year":"2016","unstructured":"Adrian M. Caulfield, Eric S. Chung, Andrew Putnam, Hari Angepat, Jeremy Fowers, Michael Haselman, Stephen Heil, Matt Humphrey, Puneet Kaur, Joo-Young Kim, et al. 2016. A cloud-scale acceleration architecture. In Proceedings of the 2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO), 1\u201313. DOI: 10.1109\/MICRO.2016.7783710"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/H2RC49586.2019.00010"},{"issue":"1","key":"e_1_3_3_18_2","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1109\/TNET.2014.2378012","article-title":"Discharging the network from its flow control headaches: Packet drops and HOL blocking","volume":"24","author":"Chrysos Nikolaos","year":"2015","unstructured":"Nikolaos Chrysos, Lydia Chen, Christoforos Kachris, and Manolis Katevenis. 2015. Discharging the network from its flow control headaches: Packet drops and HOL blocking. IEEE\/ACM Transactions on Networking 24, 1 (2015), 15\u201328.","journal-title":"IEEE\/ACM Transactions on Networking"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1504\/IJHPCN.2018.10015028"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.2172\/993908"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.420"},{"key":"e_1_3_3_22_2","first-page":"1","volume-title":"International Conference for High Performance Computing, Networking, Storage and Analysis (SC \u201920)","author":"De Sensi Daniele","year":"2020","unstructured":"Daniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth, and Torsten Hoefler. 2020. An in-depth analysis of the slingshot interconnect. In International Conference for High Performance Computing, Networking, Storage and Analysis (SC \u201920), 1\u201314. DOI: 10.1109\/SC41405.2020.00039"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/HOTI.2015.15"},{"key":"e_1_3_3_24_2","unstructured":"Digilent Inc. Fpga 2019. Digilent Inc. FPGA Microcontrollers and Instrumentation. Retrieved from http:\/\/www.digilent.com"},{"issue":"1","key":"e_1_3_3_25_2","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1177\/1094342015593158","article-title":"High-performance conjugate-gradient benchmark: A new metric for ranking high-performance computing systems","volume":"30","author":"Dongarra Jack","year":"2016","unstructured":"Jack Dongarra, Michael A. Heroux, and Piotr Luszczek. 2016. High-performance conjugate-gradient benchmark: A new metric for ranking high-performance computing systems. The International Journal of High Performance Computing Applications 30, 1 (2016), 3\u201310.","journal-title":"The International Journal of High Performance Computing Applications"},{"key":"e_1_3_3_26_2","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TPDS.2015.2407896","article-title":"Suitability analysis of FPGAs for heterogeneous platforms in HPC","volume":"27","author":"Escobar F. A.","year":"2016","unstructured":"F. A. Escobar, X. Chang, and C. Valderrama. 2016. Suitability analysis of FPGAs for heterogeneous platforms in HPC. IEEE Transactions on Parallel and Distributed Systems 27 (2016), 600\u2013612.","journal-title":"IEEE Transactions on Parallel and Distributed Systems"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE51398.2021.9474093"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-92792-3_6"},{"key":"e_1_3_3_29_2","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1109\/ISCA.2014.6853195","volume-title":"Proceedings of the 2014 ACM\/IEEE 41st International Symposium on Computer Architecture (ISCA)","author":"Putnam A.","year":"2014","unstructured":"A. Putnam, Adrian M. Caulfield, Eric S. Chung, D. Chiou, K. Constantinides, J. Demme, H. Esmaeilzadeh, J. Fowers, Gopi P. Gopal, J. Gray, et al. 2014. A reconfigurable fabric for accelerating large-scale datacenter services. In Proceedings of the 2014 ACM\/IEEE 41st International Symposium on Computer Architecture (ISCA), 13\u201324."},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/2996868"},{"key":"e_1_3_3_31_2","first-page":"1","volume-title":"IEEE Hot Chips Symposium","author":"Ouyang J.","year":"2014","unstructured":"J. Ouyang, S. Lin, W. Qi, Y. Wang, B. Yu, and S. Jiang. 2014. SDA: Software-defined accelerator for large-scale DNN systems. In IEEE Hot Chips Symposium, 1\u201323."},{"key":"e_1_3_3_32_2","unstructured":"M. Ploumidis F. Chaix N. Chrysos M. Assiminakis V. Flouris N. Kallimanis N. Kossifidis M. Nikoloudakis P. Petrakis N. Dimou et al. 2023. The ExaNeSt prototype: Evaluation of efficient HPC communication hardware in an ARM-based multi-FPGA rack. arXiv:2307.09371. Retrieved from https:\/\/arxiv.org\/abs\/2307.09371"},{"key":"e_1_3_3_33_2","doi-asserted-by":"crossref","first-page":"372","DOI":"10.1007\/978-3-642-12133-3_36","volume-title":"Proceedings of the 6th International Symposium Reconfigurable Computing: Architectures, Tools and Applications (ARC \u201910)","author":"Yoshimi M.","year":"2010","unstructured":"M. Yoshimi, Y. Nishikawa, M. Miki, T. Hiroyasu, H. Amano, and O. Mencer. 2010. A performance evaluation of CUBE: One-dimensional 512 FPGA cluster. In Proceedings of the 6th International Symposium Reconfigurable Computing: Architectures, Tools and Applications (ARC \u201910), 372\u2013381."},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.14529\/jsfi210105"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW50202.2020.00083"},{"key":"e_1_3_3_36_2","first-page":"1","volume-title":"Proceedings of the 2016 IEEE High Performance Extreme Computing Conference (HPEC)","author":"George Alan D.","year":"2016","unstructured":"Alan D. George, Martin C. Herbordt, Herman Lam, Abhijeet G. Lawande, Jiayi Sheng, and Chen Yang. 2016. Novo-G#: Large-scale reconfigurable computing with direct and programmable interconnects. In Proceedings of the 2016 IEEE High Performance Extreme Computing Conference (HPEC), 1\u20137. DOI: 10.1109\/HPEC.2016.7761639"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/NOCS.2018.8512155"},{"key":"e_1_3_3_38_2","unstructured":"SciEngines GmbH. 2019. SciEngines Hardware High Performance Reconfigurable Computing. Retrieved from https:\/\/www.sciengines.com\/technology-platform\/sciengines-hardware\/"},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTR.2006.311904"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/0167-8191(96)00024-5"},{"key":"e_1_3_3_41_2","unstructured":"Zhenhao He Dario Korolija Yu Zhu Benjamin Ramhorst Tristan Laan Lucian Petrica Michaela Blott and Gustavo Alonso. 2023. ACCL+: An FPGA-based collective engine for distributed applications. arXiv:2312.11742. Retrieved from https:\/\/arxiv.org\/abs\/2312.11742"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/H2RC54759.2021.00009"},{"key":"e_1_3_3_43_2","first-page":"1","volume-title":"Proceedings of the 2008 IEEE International Symposium on Parallel and Distributed Processing","author":"Hoefler Torsten","year":"2008","unstructured":"Torsten Hoefler, Timo Schneider, and Andrew Lumsdaine. 2008. Accurately measuring collective operations at massive scale. In Proceedings of the 2008 IEEE International Symposium on Parallel and Distributed Processing, 1\u20138. DOI: 10.1109\/IPDPS.2008.4536494"},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2016.2539167"},{"key":"e_1_3_3_45_2","unstructured":"Amazon.com Inc. 2019. Amazon EC2 F2 Instances. Retrieved from https:\/\/aws.amazon.com\/ec2\/instance-types\/f2\/"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3409115"},{"key":"e_1_3_3_47_2","volume-title":"Impacts of Operating Systems on the Scalability of Parallel Applications","author":"Jones T. R.","year":"2003","unstructured":"T. R. Jones, L. B. Brenner, J. M. Fier, Terry R. Jones, Larry B. Brenner, and Jeffrey M. Fier. 2003. Impacts of Operating Systems on the Scalability of Parallel Applications. Technical Report. Lawrence Livermore National Laboratory."},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSAMOS.2009.5289226"},{"key":"e_1_3_3_49_2","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1109\/DSD.2016.106","volume-title":"Proceedings of the 2016 Euromicro Conference on Digital System Design (DSD)","author":"Katevenis M.","year":"2016","unstructured":"M. Katevenis, N. Chrysos, M. Marazakis, I. Mavroidis, F. Chaix, N. Kallimanis, J. Navaridas, J. Goodacre, P. Vicini, A. Biagioni, et al. 2016. The ExaNeSt project: Interconnects, storage, and packaging for exascale systems. In Proceedings of the 2016 Euromicro Conference on Digital System Design (DSD), 60\u201367. DOI: 10.1109\/DSD.2016.106"},{"key":"e_1_3_3_50_2","volume-title":"The Future of Computing, Essays in Memory of Stamatis Vassiliadis","author":"Katevenis Manolis G. H.","year":"2007","unstructured":"Manolis G. H. Katevenis. 2007. Interprocessor communication seen as load-store instruction generalization. In The Future of Computing, Essays in Memory of Stamatis Vassiliadis. K. L. M. Bertels, S. Cotofana, G. N. Gaydadjiev, K. G. W. Goossens, S. Hamdioui, B. H. H. Juurlink, and H. J. van Gelderen (Eds.), Citeseer, Delft, The Netherlands."},{"key":"e_1_3_3_51_2","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1145\/3581576.3581602","volume-title":"Proceedings of the HPC Asia 2023 Workshops (HPCAsia \u201923 Workshops)","author":"Kikuchi Kohei","year":"2023","unstructured":"Kohei Kikuchi, Norihisa Fujita, Ryohei Kobayashi, and Taisuke Boku. 2023. Implementation and performance evaluation of collective communications using CIRCUS on multiple FPGAs. In Proceedings of the HPC Asia 2023 Workshops (HPCAsia \u201923 Workshops), 15\u201323. DOI: 10.1145\/3581576.3581602"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3149457.3149479"},{"issue":"1","key":"e_1_3_3_53_2","doi-asserted-by":"crossref","first-page":"30","DOI":"10.1109\/40.748794","article-title":"Implementation of ATLAS I: A single-chip ATM switch with backpressure","volume":"19","author":"Kornaros Georgios","year":"1998","unstructured":"Georgios Kornaros, Dionisios Pnevmatikatos, Panagiota Vatsolaki, Georgios Kalokerinos, Chara Xanthaki, Dimitrios Mavroidis, Dimitrios Serpanos, and Manolis Katevenis. 1998. Implementation of ATLAS I: A single-chip ATM switch with backpressure. In IEEE Micro 19, 1 (1998), 30\u201341.","journal-title":"IEEE Micro"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.3587"},{"key":"e_1_3_3_55_2","unstructured":"HiTech Global LLC. 2019. Xilinx\/Altera FPGA Boards Design Services & IP Cores. Retrieved from http:\/\/www.hitechglobal.com\/"},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.1997.569696"},{"key":"e_1_3_3_57_2","first-page":"1","volume-title":"Proceedings of the 2021 Symposium on VLSI Circuits","author":"Matsuoka Satoshi","year":"2021","unstructured":"Satoshi Matsuoka. 2021. Fugaku and A64FX: The first exascale supercomputer and its innovative arm CPU. In Proceedings of the 2021 Symposium on VLSI Circuits, 1\u20133. DOI: 10.23919\/VLSICircuits52068.2021.9492415"},{"key":"e_1_3_3_58_2","unstructured":"NVIDIA. 2024. NVIDIA Bluefield Networking Solution. Retrieved from https:\/\/www.nvidia.com\/en-us\/networking\/products\/data-processing-unit\/"},{"key":"e_1_3_3_59_2","unstructured":"NVIDIA. 2024. NVIDIA ConnectX InfiniBand Adapters. Retrieved from https:\/\/www.nvidia.com\/en-us\/networking\/infiniband-adapters\/"},{"key":"e_1_3_3_60_2","unstructured":"NVIDIA. 2024. The NVIDIA Quantum Infiniband Platform. Retrieved from https:\/\/www.nvidia.com\/en-us\/networking\/products\/infiniband\/"},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2018.08.240"},{"key":"e_1_3_3_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/1048935.1050204"},{"key":"e_1_3_3_63_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-34356-9_9"},{"key":"e_1_3_3_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2022.3175666"},{"key":"e_1_3_3_65_2","first-page":"1","volume-title":"Proceedings of the 2020 14th IEEE\/ACM International Symposium on Networks-on-Chip (NOCS)","author":"Psistakis Antonis","unstructured":"Antonis Psistakis, Nikos Chrysos, Fabien Chaix, Marios Asiminakis, Michalis Giannioudis, Pantelis Xirouchakis, Vassilis Papaefstathiou, and Manolis Katevenis. [n. d.]. PART: Pinning avoidance in RDMA technologies. In Proceedings of the 2020 14th IEEE\/ACM International Symposium on Networks-on-Chip (NOCS), 1\u20138."},{"key":"e_1_3_3_66_2","unstructured":"Atos SE. 2022. BullSequana X Supercomputers - Atos. Retrieved from https:\/\/atos.net\/en\/solutions\/high-performance-computing-hpc\/bullsequana-x-supercomputers"},{"key":"e_1_3_3_67_2","volume-title":"Proceedings of the 15th European Conference on Computer Systems (EuroSys)","author":"Sidler David","year":"2020","unstructured":"David Sidler, Zeke Wang, Monica Chiosa, Amit Kulkarni, and Gustavo Alonso. 2020. StRoM: Smart remote memory. In Proceedings of the 15th European Conference on Computer Systems (EuroSys). DOI: 10.1145\/3342195.3387519"},{"key":"e_1_3_3_68_2","doi-asserted-by":"publisher","DOI":"10.1147\/JRD.2020.2967330"},{"key":"e_1_3_3_69_2","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1109\/eScience.2019.00052","volume-title":"Proceedings of the 2019 15th International Conference on eScience (eScience)","author":"Taffoni Giuliano","year":"2019","unstructured":"Giuliano Taffoni, Luca Tornatore, David Goz, Antonio Ragagnin, Sara Bertocco, Igor Coretti, Manolis Marazakis, Fabien Chaix, Manolis Ploumidis, Manolis Katevenis, et al. 2019. Towards exascale: Measuring the energy footprint of astrophysics HPC simulations. In Proceedings of the 2019 15th International Conference on eScience (eScience), 403\u2013412. DOI: 10.1109\/eScience.2019.00052"},{"key":"e_1_3_3_70_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2021.108171"},{"key":"e_1_3_3_71_2","unstructured":"TRM 2018. Zynq UltraScale+ Device Technical Reference Manual. Retrieved from https:\/\/docs.amd.com\/r\/en-US\/ug1085-zynq-ultrascale-trm"},{"key":"e_1_3_3_72_2","doi-asserted-by":"publisher","DOI":"10.1145\/3579848"},{"key":"e_1_3_3_73_2","unstructured":"AMD Xilinx. 2016. Heterogeneous Accelerated Compute Clusters. Retrieved from https:\/\/www.amd-haccs.io\/index.html"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715152","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715152","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:18Z","timestamp":1750295898000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715152"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,12]]},"references-count":72,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,6,30]]}},"alternative-id":["10.1145\/3715152"],"URL":"https:\/\/doi.org\/10.1145\/3715152","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,12]]},"assertion":[{"value":"2024-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-19","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}