{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,11]],"date-time":"2026-08-11T19:28:05Z","timestamp":1786476485368,"version":"build-2736575974"},"publisher-location":"New York, NY, USA","reference-count":178,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,8,11]],"date-time":"2026-08-11T00:00:00Z","timestamp":1786406400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"National Science Foundation","award":["CNS-2212192"],"award-info":[{"award-number":["CNS-2212192"]}]},{"name":"National Science Foundation","award":["CAREER-2238608"],"award-info":[{"award-number":["CAREER-2238608"]}]},{"name":"National Science Foundation","award":["CAREER-2339755"],"award-info":[{"award-number":["CAREER-2339755"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,8,17]]},"DOI":"10.1145\/3789240.3829110","type":"proceedings-article","created":{"date-parts":[[2026,8,11]],"date-time":"2026-08-11T18:33:22Z","timestamp":1786473202000},"page":"696-713","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Understanding and Profiling the Accelerator Chiplet Network Using PingPoint"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-8788-7405","authenticated-orcid":false,"given":"Junyeol","family":"Ryu","sequence":"first","affiliation":[{"name":"University of Wisconsin-Madison, Madison, Wisconsin, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6509-9449","authenticated-orcid":false,"given":"Ming","family":"Liu","sequence":"additional","affiliation":[{"name":"University of Wisconsin--Madison, Madison, Wisconsin, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0189-7895","authenticated-orcid":false,"given":"Matthew D.","family":"Sinclair","sequence":"additional","affiliation":[{"name":"University of Wisconsin\u2013Madison, Madison, Wisconsin, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,8,11]]},"reference":[{"key":"e_1_3_2_1_1_1","first-page":"21","article-title":"Survey of Network on Chip (NoC) Architectures & Contributions","volume":"3","author":"Agarwal Ankur","year":"2009","unstructured":"Ankur Agarwal, Cyril Iskander, and Ravi Shankar. 2009. Survey of Network on Chip (NoC) Architectures & Contributions. Journal of Engineering, Computing and Architecture 3, 1 (2009), 21\u201327.","journal-title":"Journal of Engineering, Computing and Architecture"},{"key":"e_1_3_2_1_2_1","volume-title":"Proceedings of the IEEE INFOCOM. IEEE Press","author":"Agarwal Sugam","unstructured":"Sugam Agarwal, Murali Kodialam, and T. V. Lakshman. 2013. Traffic Engineering in Software Defined Networks. In Proceedings of the IEEE INFOCOM. IEEE Press, Piscataway, NJ, USA, 2211\u20132219."},{"key":"e_1_3_2_1_3_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery","author":"Agarwal Saksham","year":"2023","unstructured":"Saksham Agarwal, Arvind Krishnamurthy, and Rachit Agarwal. 2023. Host Congestion Control. In Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery, New York, NY, USA, 275\u2013287."},{"key":"e_1_3_2_1_4_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery","author":"Ahuja Satyajeet S.","year":"2021","unstructured":"Satyajeet S. Ahuja, Varun Gupta, Vinayak Dangui, Soshant Bali, Abishek Gopalan, Hao Zhong, Petr Lapukhov, Yiting Xia, and Ying Zhang. 2021. Capacity-Efficient and Uncertainty-Resilient Backbone Network Planning With Hose. In Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery, New York, NY, USA, 547\u2013559."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.comnet.2008.04.001"},{"key":"e_1_3_2_1_6_1","unstructured":"AMD. 2023. AMD CDNA\u2122 3 Architecture. https:\/\/www.amd.com\/content\/dam\/amd\/en\/documents\/instinct-tech-docs\/white-papers\/amd-cdna-3-white-paper.pdf. (2023)."},{"key":"e_1_3_2_1_7_1","unstructured":"AMD. 2023. AMD Instinct\u2122 MI300X Accelerators. https:\/\/www.amd.com\/en\/products\/accelerators\/instinct\/mi300\/mi300x.html. (2023)."},{"key":"e_1_3_2_1_8_1","unstructured":"AMD. 2023. AMD Instinct\u2122 MI350X Accelerators. https:\/\/www.amd.com\/en\/products\/accelerators\/instinct\/mi350\/mi350x.html. (2023)."},{"key":"e_1_3_2_1_9_1","unstructured":"AMD. 2024. AMD AI & HPC Fund: Accelerate Your Research with AMD. https:\/\/www.amd.com\/en\/corporate\/hpc-fund.html. (2024)."},{"key":"e_1_3_2_1_10_1","unstructured":"AMD. 2025. AMD CDNA\u2122 4 Architecture. https:\/\/www.amd.com\/content\/dam\/amd\/en\/documents\/instinct-tech-docs\/white-papers\/amd-cdna-4-architecture-whitepaper.pdf. (2025)."},{"key":"e_1_3_2_1_11_1","unstructured":"AMD. 2025. AMD Chiplet Ecosystem. https:\/\/www.amd.com\/content\/dam\/amd\/en\/documents\/solutions\/technologies\/chiplet-architecture-white-paper.pdf. (2025)."},{"key":"e_1_3_2_1_12_1","unstructured":"AMD. 2025. AMD Global Memory Interconnect. https:\/\/docs.amd.com\/r\/en-US\/pg292-ethernet-1-10-25g\/GMI-Transmission. (2025)."},{"key":"e_1_3_2_1_13_1","unstructured":"AMD. 2025. AMD Infinity Fabric\u2122 Link. https:\/\/docs.amd.com\/v\/u\/en-US\/AMD_Infinity_Fabric_Link_User_Guide_56978. (2025)."},{"key":"e_1_3_2_1_14_1","unstructured":"AMD. 2025. AMD Instinct MI300X GPU Partitioning Overview. https:\/\/instinct.docs.amd.com\/projects\/amdgpu-docs\/en\/latest\/gpu-partitioning\/mi300x\/overview.html. (2025)."},{"key":"e_1_3_2_1_15_1","unstructured":"AMD. 2025. AMD Instinct MI300X system optimization. https:\/\/instinct.docs.amd.com\/projects\/amdgpu-docs\/en\/docs-30. -optimization\/mi300x.html. (2025). system"},{"key":"e_1_3_2_1_16_1","unstructured":"AMD. 2025. Deep dive into the MI300 compute and memory partition modes. https:\/\/rocm.blogs.amd.com\/software-tools-optimization\/compute-memory-modes\/README.html. (2025)."},{"key":"e_1_3_2_1_17_1","unstructured":"AMD. 2025. MI300 and MI200 Series Performance Counters and Metrics. https:\/\/rocm.docs.amd.com\/en\/latest\/conceptual\/gpu-arch\/mi300-mi200-performance-counters.html. (2025)."},{"key":"e_1_3_2_1_18_1","unstructured":"AMD. 2026. Cooperative Groups. https:\/\/rocm.docs.amd.com\/projects\/HIP\/en\/latest\/tutorial\/cooperative_groups_tutorial.html. (2026)."},{"key":"e_1_3_2_1_19_1","unstructured":"AMD. 2026. MI300X L2 cache counters. https:\/\/github.com\/ROCm\/rocm-systems\/blob\/develop\/projects\/rocprofiler-compute\/src\/rocprof_compute_soc\/analysis_configs\/gfx942\/1700_l2_cache.yaml. (2026)."},{"key":"e_1_3_2_1_20_1","unstructured":"AMD. 2026. ROCm Compute Profiler documentation. https:\/\/rocm.docs.amd.com\/projects\/rocprofiler-compute\/en\/latest\/. (2026)."},{"key":"e_1_3_2_1_21_1","unstructured":"AMD. 2026. ROCprofiler-SDK documentation. https:\/\/rocm.docs.amd.com\/projects\/rocprofiler-sdk\/en\/latest\/. (2026)."},{"key":"e_1_3_2_1_22_1","unstructured":"AMD. 2026. Setting the number of compute units. https:\/\/rocm.docs.amd.com\/en\/latest\/how-to\/setting-cus.html. (2026)."},{"key":"e_1_3_2_1_23_1","volume-title":"Proceedings of the ACM Workshop on Hot Topics in Networks. Association for Computing Machinery","author":"An Seunghyun","year":"2025","unstructured":"Seunghyun An, Joontaek Oh, and Ming Liu. 2025. Server Chiplet Networking. In Proceedings of the ACM Workshop on Hot Topics in Networks. Association for Computing Machinery, New York, NY, USA, 289\u2013299."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015467.1015499"},{"key":"e_1_3_2_1_25_1","volume-title":"Proceedings of the 44th Annual International Symposium on Computer Architecture. Association for Computing Machinery","author":"Arunkumar Akhil","year":"2017","unstructured":"Akhil Arunkumar, Evgeny Bolotin, Benjamin Cho, Ugljesa Milic, Eiman Ebrahimi, Oreste Villa, Aamer Jaleel, Carole-Jean Wu, and David Nellans. 2017. MCM-GPU: Multi-Chip-Module GPUs for Continued Performance Scalability. In Proceedings of the 44th Annual International Symposium on Computer Architecture. Association for Computing Machinery, New York, NY, USA, 320\u2013332."},{"key":"e_1_3_2_1_26_1","volume-title":"Proceedings of the 29th IEEE Real-Time and Embedded Technology and Applications Symposium. IEEE Press, IEEE","author":"Bakita Joshua","unstructured":"Joshua Bakita and James H. Anderson. 2023. Hardware compute partitioning on NVIDIA GPUs. In Proceedings of the 29th IEEE Real-Time and Embedded Technology and Applications Symposium. IEEE Press, IEEE, Piscataway, NJ, USA, 54\u201366."},{"key":"e_1_3_2_1_27_1","volume-title":"Proceedings of the 20th Annual International Conference on Supercomputing. Association for Computing Machinery","author":"Balfour James","year":"2006","unstructured":"James Balfour and William J Dally. 2006. Design Tradeoffs for Tiled CMP Onchip Networks. In Proceedings of the 20th Annual International Conference on Supercomputing. Association for Computing Machinery, New York, NY, USA, 390\u2013401."},{"key":"e_1_3_2_1_28_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery","author":"Ballani Hitesh","year":"2011","unstructured":"Hitesh Ballani, Paolo Costa, Thomas Karagiannis, and Ant Rowstron. 2011. Towards Predictable Datacenter Networks. In Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery, New York, NY, USA, 242\u2013253."},{"key":"e_1_3_2_1_29_1","volume-title":"Proceedings of the 26th IEEE International Symposium on High Performance Computer Architecture. IEEE Press","author":"Baruah Trinayan","year":"2020","unstructured":"Trinayan Baruah, Yifan Sun, Ali Tolga Din\u00e7er, Saiful A. Mojumder, Jos\u00e9 L. Abell\u00e1n, Yash Ukidave, Ajay Joshi, Norman Rubin, John Kim, and David Kaeli. 2020. Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU Systems. In Proceedings of the 26th IEEE International Symposium on High Performance Computer Architecture. IEEE Press, Piscataway, NJ, USA, 596\u2013609."},{"key":"e_1_3_2_1_30_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery","author":"Basat Ran Ben","year":"2020","unstructured":"Ran Ben Basat, Sivaramakrishnan Ramanathan, Yuliang Li, Gianni Antichi, Minian Yu, and Michael Mitzenmacher. 2020. PINT: Probabilistic in-band network telemetry. In Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery, New York, NY, USA, 662\u2013680."},{"key":"e_1_3_2_1_31_1","volume-title":"Proceedings of the SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC). Association for Computing Machinery","author":"Bertolli Carlo","year":"2024","unstructured":"Carlo Bertolli, Thorsten Blass, Lynd Stringer, Nicole Aschenbrenner, Jan-Patrick Lehr, Doru Bercea, Dhruva Chakrabarti, Lawrence Meadows, and Ron Lieberman. 2024. Performance Analysis of Runtime Handling of Zero-Copy for OpenMP Programs on MI300A APUs. In Proceedings of the SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC). Association for Computing Machinery, New York, NY, USA, 1420\u20131429."},{"key":"e_1_3_2_1_32_1","volume-title":"Proceedings of the 32nd Annual ACM\/IEEE International Symposium on Microarchitecture. IEEE Press, IEEE","author":"Bharadwaj Jay","year":"1999","unstructured":"Jay Bharadwaj, Kishore Menezes, and Chris McKinsey. 1999. Wavefront Scheduling: Path-based Data Representation and Scheduling of Subgraphs. In Proceedings of the 32nd Annual ACM\/IEEE International Symposium on Microarchitecture. IEEE Press, IEEE, Piscataway, NJ, USA, 262\u2013271."},{"key":"e_1_3_2_1_33_1","first-page":"8","article-title":"AMD Next-Generation \"Zen 4","volume":"44","author":"Bhargava Ravi","year":"2024","unstructured":"Ravi Bhargava and Kai Troester. 2024. AMD Next-Generation \"Zen 4\" Core and 4th Gen AMD EPYC Server CPUs. IEEE Micro 44, 3 (2024), 8\u201317.","journal-title":"Core and 4th Gen AMD EPYC Server CPUs. IEEE Micro"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/1132952.1132953"},{"key":"e_1_3_2_1_35_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery","author":"Cai Qizhe","year":"2021","unstructured":"Qizhe Cai, Shubham Chaudhary, Midhul Vuppalapati, Jaehyun Hwang, and Rachit Agarwal. 2021. Understanding Host Network Stack Overheads. In Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery, New York, NY, USA, 65\u201377."},{"key":"e_1_3_2_1_36_1","unstructured":"Erlangen National High Performance Center. 2025. GPU benchmarks. https:\/\/github.com\/RRZE-HPC\/gpu-benches. (2025)."},{"key":"e_1_3_2_1_37_1","volume-title":"Proceedings of the 4th ACM\/IEEE International Symposium on Networks-on-Chip. IEEE Press, IEEE","author":"Chao Chih-Hao","year":"2010","unstructured":"Chih-Hao Chao, Kai-Yuan Jheng, Hao-Yu Wang, Jia-Cheng Wu, and An-Yeu Wu. 2010. Traffic-and Thermal-aware Run-time Thermal Management Scheme for 3D NoC Systems. In Proceedings of the 4th ACM\/IEEE International Symposium on Networks-on-Chip. IEEE Press, IEEE, Piscataway, NJ, USA, 223\u2013230."},{"key":"e_1_3_2_1_38_1","volume-title":"Demystifying Datapath Accelerator Enhanced Off-path SmartNIC. In 2024 IEEE 32nd International Conference on Network Protocols (ICNP'24)","author":"Chen Xuzheng","year":"2024","unstructured":"Xuzheng Chen, Jie Zhang, Ting Fu, Yifan Shen, Shu Ma, Kun Qian, Lingjun Zhu, Chao Shi, Yin Zhang, Ming Liu, and Zeke Wang. 2024. Demystifying Datapath Accelerator Enhanced Off-path SmartNIC. In 2024 IEEE 32nd International Conference on Network Protocols (ICNP'24). 1\u201312."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"crossref","first-page":"1455","DOI":"10.1016\/j.fmre.2023.04.014","article-title":"Challenges and Prospects for Advanced Packaging","volume":"4","author":"Chen Zhiwen","year":"2024","unstructured":"Zhiwen Chen, Jiaju Zhang, Shizhao Wang, and Ching-Ping Wong. 2024. Challenges and Prospects for Advanced Packaging. Fundamental Research 4, 6 (2024), 1455\u20131458.","journal-title":"Fundamental Research"},{"key":"e_1_3_2_1_40_1","unstructured":"Chips and Cheese. 2023. AMD's CDNA3 Compute Architecture. https:\/\/chipsandcheese.com\/p\/amds-cdna-3-compute-architecture. (2023)."},{"key":"e_1_3_2_1_41_1","unstructured":"Chips and Cheese. 2025. Inside the AMD Instinct MI300A's Giant Memory Subsystem. https:\/\/chipsandcheese.com\/p\/inside-the-amd-radeon-instinct-mi300as. (2025)."},{"key":"e_1_3_2_1_42_1","volume-title":"FARM: Fault-Aware Resource Management in NoC-based Multiprocessor Platforms. In Design, Automation & Test in Europe","author":"Chou Chen-Ling","year":"2011","unstructured":"Chen-Ling Chou and Radu Marculescu. 2011. FARM: Fault-Aware Resource Management in NoC-based Multiprocessor Platforms. In Design, Automation & Test in Europe. IEEE Press, Piscataway, NJ, USA, 1\u20136."},{"key":"e_1_3_2_1_43_1","volume-title":"Lisa Wu Wills, and Ganesh Dasika","author":"Choudhary Mansi","year":"2025","unstructured":"Mansi Choudhary, Karthik Sangaiah, Sonali Singh, Muhammad Osama, Lisa Wu Wills, and Ganesh Dasika. 2025. Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects. arXiv preprint arXiv:2511.02132 (2025), 11."},{"key":"e_1_3_2_1_44_1","volume-title":"Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs. arXiv preprint arXiv:2604.15379","author":"Chowdhary Sangeeta","year":"2026","unstructured":"Sangeeta Chowdhary, Ryan Swann, Sean Siddens, Muhammad Osama, Stephen Neuendorffer, Alexandru Dutu, Karthik Sangaiah, Sandeepa Bhuyan, Samuel Bayliss, and Ganesh Dasika. 2026. Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs. arXiv preprint arXiv:2604.15379 (2026), 12."},{"key":"e_1_3_2_1_45_1","unstructured":"Google Cloud. 2026. Five techniques to reach the efficient frontier of LLM inference. https:\/\/cloud.google.com\/blog\/topics\/developers-practitioners\/five-techniques-to-reach-the-efficient-frontier-of-llm-inference. (2026)."},{"key":"e_1_3_2_1_46_1","unstructured":"GitHub Commit. 2023. Remove non-functional counters for MI200 and MI300. https:\/\/github.com\/ROCm\/rocm-systems\/commit\/1471dc8a7640687c82197b2e8d0f65c63b1adfa9. (2023)."},{"key":"e_1_3_2_1_47_1","unstructured":"UALink Consortium. 2025. Ultra Accelerator Link Specification. https:\/\/ualinkconsortium.org\/specifications\/. (2025)."},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3731569.3764818"},{"key":"e_1_3_2_1_49_1","volume-title":"Principles and practices of interconnection networks","author":"Dally William James","unstructured":"William James Dally and Brian Patrick Towles. 2004. Principles and practices of interconnection networks. Elsevier, New York, NY."},{"key":"e_1_3_2_1_50_1","unstructured":"Voltron Data. 2026. Theseus. https:\/\/github.com\/voltrondata. (2026)."},{"key":"e_1_3_2_1_51_1","unstructured":"DeepSeek. 2025. DeepSeek-V3. https:\/\/www.deepseek.com\/en\/. (2025)."},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1109\/MSSC.2019.2910619","article-title":"Ultra-Short-Reach Interconnects for Die-to-Die Links: Global Bandwidth Demands in Microcosm","volume":"11","author":"Dehlaghi Behzad","year":"2019","unstructured":"Behzad Dehlaghi, Nijwm Wary, and Tony Chan Carusone. 2019. Ultra-Short-Reach Interconnects for Die-to-Die Links: Global Bandwidth Demands in Microcosm. IEEE Solid-State Circuits Magazine 11, 2 (2019), 42\u201353.","journal-title":"IEEE Solid-State Circuits Magazine"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2011.2181509"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/316188.316209"},{"key":"e_1_3_2_1_55_1","volume-title":"Proceedings of the SC23-W: Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis. Association for Computing Machinery","author":"Eiling Niklas","year":"2023","unstructured":"Niklas Eiling, Stefan Lankes, and Antonello Monti. 2023. Checkpoint\/Restart for CUDA Kernels. In Proceedings of the SC23-W: Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis. Association for Computing Machinery, New York, NY, USA, 1729\u20131737."},{"key":"e_1_3_2_1_56_1","first-page":"1","article-title":"Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity","volume":"23","author":"Fedus William","year":"2022","unstructured":"William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. Journal of Machine Learning Research 23, 120 (2022), 1\u201339.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3613424.3614310"},{"key":"e_1_3_2_1_58_1","first-page":"57","article-title":"Bufferbloat","volume":"55","author":"Nichols Jim","year":"2012","unstructured":"Gettys, Jim and Nichols, Kathleen. 2012. Bufferbloat: Dark Buffers in the Internet. Commun. ACM 55, 1 (2012), 57\u201365.","journal-title":"Dark Buffers in the Internet. Commun. ACM"},{"key":"e_1_3_2_1_59_1","volume-title":"Proceedings of the 1st International Symposium on Networks-on-Chip. IEEE Press","author":"Gindin Roman","year":"2007","unstructured":"Roman Gindin, Israel Cidon, and Idit Keidar. 2007. NoC-based FPGA: Architecture and Routing. In Proceedings of the 1st International Symposium on Networks-on-Chip. IEEE Press, Piscataway, NJ, USA, 253\u2013264."},{"key":"e_1_3_2_1_60_1","unstructured":"Google. 2025. Gemini 3. https:\/\/gemini.google.com\/. (2025)."},{"key":"e_1_3_2_1_61_1","volume-title":"Proceedings of the 15th IEEE International Symposium on High Performance Computer Architecture. IEEE Press","author":"Grot Boris","year":"2009","unstructured":"Boris Grot, Joel Hestness, Stephen W. Keckler, and Onur Mutlu. 2009. Express Cube Topologies for on-Chip Interconnects. In Proceedings of the 15th IEEE International Symposium on High Performance Computer Architecture. IEEE Press, Piscataway, NJ, USA, 163\u2013174."},{"key":"e_1_3_2_1_62_1","volume-title":"Proceedings of the ACM SIGCOMM Conference (SIGCOMM). ACM","author":"Guo Chuanxiong","year":"2016","unstructured":"Chuanxiong Guo, Haitao Wu, Zhong Deng, Gaurav Soni, Jianxi Ye, Jitu Padhye, and Marina Lipshteyn. 2016. RDMA over Commodity Ethernet at Scale. In Proceedings of the ACM SIGCOMM Conference (SIGCOMM). ACM, New York, NY, USA, 202\u2013215."},{"key":"e_1_3_2_1_63_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. ACM","author":"Guo Chuanxiong","year":"2015","unstructured":"Chuanxiong Guo, Lihua Yuan, Dong Xiang, Yingnong Dang, Ray Huang, Dave Maltz, Zhaoyi Liu, Vin Wang, Bin Pang, Hua Chen, Zhi-Wei Lin, and Varugis Kurien. 2015. Pingmesh: A Large-scale System for Data Center Network Latency Measurement and Analysis. In Proceedings of the ACM SIGCOMM Conference. ACM, New York, NY, USA, 139\u2013152."},{"key":"e_1_3_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/3613424.3614291"},{"key":"e_1_3_2_1_65_1","volume-title":"Building A CSFQ-Inspired Transport for Switched CXL Memory Pooling. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI'26)","author":"Guo Zerui","year":"2026","unstructured":"Zerui Guo, Emily Shriver, and Ming Liu. 2026. Building A CSFQ-Inspired Transport for Switched CXL Memory Pooling. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI'26). 969\u2013986."},{"key":"e_1_3_2_1_66_1","volume-title":"Proceedings of the 8th USENIX Symposium on Operating Systems Design and Implementation. USENIX Association, USA, 193\u2013208","author":"Guo Zhenyu","year":"2008","unstructured":"Zhenyu Guo, Xi Wang, Jian Tang, Xuezheng Liu, Zhilei Xu, Ming Wu, M. Frans Kaashoek, and Zheng Zhang. 2008. R2: An Application-Level Kernel for Record and Replay. In Proceedings of the 8th USENIX Symposium on Operating Systems Design and Implementation. USENIX Association, USA, 193\u2013208."},{"key":"e_1_3_2_1_67_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery","author":"Gupta Arpit","year":"2018","unstructured":"Arpit Gupta, Rob Harrison, Marco Canini, Nick Feamster, Jennifer Rexford, and Walter Willinger. 2018. Sonata: Query-Driven Streaming Network Telemetry. In Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery, New York, NY, USA, 357\u2013371."},{"key":"e_1_3_2_1_68_1","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. Association for Computing Machinery","author":"Hanindhito Bagus","year":"2025","unstructured":"Bagus Hanindhito and Bhavesh Patel. 2025. Characterizing Performance, Power, and Energy of AMD CDNA3 GPU Family. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. Association for Computing Machinery, New York, NY, USA, 905\u2013934."},{"key":"e_1_3_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3575693.3575708"},{"key":"e_1_3_2_1_70_1","volume-title":"Proceedings of the Asia and South Pacific Design Automation Conference. IEEE Press","author":"Hu Jingcao","year":"2003","unstructured":"Jingcao Hu and Radu Marculescu. 2003. Energy-aware Mapping for Tile-based NoC Architectures Under Performance Constraints. In Proceedings of the Asia and South Pacific Design Automation Conference. IEEE Press, Piscataway, NJ, USA, 233\u2013239."},{"key":"e_1_3_2_1_71_1","volume-title":"HipKittens: Fast and Furious AMD Kernels. arXiv preprint arXiv:2511.08083","author":"Hu William","year":"2025","unstructured":"William Hu, Drew Wadsworth, Sean Siddens, Stanley Winata, Daniel Y. Fu, Ryann Swann, Muhammad Osama, Christopher R\u00e9, and Simran Arora. 2025. HipKittens: Fast and Furious AMD Kernels. arXiv preprint arXiv:2511.08083 (2025), 41."},{"key":"e_1_3_2_1_72_1","volume-title":"Proceedings of the 72nd IEEE Electronic Components and Technology Conference. IEEE Press","author":"Hwang Yoonjae","year":"2022","unstructured":"Yoonjae Hwang, Sungwook Moon, Seungki Nam, and Jeong HoonAhn. 2022. Chiplet-based System PSI Optimization for 2.5D\/3D Advanced Packaging Implementation. In Proceedings of the 72nd IEEE Electronic Components and Technology Conference. IEEE Press, Piscataway, NJ, USA, 12\u201317."},{"key":"e_1_3_2_1_73_1","volume-title":"Opportunities, Applications. https:\/\/www.idtechex.com\/en\/research-report\/chiplet-technology-2025\/1041.","author":"Ex.","year":"2025","unstructured":"IDTechEx. 2025. Chiplet Technology 2025\u20132035: Technology, Opportunities, Applications. https:\/\/www.idtechex.com\/en\/research-report\/chiplet-technology-2025\/1041. (2025)."},{"key":"e_1_3_2_1_74_1","volume-title":"Chiplets: Piecing Together the Next Generation of Chips. https:\/\/www.imec-int.com\/en\/articles\/chiplets-piecing-together-next-generation-chips-part-i.","year":"2025","unstructured":"imec. 2025. Chiplets: Piecing Together the Next Generation of Chips. https:\/\/www.imec-int.com\/en\/articles\/chiplets-piecing-together-next-generation-chips-part-i. (2025)."},{"key":"e_1_3_2_1_75_1","unstructured":"Intel. 2025. Intel's Gaudi 3 AI Accelerators. https:\/\/www.intel.com\/content\/www\/us\/en\/products\/details\/processors\/ai-accelerators\/gaudi.html. (2025)."},{"key":"e_1_3_2_1_76_1","unstructured":"Intel. 2025. The Intel VTune Profiler. https:\/\/www.intel.com\/content\/www\/us\/en\/developer\/tools\/oneapi\/vtune-profiler.html. (2025)."},{"key":"e_1_3_2_1_77_1","unstructured":"GitHub Issue. 2023. Obtain CU physical ID from HIP kernel? https:\/\/github.com\/ROCm\/ROCm\/issues\/2059. (2023)."},{"key":"e_1_3_2_1_78_1","unstructured":"GitHub Issue. 2025. Matmul performance issue on MI300X. https:\/\/github.com\/triton-lang\/triton\/issues\/4959. (2025)."},{"key":"e_1_3_2_1_79_1","unstructured":"GitHub Issue. 2025. MI300X SCLK stuck around 60MHz under full load. https:\/\/github.com\/ROCm\/ROCm\/issues\/4414. (2025)."},{"key":"e_1_3_2_1_80_1","unstructured":"GitHub Issue. 2025. RCCL performance degradation on AMD Instinct MI300X GPU with AMD Pollara AI NIC. https:\/\/github.com\/ROCm\/ROCm\/issues\/5717. (2025)."},{"key":"e_1_3_2_1_81_1","unstructured":"GitHub Issue. 2025. rccl tests on two Mi300x nodes gives me poor performance using mpirun. https:\/\/github.com\/ROCm\/rccl-tests\/issues\/146. (2025)."},{"key":"e_1_3_2_1_82_1","unstructured":"Charles Jamieson Anushka Chandrashekar Ian McDougall and M. D. Sinclair. 2022. GAP: gem5 GPU Accuracy Profiler. In 4th gem5 Users' Workshop. 3."},{"key":"e_1_3_2_1_83_1","unstructured":"JEDEC. 2025. High Bandwidth Memory (HBM3) DRAM. https:\/\/www.jedec.org\/standards-documents\/docs\/jesd238b01. (2025)."},{"key":"e_1_3_2_1_84_1","volume-title":"Proceedings of the 57th IEEE\/ACM International Symposium on Microarchitecture. IEEE Press","author":"Jin Zhixian","year":"2024","unstructured":"Zhixian Jin, Christopher Rocca, Jiho Kim, Hans Kasan, Minsoo Rhu, Ali Bakhoda, Tor M. Aamodt, and John Kim. 2024. Uncovering Real GPU NoC Characteristics: Implications on Interconnect Architecture. In Proceedings of the 57th IEEE\/ACM International Symposium on Microarchitecture. IEEE Press, Piscataway, NJ, USA, 885\u2013898."},{"key":"e_1_3_2_1_85_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750392"},{"key":"e_1_3_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177728975"},{"key":"e_1_3_2_1_87_1","unstructured":"Keysight. 2025. What is a Chiplet and Why Should You Care? https:\/\/www.keysight.com\/blogs\/en\/tech\/sim-des\/what-is-a-chiplet-and-why-should-you-care. (2025)."},{"key":"e_1_3_2_1_88_1","volume-title":"Proceedings of the IEEE Hot Chips 36 Symposium. IEEE Press","author":"Khailany Brucek","year":"2024","unstructured":"Brucek Khailany and Jonah Alben. 2024. NVIDIA Blackwell Platform: Advancing Generative AI and Accelerated Computing. In Proceedings of the IEEE Hot Chips 36 Symposium. IEEE Press, Piscataway, NJ, USA, 1\u201333."},{"key":"e_1_3_2_1_89_1","volume-title":"Proceedings of the 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE Press","author":"Khairy Mahmoud","unstructured":"Mahmoud Khairy, Vadim Nikiforov, David Nellans, and Timothy G. Rogers. 2020. Locality-Centric Data and Threadblock Management for Massive GPUs. In Proceedings of the 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE Press, Piscataway, NJ, USA, 1022\u20131036."},{"key":"e_1_3_2_1_90_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00047"},{"key":"e_1_3_2_1_91_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery","author":"Kim Changhoon","unstructured":"Changhoon Kim, Anirudh Sivaraman, Naga Katta, Antonin Bas, Advait Dixit, and Lawrence J. Wobker. 2015. In-band Network Telemetry via Programmable Dataplanes. In Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery, New York, NY, USA, 1\u20132."},{"key":"e_1_3_2_1_92_1","doi-asserted-by":"publisher","DOI":"10.1109\/MCOM.2013.6461195"},{"key":"e_1_3_2_1_93_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.15"},{"key":"e_1_3_2_1_94_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2008.19"},{"key":"e_1_3_2_1_95_1","doi-asserted-by":"publisher","DOI":"10.1145\/1065579.1065726"},{"key":"e_1_3_2_1_96_1","unstructured":"Jayacharan Kolla Pedram Alizadeh and Gilbert Lee. 2025. Understanding RCCL Bandwidth and xGMI Performance on AMD Instinct MI300X. https:\/\/rocm.blogs.amd.com\/software-tools-optimization\/mi300x-rccl-xgmi\/README.html. (2025)."},{"key":"e_1_3_2_1_97_1","volume-title":"Averill","author":"Kumar Ganesh","year":"2012","unstructured":"Ganesh Kumar, Dheemanth Nagaraj, Vincent R. Freytag, Eric Delano, and Gregory S. Averill. 2012. Cache and Home Agent Memory Management. https:\/\/patents.google.com\/patent\/US8327228B2\/en. (2012)."},{"key":"e_1_3_2_1_98_1","volume-title":"Proceedings of the Conference on Communications Architectures, Protocols and Applications. Association for Computing Machinery","author":"Kung HT","year":"1994","unstructured":"HT Kung, Trevor Blackwell, and Alan Chapman. 1994. Credit-based Flow Control for ATM Networks: Credit Update Protocol, Adaptive Credit Allocation and Statistical Multiplexing. In Proceedings of the Conference on Communications Architectures, Protocols and Applications. Association for Computing Machinery, New York, NY, USA, 101\u2013114."},{"key":"e_1_3_2_1_99_1","volume-title":"Proceedings of the IEEE INFOCOM. IEEE Press","author":"Kung HT","year":"1995","unstructured":"HT Kung and Koling Chang. 1995. Receiver-Oriented Adaptive Buffer Allocation in Credit-Based Flow Control for ATM Networks. In Proceedings of the IEEE INFOCOM. IEEE Press, Piscataway, NJ, USA, 239\u2013252."},{"key":"e_1_3_2_1_100_1","doi-asserted-by":"publisher","DOI":"10.1109\/65.372658"},{"key":"e_1_3_2_1_101_1","doi-asserted-by":"crossref","first-page":"228","DOI":"10.1109\/TCPMT.2022.3144461","article-title":"Recent Advances and Trends in Advanced Packaging","volume":"12","author":"Lau John H.","year":"2022","unstructured":"John H. Lau. 2022. Recent Advances and Trends in Advanced Packaging. IEEE Transactions on Components, Packaging and Manufacturing Technology 12, 2 (2022), 228\u2013252.","journal-title":"IEEE Transactions on Components, Packaging and Manufacturing Technology"},{"key":"e_1_3_2_1_102_1","doi-asserted-by":"crossref","first-page":"651","DOI":"10.1109\/TCPMT.2025.3533926","article-title":"Current Advances and Outlooks in Hybrid Bonding","volume":"15","author":"Lau John H.","year":"2025","unstructured":"John H. Lau. 2025. Current Advances and Outlooks in Hybrid Bonding. IEEE Transactions on Components, Packaging and Manufacturing Technology 15, 4 (2025), 651\u2013681.","journal-title":"IEEE Transactions on Components, Packaging and Manufacturing Technology"},{"key":"e_1_3_2_1_103_1","volume-title":"Proceedings of the 23rd USENIX Symposium on Networked Systems Design and Implementation. USENIX Association, USA, 2515\u20132531","author":"Lei Yiran","year":"2026","unstructured":"Yiran Lei, Dongjoo Lee, Liangyu Zhao, Daniar Kurniawan, Chanmyeong Kim, Heetaek Jeong, Changsu Kim, Hyeonseong Choi, Liangcheng Yu, Arvind Krishnamurthy, Justine Sherry, and Eriko Nurvitadhi. 2026. FAST: An Efficient Scheduler for All-to-All GPU Communication. In Proceedings of the 23rd USENIX Symposium on Networked Systems Design and Implementation. USENIX Association, USA, 2515\u20132531."},{"key":"e_1_3_2_1_104_1","volume-title":"Proceedings of the 19th USENIX Symposium on Operating Systems Design and Implementation. USENIX Association, USA, 575\u2013593","author":"Li Ao","year":"2025","unstructured":"Ao Li, Marion Sudvarg, Zihan Li, Sanjoy Baruah, Chris Gill, and Ning Zhang. 2025. Tintin: A Unified Hardware Performance Profiling Infrastructure to Uncover and Manage Uncertainty. In Proceedings of the 19th USENIX Symposium on Operating Systems Design and Implementation. USENIX Association, USA, 575\u2013593."},{"key":"e_1_3_2_1_105_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2010.10.013"},{"key":"e_1_3_2_1_106_1","doi-asserted-by":"publisher","DOI":"10.1109\/OJSSCS.2024.3506694"},{"key":"e_1_3_2_1_107_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery","author":"Li Xiao","year":"2025","unstructured":"Xiao Li, Zerui Guo, Yuebin Bai, Mahesh Ketkar, Hugh Wilkinson, and Ming Liu. 2025. Understanding and Profiling CXL.mem Using PathFinder. In Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery, New York, NY, USA, 758\u2013779."},{"key":"e_1_3_2_1_108_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery","author":"Liu Bowen","year":"2025","unstructured":"Bowen Liu, Xinyang Huang, Qijing Li, Zhuobin Huang, Yijun Sun, Wenxue Li, Junxue Zhang, Ping Yin, and Kai Chen. 2025. CEIO: A Cache-Efficient Network I\/O Architecture for NIC-CPU Data Paths. In Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery, New York, NY, USA, 381\u2013394."},{"key":"e_1_3_2_1_109_1","doi-asserted-by":"publisher","DOI":"10.1145\/3676641.3715987"},{"key":"e_1_3_2_1_110_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery","author":"Liu Kefei","year":"2024","unstructured":"Kefei Liu, Zhuo Jiang, Jiao Zhang, Shixian Guo, Xuan Zhang, Yangyang Bai, Yongbin Dong, Feng Luo, Zhang Zhang, Lei Wang, Xiang Shi, Haohan Xu, Yang Bai, Dongyang Song, Haoran Wei, Bo Li, Yongchen Pan, Tian Pan, and Tao Huang. 2024. R-Pingmesh: A Service-aware ROCe Network Monitoring and Diagnostic System. In Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery, New York, NY, USA, 554\u2013567."},{"key":"e_1_3_2_1_111_1","volume-title":"Building Distributed Systems Using Programmable Networks","author":"Liu Ming","unstructured":"Ming Liu. 2020. Building Distributed Systems Using Programmable Networks. University of Washington."},{"key":"e_1_3_2_1_112_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341302.3342079"},{"key":"e_1_3_2_1_113_1","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037731"},{"key":"e_1_3_2_1_114_1","volume-title":"2019 USENIX Annual Technical Conference (USENIX ATC'19)","author":"Liu Ming","year":"2019","unstructured":"Ming Liu, Simon Peter, Arvind Krishnamurthy, and Phitchaya Mangpo Phothilimthana. 2019. E3: Energy-Efficient Microservices on SmartNIC-Accelerated Servers. In 2019 USENIX Annual Technical Conference (USENIX ATC'19). 363\u2013378."},{"key":"e_1_3_2_1_115_1","volume-title":"Proceedings of the IEEE Computer Society Annual Symposium on VLSI. IEEE Press, IEEE","author":"Liu Weichen","year":"2011","unstructured":"Weichen Liu, Jiang Xu, XiaowenWu, Yaoyao Ye, Xuan Wang, Wei Zhang, Mahdi Nikdast, and Zhehui Wang. 2011. A NoC Traffic Suite Based on Real Applications. In Proceedings of the IEEE Computer Society Annual Symposium on VLSI. IEEE Press, IEEE, Piscataway, NJ, USA, 66\u201371."},{"key":"e_1_3_2_1_116_1","unstructured":"LLVM. 2026. HW_REG_XCC_ID. https:\/\/llvm.org\/docs\/AMDGPU\/gfx940_hwreg.html. (2026)."},{"key":"e_1_3_2_1_117_1","volume-title":"Workshop on Approximate Computing Across the Stack.","author":"Luo Liang","year":"2017","unstructured":"Liang Luo, Ming Liu, Jacob Nelson, Luis Ceze, Amar Phanishayee, and Arvind Krishnamurthy. 2017. Motivating in-network aggregation for distributed deep neural network training. In Workshop on Approximate Computing Across the Stack."},{"key":"e_1_3_2_1_118_1","volume-title":"Proceedings of the 74th IEEE Electronic Components and Technology Conference. IEEE Press","author":"Mandalapu Chandra Sekhar","year":"2024","unstructured":"Chandra Sekhar Mandalapu, Chintan Buch, Priyal Shah, Roden Topacio, Patrick Cheng, Liwei Wang, Raja Swaminathan, Alan Smith, John Wuu, Kaushik Mysore, and Arsalan Alam. 2024. 3.5D Advanced Packaging Enabling Heterogenous Integration of HPC and AI Accelerators. In Proceedings of the 74th IEEE Electronic Components and Technology Conference. IEEE Press, Piscataway, NJ, USA, 798\u2013802."},{"key":"e_1_3_2_1_119_1","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1109\/TCAD.2008.2010691","article-title":"Outstanding Research Problems in NoC Design: System, Microarchitecture, and Circuit Perspectives","volume":"28","author":"Marculescu Radu","year":"2008","unstructured":"Radu Marculescu, Umit Y. Ogras, Li-Shiuan Peh, Natalie E. Jerger, and Yatin Hoskote. 2008. Outstanding Research Problems in NoC Design: System, Microarchitecture, and Circuit Perspectives. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 28, 1 (2008), 3\u201321.","journal-title":"IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems"},{"key":"e_1_3_2_1_120_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920886"},{"key":"e_1_3_2_1_121_1","volume-title":"Proceedings of the ACM symposium on Applied Computing. Association for Computing Machinery","author":"Moreira Orlando","year":"2007","unstructured":"Orlando Moreira, Jacob Jan-David Mol, and Marco Bekooij. 2007. Online Resource Management in a Multiprocessor with a Network-on-Chip. In Proceedings of the ACM symposium on Applied Computing. Association for Computing Machinery, New York, NY, USA, 1557\u20131564."},{"key":"e_1_3_2_1_122_1","volume-title":"Proceedings of the 36th Annual International Symposium on Computer Architecture. Association for Computing Machinery","author":"Mutlu Thomas","year":"2009","unstructured":"Moscibroda, Thomas and Mutlu, Onur. 2009. A Case for Bufferless Routing in On-chip Networks. In Proceedings of the 36th Annual International Symposium on Computer Architecture. Association for Computing Machinery, New York, NY, USA, 196\u2013207."},{"key":"e_1_3_2_1_123_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery","author":"Neugebauer Rolf","unstructured":"Rolf Neugebauer, Gianni Antichi, Jos\u00e9 Fernando Zazo, Yury Audzevich, Sergio L\u00f3pez-Buedo, and Andrew W. Moore. 2018. Understanding PCIe Performance for End Host Networking. In Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery, New York, NY, USA, 327\u2013341."},{"key":"e_1_3_2_1_124_1","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613163"},{"key":"e_1_3_2_1_125_1","volume-title":"Design, Automation and Test in Europe","author":"Nollet Vincent","unstructured":"Vincent Nollet, Th\u00e9odore Marescaux, Prabhat Avasare, Diederik Verkest, and J-Y Mignolet. 2005. Centralized Run-Time Resource Management in a Network-on-Chip Containing Reconfigurable Hardware Tiles. In Design, Automation and Test in Europe. IEEE Press, Piscataway, NJ, USA, 234\u2013239."},{"key":"e_1_3_2_1_126_1","unstructured":"NVIDIA. 2025. NVIDIA B200 GPU Accelerator. https:\/\/www.nvidia.com\/en-us\/data-center\/dgx-b200\/. (2025)."},{"key":"e_1_3_2_1_127_1","unstructured":"NVIDIA. 2025. NVIDIA Nsight Systems. https:\/\/developer.nvidia.com\/nsight-systems. (2025)."},{"key":"e_1_3_2_1_128_1","unstructured":"NVIDIA. 2025. NVIDIA NVLink: High-Speed GPU Interconnect. https:\/\/www.nvidia.com\/en-us\/products\/workstations\/nvlink-bridges\/. (2025)."},{"key":"e_1_3_2_1_129_1","unstructured":"NVIDIA. 2026. Cooperative Groups. https:\/\/docs.nvidia.com\/cuda\/cuda-programming-guide\/04-special-topics\/cooperative-groups.html. (2026)."},{"key":"e_1_3_2_1_130_1","unstructured":"NVIDIA. 2026. NVIDIA CUDA Profiling Tools Interface (CUPTI). https:\/\/developer.nvidia.com\/cupti. (2026)."},{"key":"e_1_3_2_1_131_1","unstructured":"NVIDIA. 2026. NVIDIA Nsight Compute. https:\/\/developer.nvidia.com\/nsight-compute. (2026)."},{"key":"e_1_3_2_1_132_1","unstructured":"OpenAI. 2025. GPT-5. https:\/\/openai.com\/gpt-5\/. (2025)."},{"key":"e_1_3_2_1_133_1","unstructured":"Muhammad Osama Ryan Swann Karthik Sangaiah Sonali Singh Ganesh Dasika and Rajneesh Bhardwaj. 2025. Deep dive into the MI300 compute and memory partition modes. https:\/\/rocm.blogs.amd.com\/software-tools-optimization\/compute-memory-modes\/README.html. (2025)."},{"key":"e_1_3_2_1_134_1","volume-title":"Proceedings of the 29th International Conference on Real-Time Networks and Systems. IEEE Press","author":"Otterness Nathan","unstructured":"Nathan Otterness and James H. Anderson. 2021. Exploring AMD GPU Scheduling Details by Experimenting With \"Worst Practices\". In Proceedings of the 29th International Conference on Real-Time Networks and Systems. IEEE Press, Piscataway, NJ, USA, 24\u201334."},{"key":"e_1_3_2_1_135_1","doi-asserted-by":"publisher","DOI":"10.1145\/3581784.3607098"},{"key":"e_1_3_2_1_136_1","doi-asserted-by":"publisher","DOI":"10.1145\/3725843.3756090"},{"key":"e_1_3_2_1_137_1","volume-title":"On-Chip Communication Architectures: System on Chip Interconnect. Morgan Kaufmann","author":"Pasricha Sudeep","unstructured":"Sudeep Pasricha and Nikil Dutt. 2010. On-Chip Communication Architectures: System on Chip Interconnect. Morgan Kaufmann, San Francisco, CA."},{"key":"e_1_3_2_1_138_1","doi-asserted-by":"publisher","DOI":"10.1145\/3730584"},{"key":"e_1_3_2_1_139_1","unstructured":"PCI-SIG. 2025. PCI Express Specification. https:\/\/pcisig.com\/specifications. (2025)."},{"key":"e_1_3_2_1_140_1","volume-title":"Floem: A Programming System for NIC-Accelerated Network Applications. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI'18)","author":"Phothilimthana Phitchaya Mangpo","year":"2018","unstructured":"Phitchaya Mangpo Phothilimthana, Ming Liu, Antoine Kaufmann, Simon Peter, Rastislav Bodik, and Thomas Anderson. 2018. Floem: A Programming System for NIC-Accelerated Network Applications. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI'18). 663\u2013679."},{"key":"e_1_3_2_1_141_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00036"},{"key":"e_1_3_2_1_142_1","doi-asserted-by":"publisher","DOI":"10.1145\/3422604.3425929"},{"key":"e_1_3_2_1_143_1","doi-asserted-by":"publisher","DOI":"10.1145\/3477132.3483583"},{"key":"e_1_3_2_1_144_1","volume-title":"Sinclair","author":"Ramadas Vishnu","year":"2023","unstructured":"Vishnu Ramadas, Daniel Kouchekinia, Ndubuisi Osuji, and Matthew D. Sinclair. 2023. Closing the Gap: Improving the Accuracy of gem5's GPU Models. In 5th gem5 Users' Workshop. 2."},{"key":"e_1_3_2_1_145_1","volume-title":"6th Young Architects' Workshop (YArch). 2.","author":"Ramadas Vishnu","unstructured":"Vishnu Ramadas, Daniel Kouchekinia, and Matthew D. Sinclair. 2024. Further Closing the GAP: Improving the Accuracy of gem5's GPU Models. In 6th Young Architects' Workshop (YArch). 2."},{"key":"e_1_3_2_1_146_1","unstructured":"reddit. 2025. Why Mi300x is technically superior but training performance was still lagging behind. Now faster then H100! https:\/\/www.reddit.com\/r\/AMD_Stock\/comments\/1hbraqd\/why_mi300x_is_technically_superior_but_training\/. (2025)."},{"key":"e_1_3_2_1_147_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2010.68"},{"key":"e_1_3_2_1_148_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.16"},{"key":"e_1_3_2_1_149_1","doi-asserted-by":"publisher","DOI":"10.1145\/3477132.3483555"},{"key":"e_1_3_2_1_150_1","volume-title":"Approximating Fair Queueing on Reconfigurable Switches. In 15th USENIX Symposium on Networked Systems Design and Implementation (NSDI'18)","author":"Sharma Naveen Kr.","year":"2018","unstructured":"Naveen Kr. Sharma, Ming Liu, Kishore Atreya, and Arvind Krishnamurthy. 2018. Approximating Fair Queueing on Reconfigurable Switches. In 15th USENIX Symposium on Networked Systems Design and Implementation (NSDI'18). 1\u201316."},{"key":"e_1_3_2_1_151_1","volume-title":"Programmable Calendar Queues for High-speed Packet Scheduling. In 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI'20)","author":"Sharma Naveen Kr.","year":"2020","unstructured":"Naveen Kr. Sharma, Chenxingyu Zhao, Ming Liu, Pravein G Kannan, Changhoon Kim, Arvind Krishnamurthy, and Anirudh Sivaraman. 2020. Programmable Calendar Queues for High-speed Packet Scheduling. In 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI'20). 685\u2013699."},{"key":"e_1_3_2_1_152_1","volume-title":"Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In International Conference on Learning Representations. OpenReview.net, 19","author":"Shazeer Noam","year":"2017","unstructured":"Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In International Conference on Learning Representations. OpenReview.net, 19."},{"key":"e_1_3_2_1_153_1","volume-title":"Proceedings of the 19th Symposium on Operating Systems Design and Implementation. USENIX Association, USA, 671\u2013692","author":"Shen Weihang","year":"2025","unstructured":"Weihang Shen, Mingcong Han, Jialong Liu, Rong Chen, and Haibo Chen. 2025. XSched: Preemptive Scheduling for Diverse XPUs. In Proceedings of the 19th Symposium on Operating Systems Design and Implementation. USENIX Association, USA, 671\u2013692."},{"key":"e_1_3_2_1_154_1","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1109\/MM.2025.3552324","article-title":"AMD Instinct\u2122 MI300X: A Generative AI Accelerator and Platform Architecture","volume":"45","author":"Smith Alan","year":"2025","unstructured":"Alan Smith and Vamsi Krishna Alla. 2025. AMD Instinct\u2122 MI300X: A Generative AI Accelerator and Platform Architecture. IEEE Micro 45, 3 (2025), 41\u201348.","journal-title":"IEEE Micro"},{"key":"e_1_3_2_1_155_1","volume-title":"Proceedings of the IEEE International Solid-State Circuits Conference","volume":"67","author":"Smith Alan","year":"2024","unstructured":"Alan Smith, Eric Chapman, Chintan Patel, Raja Swaminathan, John Wuu, Tyrone Huang, Wonjun Jung, Alexander Kaganov, Hugh McIntyre, and Ramon Mangaser. 2024. 11.1 AMD Instinct MI300 Series Modular Chiplet Package-HPC and AI accelerator for Exa-Class Systems. In Proceedings of the IEEE International Solid-State Circuits Conference, Vol. 67. IEEE Press, Piscataway, NJ, USA, 490\u2013492."},{"key":"e_1_3_2_1_156_1","volume-title":"Proceedings of the IEEE Hot Chips 34 Symposium. IEEE Press","author":"Smith Alan","year":"2022","unstructured":"Alan Smith and Norman James. 2022. AMD Instinct\u2122 MI200 Series Accelerator and Node Architectures. In Proceedings of the IEEE Hot Chips 34 Symposium. IEEE Press, Piscataway, NJ, USA, 1\u201323."},{"key":"e_1_3_2_1_157_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2024.3462351"},{"key":"e_1_3_2_1_158_1","volume-title":"Proceedings of the 51st International Symposium on Computer Architecture. IEEE Press","author":"Smith Alan","year":"2024","unstructured":"Alan Smith, Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Samuel Naffziger, Mike Mantor, Mark Fowler Nathan Kalyanasundharam, Vamsi Alla, Nicholas Malaya, Joseph L. Greathouse, Eric Chapman, and Raja Swaminathan. 2024. Realizing the AMD Exascale Heterogeneous Processor Vision: Industry Product. In Proceedings of the 51st International Symposium on Computer Architecture. IEEE Press, Piscataway, NJ, USA, 876\u2013889."},{"key":"e_1_3_2_1_159_1","volume-title":"Proceedings of the IEEE Symposium on VLSI Technology and Circuits. IEEE Press","author":"Smith Alan","year":"2024","unstructured":"Alan Smith, Gabriel H. Loh, John Wuu, Samuel Naffziger, Tyrone Huang, Hugh McIntyre, Ramon Mangaser, Wonjun Jung, and Raja Swaminathan. 2024. AMD Instinct MI300X Accelerator: Packaging and Architecture Co-Optimization. In Proceedings of the IEEE Symposium on VLSI Technology and Circuits. IEEE Press, Piscataway, NJ, USA, 1\u20132."},{"key":"e_1_3_2_1_160_1","unstructured":"Synopsys. 2025. What are Chiplets? https:\/\/www.synopsys.com\/glossary\/what-are-chiplets.html. (2025)."},{"key":"e_1_3_2_1_161_1","volume-title":"Proceedings of the ISC High Performance. Association for Computing Machinery","author":"Tandon Suyash","year":"2024","unstructured":"Suyash Tandon, Leopold Grinberg, Gheorghe-Teodor Bercea, Carlo Bertolli, Mark Olesen, Simone Bna, and Nicholas Malaya. 2024. Porting HPC Applications to AMD Instinct\u2122 MI300A using Unified Memory and OpenMP\u00ae. In Proceedings of the ISC High Performance. Association for Computing Machinery, New York, NY, USA, 1\u20139."},{"key":"e_1_3_2_1_162_1","doi-asserted-by":"publisher","DOI":"10.1145\/1150343.1150364"},{"key":"e_1_3_2_1_163_1","volume-title":"Proceedings of the SC25-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis. Association for Computing Machinery","author":"Tee Andrew","year":"2025","unstructured":"Andrew Tee, Nicholas Curtis, Noah Wolfe, and Daniel Wong. 2025. The MALL is Open: Exploring Shared Caches and Latency in AMD CDNA\u2122 3 GPUs. In Proceedings of the SC25-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis. Association for Computing Machinery, New York, NY, USA, 1110\u20131116."},{"key":"e_1_3_2_1_164_1","unstructured":"Tenstorrent. 2025. Tenstorrent Blackhole Accelerator. https:\/\/tenstorrent.com\/en\/hardware\/blackhole. (2025)."},{"key":"e_1_3_2_1_165_1","unstructured":"vLLM. 2025. Optimization and Tuning. https:\/\/docs.vllm.ai\/en\/v0.7.3\/performance\/optimization.html. (2025)."},{"key":"e_1_3_2_1_166_1","volume-title":"Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery","author":"Vuppalapati Midhul","year":"2024","unstructured":"Midhul Vuppalapati, Saksham Agarwal, Henry Schuh, Baris Kasikci, Arvind Krishnamurthy, and Rachit Agarwal. 2024. Understanding the Host Network. In Proceedings of the ACM SIGCOMM Conference. Association for Computing Machinery, New York, NY, USA, 581\u2013594."},{"key":"e_1_3_2_1_167_1","volume-title":"Proceedings of the IEEE International Symposium on Workload Characterization. IEEE Press","author":"Wahlgren Jacob","year":"2025","unstructured":"Jacob Wahlgren, Gabin Schieffer, Ruimin Shi, Edgar A Le\u00f3n, Roger Pearce, Maya Gokhale, and Ivy Peng. 2025. Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs. In Proceedings of the IEEE International Symposium on Workload Characterization. IEEE Press, Piscataway, NJ, USA, 368\u2013380."},{"key":"e_1_3_2_1_168_1","volume-title":"23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI'26). 2651\u20132669.","author":"Wang Chendong","unstructured":"Chendong Wang, Joontaek Oh, and Ming Liu. 2026. Co-Designing Traffic Control with NVMe-oF for Disaggregated Storage: A Comparative Study of Switched and Switchless SAN Architectures. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI'26). 2651\u20132669."},{"key":"e_1_3_2_1_169_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2010.5452013"},{"key":"e_1_3_2_1_170_1","doi-asserted-by":"publisher","DOI":"10.14778\/3773749.3773754"},{"key":"e_1_3_2_1_171_1","volume-title":"Sinclair","author":"Xia Yu","year":"2025","unstructured":"Yu Xia, Vishnu Ramadas, Matthew Poremba, and Matthew D. Sinclair. 2025. Narrowing the GAP: Enhancing gem5's GPU Memory Bandwidth Accuracy. In 6th gem5 Users' Workshop. 2."},{"key":"e_1_3_2_1_172_1","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613147"},{"key":"e_1_3_2_1_173_1","volume-title":"Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software. IEEE Press","author":"Yasin Ahmad","year":"2014","unstructured":"Ahmad Yasin. 2014. A Top-Down Method for Performance Analysis and Counters Architecture. In Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software. IEEE Press, Piscataway, NJ, USA, 35\u201344."},{"key":"e_1_3_2_1_174_1","doi-asserted-by":"publisher","DOI":"10.1145\/3314212.3314215"},{"key":"e_1_3_2_1_175_1","volume-title":"RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICs. In 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA'25)","author":"Zhang Jie","year":"2025","unstructured":"Jie Zhang, Hongjing Huang, Xuzheng Chen, Xiang Li, Jieru Zhao, Ming Liu, and Zeke Wang. 2025. RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICs. In 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA'25). 1379\u20131394."},{"key":"e_1_3_2_1_176_1","volume-title":"White-Boxing RDMA with Packet-Granular Software Control. In 22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI'25)","author":"Zhao Chenxingyu","year":"2025","unstructured":"Chenxingyu Zhao, Jaehong Min, Ming Liu, and Arvind Krishnamurthy. 2025. White-Boxing RDMA with Packet-Granular Software Control. In 22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI'25). 427\u2013449."},{"key":"e_1_3_2_1_177_1","volume-title":"Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems","volume":"2","author":"Zhao Chenxingyu","year":"2026","unstructured":"Chenxingyu Zhao, Hongtao Zhang, Jaehong Min, Shengkai Lin, Wei Zhang, Kaiyuan Zhang, Ming Liu, and Arvind Krishnamurthy. 2026. SG-IOV: Socket-Granular I\/O Virtualization for SmartNIC-Based Container Networks. In Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. 1727\u20131748."},{"key":"e_1_3_2_1_178_1","doi-asserted-by":"publisher","DOI":"10.52202\/068431-0515"}],"event":{"name":"SIGCOMM '26: ACM SIGCOMM 2026 Conference","location":"Colorado Convention Center Denver CO USA","acronym":"SIGCOMM '26","sponsor":["SIGCOMM ACM Special Interest Group on Data Communication"]},"container-title":["Proceedings of the ACM SIGCOMM 2026 Conference"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3789240.3829110","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,11]],"date-time":"2026-08-11T18:44:19Z","timestamp":1786473859000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3789240.3829110"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,8,11]]},"references-count":178,"alternative-id":["10.1145\/3789240.3829110","10.1145\/3789240"],"URL":"https:\/\/doi.org\/10.1145\/3789240.3829110","relation":{},"subject":[],"published":{"date-parts":[[2026,8,11]]},"assertion":[{"value":"2026-08-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}