{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T00:42:50Z","timestamp":1760056970992,"version":"build-2065373602"},"publisher-location":"New York, NY, USA","reference-count":21,"publisher":"ACM","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,9,8]]},"DOI":"10.1145\/3748273.3749206","type":"proceedings-article","created":{"date-parts":[[2025,9,2]],"date-time":"2025-09-02T16:19:34Z","timestamp":1756829974000},"page":"64-66","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Sparse Collectives: Exploiting Data Sparsity to Improve Communication Efficiency"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3181-2514","authenticated-orcid":false,"given":"Dhananjaya","family":"Wijerathne","sequence":"first","affiliation":[{"name":"AMD Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3472-0803","authenticated-orcid":false,"given":"Haris","family":"Javaid","sequence":"additional","affiliation":[{"name":"AMD Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0746-2921","authenticated-orcid":false,"given":"Guanwen","family":"Zhong","sequence":"additional","affiliation":[{"name":"AMD Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-5260-0980","authenticated-orcid":false,"given":"Dan","family":"Wu","sequence":"additional","affiliation":[{"name":"AMD Singapore and AMD"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-6624-8476","authenticated-orcid":false,"given":"Xing Yuan","family":"Kom","sequence":"additional","affiliation":[{"name":"AMD Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1964-7165","authenticated-orcid":false,"given":"Mario","family":"Baldi","sequence":"additional","affiliation":[{"name":"AMD, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,8]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Abhinav Agarwalla Abhay Gupta Alexandre Marques Shubhra Pandit Michael Goin Eldar Kurtic Kevin Leong Tuan Nguyen Mahmoud Salem Dan Alistarh et al. 2024. Enabling high-sparsity foundational llama models with efficient pretraining and deployment. arXiv preprint arXiv:2405.03594 (2024)."},{"key":"e_1_3_2_1_2_1","unstructured":"AMD. 2024. ROCm Communication Collectives Library (RCCL). https:\/\/github.com\/ROCm\/rccl."},{"key":"e_1_3_2_1_3_1","unstructured":"AMD ROCm Team. 2024. RCCL-Tests: Benchmarking Suite for ROCm Collective Communication Library. https:\/\/github.com\/ROCm\/rccl-tests."},{"key":"e_1_3_2_1_4_1","volume-title":"Demystifying the Communication Characteristics for Distributed Transformer Models. In 2024 IEEE Symposium on High-Performance Interconnects (HOTI). IEEE, 57--65","author":"Anthony Quentin","year":"2024","unstructured":"Quentin Anthony, Benjamin Michalowicz, Jacob Hatef, Lang Xu, Mustafa Abduljabbai, Aamir Shafi, Hari Subramoni, and Dhabaleswar K Panda. 2024. Demystifying the Communication Characteristics for Distributed Transformer Models. In 2024 IEEE Symposium on High-Performance Interconnects (HOTI). IEEE, 57--65."},{"key":"e_1_3_2_1_5_1","unstructured":"Li-Wen Chang Wenlei Bao Qi Hou Chengquan Jiang Ningxin Zheng Yinmin Zhong Xuanrun Zhang Zuquan Song Chengji Yao Ziheng Jiang et al. 2024. FLUX: fast software-based communication overlap on gpus through kernel fusion. arXiv preprint arXiv:2406.06858 (2024)."},{"key":"e_1_3_2_1_6_1","volume-title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv preprint arXiv:2010.11929","author":"Dosovitskiy Alexey","year":"2020","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv preprint arXiv:2010.11929 (2020). https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3452296.3472904"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3650200.3656636"},{"key":"e_1_3_2_1_9_1","unstructured":"Hugging Face. [n. d.]. SparseLLM: Sparse Large Language Models. https:\/\/huggingface.co\/SparseLLM\/prosparse-llama-2-7b."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"crossref","unstructured":"S. Jin P. Grosset C. M. Biwer J. Pulido J. Tian D. Tao and J. P. Ahrens. 2020. Understanding GPU-Based Lossy Compression for Extreme-Scale Cosmological Simulations. arXiv preprint arXiv:2004.00224 (2020).","DOI":"10.1109\/IPDPS47924.2020.00021"},{"key":"e_1_3_2_1_11_1","unstructured":"Shen Li Yanli Zhao Rohan Varma Omkar Salpekar Pieter Noordhuis Teng Li Adam Paszke Jeff Smith Brian Vaughan Pritam Damania et al. 2020. Pytorch distributed: Experiences on accelerating data parallel training. arXiv preprint arXiv:2006.15704 (2020)."},{"key":"e_1_3_2_1_12_1","volume-title":"Sashank J Reddi, Ke Ye, Felix Chern, Felix Yu, Ruiqi Guo, et al.","author":"Li Zonglin","year":"2022","unstructured":"Zonglin Li, Chong You, Srinadh Bhojanapalli, Daliang Li, Ankit Singh Rawat, Sashank J Reddi, Ke Ye, Felix Chern, Felix Yu, Ruiqi Guo, et al. 2022. The lazy neuron phenomenon: On emergence of activation sparsity in transformers. arXiv preprint arXiv:2210.06313 (2022)."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_2_1_14_1","volume-title":"Oncel Tuzel, Golnoosh Samei, Mohammad Rastegari, and Mehrdad Farajtabar.","author":"Mirzadeh Iman","year":"2023","unstructured":"Iman Mirzadeh, Keivan Alizadeh, Sachin Mehta, Carlo C Del Mundo, Oncel Tuzel, Golnoosh Samei, Mohammad Rastegari, and Mehrdad Farajtabar. 2023. Relu strikes back: Exploiting activation sparsity in large language models. arXiv preprint arXiv:2310.04564 (2023)."},{"key":"e_1_3_2_1_15_1","unstructured":"NVIDIA. 2024. NVIDIA Collective Communication Library (NCCL). https:\/\/developer.nvidia.com\/nccl."},{"volume-title":"Fourth Workshop on General Purpose Processing on Graphics Processing Units.","author":"O'Neil M. A.","key":"e_1_3_2_1_16_1","unstructured":"M. A. O'Neil and M. Burtscher. 2011. Floating-point data compression at 75 Gb\/s on a GPU. In Fourth Workshop on General Purpose Processing on Graphics Processing Units."},{"key":"e_1_3_2_1_17_1","volume-title":"Very Deep Convolutional Networks for Large-Scale Image Recognition. CoRR abs\/1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. CoRR abs\/1409.1556 (2014). arXiv:1409.1556 http:\/\/arxiv.org\/abs\/1409.1556"},{"volume-title":"IEEE Cluster Conference.","author":"Yang A.","key":"e_1_3_2_1_18_1","unstructured":"A. Yang, H. Mukka, F. Hesaaraki, and M. Burtscher. 2015. MPC: A massively parallel compression algorithm for scientific data. In IEEE Cluster Conference."},{"volume-title":"Proceedings of the 29th International Symposium on High-Performance Parallel and Distributed Computing. 89--100","author":"Zhao K.","key":"e_1_3_2_1_19_1","unstructured":"K. Zhao, S. Di, X. Liang, S. Li, D. Tao, Z. Chen, and F. Cappello. 2020. Significantly improving lossy compression for HPC datasets with second-order prediction and parameter optimization. In Proceedings of the 29th International Symposium on High-Performance Parallel and Distributed Computing. 89--100."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS54959.2023.00023"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS49936.2021.00053"}],"event":{"name":"SIGCOMM '25: ACM SIGCOMM 2025 Conference","sponsor":["SIGCOMM ACM Special Interest Group on Data Communication"],"location":"Coimbra Portugal","acronym":"SIGCOMM '25"},"container-title":["Proceedings of the 2nd Workshop on Networks for AI Computing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3748273.3749206","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T14:10:17Z","timestamp":1760019017000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3748273.3749206"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,8]]},"references-count":21,"alternative-id":["10.1145\/3748273.3749206","10.1145\/3748273"],"URL":"https:\/\/doi.org\/10.1145\/3748273.3749206","relation":{},"subject":[],"published":{"date-parts":[[2025,9,8]]},"assertion":[{"value":"2025-09-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}