{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,24]],"date-time":"2025-08-24T01:44:31Z","timestamp":1755999871056,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":58,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,11,7]],"date-time":"2022-11-07T00:00:00Z","timestamp":1667779200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100000781","name":"European Research Council","doi-asserted-by":"publisher","award":["805223"],"award-info":[{"award-number":["805223"]}],"id":[{"id":"10.13039\/501100000781","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,11,7]]},"DOI":"10.1145\/3528535.3565248","type":"proceedings-article","created":{"date-parts":[[2022,12,20]],"date-time":"2022-12-20T13:40:01Z","timestamp":1671543601000},"page":"241-254","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["CGX"],"prefix":"10.1145","author":[{"given":"Ilia","family":"Markov","sequence":"first","affiliation":[{"name":"Institute of Science and Technology, Austria"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hamidreza","family":"Ramezanikebrya","sequence":"additional","affiliation":[{"name":"University of British Columbia, Vancouver, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dan","family":"Alistarh","sequence":"additional","affiliation":[{"name":"Institute of Science and Technology, Austria"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,11,8]]},"reference":[{"volume-title":"Retrieved","year":"2022","key":"e_1_3_2_1_1_1","unstructured":"2021. NVIDIA AMPERE GA102 GPU ARCHITECTURE. Retrieved September 30, 2022 from https:\/\/images.nvidia.com\/aem-dam\/en-zz\/Solutions\/geforce\/ampere\/pdf\/NVIDIA-ampere-GA102-GPU-Architecture-Whitepaper-V1.pdf"},{"key":"e_1_3_2_1_2_1","volume-title":"Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. 2016. Tensorflow: A system for large-scale machine learning. In Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation (Savannah, GA, USA, November 2 - 4, 2016) (OSDI'16). USENIX Association, USA, 265--283."},{"key":"e_1_3_2_1_3_1","volume-title":"Proceedings of Machine Learning and Systems (Virtual event, USA, April 5 - 9","volume":"3","author":"Agarwal Saurabh","year":"2021","unstructured":"Saurabh Agarwal, Hongyi Wang, Kangwook Lee, Shivaram Venkataraman, and Dimitris Papailiopoulos. 2021. Adaptive Gradient Communication via Critical Learning Regime Identification. In Proceedings of Machine Learning and Systems (Virtual event, USA, April 5 - 9, 2021), Vol. 3. 55--80."},{"volume-title":"Proceedings of Machine Learning and Systems.","author":"Agarwal Saurabh","key":"e_1_3_2_1_4_1","unstructured":"Saurabh Agarwal, Hongyi Wang, Shivaram Venkataraman, and Dimitris Papailiopoulos. [n.d.]. In Proceedings of Machine Learning and Systems."},{"key":"e_1_3_2_1_5_1","volume-title":"QSGD: Communication-efficient SGD via gradient quantization and encoding. In Advances in Neural Information Processing Systems (Long Beach, CA, USA, December 4 - 7","author":"Alistarh Dan","year":"2017","unstructured":"Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic. 2017. QSGD: Communication-efficient SGD via gradient quantization and encoding. In Advances in Neural Information Processing Systems (Long Beach, CA, USA, December 4 - 7, 2017), Vol. 30. 1709--1720."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3477132.3483553"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11728"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid49817.2020.00-40"},{"key":"e_1_3_2_1_9_1","volume-title":"Proceedings of the 11th USENIX Conference on Operating Systems Design and Implementation","volume":"14","author":"Chilimbi Trishul M","year":"2014","unstructured":"Trishul M Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman. 2014. Project Adam: Building an Efficient and Scalable Deep Learning Training System. In Proceedings of the 11th USENIX Conference on Operating Systems Design and Implementation (Broomfield, CO, USA, October 6--8, 2014), Vol. 14. 571--582."},{"key":"e_1_3_2_1_10_1","volume-title":"Transformer-xl: Attentive language models beyond a fixed-length context.","author":"Dai Zihang","year":"2019","unstructured":"Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019. Transformer-xl: Attentive language models beyond a fixed-length context. (2019). arXiv:arXiv:1901.02860"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_1_12_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding.","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. (2018). arXiv:arXiv:1810.04805"},{"key":"e_1_3_2_1_13_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/3018874.3018875"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5793"},{"key":"e_1_3_2_1_16_1","volume-title":"Adaptive gradient quantization for data-parallel sgd. Advances in neural information processing systems 33","author":"Faghri Fartash","year":"2020","unstructured":"Fartash Faghri, Iman Tabrizian, Ilia Markov, Dan Alistarh, Daniel M Roy, and Ali Ramezani-Kebrya. 2020. Adaptive gradient quantization for data-parallel sgd. Advances in neural information processing systems 33 (2020), 3174--3185."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3452296.3472904"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.14778\/3503585.3503590"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2017.36"},{"volume-title":"Retrieved","year":"2021","key":"e_1_3_2_1_20_1","unstructured":"Genesis. 2021. Genesis GPU Cloud Offering. Retrieved September 30, 2022 from https:\/genesiscloud.com"},{"key":"e_1_3_2_1_21_1","unstructured":"Priya Goyal Piotr Doll\u00e1r Ross Girshick Pieter Noordhuis Lukasz Wesolowski Aapo Kyrola Andrew Tulloch Yangqing Jia and Kaiming He. 2017. Accurate large minibatch sgd: Training imagenet in 1 hour. (2017). arXiv:arXiv:1706.02677"},{"key":"e_1_3_2_1_22_1","volume-title":"Proceedings of the 21st International Conference on Extending Database Technology","author":"Grubic Demjan","year":"2018","unstructured":"Demjan Grubic, Leo K Tam, Dan Alistarh, and Ce Zhang. 2018. Synchronous multi-gpu deep learning with low-precision communication: An experimental study. In Proceedings of the 21st International Conference on Extending Database Technology (Vienna, Austria, March 26--29, 2018). OpenProceedings, 145--156."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9054164"},{"key":"e_1_3_2_1_24_1","volume-title":"Retrieved","author":"Harmon William","year":"2021","unstructured":"William Harmon. 2021. Dual NVIDIA GeForce RTX 3090 NVLink Performance Review. Retrieved September 30, 2022 from https:\/\/www.servethehome.com\/dual-nvidia-geforce-rtx-3090-nvlink-performance-review-asus-zotac\/"},{"key":"e_1_3_2_1_25_1","volume-title":"Retrieved","author":"Huggingface Inc","year":"2022","unstructured":"Inc Huggingface. 2022. Huggingface Transformers Repository. Retrieved April 30, 2022 from https:\/\/huggingface.co\/models"},{"key":"e_1_3_2_1_26_1","unstructured":"Anand Jayarajan Jinliang Wei Garth Gibson Alexandra Fedorova and Gennady Pekhimenko. 2019. Priority-based parameter propagation for distributed DNN training. (2019). arXiv:arXiv:1905.03960"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/3488766.3488792"},{"key":"e_1_3_2_1_28_1","volume-title":"International Conference on Machine Learning","author":"Karimireddy Sai Praneeth","year":"2019","unstructured":"Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi. 2019. Error feedback fixes signsgd and other gradient compression schemes. In International Conference on Machine Learning (Long Beach, CA, USA, Jun 10 - 15, 2019). PMLR, 3252--3261."},{"volume-title":"Retrieved","year":"2021","key":"e_1_3_2_1_29_1","unstructured":"LambdaLabs. 2021. LambdaLabs GPU Cloud Offering. Retrieved September 30, 2022 from https:\/\/lambdalabs.com\/cloud"},{"key":"e_1_3_2_1_30_1","volume-title":"Retrieved","author":"GPU.","year":"2021","unstructured":"LeaderGPU. 2021. LeaderGPU Cloud Offering. Retrieved September 30, 2022 from https:\/\/www.leadergpu.com\/"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2928289"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2018.8573483"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2640087.2644155"},{"key":"e_1_3_2_1_34_1","unstructured":"Shijian Li Oren Mangoubi Lijie Xu and Tian Guo. 2021. Sync-Switch: Hybrid Parameter Synchronization for Distributed Deep Learning. (2021). arXiv:arXiv:2104.08364"},{"key":"e_1_3_2_1_35_1","unstructured":"Hyeontaek Lim David G Andersen and Michael Kaminsky. 2018. 3lc: Lightweight and effective traffic compression for distributed machine learning. (2018). arXiv:arXiv:1802.07389"},{"key":"e_1_3_2_1_36_1","unstructured":"Yujun Lin Song Han Huizi Mao Yu Wang and William J Dally. 2017. Deep gradient compression: Reducing the communication bandwidth for distributed training. (2017). arXiv:arXiv:1712.01887"},{"key":"e_1_3_2_1_37_1","volume-title":"Proceedings of Machine Learning and Systems (Virtual event, USA, April 5 - 9","volume":"3","author":"Abdelmoniem Ahmed M","year":"2021","unstructured":"Ahmed M Abdelmoniem, Ahmed Elzanaty, Mohamed-Slim Alouini, and Marco Canini. 2021. An efficient statistical-based gradient compression technique for distributed training systems. In Proceedings of Machine Learning and Systems (Virtual event, USA, April 5 - 9, 2021), Vol. 3. 297--322."},{"key":"e_1_3_2_1_38_1","volume-title":"Proc. of the fifth Berkeley Symposium on Mathematical Statistics and Probability","volume":"1","author":"MacQueen J. B.","year":"1967","unstructured":"J. B. MacQueen. 1967. Some Methods for Classification and Analysis of MultiVariate Observations. In Proc. of the fifth Berkeley Symposium on Mathematical Statistics and Probability (Berkeley, CA, USA, June 21-July 18 1965), Vol. 1. University of California Press, 281--297."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2020.2974843"},{"volume-title":"Retrieved","year":"2020","key":"e_1_3_2_1_40_1","unstructured":"Nvidia. 2020. NVIDIA Deep Learning Examples for Tensor Cores. Retrieved April 30, 2022 from https:\/\/github.com\/NVIDIA\/DeepLearningExamples"},{"key":"e_1_3_2_1_41_1","volume-title":"International Conference on Machine Learning","author":"Parmar Niki","year":"2018","unstructured":"Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018. Image transformer. In International Conference on Machine Learning (Stockholm, Sweden, July 10--15, 2018). PMLR, 4055--4064."},{"key":"e_1_3_2_1_42_1","volume-title":"Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32 (2019), 8026--8037."},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359642"},{"key":"e_1_3_2_1_44_1","first-page":"1","article-title":"NUQSGD: Provably Communication-efficient Data-parallel SGD via Nonuniform Quantization","volume":"22","author":"Ramezani-Kebrya Ali","year":"2021","unstructured":"Ali Ramezani-Kebrya, Fartash Faghri, Ilya Markov, Vitalii Aksenov, Dan Alistarh, and Daniel M Roy. 2021. NUQSGD: Provably Communication-efficient Data-parallel SGD via Nonuniform Quantization. Journal of Machine Learning Research 22, 114 (2021), 1--43.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356222"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2014-274"},{"key":"e_1_3_2_1_47_1","unstructured":"Alexander Sergeev and Mike Del Balso. 2018. Horovod: fast and easy distributed deep learning in TensorFlow. (2018). arXiv:arXiv:1802.05799"},{"key":"e_1_3_2_1_48_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. (2014). arXiv:arXiv:1409.1556"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2015-354"},{"key":"e_1_3_2_1_50_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008."},{"key":"e_1_3_2_1_51_1","volume-title":"Sai Praneeth Karinireddy, and Martin Jaggi","author":"Vogels Thijs","year":"2019","unstructured":"Thijs Vogels, Sai Praneeth Karinireddy, and Martin Jaggi. 2019. PowerSGD: Practical low-rank gradient compression for distributed optimization. Advances In Neural Information Processing Systems 32 (Nips 2019) 32 (2019)."},{"key":"e_1_3_2_1_52_1","volume-title":"ATOMO: Communication-efficient learning via atomic sparsification.","author":"Wang Hongyi","year":"2018","unstructured":"Hongyi Wang, Scott Sievert, Zachary Charles, Shengchao Liu, Stephen Wright, and Dimitris Papailiopoulos. 2018. ATOMO: Communication-efficient learning via atomic sparsification. (2018). arXiv:arXiv:1806.04090"},{"key":"e_1_3_2_1_53_1","volume-title":"Terngrad: Ternary gradients to reduce communication in distributed deep learning.","author":"Wen Wei","year":"2017","unstructured":"Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2017. Terngrad: Ternary gradients to reduce communication in distributed deep learning. (2017). arXiv:arXiv preprint arXiv:1705.07878"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","unstructured":"Ross Wightman. 2019. PyTorch Image Models. https:\/\/github.com\/rwightman\/pytorch-image-models. 10.5281\/zenodo.4414861","DOI":"10.5281\/zenodo.4414861"},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS51616.2021.00060"},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3423211.3425675"},{"key":"e_1_3_2_1_57_1","unstructured":"Yang You Jing Li Sashank Reddi Jonathan Hseu Sanjiv Kumar Srinadh Bhojanapalli Xiaodan Song James Demmel Kurt Keutzer and Cho-Jui Hsieh. 2019. Large batch optimization for deep learning: Training bert in 76 minutes. (2019). arXiv:arXiv:1904.00962"},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2020.3040601"}],"event":{"name":"Middleware '22: 23rd International Middleware Conference","sponsor":["ACM Association for Computing Machinery","IFIP"],"location":"Quebec QC Canada","acronym":"Middleware '22"},"container-title":["Proceedings of the 23rd ACM\/IFIP International Middleware Conference"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3528535.3565248","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3528535.3565248","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:02:43Z","timestamp":1750186963000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3528535.3565248"}},"subtitle":["adaptive system support for communication-efficient deep learning"],"short-title":[],"issued":{"date-parts":[[2022,11,7]]},"references-count":58,"alternative-id":["10.1145\/3528535.3565248","10.1145\/3528535"],"URL":"https:\/\/doi.org\/10.1145\/3528535.3565248","relation":{},"subject":[],"published":{"date-parts":[[2022,11,7]]},"assertion":[{"value":"2022-11-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}