{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T19:41:16Z","timestamp":1787341276201,"version":"build-2736575974"},"reference-count":49,"publisher":"Society for Industrial & Applied Mathematics (SIAM)","issue":"2","funder":[{"DOI":"10.13039\/100000001","name":"NSF","doi-asserted-by":"publisher","award":["ECCS 2330196"],"award-info":[{"award-number":["ECCS 2330196"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Ministry of Science, Technological Development and Innovation, Republic of Serbia","award":["7359"],"award-info":[{"award-number":["7359"]}]},{"name":"Provincial Secretariat for Higher Education and Scientific Research","award":["142-451- 2593\/2021-01\/2"],"award-info":[{"award-number":["142-451- 2593\/2021-01\/2"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["SIAM J. Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Abstract.<\/jats:p>\n                  <jats:p>Motivated by understanding and analysis of large-scale machine learning under heavy-tailed gradient noise, we study decentralized optimization with gradient clipping, i.e., in which certain clipping operators are applied to the gradients or gradient estimates computed from local nodes prior to further processing. While vanilla gradient clipping has proven effective in mitigating the impact of heavy-tailed gradient noise in nondistributed setups, it incurs bias that causes convergence issues in heterogeneous distributed settings. To address the inherent bias introduced by gradient clipping, we develop a smoothed clipping operator, and propose a decentralized gradient method equipped with an error feedback mechanism, i.e., the clipping operator is applied on the difference between some local gradient estimator and local stochastic gradient. We consider strongly convex and smooth local functions under symmetric heavy-tailed gradient noise that may not have finite moments of order greater than one. We show that the proposed decentralized gradient clipping method achieves a mean-square error (MSE) convergence rate of [Formula: see text], [Formula: see text], where the exponent [Formula: see text] is independent of the existence of higher order gradient noise moments [Formula: see text] and lower bounded by some constant dependent on condition number. To the best of our knowledge, this is the first MSE convergence result for decentralized gradient clipping under heavy-tailed noise without assuming bounded gradient. Numerical experiments validate our theoretical findings.<\/jats:p>","DOI":"10.1137\/24m170747x","type":"journal-article","created":{"date-parts":[[2026,4,23]],"date-time":"2026-04-23T09:01:02Z","timestamp":1776934862000},"page":"703-728","source":"Crossref","is-referenced-by-count":2,"title":["Smoothed Gradient Clipping and Error Feedback for Decentralized Optimization under Symmetric Heavy-Tailed Noise"],"prefix":"10.1137","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6054-673X","authenticated-orcid":true,"given":"Shuhua","family":"Yu","sequence":"first","affiliation":[{"name":"Department of Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PA 15213 USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Du\u0161an","family":"Jakovetic","sequence":"additional","affiliation":[{"name":"Faculty of Sciences, Department of Mathematics and Informatics, University of Novi Sad, Novi Sad, 21000 Serbia."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Soummya","family":"Kar","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, Carnegie Mellon University, Pittsburgh, PA 15213 USA."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"351","published-online":{"date-parts":[[2026,4,23]]},"reference":[{"key":"ref1","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2016.2535239"},{"key":"ref2","unstructured":"A. Armacki, S. Yu, P. Sharma, G. Joshi, D. Bajovic, D. Jakovetic, and S. Kar, Nonlinear Stochastic Gradient Descent and Heavy-Tailed Noise: A Unified Framework and High-Probability Guarantees, preprint, arXiv:2410.13954, 2024."},{"key":"ref3","first-page":"29364","volume":"34","author":"Barsbey M.","year":"2021","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"ref4","unstructured":"B. Battash, L. Wolf, and O. Lindenbaum, Revisiting the noise model of stochastic gradient descent, in International Conference on Artificial Intelligence and Statistics, PMLR, 2024, pp. 4780\u20134788."},{"key":"ref5","doi-asserted-by":"publisher","DOI":"10.2307\/121080"},{"key":"ref6","unstructured":"M. Crawshaw, Y. Bao, and M. Liu, Federated learning with client subsampling, data heterogeneity, and unbounded smoothness: A new algorithm and lower bounds, in Advances in Neural Information Processing Systems 36 (NeurIPS 2023) 36, 2023. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2023\/hash\/14ecbfb2216bab76195b60bfac7efb1f-Abstract-Conference.html."},{"key":"ref7","first-page":"4883","volume":"34","author":"Cutkosky A.","year":"2021","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"ref8","first-page":"1","volume":"22","author":"Davis D.","year":"2021","journal-title":"J. Mach. Learn. Res."},{"key":"ref9","doi-asserted-by":"publisher","DOI":"10.1109\/TSIPN.2016.2524588"},{"key":"ref10","doi-asserted-by":"publisher","DOI":"10.29012\/jpc.v7i3.405"},{"key":"ref11","doi-asserted-by":"publisher","DOI":"10.3166\/EJC.18.539-557"},{"key":"ref12","first-page":"15042","volume":"33","author":"Gorbunov E.","year":"2020","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"ref13","unstructured":"E. Gorbunov, M. Danilova, I. Shibaev, P. Dvurechensky, and A. Gasnikov, Near-optimal High Probability Complexity Bounds for Non-Smooth Stochastic Optimization with Heavy-Tailed Noise, preprint, arXiv:2106.05958, 2021."},{"key":"ref14","unstructured":"E. Gorbunov, A. Sadiev, M. Danilova, S. Horv\u00e1th, G. Gidel, P. Dvurechensky, A. Gasnikov, and P. Richt\u00e1rik, High-probability convergence for composite and distributed stochastic minimization and variational inequalities with heavy-tailed noise, in Forty-first International Conference on Machine Learning, 2024."},{"key":"ref15","unstructured":"M. Gurbuzbalaban, U. Simsekli, and L. Zhu, The heavy-tail phenomenon in SGD, in International Conference on Machine Learning, PMLR, 2021, pp. 3964\u20133975."},{"key":"ref16","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9781139020411"},{"key":"ref17","doi-asserted-by":"publisher","DOI":"10.1137\/21M145896X"},{"key":"ref18","doi-asserted-by":"publisher","DOI":"10.1137\/22M1477015"},{"key":"ref19","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.2014.2298712"},{"key":"ref20","doi-asserted-by":"publisher","DOI":"10.1561\/2200000083"},{"key":"ref21","unstructured":"S. Khirirat, E. Gorbunov, S. Horv\u00e1th, R. Islamov, F. Karray, and P. Richt\u00e1rik, Clip21: Error Feedback for Gradient Clipping, preprint, arXiv:2305.18929, 2023."},{"key":"ref22","unstructured":"A. Koloskova, H. Hendrikx, and S. U. Stich, Revisiting gradient clipping: Stochastic bias and tight convergence guarantees, in ICML 2023-40th International Conference on Machine Learning, 2023."},{"key":"ref23","unstructured":"A. Koloskova, N. Loizou, S. Boreiri, M. Jaggi, and S. Stich, A unified theory of decentralized SGD with changing topology and local updates, in International Conference on Machine Learning, PMLR, 2020, pp. 5381\u20135393."},{"key":"ref24","unstructured":"B. Li and Y. Chi, Convergence and Privacy of Decentralized Nonconvex Optimization with Gradient Clipping and Communication Compression, preprint, arXiv:2305.09896, 2023."},{"key":"ref25","doi-asserted-by":"publisher","DOI":"10.1017\/9781009053730"},{"key":"ref26","doi-asserted-by":"publisher","DOI":"10.1134\/S0005117919090042"},{"key":"ref27","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.2008.2009515"},{"key":"ref28","doi-asserted-by":"crossref","first-page":"24191","DOI":"10.52202\/075280-1052","volume":"36","author":"Nguyen T. D.","year":"2023","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"ref29","doi-asserted-by":"publisher","DOI":"10.1016\/0024-3795(79)90157-5"},{"key":"ref30","unstructured":"S. Peluchetti, S. Favaro, and S. Fortini, Stable behaviour of infinitely wide deep neural networks, in International Conference on Artificial Intelligence and Statistics, PMLR, 2020, pp. 1137\u20131146."},{"key":"ref31","doi-asserted-by":"crossref","unstructured":"I. Pinelis, Exact Lower and Upper Bounds on the Incomplete Gamma Function, preprint, arXiv:2005.06384, 2020.","DOI":"10.7153\/mia-2020-23-95"},{"key":"ref32","unstructured":"B. T. Polyak, Introduction to optimization, 1987."},{"key":"ref33","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-020-01487-0"},{"key":"ref34","unstructured":"N. Puchkin, E. Gorbunov, N. Kutuzov, and A. Gasnikov, Breaking the Heavy-Tailed Noise Barrier in Stochastic Optimization Problems, preprint, arXiv:2311.04161, 2023."},{"key":"ref35","unstructured":"A. Sadiev, M. Danilova, E. Gorbunov, S. Horv\u00e1th, G. Gidel, P. Dvurechensky, A. Gasnikov, and P. Richt\u00e1rik, High-probability bounds for stochastic optimization and variational inequalities: The case of unbounded variance, in International Conference on Machine Learning, PMLR, 2023, pp. 29563\u201329648."},{"key":"ref36","doi-asserted-by":"publisher","DOI":"10.1137\/14096668X"},{"key":"ref37","unstructured":"U. \u015eim\u015fekli, M. G\u00fcrb\u00fczbalaban, T. H. Nguyen, G. Richard, and L. Sagun, On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks, preprint, arXiv:1912.00018, 2019."},{"key":"ref38","unstructured":"U. Simsekli, L. Sagun, and M. Gurbuzbalaban, A tail-index analysis of stochastic gradient noise in deep neural networks, in International Conference on Machine Learning, PMLR, 2019, pp. 5827\u20135837."},{"key":"ref39","doi-asserted-by":"crossref","unstructured":"C. Sun and B. Chen, Distributed stochastic strongly convex optimization under heavy-tailed noises, in 2024 IEEE International Conference on Cybernetics and Intelligent Systems (CIS) and IEEE International Conference on Robotics, Automation and Mechatronics (RAM), IEEE, 2024, pp. 150\u2013155.","DOI":"10.1109\/CIS-RAM61939.2024.10673219"},{"key":"ref40","doi-asserted-by":"publisher","DOI":"10.1007\/s10957-010-9737-7"},{"key":"ref41","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.1986.1104412"},{"key":"ref42","doi-asserted-by":"publisher","DOI":"10.1137\/22M1543197"},{"key":"ref43","doi-asserted-by":"crossref","unstructured":"R. Xin, S. Pu, A. Nedi\u0107, and U. A. Khan, A general framework for decentralized optimization with first-order methods, Proc. IEEE, 108 (2020), pp. 1869\u20131889.","DOI":"10.1109\/JPROC.2020.3024266"},{"key":"ref44","doi-asserted-by":"crossref","first-page":"17017","DOI":"10.52202\/068431-1238","volume":"35","author":"Yang H.","year":"2022","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"ref45","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2023.3277211"},{"key":"ref46","doi-asserted-by":"crossref","first-page":"8000","DOI":"10.52202\/068431-0581","volume":"35","author":"Zhang J.","year":"2022","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"ref47","unstructured":"J. Zhang, T. He, S. Sra, and A. Jadbabaie, Why gradient clipping accelerates training: A theoretical justification for adaptivity, in International Conference on Learning Representations, 2020, https:\/\/openreview.net\/forum?id=BJgnXpVYwS."},{"key":"ref48","first-page":"15383","volume":"33","author":"Zhang J.","year":"2020","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"ref49","unstructured":"X. Zhang, X. Chen, M. Hong, Z. S. Wu, and J. Yi, Understanding clipping for federated learning: Convergence and client-level differential privacy, in International Conference on Machine Learning, ICML 2022, 2022."}],"container-title":["SIAM Journal on Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/epubs.siam.org\/doi\/pdf\/10.1137\/24M170747X","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T19:12:10Z","timestamp":1787339530000},"score":1,"resource":{"primary":{"URL":"https:\/\/epubs.siam.org\/doi\/10.1137\/24M170747X"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,23]]},"references-count":49,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1137\/24M170747X"],"URL":"https:\/\/doi.org\/10.1137\/24m170747x","relation":{},"ISSN":["1052-6234","1095-7189"],"issn-type":[{"value":"1052-6234","type":"print"},{"value":"1095-7189","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,23]]}}}