{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T19:25:42Z","timestamp":1780773942719,"version":"3.54.1"},"reference-count":44,"publisher":"MIT Press","issue":"9","content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,8,19]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Adam-type algorithms have become a preferred choice for optimization in the deep learning setting; however, despite their success, their convergence is still not well understood. To this end, we introduce a unified framework for Adam-type algorithms, termed UAdam. It is equipped with a general form of the second-order moment, which makes it possible to include Adam and its existing and future variants as special cases, such as NAdam, AMSGrad, AdaBound, AdaFom, and Adan. The approach is supported by a rigorous convergence analysis of UAdam in the general nonconvex stochastic setting, showing that UAdam converges to the neighborhood of stationary points with a rate of O(1\/T). Furthermore, the size of the neighborhood decreases as the parameter \u03b21 increases. Importantly, our analysis only requires the first-order momentum factor to be close enough to 1, without any restrictions on the second-order momentum factor. Theoretical results also reveal the convergence conditions of vanilla Adam, together with the selection of appropriate hyperparameters. This provides a theoretical guarantee for the analysis, applications, and further developments of the whole general class of Adam-type algorithms. Finally, several numerical experiments are provided to support our theoretical findings.<\/jats:p>","DOI":"10.1162\/neco_a_01692","type":"journal-article","created":{"date-parts":[[2024,8,6]],"date-time":"2024-08-06T20:28:40Z","timestamp":1722976120000},"page":"1912-1938","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":4,"title":["UAdam: Unified Adam-Type Algorithmic Framework for Nonconvex Optimization"],"prefix":"10.1162","volume":"36","author":[{"given":"Yiming","family":"Jiang","sequence":"first","affiliation":[{"name":"Key Laboratory for Applied Statistics of MOE, School of Mathematics and Statistics, Northeast Normal University, Changchun 130024, China jiangym048@nenu.edu.cn"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinlan","family":"Liu","sequence":"additional","affiliation":[{"name":"Department of Mathematics, Changchun Normal University, Changchun 130032, China liujinlan@ccsfu.edu.cn"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dongpo","family":"Xu","sequence":"additional","affiliation":[{"name":"Key Laboratory for Applied Statistics of MOE, School of Mathematics and Statistics, Northeast Normal University, Changchun 130024, China xudp100@nenu.edu.cn"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Danilo P.","family":"Mandic","sequence":"additional","affiliation":[{"name":"Department of Electrical and Electronic Engineering, Imperial College London, SW7 2AZ London, U.K. d.mandic@imperial.ac.uk"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,8,19]]},"reference":[{"issue":"2","key":"2024082018250622900_bib1","doi-asserted-by":"publisher","first-page":"223","DOI":"10.1137\/16M1080173","article-title":"Optimization methods for large-scale machine learning","volume":"60","author":"Bottou","year":"2018","journal-title":"SIAM Review"},{"issue":"229","key":"2024082018250622900_bib2","first-page":"1","article-title":"Towards practical Adam: Nonconvexity, convergence theory, and mini-batch acceleration","volume":"23","author":"Chen","year":"2022","journal-title":"Journal of Machine Learning Research"},{"key":"2024082018250622900_bib3","first-page":"1","article-title":"On the convergence of a class of Adam-type algorithms for non-convex optimization","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Chen","year":"2019"},{"key":"2024082018250622900_bib4","first-page":"1","article-title":"An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dosovitskiy","year":"2021"},{"key":"2024082018250622900_bib5","first-page":"1","article-title":"Incorporating Nesterov momentum into Adam","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dozat","year":"2016"},{"key":"2024082018250622900_bib6","first-page":"2121","article-title":"Adaptive subgradient methods for on line learning and stochastic optimization","volume":"12","author":"Duchi","year":"2011","journal-title":"Journal of Machine Learning Research"},{"key":"2024082018250622900_bib7","doi-asserted-by":"publisher","first-page":"70","DOI":"10.1016\/j.ins.2021.11.044","article-title":"Convergence analysis for sigma-pi-sigma neural network based on some relaxed conditions","volume":"585","author":"Fan","year":"2022","journal-title":"Information Sciences"},{"issue":"4","key":"2024082018250622900_bib8","doi-asserted-by":"publisher","first-page":"2341","DOI":"10.1137\/120880811","article-title":"Stochastic first- and zeroth-order methods for nonconvex stochastic programming","volume":"23","author":"Ghadimi","year":"2013","journal-title":"SIAM Journal on Optimization"},{"issue":"12","key":"2024082018250622900_bib9","doi-asserted-by":"publisher","first-page":"2699","DOI":"10.1162\/0899766042321779","article-title":"A complex-valued RTRL algorithm for recurrent neural networks","volume":"16","author":"Goh","year":"2004","journal-title":"Neural Computation"},{"issue":"4","key":"2024082018250622900_bib10","doi-asserted-by":"publisher","first-page":"1039","DOI":"10.1162\/neco.2007.19.4.1039","article-title":"An augmented extended Kalman filter algorithm for complex-valued recurrent neural networks","volume":"19","author":"Goh","year":"2007","journal-title":"Neural Computation"},{"key":"2024082018250622900_bib11","first-page":"1","article-title":"A novel convergence analysis for algorithms of the Adam family","volume-title":"Proceedings of the 13th Annual Workshop on Optimization for Machine Learning","author":"Guo","year":"2021"},{"key":"2024082018250622900_bib12","first-page":"770","article-title":"Deep residual learning for image recognition","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"He","year":"2016"},{"issue":"2","key":"2024082018250622900_bib13","doi-asserted-by":"crossref","first-page":"634","DOI":"10.1137\/21M1394308","article-title":"Adaptivity of stochastic gradient methods for nonconvex optimization","volume":"4","author":"Horv\u00e1th","year":"2022","journal-title":"SIAM Journal on Mathematics of Data Science"},{"key":"2024082018250622900_bib14","first-page":"328","article-title":"Universal language model fine-tuning for text classification","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics","author":"Howard","year":"2018"},{"key":"2024082018250622900_bib15","doi-asserted-by":"publisher","first-page":"1143","DOI":"10.1109\/TSP.2023.3262181","article-title":"Distributed random reshuffling over networks","volume":"71","author":"Huang","year":"2023","journal-title":"IEEE Transactions on Signal Processing"},{"key":"2024082018250622900_bib16","first-page":"1","article-title":"Better theory for SGD in the nonconvex world","author":"Khaled","year":"2023","journal-title":"Transactions on Machine Learning Research"},{"key":"2024082018250622900_bib17","first-page":"1","article-title":"Adam: A method for stochastic optimization","volume-title":"Proceedings of the 3rd International Conference on Learning Representations","author":"Kingma","year":"2015"},{"key":"2024082018250622900_bib18","doi-asserted-by":"crossref","DOI":"10.1007\/978-981-15-2910-8","volume-title":"Accelerated optimization for machine learning","author":"Lin","year":"2020"},{"key":"2024082018250622900_bib19","doi-asserted-by":"publisher","first-page":"300","DOI":"10.1016\/j.neunet.2021.10.026","article-title":"Convergence analysis of AdaBound with relaxed bound functions for non-convex optimization","volume":"145","author":"Liu","year":"2022","journal-title":"Neural Networks"},{"key":"2024082018250622900_bib20","doi-asserted-by":"publisher","first-page":"27","DOI":"10.1016\/j.neucom.2023.01.032","article-title":"Last-iterate convergence analysis of stochastic momentum methods for neural networks","volume":"527","author":"Liu","year":"2023","journal-title":"Neurocomputing"},{"key":"2024082018250622900_bib21","first-page":"1","article-title":"On hyper-parameter selection for guaranteed convergence of RMSProp","author":"Liu","year":"2022","journal-title":"Cognitive Neurodynamics"},{"key":"2024082018250622900_bib22","first-page":"1","article-title":"Adaptive gradient methods with dynamic bound of learning rate","volume-title":"Proceedings of the 7th International Conference on Learning Representations","author":"Luo","year":"2019"},{"issue":"11","key":"2024082018250622900_bib23","doi-asserted-by":"publisher","first-page":"2693","DOI":"10.1162\/089976602760408026","article-title":"Data-reusing recurrent neural adaptive filters","volume":"14","author":"Mandic","year":"2002","journal-title":"Neural Computation"},{"key":"2024082018250622900_bib24","first-page":"372","article-title":"A method of solving a convex programming problem with convergence rate O(1\/k2)","volume":"27","author":"Nesterov","year":"1983","journal-title":"Soviet Mathematics Doklady"},{"key":"2024082018250622900_bib25","volume-title":"Introductory lectures on convex optimization: A basic course","author":"Nesterov","year":"2013"},{"issue":"4","key":"2024082018250622900_bib26","doi-asserted-by":"publisher","first-page":"1042","DOI":"10.1162\/neco.2008.12-06-418","article-title":"A homomorphic neural network for modeling and prediction","volume":"20","author":"Pedzisz","year":"2008","journal-title":"Neural Computation"},{"issue":"5","key":"2024082018250622900_bib27","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/0041-5553(64)90137-5","article-title":"Some methods of speeding up the convergence of iteration methods","volume":"4","author":"Polyak","year":"1964","journal-title":"USSR Computational Mathematics and Mathematical Physics"},{"key":"2024082018250622900_bib28","first-page":"1","article-title":"On the convergence of Adam and beyond","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Reddi","year":"2018"},{"key":"2024082018250622900_bib29","doi-asserted-by":"publisher","first-page":"400","DOI":"10.1214\/aoms\/1177729586","article-title":"A stochastic approximation method","volume":"22","author":"Robbins","year":"1951","journal-title":"Annals of Mathematical Statistics"},{"key":"2024082018250622900_bib30","article-title":"RMSprop converges with proper hyperparameter","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Shi","year":"2021"},{"issue":"2","key":"2024082018250622900_bib31","doi-asserted-by":"publisher","first-page":"249","DOI":"10.1007\/s40305-020-00309-6","article-title":"Optimization for deep learning: An overview","volume":"8","author":"Sun","year":"2020","journal-title":"Journal of the Operations Research Society of China"},{"key":"2024082018250622900_bib32","first-page":"1139","article-title":"On the importance of initialization and momentum in deep learning","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Sutskever","year":"2013"},{"issue":"2","key":"2024082018250622900_bib33","first-page":"26","article-title":"Lecture 6.5-RMSProp: Divide the gradient by a running average of its recent magnitude","volume":"4","author":"Tieleman","year":"2012","journal-title":"COURSERA: Neural Networks for Machine Learning"},{"key":"2024082018250622900_bib34","first-page":"1195","article-title":"Fast and faster convergence of SGD for over-parameterized models and an accelerated perceptron","volume-title":"Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics","author":"Vaswani","year":"2019"},{"key":"2024082018250622900_bib35","first-page":"1","article-title":"Speech emotion recognition using convolution neural networks and deep stride convolutional neural networks","volume-title":"Proceedings of the 6th International Conference on Wireless and Telematics","author":"Wani","year":"2020"},{"key":"2024082018250622900_bib36","author":"Xie","year":"2022","journal-title":"Adan: Adaptive Nesterov momentum algorithm for faster optimizing deep models"},{"issue":"10","key":"2024082018250622900_bib37","doi-asserted-by":"publisher","first-page":"2655","DOI":"10.1162\/NECO_a_00021","article-title":"Convergence analysis of three classes of split-complex gradient algorithms for complex-valued recurrent neural networks","volume":"22","author":"Xu","year":"2010","journal-title":"Neural Computation"},{"key":"2024082018250622900_bib38","doi-asserted-by":"publisher","first-page":"17","DOI":"10.1016\/j.neunet.2021.02.011","article-title":"Convergence of the RMSProp deep learning method with penalty for nonconvex optimization","volume":"139","author":"Xu","year":"2021","journal-title":"Neural Networks"},{"issue":"8","key":"2024082018250622900_bib39","doi-asserted-by":"publisher","first-page":"1404","DOI":"10.1162\/neco_a_01598","article-title":"Graph-regularized tensor regression: A domain-aware framework for interpretable modeling of multiway data on graphs","volume":"35","author":"Xu","year":"2023","journal-title":"Neural Computation"},{"key":"2024082018250622900_bib40","doi-asserted-by":"publisher","first-page":"119034","DOI":"10.1016\/j.ins.2023.119034","article-title":"A novel parallel merge neural network with streams of spiking neural network and artificial neural network","volume":"642","author":"Yang","year":"2023","journal-title":"Information Sciences"},{"key":"2024082018250622900_bib41","doi-asserted-by":"publisher","first-page":"119648","DOI":"10.1016\/j.ins.2023.119648","article-title":"Pseudo inverse versus iterated projection: Novel learning approach and its application on broad learning system","volume":"649","author":"Yin","year":"2023","journal-title":"Information Sciences"},{"key":"2024082018250622900_bib42","first-page":"9815","article-title":"Adaptive methods for nonconvex optimization","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Zaheer","year":"2018"},{"key":"2024082018250622900_bib43","first-page":"1","article-title":"Adam can converge without any modification on update rules","volume-title":"Advances in neural information processing systems, 35","author":"Zhang","year":"2022"},{"key":"2024082018250622900_bib44","first-page":"11119","article-title":"A sufficient condition for convergences of Adam and RMSProp","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zou","year":"2019"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/36\/9\/1912\/2465946\/neco_a_01692.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/36\/9\/1912\/2465946\/neco_a_01692.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,20]],"date-time":"2024-08-20T18:26:18Z","timestamp":1724178378000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/36\/9\/1912\/123691\/UAdam-Unified-Adam-Type-Algorithmic-Framework-for"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,19]]},"references-count":44,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2024,8,19]]},"published-print":{"date-parts":[[2024,8,19]]}},"URL":"https:\/\/doi.org\/10.1162\/neco_a_01692","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024,9]]},"published":{"date-parts":[[2024,8,19]]}}}