{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T03:08:38Z","timestamp":1782184118447,"version":"3.54.5"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2023,1,4]],"date-time":"2023-01-04T00:00:00Z","timestamp":1672790400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,1,4]],"date-time":"2023-01-04T00:00:00Z","timestamp":1672790400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100018960","name":"National Technical University of Athens","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100018960","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Process Lett"],"published-print":{"date-parts":[[2023,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Various works have been published around the optimization of Neural Networks that emphasize the significance of the learning rate. In this study we analyze the need for a different treatment for each layer and how this affects training. We propose a novel optimization technique, called AdaLip, that utilizes an estimation of the Lipschitz constant of the gradients in order to construct an adaptive learning rate per layer that can work on top of already existing optimizers, like SGD or Adam. A detailed experimental framework was used to prove the usefulness of the optimizer on three benchmark datasets. It showed that AdaLip improves the training performance and the convergence speed, but also made the training process more robust to the selection of the initial global learning rate.<\/jats:p>","DOI":"10.1007\/s11063-022-11140-w","type":"journal-article","created":{"date-parts":[[2023,1,4]],"date-time":"2023-01-04T19:02:32Z","timestamp":1672858952000},"page":"6311-6338","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["AdaLip: An Adaptive Learning Rate Method per Layer for Stochastic Optimization"],"prefix":"10.1007","volume":"55","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6153-7712","authenticated-orcid":false,"given":"George","family":"Ioannou","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thanos","family":"Tagaris","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andreas","family":"Stafylopatis","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,1,4]]},"reference":[{"key":"11140_CR1","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1016\/j.media.2017.07.005","volume":"42","author":"G Litjens","year":"2017","unstructured":"Litjens G, Kooi T, Bejnordi BE, Setio AAA, Ciompi F, Ghafoorian M, Van Der Laak JA, Van Ginneken B, S\u00e1nchez CI (2017) A survey on deep learning in medical image analysis. Med Image Anal 42:60\u201388","journal-title":"Med Image Anal"},{"key":"11140_CR2","first-page":"668","volume":"2","author":"S Minaee","year":"2021","unstructured":"Minaee S, Boykov YY, Porikli F, Plaza AJ, Kehtarnavaz N, Terzopoulos D (2021) Image segmentation using deep learning: a survey. IEEE Trans Pattern Anal Mach Intell 2:668","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"4","key":"11140_CR3","doi-asserted-by":"publisher","first-page":"240","DOI":"10.1080\/02564602.2015.1010611","volume":"32","author":"J Padmanabhan","year":"2015","unstructured":"Padmanabhan J, Johnson Premkumar MJ (2015) Machine learning in automatic speech recognition: a survey. IETE Tech Rev 32(4):240\u2013251","journal-title":"IETE Tech Rev"},{"key":"11140_CR4","doi-asserted-by":"crossref","unstructured":"Kumar A, Verma S, Mangla H( 2018) A survey of deep learning techniques in speech recognition. In: 2018 international conference on advances in computing, communication control and networking (ICACCCN), pp 179\u2013 185. IEEE","DOI":"10.1109\/ICACCCN.2018.8748399"},{"key":"11140_CR5","unstructured":"Yang S, Wang Y, Chu X (2020) A survey of deep learning techniques for neural machine translation. arXiv preprint arXiv:2002.07526"},{"issue":"3","key":"11140_CR6","doi-asserted-by":"publisher","first-page":"362","DOI":"10.1002\/rob.21918","volume":"37","author":"S Grigorescu","year":"2020","unstructured":"Grigorescu S, Trasnea B, Cocias T, Macesanu G (2020) A survey of deep learning techniques for autonomous driving. J Field Robot 37(3):362\u2013386","journal-title":"J Field Robot"},{"key":"11140_CR7","first-page":"998","volume":"6","author":"T Iqbal","year":"2020","unstructured":"Iqbal T, Qureshi S (2020) The survey: text generation models in deep learning. J King Saud Univ Comput Inf Sci 6:998","journal-title":"J King Saud Univ Comput Inf Sci"},{"key":"11140_CR8","unstructured":"Loshchilov I, Hutter F (2016) Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983"},{"key":"11140_CR9","unstructured":"Huang G, Li Y, Pleiss G, Liu Z, Hopcroft JE, Weinberger KQ (2017) Snapshot ensembles: Train 1, get m for free. arXiv preprint arXiv:1704.00109"},{"issue":"3","key":"11140_CR10","doi-asserted-by":"publisher","first-page":"400","DOI":"10.1214\/aoms\/1177729586","volume":"22","author":"H Robbins","year":"1951","unstructured":"Robbins H, Monro S (1951) A stochastic approximation method. Ann Math Stat 22(3):400\u2013407. https:\/\/doi.org\/10.1214\/aoms\/1177729586","journal-title":"Ann Math Stat"},{"key":"11140_CR11","unstructured":"Kleinberg R, Li Y, Yuan Y (2018) An alternative view: when does SGD escape local minima? In: Dy JG, Krause A (eds.) Proceedings of the 35th international conference on machine learning, ICML 2018, Stockholmsm\u00e4ssan, Stockholm, Sweden, July 10\u201315. Proceedings of Machine Learning Research, vol 80, pp 2703\u20132712. PMLR. http:\/\/proceedings.mlr.press\/v80\/kleinberg18a.html"},{"key":"11140_CR12","doi-asserted-by":"publisher","unstructured":"Smith LN (2017) Cyclical learning rates for training neural networks. In: 2017 IEEE winter conference on applications of computer vision, WACV 2017, Santa Rosa, CA, USA, March 24\u201331, pp 464\u2013472. IEEE Computer Society. https:\/\/doi.org\/10.1109\/WACV.2017.58","DOI":"10.1109\/WACV.2017.58"},{"issue":"2","key":"11140_CR13","doi-asserted-by":"publisher","first-page":"223","DOI":"10.1137\/16M1080173","volume":"60","author":"L Bottou","year":"2018","unstructured":"Bottou L, Curtis FE, Nocedal J (2018) Optimization methods for large-scale machine learning. SIAM Rev 60(2):223\u2013311","journal-title":"SIAM Rev"},{"key":"11140_CR14","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J ( 2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 770\u2013 778","DOI":"10.1109\/CVPR.2016.90"},{"key":"11140_CR15","volume-title":"Deep learning with python","author":"F Chollet","year":"2017","unstructured":"Chollet F (2017) Deep learning with python, 1st edn. Manning Publications Co., New York","edition":"1"},{"key":"11140_CR16","unstructured":"Shamir O, Zhang T (2013) Stochastic gradient descent for non-smooth optimization: convergence results and optimal averaging schemes. In: International conference on machine learning, pp 71\u2013 79"},{"key":"11140_CR17","unstructured":"Zinkevich M (xxxx) Online convex programming and generalized infinitesimal gradient ascent. In: Proceedings of the twentieth international conference on international conference on machine learning. ICML\u201903, pp 928\u2013935. AAAI Press"},{"key":"11140_CR18","unstructured":"Wu X, Ward R, Bottou L(2018) Wngrad: learn the learning rate in gradient descent. CoRR abs\/1803.02865. arXiv:1803.02865"},{"key":"11140_CR19","first-page":"2121","volume":"12","author":"JC Duchi","year":"2011","unstructured":"Duchi JC, Hazan E, Singer Y (2011) Adaptive subgradient methods for online learning and stochastic optimization. J Mach Learn Res 12:2121\u20132159","journal-title":"J Mach Learn Res"},{"key":"11140_CR20","first-page":"58","volume":"2","author":"T Tieleman","year":"2012","unstructured":"Tieleman T, Hinton G (2012) Lecture 6.5\u2013RmsProp: divide the gradient by a running average of its recent magnitude. COURSERA Neural Netw Mach Learn 2:58","journal-title":"COURSERA Neural Netw Mach Learn"},{"key":"11140_CR21","unstructured":"Kingma DP, Ba J (2015) Adam: a method for stochastic optimization. In: Bengio Y, LeCun Y (eds.) 3rd International conference on learning representations, ICLR 2015, San Diego, CA, USA, May 7\u20139, 2015, Conference Track Proceedings. arXiv:1412.6980"},{"key":"11140_CR22","unstructured":"Wilson AC, Roelofs R, Stern M, Srebro N, Recht B ( 2017) The marginal value of adaptive gradient methods in machine learning. In: Advances in neural information processing systems, pp 4148\u2013 4158"},{"key":"11140_CR23","unstructured":"Reddi SJ, Kale S, Kumar S (2018) On the convergence of adam and beyond. In: 6th international conference on learning representations, ICLR 2018, Vancouver, BC, Canada, April 30\u2013May 3, 2018, Conference Track Proceedings. OpenReview.net. https:\/\/openreview.net\/forum?id=ryQu7f-RZ"},{"key":"11140_CR24","unstructured":"Luo L, Xiong Y, Liu Y, Sun X (2019) Adaptive gradient methods with dynamic bound of learning rate. In: 7th international conference on learning representations, ICLR 2019, New Orleans, LA, USA, May 6\u20139, 2019. OpenReview.net. https:\/\/openreview.net\/forum?id=Bkg3g2R9FX"},{"key":"11140_CR25","doi-asserted-by":"crossref","unstructured":"Yedida R, Saha S (2019) LipschitzLR: using theoretically computed adaptive learning rates for fast convergence","DOI":"10.1007\/s10489-020-01892-0"},{"key":"11140_CR26","unstructured":"Fazlyab M, Robey A, Hassani H, Morari M, Pappas GJ (2019) Efficient and accurate estimation of lipschitz constants for deep neural networks. In: Wallach HM, Larochelle H, Beygelzimer A, d\u2019Alch\u00e9-Buc F, Fox EB, Garnett R (eds.) Advances in neural information processing systems 32: annual conference on neural information processing systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp 11423\u2013 11434. https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/95e1533eb1b20a97777749fb94fdb944-Abstract.html"},{"key":"11140_CR27","unstructured":"Baydin AG, Cornish R, Mart\u00ednez-Rubio D, Schmidt M, Wood F (2018) Online learning rate adaptation with hypergradient descent. In: 6th international conference on learning representations, ICLR 2018, Vancouver, BC, Canada, April 30\u2013May 3, Conference Track Proceedings. OpenReview.net. https:\/\/openreview.net\/forum?id=BkrsAzWAb"},{"key":"11140_CR28","unstructured":"Ioffe S, Szegedy C (2015) Batch normalization: accelerating deep network training by reducing internal covariate shift. In: Bach FR, Blei DM (eds.) Proceedings of the 32nd international conference on machine learning, ICML 2015, Lille, France, 6\u201311 July. JMLR Workshop and Conference Proceedings, vol 37, pp 448\u2013456. JMLR.org. http:\/\/proceedings.mlr.press\/v37\/ioffe15.html"},{"key":"11140_CR29","unstructured":"Nair V, Hinton GE ( 2010) Rectified linear units improve restricted boltzmann machines. In: Proceedings of the 27th international conference on machine learning (ICML-10), pp 807\u2013 814"},{"key":"11140_CR30","unstructured":"Glorot X, Bengio Y (2010) Understanding the difficulty of training deep feedforward neural networks. In: Proceedings of the thirteenth international conference on artificial intelligence and statistics, pp 249\u2013 256"},{"key":"11140_CR31","unstructured":"LeCun Y, Cortes C (2010) MNIST handwritten digit database"},{"key":"11140_CR32","unstructured":"Krizhevsky A (2009) Learning multiple layers of features from tiny images"},{"key":"11140_CR33","unstructured":"Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition"},{"issue":"1","key":"11140_CR34","first-page":"1929","volume":"15","author":"N Srivastava","year":"2014","unstructured":"Srivastava N, Hinton GE, Krizhevsky A, Sutskever I, Salakhutdinov R (2014) Dropout: a simple way to prevent neural networks from overfitting. J Mach Learn Res 15(1):1929\u20131958","journal-title":"J Mach Learn Res"},{"key":"11140_CR35","unstructured":"Choromanska A, LeCun Y, Arous GB (2015) Open problem: the landscape of the loss surfaces of multilayer networks. In: Gr\u00fcnwald P, Hazan E, Kale S (eds.) Proceedings of The 28th conference on learning theory, COLT 2015, Paris, France, July 3\u20136. JMLR Workshop and Conference Proceedings, vol 40, pp 1756\u20131760. JMLR.org. http:\/\/proceedings.mlr.press\/v40\/Choromanska15.html"},{"key":"11140_CR36","unstructured":"Li H, Xu Z, Taylor G, Studer C, Goldstein T (2018) Visualizing the loss landscape of neural nets. In: Advances in neural information processing systems, pp 6389\u2013 6399"},{"key":"11140_CR37","unstructured":"Ge R, Huang F, Jin C, Yuan Y (2015) Escaping from saddle points-online stochastic gradient for tensor decomposition. In: Gr\u00fcnwald P, Hazan E, Kale S (eds.) Proceedings of The 28th conference on learning theory, COLT 2015, Paris, France, July 3\u20136, . JMLR workshop and conference proceedings, vol 40, pp 797\u2013842. JMLR.org. http:\/\/proceedings.mlr.press\/v40\/Ge15.html"},{"key":"11140_CR38","unstructured":"Jin C, Ge R, Netrapalli,P, Kakade SM, Jordan MI (2017) How to escape saddle points efficiently. In: Precup D, Teh YW (eds.) Proceedings of the 34th international conference on machine learning, ICML 2017, Sydney, NSW, Australia, 6-11 August. Proceedings of machine learning research, vol 70, pp 1724\u20131732. PMLR. http:\/\/proceedings.mlr.press\/v70\/jin17a.html"}],"container-title":["Neural Processing Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-022-11140-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11063-022-11140-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-022-11140-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,29]],"date-time":"2023-09-29T16:20:03Z","timestamp":1696004403000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11063-022-11140-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,4]]},"references-count":38,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2023,10]]}},"alternative-id":["11140"],"URL":"https:\/\/doi.org\/10.1007\/s11063-022-11140-w","relation":{},"ISSN":["1370-4621","1573-773X"],"issn-type":[{"value":"1370-4621","type":"print"},{"value":"1573-773X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,4]]},"assertion":[{"value":"22 December 2022","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 January 2023","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial or non-financial interests to disclose. The authors give consent for publication.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"The code can be found at https:\/\/github.com\/geoioannou\/AdaLip The authors have equal contributed to the work.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Code availability"}}]}}