{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T17:43:54Z","timestamp":1776879834831,"version":"3.51.2"},"reference-count":20,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2024,6,20]],"date-time":"2024-06-20T00:00:00Z","timestamp":1718841600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,6,20]],"date-time":"2024-06-20T00:00:00Z","timestamp":1718841600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001871","name":"Funda\u00e7\u00e3o para a Ci\u00eancia e a Tecnologia","doi-asserted-by":"publisher","award":["UI\/BD\/151053\/2021"],"award-info":[{"award-number":["UI\/BD\/151053\/2021"]}],"id":[{"id":"10.13039\/501100001871","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001871","name":"Funda\u00e7\u00e3o para a Ci\u00eancia e a Tecnologia","doi-asserted-by":"publisher","award":["UID\/CEC\/00326\/2020"],"award-info":[{"award-number":["UID\/CEC\/00326\/2020"]}],"id":[{"id":"10.13039\/501100001871","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["SN COMPUT. SCI."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Artificial neural networks are a staple of modern artificial intelligence. These systems must often undergo a training procedure to learn how to solve a designated task. Properly choosing and tuning an optimizer for a problem can significantly improve training speed and quality. Research into optimizers focuses on creating solutions that are generally applicable to any task, relying on parameter tuning for specialization. While parameter tuning is successful in adapting optimizers for a specific problem, the benefits of creating specialized optimizers from scratch is still underdeveloped. We propose an evolutionary framework called AutoLR, capable of evolving optimizers for specific tasks. We use the framework to evolve optimizers for a popular image classification problem and found that evolved optimizers are competitive with human-made optimizers. Furthermore, we find that the evolved solutions remain competitive when moved to a different dataset. Analysis of the best performing optimizers reveals that the system evolved novel behavior undiscovered by humans. The results achieved in this work suggest evolutionary algorithms can improve the quality of neural network training, motivating further research into the framework and its applications.<\/jats:p>","DOI":"10.1007\/s42979-024-02972-5","type":"journal-article","created":{"date-parts":[[2024,6,20]],"date-time":"2024-06-20T06:01:31Z","timestamp":1718863291000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["How to Improve Neural Network Training Using Evolutionary Algorithms"],"prefix":"10.1007","volume":"5","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3845-4617","authenticated-orcid":false,"given":"Pedro","family":"Carvalho","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nuno","family":"Louren\u00e7o","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Penousal","family":"Machado","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,6,20]]},"reference":[{"issue":"8","key":"2972_CR1","doi-asserted-by":"publisher","first-page":"1915","DOI":"10.1109\/TPAMI.2012.231","volume":"35","author":"C Farabet","year":"2012","unstructured":"Farabet C, Couprie C, Najman L, LeCun Y. Learning hierarchical features for scene labeling. IEEE Trans Pattern Anal Mach Intell. 2012;35(8):1915\u201329.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"6","key":"2972_CR2","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1145\/3065386","volume":"60","author":"A Krizhevsky","year":"2017","unstructured":"Krizhevsky A, Sutskever I, Hinton GE. Imagenet classification with deep convolutional neural networks. Commun ACM. 2017;60(6):84\u201390. https:\/\/doi.org\/10.1145\/3065386.","journal-title":"Commun ACM"},{"issue":"8","key":"2972_CR3","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter S, Schmidhuber J. Long short-term memory. Neural Comput. 1997;9(8):1735\u201380.","journal-title":"Neural Comput"},{"key":"2972_CR4","unstructured":"Lopez MM, Kalita J. Deep learning applied to nlp. 2017. arXiv:1703.03091."},{"key":"2972_CR5","unstructured":"Pedro C. AutoLR. GitHub. 2020. https:\/\/github.com\/soren5\/autolr."},{"key":"2972_CR6","doi-asserted-by":"crossref","unstructured":"Louren\u00e7o N, Assun\u00e7\u00e3o F, Pereira FB, Costa E, Machado P. Structured grammatical evolution: a dynamic approach. In: Handbook of grammatical evolution. Springer; 2018. p. 137\u201361.","DOI":"10.1007\/978-3-319-78717-6_6"},{"key":"2972_CR7","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1007\/978-3-031-02056-8_1","volume-title":"Genetic programming","author":"P Carvalho","year":"2022","unstructured":"Carvalho P, Louren\u00e7o N, Machado P. Evolving adaptive neural network optimizers for image classification. In: Medvet E, Pappa G, Xue B, editors. Genetic programming. Cham: Springer; 2022. p. 3\u201318."},{"key":"2972_CR8","unstructured":"Abadi M, Agarwal A, Barham P, Brevdo E, Chen Z, Citro C, Corrado GS, Davis A, Dean J, Devin M, Ghemawat S, Goodfellow I, Harp A, Irving G, Isard M, Jia Y, Jozefowicz R, Kaiser L, Kudlur M, Levenberg J, Man\u00e9 D, Monga R, Moore S, Murray D, Olah C, Schuster M, Shlens J, Steiner B, Sutskever I, Talwar K, Tucker P, Vanhoucke V, Vasudevan V, Vi\u00e9gas F, Vinyals O, Warden P, Wattenberg M, Wicke M, Yu Y, Zheng X. TensorFlow: large-scale machine learning on heterogeneous systems. Software available from tensorflow.org. 2015. https:\/\/www.tensorflow.org\/."},{"key":"2972_CR9","unstructured":"Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, Killeen T, Lin Z, Gimelshein N, Antiga L, Desmaison A, Kopf A, Yang E, DeVito Z, Raison M, Tejani A, Chilamkurthy S, Steiner B, Fang L, Bai J, Chintala S. Pytorch: an imperative style, high-performance deep learning library. In: Advances in neural information processing systems, vol 32. Curran Associates, Inc. p. 8024\u201335. http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf."},{"key":"2972_CR10","doi-asserted-by":"crossref","unstructured":"Rumelhart DE, Hinton GE, Williams RJ. Learning internal representations by error propagation. Technical report, California Univ San Diego La Jolla Inst for Cognitive Science; 1985.","DOI":"10.21236\/ADA164453"},{"key":"2972_CR11","first-page":"9","volume-title":"On-line learning and stochastic approximations","author":"L Bottou","year":"1999","unstructured":"Bottou L. On-line learning and stochastic approximations. Cambridge: Cambridge University Press; 1999. p. 9\u201342."},{"key":"2972_CR12","doi-asserted-by":"crossref","unstructured":"Senior A, Heigold G, Ranzato M, Yang K. An empirical study of learning rates in deep neural networks for speech recognition. In: 2013 IEEE international conference on acoustics, speech and signal processing. IEEE; 2013. p. 6724\u201328.","DOI":"10.1109\/ICASSP.2013.6638963"},{"issue":"4","key":"2972_CR13","doi-asserted-by":"publisher","first-page":"295","DOI":"10.1016\/0893-6080(88)90003-2","volume":"1","author":"RA Jacobs","year":"1988","unstructured":"Jacobs RA. Increased rates of convergence through learning rate adaptation. Neural Netw. 1988;1(4):295\u2013307.","journal-title":"Neural Netw"},{"key":"2972_CR14","first-page":"543","volume":"269","author":"Y Nesterov","year":"1983","unstructured":"Nesterov Y. A method for unconstrained convex minimization problem with the rate of convergence o (1\/k2). Dokl Ussr. 1983;269:543\u20137.","journal-title":"Dokl Ussr"},{"key":"2972_CR15","unstructured":"Geoffrey\u00a0H, Nitish\u00a0Srivastava KS. Overview of mini-batch gradient descent. University Lecture; 2015. https:www.cs.toronto.edu$$\\backslash $$texttildelowtijmencsc321slideslecture_ slides_lec6.pdf."},{"key":"2972_CR16","unstructured":"Kingma DP, Ba J. Adam: a method for stochastic optimization. 2014. arXiv:1412.6980."},{"key":"2972_CR17","unstructured":"Koza J. On the programming of computers by means of natural selection. Genetic programming. 1992."},{"key":"2972_CR18","doi-asserted-by":"crossref","unstructured":"Louren\u00e7o N, Pereira FB, Costa E. Sge: a structured representation for grammatical evolution. In: International conference on artificial evolution (evolution Artificielle). Springer; 2015. p. 136\u201348.","DOI":"10.1007\/978-3-319-31471-6_11"},{"key":"2972_CR19","unstructured":"Mockus J, Tiesis V, Zilinskas A. The application of Bayesian methods for seeking the extremum, vol 2. 2014. p. 117\u201329."},{"key":"2972_CR20","unstructured":"Pedro C. AutoLR. Google Colab. 2022. https:\/\/colab.research.google.com\/drive\/1N06JBM-7DMj91OB9v3aeqXYxHXeS6azx?usp=sharing."}],"container-title":["SN Computer Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s42979-024-02972-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s42979-024-02972-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s42979-024-02972-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,20]],"date-time":"2024-06-20T06:22:44Z","timestamp":1718864564000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s42979-024-02972-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,20]]},"references-count":20,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2024,8]]}},"alternative-id":["2972"],"URL":"https:\/\/doi.org\/10.1007\/s42979-024-02972-5","relation":{},"ISSN":["2661-8907"],"issn-type":[{"value":"2661-8907","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,20]]},"assertion":[{"value":"4 October 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 May 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 June 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"On behalf of all authors, the corresponding author states that there is no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"664"}}