{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,11]],"date-time":"2026-02-11T14:05:44Z","timestamp":1770818744383,"version":"3.50.1"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2021,6,30]],"date-time":"2021-06-30T00:00:00Z","timestamp":1625011200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Evol. Learn. Optim."],"published-print":{"date-parts":[[2021,6,30]]},"abstract":"<jats:p>The hyper-parameters of a neural network are traditionally designed through a time-consuming process of trial and error that requires substantial expert knowledge. Neural Architecture Search algorithms aim to take the human out of the loop by automatically finding a good set of hyper-parameters for the problem at hand. These algorithms have mostly focused on hyper-parameters such as the architectural configurations of the hidden layers and the connectivity of the hidden neurons, but there has been relatively little work on automating the search for completely new activation functions, which are one of the most crucial hyperparameters to choose. There are some widely used activation functions nowadays that are simple and work well, but nonetheless, there has been some interest in finding better activation functions. The work in the literature has mostly focused on designing new activation functions by hand or choosing from a set of predefined functions while this work presents an evolutionary algorithm to automate the search for completely new activation functions. We compare these new evolved activation functions to other existing and commonly used activation functions. The results are favorable and are obtained from averaging the performance of the activation functions found over 30 runs, with experiments being conducted on 10 different datasets and architectures to ensure the statistical robustness of the study.<\/jats:p>","DOI":"10.1145\/3464384","type":"journal-article","created":{"date-parts":[[2021,7,29]],"date-time":"2021-07-29T14:57:04Z","timestamp":1627570624000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["Evolution of Activation Functions: An Empirical Investigation"],"prefix":"10.1145","volume":"1","author":[{"given":"Andrew","family":"Nader","sequence":"first","affiliation":[{"name":"Department of Computer Science and Mathematics,Lebanese American University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Danielle","family":"Azar","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Mathematics,Lebanese American University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,7,29]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. Retrieved from https:\/\/www.tensorflow.org\/.  Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. Retrieved from https:\/\/www.tensorflow.org\/."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1097\/00004647-200110000-00001"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377930.3389841"},{"key":"e_1_2_1_4_1","unstructured":"Fran\u00e7ois Chollet et\u00a0al. 2015. Keras. Retrieved from https:\/\/github.com\/fchollet\/keras.  Fran\u00e7ois Chollet et\u00a0al. 2015. Keras. Retrieved from https:\/\/github.com\/fchollet\/keras."},{"key":"e_1_2_1_5_1","unstructured":"Djork-Arn\u00e9 Clevert Thomas Unterthiner and Sepp Hochreiter. 2015. Fast and accurate deep network learning by exponential linear units (elus). arXiv:1511.07289. Retrieved from https:\/\/arxiv.org\/abs\/1511.07289.  Djork-Arn\u00e9 Clevert Thomas Unterthiner and Sepp Hochreiter. 2015. Fast and accurate deep network learning by exponential linear units (elus). arXiv:1511.07289. Retrieved from https:\/\/arxiv.org\/abs\/1511.07289."},{"key":"e_1_2_1_6_1","first-page":"7","article-title":"Approximation with artificial neural networks. Faculty of Sciences, Etvs Lornd University","volume":"24","author":"Cs\u00e1ji Bal\u00e1zs Csan\u00e1d","year":"2001","unstructured":"Bal\u00e1zs Csan\u00e1d Cs\u00e1ji . 2001 . Approximation with artificial neural networks. Faculty of Sciences, Etvs Lornd University , Hungary 24 , 48 (2001), 7 . Bal\u00e1zs Csan\u00e1d Cs\u00e1ji. 2001. Approximation with artificial neural networks. Faculty of Sciences, Etvs Lornd University, Hungary 24, 48 (2001), 7.","journal-title":"Hungary"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the 2019 IEEE Congress on Evolutionary Computation (CEC\u201919)","author":"Cui Peiyu","year":"2019","unstructured":"Peiyu Cui , Boris Shabash , and Kay C Wiese . 2019 . EvoDNN-an evolutionary deep neural network with heterogeneous activation functions . In Proceedings of the 2019 IEEE Congress on Evolutionary Computation (CEC\u201919) . IEEE, 2362\u20132369. Peiyu Cui, Boris Shabash, and Kay C Wiese. 2019. EvoDNN-an evolutionary deep neural network with heterogeneous activation functions. In Proceedings of the 2019 IEEE Congress on Evolutionary Computation (CEC\u201919). IEEE, 2362\u20132369."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/3322706.3361996"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969442.2969547"},{"key":"e_1_2_1_10_1","volume-title":"OpenML-Python: An extensible python API for OpenML. arXiv","author":"Feurer Matthias","year":"1911","unstructured":"Matthias Feurer , Jan N. van Rijn , Arlind Kadra , Pieter Gijsbers , Neeratyoy Mallik , Sahithya Ravi , Andreas Mueller , Joaquin Vanschoren , and Frank Hutter . 2019. OpenML-Python: An extensible python API for OpenML. arXiv 1911 .02490. Retrieved from https:\/\/arxiv.org\/abs\/1911.02490. Matthias Feurer, Jan N. van Rijn, Arlind Kadra, Pieter Gijsbers, Neeratyoy Mallik, Sahithya Ravi, Andreas Mueller, Joaquin Vanschoren, and Frank Hutter. 2019. OpenML-Python: An extensible python API for OpenML. arXiv 1911.02490. Retrieved from https:\/\/arxiv.org\/abs\/1911.02490."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/2503308.2503311"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 13th International Conference on Artificial Intelligence and Statistics. 249\u2013256","author":"Glorot Xavier","year":"2010","unstructured":"Xavier Glorot and Yoshua Bengio . 2010 . Understanding the difficulty of training deep feedforward neural networks . In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics. 249\u2013256 . Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics. 249\u2013256."},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the 14th International Conference on Artificial Intelligence and Statistics. 315\u2013323","author":"Glorot Xavier","year":"2011","unstructured":"Xavier Glorot , Antoine Bordes , and Yoshua Bengio . 2011 . Deep sparse rectifier neural networks . In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics. 315\u2013323 . Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011. Deep sparse rectifier neural networks. In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics. 315\u2013323."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.123"},{"key":"e_1_2_1_15_1","unstructured":"Dan Hendrycks and Kevin Gimpel. 2016. Bridging nonlinearities and stochastic regularizers with gaussian error linear units. CoRR abs\/1606.08415 2016.  Dan Hendrycks and Kevin Gimpel. 2016. Bridging nonlinearities and stochastic regularizers with gaussian error linear units. CoRR abs\/1606.08415 2016."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/3360092"},{"key":"e_1_2_1_17_1","volume-title":"Adam: A method for stochastic optimization. arxiv:1412.6980.","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Jimmy Ba . 2014 . Adam: A method for stochastic optimization. arxiv:1412.6980. Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arxiv:1412.6980."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/3294771.3294864"},{"key":"e_1_2_1_19_1","volume-title":"et\u00a0al","author":"Krizhevsky Alex","year":"2009","unstructured":"Alex Krizhevsky , Geoffrey Hinton , et\u00a0al . 2009 . Learning Multiple Layers of Features from Tiny Images. Technical Report. Citeseer . Alex Krizhevsky, Geoffrey Hinton, et\u00a0al. 2009. Learning Multiple Layers of Features from Tiny Images. Technical Report. Citeseer."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/2999134.2999257"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevLett.66.2396"},{"key":"e_1_2_1_22_1","unstructured":"Yann LeCun Corinna Cortes and C. J. Burges. 2010. MNIST handwritten digit database. ATT Labs 2. Retrieved from http:\/\/yann.lecun.com\/exdb\/mnist.  Yann LeCun Corinna Cortes and C. J. Burges. 2010. MNIST handwritten digit database. ATT Labs 2. Retrieved from http:\/\/yann.lecun.com\/exdb\/mnist."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/645754.668382"},{"key":"e_1_2_1_24_1","unstructured":"Lu Lu Yeonjong Shin Yanhui Su and George Em Karniadakis. 2019. Dying ReLU and initialization: Theory and numerical examples. arXiv:1903.06733. Retrieved from https:\/\/arxiv.org\/abs\/1903.06733.  Lu Lu Yeonjong Shin Yanhui Su and George Em Karniadakis. 2019. Dying ReLU and initialization: Theory and numerical examples. arXiv:1903.06733. Retrieved from https:\/\/arxiv.org\/abs\/1903.06733."},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the International Conference on Machine Learning (ICML\u201913)","volume":"30","author":"Maas Andrew L.","unstructured":"Andrew L. Maas , Awni Y. Hannun , and Andrew Y. Ng . 2013. Rectifier nonlinearities improve neural network acoustic models . In Proceedings of the International Conference on Machine Learning (ICML\u201913) , Vol. 30 . 3. Andrew L. Maas, Awni Y. Hannun, and Andrew Y. Ng. 2013. Rectifier nonlinearities improve neural network acoustic models. In Proceedings of the International Conference on Machine Learning (ICML\u201913), Vol. 30. 3."},{"key":"e_1_2_1_26_1","volume-title":"Matthias Urban, Michael Burkart, Max Dippel, Marius Lindauer, and Frank Hutter.","author":"Mendoza Hector","year":"2018","unstructured":"Hector Mendoza , Aaron Klein , Matthias Feurer , Jost Tobias Springenberg , Matthias Urban, Michael Burkart, Max Dippel, Marius Lindauer, and Frank Hutter. 2018 . Towards automatically-tuned deep neural networks. In AutoML: Methods, Sytems, Challenges, Frank Hutter, Lars Kotthoff, and Joaquin Vanschoren (Eds.). Springer , Chapter 7, 141\u2013156. To appear. Hector Mendoza, Aaron Klein, Matthias Feurer, Jost Tobias Springenberg, Matthias Urban, Michael Burkart, Max Dippel, Marius Lindauer, and Frank Hutter. 2018. Towards automatically-tuned deep neural networks. In AutoML: Methods, Sytems, Challenges, Frank Hutter, Lars Kotthoff, and Joaquin Vanschoren (Eds.). Springer, Chapter 7, 141\u2013156. To appear."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377929.3389942"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/3104322.3104425"},{"key":"e_1_2_1_29_1","unstructured":"Tapani Raiko Harri Valpola and Yann LeCun. 2012. Deep learning made easier by linear transformations in perceptrons. In Artificial Intelligence and Statistics. 924\u2013932.  Tapani Raiko Harri Valpola and Yann LeCun. 2012. Deep learning made easier by linear transformations in perceptrons. In Artificial Intelligence and Statistics. 924\u2013932."},{"key":"e_1_2_1_30_1","volume-title":"Le","author":"Ramachandran Prajit","year":"2017","unstructured":"Prajit Ramachandran , Barret Zoph , and Quoc V . Le . 2017 . Searching for activation functions. arXiv:1710.05941. Retrieved from https:\/\/arxiv.org\/abs\/1710.05941. Prajit Ramachandran, Barret Zoph, and Quoc V. Le. 2017. Searching for activation functions. arXiv:1710.05941. Retrieved from https:\/\/arxiv.org\/abs\/1710.05941."},{"key":"e_1_2_1_31_1","unstructured":"Andrew M. Saxe James L. McClelland and Surya Ganguli. 2013. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. arXiv:1312.6120. Retrieved from https:\/\/arxiv.org\/abs\/1312.6120.  Andrew M. Saxe James L. McClelland and Surya Ganguli. 2013. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. arXiv:1312.6120. Retrieved from https:\/\/arxiv.org\/abs\/1312.6120."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.5555\/645754.668388"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3205651.3208282"},{"key":"e_1_2_1_34_1","volume-title":"IET Conference Proceedings. 323\u2013328","author":"Sopena Josep M.","year":"1999","unstructured":"Josep M. Sopena , Enrique Romero , and Rene Alquezar . 1999 . Neural networks with periodic and monotonic activation functions: A comparative study in classification problems . In IET Conference Proceedings. 323\u2013328 . Josep M. Sopena, Enrique Romero, and Rene Alquezar. 1999. Neural networks with periodic and monotonic activation functions: A comparative study in classification problems. In IET Conference Proceedings. 323\u2013328."},{"key":"e_1_2_1_35_1","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1016\/j.tins.2015.05.005","article-title":"Questioning the role of sparse coding in the brain","volume":"38","author":"Spanne Anton","year":"2015","unstructured":"Anton Spanne and Henrik J\u00f6rntell . 2015 . Questioning the role of sparse coding in the brain . Trends Neurosci. 38 , 7 (2015), 417 \u2013 427 . Anton Spanne and Henrik J\u00f6rntell. 2015. Questioning the role of sparse coding in the brain. Trends Neurosci. 38, 7 (2015), 417\u2013427.","journal-title":"Trends Neurosci."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2670313"},{"key":"e_1_2_1_37_1","doi-asserted-by":"crossref","first-page":"89","DOI":"10.1109\/TEVC.2018.2808689","article-title":"Evolving unsupervised deep neural networks for learning meaningful representations","volume":"23","author":"Sun Yanan","year":"2018","unstructured":"Yanan Sun , Gary G. Yen , and Zhang Yi . 2018 . Evolving unsupervised deep neural networks for learning meaningful representations . IEEE Trans. Evol. Comput. 23 , 1 (2018), 89 \u2013 103 . Yanan Sun, Gary G. Yen, and Zhang Yi. 2018. Evolving unsupervised deep neural networks for learning meaningful representations. IEEE Trans. Evol. Comput. 23, 1 (2018), 89\u2013103.","journal-title":"IEEE Trans. Evol. Comput."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/2641190.2641198"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2009.06.006"},{"key":"e_1_2_1_40_1","unstructured":"Han Xiao Kashif Rasul and Roland Vollgraf. 2017. Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms. arXiv:cs.LG\/cs.LG\/1708.07747. Retrieved from https:\/\/arxiv.org\/abs\/cs.LG\/cs.LG\/1708.07747.  Han Xiao Kashif Rasul and Roland Vollgraf. 2017. Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms. arXiv:cs.LG\/cs.LG\/1708.07747. Retrieved from https:\/\/arxiv.org\/abs\/cs.LG\/cs.LG\/1708.07747."},{"key":"e_1_2_1_41_1","unstructured":"Bing Xu Naiyan Wang Tianqi Chen and Mu Li. 2015. Empirical evaluation of rectified activations in convolutional network. arXiv:1505.00853. Retrieved from https:\/\/arxiv.org\/abs\/1505.00853.  Bing Xu Naiyan Wang Tianqi Chen and Mu Li. 2015. Empirical evaluation of rectified activations in convolutional network. arXiv:1505.00853. Retrieved from https:\/\/arxiv.org\/abs\/1505.00853."},{"key":"e_1_2_1_42_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 8697\u20138710","author":"Zoph Barret","unstructured":"Barret Zoph , Vijay Vasudevan , Jonathon Shlens , and Quoc V. Le . 2018. Learning transferable architectures for scalable image recognition . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 8697\u20138710 . Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V. Le. 2018. Learning transferable architectures for scalable image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 8697\u20138710."}],"container-title":["ACM Transactions on Evolutionary Learning and Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3464384","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3464384","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:17:10Z","timestamp":1750191430000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3464384"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,6,30]]},"references-count":42,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,6,30]]}},"alternative-id":["10.1145\/3464384"],"URL":"https:\/\/doi.org\/10.1145\/3464384","relation":{},"ISSN":["2688-299X","2688-3007"],"issn-type":[{"value":"2688-299X","type":"print"},{"value":"2688-3007","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,6,30]]},"assertion":[{"value":"2020-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-07-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}