{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T04:43:52Z","timestamp":1787028232338,"version":"3.56.0"},"reference-count":33,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2019,6,26]],"date-time":"2019-06-26T00:00:00Z","timestamp":1561507200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100004358","name":"Samsung","doi-asserted-by":"publisher","award":["SSTF-BA1601-02"],"award-info":[{"award-number":["SSTF-BA1601-02"]}],"id":[{"id":"10.13039\/100004358","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>There has been a growing interest in expressivity of deep neural networks. However, most of the existing work about this topic focuses only on the specific activation function such as ReLU or sigmoid. In this paper, we investigate the approximation ability of deep neural networks with a broad class of activation functions. This class of activation functions includes most of frequently used activation functions. We derive the required depth, width and sparsity of a deep neural network to approximate any H\u00f6lder smooth function upto a given approximation error for the large class of activation functions. Based on our approximation error analysis, we derive the minimax optimality of the deep neural network estimators with the general activation functions in both regression and classification problems.<\/jats:p>","DOI":"10.3390\/e21070627","type":"journal-article","created":{"date-parts":[[2019,6,26]],"date-time":"2019-06-26T07:24:17Z","timestamp":1561533857000},"page":"627","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":64,"title":["Smooth Function Approximation by Deep Neural Networks with General Activation Functions"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3953-5953","authenticated-orcid":false,"given":"Ilsang","family":"Ohn","sequence":"first","affiliation":[{"name":"Department of Statistics, Seoul National University, Seoul 08826, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yongdai","family":"Kim","sequence":"additional","affiliation":[{"name":"Department of Statistics, Seoul National University, Seoul 08826, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,6,26]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"ref_2","unstructured":"Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep Learning, MIT Press."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/BF02551274","article-title":"Approximation by superpositions of a sigmoidal function","volume":"2","author":"Cybenko","year":"1989","journal-title":"Math. Control Signals Syst."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1016\/0893-6080(89)90020-8","article-title":"Multilayer feedforward networks are universal approximators","volume":"2","author":"Hornik","year":"1989","journal-title":"Neural Netw."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1016\/0893-6080(89)90003-8","article-title":"On the approximate realization of continuous mappings by neural networks","volume":"2","author":"Funahashi","year":"1989","journal-title":"Neural Netw."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"131","DOI":"10.1016\/0021-9045(92)90081-X","article-title":"Approximation by ridge functions and neural networks with one hidden layer","volume":"70","author":"Chui","year":"1992","journal-title":"J. Approx. Theory"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"861","DOI":"10.1016\/S0893-6080(05)80131-5","article-title":"Multilayer feedforward networks with a nonpolynomial activation function can approximate any function","volume":"6","author":"Leshno","year":"1993","journal-title":"Neural Netw."},{"key":"ref_8","unstructured":"Telgarsky, M. (2017, January 6\u201311). Neural networks and rational functions. Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1016\/j.neunet.2017.07.002","article-title":"Error bounds for approximations with deep ReLU networks","volume":"94","author":"Yarotsky","year":"2017","journal-title":"Neural Netw."},{"key":"ref_10","unstructured":"Schmidt-Hieber, J. (2017). Nonparametric regression using deep neural networks with ReLU activation function. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Bauer, B., and Kohler, M. (2019). On deep learning as a remedy for the curse of dimensionality in nonparametric regression. Ann. Stat., accepted.","DOI":"10.1214\/18-AOS1747"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Li, B., Tang, S., and Yu, H. (2019). Better Approximations of High Dimensional Smooth Functions by Deep Neural Networks with Rectified Power Units. arXiv.","DOI":"10.4208\/cicp.OA-2019-0168"},{"key":"ref_13","unstructured":"Suzuki, T. (2018). Adaptivity of deep ReLU network for learning in Besov and mixed smooth Besov spaces: Optimal rate and curse of dimensionality. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"296","DOI":"10.1016\/j.neunet.2018.08.019","article-title":"Optimal approximation of piecewise smooth functions using deep ReLU neural networks","volume":"108","author":"Petersen","year":"2018","journal-title":"Neural Netw."},{"key":"ref_15","unstructured":"Imaizumi, M., and Fukumizu, K. (2018). Deep Neural Networks Learn Non-Smooth Functions Effectively. arXiv."},{"key":"ref_16","unstructured":"Bergstra, J., Desjardins, G., Lamblin, P., and Bengio, Y. (2009). Quadratic Polynomials Learn Better Image Features, D\u00e9partement d\u2019Informatique et de Recherche Operationnelle, Universit\u00e9 de Montr\u00e9al. Technical Report 1337."},{"key":"ref_17","unstructured":"Clevert, D.A., Unterthiner, T., and Hochreiter, S. (2015). Fast and accurate deep network learning by exponential linear units (elus). arXiv."},{"key":"ref_18","unstructured":"Carlile, B., Delamarter, G., Kinney, P., Marti, A., and Whitney, B. (2017). Improving deep learning by inverse square root linear units (ISRLUs). arXiv."},{"key":"ref_19","unstructured":"Klimek, M.D., and Perelstein, M. (2018). Neural Network-Based Approach to Phase Space Integration. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wuraola, A., and Patel, N. (2018, January 8\u201313). SQNL: A New Computationally Efficient Activation Function. Proceedings of the 2018 International Joint Conference on Neural Networks (IJCNN), Rio de Janeiro, Brazil.","DOI":"10.1109\/IJCNN.2018.8489043"},{"key":"ref_21","unstructured":"Ramachandran, P., Zoph, B., and Le, Q.V. (2017). Searching for activation functions. arXiv."},{"key":"ref_22","unstructured":"Glorot, X., Bordes, A., and Bengio, Y. (2011, January 11\u201313). Deep sparse rectifier neural networks. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, Ft. Lauderdale, FL, USA."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1007\/BF02070821","article-title":"Approximation properties of a multilayered feedforward artificial neural network","volume":"1","author":"Mhaskar","year":"1993","journal-title":"Adv. Comput. Math."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1555","DOI":"10.1007\/s00025-017-0692-6","article-title":"Saturation classes for max-product neural network operators activated by sigmoidal functions","volume":"72","author":"Costarelli","year":"2017","journal-title":"Results Math."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1007\/s40314-016-0334-8","article-title":"Solving numerically nonlinear systems of balance laws by multivariate sigmoidal functions approximation","volume":"37","author":"Costarelli","year":"2018","journal-title":"Comput. Appl. Math."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1007\/s00025-018-0790-0","article-title":"Estimates for the neural network operators of the max-product type with continuous and p-integrable functions","volume":"73","author":"Costarelli","year":"2018","journal-title":"Results Math."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1007\/s00025-018-0799-4","article-title":"Approximation results in Orlicz spaces for sequences of Kantorovich max-product neural network operators","volume":"73","author":"Costarelli","year":"2018","journal-title":"Results Math."},{"key":"ref_28","unstructured":"Kim, Y., Ohn, I., and Kim, D. (2018). Fast convergence rates of deep neural networks for classification. arXiv."},{"key":"ref_29","unstructured":"Anthony, M., and Bartlett, P.L. (2001). Neural Network Learning: Theoretical Foundations, Cambridge University Press."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1808","DOI":"10.1214\/aos\/1017939240","article-title":"Smooth discrimination analysis","volume":"27","author":"Mammen","year":"1999","journal-title":"Ann. Stat."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1214\/aos\/1079120131","article-title":"Optimal aggregation of classifiers in statistical learning","volume":"32","author":"Tsybakov","year":"2004","journal-title":"Ann. Stat."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1016\/j.spl.2004.03.002","article-title":"A note on margin-based loss functions in classification","volume":"68","author":"Lin","year":"2004","journal-title":"Stat. Probab. Lett."},{"key":"ref_33","unstructured":"Steinwart, I., and Christmann, A. (2008). Support Vector Machines, Springer Science & Business Media."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/21\/7\/627\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:01:22Z","timestamp":1760187682000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/21\/7\/627"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,6,26]]},"references-count":33,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2019,7]]}},"alternative-id":["e21070627"],"URL":"https:\/\/doi.org\/10.3390\/e21070627","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,6,26]]}}}