{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,22]],"date-time":"2026-02-22T06:43:03Z","timestamp":1771742583980,"version":"3.50.1"},"reference-count":46,"publisher":"MIT Press - Journals","issue":"2","content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,1,14]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Recent theoretical studies proved that deep neural network (DNN) estimators obtained by minimizing empirical risk with a certain sparsity constraint can attain optimal convergence rates for regression and classification problems. However, the sparsity constraint requires knowing certain properties of the true model, which are not available in practice. Moreover, computation is difficult due to the discrete nature of the sparsity constraint. In this letter, we propose a novel penalized estimation method for sparse DNNs that resolves the problems existing in the sparsity constraint. We establish an oracle inequality for the excess risk of the proposed sparse-penalized DNN estimator and derive convergence rates for several learning tasks. In particular, we prove that the sparse-penalized estimator can adaptively attain minimax convergence rates for various nonparametric regression problems. For computation, we develop an efficient gradient-based optimization algorithm that guarantees the monotonic reduction of the objective function.<\/jats:p>","DOI":"10.1162\/neco_a_01457","type":"journal-article","created":{"date-parts":[[2021,11,11]],"date-time":"2021-11-11T01:09:48Z","timestamp":1636592988000},"page":"476-517","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":15,"title":["Nonconvex Sparse Regularization for Deep Neural Networks and Its Optimality"],"prefix":"10.1162","volume":"34","author":[{"given":"Ilsang","family":"Ohn","sequence":"first","affiliation":[{"name":"Department of Applied and Computational Mathematics and Statistics, University of Notre Dame, Notre Dame, IN 46556, U.S.A. iohn@nd.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongdai","family":"Kim","sequence":"additional","affiliation":[{"name":"Department of Statistics, Seoul National University, Seoul 08826, Republic of Korea ydkim0903@gmail.com"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2022,1,14]]},"reference":[{"issue":"4","key":"2022040618223930800_B1","doi-asserted-by":"publisher","first-page":"2261","DOI":"10.1214\/18-AOS1747","article-title":"On deep learning as a remedy for the curse of dimensionality in nonparametric regression","volume":"47","author":"Bauer","year":"2019","journal-title":"Annals of Statistics"},{"key":"2022040618223930800_B2","author":"Bergstra","year":"2009","journal-title":"Quadratic polynomials learn better image features"},{"key":"2022040618223930800_B3","author":"Briol","year":"2019","journal-title":"Statistical inference for generative models with maximum mean discrepancy"},{"key":"2022040618223930800_B4","first-page":"453","article-title":"Linear smoothers and additive models","volume":"17","author":"Buja","year":"1989","journal-title":"Annals of Statistics"},{"key":"2022040618223930800_B5","author":"Carlile","year":"2017","journal-title":"Improving deep learning by inverse square root linear units (ISRLUs)"},{"key":"2022040618223930800_B6","author":"Clevert","year":"2015","journal-title":"Fast and accurate deep network learning by exponential linear units (ELUs)"},{"key":"2022040618223930800_B7","first-page":"2121","article-title":"Adaptive subgradient methods for online learning and stochastic optimization","volume":"12","author":"Duchi","year":"2011","journal-title":"Journal of Machine Learning Research"},{"issue":"456","key":"2022040618223930800_B8","doi-asserted-by":"publisher","first-page":"1348","DOI":"10.1198\/016214501753382273","article-title":"Variable selection via nonconcave penalized likelihood and its oracle properties","volume":"96","author":"Fan","year":"2001","journal-title":"Journal of the American Statistical Association"},{"key":"2022040618223930800_B9","author":"Frankle","year":"2018"},{"key":"2022040618223930800_B10","doi-asserted-by":"publisher","first-page":"538","DOI":"10.1214\/07-EJS077","article-title":"Optimal rates and adaptation in the single-index model using aggregation","volume":"1","author":"Gaiffas","year":"2007","journal-title":"Electronic Journal of Statistics"},{"key":"2022040618223930800_B11","first-page":"315","article-title":"Deep sparse rectifier neural networks.","author":"Glorot","year":"2011","journal-title":"Proceedings of the International Conference on Artificial Intelligence and Statistics"},{"key":"2022040618223930800_B12","first-page":"2672","volume-title":"Advances in neural information processing systems","author":"Goodfellow","year":"2014"},{"key":"2022040618223930800_B13","author":"Gy\u00f6rfi","year":"2006"},{"key":"2022040618223930800_B14","first-page":"1135","volume-title":"Advances in neural information processing systems","author":"Han","year":"2015"},{"issue":"6","key":"2022040618223930800_B15","doi-asserted-by":"publisher","first-page":"2589","DOI":"10.1214\/009053607000000415","article-title":"Rate-optimal estimation for a general class of nonparametric regression models with unknown link functions","volume":"35","author":"Horowitz","year":"2007","journal-title":"Annals of Statistics"},{"key":"2022040618223930800_B16","article-title":"Deep neural networks learn non-smooth functions effectively.","author":"Imaizumi","year":"2019","journal-title":"Proceedings of the International Conference on Artificial Intelligence and Statistics"},{"key":"2022040618223930800_B17","author":"Imaizumi","year":"2020"},{"key":"2022040618223930800_B18","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1016\/j.neunet.2021.02.012","article-title":"Fast convergence rates of deep neural networks for classification","volume":"138","author":"Kim","year":"2021","journal-title":"Neural Networks"},{"key":"2022040618223930800_B19","author":"Kingma","year":"2014","journal-title":"Adam: A method for stochastic optimization"},{"key":"2022040618223930800_B20","author":"Kingma","year":"2013"},{"key":"2022040618223930800_B21","author":"Klimek","year":"2018","journal-title":"Neural network\u2013based approach to phase space integration"},{"issue":"4","key":"2022040618223930800_B22","doi-asserted-by":"publisher","first-page":"2231","DOI":"10.1214\/20-AOS2034","article-title":"On the rate of convergence of fully connected deep neural network regression estimates","volume":"49","author":"Kohler","year":"2021","journal-title":"Annals of Statistics"},{"key":"2022040618223930800_B23","article-title":"The MM algorithm.","author":"Lange","year":"2013","journal-title":"Optimization"},{"key":"2022040618223930800_B24","author":"Li","year":"2016","journal-title":"Pruning filters for efficient ConvNets"},{"key":"2022040618223930800_B25","article-title":"On how well generative adversarial networks learn densities: Nonparametric and parametric results.","author":"Liang","year":"2018"},{"key":"2022040618223930800_B26","first-page":"806","article-title":"Sparse convolutional neural networks.","author":"Liu","year":"2015","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"2022040618223930800_B27","author":"Liu","year":"2019","journal-title":"On the variance of the adaptive learning rate and beyond."},{"key":"2022040618223930800_B28","article-title":"Learning sparse neural networks through l0 regularization.","author":"Louizos","year":"2018","journal-title":"Proceedings of the International Conference on Learning Representations"},{"key":"2022040618223930800_B29","author":"Luo","year":"2019","journal-title":"Adaptive gradient methods with dynamic bound of learning rate"},{"issue":"7","key":"2022040618223930800_B30","doi-asserted-by":"publisher","DOI":"10.3390\/e21070627","article-title":"Smooth function approximation by deep neural networks with general activation functions","volume":"21","author":"Ohn","year":"2019","journal-title":"Entropy"},{"issue":"3","key":"2022040618223930800_B31","doi-asserted-by":"publisher","first-page":"127","DOI":"10.1561\/2400000003","article-title":"Proximal algorithms","volume":"1","author":"Parikh","year":"2014","journal-title":"Foundations and Trends, in Optimization"},{"issue":"8","key":"2022040618223930800_B32","doi-asserted-by":"publisher","first-page":"2543","DOI":"10.1016\/j.jspi.2008.11.011","article-title":"Convergence rates of generalization errors for margin-based classification","volume":"139","author":"Park","year":"2009","journal-title":"Journal of Statistical Planning and Inference"},{"key":"2022040618223930800_B33","doi-asserted-by":"publisher","first-page":"296","DOI":"10.1016\/j.neunet.2018.08.019","article-title":"Optimal approximation of piecewise smooth functions using deep ReLU neural networks","volume":"108","author":"Petersen","year":"2018","journal-title":"Neural Networks"},{"key":"2022040618223930800_B34","author":"Ramachandran","year":"2017","journal-title":"Searching for activation functions."},{"issue":"4","key":"2022040618223930800_B35","first-page":"1875","article-title":"Nonparametric regression using deep neural networks with ReLU activation function","volume":"48","author":"Schmidt-Hieber","year":"2020","journal-title":"Annals of Statistics"},{"issue":"2","key":"2022040618223930800_B36","doi-asserted-by":"publisher","first-page":"689","DOI":"10.1214\/aos\/1176349548","article-title":"Additive regression and other nonparametric models","volume":"13","author":"Stone","year":"1985","journal-title":"Annals of Statistics"},{"key":"2022040618223930800_B37","article-title":"Adaptivity of deep ReLU network for learning in Besov and mixed smooth Besov spaces: Optimal rate and curse of dimensionality.","author":"Suzuki","year":"2019","journal-title":"Proceedings of the International Conference on Learning Representations"},{"issue":"1","key":"2022040618223930800_B38","first-page":"289","article-title":"Convex analysis approach to DC programming: Theory, algorithms and applications","volume":"22","author":"Tao","year":"1997","journal-title":"Acta Mathematica Vietnamica"},{"issue":"1","key":"2022040618223930800_B39","doi-asserted-by":"publisher","first-page":"1869","DOI":"10.1214\/21-EJS1828","article-title":"Estimation error analysis of deep learning on the regression problem on the variable exponent Besov space","volume":"15","author":"Tsuji","year":"2021","journal-title":"Electronic Journal of Statistics"},{"key":"2022040618223930800_B40","first-page":"9086","volume-title":"Advances in neural information processing systems","author":"Uppal","year":"2019"},{"key":"2022040618223930800_B41","author":"Wainwright","year":"2019","journal-title":"High-dimensional statistics: A non-asymptotic viewpoint"},{"key":"2022040618223930800_B42","first-page":"2074","volume-title":"Advances in neural information processing systems","author":"Wen","year":"2016"},{"key":"2022040618223930800_B43","doi-asserted-by":"publisher","first-page":"103","DOI":"10.1016\/j.neunet.2017.07.002","article-title":"Error bounds for approximations with deep ReLU networks","volume":"94","author":"Yarotsky","year":"2017","journal-title":"Neural Networks"},{"issue":"4","key":"2022040618223930800_B44","doi-asserted-by":"publisher","first-page":"915","DOI":"10.1162\/08997660360581958","article-title":"The concave-convex procedure","volume":"15","author":"Yuille","year":"2003","journal-title":"Neural Computation"},{"issue":"2","key":"2022040618223930800_B45","doi-asserted-by":"publisher","first-page":"894","DOI":"10.1214\/09-AOS729","article-title":"Nearly unbiased variable selection under minimax concave penalty","volume":"38","author":"Zhang","year":"2010","journal-title":"Annals of Statistics"},{"key":"2022040618223930800_B46","first-page":"1081","article-title":"Analysis of multi-stage convex relaxation for sparse regularization","volume":"11","author":"Zhang","year":"2010","journal-title":"Journal of Machine Learning Research"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/34\/2\/476\/2006817\/neco_a_01457.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/34\/2\/476\/2006817\/neco_a_01457.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,4,6]],"date-time":"2022-04-06T18:23:14Z","timestamp":1649269394000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/34\/2\/476\/107907\/Nonconvex-Sparse-Regularization-for-Deep-Neural"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,14]]},"references-count":46,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2022,1,14]]},"published-print":{"date-parts":[[2022,1,14]]}},"URL":"https:\/\/doi.org\/10.1162\/neco_a_01457","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022,2]]},"published":{"date-parts":[[2022,1,14]]}}}