{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,15]],"date-time":"2026-01-15T05:44:49Z","timestamp":1768455889104,"version":"3.49.0"},"reference-count":34,"publisher":"MIT Press - Journals","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Neural Computation"],"published-print":{"date-parts":[[2019,3]]},"abstract":"<jats:p> We analyze algorithms for approximating a function [Formula: see text] mapping [Formula: see text] to [Formula: see text] using deep linear neural networks, that is, that learn a function [Formula: see text] parameterized by matrices [Formula: see text] and defined by [Formula: see text]. We focus on algorithms that learn through gradient descent on the population quadratic loss in the case that the distribution over the inputs is isotropic. We provide polynomial bounds on the number of iterations for gradient descent to approximate the least-squares matrix [Formula: see text], in the case where the initial hypothesis [Formula: see text] has excess loss bounded by a small enough constant. We also show that gradient descent fails to converge for [Formula: see text] whose distance from the identity is a larger constant, and we show that some forms of regularization toward the identity in each layer do not help. If [Formula: see text] is symmetric positive definite, we show that an algorithm that initializes [Formula: see text] learns an [Formula: see text]-approximation of [Formula: see text] using a number of updates polynomial in [Formula: see text], the condition number of [Formula: see text], and [Formula: see text]. In contrast, we show that if the least-squares matrix [Formula: see text] is symmetric and has a negative eigenvalue, then all members of a class of algorithms that perform gradient descent with identity initialization, and optionally regularize toward the identity in each layer, fail to converge. We analyze an algorithm for the case that [Formula: see text] satisfies [Formula: see text] for all [Formula: see text] but may not be symmetric. This algorithm uses two regularizers: one that maintains the invariant [Formula: see text] for all [Formula: see text] and the other that \u201cbalances\u201d [Formula: see text] so that they have the same singular values. <\/jats:p>","DOI":"10.1162\/neco_a_01164","type":"journal-article","created":{"date-parts":[[2019,1,15]],"date-time":"2019-01-15T18:18:06Z","timestamp":1547576286000},"page":"477-502","source":"Crossref","is-referenced-by-count":17,"title":["Gradient Descent with Identity Initialization Efficiently Learns Positive-Definite Linear Transformations by Deep Residual Networks"],"prefix":"10.1162","volume":"31","author":[{"given":"Peter L.","family":"Bartlett","sequence":"first","affiliation":[{"name":"Department of Statistics, University of California, Berkeley, Berkeley, CA 94720-3860, U.S.A."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David P.","family":"Helmbold","sequence":"additional","affiliation":[{"name":"Computer Science Department, University of California Santa Cruz, Santa Cruz, CA 95064, U.S.A."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Philip M.","family":"Long","sequence":"additional","affiliation":[{"name":"Google, Mountain View, CA 94043, U.S.A."}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","reference":[{"key":"B1","author":"Andoni A.","year":"2014","journal-title":"Proceedings of Machine Learning Research"},{"key":"B2","first-page":"584","author":"Arora S.","year":"2014","journal-title":"Proceedings of the International Conference on Machine Learning Research"},{"key":"B3","author":"Bartlett P. L.","year":"2018","journal-title":"Representing smooth functions as compositions of near-identity functions with implications for deep network optimization"},{"key":"B4","first-page":"520","author":"Bartlett P. L.","year":"2018","journal-title":"JMLR Workshop and Conference Proceedings"},{"key":"B5","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511804441"},{"key":"B6","first-page":"605","author":"Brutzkus A.","year":"2017","journal-title":"Proceedings of the 34th International Conference on Machine Learning"},{"key":"B7","author":"Brutzkus A.","year":"2018","journal-title":"Proceedings of the International Conference on Learning Representations"},{"key":"B8","doi-asserted-by":"publisher","DOI":"10.1090\/S0002-9939-1966-0202740-6"},{"key":"B9","author":"Daniely A.","year":"2017","journal-title":"Presented at the Annual Conference on Neural Information Processing Systems"},{"key":"B10","author":"Ge R.","year":"2017","journal-title":"No spurious local minima in nonconvex low rank problems: A unified geometric analysis"},{"key":"B11","author":"Ge R.","year":"2017","journal-title":"Learning one-hidden-layer neural networks with landscape design"},{"key":"B12","author":"Ge R.","year":"2018","journal-title":"Proceedings of the International Conference on Learning Representations"},{"key":"B13","author":"Hardt M.","year":"2017","journal-title":"Proceedings of the International Conference on Learning Representations"},{"key":"B14","doi-asserted-by":"publisher","DOI":"10.1007\/b98818"},{"key":"B15","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"B16","volume-title":"Topics in matrix analysis","author":"Horn R. A.","year":"1986"},{"key":"B17","volume-title":"Matrix analysis","author":"Horn R. A.","year":"2013","edition":"2"},{"key":"B18","first-page":"479","author":"Jain P.","year":"2017","journal-title":"Proceedings of the 20th International Conference on Artificial Intelligence and Statistics"},{"key":"B19","author":"Janzamin M.","year":"2015","journal-title":"Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods"},{"key":"B20","first-page":"586","volume-title":"Advances in neural information processing systems","author":"Kawaguchi K.","year":"2016"},{"key":"B21","first-page":"1246","author":"Lee J. D.","year":"2016","journal-title":"Proceedings of the 29th Annual Conference on Learning Theory"},{"key":"B22","doi-asserted-by":"publisher","DOI":"10.1109\/18.556601"},{"key":"B23","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2013.2237919"},{"key":"B24","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-70139-4"},{"key":"B25","first-page":"855","volume-title":"Advances in neural information processing systems","author":"Livni R.","year":"2014"},{"key":"B26","first-page":"2603","author":"Nguyen Q.","year":"2017","journal-title":"Proceedings of Machine Learning Research"},{"key":"B27","author":"Orhan A. E.","year":"2018","journal-title":"Proceedings of the International Conference on Learning Representations"},{"key":"B28","first-page":"774","author":"Safran I.","year":"2016","journal-title":"Proceedings of the International Conference on Machine Learning"},{"key":"B29","author":"Saxe A. M.","year":"2013","journal-title":"Exact solutions to the nonlinear dynamics of learning in deep linear neural networks"},{"key":"B31","first-page":"2502","volume-title":"Advances in neural information processing systems","volume":"30","author":"Taghvaei A.","year":"2017"},{"key":"B32","author":"Zhang Q.","year":"2018","journal-title":"Proceedings of the Conference on Innovations in Theoretical Computer Science"},{"key":"B33","first-page":"993","author":"Zhang Y.","year":"2016","journal-title":"Proceedings of the International Conference on Machine Learning"},{"key":"B34","first-page":"83","author":"Zhang Y.","year":"2017","journal-title":"Proceedings of the 20th International Conference on Artificial Intelligence and Statistics"},{"key":"B35","first-page":"4140","author":"Zhong K.","year":"2017","journal-title":"Proceedings of the Conference on Machine Learning"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mitpressjournals.org\/doi\/pdf\/10.1162\/neco_a_01164","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,3,12]],"date-time":"2021-03-12T21:43:00Z","timestamp":1615585380000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/31\/3\/477-502\/8456"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,3]]},"references-count":34,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2019,3]]}},"alternative-id":["10.1162\/neco_a_01164"],"URL":"https:\/\/doi.org\/10.1162\/neco_a_01164","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,3]]}}}