{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T19:16:30Z","timestamp":1780082190883,"version":"3.54.0"},"reference-count":47,"publisher":"MIT Press - Journals","issue":"6","content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,5,13]]},"abstract":"<jats:p>Despite the fact that the loss functions of deep neural networks are highly nonconvex, gradient-based optimization algorithms converge to approximately the same performance from many random initial points. One thread of work has focused on explaining this phenomenon by numerically characterizing the local curvature near critical points of the loss function, where the gradients are near zero. Such studies have reported that neural network losses enjoy a no-bad-local-minima property, in disagreement with more recent theoretical results. We report here that the methods used to find these putative critical points suffer from a bad local minima problem of their own: they often converge to or pass through regions where the gradient norm has a stationary point. We call these gradient-flat regions, since they arise when the gradient is approximately in the kernel of the Hessian, such that the loss is locally approximately linear, or flat, in the direction of the gradient. We describe how the presence of these regions necessitates care in both interpreting past results that claimed to find critical points of neural network losses and in designing second-order methods for optimizing neural networks.<\/jats:p>","DOI":"10.1162\/neco_a_01388","type":"journal-article","created":{"date-parts":[[2021,4,22]],"date-time":"2021-04-22T23:20:41Z","timestamp":1619133641000},"page":"1469-1497","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":6,"title":["Critical Point-Finding Methods Reveal Gradient-Flat Regions of Deep Network Losses"],"prefix":"10.1162","volume":"33","author":[{"given":"Charles G.","family":"Frye","sequence":"first","affiliation":[{"name":"Redwood Center for Theoretical Neuroscience and Helen Wills Neuroscience Institute, University of California, Berkeley, CA 94720, U.S.A. cfrye59@gmail.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"James","family":"Simon","sequence":"additional","affiliation":[{"name":"Redwood Center for Theoretical Neuroscience and Department of Physics, University of California, Berkeley, CA 94720, U.S.A. james.simon@berkeley.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Neha S.","family":"Wadia","sequence":"additional","affiliation":[{"name":"Redwood Center for Theoretical Neuroscience and Biophysics Graduate Group, University of California, Berkeley, CA 94720, U.S.A. neha.wadia@berkeley.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrew","family":"Ligeralde","sequence":"additional","affiliation":[{"name":"Redwood Center for Theoretical Neuroscience and Biophysics Graduate Group, University of California, Berkeley, CA 94720, U.S.A. ligeralde@berkeley.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Michael R.","family":"DeWeese","sequence":"additional","affiliation":[{"name":"Redwood Center for Theoretical Neuroscience, Helen Wills Neuroscience Institute, Department of Physics, and Biophysics Graduate Group, University of California, Berkeley, CA 94720, U.S.A. deweese@berkeley.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kristofer E.","family":"Bouchard","sequence":"additional","affiliation":[{"name":"Redwood Center for Theoretical Neuroscience and Helen Wills Neuroscience Institute, University of California, Berkeley, CA 94720, USA; and Biological Systems and Engineering Division and Computational Research Division, Lawrence Berkeley National Lab, Berkeley, CA 94720, U.S.A. kebouchard@lbl.gov"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2021,5,13]]},"reference":[{"issue":"25","key":"2021051720034188700_B1","doi-asserted-by":"crossref","first-page":"5356","DOI":"10.1103\/PhysRevLett.85.5356","article-title":"Saddles in the energy landscape probed by supercooled liquids","volume":"85","author":"Angelani","year":"2000","journal-title":"Physical Review Letters"},{"issue":"1","key":"2021051720034188700_B2","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1016\/0893-6080(89)90014-2","article-title":"Neural networks and principal component analysis: Learning from examples without local minima","volume":"2","author":"Baldi","year":"1989","journal-title":"Neural Networks"},{"key":"2021051720034188700_B3","doi-asserted-by":"crossref","first-page":"12585","DOI":"10.1039\/C7CP01108C","article-title":"Energy landscapes for machine learning","volume":"19","author":"Ballard","year":"2017","journal-title":"Phys. Chemistry Chemical Physics"},{"key":"2021051720034188700_B4","doi-asserted-by":"crossref","DOI":"10.1137\/1.9781611972702","author":"Bates","year":"2013","journal-title":"Numerically solving polynomial systems with Bertini (software, environments and tools)"},{"key":"2021051720034188700_B5","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511804441","author":"Boyd","year":"2004","journal-title":"Convex optimization"},{"issue":"25","key":"2021051720034188700_B6","doi-asserted-by":"crossref","first-page":"5360","DOI":"10.1103\/PhysRevLett.85.5360","article-title":"Energy landscape of a Lennard-Jones liquid: Statistics of stationary points","volume":"85","author":"Broderix","year":"2000","journal-title":"Physical Review Letters"},{"issue":"1","key":"2021051720034188700_B7","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1007\/s10107-003-0376-8","article-title":"On the convergence of Newton iterations to non-stationary points.","volume":"99","author":"Byrd","year":"2004","journal-title":"Mathematical Programming"},{"issue":"6","key":"2021051720034188700_B8","doi-asserted-by":"crossref","first-page":"2800","DOI":"10.1063\/1.442352","article-title":"On finding transition states","volume":"75","author":"Cerjan","year":"1981","journal-title":"Journal of Chemical Physics"},{"issue":"4","key":"2021051720034188700_B9","doi-asserted-by":"crossref","first-page":"1810","DOI":"10.1137\/100787921","article-title":"MINRES-QLP: A Krylov subspace method for indefinite or singular symmetric systems","volume":"33","author":"Choi","year":"2011","journal-title":"SIAM Journal on Scientific Computing"},{"key":"2021051720034188700_B10","first-page":"410","volume-title":"Advances in neural information processing systems","author":"Coetzee","year":"1997"},{"key":"2021051720034188700_B11","author":"Dauphin","year":"2014","journal-title":"Identifying and attacking the saddle point problem in high- dimensional non-convex optimization"},{"key":"2021051720034188700_B12","author":"Ding","year":"2019","journal-title":"Sub-optimal local minima exist for almost all over-parameterized neural networks."},{"issue":"9","key":"2021051720034188700_B13","doi-asserted-by":"crossref","first-page":"3777","DOI":"10.1063\/1.1436470","article-title":"Saddle points and dynamics of Lennard-Jones clusters, solids, and supercooled liquids","volume":"116","author":"Doye","year":"2002","journal-title":"Journal of Chemical Physics"},{"key":"2021051720034188700_B14","article-title":"Gradient descent provably optimizes over-parameterized neural networks.","author":"Du","year":"2019","journal-title":"Proceedings of the International Conference on Learning Representations."},{"key":"2021051720034188700_B15","first-page":"2121","article-title":"Adaptive subgradient methods for online learning and stochastic optimization","volume":"12","author":"Duchi","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"2021051720034188700_B16","author":"Frye","year":"2019","journal-title":"Numerically recovering the critical points of a deep linear autoencoder."},{"key":"2021051720034188700_B17","author":"Garipov","year":"2018","journal-title":"Loss surfaces, mode connectivity, and fast ensembling of DNNs"},{"key":"2021051720034188700_B18","article-title":"An investigation into neural net optimization via Hessian eigenvalue density","author":"Ghorbani","year":"2019","journal-title":"Proceedings of Machine Learning Research"},{"key":"2021051720034188700_B19","author":"Goodfellow","year":"2014","journal-title":"Qualitatively characterizing neural network optimization problems."},{"issue":"4","key":"2021051720034188700_B20","doi-asserted-by":"crossref","first-page":"747","DOI":"10.1137\/0720050","article-title":"Analysis of Newton's method at irregular singularities","volume":"20","author":"Griewank","year":"1983","journal-title":"SIAM Journal on Numerical Analysis"},{"key":"2021051720034188700_B21","author":"Holzm\u00fcller","year":"2020","journal-title":"Training two-layer RELU networks with gradient descent is inconsistent"},{"key":"2021051720034188700_B22","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-04247-3","author":"Izmailov","year":"2014","journal-title":"Newton-type methods for optimization and variational problems"},{"key":"2021051720034188700_B23","author":"Jin","year":"2017","journal-title":"How to escape saddle points efficiently"},{"key":"2021051720034188700_B24","author":"Kingma","year":"2014","journal-title":"Adam: A method for stochastic optimization."},{"key":"2021051720034188700_B25","article-title":"The multilinear structure of ReLU networks.","author":"Laurent","year":"2017"},{"key":"2021051720034188700_B26","author":"LeCun","year":"2010","journal-title":"MNIST handwritten digit database"},{"key":"2021051720034188700_B27","first-page":"1246","article-title":"Gradient descent only converges to minimizers.","volume":"49","author":"Lee","year":"2016","journal-title":"Proceedings of the 29th Annual Conference on Learning Theory"},{"key":"2021051720034188700_B28","author":"Li","year":"2018","journal-title":"On the benefit of width for neural networks: Disappearance of bad basins"},{"key":"2021051720034188700_B29","author":"Maclaurin","year":"2016","journal-title":"Modeling, inference and optimization with composable differentiable procedures"},{"key":"2021051720034188700_B30","author":"Martens","year":"2015","journal-title":"Optimizing neural networks with Kronecker-factored approximate curvature"},{"issue":"8","key":"2021051720034188700_B31","doi-asserted-by":"crossref","first-page":"2625","DOI":"10.1021\/ja00763a011","article-title":"Structure of transition states in organic reactions: General theory and an application to the cyclobutene-butadiene isomerization using a semiempirical molecular orbital method","volume":"94","author":"McIver","year":"1972","journal-title":"Journal of the American Chemical Society"},{"key":"2021051720034188700_B32","article-title":"The loss surface of deep linear networks viewed through the algebraic geometry lens.","author":"Mehta","year":"2018"},{"issue":"5","key":"2021051720034188700_B33","doi-asserted-by":"crossref","DOI":"10.1103\/PhysRevE.97.052307","article-title":"Loss surface of XOR artificial neural networks.","volume":"97","author":"Mehta","year":"2018","journal-title":"Physical Review E"},{"key":"2021051720034188700_B34","article-title":"Center for Operations Research and Econometrics Universit\u00e9 catholique de Louvain.","author":"Nesterov","year":"2018","journal-title":"Implementable tensor methods in unconstrained convex optimization"},{"key":"2021051720034188700_B35","volume-title":"Numerical optimization","author":"Nocedal","year":"2006","edition":"2nd"},{"issue":"6","key":"2021051720034188700_B36","doi-asserted-by":"crossref","first-page":"1898","DOI":"10.1137\/S1064827500381239","article-title":"Residual and backward error bounds in minimum residual Krylov subspace methods.","volume":"23","author":"Paige","year":"2002","journal-title":"SIAM Journal on Scientific Computing"},{"key":"2021051720034188700_B37","first-page":"8024","volume-title":"Advances in neural information processing systems","author":"Paszke","year":"2019"},{"key":"2021051720034188700_B38","first-page":"2825","article-title":"Scikit-learn: Machine learning in Python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"Journal of Machine Learning Research"},{"key":"2021051720034188700_B39","article-title":"Geometry of neural network loss surfaces via random matrix theory.","author":"Pennington","year":"2017","journal-title":"Proceedings of the International Conference on Learning Representations"},{"issue":"1","key":"2021051720034188700_B40","doi-asserted-by":"crossref","DOI":"10.1038\/s41467-020-14663-9","article-title":"Complexity control by gradient descent in deep networks.","volume":"11","author":"Poggio","year":"2020","journal-title":"Nature Communications"},{"key":"2021051720034188700_B41","author":"Powell","year":"1970","journal-title":"A hybrid method for nonlinear equations. Numerical methods for nonlinear algebraic equations."},{"key":"2021051720034188700_B42","author":"Ramachandran","year":"2017","journal-title":"Searching for activation functions"},{"key":"2021051720034188700_B43","author":"Roosta","year":"2018","journal-title":"Newton-MR: Newton's method without smoothness or convexity."},{"key":"2021051720034188700_B44","author":"Sagun","year":"2017","journal-title":"Empirical analysis of the Hessian of over-parameterized neural networks"},{"issue":"9","key":"2021051720034188700_B45","doi-asserted-by":"crossref","DOI":"10.1080\/00029890.1993.11990500","article-title":"The fundamental theorem of linear algebra.","volume":"100","author":"Strang","year":"1993","journal-title":"American Mathematical Monthly"},{"key":"2021051720034188700_B46","article-title":"Optimization for deep learning: Theory and algorithms","author":"Sun","year":"2019"},{"key":"2021051720034188700_B47","author":"Zhang","year":"2016","journal-title":"Understanding deep learning requires rethinking generalization."}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/direct.mit.edu\/neco\/article-pdf\/33\/6\/1469\/1916370\/neco_a_01388.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"http:\/\/direct.mit.edu\/neco\/article-pdf\/33\/6\/1469\/1916370\/neco_a_01388.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,5,17]],"date-time":"2021-05-17T20:04:16Z","timestamp":1621281856000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/33\/6\/1469\/100574\/Critical-Point-Finding-Methods-Reveal-Gradient"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,13]]},"references-count":47,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2021,5,13]]},"published-print":{"date-parts":[[2021,5,13]]}},"URL":"https:\/\/doi.org\/10.1162\/neco_a_01388","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2021,6]]},"published":{"date-parts":[[2021,5,13]]}}}