{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T04:19:18Z","timestamp":1772252358539,"version":"3.50.1"},"reference-count":36,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2021,12,24]],"date-time":"2021-12-24T00:00:00Z","timestamp":1640304000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Gradient-based methods are popularly used in training neural networks and can be broadly categorized into first and second order methods. Second order methods have shown to have better convergence compared to first order methods, especially in solving highly nonlinear problems. The BFGS quasi-Newton method is the most commonly studied second order method for neural network training. Recent methods have been shown to speed up the convergence of the BFGS method using the Nesterov\u2019s acclerated gradient and momentum terms. The SR1 quasi-Newton method, though less commonly used in training neural networks, is known to have interesting properties and provide good Hessian approximations when used with a trust-region approach. Thus, this paper aims to investigate accelerating the Symmetric Rank-1 (SR1) quasi-Newton method with the Nesterov\u2019s gradient for training neural networks, and to briefly discuss its convergence. The performance of the proposed method is evaluated on a function approximation and image classification problem.<\/jats:p>","DOI":"10.3390\/a15010006","type":"journal-article","created":{"date-parts":[[2021,12,24]],"date-time":"2021-12-24T08:38:46Z","timestamp":1640335126000},"page":"6","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Accelerating Symmetric Rank-1 Quasi-Newton Method with Nesterov\u2019s Gradient for Training Neural Networks"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7714-1659","authenticated-orcid":false,"given":"S.","family":"Indrapriyadarsini","sequence":"first","affiliation":[{"name":"Graduate School of Science and Technology, Shizuoka University, Hamamatsu 432-8561, Shizuoka, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9577-3721","authenticated-orcid":false,"given":"Shahrzad","family":"Mahboubi","sequence":"additional","affiliation":[{"name":"Graduate School of Electrical and Information Engineering, Shonan Institute of Technology, Fujisawa 251-8511, Kanagawa, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3418-3178","authenticated-orcid":false,"given":"Hiroshi","family":"Ninomiya","sequence":"additional","affiliation":[{"name":"Graduate School of Electrical and Information Engineering, Shonan Institute of Technology, Fujisawa 251-8511, Kanagawa, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Takeshi","family":"Kamio","sequence":"additional","affiliation":[{"name":"Graduate School of Information Sciences, Hiroshima City University, Hiroshima 731-3194, Shizuoka, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7019-2712","authenticated-orcid":false,"given":"Hideki","family":"Asai","sequence":"additional","affiliation":[{"name":"Research Institute of Electronics, Shizuoka University, Hamamatsu 432-8561, Shizuoka, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,12,24]]},"reference":[{"key":"ref_1","first-page":"217","article-title":"Large scale online learning","volume":"16","author":"Bottou","year":"2004","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Bottou, L. (2010). Large-scale machine learning with stochastic gradient descent. Proceedings of COMPSTAT\u20192010, Springer.","DOI":"10.1007\/978-3-7908-2604-3_16"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"400","DOI":"10.1214\/aoms\/1177729586","article-title":"A stochastic approximation method","volume":"22","author":"Robbins","year":"1951","journal-title":"Ann. Math. Stat."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"4649","DOI":"10.1109\/TNNLS.2019.2957003","article-title":"Accelerating minibatch stochastic gradient descent using typicality sampling","volume":"31","author":"Peng","year":"2019","journal-title":"IEEE Trans. Neural Networks Learn. Syst."},{"key":"ref_5","first-page":"315","article-title":"Accelerating stochastic gradient descent using predictive variance reduction","volume":"26","author":"Johnson","year":"2013","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_6","first-page":"543","article-title":"A method for solving the convex programming problem with convergence rate O(1\/k\u02c62)","volume":"269","author":"Nesterov","year":"1983","journal-title":"Dokl. Akad. Nauk Sssr"},{"key":"ref_7","first-page":"2121","article-title":"Adaptive subgradient methods for online learning and stochastic optimization","volume":"12","author":"Duchi","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref_8","first-page":"26","article-title":"Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude","volume":"4","author":"Tieleman","year":"2012","journal-title":"Neural Netw. Mach. Learn."},{"key":"ref_9","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_10","first-page":"735","article-title":"Deep learning via Hessian-free optimization","volume":"27","author":"Martens","year":"2010","journal-title":"ICML"},{"key":"ref_11","unstructured":"Roosta-Khorasani, F., and Mahoney, M.W. (2016). Sub-sampled Newton methods I: Globally convergent algorithms. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"46","DOI":"10.1137\/1019005","article-title":"Quasi-Newton methods, motivation and theory","volume":"19","author":"Dennis","year":"1977","journal-title":"SIAM Rev."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"6089","DOI":"10.1109\/TSP.2014.2357775","article-title":"RES: Regularized stochastic BFGS algorithm","volume":"62","author":"Mokhtari","year":"2014","journal-title":"IEEE Trans. Signal Process."},{"key":"ref_14","first-page":"3151","article-title":"Global convergence of online limited memory BFGS","volume":"16","author":"Mokhtari","year":"2015","journal-title":"J. Mach. Learn. Res."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1008","DOI":"10.1137\/140954362","article-title":"A stochastic quasi-Newton method for large-scale optimization","volume":"26","author":"Byrd","year":"2016","journal-title":"SIAM J. Optim."},{"key":"ref_16","first-page":"436","article-title":"A stochastic quasi-Newton method for online convex optimization","volume":"26","author":"Schraudolph","year":"2007","journal-title":"Artif. Intell. Stat."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1025","DOI":"10.1137\/S1052623493252985","article-title":"Analysis of a symmetric rank-one trust region method","volume":"6","author":"Byrd","year":"1996","journal-title":"SIAM J. Optim."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"245","DOI":"10.1007\/s10589-016-9868-3","article-title":"On solving L-SR1 trust-region subproblems","volume":"66","author":"Brust","year":"2017","journal-title":"Comput. Optim. Appl."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"273","DOI":"10.1023\/A:1011259905470","article-title":"A modified rank one update which converges Q-superlinearly","volume":"19","author":"Spellucci","year":"2001","journal-title":"Comput. Optim. Appl."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"392","DOI":"10.1016\/j.camwa.2011.05.022","article-title":"A symmetric rank-one method based on extra updating techniques for unconstrained optimization","volume":"62","author":"Modarres","year":"2011","journal-title":"Comput. Math. Appl."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1137\/0803001","article-title":"A theoretical and experimental study of the symmetric rank-one update","volume":"3","author":"Khalfan","year":"1993","journal-title":"SIAM J. Optim."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Jahani, M., Nazari, M., Rusakov, S., Berahas, A.S., and Tak\u00e1\u010d, M. (2020, January 19\u201323). Scaling up quasi-newton algorithms: Communication efficient distributed sr1. Proceedings of the International Conference on Machine Learning, Optimization, and Data Science, Siena, Italy.","DOI":"10.1007\/978-3-030-64583-0_5"},{"key":"ref_23","first-page":"1","article-title":"Quasi-Newton methods for machine learning: Forget the past, just sample","volume":"36","author":"Berahas","year":"2021","journal-title":"Optim. Methods Softw."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1587\/nolta.8.289","article-title":"A novel quasi-Newton-based optimization for neural network training incorporating Nesterov\u2019s accelerated gradient","volume":"8","author":"Ninomiya","year":"2017","journal-title":"Nonlinear Theory Its Appl. IEICE"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"554","DOI":"10.1587\/nolta.12.554","article-title":"Momentum acceleration of quasi-Newton based optimization technique for neural network training","volume":"12","author":"Mahboubi","year":"2021","journal-title":"Nonlinear Theory Its Appl. IEICE"},{"key":"ref_26","unstructured":"Sutskever, I., Martens, J., Dahl, G.E., and Hinton, G.E. (2013, January 16\u201321). On the importance of initialization and momentum in deep learning. Proceedings of the 30th International Conference on Machine Learning, Atlanta, GA, USA."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"715","DOI":"10.1007\/s10208-013-9150-3","article-title":"Adaptive restart for accelerated gradient schemes","volume":"15","author":"Candes","year":"2015","journal-title":"Found. Comput. Math."},{"key":"ref_28","unstructured":"Nocedal, J., and Wright, S.J. (2006). Numerical Optimization, Springer. [2nd ed.]."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Mahboubi, S., Indrapriyadarsini, S., Ninomiya, H., and Asai, H. (2019). Momentum Acceleration of Quasi-Newton Training for Neural Networks. Pacific Rim International Conference on Artificial Intelligence, Springer.","DOI":"10.1007\/978-3-030-29911-8_21"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"129","DOI":"10.1007\/BF01582063","article-title":"Representations of quasi-Newton matrices and their use in limited memory methods","volume":"63","author":"Byrd","year":"1994","journal-title":"Math. Program."},{"key":"ref_31","unstructured":"Lu, X., and Byrd, R.H. (1996). A Study of the Limited Memory Sr1 Method in Practice. [Ph.D. Thesis, University of Colorado at Boulder]."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"47","DOI":"10.1137\/0722003","article-title":"A family of trust-region-based algorithms for unconstrained minimization with strong global convergence properties","volume":"22","author":"Shultz","year":"1985","journal-title":"SIAM J. Numer. Anal."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Indrapriyadarsini, S., Mahboubi, S., Ninomiya, H., and Asai, H. (2019). A Stochastic Quasi-Newton Method with Nesterov\u2019s Accelerated Gradient. ECML-PKDD, Springer.","DOI":"10.1007\/978-3-030-46150-8_43"},{"key":"ref_34","first-page":"323","article-title":"A Novel Training Algorithm based on Limited-Memory quasi-Newton method with Nesterov\u2019s Accelerated Gradient in Neural Networks and its Application to Highly-Nonlinear Modeling of Microwave Circuit","volume":"11","author":"Mahboubi","year":"2018","journal-title":"IARIA Int. J. Adv. Softw."},{"key":"ref_35","unstructured":"Indrapriyadarsini, S., Mahboubi, S., Ninomiya, H., Takeshi, K., and Asai, H. (2021, January 6\u20138). A modified limited memory Nesterov\u2019s accelerated quasi-Newton. Proceedings of the NOLTA Society Conference, IEICE, Online."},{"key":"ref_36","first-page":"414","article-title":"Adaptive regularization of weight vectors","volume":"22","author":"Crammer","year":"2009","journal-title":"Adv. Neural Inf. Process. Syst."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/15\/1\/6\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:52:52Z","timestamp":1760169172000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/15\/1\/6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,24]]},"references-count":36,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,1]]}},"alternative-id":["a15010006"],"URL":"https:\/\/doi.org\/10.3390\/a15010006","relation":{"has-preprint":[{"id-type":"doi","id":"10.20944\/preprints202112.0097.v1","asserted-by":"object"},{"id-type":"doi","id":"10.20944\/preprints202112.0097.v2","asserted-by":"object"}]},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12,24]]}}}