{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:12:55Z","timestamp":1750219975251,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":14,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,11,25]],"date-time":"2022-11-25T00:00:00Z","timestamp":1669334400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"the National Key Research and Development Program of China","award":["2021ZD0113004"],"award-info":[{"award-number":["2021ZD0113004"]}]},{"name":"Shandong Provincial Natural Science Foundation","award":["ZR2021QF067"],"award-info":[{"award-number":["ZR2021QF067"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,11,25]]},"DOI":"10.1145\/3573834.3574539","type":"proceedings-article","created":{"date-parts":[[2023,1,18]],"date-time":"2023-01-18T03:29:17Z","timestamp":1674012557000},"page":"1-6","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Layer-wise based Adabelief Optimization Algorithm for Deep Learning"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2488-0060","authenticated-orcid":false,"given":"Zhiyong","family":"Qiu","sequence":"first","affiliation":[{"name":"State Key Laboratory of High-end Server and Storage Technology, Shandong massive information technology research institute, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1303-6681","authenticated-orcid":false,"given":"Zhenhua","family":"Guo","sequence":"additional","affiliation":[{"name":"State Key Laboratory of High-end Server and Storage Technology, Shandong massive information technology research institute, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3975-104X","authenticated-orcid":false,"given":"Li","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of High-end Server and Storage Technology, Shandong massive information technology research institute, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9170-0090","authenticated-orcid":false,"given":"Yaqian","family":"Zhao","sequence":"additional","affiliation":[{"name":"State Key Laboratory of High-end Server and Storage Technology, Inspur Electronic Information Industry Co., Ltd., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4297-4335","authenticated-orcid":false,"given":"Rengang","family":"Li","sequence":"additional","affiliation":[{"name":"State Key Laboratory of High-end Server and Storage Technology, Inspur Electronic Information Industry Co., Ltd., China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,1,17]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Soviet Mathematics Doklady","volume":"27","author":"Nesterov Y.","year":"1983","unstructured":"[1] Nesterov , Y. A method of solving a convex programming problem with convergence rate o(1\/k2) . In Soviet Mathematics Doklady , vol. 27 , 1983 . [1]Nesterov, Y. A method of solving a convex programming problem with convergence rate o(1\/k2). In Soviet Mathematics Doklady, vol. 27, 1983."},{"key":"e_1_3_2_1_2_1","first-page":"2121","volume":"12","author":"John D.","year":"2011","unstructured":"[ 2 ] John D. , Elad H. , and Yoram S .. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization. Journal of Machine Learning Research , Vol. 12 , pp. 2121 \u2013 2159 , 2011 . [2] John D., Elad H., and Yoram S.. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization. Journal of Machine Learning Research, Vol.12, pp. 2121\u20132159, 2011.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_1_3_1","volume-title":"International Conference on Learning Representations","author":"Kingma D.P.","year":"2015","unstructured":"[ 3 ] Kingma , D.P. , Ba , J.L. Adam : A method for stochastic optimization . In International Conference on Learning Representations , 2015 . [3] Kingma, D.P., Ba, J.L. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015."},{"key":"e_1_3_2_1_4_1","volume-title":"International Conference on Learning Representations","author":"Sashank J.","year":"2018","unstructured":"[ 4 ] Sashank J. Reddi , Satyen Kale, and Sanjiv Kumar . On the convergence of adam and beyond . International Conference on Learning Representations , 2018 . [4] Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar. On the convergence of adam and beyond. International Conference on Learning Representations, 2018."},{"key":"e_1_3_2_1_5_1","volume-title":"F. Decoupled Weight Decay Regularization. 7th International Conference on Learning Representations","author":"Loshchilov I.","year":"2019","unstructured":"[ 5 ] Loshchilov , I. , Hutter , F. Decoupled Weight Decay Regularization. 7th International Conference on Learning Representations , 2019 . [5] Loshchilov, I., Hutter, F. Decoupled Weight Decay Regularization. 7th International Conference on Learning Representations, 2019."},{"key":"e_1_3_2_1_6_1","volume-title":"On the variance of the adaptive learning rate and beyond. arXiv preprint arXiv:1908.03265","author":"Liu H.","year":"2019","unstructured":"[ 6 ] L. Liu , H. Jiang , P. He , et al. On the variance of the adaptive learning rate and beyond. arXiv preprint arXiv:1908.03265 , 2019 . [6] L. Liu, H. Jiang, P. He, et al. On the variance of the adaptive learning rate and beyond. arXiv preprint arXiv:1908.03265, 2019."},{"key":"e_1_3_2_1_7_1","volume-title":"SAdam: A Variant of Adam for Strongly Convex Functions. International Conference on Learning Representations.","author":"Wang G.","year":"2020","unstructured":"[ 7 ] Wang G. , Lu S. , Cheng Q. , SAdam: A Variant of Adam for Strongly Convex Functions. International Conference on Learning Representations. 2020 . [7] Wang G., Lu S., Cheng Q., et al. SAdam: A Variant of Adam for Strongly Convex Functions. International Conference on Learning Representations. 2020."},{"key":"e_1_3_2_1_8_1","volume-title":"You Y. Large Batch Training of Convolutional Networks with Layer-wise Adaptive Rate Scaling. International Conference on Learning Representations","author":"Ginsburg B","year":"2018","unstructured":"[ 8 ] Ginsburg B , Gitman I , You Y. Large Batch Training of Convolutional Networks with Layer-wise Adaptive Rate Scaling. International Conference on Learning Representations , 2018 . [8] Ginsburg B, Gitman I, You Y. Large Batch Training of Convolutional Networks with Layer-wise Adaptive Rate Scaling. International Conference on Learning Representations, 2018."},{"key":"e_1_3_2_1_9_1","volume-title":"Large batch optimization for deep learning: Training bert in 76 minutes. arXiv preprint arXiv:1904.00962","author":"You Y.","year":"2019","unstructured":"[ 9 ] You , Y. , Li , J. , Reddi , S. , Large batch optimization for deep learning: Training bert in 76 minutes. arXiv preprint arXiv:1904.00962 , 2019 . [9] You, Y., Li, J., Reddi, S., et al. Large batch optimization for deep learning: Training bert in 76 minutes. arXiv preprint arXiv:1904.00962, 2019."},{"key":"e_1_3_2_1_10_1","volume-title":"Adabelief optimizer: Adapting stepsizes by the belief in observed gradients. Advances in neural information processing systems","author":"Zhuang J.","year":"2020","unstructured":"[ 10 ] Zhuang J. , Tang T. , Ding Y. , Adabelief optimizer: Adapting stepsizes by the belief in observed gradients. Advances in neural information processing systems , Vol. 33 , 2020 . [10] Zhuang J., Tang T., Ding Y., et al. Adabelief optimizer: Adapting stepsizes by the belief in observed gradients. Advances in neural information processing systems, Vol.33, 2020."},{"key":"e_1_3_2_1_11_1","article-title":"FastAdaBelief: improving convergence rate for belief-based adaptive optimizers by exploiting strong convexity","author":"Zhou Y.","year":"2022","unstructured":"[ 11 ] Zhou Y. , Huang K. , Cheng C. , FastAdaBelief: improving convergence rate for belief-based adaptive optimizers by exploiting strong convexity . IEEE Transactions on Neural Networks and Learning Systems , 2022 . [11] Zhou Y., Huang K., Cheng C., et al. FastAdaBelief: improving convergence rate for belief-based adaptive optimizers by exploiting strong convexity. IEEE Transactions on Neural Networks and Learning Systems, 2022.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_2_1_12_1","first-page":"2417","volume-title":"International conference on machine learning","author":"Martens R.","year":"2015","unstructured":"[ 12 ] J. Martens , R. Grosse . Optimizing neural networks with kronecker-factored approximate curvature . In International conference on machine learning , pp. 2408\u2013 2417 , 2015 . [12] J. Martens, R. Grosse. Optimizing neural networks with kronecker-factored approximate curvature. In International conference on machine learning, pp. 2408\u20132417, 2015."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3004354"},{"key":"e_1_3_2_1_14_1","first-page":"7054","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Chen K.","year":"2021","unstructured":"[ 14 ] M. Chen , K. Gao , X. Liu , et al., THOR , trace-based hardwaredriven layer-oriented natural gradient descent computation . In Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35 , pp. 7046\u2013 7054 , 2021 . [14] M. Chen, K. Gao, X. Liu, et al., THOR, trace-based hardwaredriven layer-oriented natural gradient descent computation. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 7046\u20137054, 2021."}],"event":{"name":"AISS 2022: 2022 4th International Conference on Advanced Information Science and System","acronym":"AISS 2022","location":"Sanya China"},"container-title":["Proceedings of the 4th International Conference on Advanced Information Science and System"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3573834.3574539","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3573834.3574539","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:21Z","timestamp":1750182561000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3573834.3574539"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,11,25]]},"references-count":14,"alternative-id":["10.1145\/3573834.3574539","10.1145\/3573834"],"URL":"https:\/\/doi.org\/10.1145\/3573834.3574539","relation":{},"subject":[],"published":{"date-parts":[[2022,11,25]]},"assertion":[{"value":"2023-01-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}