{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T01:52:45Z","timestamp":1761357165692,"version":"build-2065373602"},"reference-count":46,"publisher":"Springer Science and Business Media LLC","issue":"11","license":[{"start":{"date-parts":[[2025,8,30]],"date-time":"2025-08-30T00:00:00Z","timestamp":1756512000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0"},{"start":{"date-parts":[[2025,8,30]],"date-time":"2025-08-30T00:00:00Z","timestamp":1756512000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61771001"],"award-info":[{"award-number":["61771001"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"DOI":"10.1007\/s10462-025-11362-z","type":"journal-article","created":{"date-parts":[[2025,8,30]],"date-time":"2025-08-30T04:36:09Z","timestamp":1756528569000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["DAoG: decayed adaptation over gradients for parameter-free step size control"],"prefix":"10.1007","volume":"58","author":[{"given":"Yifan","family":"Zhang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Di","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongyi","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengwei","family":"Pan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,8,30]]},"reference":[{"key":"11362_CR1","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A et al (2020) Language models are few-shot learners. Adv Neural Inf Process Syst 33:1877\u20131901","journal-title":"Adv Neural Inf Process Syst"},{"key":"11362_CR2","unstructured":"Carmon Y, Hinder O (2022) Making sgd parameter-free. In: Conference on learning theory, vol. 178, pp. 2360\u20132389 PMLR"},{"key":"11362_CR4","unstructured":"Cohen J, Kaur S, Li Y, Kolter JZ, Talwalkar A (2021) Gradient descent on neural networks typically occurs at the edge of stability. In: International conference on learning representations"},{"key":"11362_CR3","unstructured":"Cohen J, Ghorbani B, Krishnan S, Agarwal N, Medapati S, Badura M, Suo D, Nado Z, Dahl GE, Gilmer J (2023) Adaptive gradient methods at the edge of stability. In: NeurIPS 2023 workshop heavy tails in machine learning"},{"key":"11362_CR5","unstructured":"Defazio A, Mishchenko K (2022) Parameter free dual averaging: Optimizing lipschitz functions in a single pass. In: OPT 2022: optimization for machine learning (NeurIPS 2022 Workshop), pp. 1\u201316"},{"key":"11362_CR6","unstructured":"Defazio A, Mishchenko K (2023) Learning-rate-free learning by d-adaptation. In: International conference on machine learning, vol. 202, pp. 7449\u20137479 PMLR"},{"key":"11362_CR7","unstructured":"Devlin J, Chang M-W, Lee K, Toutanova K (2019) Bert: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, Volume 1 (long and Short Papers), pp. 4171\u20134186"},{"key":"11362_CR8","unstructured":"Gower RM, Loizou N, Qian X, Sailanbayev A, Shulgin E, Richt\u00e1rik P (2019) Sgd: General analysis and improved rates. In: International conference on machine learning, vol. 97, pp. 5200\u20135209 PMLR"},{"key":"11362_CR9","unstructured":"Henaff O (2020) Data-efficient image recognition with contrastive predictive coding. In: International conference on machine learning, 119, 4182\u20134192 PMLR"},{"key":"11362_CR10","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"issue":"3","key":"11362_CR11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s10915-025-02798-0","volume":"102","author":"X Hu","year":"2025","unstructured":"Hu X, Xiao N, Liu X, Toh K-C (2025) Learning-rate-free momentum sgd with reshuffling converges in nonsmooth nonconvex optimization. J Sci Comput 102(3):1\u201331","journal-title":"J Sci Comput"},{"key":"11362_CR12","doi-asserted-by":"crossref","unstructured":"Huang G, Liu Z, Van Der\u00a0Maaten L, Weinberger KQ (2017) Densely connected convolutional networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4700\u20134708","DOI":"10.1109\/CVPR.2017.243"},{"issue":"5","key":"11362_CR13","doi-asserted-by":"publisher","first-page":"6311","DOI":"10.1007\/s11063-022-11140-w","volume":"55","author":"G Ioannou","year":"2023","unstructured":"Ioannou G, Tagaris T, Stafylopatis A (2023) Adalip: An adaptive learning rate method per layer for stochastic optimization. Neural Process Lett 55(5):6311\u20136338","journal-title":"Neural Process Lett"},{"key":"11362_CR14","unstructured":"Ivgi M, Hinder O, Carmon Y (2023) Dog is sgd\u2019s best friend: A parameter-free dynamic step size schedule. In: International conference on machine learning, vol. 202, pp. 14465\u201314499 PMLR"},{"key":"11362_CR15","unstructured":"Jacobsen A, Cutkosky A (2022) Parameter-free mirror descent. In: Conference on learning theory, vol. 178, pp. 4160\u20134211 PMLR"},{"issue":"3","key":"11362_CR16","doi-asserted-by":"publisher","first-page":"2361","DOI":"10.1007\/s10489-024-05303-6","volume":"54","author":"W Jiang","year":"2024","unstructured":"Jiang W, Liang Y, Jiang Z, Xu D, Zhou L (2024) Abngrad: adaptive step size gradient descent for optimizing neural networks. Appl Intell 54(3):2361\u20132378","journal-title":"Appl Intell"},{"key":"11362_CR17","doi-asserted-by":"crossref","unstructured":"Johnson R, Zhang T (2017) Deep pyramid convolutional neural networks for text categorization. In: Proceedings of the 55th annual meeting of the association for computational linguistics (Volume 1: Long Papers), pp. 562\u2013570","DOI":"10.18653\/v1\/P17-1052"},{"key":"11362_CR18","first-page":"6748","volume":"36","author":"A Khaled","year":"2023","unstructured":"Khaled A, Mishchenko K, Jin C (2023) Dowg unleashed: an efficient universal parameter-free gradient descent method. Adv Neural Inf Process Syst 36:6748\u20136769","journal-title":"Adv Neural Inf Process Syst"},{"issue":"3","key":"11362_CR19","doi-asserted-by":"publisher","first-page":"3713","DOI":"10.1007\/s11042-022-13428-4","volume":"82","author":"D Khurana","year":"2023","unstructured":"Khurana D, Koli A, Khatter K, Singh S (2023) Natural language processing: state of the art, current trends and challenges. Multimedia Tools Appl 82(3):3713\u20133744","journal-title":"Multimedia Tools Appl"},{"key":"11362_CR20","unstructured":"Kingma DP, Ba J (2014) Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980, pp 1\u201315"},{"key":"11362_CR21","unstructured":"Krizhevsky A, Hinton G, et al. (2009) Learning multiple layers of features from tiny images, 1\u201360"},{"key":"11362_CR22","doi-asserted-by":"publisher","first-page":"443","DOI":"10.1016\/j.neucom.2021.05.103","volume":"470","author":"I Lauriola","year":"2022","unstructured":"Lauriola I, Lavelli A, Aiolli F (2022) An introduction to deep learning in natural language processing: models, techniques, and tools. Neurocomputing 470:443\u2013456","journal-title":"Neurocomputing"},{"issue":"11","key":"11362_CR23","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","volume":"86","author":"Y LeCun","year":"1998","unstructured":"LeCun Y, Bottou L, Bengio Y, Haffner P (1998) Gradient-based learning applied to document recognition. Proc IEEE 86(11):2278\u20132324","journal-title":"Proc IEEE"},{"issue":"3","key":"11362_CR24","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3506695","volume":"31","author":"L Liao","year":"2022","unstructured":"Liao L, Li H, Shang W, Ma L (2022) An empirical study of the impact of hyperparameter tuning and model optimization on the performance properties of deep neural networks. ACM Trans Softw Eng Methodol (TOSEM) 31(3):1\u201340","journal-title":"ACM Trans Softw Eng Methodol (TOSEM)"},{"key":"11362_CR25","doi-asserted-by":"crossref","unstructured":"Li Z, Tang H, Peng Z, Qi G-J, Tang J (2023) Knowledge-guided semantic transfer network for few-shot image recognition. IEEE Trans Neural Netw Learn Syst, 1\u201315","DOI":"10.1109\/TNNLS.2023.3240195"},{"issue":"15","key":"11362_CR26","doi-asserted-by":"publisher","first-page":"17691","DOI":"10.1007\/s11227-023-05338-5","volume":"79","author":"Y Liu","year":"2023","unstructured":"Liu Y, Li D (2023) Adaxod: a new adaptive and momental bound algorithm for training deep neural networks. J Supercomput 79(15):17691\u201317715","journal-title":"J Supercomput"},{"key":"11362_CR27","unstructured":"Loizou N, Vaswani S, Laradji IH, Lacoste-Julien S (2021) Stochastic polyak step-size for sgd: an adaptive learning rate for fast convergence. In: International conference on artificial intelligence and statistics, vol. 130, pp. 1306\u20131314 PMLR"},{"key":"11362_CR28","unstructured":"Maas A, Daly RE, Pham PT, Huang D, Ng AY, Potts C (2011) Learning word vectors for sentiment analysis. In: Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pp. 142\u2013150"},{"key":"11362_CR29","first-page":"100670","volume":"37","author":"Y Malitsky","year":"2024","unstructured":"Malitsky Y, Mishchenko K (2024) Adaptive proximal gradient method for convex optimization. Adv Neural Inf Process Syst 37:100670\u2013100697","journal-title":"Adv Neural Inf Process Syst"},{"key":"11362_CR30","unstructured":"Mishchenko K, Defazio A (2024) Prodigy: an expeditiously adaptive parameter-free learner. In: International conference on machine learning, vol. 235, pp. 35779\u201335804 PMLR"},{"key":"11362_CR31","doi-asserted-by":"crossref","unstructured":"Nilsback M-E, Zisserman A (2008) Automated flower classification over a large number of classes. In: 2008 Sixth Indian conference on computer vision, graphics & image processing. IEEE, pp. 722\u2013729","DOI":"10.1109\/ICVGIP.2008.47"},{"issue":"12","key":"11362_CR32","doi-asserted-by":"publisher","first-page":"13911","DOI":"10.1007\/s11227-021-03838-w","volume":"77","author":"I Priyadarshini","year":"2021","unstructured":"Priyadarshini I, Cotton C (2021) A novel lstm-cnn-grid search-based deep neural network for sentiment analysis. J Supercomput 77(12):13911\u201313932","journal-title":"J Supercomput"},{"key":"11362_CR33","first-page":"1","volume":"2021","author":"S Pyysalo","year":"2021","unstructured":"Pyysalo S, Kanerva J, Virtanen A, Ginter F (2021) Wikibert models: Deep transfer learning for many languages. NoDaLiDa 2021:1\u201310","journal-title":"NoDaLiDa"},{"key":"11362_CR34","unstructured":"Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 1\u201314"},{"key":"11362_CR35","doi-asserted-by":"crossref","unstructured":"Socher R, Perelygin A, Wu J, Chuang J, Manning CD, Ng AY, Potts C (2013) Recursive deep models for semantic compositionality over a sentiment treebank. In: Proceedings of the 2013 conference on empirical methods in natural language processing, pp. 1631\u20131642","DOI":"10.18653\/v1\/D13-1170"},{"key":"11362_CR36","doi-asserted-by":"crossref","unstructured":"Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z (2016) Rethinking the inception architecture for computer vision. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818\u20132826","DOI":"10.1109\/CVPR.2016.308"},{"key":"11362_CR37","first-page":"44575","volume":"37","author":"Y Takezawa","year":"2024","unstructured":"Takezawa Y, Bao H, Sato R, Niwa K, Yamada M (2024) Parameter-free clipped gradient descent meets polyak. Adv Neural Inf Process Syst 37:44575\u201344599","journal-title":"Adv Neural Inf Process Syst"},{"key":"11362_CR38","unstructured":"Tao Y, Yuan H, Zhou X, Cao Y, Gu Q (2024) Towards simple and provable parameter-free adaptive gradient methods. arXiv preprint arXiv:2412.19444, 1\u201334"},{"key":"11362_CR39","doi-asserted-by":"crossref","unstructured":"Wang X (2024a) Gyro fireworks algorithm: a new metaheuristic algorithm. AIP Ad 14(8)","DOI":"10.1063\/5.0213886"},{"issue":"12","key":"11362_CR40","doi-asserted-by":"publisher","DOI":"10.1088\/1402-4896\/ad8e0e","volume":"99","author":"X Wang","year":"2024","unstructured":"Wang X (2024b) Frigatebird optimizer: a novel metaheuristic algorithm. Phys Scr 99(12):125233","journal-title":"Phys Scr"},{"issue":"2","key":"11362_CR41","doi-asserted-by":"publisher","first-page":"780","DOI":"10.1108\/EC-10-2024-0904","volume":"42","author":"X Wang","year":"2025","unstructured":"Wang X (2025) Fishing cat optimizer: a novel metaheuristic technique. Eng Comput 42(2):780\u2013833","journal-title":"Eng Comput"},{"key":"11362_CR42","doi-asserted-by":"crossref","unstructured":"Wang A, Singh A, Michael J, Hill F, Levy O, Bowman SR (2019) Glue: A multi-task benchmark and analysis platform for natural language understanding. In: International conference on learning representations, pp. 1\u201320","DOI":"10.18653\/v1\/W18-5446"},{"key":"11362_CR43","first-page":"1","volume":"30","author":"AC Wilson","year":"2017","unstructured":"Wilson AC, Roelofs R, Stern M, Srebro N, Recht B (2017) The marginal value of adaptive gradient methods in machine learning. Adv Neural Inf Process Syst 30:1\u201311","journal-title":"Adv Neural Inf Process Syst"},{"key":"11362_CR44","doi-asserted-by":"publisher","first-page":"51522","DOI":"10.1109\/ACCESS.2019.2909919","volume":"7","author":"G Xu","year":"2019","unstructured":"Xu G, Meng Y, Qiu X, Yu Z, Wu X (2019) Sentiment analysis of comment texts based on bilstm. IEEE Access 7:51522\u201351532","journal-title":"IEEE Access"},{"issue":"4","key":"11362_CR45","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s10489-021-02224-6","volume":"52","author":"W Yuan","year":"2022","unstructured":"Yuan W, Hu F, Lu L (2022) A new non-adaptive optimization method: stochastic gradient descent with momentum and difference. Appl Intell 52(4):1\u201315","journal-title":"Appl Intell"},{"key":"11362_CR46","first-page":"18795","volume":"33","author":"J Zhuang","year":"2020","unstructured":"Zhuang J, Tang T, Ding Y, Tatikonda SC, Dvornek N, Papademetris X, Duncan J (2020) Adabelief optimizer: adapting stepsizes by the belief in observed gradients. Adv Neural Inf Process Syst 33:18795\u201318806","journal-title":"Adv Neural Inf Process Syst"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11362-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-025-11362-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11362-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T01:48:16Z","timestamp":1761356896000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-025-11362-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,30]]},"references-count":46,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2025,11]]}},"alternative-id":["11362"],"URL":"https:\/\/doi.org\/10.1007\/s10462-025-11362-z","relation":{},"ISSN":["1573-7462"],"issn-type":[{"type":"electronic","value":"1573-7462"}],"subject":[],"published":{"date-parts":[[2025,8,30]]},"assertion":[{"value":"14 August 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 August 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that there is no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"The data used in this study were obtained from publicly available datasets. The dataset is anonymized and does not contain any personally identifiable information.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"Informed consent was obtained from all individual participants included in the study.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Informed consent"}}],"article-number":"364"}}