{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T15:20:09Z","timestamp":1774279209968,"version":"3.50.1"},"reference-count":62,"publisher":"Springer Science and Business Media LLC","issue":"19","license":[{"start":{"date-parts":[[2025,5,10]],"date-time":"2025-05-10T00:00:00Z","timestamp":1746835200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,5,10]],"date-time":"2025-05-10T00:00:00Z","timestamp":1746835200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2025,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>In this paper, we propose a novel algorithm called adaptive learning &amp; momentum sharpness-aware minimization MAML (AS-MAML) that builds upon the concept of sharpness-aware minimization and adaptive learning and momentum to improve generalization. We prove theoretically by performing convergence analysis and PAC-Bayes analysis that AS-MAML performs better than the state-of-the-art algorithms in model-agnostic meta-learning. We draw the same conclusion through extensive experimental analysis using benchmark datasets.<\/jats:p>","DOI":"10.1007\/s00521-025-11220-7","type":"journal-article","created":{"date-parts":[[2025,5,10]],"date-time":"2025-05-10T17:07:48Z","timestamp":1746896868000},"page":"14399-14426","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Using adaptive learning and momentum to improve generalization"],"prefix":"10.1007","volume":"37","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9280-772X","authenticated-orcid":false,"given":"Usman","family":"Anjum","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chris","family":"Stockman","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cat","family":"Luong","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Felix","family":"Zhan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,5,10]]},"reference":[{"key":"11220_CR1","unstructured":"Finn C, Abbeel P, Levine S (2017) Model-agnostic meta-learning for fast adaptation of deep networks. In: International Conference on Machine Learning, pp. 1126\u20131135. PMLR"},{"issue":"2","key":"11220_CR2","doi-asserted-by":"publisher","first-page":"223","DOI":"10.1137\/16M1080173","volume":"60","author":"L Bottou","year":"2018","unstructured":"Bottou L, Curtis FE, Nocedal J (2018) Optimization methods for large-scale machine learning. SIAM review 60(2):223\u2013311","journal-title":"SIAM review"},{"key":"11220_CR3","unstructured":"Foret P, Kleiner A, Mobahi H, Neyshabur B (2020) Sharpness-aware minimization for efficiently improving generalization. arXiv preprint arXiv:2010.01412"},{"key":"11220_CR4","unstructured":"Abbas M, Xiao Q, Chen L, Chen P-Y, Chen T (2022) Sharp-maml: Sharpness-aware model-agnostic meta learning. In: International Conference on Machine Learning, pp. 10\u201332. PMLR"},{"key":"11220_CR5","doi-asserted-by":"crossref","unstructured":"Wang P, Zhang Z, Lei Z, Zhang L (2023) Sharpness-aware gradient matching for domain generalization. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 3769\u20133778","DOI":"10.1109\/CVPR52729.2023.00367"},{"issue":"7","key":"11220_CR6","first-page":"2121","volume":"12","author":"J Duchi","year":"2011","unstructured":"Duchi J, Hazan E, Singer Y (2011) Adaptive subgradient methods for online learning and stochastic optimization. J Mach Learn Res 12(7):2121\u20132159","journal-title":"J Mach Learn Res"},{"key":"11220_CR7","unstructured":"Kingma DP, Ba J (2014) Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980"},{"key":"11220_CR8","unstructured":"Reddi SJ, Kale S, Kumar S (2019) On the convergence of adam and beyond. arXiv preprint arXiv:1904.09237"},{"key":"11220_CR9","doi-asserted-by":"crossref","unstructured":"Sun H, Shen L, Zhong Q, Ding L, Chen S, Sun J, Li J, Sun G, Tao D (2023) Adasam: boosting sharpness-aware minimization with adaptive learning rate and momentum for training deep neural networks. arXiv preprint arXiv:2303.00565","DOI":"10.1016\/j.neunet.2023.10.044"},{"issue":"5","key":"11220_CR10","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/0041-5553(64)90137-5","volume":"4","author":"BT Polyak","year":"1964","unstructured":"Polyak BT (1964) Some methods of speeding up the convergence of iteration methods. Ussr Comput Math Math Phys 4(5):1\u201317","journal-title":"Ussr Comput Math Math Phys"},{"key":"11220_CR11","volume-title":"Introductory Lectures on Convex Optimization: A Basic Course","author":"Y Nesterov","year":"2003","unstructured":"Nesterov Y (2003) Introductory Lectures on Convex Optimization: A Basic Course, vol 87. Springer"},{"key":"11220_CR12","doi-asserted-by":"crossref","unstructured":"O\u2019donoghue B, Candes E (2015) Adaptive restart for accelerated gradient schemes Foundations of computational mathematics. 15: 715\u2013732","DOI":"10.1007\/s10208-013-9150-3"},{"key":"11220_CR13","first-page":"2173","volume":"34","author":"A Farid","year":"2021","unstructured":"Farid A, Majumdar A (2021) Generalization bounds for meta-learning via pac-bayes and uniform stability. Adv Neural Inf Process Syst 34:2173\u20132186","journal-title":"Adv Neural Inf Process Syst"},{"key":"11220_CR14","unstructured":"Yoon J, Kim T, Dia O, Kim S, Bengio Y, Ahn S (2018) Bayesian model-agnostic meta-learning. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems, pp. 7343\u20137353. NIPS'18"},{"key":"11220_CR15","unstructured":"Finn C, Xu K, Levine S (2018) Probabilistic model-agnostic meta-learning. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems, pp. 9537\u2013954. NIPS\u201918"},{"key":"11220_CR16","unstructured":"Grant E, Finn C, Levine S, Darrell T, Griffiths T (2018) Recasting gradient-based meta-learning as hierarchical bayes. arXiv preprint arXiv:1801.08930"},{"key":"11220_CR17","unstructured":"Chen L, Chen T (2022) Is bayesian model-agnostic meta learning better than model-agnostic meta learning, provably? In: International Conference on Artificial Intelligence and Statistics, pp. 1733\u20131774. PMLR"},{"key":"11220_CR18","unstructured":"Denevi G, Ciliberto C, Stamos D, Pontil M (2018) Learning to learn around a common mean. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems, pp. 10190\u201310200. NIPS\u201918"},{"key":"11220_CR19","unstructured":"Wang Z, Grigsby J, Sekhon A, Qi Y (2022) St-maml: A stochastic-task based method for task-heterogeneous meta-learning. In: Uncertainty in Artificial Intelligence, pp. 2066\u20132074. PMLR"},{"key":"11220_CR20","unstructured":"Wang R, Xu K, Liu S, Chen P-Y, Weng T-W, Gan C, Wang M (2021) On fast adversarial robustness adaptation in model-agnostic meta-learning. arXiv preprint arXiv:2102.10454"},{"key":"11220_CR21","first-page":"17886","volume":"33","author":"M Goldblum","year":"2020","unstructured":"Goldblum M, Fowl L, Goldstein T (2020) Adversarially robust few-shot learning: a meta-learning approach. Adv Neural Inform Process Syst 33:17886\u201317895","journal-title":"Adv Neural Inform Process Syst"},{"key":"11220_CR22","doi-asserted-by":"crossref","unstructured":"Xu H, Li Y, Liu X, Liu H, Tang J (2021) Yet meta learning can adapt fast, it can also break easily. In: Proceedings of the 2021 SIAM International Conference on Data Mining (SDM), pp. 540\u2013548. SIAM","DOI":"10.1137\/1.9781611976700.61"},{"key":"11220_CR23","first-page":"3096","volume":"34","author":"A Fallah","year":"2021","unstructured":"Fallah A, Georgiev K, Mokhtari A, Ozdaglar A (2021) On the convergence theory of debiased model-agnostic meta-reinforcement learning. Adv Neural Inform Process Syst 34:3096\u20133107","journal-title":"Adv Neural Inform Process Syst"},{"key":"11220_CR24","doi-asserted-by":"crossref","unstructured":"Nguyen C, Do T-T, Carneiro G (2020) Uncertainty in model-agnostic meta-learning using variational inference. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, pp. 3090\u20133100","DOI":"10.1109\/WACV45572.2020.9093536"},{"issue":"15","key":"11220_CR25","doi-asserted-by":"publisher","first-page":"8653","DOI":"10.3390\/app13158653","volume":"13","author":"Z Zhang","year":"2023","unstructured":"Zhang Z, Li X, Wang S (2023) Amortized bayesian meta-learning with accelerated gradient descent steps. Appl Sci 13(15):8653","journal-title":"Appl Sci"},{"key":"11220_CR26","unstructured":"Zintgraf L, Shiarli K, Kurin V, Hofmann K, Whiteson S (2019) Fast context adaptation via meta-learning. In: International Conference on Machine Learning, pp. 7693\u20137702. PMLR"},{"key":"11220_CR27","unstructured":"Rajeswaran A, Finn C, Kakade SM, Levine S (2019) Meta-learning with implicit gradients. In: Proceedings of the 33nd International Conference on Neural Information Processing Systems, pp. 113\u2013124. NIPS\u201919"},{"key":"11220_CR28","unstructured":"Raghu A, Raghu M, Bengio S, Vinyals O (2019) Rapid learning or feature reuse? towards understanding the effectiveness of maml. arXiv preprint arXiv:1909.09157"},{"key":"11220_CR29","unstructured":"Vinyals O, Blundell C, Lillicrap T, Wierstra D et al (2016) Matching networks for one shot learning. In: Proceedings of the 30th International Conference on Neural Information Processing Systems, pp. 3637\u20133645. NIPS'16"},{"key":"11220_CR30","unstructured":"Snell J, Swersky K, Zemel R (2017) Prototypical networks for few-shot learning. In: Proceedings of the 31st International Conference on Neural Information Processing Systems, pp. 4080\u2013409. NIPS\u201917"},{"key":"11220_CR31","unstructured":"Nichol A, Schulman J (2018) Reptile: a scalable metalearning algorithm. arXiv preprint arXiv:1803.02999 2(3),4"},{"key":"11220_CR32","doi-asserted-by":"crossref","unstructured":"Anjum U, Stockman C, Luong C, Zhan J (2024) Domain-generalization to improve learning in meta-learning algorithms. Available at SSRN 4992443","DOI":"10.2139\/ssrn.4992443"},{"issue":"12","key":"11220_CR33","doi-asserted-by":"publisher","first-page":"1213","DOI":"10.3390\/diagnostics14121213","volume":"14","author":"AM Alsaleh","year":"2024","unstructured":"Alsaleh AM, Albalawi E, Algosaibi A, Albakheet SS, Khan SB (2024) Few-shot learning for medical image segmentation using 3d u-net and model-agnostic meta-learning (maml). Diagnostics 14(12):1213","journal-title":"Diagnostics"},{"key":"11220_CR34","doi-asserted-by":"crossref","unstructured":"Chen S, Zheng Y, Lin D, Cai P, Xiao Y, Wang S (2025) Maml-kalmannet: A neural network-assisted kalman filter based on model-agnostic meta-learning. IEEE Transactions on Signal Processing","DOI":"10.1109\/TSP.2025.3540018"},{"key":"11220_CR35","unstructured":"Daaboul K, Kuhm F, Joseph T, Zoellner JM (2024) Constrained meta-agnostic reinforcement learning. arXiv arXiv:2406.14047"},{"key":"11220_CR36","unstructured":"Shabbar A (2024) Model-agnostic meta-learning with open-ended reinforcement learning. In: NeurIPS. https:\/\/neurips.cc\/virtual\/2024\/101072"},{"key":"11220_CR37","doi-asserted-by":"crossref","unstructured":"Sinha S, Yue Y, Soto V, Kulkarni M, Lu J, Zhang A (2024) Maml-en-llm: Model agnostic meta-training of llms for improved in-context learning. In: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 2711\u20132720","DOI":"10.1145\/3637528.3671905"},{"issue":"3","key":"11220_CR38","doi-asserted-by":"publisher","first-page":"535","DOI":"10.3390\/electronics13030535","volume":"13","author":"J-H Tak","year":"2024","unstructured":"Tak J-H, Hong B-W (2024) Enhancing model agnostic meta-learning via gradient similarity loss. Electronics 13(3):535","journal-title":"Electronics"},{"key":"11220_CR39","unstructured":"Behdin K, Song Q, Gupta A, Acharya A, Durfee D, Ocejo B, Keerthi S, Mazumder R (2023) msam: Micro-batch-averaged sharpness-aware minimization. arXiv preprint arXiv:2302.09693"},{"key":"11220_CR40","unstructured":"Ni R, Chiang P-y, Geiping J, Goldblum M, Wilson AG, Goldstein T (2022) K-sam: sharpness-aware minimization at the speed of sgd. arXiv preprint arXiv:2210.12864"},{"key":"11220_CR41","unstructured":"Du J, Yan H, Feng J, Zhou JT, Zhen L, Goh RSM, Tan VY (2021) Efficient sharpness-aware minimization for improved training of neural networks. arXiv preprint arXiv:2110.03141"},{"key":"11220_CR42","unstructured":"Zhuang J, Gong B, Yuan L, Cui Y, Adam H, Dvornek N, Tatikonda S, Duncan J, Liu T (2022) Surrogate gap minimization improves sharpness-aware training. arXiv preprint arXiv:2203.08065"},{"key":"11220_CR43","first-page":"16577","volume":"35","author":"J Kaddour","year":"2022","unstructured":"Kaddour J, Liu L, Silva R, Kusner MJ (2022) When do flat minima optimizers work? Adv Neural Inform Process Syst 35:16577\u201316595","journal-title":"Adv Neural Inform Process Syst"},{"key":"11220_CR44","unstructured":"Kwon J, Kim J, Park H, Choi IK (2021) Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks. In: International Conference on Machine Learning, pp. 5905\u20135914. PMLR"},{"issue":"179","key":"11220_CR45","first-page":"1","volume":"25","author":"PM Long","year":"2024","unstructured":"Long PM, Bartlett PL (2024) Sharpness-aware minimization and the edge of stability. J Mach Learn Res 25(179):1\u201320","journal-title":"J Mach Learn Res"},{"key":"11220_CR46","doi-asserted-by":"crossref","unstructured":"Li T, Zhou P, He Z, Cheng X, Huang X (2024) Friendly sharpness-aware minimization. In: CVPR. https:\/\/openaccess.thecvf.com\/content\/CVPR2024\/html\/Li_Friendly_Sharpness-Aware_Minimization_CVPR_2024_paper.html","DOI":"10.1109\/CVPR52733.2024.00538"},{"key":"11220_CR47","doi-asserted-by":"crossref","unstructured":"Wu T, Luo T, Wunsch\u00a0II DC (2024) Cr-sam: Curvature regularized sharpness-aware minimization. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, pp. 6144\u20136152","DOI":"10.1609\/aaai.v38i6.28431"},{"key":"11220_CR48","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4615-5529-2","volume-title":"Learning to Learn","author":"S Thrun","year":"1998","unstructured":"Thrun S, Pratt L (1998) Learning to Learn. Springer"},{"key":"11220_CR49","doi-asserted-by":"publisher","first-page":"131","DOI":"10.1162\/neco.1992.4.1.131","volume":"4","author":"J Schmidhuber","year":"1992","unstructured":"Schmidhuber J (1992) Learning to control fast-weight memories: an alternative to dynamic recurrent networks. Neural Comput 4(1):131\u2013139","journal-title":"Neural Comput"},{"key":"11220_CR50","unstructured":"Yao H, Wei Y, Huang J, Li Z (2019) Hierarchically structured meta-learning. In: International conference on machine learning, pp. 7045\u20137054. PMLR'19"},{"key":"11220_CR51","unstructured":"Lake B, Salakhutdinov R, Gross J, Tenenbaum J (2011) One shot learning of simple visual concepts. In: Proceedings of the Annual Meeting of the Cognitive Science Society, vol. 33. https:\/\/escholarship.org\/uc\/item\/4ht821jx"},{"key":"11220_CR52","unstructured":"Ravi S, Larochelle H (2016) Optimization as a model for few-shot learning. In: 5th International Conference on Learning Representations, pp. 2936\u20132946. ICLR'17"},{"key":"11220_CR53","unstructured":"Sun S-H (2019) Multi-digit MNIST for Few-shot Learning. https:\/\/github.com\/shaohua0116\/MultiDigitMNIST"},{"key":"11220_CR54","unstructured":"Deleu T, W\u00fcrfl T, Samiei M, Cohen JP, Bengio Y (2019) Torchmeta: A Meta-Learning library for PyTorch. Available at: https:\/\/github.com\/tristandeleu\/pytorch-meta. https:\/\/arxiv.org\/abs\/1909.06576"},{"key":"11220_CR55","unstructured":"Nichol A, Achiam J, Schulman J (2018) On first-order meta-learning algorithms. arXiv preprint arXiv:1803.02999"},{"key":"11220_CR56","unstructured":"Lazreg MB, Anjum U, Zadorozhny V, Goodwin M (2020) Semantic decay filter for event detection. In: Hughes, A.L., McNeill, F., Zobel, C.W. (eds.) 17th International Conference on Information Systems for Crisis Response and Management, ISCRAM 2020, May 2020, pp. 14\u201326. ISCRAM Digital Library. https:\/\/idl.iscram.org\/show.php?record=2203"},{"key":"11220_CR57","unstructured":"Anjum U, Zadorozhny V, Krishnamurthy P (2021) TBAM: towards an agent-based model to enrich twitter data. In: Adrot, A., Grace, R., Moore, K.A., Zobel, C.W. (eds.) 18th International Conference on Information Systems for Crisis Response and Management, ISCRAM 2021, Blacksburg, VA, USA, May 2021, pp. 146\u2013158. ISCRAM Digital Library. https:\/\/idl.iscram.org\/show.php?record=2321"},{"key":"11220_CR58","doi-asserted-by":"publisher","first-page":"100209","DOI":"10.1016\/j.osnem.2022.100209","volume":"29","author":"U Anjum","year":"2022","unstructured":"Anjum U, Zadorozhny V, Krishnamurthy P (2022) Localization of unidentified events with raw microblogging data. Online Soc Netw Med 29:100209","journal-title":"Online Soc Netw Med"},{"key":"11220_CR59","unstructured":"Nkhata G, Anjum U, Zhan J (2023) Sentiment analysis of movie reviews using bert. In: The Fifteenth International Conference on Information, Process, and Knowledge Management, eKNOW23"},{"key":"11220_CR60","doi-asserted-by":"publisher","unstructured":"Guo X, Anjum U, Zhan J (2022) Cyberbully detection using BERT with augmented texts. In: Tsumoto, S., Ohsawa, Y., Chen, L., Poel, D.V., Hu, X., Motomura, Y., Takagi, T., Wu, L., Xie, Y., Abe, A., Raghavan, V. (eds.) IEEE International Conference on Big Data, Big Data 2022, Osaka, Japan, December 17-20, 2022, pp. 1246\u20131253. IEEE. https:\/\/doi.org\/10.1109\/BIGDATA55660.2022.10020581 . https:\/\/doi.org\/10.1109\/BigData55660.2022.10020581","DOI":"10.1109\/BIGDATA55660.2022.10020581"},{"issue":"19","key":"11220_CR61","doi-asserted-by":"publisher","first-page":"3825","DOI":"10.3390\/electronics13193825","volume":"13","author":"U Anjum","year":"2024","unstructured":"Anjum U, Zhan J (2024) A novel tsetlin machine with enhanced generalization. Electronics 13(19):3825","journal-title":"Electronics"},{"key":"11220_CR62","doi-asserted-by":"crossref","unstructured":"Wu Z, Li J, Cai M, Lin Y, Zhang W (2016) On membership of black-box or white-box of artificial neural network models. In: 2016 IEEE 11th Conference on Industrial Electronics and Applications (ICIEA), pp. 1400\u20131404. IEEE","DOI":"10.1109\/ICIEA.2016.7603804"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-025-11220-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-025-11220-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-025-11220-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,27]],"date-time":"2025-06-27T08:30:18Z","timestamp":1751013018000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-025-11220-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,10]]},"references-count":62,"journal-issue":{"issue":"19","published-print":{"date-parts":[[2025,7]]}},"alternative-id":["11220"],"URL":"https:\/\/doi.org\/10.1007\/s00521-025-11220-7","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,10]]},"assertion":[{"value":"9 October 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 March 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 May 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"There are no conflicts of interest or conflicts.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}