{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,10]],"date-time":"2026-02-10T17:33:03Z","timestamp":1770744783742,"version":"3.49.0"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"9","license":[{"start":{"date-parts":[[2022,7,13]],"date-time":"2022-07-13T00:00:00Z","timestamp":1657670400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,7,13]],"date-time":"2022-07-13T00:00:00Z","timestamp":1657670400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2022,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Meta-learning can be used to learn a good prior that facilitates quick learning; two popular approaches are MAML and the meta-learner LSTM. These two methods represent important and different approaches in meta-learning. In this work, we study the two and formally show that the meta-learner LSTM subsumes MAML, although MAML, which is in this sense less general, outperforms the other. We suggest the reason for this surprising performance gap is related to second-order gradients. We construct a new algorithm (named TURTLE) to gain more insight into the importance of second-order gradients. TURTLE is simpler than the meta-learner LSTM yet more expressive than MAML and outperforms both techniques at few-shot sine wave regression and 50% of the tested image classification settings (without any additional hyperparameter tuning) and is competitive otherwise, at a computational cost that is comparable to second-order MAML. We find that second-order gradients also significantly increase the accuracy of the meta-learner LSTM. When MAML was introduced, one of its remarkable features was the use of second-order gradients. Subsequent work focused on cheaper first-order approximations. On the basis of our findings, we argue for more attention for second-order gradients.<\/jats:p>","DOI":"10.1007\/s10994-022-06210-y","type":"journal-article","created":{"date-parts":[[2022,7,13]],"date-time":"2022-07-13T23:05:07Z","timestamp":1657753507000},"page":"3227-3244","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Stateless neural meta-learning using second-order gradients"],"prefix":"10.1007","volume":"111","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9215-2973","authenticated-orcid":false,"given":"Mike","family":"Huisman","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aske","family":"Plaat","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jan N.","family":"van Rijn","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,7,13]]},"reference":[{"key":"6210_CR1","unstructured":"Andrychowicz, M., Denil, M., Colmenarejo, S.G., Hoffman, M.W., Pfau, D., Schaul, T., Shillingford, B., & de Freitas, N. (2016). Learning to learn by gradient descent by gradient descent. In Advances in neural information processing systems 29, NIPS\u201916, pp. 3988\u20133996. Curran Associates Inc.."},{"key":"6210_CR2","unstructured":"Baik, S., Choi, M., Choi, J., Kim, H., & Lee, K.M. (2020). Meta-learning with adaptive hyperparameters. In Advances in neural information processing systems 33, NIPS\u201920."},{"key":"6210_CR3","doi-asserted-by":"crossref","unstructured":"Bottou, L. (2004). Stochastic learning. In Advanced lectures on machine learning, pp. 146\u2013168. Springer.","DOI":"10.1007\/978-3-540-28650-9_7"},{"key":"6210_CR4","unstructured":"Chen, W.-Y., Liu, Y.-C., Kira, Z., Wang, Y.-C.F., & Huang, J.-B. (2019). A closer look at few-shot classification. In International conference on learning representations, ICLR\u201919."},{"key":"6210_CR5","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). ImageNet: A large-scale hierarchical image database. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 248\u2013255. IEEE.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"6210_CR6","unstructured":"Finn, C., & Levine, S. (2018). Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm. In International conference on learning representations, ICLR\u201918."},{"key":"6210_CR7","unstructured":"Finn, C., Abbeel, P., & Levine, S. (2017). Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the 34th International conference on machine learning, ICML\u201917, pp. 1126-1135. PMLR."},{"key":"6210_CR8","unstructured":"Finn, C., Xu, K., & Levine, S. (2018). Probabilistic model-agnostic meta-learning. In Advances in neural information processing systems 31, NIPS\u201918, pp. 9516\u20139527."},{"key":"6210_CR9","unstructured":"Grant, E., Finn, C., Levine, S., Darrell, T., & Griffiths, T. (2018). Recasting gradient-based meta-learning as hierarchical bayes. In International conference on learning representations, ICLR\u201918."},{"key":"6210_CR10","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2015). Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, pp. 1026\u20131034.","DOI":"10.1109\/ICCV.2015.123"},{"key":"6210_CR11","doi-asserted-by":"crossref","unstructured":"Hospedales, T., Antoniou, A., Micaelli, P., & Storkey, A. (2020). Meta-learning in neural networks: A survey. arXiv:2004.05439.","DOI":"10.1109\/TPAMI.2021.3079209"},{"key":"6210_CR12","doi-asserted-by":"publisher","unstructured":"Huisman, M., van Rijn, J.N., & Plaat, A. (2021). A survey of deep meta-learning. Artificial Intelligence Review. ISSN 0269-2821. https:\/\/doi.org\/10.1007\/s10462-021-10004-4.","DOI":"10.1007\/s10462-021-10004-4"},{"key":"6210_CR13","unstructured":"Kingma, D.P., & Ba, J.L. (2015). Adam: A method for stochastic gradient descent. In International conference on learning representations, ICLR\u201915."},{"key":"6210_CR14","unstructured":"Krizhevsky, A., Sutskever, I., & Hinton, G.\u00a0E. (2012). ImageNet classification with deep convolutional neural networks. In Advances in neural information processing systems 25, NIPS\u201912, pp. 1097\u20131105."},{"issue":"7553","key":"6210_CR15","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y LeCun","year":"2015","unstructured":"LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436\u2013444.","journal-title":"Nature"},{"key":"6210_CR16","unstructured":"Lee, Y., & Choi, S. (2018). Gradient-based meta-learning with learned layerwise metric and subspace. In Proceedings of the 35th international conference on machine learning, ICML\u201918, pp. 2927\u20132936. PMLR."},{"key":"6210_CR17","doi-asserted-by":"crossref","unstructured":"Lee, K., Maji, S., Ravichandran, A., & Soatto, S. (2019). Meta-learning with differentiable convex optimization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 10657\u201310665.","DOI":"10.1109\/CVPR.2019.01091"},{"key":"6210_CR18","unstructured":"Li, Z., Zhou, F., Chen, F., & Li, H. (2017). Meta-SGD: Learning to learn quickly for few-shot learning. arXiv:1707.09835."},{"key":"6210_CR19","unstructured":"Lu, J., Gong, P., Ye, J., & Zhang, C. (2020). Learning from very few samples: A survey. arXiv:2009.02653."},{"key":"6210_CR20","unstructured":"Metz, L., Maheswaranathan, N., Nixon, J., Freeman, D., & Sohl-Dickstein, J. (2019). Understanding and correcting pathologies in the training of learned optimizers. In Proceedings of the 36th international conference on machine learning, ICML\u201919, pp. 4556\u20134565. PMLR."},{"issue":"7540","key":"6210_CR21","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529\u2013533.","journal-title":"Nature"},{"key":"6210_CR22","unstructured":"Nichol, A., Achiam, J., & Schulman, J. (2018). On first-order meta-learning algorithms. arXiv:1803.02999."},{"key":"6210_CR23","unstructured":"Park, E., & Oliva, J.\u00a0B. (2019). Meta-curvature. In Advances in neural information processing systems 32, NIPS\u201919, pp 3314\u20133324."},{"key":"6210_CR24","unstructured":"Rajeswaran, A., Finn, C., Kakade, S.M., & Levine, S. (2019). Meta-Learning with Implicit Gradients. In Advances in neural information processing systems 32, NIPS\u201919, pp. 113\u2013124."},{"key":"6210_CR25","unstructured":"Ravi, S., & Larochelle, H. (2017). Optimization as a model for few-shot learning. In International conference on learning representations, ICLR\u201917."},{"key":"6210_CR26","unstructured":"Rusu, A.A., Rao, D., Sygnowski, J. Vinyals, O., Pascanu, R., Osindero, S., & Hadsell, R. (2019). Meta-learning with latent embedding optimization. In International conference on learning representations, ICLR\u201919."},{"key":"6210_CR27","doi-asserted-by":"crossref","unstructured":"Schaul, T. (2010). J. Schmidhuber. Metalearning. Scholarpedia, 5(6), 4650.","DOI":"10.4249\/scholarpedia.4650"},{"key":"6210_CR28","unstructured":"Schmidhuber, J. (1987). Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-... hook. Master\u2019s thesis, Technische Universit\u00e4t M\u00fcnchen."},{"issue":"7587","key":"6210_CR29","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., et al. (2016). Mastering the game of go with deep neural networks and tree search. Nature, 529(7587), 484\u2013489.","journal-title":"Nature"},{"key":"6210_CR30","unstructured":"Snell, J., Swersky, K., & Zemel, R. (2017). Prototypical Networks for Few-shot Learning. In Advances in neural information processing systems 30, NIPS\u201917, pp. 4077\u20134087. Curran Associates Inc."},{"key":"6210_CR31","doi-asserted-by":"crossref","unstructured":"Sun, Q., Liu, Y., Chua, T.-S., & Schiele, B. (2019). Meta-transfer learning for few-shot learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp 403\u2013412.","DOI":"10.1109\/CVPR.2019.00049"},{"key":"6210_CR32","unstructured":"Tieleman, T., & Hinton, G. (2017). Divide the gradient by a running average of its recent magnitude. Coursera: Neural networks for machine learning. Technical Report.."},{"key":"6210_CR33","doi-asserted-by":"crossref","unstructured":"Vanschoren, J. (2018). Meta-learning: A Survey. arXiv:1810.03548.","DOI":"10.1007\/978-3-030-05318-5_2"},{"issue":"2","key":"6210_CR34","doi-asserted-by":"publisher","first-page":"77","DOI":"10.1023\/A:1019956318069","volume":"18","author":"R Vilalta","year":"2002","unstructured":"Vilalta, R., & Drissi, Y. (2002). A perspective view and survey of meta-learning. Artificial Intelligence Review, 18(2), 77\u201395.","journal-title":"Artificial Intelligence Review"},{"key":"6210_CR35","unstructured":"Vinyals, O., Blundell, C., Lillicrap, T., Kavukcuoglu, K., & Wierstra, D. (2016). Matching networks for one shot learning. In Advances in neural information processing systems 29, NIPS\u201916, pp. 3637\u20133645."},{"key":"6210_CR36","unstructured":"Wah, C., Branson, S., Welinder, P., Perona, P., & Belongie, S. (2011). The Caltech-UCSD Birds-200-2011 Dataset. Technical Report CNS-TR-2011-001, California Institute of Technology."}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06210-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-022-06210-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06210-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,9,15]],"date-time":"2022-09-15T23:11:38Z","timestamp":1663283498000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-022-06210-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,13]]},"references-count":36,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2022,9]]}},"alternative-id":["6210"],"URL":"https:\/\/doi.org\/10.1007\/s10994-022-06210-y","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,7,13]]},"assertion":[{"value":"8 October 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 May 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 June 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 July 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"All authors certify that they have no affiliations with or involvement in any organization or entity with any financial interest or non-financial interest in the subject matter or materials discussed in this manuscript.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"Not applicable: this research does not involve personal data, and publishing of this manuscript will not result in the disruption of any individual\u2019s privacy.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}]}}