{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T20:29:43Z","timestamp":1785356983544,"version":"3.55.0"},"reference-count":48,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2025,2,5]],"date-time":"2025-02-05T00:00:00Z","timestamp":1738713600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,2,5]],"date-time":"2025-02-05T00:00:00Z","timestamp":1738713600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U22B2042"],"award-info":[{"award-number":["U22B2042"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Bavarian State Ministry for Economic Affairs, Regional Development and Energy (StMWi) for the Lighthouse Initiative KI.FABRIK"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Nat Mach Intell"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Humans can continually accumulate knowledge and develop increasingly complex behaviours and skills throughout their lives, which is a capability known as \u2018lifelong learning\u2019. Although this lifelong learning capability is considered an essential mechanism that makes up general intelligence, recent advancements in artificial intelligence predominantly excel in narrow, specialized domains and generally lack this lifelong learning capability. Here we introduce a robotic lifelong reinforcement learning framework that addresses this gap by developing a knowledge space inspired by the Bayesian non-parametric domain. In addition, we enhance the agent\u2019s semantic understanding of tasks by integrating language embeddings into the framework. Our proposed embodied agent can consistently accumulate knowledge from a continuous stream of one-time feeding tasks. Furthermore, our agent can tackle challenging real-world long-horizon tasks by combining and reapplying its acquired knowledge from the original tasks stream. The proposed framework advances our understanding of the robotic lifelong learning process and may inspire the development of more broadly applicable intelligence.<\/jats:p>","DOI":"10.1038\/s42256-025-00983-2","type":"journal-article","created":{"date-parts":[[2025,2,5]],"date-time":"2025-02-05T05:02:36Z","timestamp":1738731756000},"page":"256-269","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":38,"title":["Preserving and combining knowledge in robotic lifelong reinforcement learning"],"prefix":"10.1038","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-5949-8793","authenticated-orcid":false,"given":"Yuan","family":"Meng","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0896-2517","authenticated-orcid":false,"given":"Zhenshan","family":"Bing","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2556-3072","authenticated-orcid":false,"given":"Xiangtong","family":"Yao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kejia","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kai","family":"Huang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yang","family":"Gao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fuchun","family":"Sun","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alois","family":"Knoll","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,2,5]]},"reference":[{"key":"983_CR1","doi-asserted-by":"publisher","first-page":"1401","DOI":"10.1613\/jair.1.13673","volume":"75","author":"K Khetarpal","year":"2022","unstructured":"Khetarpal, K., Riemer, M., Rish, I. & Precup, D. Towards continual reinforcement learning: a review and perspectives. J. Artif. Intell. Res. 75, 1401\u20131476 (2022).","journal-title":"J. Artif. Intell. Res."},{"key":"983_CR2","doi-asserted-by":"publisher","first-page":"10850","DOI":"10.1109\/TPAMI.2023.3261988","volume":"45","author":"F-A Croitoru","year":"2023","unstructured":"Croitoru, F.-A., Hondru, V., Ionescu, R. T. & Shah, M. Diffusion models in vision: a survey. IEEE Trans. Pattern Anal. Mach. Intell. 45, 10850\u201310869 (2023).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"983_CR3","unstructured":"Achiam, J. et al. GPT-4 technical report. Preprint at https:\/\/arxiv.org\/abs\/2303.08774 (2023)."},{"key":"983_CR4","doi-asserted-by":"publisher","first-page":"122836","DOI":"10.1016\/j.eswa.2023.122836","volume":"242","author":"J Zhao","year":"2024","unstructured":"Zhao, J. et al. Autonomous driving system: a comprehensive survey. Expert Syst. Appl. 242, 122836 (2024).","journal-title":"Expert Syst. Appl."},{"key":"983_CR5","doi-asserted-by":"crossref","unstructured":"Aljundi, R., Babiloni, F., Elhoseiny, M., Rohrbach, M. & Tuytelaars, T. Memory aware synapses: learning what (not) to forget. In Proc. European Conference on Computer Vision (ECCV) 139\u2013154 (2018).","DOI":"10.1007\/978-3-030-01219-9_9"},{"key":"983_CR6","doi-asserted-by":"crossref","unstructured":"Chaudhry, A., Dokania, P. K., Ajanthan, T. & Torr, P. H. Riemannian walk for incremental learning: understanding forgetting and intransigence. In Proc. European Conference on Computer Vision (ECCV) 532\u2013547 (2018).","DOI":"10.1007\/978-3-030-01252-6_33"},{"key":"983_CR7","unstructured":"Chen, Z. & Liu, B. Lifelong Machine Learning Vol. 1 (Springer Nature, 2022)."},{"key":"983_CR8","doi-asserted-by":"crossref","unstructured":"Delange, M. et al. A continual learning survey: sefying forgetting in classification tasks. In IEEE Transactions on Pattern Analysis and Machine Intelligence 1\u20131 (IEEE, 2021).","DOI":"10.1109\/TPAMI.2021.3057446"},{"key":"983_CR9","doi-asserted-by":"publisher","first-page":"3521","DOI":"10.1073\/pnas.1611835114","volume":"114","author":"J Kirkpatrick","year":"2017","unstructured":"Kirkpatrick, J. et al. Overcoming catastrophic forgetting in neural networks. Proc. Natl Acad. Sci. USA 114, 3521\u20133526 (2017).","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"983_CR10","unstructured":"Nguyen, C. V., Li, Y., Bui, T. D. & Turner, R. E. Variational continual learning. In Proc. 6th International Conference on Learning Representations (OpenReview.net, 2018)."},{"key":"983_CR11","doi-asserted-by":"crossref","unstructured":"Mallya, A. & Lazebnik, S. PackNet: adding multiple tasks to a single network by iterative pruning. In Proc. IEEE Conference on Computer Vision and Pattern Recognition 7765\u20137773 (IEEE, 2018).","DOI":"10.1109\/CVPR.2018.00810"},{"key":"983_CR12","unstructured":"Fernando, C. et al. PathNet: evolution channels gradient descent in super neural networks. Preprint at https:\/\/arxiv.org\/abs\/1701.08734 (2017)."},{"key":"983_CR13","doi-asserted-by":"publisher","first-page":"651","DOI":"10.1109\/TPAMI.2018.2884462","volume":"42","author":"A Rosenfeld","year":"2018","unstructured":"Rosenfeld, A. & Tsotsos, J. K. Incremental learning through deep adaptation. IEEE Trans. Pattern Anal. Mach. Intell. 42, 651\u2013663 (2018).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"983_CR14","unstructured":"Chaudhry, A., Ranzato, M., Rohrbach, M. & Elhoseiny, M. Efficient lifelong learning with A-GEM. In Proc. 7th International Conference on Learning Representations (OpenReview.net, 2019)."},{"key":"983_CR15","unstructured":"Lopez-Paz, D. & Ranzato, M. Gradient episodic memory for continual learning. In Advances in neural information processing systems (eds Guyon, I. et al.) Vol. 30 (Curran Associates, Inc., 2017)."},{"key":"983_CR16","doi-asserted-by":"crossref","unstructured":"Rebuffi, S.-A., Kolesnikov, A., Sperl, G. & Lampert, C. H. iCaRL: incremental classifier and representation learning. In Proc. IEEE Conference on Computer Vision and Pattern Recognition 2001\u20132010 (IEEE, 2017).","DOI":"10.1109\/CVPR.2017.587"},{"key":"983_CR17","doi-asserted-by":"publisher","first-page":"54","DOI":"10.1016\/j.neunet.2019.01.012","volume":"113","author":"GI Parisi","year":"2019","unstructured":"Parisi, G. I., Kemker, R., Part, J. L., Kanan, C. & Wermter, S. Continual lifelong learning with neural networks: a review. Neural Netw. 113, 54\u201371 (2019).","journal-title":"Neural Netw."},{"key":"983_CR18","unstructured":"Yu, T. et al. Meta-world: a benchmark and evaluation for multi-task and meta reinforcement learning. In Proc. Conference on Robot Learning 1094\u20131100 (PMLR, 2020)."},{"key":"983_CR19","doi-asserted-by":"publisher","first-page":"1363","DOI":"10.3390\/electronics9091363","volume":"9","author":"N Vithayathil Varghese","year":"2020","unstructured":"Vithayathil Varghese, N. & Mahmoud, Q. H. A survey of multi-task deep reinforcement learning. Electronics 9, 1363 (2020).","journal-title":"Electronics"},{"key":"983_CR20","first-page":"5824","volume":"33","author":"T Yu","year":"2020","unstructured":"Yu, T. et al. Gradient surgery for multi-task learning. Adv. Neural Inf. Process. Syst. 33, 5824\u20135836 (2020).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"983_CR21","first-page":"4767","volume":"33","author":"R Yang","year":"2020","unstructured":"Yang, R., Xu, H., Wu, Y. & Wang, X. Multi-task reinforcement learning with soft modularization. Adv. Neural Inf. Process. Syst. 33, 4767\u20134777 (2020).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"983_CR22","unstructured":"Sodhani, S., Zhang, A. & Pineau, J. Multi-task reinforcement learning with context-based representations. In Proc. 38th International Conference on Machine Learning 9767\u20139779 (PMLR, 2021)."},{"key":"983_CR23","doi-asserted-by":"crossref","unstructured":"Perez, E., Strub, F., De Vries, H., Dumoulin, V. & Courville, A. FiLM: visual reasoning with a general conditioning layer. In Proc. AAAI Conference on Artificial Intelligence 32 (AAAI, 2018).","DOI":"10.1609\/aaai.v32i1.11671"},{"key":"983_CR24","unstructured":"Borsa, D., Graepel, T. & Shawe-Taylor, J. Learning shared representations in multi-task reinforcement learning. Preprint at https:\/\/arxiv.org\/abs\/1603.02041 (2016)."},{"key":"983_CR25","first-page":"3476","volume":"45","author":"Z Bing","year":"2022","unstructured":"Bing, Z., Lerch, D., Huang, K. & Knoll, A. Meta-reinforcement learning in non-stationary and dynamic environments. IEEE Trans. Pattern Anal. Mach. Intell. 45, 3476\u20133491 (2022).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"983_CR26","unstructured":"Rakelly, K., Zhou, A., Finn, C., Levine, S. & Quillen, D. Efficient off-policy meta-reinforcement learning via probabilistic context variables. In Proc. 36th International Conference on Machine Learning, Proc. Machine Learning Research Vol. 97 (eds Chaudhuri, K. & Salakhutdinov, R.) 5331\u20135340 (PMLR, 2019)."},{"key":"983_CR27","unstructured":"Finn, C., Abbeel, P. & Levine, S. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning 1126\u20131135 (PMLR, 2017)."},{"key":"983_CR28","unstructured":"Melo, L. C. Transformers are meta-reinforcement learners In Proc. 39th International Conference on Machine Learning, Proc. Machine Learning Research Vol. 162 (eds Chaudhuriet, K. et al.) 15340\u201315359 (PMLR, 2022)."},{"key":"983_CR29","unstructured":"Nam, T., Sun, S.-H., Pertsch, K., Hwang, S. J. & Lim, J. J. Skill-based meta-reinforcement learning. In Proc. 10th International Conference on Learning Representations (OpenReview.net, 2022)."},{"key":"983_CR30","unstructured":"Hughes, M. C. & Sudderth, E. Memoized online variational inference for dirichlet process mixture models. Adv. Neural Inf. Process. Syst. 26, (2013)."},{"key":"983_CR31","unstructured":"Liu, Y. RoBERTa: a robustly optimized bert pretraining approach. Preprint at https:\/\/arxiv.org\/abs\/1907.11692 (2019)."},{"key":"983_CR32","first-page":"6993","volume":"35","author":"A Chaudhry","year":"2021","unstructured":"Chaudhry, A., Gordo, A., Dokania, P., Torr, P. & Lopez-Paz, D. Using hindsight to anchor past knowledge in continual learning. Proc. AAAI Conf. Artif. Intell. 35, 6993\u20137001 (2021).","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"983_CR33","unstructured":"Rolnick, D., Ahuja, A., Schwarz, J., Lillicrap, T. & Wayne, G. Experience replay for continual learning. In Proc. 33rd International Conference on Neural Information Processing Systems 350\u2013360 (Curran Associates Inc., 2019)."},{"key":"983_CR34","doi-asserted-by":"publisher","first-page":"196","DOI":"10.1038\/s42256-022-00452-0","volume":"4","author":"D Kudithipudi","year":"2022","unstructured":"Kudithipudi, D. et al. Biological underpinnings for lifelong learning machines. Nat. Mach. Intell. 4, 196\u2013210 (2022).","journal-title":"Nat. Mach. Intell."},{"key":"983_CR35","doi-asserted-by":"publisher","first-page":"1028","DOI":"10.1016\/j.tics.2020.09.004","volume":"24","author":"R Hadsell","year":"2020","unstructured":"Hadsell, R., Rao, D., Rusu, A. A. & Pascanu, R. Embracing change: continual learning in deep neural networks. Trends Cogn. Sci. 24, 1028\u20131040 (2020).","journal-title":"Trends Cogn. Sci."},{"key":"983_CR36","doi-asserted-by":"publisher","first-page":"676","DOI":"10.1126\/science.8036517","volume":"265","author":"MA Wilson","year":"1994","unstructured":"Wilson, M. A. & McNaughton, B. L. Reactivation of hippocampal ensemble memories during sleep. Science 265, 676\u2013679 (1994).","journal-title":"Science"},{"key":"983_CR37","doi-asserted-by":"publisher","first-page":"100","DOI":"10.1038\/nn1825","volume":"10","author":"D Ji","year":"2007","unstructured":"Ji, D. & Wilson, M. Coordinated memory replay in the visual cortex and hippocampus during sleep. Nat. Neurosci. 10, 100\u2013107 (2007).","journal-title":"Nat. Neurosci."},{"key":"983_CR38","doi-asserted-by":"publisher","first-page":"698","DOI":"10.1016\/j.conb.2007.11.007","volume":"17","author":"B Rasch","year":"2007","unstructured":"Rasch, B. & Born, J. Maintaining memories by reactivation. Curr. Opin. Neurobiol. 17, 698\u2013703 (2007).","journal-title":"Curr. Opin. Neurobiol."},{"key":"983_CR39","doi-asserted-by":"publisher","first-page":"1264","DOI":"10.1038\/nn.2205","volume":"11","author":"JL Lee","year":"2008","unstructured":"Lee, J. L. Memory reconsolidation mediates the strengthening of memories by additional learning. Nat. Neurosci. 11, 1264\u20131266 (2008).","journal-title":"Nat. Neurosci."},{"key":"983_CR40","doi-asserted-by":"publisher","first-page":"561","DOI":"10.1016\/S0022-5371(72)80039-2","volume":"11","author":"LL Jacoby","year":"1972","unstructured":"Jacoby, L. L. & Bartz, W. H. Rehearsal and transfer to LTM. J. Verbal Learning Verbal Behav. 11, 561\u2013565 (1972).","journal-title":"J. Verbal Learning Verbal Behav."},{"key":"983_CR41","doi-asserted-by":"publisher","first-page":"479","DOI":"10.1016\/S0022-5371(76)90043-8","volume":"15","author":"VJ Dark","year":"1976","unstructured":"Dark, V. J. & Loftus, G. R. The role of rehearsal in long-term memory performance. J. Verbal Learning Verbal Behav. 15, 479\u2013490 (1976).","journal-title":"J. Verbal Learning Verbal Behav."},{"key":"983_CR42","unstructured":"Ma, Y. J. et al. Eureka: human-level reward design via coding large language models. In Proc. 12th International Conference on Learning Representations (OpenReview.net, 2024)."},{"key":"983_CR43","doi-asserted-by":"crossref","unstructured":"Ma, Y. J. et al. DrEureka: language model guided sim-to-real transfer. In Robotics: Science and Systems (RSS) (2024).","DOI":"10.15607\/RSS.2024.XX.094"},{"key":"983_CR44","unstructured":"Yu, W. et al. Language to rewards for robotic skill synthesis. In Proc. 7th Conference on Robot Learning (eds Tan, J. et al) 374\u2013404 (PMLR, 2023)."},{"key":"983_CR45","unstructured":"Haarnoja, T. et al. Soft actor\u2013critic algorithms and applications. Preprint at https:\/\/arxiv.org\/abs\/1812.05905 (2018)."},{"key":"983_CR46","first-page":"28496","volume":"34","author":"M Wo\u0142czyk","year":"2021","unstructured":"Wo\u0142czyk, M., Zaj\u0105c, M., Pascanu, R., Kuci\u0144ski, \u0141. & Mi\u0142o\u015b, P. Continual world: a robotic benchmark for continual reinforcement learning. Adv. Neural Inf. Process. Syst. 34, 28496\u201328510 (2021).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"983_CR47","unstructured":"Higgins, I. et al. Beta-VAE: learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations (Poster) 3, (2017)."},{"key":"983_CR48","doi-asserted-by":"publisher","unstructured":"Meng, Y. et al. LEGION. Zenodo https:\/\/doi.org\/10.5281\/zenodo.14265088 (2024).","DOI":"10.5281\/zenodo.14265088"}],"container-title":["Nature Machine Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s42256-025-00983-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s42256-025-00983-2","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s42256-025-00983-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,23]],"date-time":"2025-02-23T18:03:35Z","timestamp":1740333815000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s42256-025-00983-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,5]]},"references-count":48,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,2]]}},"alternative-id":["983"],"URL":"https:\/\/doi.org\/10.1038\/s42256-025-00983-2","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-4353532\/v1","asserted-by":"object"}]},"ISSN":["2522-5839"],"issn-type":[{"value":"2522-5839","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,5]]},"assertion":[{"value":"1 May 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 January 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 February 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}]}}