{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T14:40:37Z","timestamp":1777560037986,"version":"3.51.4"},"reference-count":68,"publisher":"SAGE Publications","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["AIC"],"published-print":{"date-parts":[[2024,3,21]]},"abstract":"<jats:p>A long-standing challenge in artificial intelligence is lifelong reinforcement learning, where learners are given many tasks in sequence and must transfer knowledge between tasks while avoiding catastrophic forgetting. Policy reuse and other multi-policy reinforcement learning techniques can learn multiple tasks but may generate many policies. This paper presents two novel contributions, namely 1) Lifetime Policy Reuse, a model-agnostic policy reuse algorithm that avoids generating many policies by optimising a fixed number of near-optimal policies through a combination of policy optimisation and adaptive policy selection; and 2) the task capacity, a measure for the maximal number of tasks that a policy can accurately solve. Comparing two state-of-the-art base-learners, the results demonstrate the importance of Lifetime Policy Reuse and task capacity based pre-selection on an 18-task partially observable Pacman domain and a Cartpole domain of up to 125 tasks.<\/jats:p>","DOI":"10.3233\/aic-230040","type":"journal-article","created":{"date-parts":[[2023,10,24]],"date-time":"2023-10-24T11:30:19Z","timestamp":1698147019000},"page":"115-148","source":"Crossref","is-referenced-by-count":0,"title":["Lifetime policy reuse and the importance of task capacity"],"prefix":"10.1177","volume":"37","author":[{"given":"David M.","family":"Bossens","sequence":"first","affiliation":[{"name":"Maritime Engineering group, University of Southampton, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adam J.","family":"Sobey","sequence":"additional","affiliation":[{"name":"Maritime Engineering group, University of Southampton, United Kingdom"},{"name":"Marine and Maritime Group, The Alan Turing Institute, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"key":"10.3233\/AIC-230040_ref1","unstructured":"D.\u00a0Abel, Y.\u00a0Jinna, Y.\u00a0Guo, G.\u00a0Konidaris and M.L.\u00a0Littman, Policy and value transfer in lifelong reinforcement learning, in: Proceedings of the International Conference on Machine Learning (ICML 2018), Stockholm, Sweden, 2018, pp.\u00a01\u201310."},{"key":"10.3233\/AIC-230040_ref2","unstructured":"M.\u00a0Andrychowicz, M.\u00a0Denil, S.G.\u00a0Colmenarejo and M.W.\u00a0Hoffman, Learning to learn by gradient descent by gradient descent, in: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS 2016), 2016, pp.\u00a01\u201317."},{"issue":"2\u20133","key":"10.3233\/AIC-230040_ref3","doi-asserted-by":"publisher","first-page":"235","DOI":"10.1023\/A:1013689704352","article-title":"Finite-time analysis of the multiarmed bandit problem","volume":"47","author":"Auer","year":"2002","journal-title":"Machine Learning"},{"key":"10.3233\/AIC-230040_ref4","doi-asserted-by":"publisher","first-page":"288","DOI":"10.1016\/j.neunet.2019.04.009","article-title":"The capacity of feedforward neural networks","volume":"116","author":"Baldi","year":"2019","journal-title":"Neural Networks"},{"key":"10.3233\/AIC-230040_ref5","unstructured":"J.\u00a0Bieger, K.R.\u00a0Thorisson, B.R.\u00a0Steunebrink, T.\u00a0Thorarensen and J.S.\u00a0Sigurdardottir, Evaluation of general-purpose artificial intelligence: Why, what & how, in: Evaluating General-Purpose A.I. Workshop in the European Conference on Artificial Intelligence (ECAI 2016), The Hague, The Netherlands, 2016."},{"key":"10.3233\/AIC-230040_ref6","doi-asserted-by":"publisher","first-page":"30","DOI":"10.1016\/j.neunet.2019.03.006","article-title":"Learning to learn with active adaptive perception","volume":"115","author":"Bossens","year":"2019","journal-title":"Neural Networks"},{"key":"10.3233\/AIC-230040_ref8","unstructured":"E.\u00a0Brunskill and L.\u00a0Li, PAC-inspired option discovery in lifelong reinforcement learning, in: Proceedings of the International Conference on Machine Learning (ICML 2014), Vol.\u00a032, JMLR: W{&}CP, Beijing, China, 2014, pp.\u00a0316\u2013324."},{"key":"10.3233\/AIC-230040_ref9","unstructured":"Y.\u00a0Burda, A.\u00a0Storkey, T.\u00a0Darrell and A.A.\u00a0Efros, Large-scale study of curiosity-driven learning, in: Proceedings of the International Conference on Learning Representations (ICLR 2019), 2019, pp.\u00a01\u201317."},{"key":"10.3233\/AIC-230040_ref10","doi-asserted-by":"crossref","unstructured":"D.S.\u00a0Chaplot and G.\u00a0Lample, Arnold: An autonomous agent to play FPS games, in: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2017), 2017, pp.\u00a02\u20133.","DOI":"10.1609\/aaai.v31i1.10534"},{"key":"10.3233\/AIC-230040_ref11","doi-asserted-by":"crossref","unstructured":"Z.\u00a0Chen and B.\u00a0Liu, Lifelong Machine Learning, Morgan & Claypool Publishers, 2016.","DOI":"10.1007\/978-3-031-01575-5"},{"key":"10.3233\/AIC-230040_ref12","unstructured":"W.C.\u00a0Cheung, D.\u00a0Simchi-Levi and R.\u00a0Zhu, Reinforcement learning for non-stationary Markov decision processes: The blessing of (more) optimism, in: Proceedings of the International Conference on Machine Learning (ICML 2020), 2020."},{"key":"10.3233\/AIC-230040_ref13","unstructured":"N.\u00a0Cohen, O.\u00a0Sharir, R.\u00a0Tamari and A.\u00a0Shashua, Analysis and design of convolutional networks, in: Why & when Deep Learning Works \u2013 Looking Inside Deep Learning, ICRI-CI paper bundle, Intel Collaborative Research Institute for Computational Intelligence (ICRI-CI), 2017."},{"issue":"7553","key":"10.3233\/AIC-230040_ref14","doi-asserted-by":"publisher","first-page":"503","DOI":"10.1038\/nature14422","article-title":"Robots that can adapt like animals","volume":"521","author":"Cully","year":"2015","journal-title":"Nature"},{"key":"10.3233\/AIC-230040_ref15","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2014.6907421"},{"key":"10.3233\/AIC-230040_ref16","doi-asserted-by":"publisher","DOI":"10.1145\/1160633.1160762"},{"key":"10.3233\/AIC-230040_ref17","unstructured":"C.\u00a0Finn, P.\u00a0Abbeel and S.\u00a0Levine, Model-agnostic meta-learning for fast adaptation of deep networks, in: Proceedings of the International Conference on Machine Learning (ICML 2017), Sydney, Australia, 2017."},{"key":"10.3233\/AIC-230040_ref18","doi-asserted-by":"publisher","first-page":"1","DOI":"10.3389\/fncom.2016.00144","article-title":"On the maximum storage capacity of the Hopfield model","volume":"10","author":"Folli","year":"2017","journal-title":"Frontiers in Computational Neuroscience"},{"issue":"3\u20134","key":"10.3233\/AIC-230040_ref19","doi-asserted-by":"publisher","first-page":"365","DOI":"10.1080\/09540099208946624","article-title":"Semi-distributed representations and catastrophic forgetting in connectionist networks","volume":"4","author":"French","year":"1992","journal-title":"Connection Science"},{"key":"10.3233\/AIC-230040_ref21","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-71682-4_5"},{"issue":"6","key":"10.3233\/AIC-230040_ref22","doi-asserted-by":"publisher","first-page":"407","DOI":"10.1016\/j.tics.2017.04.001","article-title":"Avoiding catastrophic forgetting","volume":"21","author":"Hasselmo","year":"2017","journal-title":"Trends in Cognitive Sciences"},{"key":"10.3233\/AIC-230040_ref23","unstructured":"M.\u00a0Hausknecht and P.\u00a0Stone, Deep recurrent Q-learning for partially observable MDPs, in: Proceedings of the AAAI Fall Symposium Series (FSS 2021), 2015, pp.\u00a029\u201337."},{"key":"10.3233\/AIC-230040_ref25","unstructured":"P.\u00a0Hernandez-Leal, B.\u00a0Rosman, M.E.\u00a0Taylor, L.E.\u00a0Sucar and E.M.\u00a0De Cote, A Bayesian approach for learning and tracking switching, non-stationary opponents, in: Proceedings of the International Joint Conference on Autonomous Agents and Multiagent Systems (AAMAS 2016), 2016, pp.\u00a01315\u20131316."},{"issue":"8","key":"10.3233\/AIC-230040_ref26","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Computation"},{"key":"10.3233\/AIC-230040_ref27","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-44668-0"},{"issue":"5","key":"10.3233\/AIC-230040_ref28","doi-asserted-by":"publisher","first-page":"359","DOI":"10.1016\/0893-6080(89)90020-8","article-title":"Multilayer feedforward networks are universal approximators","volume":"2","author":"Hornik","year":"1989","journal-title":"Neural Networks"},{"key":"10.3233\/AIC-230040_ref29","doi-asserted-by":"crossref","unstructured":"D.\u00a0Isele and A.\u00a0Cosgun, Selective experience replay for lifelong learning, in: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2018), 2018, pp.\u00a03302\u20133309.","DOI":"10.1609\/aaai.v32i1.11595"},{"key":"10.3233\/AIC-230040_ref30","doi-asserted-by":"crossref","unstructured":"H.\u00a0Jung, J.\u00a0Ju, M.\u00a0Jung and J.\u00a0Kim, Less-forgetful learning for domain expansion in deep neural networks, in: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u00a018), 2017, pp.\u00a03358\u20133365.","DOI":"10.1609\/aaai.v32i1.11769"},{"key":"10.3233\/AIC-230040_ref31","unstructured":"S.\u00a0Kapturowski, G.\u00a0Ostrovski, J.\u00a0Quan, R.\u00a0Munos and W.\u00a0Dabney, Recurrent experience replay in distributed reinforcement learning, in: Proceedings of the International Conference on Learning Representations (ICLR 2019), 2019, pp.\u00a01\u201319."},{"key":"10.3233\/AIC-230040_ref32","doi-asserted-by":"crossref","unstructured":"R.\u00a0Kemker, M.\u00a0Mcclure, A.\u00a0Abitino, T.L.\u00a0Hayes and C.\u00a0Kanan, Measuring catastrophic forgetting in neural networks, in: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u00a018), 2018, pp.\u00a03390\u20133398.","DOI":"10.1609\/aaai.v32i1.11651"},{"key":"10.3233\/AIC-230040_ref33","unstructured":"D.P.\u00a0Kingma and J.L.\u00a0Ba, Adam: A method for stochastic optimisation, in: Proceedings of the International Conference on Learning Representations (ICLR 2015), 2015, pp.\u00a01\u201315."},{"issue":"13","key":"10.3233\/AIC-230040_ref34","doi-asserted-by":"publisher","first-page":"3521","DOI":"10.1073\/pnas.1611835114","article-title":"Overcoming catastrophic forgetting in neural networks","volume":"114","author":"Kirkpatrick","year":"2017","journal-title":"Proceedings of the National Academy of Sciences of the United States of America (PNAS 2017)"},{"issue":"1","key":"10.3233\/AIC-230040_ref35","first-page":"1333","article-title":"Transfer in reinforcement learning via shared features","volume":"13","author":"Konidaris","year":"2012","journal-title":"Journal of Machine Learning Research"},{"key":"10.3233\/AIC-230040_ref36","doi-asserted-by":"crossref","unstructured":"G.\u00a0Lample and D.S.\u00a0Chaplot, Playing FPS games with deep reinforcement learning, in: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2016), 2017, pp.\u00a02140\u20132146.","DOI":"10.1609\/aaai.v31i1.10827"},{"key":"10.3233\/AIC-230040_ref37","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-27645-3_5"},{"key":"10.3233\/AIC-230040_ref38","unstructured":"E.\u00a0Lecarpentier and E.\u00a0Rachelson, Non-stationary Markov decision processes: A worst-case approach using model-based reinforcement learning, in: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS 2019), 2019."},{"key":"10.3233\/AIC-230040_ref39","unstructured":"A.\u00a0Levy, R.\u00a0Platt, G.\u00a0Konidaris and K.\u00a0Saenko, Learning multi-level hierarchies with hindsight, in: Proceedings of the International Conference on Learning Representations (ICLR 2019), 2019, pp.\u00a01\u201316."},{"key":"10.3233\/AIC-230040_ref40","unstructured":"S.\u00a0Li, F.\u00a0Gu, G.\u00a0Zhu and C.\u00a0Zhang, Context-aware policy reuse, in: Proceedings of the International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS 2018), 2018."},{"key":"10.3233\/AIC-230040_ref42","doi-asserted-by":"crossref","unstructured":"S.\u00a0Li and C.\u00a0Zhang, An optimal online method of selecting source policies for reinforcement learning, in: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2018), 2018, pp.\u00a03562\u20133570.","DOI":"10.1609\/aaai.v32i1.11718"},{"issue":"12","key":"10.3233\/AIC-230040_ref43","doi-asserted-by":"publisher","first-page":"2935","DOI":"10.1109\/TPAMI.2017.2773081","article-title":"Learning without forgetting","volume":"40","author":"Li","year":"2018","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"10.3233\/AIC-230040_ref44","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW50498.2020.00132"},{"issue":"7540","key":"10.3233\/AIC-230040_ref46","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"10.3233\/AIC-230040_ref47","unstructured":"A.\u00a0Naik, R.\u00a0Shariff, N.\u00a0Yasui, H.\u00a0Yao and R.S.\u00a0Sutton, Discounted reinforcement learning is not an optimization problem, in: Optimization Foundations for Reinforcement Learning Workshop at the Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada, 2019, pp.\u00a01\u20137."},{"key":"10.3233\/AIC-230040_ref48","unstructured":"A.\u00a0Nair, V.\u00a0Pong, M.\u00a0Dalal, S.\u00a0Bahl, S.\u00a0Lin and S.\u00a0Levine, Visual reinforcement learning with imagined goals, in: Advances in Neural Information Processing Systems (NeurIPS 2018), 2018, pp.\u00a09191\u20139200."},{"issue":"10","key":"10.3233\/AIC-230040_ref50","doi-asserted-by":"publisher","first-page":"1345","DOI":"10.1109\/TKDE.2009.191","article-title":"A survey on transfer learning","volume":"22","author":"Pan","year":"2010","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"10.3233\/AIC-230040_ref51","unstructured":"M.\u00a0Riemer, I.\u00a0Cases, R.\u00a0Ajemian, M.\u00a0Liu, I.\u00a0Rish, Y.\u00a0Tu and G.\u00a0Tesauro, Learning to learn without forgetting by maximizing transfer and minimizing interference, in: Proceedings of the International Conference on Learning Representations (ICLR 2019), 2019."},{"key":"10.3233\/AIC-230040_ref52","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1007\/s10994-016-5547-y","article-title":"Bayesian policy reuse","volume":"104","author":"Rosman","year":"2016","journal-title":"Machine Learning"},{"key":"10.3233\/AIC-230040_ref53","doi-asserted-by":"publisher","first-page":"673","DOI":"10.1613\/JAIR.1.11304","article-title":"Using task descriptions in lifelong machine learning for improved performance and zero-shot transfer","volume":"67","author":"Rostami","year":"2020","journal-title":"Journal of Artificial Intelligence Research"},{"issue":"5","key":"10.3233\/AIC-230040_ref54","doi-asserted-by":"publisher","first-page":"533","DOI":"10.1038\/323533a0","article-title":"Learning representations by back-propagating errors","volume":"323","author":"Rumelhart","year":"1986","journal-title":"Nature"},{"key":"10.3233\/AIC-230040_ref56","doi-asserted-by":"publisher","DOI":"10.1142\/s0129065707001111"},{"key":"10.3233\/AIC-230040_ref57","unstructured":"T.\u00a0Schaul, D.\u00a0Horgan, K.\u00a0Gregor and D.\u00a0Silver, Universal value function approximators, in: Proceedings of the International Conference on Machine Learning (ICML 2015), Lille, France, 2015, pp.\u00a01312\u20131320."},{"key":"10.3233\/AIC-230040_ref58","doi-asserted-by":"publisher","DOI":"10.1142\/9789812817471_0003"},{"key":"10.3233\/AIC-230040_ref59","unstructured":"J.\u00a0Schulman, P.\u00a0Moritz, S.\u00a0Levine, M.I.\u00a0Jordan and P.\u00a0Abbeel, High-dimensional continuous control using generalised advantage estimation, in: Proceedings of the International Conference on Learning Representations (ICLR 2016), 2016."},{"key":"10.3233\/AIC-230040_ref61","doi-asserted-by":"crossref","unstructured":"C.\u00a0Schulze and M.\u00a0Schulze, ViZDoom: DRQN with prioritized experience replay, double-q learning, & snapshot ensembling, in: Proceedings of the SAI Intelligent Systems Conference (IntelliSys 2018), 2018, pp.\u00a01\u201317.","DOI":"10.1007\/978-3-030-01054-6_1"},{"key":"10.3233\/AIC-230040_ref62","unstructured":"D.L.\u00a0Silver, Q.\u00a0Yang and L.\u00a0Li, Lifelong machine learning systems: Beyond learning algorithms, in: AAAI Spring Symposium Series (SSS 2013), 2013, pp.\u00a049\u201355."},{"key":"10.3233\/AIC-230040_ref63","first-page":"69","article-title":"VC dimension of neural networks","volume":"168","author":"Sontag","year":"1998","journal-title":"NATO ASI Series F Computer and Systems Sciences"},{"key":"10.3233\/AIC-230040_ref64","doi-asserted-by":"publisher","DOI":"10.1016\/s1364-6613(99)01331-5"},{"key":"10.3233\/AIC-230040_ref65","first-page":"1633","article-title":"Transfer learning for reinforcement learning domains: A survey","volume":"10","author":"Taylor","year":"2009","journal-title":"Journal of Machine Learning Research"},{"key":"10.3233\/AIC-230040_ref67","doi-asserted-by":"crossref","unstructured":"C.\u00a0Tessler, S.\u00a0Givony, T.\u00a0Zahavy, D.J.\u00a0Mankowitz and S.\u00a0Mannor, A deep hierarchical approach to lifelong learning in minecraft, in: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2016), 2016, pp.\u00a055\u20131561.","DOI":"10.1609\/aaai.v31i1.10744"},{"key":"10.3233\/AIC-230040_ref68","unstructured":"S.\u00a0Thrun and A.\u00a0Schwartz, Finding structure in reinforcement learning, in: Advances in Neural Information Processing Systems (NeurIPS 1995), 1995, pp.\u00a0385\u2013392."},{"key":"10.3233\/AIC-230040_ref70","unstructured":"University of Southampton, The Iridis Compute Cluster, 2017. https:\/\/www.southampton.ac.uk\/isolutions\/staff\/iridis.page."},{"issue":"2","key":"10.3233\/AIC-230040_ref71","doi-asserted-by":"publisher","first-page":"264","DOI":"10.1137\/1116025","article-title":"On the uniform convergence of relative frequencies of events to their probabilities","volume":"16","author":"Vapnik","year":"1971","journal-title":"Theory of probability and its applications"},{"key":"10.3233\/AIC-230040_ref72","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1613\/jair.3125","article-title":"A Monte-Carlo AIXI approximation","volume":"40","author":"Veness","year":"2011","journal-title":"Journal of Artificial Intelligence Research"},{"key":"10.3233\/AIC-230040_ref73","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1016\/j.neucom.2020.02.117","article-title":"Target transfer Q-learning and its convergence analysis","volume":"392","author":"Wang","year":"2020","journal-title":"Neurocomputing"},{"issue":"3\u20134","key":"10.3233\/AIC-230040_ref74","doi-asserted-by":"publisher","first-page":"279","DOI":"10.1007\/BF00992698","volume":"8","author":"Watkins","year":"1992","journal-title":"Q-learning, Machine Learning"},{"key":"10.3233\/AIC-230040_ref75","doi-asserted-by":"publisher","DOI":"10.1145\/1273496.1273624"},{"key":"10.3233\/AIC-230040_ref76","unstructured":"T.\u00a0Xu, Q.\u00a0Liu, L.\u00a0Zhao and J.\u00a0Peng, Learning to explore via meta-policy gradient, in: Proceedings of the International Conference on Machine Learning (ICML 2018), Vol.\u00a012, Stockholm, Sweden, 2018, pp.\u00a08686\u20138706."},{"key":"10.3233\/AIC-230040_ref77","unstructured":"T.\u00a0Yu, D.\u00a0Quillen, Z.\u00a0He, R.\u00a0Julian, K.\u00a0Hausman, C.\u00a0Finn and S.\u00a0Levine, Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning, in: Proceedings of the Conference on Robot Learning (CoRL 2019), 2019, pp.\u00a01\u201318."},{"key":"10.3233\/AIC-230040_ref79","unstructured":"Y.\u00a0Zheng, Z.\u00a0Meng, J.\u00a0Hao, Z.\u00a0Zhang, T.\u00a0Yang and C.\u00a0Fan, A deep Bayesian policy reuse approach against non-stationary agents, in: Advances in Neural Information Processing Systems (NeurIPS 2018), 2018, pp.\u00a0954\u2013964."}],"container-title":["AI Communications"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/AIC-230040","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T18:28:12Z","timestamp":1777400892000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/full\/10.3233\/AIC-230040"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,21]]},"references-count":68,"journal-issue":{"issue":"1"},"URL":"https:\/\/doi.org\/10.3233\/aic-230040","relation":{},"ISSN":["1875-8452","0921-7126"],"issn-type":[{"value":"1875-8452","type":"electronic"},{"value":"0921-7126","type":"print"}],"subject":[],"published":{"date-parts":[[2024,3,21]]}}}