{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,13]],"date-time":"2026-02-13T07:20:02Z","timestamp":1770967202195,"version":"3.50.1"},"reference-count":64,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2020,9,16]],"date-time":"2020-09-16T00:00:00Z","timestamp":1600214400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,9,16]],"date-time":"2020-09-16T00:00:00Z","timestamp":1600214400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Appl Intell"],"published-print":{"date-parts":[[2021,2]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Deep reinforcement learning (DRL) algorithms rely on carefully designed environment rewards that are extrinsic to the agent. However, in many real-world scenarios rewards are sparse or delayed, motivating the need for discovering efficient exploration strategies. While intrinsically motivated agents hold promise of better local exploration, solving problems that require coordinated decisions over long-time horizons remains an open problem. We postulate that to discover such strategies, a DRL agent should be able to combine local and high-level exploration behaviors. To this end, we introduce the concept of fast and slow curiosity that aims to incentivize long-time horizon exploration. Our method decomposes the curiosity bonus into a fast reward that deals with local exploration and a slow reward that encourages global exploration. We formulate this bonus as the error in an agent\u2019s ability to reconstruct the observations given their contexts. We further propose to dynamically weight local and high-level strategies by measuring state diversity. We evaluate our method on a variety of benchmark environments, including Minigrid, Super Mario Bros, and Atari games. Experimental results show that our agent outperforms prior approaches in most tasks in terms of exploration efficiency and mean scores.<\/jats:p>","DOI":"10.1007\/s10489-020-01849-3","type":"journal-article","created":{"date-parts":[[2020,9,16]],"date-time":"2020-09-16T02:02:29Z","timestamp":1600221749000},"page":"1086-1107","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":24,"title":["Fast and slow curiosity for high-level exploration in reinforcement learning"],"prefix":"10.1007","volume":"51","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9856-0038","authenticated-orcid":false,"given":"Nicolas","family":"Bougie","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ryutaro","family":"Ichise","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2020,9,16]]},"reference":[{"key":"1849_CR1","unstructured":"Abel D, Agarwal A, Diaz F, Krishnamurthy A, Schapire RE (2016) Exploratory gradient boosting for reinforcement learning in complex domains. ICML Workshop on Abstraction in Reinforcement Learning"},{"key":"1849_CR2","unstructured":"Achiam J, Sastry S (2017) Surprise-based intrinsic motivation for deep reinforcement learning. arXiv:170301732"},{"key":"1849_CR3","unstructured":"Achiam J, Edwards H, Amodei D, Abbeel P (2018) Variational option discovery algorithms. arXiv:180710299"},{"key":"1849_CR4","unstructured":"Andrychowicz M, Wolski F, Ray A, Schneider J, Fong R, Welinder P, McGrew B, Tobin J, Abbeel OP, Zaremba W (2017) Hindsight experience replay. In: Proceedings of advances in neural information processing systems, pp 5048\u20135058"},{"key":"1849_CR5","unstructured":"Baldi P (2012) Autoencoders, unsupervised learning, and deep architectures. In: Proceedings of International conference on machine learning workshop on unsupervised and transfer learning, pp 37\u201349"},{"issue":"1","key":"1849_CR6","doi-asserted-by":"crossref","first-page":"49","DOI":"10.1016\/j.robot.2012.05.008","volume":"61","author":"A Baranes","year":"2013","unstructured":"Baranes A, Oudeyer PY (2013) Active learning of inverse models with intrinsically motivated goal exploration in robots. Robot. Auton. Syst. 61(1):49\u201373","journal-title":"Robot. Auton. Syst."},{"key":"1849_CR7","unstructured":"Bellemare M, Srinivasan S, Ostrovski G, Schaul T, Saxton D, Munos R (2016) Unifying count-based exploration and intrinsic motivation. In: Proceedings of advances in neural information processing systems, pp 1471\u20131479"},{"key":"1849_CR8","doi-asserted-by":"crossref","first-page":"253","DOI":"10.1613\/jair.3912","volume":"47","author":"MG Bellemare","year":"2013","unstructured":"Bellemare MG, Naddaf Y, Veness J, Bowling M (2013) The arcade learning environment: An evaluation platform for general agents. J. Artif. Intell. Res. 47:253\u2013279","journal-title":"J. Artif. Intell. Res."},{"key":"1849_CR9","doi-asserted-by":"crossref","unstructured":"Botvinick M, Ritter S, Wang JX, Kurth-Nelson Z, Blundell C, Hassabis D (2019) Reinforcement learning, fast and slow. Trends in cognitive sciences","DOI":"10.1016\/j.tics.2019.02.006"},{"key":"1849_CR10","doi-asserted-by":"crossref","first-page":"493","DOI":"10.1007\/s10994-019-05845-8","volume":"109","author":"N Bougie","year":"2019","unstructured":"Bougie N, Ichise R (2019) Skill-based curiosity for intrinsically motivated reinforcement learning. Mach Learn 109:493\u2013512","journal-title":"Mach Learn"},{"key":"1849_CR11","unstructured":"Burda Y, Edwards H, Storkey A, Klimov O (2018) Exploration by random network distillation. arXiv:181012894"},{"key":"1849_CR12","unstructured":"Burda Y, Edwards H, Pathak D, Storkey A, Darrell T, Efros AA (2019) Large-scale study of curiosity-driven learning. In: Proceedings of the The International Conference on Learning Representations"},{"key":"1849_CR13","unstructured":"Burgess CP, Higgins I, Pal A, Matthey L, Watters N, Desjardins G, Lerchner A (2018) Understanding disentangling in beta-vae. arXiv:180403599"},{"key":"1849_CR14","unstructured":"Chevalier-Boisvert M, Willems L, Pal S (2018) Minimalistic gridworld environment for openai gym. https:\/\/github.com\/maximecb\/gym-minigrid"},{"key":"1849_CR15","unstructured":"Cho K (2013) Simple sparsification improves sparse denoising autoencoders in denoising highly corrupted images. In: International conference on machine learning, pp 432\u2013440"},{"key":"1849_CR16","unstructured":"Dhariwal P, Hesse C, Klimov O, Nichol A, Plappert M, Radford A, Schulman J, Sidor S, Wu Y, Zhokhov P (2017) Openai baselines. https:\/\/github.com\/openai\/baselines"},{"issue":"12","key":"1849_CR17","doi-asserted-by":"crossref","first-page":"3736","DOI":"10.1109\/TIP.2006.881969","volume":"15","author":"M Elad","year":"2006","unstructured":"Elad M, Aharon M (2006) Image denoising via sparse and redundant representations over learned dictionaries. IEEE Trans. Image Process. 15(12):3736\u20133745","journal-title":"IEEE Trans. Image Process."},{"key":"1849_CR18","unstructured":"Florensa C, Held D, Geng X, Abbeel P (2018) Automatic goal generation for reinforcement learning agents. In: Proceedings of the International Conference on Machine Learning"},{"key":"1849_CR19","unstructured":"Forestier S, Mollard Y, Oudeyer PY (2017) Intrinsically motivated goal exploration processes with automatic curriculum learning. arXiv:170802190"},{"key":"1849_CR20","unstructured":"Haarnoja T, Zhou A, Abbeel P, Levine S (2018) Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. Machine Learning Research"},{"key":"1849_CR21","doi-asserted-by":"crossref","unstructured":"Han D (2013) Comparison of commonly used image interpolation methods. In: Proceedings of the international conference on computer science and electronics engineering","DOI":"10.2991\/iccsee.2013.391"},{"issue":"2","key":"1849_CR22","doi-asserted-by":"crossref","first-page":"245","DOI":"10.1016\/j.neuron.2017.06.011","volume":"95","author":"D Hassabis","year":"2017","unstructured":"Hassabis D, Kumaran D, Summerfield C, Botvinick M (2017) Neuroscience-inspired artificial intelligence. Neuron 95(2):245\u2013258","journal-title":"Neuron"},{"key":"1849_CR23","unstructured":"Higgins I, Matthey L, Glorot X, Pal A, Uria B, Blundell C, Mohamed S, Lerchner A (2016) Early visual concept learning with unsupervised deep learning. arXiv:160605579"},{"key":"1849_CR24","doi-asserted-by":"crossref","first-page":"106945","DOI":"10.1016\/j.patcog.2019.06.011","volume":"96","author":"I Hong","year":"2019","unstructured":"Hong I, Hwang Y, Kim D (2019) Efficient deep learning of image denoising using patch complexity local divide and deep conquer. Pattern Recogn. 96:106945","journal-title":"Pattern Recogn."},{"key":"1849_CR25","unstructured":"Hong ZW, Shann TY, Su SY, Chang YH, Fu TJ, Lee CY (2018) Diversity-driven exploration strategy for deep reinforcement learning. In: Proceedings of Advances in neural information processing systems"},{"key":"1849_CR26","unstructured":"Houthooft R, Chen X, Chen X, Duan Y, Schulman J, De Turck F, Abbeel P (2016) Vime: Variational information maximizing exploration. In: Proceedings of advances in neural information processing systems, pp 1109\u20131117"},{"key":"1849_CR27","unstructured":"Jinnai Y, Park JW, Abel D, Konidaris G (2019) Discovering options for exploration by minimizing cover time. In: Proceedings of the International Conference on Machine Learning"},{"key":"1849_CR28","unstructured":"Kaelbling LP (1993) Learning to achieve goals. In: Proceedings of the International Joint Conferences on Artificial Intelligence, pp 1094\u20131098"},{"key":"1849_CR29","unstructured":"Kauten C (2018) Super Mario Bros for OpenAI Gym. https:\/\/github.com\/Kautenja\/gym-super-mario-bros"},{"key":"1849_CR30","unstructured":"Kingma DP, Ba J (2014) Adam: A method for stochastic optimization. arXiv:14126980"},{"key":"1849_CR31","unstructured":"Kingma DP, Welling M (2014) Auto-encoding variational Bayes. In: Proceedings of the international conference on learning representations"},{"key":"1849_CR32","doi-asserted-by":"crossref","unstructured":"Klyubin AS, Polani D, Nehaniv CL (2005) Empowerment: A universal agent-centric measure of control. In: IEEE Congress on Evolutionary Computation, vol 1, pp 128\u2013135","DOI":"10.1109\/CEC.2005.1554676"},{"key":"1849_CR33","doi-asserted-by":"crossref","unstructured":"Kuderer M, Gulati S, Burgard W (2015) Learning driving styles for autonomous vehicles from demonstration. In: IEEE International Conference on Robotics and Automation, pp 2641\u20132646","DOI":"10.1109\/ICRA.2015.7139555"},{"key":"1849_CR34","doi-asserted-by":"crossref","unstructured":"Lehman J, Stanley KO (2011) Abandoning objectives: Evolution through the search for novelty alone. Evolutionary computation, 189\u2013223","DOI":"10.1162\/EVCO_a_00025"},{"key":"1849_CR35","unstructured":"Lillicrap TP, Hunt JJ, Pritzel A, Heess N, Erez T, Tassa Y, Silver D, Wierstra D (2016) Continuous control with deep reinforcement learning. In: Proceedings of international conference on learning representations"},{"key":"1849_CR36","unstructured":"Machado MC, Bellemare MG, Bowling M (2017) A laplacian framework for option discovery in reinforcement learning. In: Proceedings of the International Conference on Machine Learning, pp 2295\u20132304"},{"key":"1849_CR37","unstructured":"Machado MC, Bellemare MG, Bowling M (2018) Count-based exploration with the successor representation. arXiv:180711622"},{"key":"1849_CR38","unstructured":"Mao XJ, Shen C, Yang YB (2016) Image restoration using convolutional auto-encoders with symmetric skip connections. arXiv:160608921"},{"key":"1849_CR39","doi-asserted-by":"crossref","unstructured":"Martin J, Sasikumar SN, Everitt T, Hutter M (2017) Count-based exploration in feature space for reinforcement learning. In: Proceedings of the International Joint Conference on Artificial Intelligence","DOI":"10.24963\/ijcai.2017\/344"},{"issue":"7540","key":"1849_CR40","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, Graves A, Riedmiller M, Fidjeland AK, Ostrovski G, et al. (2015) Human-level control through deep reinforcement learning. Nature 518(7540):529","journal-title":"Nature"},{"key":"1849_CR41","unstructured":"Mnih V, Badia AP, Mirza M, Graves A, Lillicrap T, Harley T, Silver D, Kavukcuoglu K (2016) Asynchronous methods for deep reinforcement learning. In: Proceedings of international conference on machine learning, pp 1928\u20131937"},{"key":"1849_CR42","unstructured":"Nair AV, Pong V, Dalal M, Bahl S, Lin S, Levine S (2018) Visual reinforcement learning with imagined goals. In: Proceedings of advances in neural information processing systems, pp 9191\u20139200"},{"key":"1849_CR43","unstructured":"Ostrovski G, Bellemare MG, van den Oord A, Munos R (2017) Count-based exploration with neural density models. In: Proceedings of the international conference on machine learning, pp 2721\u20132730"},{"key":"1849_CR44","doi-asserted-by":"crossref","unstructured":"Pathak D, Agrawal P, Efros AA, Darrell T (2017) Curiosity-driven exploration by self-supervised prediction: In Proceedings of the international conference on international conference on machine learning","DOI":"10.1109\/CVPRW.2017.70"},{"key":"1849_CR45","unstructured":"Pere A, Forestier S, Sigaud O, Oudeyer PY (2018) Unsupervised learning of goal spaces for intrinsically motivated goal exploration. In Proceedings of the international conference on learning representations"},{"key":"1849_CR46","unstructured":"Plappert M, Andrychowicz M, Ray A, McGrew B, Baker B, Powell G, Schneider J, Tobin J, Chociej M, Welinder P, et al. (2018) Multi-goal reinforcement learning: Challenging robotics environments and request for research. arXiv:180209464"},{"key":"1849_CR47","unstructured":"Pong VH, Dalal M, Lin S, Nair A, Bahk S, Levine S (2019) Skew-fit: State-covering self-supervised reinforcement learning. arXiv:190303698"},{"key":"1849_CR48","unstructured":"Savinov N, Raichuk A, Marinier R, Vincent D, Pollefeys M, Lillicrap T, Gelly S (2019) Episodic curiosity through reachability. In: Proceedings of the international conference on learning representations"},{"key":"1849_CR49","unstructured":"Schaul T, Horgan D, Gregor K, Silver D (2015) Universal value function approximators. In: Proceedings of the International conference on machine learning"},{"key":"1849_CR50","unstructured":"Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O (2017) Proximal policy optimization algorithms. arXiv:170706347"},{"issue":"7587","key":"1849_CR51","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver D, Huang A, Maddison CJ, Guez A, Sifre L, Van Den Driessche G, Schrittwieser J, Antonoglou I, Panneershelvam V, Lanctot M, et al. (2016) Mastering the game of go with deep neural networks and tree search. Nature 529(7587):484","journal-title":"Nature"},{"key":"1849_CR52","doi-asserted-by":"crossref","unstructured":"Snell J, Ridgeway K, Liao R, Roads BD, Mozer MC, Zemel RS (2017) Learning to generate images with perceptual similarity metrics. In: IEEE International Conference on Image Processing (ICIP), vol 2017. IEEE, pp 4277\u20134281","DOI":"10.1109\/ICIP.2017.8297089"},{"key":"1849_CR53","unstructured":"Stadie BC, Levine S, Abbeel P (2015) Incentivizing exploration in reinforcement learning with deep predictive models"},{"key":"1849_CR54","unstructured":"Stanton C, Clune J (2019) Deep curiosity search: Intra-life exploration improves performance on challenging deep reinforcement learning problems. In: Proceedings of the international conference on international conference on machine learning"},{"issue":"8","key":"1849_CR55","doi-asserted-by":"crossref","first-page":"1309","DOI":"10.1016\/j.jcss.2007.08.009","volume":"74","author":"AL Strehl","year":"2008","unstructured":"Strehl AL, Littman ML (2008) An analysis of model-based interval estimation for Markov decision processes. J. Comput. Syst. Sci. 74(8):1309\u20131331","journal-title":"J. Comput. Syst. Sci."},{"issue":"1","key":"1849_CR56","first-page":"9","volume":"3","author":"RS Sutton","year":"1988","unstructured":"Sutton RS (1988) Learning to predict by the methods of temporal differences. Machine Learning 3(1):9\u201344","journal-title":"Machine Learning"},{"key":"1849_CR57","volume-title":"Reinforcement learning: an introduction","author":"RS Sutton","year":"1998","unstructured":"Sutton RS, Barto AG (1998) Reinforcement learning: an introduction. MIT Press, Cambridge"},{"key":"1849_CR58","unstructured":"Tang H, Houthooft R, Foote D, Stooke A, Chen OX, Duan Y, Schulman J, DeTurck F, Abbeel P (2017) # exploration: A study of count-based exploration for deep reinforcement learning. In: Proceedings of Advances in neural information processing systems, pp 2753\u20132762"},{"key":"1849_CR59","doi-asserted-by":"crossref","unstructured":"Todorov E, Erez T, Tassa Y (2012) Mujoco: A physics engine for model-based control. In: IEEE\/RSJ international conference on intelligent robots and systems, pp 5026\u20135033","DOI":"10.1109\/IROS.2012.6386109"},{"key":"1849_CR60","unstructured":"Vezhnevets AS, Osindero S, Schaul T, Heess N, Jaderberg M, Silver D, Kavukcuoglu K (2017) Feudal networks for hierarchical reinforcement learning. In: Proceedings of the international conference on machine learning, pp 3540\u20133549"},{"key":"1849_CR61","doi-asserted-by":"crossref","unstructured":"Wang Z, Simoncelli EP, Bovik AC (2003) Multiscale structural similarity for image quality assessment. In: Proceedings of the Conference on Signals, Systems & Computers, vol 2. IEEE, pp 1398\u20131402","DOI":"10.1109\/ACSSC.2003.1292216"},{"issue":"4","key":"1849_CR62","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","volume":"13","author":"Z Wang","year":"2004","unstructured":"Wang Z, Bovik AC, Sheikh HR, Simoncelli EP, et al. (2004) Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing 13(4):600\u2013612","journal-title":"IEEE Transactions on Image Processing"},{"key":"1849_CR63","unstructured":"Wang Z, Schaul T, Hessel M, Van Hasselt H, Lanctot M, De Freitas N (2016) Dueling network architectures for deep reinforcement learning"},{"key":"1849_CR64","unstructured":"Yang HK, Chiang PH, Hong MF, Lee CY (2019) Exploration via flow-based intrinsic rewards. arXiv:190510071"}],"container-title":["Applied Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-020-01849-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10489-020-01849-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-020-01849-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,16]],"date-time":"2021-09-16T00:22:46Z","timestamp":1631751766000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10489-020-01849-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,16]]},"references-count":64,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,2]]}},"alternative-id":["1849"],"URL":"https:\/\/doi.org\/10.1007\/s10489-020-01849-3","relation":{},"ISSN":["0924-669X","1573-7497"],"issn-type":[{"value":"0924-669X","type":"print"},{"value":"1573-7497","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,9,16]]},"assertion":[{"value":"16 September 2020","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}