{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:08:42Z","timestamp":1750219722409,"version":"3.41.0"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,9,20]],"date-time":"2023-09-20T00:00:00Z","timestamp":1695168000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Evol. Learn. Optim."],"published-print":{"date-parts":[[2023,9,30]]},"abstract":"<jats:p>\n            When searching for policies, reward-sparse environments often lack sufficient information about which behaviors to improve upon or avoid. In such environments, the policy search process is bound to blindly search for reward-yielding transitions and no early reward can bias this search in one direction or another. A way to overcome this is to use intrinsic motivation in order to explore new transitions until a reward is found. In this work, we use a recently proposed definition of intrinsic motivation, Curiosity, in an evolutionary policy search method. We propose Curiosity-ES,\n            <jats:xref ref-type=\"fn\">\n              <jats:sup>1<\/jats:sup>\n            <\/jats:xref>\n            an evolutionary strategy adapted to use Curiosity as a fitness metric. We compare Curiosity-ES with other evolutionary algorithms intended for exploration, as well as with Curiosity-based reinforcement learning, and find that Curiosity-ES can generate higher diversity without the need for an explicit diversity criterion and leads to more policies which find reward.\n          <\/jats:p>","DOI":"10.1145\/3605782","type":"journal-article","created":{"date-parts":[[2023,6,30]],"date-time":"2023-06-30T11:56:45Z","timestamp":1688126205000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Curiosity Creates Diversity in Policy Search"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4555-6718","authenticated-orcid":false,"given":"Paul-Antoine","family":"Le Tolguenec","sequence":"first","affiliation":[{"name":"ISAE-Supaero, Universit\u00e9 de Toulouse, Airbus, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8559-1617","authenticated-orcid":false,"given":"Emmanuel","family":"Rachelson","sequence":"additional","affiliation":[{"name":"ISAE-Supaero, Universit\u00e9 de Toulouse, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3639-6090","authenticated-orcid":false,"given":"Yann","family":"Besse","sequence":"additional","affiliation":[{"name":"Airbus, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2414-0051","authenticated-orcid":false,"given":"Dennis G.","family":"Wilson","sequence":"additional","affiliation":[{"name":"ISAE-Supaero, Universit\u00e9 de Toulouse, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,9,20]]},"reference":[{"key":"e_1_3_2_2_1","unstructured":"Ferran Alet Martin F. Schneider Tom\u00e1s Lozano-P\u00e9rez and Leslie Pack Kaelbling. 2020. Meta-learning curiosity algorithms. (2020)."},{"key":"e_1_3_2_3_1","article-title":"Hindsight experience replay","volume":"30","author":"Andrychowicz Marcin","year":"2017","unstructured":"Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba. 2017. Hindsight experience replay. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_4_1","article-title":"A survey on intrinsic motivation in reinforcement learning","author":"Aubret Arthur","year":"2019","unstructured":"Arthur Aubret, Laetitia Matignon, and Salima Hassas. 2019. A survey on intrinsic motivation in reinforcement learning. arXiv preprint arXiv:1908.06976 (2019).","journal-title":"arXiv preprint arXiv:1908.06976"},{"key":"e_1_3_2_5_1","volume-title":"International Conference on Learning Representations","author":"Badia Adri\u00e0 Puigdom\u00e8nech","year":"2020","unstructured":"Adri\u00e0 Puigdom\u00e8nech Badia, Pablo Sprechmann, Alex Vitvitskyi, Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martin Arjovsky, Alexander Pritzel, Andrew Bolt, and Charles Blundell. 2020. Never give up: Learning directed exploration strategies. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=Sye57xStvB."},{"key":"e_1_3_2_6_1","first-page":"213","article-title":"R-max-a general polynomial time algorithm for near-optimal reinforcement learning","volume":"3","author":"Brafman Ronen I.","year":"2002","unstructured":"Ronen I. Brafman and Moshe Tennenholtz. 2002. R-max-a general polynomial time algorithm for near-optimal reinforcement learning. Journal of Machine Learning Research 3, Oct. (2002), 213\u2013231.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_7_1","volume-title":"International Conference on Learning Representations","author":"Burda Yuri","year":"2018","unstructured":"Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov. 2018. Exploration by random network distillation. In International Conference on Learning Representations."},{"key":"e_1_3_2_8_1","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1007\/978-3-030-66515-9_4","volume-title":"Black Box Optimization, Machine Learning, and No-Free Lunch Theorems","author":"Chatzilygeroudis Konstantinos","year":"2021","unstructured":"Konstantinos Chatzilygeroudis, Antoine Cully, Vassilis Vassiliades, and Jean-Baptiste Mouret. 2021. Quality-diversity optimization: A novel branch of stochastic optimization. In Black Box Optimization, Machine Learning, and No-Free Lunch Theorems. Springer, 109\u2013135."},{"key":"e_1_3_2_9_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2018\/197"},{"key":"e_1_3_2_10_1","article-title":"Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents","volume":"31","author":"Conti Edoardo","year":"2018","unstructured":"Edoardo Conti, Vashisht Madhavan, Felipe Petroski Such, Joel Lehman, Kenneth Stanley, and Jeff Clune. 2018. Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents. Advances in Neural Information Processing Systems 31 (2018).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3321707.3321804"},{"key":"e_1_3_2_12_1","doi-asserted-by":"publisher","unstructured":"Antoine Cully Jeff Clune Danesh Tarapore and Jean-Baptiste Mouret. [n. d.]. Robots that can adapt like animals. 521 7553 ([n. d.]) 503\u2013507. Issue 7553. 10.1038\/nature14422","DOI":"10.1038\/nature14422"},{"key":"e_1_3_2_13_1","volume-title":"Proceedings of the Eleventh International Conference on Learning Representations (ICLR 2023)","author":"Eberhard Onno","year":"2023","unstructured":"Onno Eberhard, Jakob Hollenstein, Cristina Pinneri, and Georg Martius. 2023. Pink noise is all you need: Colored noise exploration in deep reinforcement learning. In Proceedings of the Eleventh International Conference on Learning Representations (ICLR 2023). https:\/\/openreview.net\/forum?id=hQ9V5QN27eS."},{"key":"e_1_3_2_14_1","first-page":"arXiv\u20131901","article-title":"Go-explore: A new approach for hard-exploration problems","author":"Ecoffet Adrien","year":"2019","unstructured":"Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O. Stanley, and Jeff Clune. 2019. Go-explore: A new approach for hard-exploration problems. arXiv e-prints (2019), arXiv\u20131901.","journal-title":"arXiv e-prints"},{"key":"e_1_3_2_15_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-020-03157-9"},{"key":"e_1_3_2_16_1","volume-title":"International Conference on Learning Representations","author":"Eysenbach Benjamin","year":"2019","unstructured":"Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine. 2019. Diversity is all you need: Learning skills without a reward function. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=SJx63jRqFm."},{"key":"e_1_3_2_17_1","volume-title":"International Conference on Learning Representations","author":"Flet-Berliac Yannis","year":"2020","unstructured":"Yannis Flet-Berliac, Johan Ferret, Olivier Pietquin, Philippe Preux, and Matthieu Geist. 2020. Adversarially guided actor-critic. In International Conference on Learning Representations."},{"key":"e_1_3_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377930.3390232"},{"key":"e_1_3_2_19_1","article-title":"Ex2: Exploration with exemplar models for deep reinforcement learning","volume":"30","author":"Fu Justin","year":"2017","unstructured":"Justin Fu, John Co-Reyes, and Sergey Levine. 2017. Ex2: Exploration with exemplar models for deep reinforcement learning. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_20_1","first-page":"1587","volume-title":"International Conference on Machine Learning","author":"Fujimoto Scott","year":"2018","unstructured":"Scott Fujimoto, Herke Hoof, and David Meger. 2018. Addressing function approximation error in actor-critic methods. In International Conference on Machine Learning. PMLR, 1587\u20131596."},{"key":"e_1_3_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3449726.3459490"},{"key":"e_1_3_2_22_1","article-title":"Variational intrinsic control","author":"Gregor Karol","year":"2016","unstructured":"Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra. 2016. Variational intrinsic control. arXiv preprint arXiv:1611.07507 (2016).","journal-title":"arXiv preprint arXiv:1611.07507"},{"key":"e_1_3_2_23_1","doi-asserted-by":"publisher","unstructured":"Luca Grillotti and Antoine Cully. [n. d.]. Unsupervised behaviour discovery with quality-diversity optimisation. ([n. d.]) 1\u20131. 10.1109\/TEVC.2022.3159855","DOI":"10.1109\/TEVC.2022.3159855"},{"key":"e_1_3_2_24_1","article-title":"Dream to control: Learning behaviors by latent imagination","author":"Hafner Danijar","year":"2019","unstructured":"Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. 2019. Dream to control: Learning behaviors by latent imagination. arXiv preprint arXiv:1912.01603 (2019).","journal-title":"arXiv preprint arXiv:1912.01603"},{"key":"e_1_3_2_25_1","doi-asserted-by":"publisher","DOI":"10.1162\/106365601750190398"},{"key":"e_1_3_2_26_1","doi-asserted-by":"publisher","DOI":"10.1038\/scientificamerican0792-66"},{"key":"e_1_3_2_27_1","first-page":"1","volume-title":"2018 IEEE Intelligent Vehicles Symposium (IV)","author":"Koren Mark","year":"2018","unstructured":"Mark Koren, Saud Alsaif, Ritchie Lee, and Mykel J. Kochenderfer. 2018. Adaptive stress testing for autonomous vehicles. In 2018 IEEE Intelligent Vehicles Symposium (IV). IEEE, 1\u20137."},{"key":"e_1_3_2_28_1","first-page":"329","volume-title":"ALIFE","author":"Lehman Joel","year":"2008","unstructured":"Joel Lehman and Kenneth O. Stanley. [n. d.]. Exploiting open-endedness to solve problems through the search for novelty. In ALIFE (2008). 329\u2013336."},{"key":"e_1_3_2_29_1","doi-asserted-by":"publisher","DOI":"10.1162\/EVCO_a_00025"},{"key":"e_1_3_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2001576.2001606"},{"key":"e_1_3_2_31_1","first-page":"3053","volume-title":"International Conference on Machine Learning","author":"Liang Eric","year":"2018","unstructured":"Eric Liang, Richard Liaw, Robert Nishihara, Philipp Moritz, Roy Fox, Ken Goldberg, Joseph Gonzalez, Michael Jordan, and Ion Stoica. 2018. RLlib: Abstractions for distributed reinforcement learning. In International Conference on Machine Learning. PMLR, 3053\u20133062."},{"key":"e_1_3_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/IROS51168.2021.9636234"},{"key":"e_1_3_2_33_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_2_34_1","first-page":"561","volume-title":"13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Moritz Philipp","year":"2018","unstructured":"Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, et\u00a0al. 2018. Ray: A distributed framework for emerging \\(\\lbrace\\) AI \\(\\rbrace\\) applications. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 561\u2013577."},{"key":"e_1_3_2_35_1","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1007\/978-3-642-18272-3_10","volume-title":"New Horizons in Evolutionary Robotics: Extended Contributions from the 2009 EvoDeRob Workshop","author":"Mouret Jean-Baptiste","year":"2011","unstructured":"Jean-Baptiste Mouret. 2011. Novelty-based multiobjectivization. In New Horizons in Evolutionary Robotics: Extended Contributions from the 2009 EvoDeRob Workshop. Springer, 139\u2013154."},{"key":"e_1_3_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3449639.3459314"},{"key":"e_1_3_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA40945.2020.9196819"},{"key":"e_1_3_2_38_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305890.3305968"},{"key":"e_1_3_2_39_1","doi-asserted-by":"crossref","first-page":"83","DOI":"10.1007\/978-3-642-81283-5_8","volume-title":"Simulationsmethoden in Der Medizin und Biologie","author":"Rechenberg Ingo","year":"1978","unstructured":"Ingo Rechenberg. 1978. Evolutionsstrategien. In Simulationsmethoden in Der Medizin und Biologie. Springer, 83\u2013114."},{"key":"e_1_3_2_40_1","article-title":"Evolution strategies as a scalable alternative to reinforcement learning","author":"Salimans Tim","year":"2017","unstructured":"Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. 2017. Evolution strategies as a scalable alternative to reinforcement learning. arXiv preprint arXiv:1703.03864 (2017).","journal-title":"arXiv preprint arXiv:1703.03864"},{"key":"e_1_3_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.3008735"},{"key":"e_1_3_2_42_1","article-title":"Combining evolution and deep reinforcement learning for policy search: A survey","author":"Sigaud Olivier","year":"2022","unstructured":"Olivier Sigaud. 2022. Combining evolution and deep reinforcement learning for policy search: A survey. arXiv preprint arXiv:2203.14009 (2022).","journal-title":"arXiv preprint arXiv:2203.14009"},{"key":"e_1_3_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3520304.3528770"},{"key":"e_1_3_2_44_1","article-title":"DeepMind control suite","author":"Tassa Yuval","year":"2018","unstructured":"Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et\u00a0al. 2018. DeepMind control suite. arXiv preprint arXiv:1801.00690 (2018).","journal-title":"arXiv preprint arXiv:1801.00690"},{"key":"e_1_3_2_45_1","unstructured":"Bryon Tjanaka Matthew C. Fontaine Yulun Zhang Sam Sommerer Nathan Dennler and Stefanos Nikolaidis. 2021. pyribs: A bare-bones Python library for quality diversity optimization. https:\/\/github.com\/icaros-usc\/pyribs."},{"key":"e_1_3_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185553"}],"container-title":["ACM Transactions on Evolutionary Learning and Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3605782","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3605782","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:36:19Z","timestamp":1750178179000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3605782"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,20]]},"references-count":45,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,9,30]]}},"alternative-id":["10.1145\/3605782"],"URL":"https:\/\/doi.org\/10.1145\/3605782","relation":{},"ISSN":["2688-299X","2688-3007"],"issn-type":[{"type":"print","value":"2688-299X"},{"type":"electronic","value":"2688-3007"}],"subject":[],"published":{"date-parts":[[2023,9,20]]},"assertion":[{"value":"2022-08-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-04","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}