{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T02:48:07Z","timestamp":1785898087573,"version":"3.56.0"},"publisher-location":"New York, NY, USA","reference-count":46,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,7,8]],"date-time":"2022-07-08T00:00:00Z","timestamp":1657238400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,7,8]]},"DOI":"10.1145\/3512290.3528845","type":"proceedings-article","created":{"date-parts":[[2022,8,17]],"date-time":"2022-08-17T15:32:35Z","timestamp":1660750355000},"page":"1075-1083","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":37,"title":["Diversity policy gradient for sample efficient quality-diversity optimization"],"prefix":"10.1145","author":[{"given":"Thomas","family":"Pierrot","sequence":"first","affiliation":[{"name":"InstaDeep, Paris, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Valentin","family":"Mac\u00e9","sequence":"additional","affiliation":[{"name":"InstaDeep, Paris, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Felix","family":"Chalumeau","sequence":"additional","affiliation":[{"name":"InstaDeep, Paris, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Arthur","family":"Flajolet","sequence":"additional","affiliation":[{"name":"InstaDeep, Paris, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Geoffrey","family":"Cideron","sequence":"additional","affiliation":[{"name":"InstaDeep, Paris, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Karim","family":"Beguir","sequence":"additional","affiliation":[{"name":"InstaDeep, London, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Antoine","family":"Cully","sequence":"additional","affiliation":[{"name":"Imperial College London, London, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Olivier","family":"Sigaud","sequence":"additional","affiliation":[{"name":"Sorbonne Universit\u00e9, Paris, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nicolas","family":"Perrin-Gilbert","sequence":"additional","affiliation":[{"name":"Sorbonne Universit\u00e9, Paris, France"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,7,8]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CIG.2019.8848022"},{"key":"e_1_3_2_2_2_1","volume-title":"Openai gym. arXiv preprint arXiv:1606.01540","author":"Brockman Greg","year":"2016","unstructured":"Greg Brockman , Vicki Cheung , Ludwig Pettersson , Jonas Schneider , John Schulman , Jie Tang , and Wojciech Zaremba . 2016. Openai gym. arXiv preprint arXiv:1606.01540 ( 2016 ). Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016. Openai gym. arXiv preprint arXiv:1606.01540 (2016)."},{"key":"e_1_3_2_2_3_1","volume-title":"Exploration by random network distillation. arXiv preprint arXiv:1810.12894","author":"Burda Yuri","year":"2018","unstructured":"Yuri Burda , Harrison Edwards , Amos Storkey , and Oleg Klimov . 2018. Exploration by random network distillation. arXiv preprint arXiv:1810.12894 ( 2018 ). Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov. 2018. Exploration by random network distillation. arXiv preprint arXiv:1810.12894 (2018)."},{"key":"e_1_3_2_2_4_1","volume-title":"Exploring Self-Assembling Behaviors in a Swarm of Bio-micro-robots using Surrogate-Assisted MAP-Elites. arXiv preprint arXiv:1910.00230","author":"Cazenille Leo","year":"2019","unstructured":"Leo Cazenille , Nicolas Bredeche , and Nathanael Aubert-Kato . 2019. Exploring Self-Assembling Behaviors in a Swarm of Bio-micro-robots using Surrogate-Assisted MAP-Elites. arXiv preprint arXiv:1910.00230 ( 2019 ). Leo Cazenille, Nicolas Bredeche, and Nathanael Aubert-Kato. 2019. Exploring Self-Assembling Behaviors in a Swarm of Bio-micro-robots using Surrogate-Assisted MAP-Elites. arXiv preprint arXiv:1910.00230 (2019)."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377930.3390217"},{"key":"e_1_3_2_2_6_1","volume-title":"GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms. arXiv preprint arXiv:1802.05054","author":"Colas C\u00e9dric","year":"2018","unstructured":"C\u00e9dric Colas , Olivier Sigaud , and Pierre-Yves Oudeyer . 2018. GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms. arXiv preprint arXiv:1802.05054 ( 2018 ). C\u00e9dric Colas, Olivier Sigaud, and Pierre-Yves Oudeyer. 2018. GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms. arXiv preprint arXiv:1802.05054 (2018)."},{"key":"e_1_3_2_2_7_1","volume-title":"Joel Lehman, Kenneth Stanley, and Jeff Clune.","author":"Conti Edoardo","year":"2018","unstructured":"Edoardo Conti , Vashisht Madhavan , Felipe Petroski Such , Joel Lehman, Kenneth Stanley, and Jeff Clune. 2018 . Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents. In Advances in neural information processing systems. 5027--5038. Edoardo Conti, Vashisht Madhavan, Felipe Petroski Such, Joel Lehman, Kenneth Stanley, and Jeff Clune. 2018. Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents. In Advances in neural information processing systems. 5027--5038."},{"key":"e_1_3_2_2_8_1","volume-title":"Robots that can adapt like animals. Nature 521, 7553","author":"Cully Antoine","year":"2015","unstructured":"Antoine Cully , Jeff Clune , Danesh Tarapore , and Jean-Baptiste Mouret . 2015. Robots that can adapt like animals. Nature 521, 7553 ( 2015 ), 503--507. Antoine Cully, Jeff Clune, Danesh Tarapore, and Jean-Baptiste Mouret. 2015. Robots that can adapt like animals. Nature 521, 7553 (2015), 503--507."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TEVC.2017.2704781"},{"key":"e_1_3_2_2_10_1","volume-title":"Attraction-Repulsion Actor-Critic for Continuous Control Reinforcement Learning. arXiv preprint arXiv:1909.07543","author":"Doan Thang","year":"2019","unstructured":"Thang Doan , Bogdan Mazoure , Audrey Durand , Joelle Pineau , and R Devon Hjelm . 2019. Attraction-Repulsion Actor-Critic for Continuous Control Reinforcement Learning. arXiv preprint arXiv:1909.07543 ( 2019 ). Thang Doan, Bogdan Mazoure, Audrey Durand, Joelle Pineau, and R Devon Hjelm. 2019. Attraction-Repulsion Actor-Critic for Continuous Control Reinforcement Learning. arXiv preprint arXiv:1909.07543 (2019)."},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3321707.3321752"},{"key":"e_1_3_2_2_12_1","volume-title":"Go-explore: a new approach for hard-exploration problems. arXiv preprint arXiv:1901.10995","author":"Ecoffet Adrien","year":"2019","unstructured":"Adrien Ecoffet , Joost Huizinga , Joel Lehman , Kenneth O Stanley , and Jeff Clune . 2019. Go-explore: a new approach for hard-exploration problems. arXiv preprint arXiv:1901.10995 ( 2019 ). Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune. 2019. Go-explore: a new approach for hard-exploration problems. arXiv preprint arXiv:1901.10995 (2019)."},{"key":"e_1_3_2_2_13_1","volume-title":"Diversity is All You Need: Learning Skills without a Reward Function. arXiv preprint arXiv:1802.06070","author":"Eysenbach Benjamin","year":"2018","unstructured":"Benjamin Eysenbach , Abhishek Gupta , Julian Ibarz , and Sergey Levine . 2018. Diversity is All You Need: Learning Skills without a Reward Function. arXiv preprint arXiv:1802.06070 ( 2018 ). Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine. 2018. Diversity is All You Need: Learning Skills without a Reward Function. arXiv preprint arXiv:1802.06070 (2018)."},{"key":"e_1_3_2_2_14_1","unstructured":"Yannis Flet-Berliac Johan Ferret Olivier Pietquin and Philippe Preux. [n. d.]. ADVERSARIALLY GUIDED ACTOR-CRITIC. ([n. d.]).  Yannis Flet-Berliac Johan Ferret Olivier Pietquin and Philippe Preux. [n. d.]. ADVERSARIALLY GUIDED ACTOR-CRITIC. ([n. d.])."},{"key":"e_1_3_2_2_15_1","volume-title":"Fontaine and Stefanos Nikolaidis","author":"Matthew","year":"2021","unstructured":"Matthew C. Fontaine and Stefanos Nikolaidis . 2021 . Differentiable Quality Diversity. CoRR abs\/2106.03894 (2021). arXiv:2106.03894 https:\/\/arxiv.org\/abs\/2106.03894 Matthew C. Fontaine and Stefanos Nikolaidis. 2021. Differentiable Quality Diversity. CoRR abs\/2106.03894 (2021). arXiv:2106.03894 https:\/\/arxiv.org\/abs\/2106.03894"},{"key":"e_1_3_2_2_16_1","volume-title":"Intrinsically motivated goal exploration processes with automatic curriculum learning. arXiv preprint arXiv:1708.02190","author":"Forestier S\u00e9bastien","year":"2017","unstructured":"S\u00e9bastien Forestier , Yoan Mollard , and Pierre-Yves Oudeyer . 2017. Intrinsically motivated goal exploration processes with automatic curriculum learning. arXiv preprint arXiv:1708.02190 ( 2017 ). S\u00e9bastien Forestier, Yoan Mollard, and Pierre-Yves Oudeyer. 2017. Intrinsically motivated goal exploration processes with automatic curriculum learning. arXiv preprint arXiv:1708.02190 (2017)."},{"key":"e_1_3_2_2_17_1","volume-title":"Proc. of ICLR","author":"Frans Kevin","year":"2018","unstructured":"Kevin Frans , Jonathan Ho , Xi Chen , Pieter Abbeel , and John Schulman . 2018 . Meta learning shared hierarchies . Proc. of ICLR (2018). Kevin Frans, Jonathan Ho, Xi Chen, Pieter Abbeel, and John Schulman. 2018. Meta learning shared hierarchies. Proc. of ICLR (2018)."},{"key":"e_1_3_2_2_18_1","volume-title":"Herke Van Hoof, and David Meger","author":"Fujimoto Scott","year":"2018","unstructured":"Scott Fujimoto , Herke Van Hoof, and David Meger . 2018 . Addressing function approximation error in actor-critic methods. arXiv preprint arXiv:1802.09477 (2018). Scott Fujimoto, Herke Van Hoof, and David Meger. 2018. Addressing function approximation error in actor-critic methods. arXiv preprint arXiv:1802.09477 (2018)."},{"key":"e_1_3_2_2_19_1","unstructured":"Tuomas Haarnoja Aurick Zhou Kristian Hartikainen George Tucker Sehoon Ha Jie Tan Vikash Kumar Henry Zhu Abhishek Gupta Pieter Abbeel etal 2018. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905 (2018).  Tuomas Haarnoja Aurick Zhou Kristian Hartikainen George Tucker Sehoon Ha Jie Tan Vikash Kumar Henry Zhu Abhishek Gupta Pieter Abbeel et al. 2018. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905 (2018)."},{"key":"e_1_3_2_2_20_1","volume-title":"Marginalized State Distribution Entropy Regularization in Policy Optimization. arXiv preprint arXiv:1912.05128","author":"Islam Riashat","year":"2019","unstructured":"Riashat Islam , Zafarali Ahmed , and Doina Precup . 2019. Marginalized State Distribution Entropy Regularization in Policy Optimization. arXiv preprint arXiv:1912.05128 ( 2019 ). Riashat Islam, Zafarali Ahmed, and Doina Precup. 2019. Marginalized State Distribution Entropy Regularization in Policy Optimization. arXiv preprint arXiv:1912.05128 (2019)."},{"key":"e_1_3_2_2_21_1","unstructured":"Max Jaderberg Valentin Dalibard Simon Osindero Wojciech M Czarnecki Jeff Donahue Ali Razavi Oriol Vinyals Tim Green Iain Dunning Karen Simonyan etal 2017. Population-based training of neural networks. arXiv preprint arXiv:1711.09846 (2017).  Max Jaderberg Valentin Dalibard Simon Osindero Wojciech M Czarnecki Jeff Donahue Ali Razavi Oriol Vinyals Tim Green Iain Dunning Karen Simonyan et al. 2017. Population-based training of neural networks. arXiv preprint arXiv:1711.09846 (2017)."},{"key":"e_1_3_2_2_22_1","volume-title":"Population-Guided Parallel Policy Search for Reinforcement Learning. In International Conference on Learning Representations.","author":"Jung Whiyoung","year":"2020","unstructured":"Whiyoung Jung , Giseung Park , and Youngchul Sung . 2020 . Population-Guided Parallel Policy Search for Reinforcement Learning. In International Conference on Learning Representations. Whiyoung Jung, Giseung Park, and Youngchul Sung. 2020. Population-Guided Parallel Policy Search for Reinforcement Learning. In International Conference on Learning Representations."},{"key":"e_1_3_2_2_23_1","volume-title":"Collaborative evolutionary reinforcement learning. arXiv preprint arXiv:1905.00976","author":"Khadka Shauharda","year":"2019","unstructured":"Shauharda Khadka , Somdeb Majumdar , Santiago Miret , Evren Tumer , Tarek Nassar , Zach Dwiel , Yinyin Liu , and Kagan Tumer . 2019. Collaborative evolutionary reinforcement learning. arXiv preprint arXiv:1905.00976 ( 2019 ). Shauharda Khadka, Somdeb Majumdar, Santiago Miret, Evren Tumer, Tarek Nassar, Zach Dwiel, Yinyin Liu, and Kagan Tumer. 2019. Collaborative evolutionary reinforcement learning. arXiv preprint arXiv:1905.00976 (2019)."},{"key":"e_1_3_2_2_24_1","unstructured":"Shauharda Khadka Somdeb Majumdar Tarek Nassar Zach Dwiel Evren Tumer Santiago Miret Yinyin Liu and Kagan Tumer. [n. d.]. Collaborative Evolutionary Reinforcement Learning. ([n. d.]).  Shauharda Khadka Somdeb Majumdar Tarek Nassar Zach Dwiel Evren Tumer Santiago Miret Yinyin Liu and Kagan Tumer. [n. d.]. Collaborative Evolutionary Reinforcement Learning. ([n. d.])."},{"key":"e_1_3_2_2_25_1","unstructured":"Shauharda Khadka and Kagan Tumer. 2018. Evolution-Guided Policy Gradient in Reinforcement Learning. In Neural Information Processing Systems.  Shauharda Khadka and Kagan Tumer. 2018. Evolution-Guided Policy Gradient in Reinforcement Learning. In Neural Information Processing Systems."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TEVC.2012.2185849"},{"key":"e_1_3_2_2_27_1","volume-title":"Efficient exploration via state marginal matching. arXiv preprint arXiv:1906.05274","author":"Lee Lisa","year":"2019","unstructured":"Lisa Lee , Benjamin Eysenbach , Emilio Parisotto , Eric Xing , Sergey Levine , and Ruslan Salakhutdinov . 2019. Efficient exploration via state marginal matching. arXiv preprint arXiv:1906.05274 ( 2019 ). Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov. 2019. Efficient exploration via state marginal matching. arXiv preprint arXiv:1906.05274 (2019)."},{"key":"e_1_3_2_2_28_1","volume-title":"Abandoning objectives: Evolution through the search for novelty alone. Evolutionary computation 19, 2","author":"Lehman Joel","year":"2011","unstructured":"Joel Lehman and Kenneth O Stanley . 2011. Abandoning objectives: Evolution through the search for novelty alone. Evolutionary computation 19, 2 ( 2011 ), 189--223. Joel Lehman and Kenneth O Stanley. 2011. Abandoning objectives: Evolution through the search for novelty alone. Evolutionary computation 19, 2 (2011), 189--223."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2001576.2001606"},{"key":"e_1_3_2_2_30_1","volume-title":"Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971","author":"Lillicrap Timothy P","year":"2015","unstructured":"Timothy P Lillicrap , Jonathan J Hunt , Alexander Pritzel , Nicolas Heess , Tom Erez , Yuval Tassa , David Silver , and Daan Wierstra . 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971 ( 2015 ). Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971 (2015)."},{"key":"e_1_3_2_2_31_1","volume-title":"Illuminating search spaces by mapping elites. arXiv preprint arXiv:1504.04909","author":"Mouret Jean-Baptiste","year":"2015","unstructured":"Jean-Baptiste Mouret and Jeff Clune . 2015. Illuminating search spaces by mapping elites. arXiv preprint arXiv:1504.04909 ( 2015 ). Jean-Baptiste Mouret and Jeff Clune. 2015. Illuminating search spaces by mapping elites. arXiv preprint arXiv:1504.04909 (2015)."},{"key":"e_1_3_2_2_32_1","volume-title":"Planning with goal-conditioned policies. arXiv preprint arXiv:1911.08453","author":"Nasiriany Soroush","year":"2019","unstructured":"Soroush Nasiriany , Vitchyr H Pong , Steven Lin , and Sergey Levine . 2019. Planning with goal-conditioned policies. arXiv preprint arXiv:1911.08453 ( 2019 ). Soroush Nasiriany, Vitchyr H Pong, Steven Lin, and Sergey Levine. 2019. Planning with goal-conditioned policies. arXiv preprint arXiv:1911.08453 (2019)."},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3449639.3459304"},{"key":"e_1_3_2_2_34_1","unstructured":"Jack Parker-Holder Aldo Pacchiano Krzysztof Choromanski and Stephen Roberts. 2020. Effective Diversity in Population-Based Reinforcement Learning. In Neural Information Processing Systems.  Jack Parker-Holder Aldo Pacchiano Krzysztof Choromanski and Stephen Roberts. 2020. Effective Diversity in Population-Based Reinforcement Learning. In Neural Information Processing Systems."},{"key":"e_1_3_2_2_35_1","volume-title":"Skew-fit: State-covering self-supervised reinforcement learning. arXiv preprint arXiv:1903.03698","author":"Pong Vitchyr H","year":"2019","unstructured":"Vitchyr H Pong , Murtaza Dalal , Steven Lin , Ashvin Nair , Shikhar Bahl , and Sergey Levine . 2019 . Skew-fit: State-covering self-supervised reinforcement learning. arXiv preprint arXiv:1903.03698 (2019). Vitchyr H Pong, Murtaza Dalal, Steven Lin, Ashvin Nair, Shikhar Bahl, and Sergey Levine. 2019. Skew-fit: State-covering self-supervised reinforcement learning. arXiv preprint arXiv:1903.03698 (2019)."},{"key":"e_1_3_2_2_36_1","volume-title":"CEM-RL: Combining evolutionary and gradient-based methods for policy search. arXiv preprint arXiv:1810.01222","author":"Pourchot Alo\u00efs","year":"2018","unstructured":"Alo\u00efs Pourchot and Olivier Sigaud . 2018. CEM-RL: Combining evolutionary and gradient-based methods for policy search. arXiv preprint arXiv:1810.01222 ( 2018 ). Alo\u00efs Pourchot and Olivier Sigaud. 2018. CEM-RL: Combining evolutionary and gradient-based methods for policy search. arXiv preprint arXiv:1810.01222 (2018)."},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.3389\/frobt.2016.00040"},{"key":"e_1_3_2_2_38_1","volume-title":"Evolution strategies as a scalable alternative to reinforcement learning. arXiv preprint arXiv:1703.03864","author":"Salimans Tim","year":"2017","unstructured":"Tim Salimans , Jonathan Ho , Xi Chen , Szymon Sidor , and Ilya Sutskever . 2017. Evolution strategies as a scalable alternative to reinforcement learning. arXiv preprint arXiv:1703.03864 ( 2017 ). Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. 2017. Evolution strategies as a scalable alternative to reinforcement learning. arXiv preprint arXiv:1703.03864 (2017)."},{"key":"e_1_3_2_2_39_1","volume-title":"Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347","author":"Schulman John","year":"2017","unstructured":"John Schulman , Filip Wolski , Prafulla Dhariwal , Alec Radford , and Oleg Klimov . 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 ( 2017 ). John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017)."},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.3008735"},{"key":"e_1_3_2_2_41_1","volume-title":"Proceedings of the 30th International Conference in Machine Learning.","author":"Silver David","year":"2014","unstructured":"David Silver , Guy Lever , Nicolas Heess , Thomas Degris , Daan Wierstra , and Martin Riedmiller . 2014 . Deterministic policy gradient algorithms . In Proceedings of the 30th International Conference in Machine Learning. David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014. Deterministic policy gradient algorithms. In Proceedings of the 30th International Conference in Machine Learning."},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0162235"},{"key":"e_1_3_2_2_43_1","volume-title":"NIPs","volume":"99","author":"Sutton Richard S","year":"1999","unstructured":"Richard S Sutton , David A McAllester , Satinder P Singh , Yishay Mansour , 1999 . Policy gradient methods for reinforcement learning with function approximation .. In NIPs , Vol. 99 . Citeseer, 1057--1063. Richard S Sutton, David A McAllester, Satinder P Singh, Yishay Mansour, et al. 1999. Policy gradient methods for reinforcement learning with function approximation.. In NIPs, Vol. 99. Citeseer, 1057--1063."},{"key":"e_1_3_2_2_44_1","volume-title":"Scaling Up MAP-Elites Using Centroidal Voronoi Tessellations. CoRR abs\/1610.05729","author":"Vassiliades Vassilis","year":"2016","unstructured":"Vassilis Vassiliades , Konstantinos I. Chatzilygeroudis , and Jean-Baptiste Mouret . 2016. Scaling Up MAP-Elites Using Centroidal Voronoi Tessellations. CoRR abs\/1610.05729 ( 2016 ). arXiv:1610.05729 http:\/\/arxiv.org\/abs\/1610.05729 Vassilis Vassiliades, Konstantinos I. Chatzilygeroudis, and Jean-Baptiste Mouret. 2016. Scaling Up MAP-Elites Using Centroidal Voronoi Tessellations. CoRR abs\/1610.05729 (2016). arXiv:1610.05729 http:\/\/arxiv.org\/abs\/1610.05729"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3205455.3205602"},{"key":"e_1_3_2_2_46_1","volume-title":"The 22nd International Conference on Artificial Intelligence and Statistics. PMLR, 2926--2935","author":"Vemula Anirudh","year":"2019","unstructured":"Anirudh Vemula , Wen Sun , and J Bagnell . 2019 . Contrasting exploration in parameter and action space: A zeroth-order optimization perspective . In The 22nd International Conference on Artificial Intelligence and Statistics. PMLR, 2926--2935 . Anirudh Vemula, Wen Sun, and J Bagnell. 2019. Contrasting exploration in parameter and action space: A zeroth-order optimization perspective. In The 22nd International Conference on Artificial Intelligence and Statistics. PMLR, 2926--2935."}],"event":{"name":"GECCO '22: Genetic and Evolutionary Computation Conference","location":"Boston Massachusetts","acronym":"GECCO '22","sponsor":["SIGEVO ACM Special Interest Group on Genetic and Evolutionary Computation"]},"container-title":["Proceedings of the Genetic and Evolutionary Computation Conference"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3512290.3528845","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3512290.3528845","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:09:57Z","timestamp":1750183797000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3512290.3528845"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,8]]},"references-count":46,"alternative-id":["10.1145\/3512290.3528845","10.1145\/3512290"],"URL":"https:\/\/doi.org\/10.1145\/3512290.3528845","relation":{},"subject":[],"published":{"date-parts":[[2022,7,8]]},"assertion":[{"value":"2022-07-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}