{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T13:49:48Z","timestamp":1780408188631,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":61,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,7,8]],"date-time":"2022-07-08T00:00:00Z","timestamp":1657238400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["1053128,DGE-1842487"],"award-info":[{"award-number":["1053128,DGE-1842487"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,7,8]]},"DOI":"10.1145\/3512290.3528705","type":"proceedings-article","created":{"date-parts":[[2022,7,18]],"date-time":"2022-07-18T13:59:57Z","timestamp":1658152797000},"page":"1102-1111","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":25,"title":["Approximating gradients for differentiable quality diversity in reinforcement learning"],"prefix":"10.1145","author":[{"given":"Bryon","family":"Tjanaka","sequence":"first","affiliation":[{"name":"University of Southern California"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Matthew C.","family":"Fontaine","sequence":"additional","affiliation":[{"name":"University of Southern California"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Julian","family":"Togelius","sequence":"additional","affiliation":[{"name":"New York University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stefanos","family":"Nikolaidis","sequence":"additional","affiliation":[{"name":"University of Southern California"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,7,8]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Parallel Problem Solving from Nature, PPSN XI, Robert Schaefer, Carlos Cotta, Joanna Ko\u0142odziej, and G\u00fcnter Rudolph (Eds.)","author":"Akimoto Youhei","unstructured":"Youhei Akimoto , Yuichi Nagata , Isao Ono , and Shigenobu Kobayashi . 2010. Bidirectional Relation between CMA Evolution Strategies and Natural Evolution Strategies . In Parallel Problem Solving from Nature, PPSN XI, Robert Schaefer, Carlos Cotta, Joanna Ko\u0142odziej, and G\u00fcnter Rudolph (Eds.) . Springer Berlin Heidelberg , Berlin, Heidelberg , 154--163. Youhei Akimoto, Yuichi Nagata, Isao Ono, and Shigenobu Kobayashi. 2010. Bidirectional Relation between CMA Evolution Strategies and Natural Evolution Strategies. In Parallel Problem Solving from Nature, PPSN XI, Robert Schaefer, Carlos Cotta, Joanna Ko\u0142odziej, and G\u00fcnter Rudolph (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 154--163."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1162\/089976698300017746"},{"key":"e_1_3_2_2_3_1","volume-title":"Garnett (Eds.)","volume":"30","author":"Andrychowicz Marcin","year":"2017","unstructured":"Marcin Andrychowicz , Filip Wolski , Alex Ray , Jonas Schneider , Rachel Fong , Peter Welinder , Bob McGrew , Josh Tobin , Open AI Pieter Abbeel , and Wojciech Zaremba . 2017 . Hindsight Experience Replay. In Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R . Garnett (Eds.) , Vol. 30 . Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper\/ 2017\/file\/453fadbd8a1a3af50a9df4df899537b5-Paper.pdf Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba. 2017. Hindsight Experience Replay. In Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper\/2017\/file\/453fadbd8a1a3af50a9df4df899537b5-Paper.pdf"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1015059928466"},{"key":"e_1_3_2_2_5_1","volume-title":"Parallel Problem Solving from Nature, PPSN XI, Robert Schaefer, Carlos Cotta, Joanna Ko\u0142odziej, and G\u00fcnter Rudolph (Eds.)","author":"Brockhoff Dimo","unstructured":"Dimo Brockhoff , Anne Auger , Nikolaus Hansen , Dirk V. Arnold , and Tim Hohm . 2010. Mirrored Sampling and Sequential Selection for Evolution Strategies . In Parallel Problem Solving from Nature, PPSN XI, Robert Schaefer, Carlos Cotta, Joanna Ko\u0142odziej, and G\u00fcnter Rudolph (Eds.) . Springer Berlin Heidelberg , Berlin, Heidelberg , 11--21. Dimo Brockhoff, Anne Auger, Nikolaus Hansen, Dirk V. Arnold, and Tim Hohm. 2010. Mirrored Sampling and Sequential Selection for Evolution Strategies. In Parallel Problem Solving from Nature, PPSN XI, Robert Schaefer, Carlos Cotta, Joanna Ko\u0142odziej, and G\u00fcnter Rudolph (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 11--21."},{"key":"e_1_3_2_2_6_1","volume-title":"CoRR abs\/1606.01540","author":"Brockman Greg","year":"2016","unstructured":"Greg Brockman , Vicki Cheung , Ludwig Pettersson , Jonas Schneider , John Schulman , Jie Tang , and Wojciech Zaremba . 2016. Open AI Gym . CoRR abs\/1606.01540 ( 2016 ). arXiv:1606.01540 http:\/\/arxiv.org\/abs\/1606.01540 Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016. OpenAI Gym. CoRR abs\/1606.01540 (2016). arXiv:1606.01540 http:\/\/arxiv.org\/abs\/1606.01540"},{"key":"e_1_3_2_2_7_1","volume-title":"QD-RL: Efficient Mixing of Quality and Diversity in Reinforcement Learning. CoRR abs\/2006.08505","author":"Cideron Geoffrey","year":"2020","unstructured":"Geoffrey Cideron , Thomas Pierrot , Nicolas Perrin , Karim Beguir , and Olivier Sigaud . 2020. QD-RL: Efficient Mixing of Quality and Diversity in Reinforcement Learning. CoRR abs\/2006.08505 ( 2020 ). arXiv:2006.08505 https:\/\/arxiv.org\/abs\/2006.08505 Geoffrey Cideron, Thomas Pierrot, Nicolas Perrin, Karim Beguir, and Olivier Sigaud. 2020. QD-RL: Efficient Mixing of Quality and Diversity in Reinforcement Learning. CoRR abs\/2006.08505 (2020). arXiv:2006.08505 https:\/\/arxiv.org\/abs\/2006.08505"},{"key":"e_1_3_2_2_8_1","unstructured":"Jack Clark and Dario Amodei. 2016. Faulty Reward Functions in the Wild. https:\/\/openai.com\/blog\/faulty-reward-functions\/.  Jack Clark and Dario Amodei. 2016. Faulty Reward Functions in the Wild. https:\/\/openai.com\/blog\/faulty-reward-functions\/."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377930.3390217"},{"key":"e_1_3_2_2_10_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"1048","author":"Colas C\u00e9dric","year":"2018","unstructured":"C\u00e9dric Colas , Olivier Sigaud , and Pierre-Yves Oudeyer . 2018 . GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms . In Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research , Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, 1039-- 1048 . https:\/\/proceedings.mlr.press\/v80\/colas18a.html C\u00e9dric Colas, Olivier Sigaud, and Pierre-Yves Oudeyer. 2018. GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms. In Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, 1039--1048. https:\/\/proceedings.mlr.press\/v80\/colas18a.html"},{"key":"e_1_3_2_2_11_1","volume-title":"Joel Lehman, Kenneth Stanley, and Jeff Clune.","author":"Conti Edoardo","year":"2018","unstructured":"Edoardo Conti , Vashisht Madhavan , Felipe Petroski Such , Joel Lehman, Kenneth Stanley, and Jeff Clune. 2018 . Improving Exploration in Evolution Strategies for Deep Reinforcement Learning via a Population of Novelty-Seeking Agents. In Advances in Neural Information Processing Systems 31, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.). Curran Associates, Inc ., 5027--5038. http:\/\/papers.nips.cc\/paper\/7750-improving-exploration-in-evolution-strategies-for-deep-reinforcement-learning-via-a-population-of-novelty-seeking-agents.pdf Edoardo Conti, Vashisht Madhavan, Felipe Petroski Such, Joel Lehman, Kenneth Stanley, and Jeff Clune. 2018. Improving Exploration in Evolution Strategies for Deep Reinforcement Learning via a Population of Novelty-Seeking Agents. In Advances in Neural Information Processing Systems 31, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.). Curran Associates, Inc., 5027--5038. http:\/\/papers.nips.cc\/paper\/7750-improving-exploration-in-evolution-strategies-for-deep-reinforcement-learning-via-a-population-of-novelty-seeking-agents.pdf"},{"key":"e_1_3_2_2_12_1","unstructured":"Erwin Coumans and Yunfei Bai. 2016--2020. PyBullet a Python module for physics simulation for games robotics and machine learning. http:\/\/pybullet.org.  Erwin Coumans and Yunfei Bai. 2016--2020. PyBullet a Python module for physics simulation for games robotics and machine learning. http:\/\/pybullet.org."},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature14422"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10479-005-5724-z"},{"key":"e_1_3_2_2_15_1","unstructured":"Benjamin Ellenberger. 2018--2019. PyBullet Gymperium. https:\/\/github.com\/benelot\/pybullet-gym.  Benjamin Ellenberger. 2018--2019. PyBullet Gymperium. https:\/\/github.com\/benelot\/pybullet-gym."},{"key":"e_1_3_2_2_16_1","volume-title":"International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=SJx63jRqFm","author":"Eysenbach Benjamin","year":"2019","unstructured":"Benjamin Eysenbach , Abhishek Gupta , Julian Ibarz , and Sergey Levine . 2019 . Diversity is All You Need: Learning Skills without a Reward Function . In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=SJx63jRqFm Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine. 2019. Diversity is All You Need: Learning Skills without a Reward Function. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=SJx63jRqFm"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2021.XVII.036"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2021.XVII.038"},{"key":"e_1_3_2_2_19_1","volume-title":"Fontaine and Stefanos Nikolaidis","author":"Matthew","year":"2021","unstructured":"Matthew C. Fontaine and Stefanos Nikolaidis . 2021 . Differentiable Quality Diversity. Advances in Neural Information Processing Systems 34 (2021). https:\/\/proceedings.neurips.cc\/paper\/2021\/file\/532923f11ac97d3e7cb0130315b067dc-Paper.pdf Matthew C. Fontaine and Stefanos Nikolaidis. 2021. Differentiable Quality Diversity. Advances in Neural Information Processing Systems 34 (2021). https:\/\/proceedings.neurips.cc\/paper\/2021\/file\/532923f11ac97d3e7cb0130315b067dc-Paper.pdf"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377930.3390232"},{"key":"e_1_3_2_2_21_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"1596","author":"Fujimoto Scott","year":"2018","unstructured":"Scott Fujimoto , Herke van Hoof , and David Meger . 2018 . Addressing Function Approximation Error in Actor-Critic Methods . In Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research , Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, 1587-- 1596 . http:\/\/proceedings.mlr.press\/v80\/fujimoto18a.html Scott Fujimoto, Herke van Hoof, and David Meger. 2018. Addressing Function Approximation Error in Actor-Critic Methods. In Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, 1587--1596. http:\/\/proceedings.mlr.press\/v80\/fujimoto18a.html"},{"key":"e_1_3_2_2_22_1","volume-title":"Data-efficient design exploration through surrogate-assisted illumination. Evolutionary computation 26, 3","author":"Gaier Adam","year":"2018","unstructured":"Adam Gaier , Alexander Asteroth , and Jean-Baptiste Mouret . 2018. Data-efficient design exploration through surrogate-assisted illumination. Evolutionary computation 26, 3 ( 2018 ), 381--410. Adam Gaier, Alexander Asteroth, and Jean-Baptiste Mouret. 2018. Data-efficient design exploration through surrogate-assisted illumination. Evolutionary computation 26, 3 (2018), 381--410."},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377930.3390221"},{"key":"e_1_3_2_2_24_1","volume-title":"Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research","volume":"256","author":"Glorot Xavier","year":"2010","unstructured":"Xavier Glorot and Yoshua Bengio . 2010 . Understanding the difficulty of training deep feedforward neural networks . In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research , Vol. 9), Yee Whye Teh and Mike Titterington (Eds.). PMLR, Chia Laguna Resort, Sardinia, Italy, 249-- 256 . https:\/\/proceedings.mlr.press\/v9\/glorot10a.html Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research, Vol. 9), Yee Whye Teh and Mike Titterington (Eds.). PMLR, Chia Laguna Resort, Sardinia, Italy, 249--256. https:\/\/proceedings.mlr.press\/v9\/glorot10a.html"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CIG.2019.8848053"},{"key":"e_1_3_2_2_26_1","volume-title":"A Visual Guide to Evolution Strategies. blog.otoro.net","author":"David Ha.","year":"2017","unstructured":"David Ha. 2017. A Visual Guide to Evolution Strategies. blog.otoro.net ( 2017 ). https:\/\/blog.otoro.net\/2017\/10\/29\/visual-evolution-strategies\/ David Ha. 2017. A Visual Guide to Evolution Strategies. blog.otoro.net (2017). https:\/\/blog.otoro.net\/2017\/10\/29\/visual-evolution-strategies\/"},{"key":"e_1_3_2_2_27_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"1870","author":"Haarnoja Tuomas","year":"2018","unstructured":"Tuomas Haarnoja , Aurick Zhou , Pieter Abbeel , and Sergey Levine . 2018 . Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor . In Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research , Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, 1861-- 1870 . https:\/\/proceedings.mlr.press\/v80\/haarnoja18b.html Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, 1861--1870. https:\/\/proceedings.mlr.press\/v80\/haarnoja18b.html"},{"key":"e_1_3_2_2_28_1","volume-title":"The CMA Evolution Strategy: A Tutorial. CoRR abs\/1604.00772","author":"Hansen Nikolaus","year":"2016","unstructured":"Nikolaus Hansen . 2016. The CMA Evolution Strategy: A Tutorial. CoRR abs\/1604.00772 ( 2016 ). arXiv:1604.00772 http:\/\/arxiv.org\/abs\/1604.00772 Nikolaus Hansen. 2016. The CMA Evolution Strategy: A Tutorial. CoRR abs\/1604.00772 (2016). arXiv:1604.00772 http:\/\/arxiv.org\/abs\/1604.00772"},{"key":"e_1_3_2_2_29_1","unstructured":"Alex Irpan. 2018. Deep Reinforcement Learning Doesn't Work Yet. https:\/\/www.alexirpan.com\/2018\/02\/14\/rl-hard.html.  Alex Irpan. 2018. Deep Reinforcement Learning Doesn't Work Yet. https:\/\/www.alexirpan.com\/2018\/02\/14\/rl-hard.html."},{"key":"e_1_3_2_2_30_1","volume-title":"Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"3350","author":"Khadka Shauharda","year":"2019","unstructured":"Shauharda Khadka , Somdeb Majumdar , Tarek Nassar , Zach Dwiel , Evren Tumer , Santiago Miret , Yinyin Liu , and Kagan Tumer . 2019 . Collaborative Evolutionary Reinforcement Learning . In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research , Vol. 97), Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, 3341-- 3350 . https:\/\/proceedings.mlr.press\/v97\/khadka19a.html Shauharda Khadka, Somdeb Majumdar, Tarek Nassar, Zach Dwiel, Evren Tumer, Santiago Miret, Yinyin Liu, and Kagan Tumer. 2019. Collaborative Evolutionary Reinforcement Learning. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97), Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, 3341--3350. https:\/\/proceedings.mlr.press\/v97\/khadka19a.html"},{"key":"e_1_3_2_2_31_1","volume-title":"Garnett (Eds.)","volume":"31","author":"Khadka Shauharda","year":"2018","unstructured":"Shauharda Khadka and Kagan Tumer . 2018 . Evolution-Guided Policy Gradient in Reinforcement Learning. In Advances in Neural Information Processing Systems, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R . Garnett (Eds.) , Vol. 31 . Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper\/ 2018\/file\/85fc37b18c57097425b52fc7afbb6969-Paper.pdf Shauharda Khadka and Kagan Tumer. 2018. Evolution-Guided Policy Gradient in Reinforcement Learning. In Advances in Neural Information Processing Systems, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.), Vol. 31. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper\/2018\/file\/85fc37b18c57097425b52fc7afbb6969-Paper.pdf"},{"key":"e_1_3_2_2_32_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba . 2015 . Adam : A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7--9, 2015, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds .). http:\/\/arxiv.org\/abs\/1412.6980 Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7--9, 2015, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_2_2_33_1","volume-title":"One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RL. Advances in Neural Information Processing Systems 33","author":"Kumar Saurabh","year":"2020","unstructured":"Saurabh Kumar , Aviral Kumar , Sergey Levine , and Chelsea Finn . 2020. One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RL. Advances in Neural Information Processing Systems 33 ( 2020 ). Saurabh Kumar, Aviral Kumar, Sergey Levine, and Chelsea Finn. 2020. One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RL. Advances in Neural Information Processing Systems 33 (2020)."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3205455.3205474"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1162\/EVCO_a_00025"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2001576.2001606"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-70139-4"},{"key":"e_1_3_2_2_38_1","volume-title":"4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2--4, 2016, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1509","author":"Lillicrap Timothy P.","year":"2016","unstructured":"Timothy P. Lillicrap , Jonathan J. Hunt , Alexander Pritzel , Nicolas Heess , Tom Erez , Yuval Tassa , David Silver , and Daan Wierstra . 2016 . Continuous control with deep reinforcement learning . In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2--4, 2016, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1509 .02971 Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2016. Continuous control with deep reinforcement learning. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2--4, 2016, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1509.02971"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.5555\/3326943.3327109"},{"key":"e_1_3_2_2_40_1","volume-title":"Illuminating search spaces by mapping elites. CoRR abs\/1504.04909","author":"Mouret Jean-Baptiste","year":"2015","unstructured":"Jean-Baptiste Mouret and Jeff Clune . 2015. Illuminating search spaces by mapping elites. CoRR abs\/1504.04909 ( 2015 ). arXiv:1504.04909 http:\/\/arxiv.org\/abs\/1504.04909 Jean-Baptiste Mouret and Jeff Clune. 2015. Illuminating search spaces by mapping elites. CoRR abs\/1504.04909 (2015). arXiv:1504.04909 http:\/\/arxiv.org\/abs\/1504.04909"},{"key":"e_1_3_2_2_41_1","volume-title":"Proceedings of the Sixteenth International Conference on Machine Learning (ICML '99)","author":"Ng Andrew Y.","unstructured":"Andrew Y. Ng , Daishi Harada , and Stuart J. Russell . 1999. Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping . In Proceedings of the Sixteenth International Conference on Machine Learning (ICML '99) . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 278--287. Andrew Y. Ng, Daishi Harada, and Stuart J. Russell. 1999. Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping. In Proceedings of the Sixteenth International Conference on Machine Learning (ICML '99). Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 278--287."},{"key":"e_1_3_2_2_42_1","unstructured":"Olle Nilsson. 2021. QDgym. https:\/\/github.com\/ollenilsson19\/QDgym.  Olle Nilsson. 2021. QDgym. https:\/\/github.com\/ollenilsson19\/QDgym."},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3449639.3459304"},{"key":"e_1_3_2_2_44_1","volume-title":"Solving Rubik's Cube with a Robot Hand. arXiv preprint","author":"Ilge Akkaya AI","year":"2019","unstructured":"Open AI , Ilge Akkaya , Marcin Andrychowicz , Maciek Chociej , Mateusz Litwin , Bob McGrew , Arthur Petron , Alex Paino , Matthias Plappert , Glenn Powell , Raphael Ribas , Jonas Schneider , Nikolas Tezak , Jerry Tworek , Peter Welinder , Lilian Weng , Qiming Yuan , Wojciech Zaremba , and Lei Zhang . 2019. Solving Rubik's Cube with a Robot Hand. arXiv preprint ( 2019 ). OpenAI, Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, Jonas Schneider, Nikolas Tezak, Jerry Tworek, Peter Welinder, Lilian Weng, Qiming Yuan, Wojciech Zaremba, and Lei Zhang. 2019. Solving Rubik's Cube with a Robot Hand. arXiv preprint (2019)."},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.3389\/frobt.2020.00098"},{"key":"e_1_3_2_2_46_1","volume-title":"Lin (Eds.)","volume":"33","author":"Parker-Holder Jack","year":"2020","unstructured":"Jack Parker-Holder , Aldo Pacchiano , Krzysztof M Choromanski , and Stephen J Roberts . 2020 . Effective Diversity in Population Based Reinforcement Learning. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H . Lin (Eds.) , Vol. 33 . Curran Associates, Inc. , 18050--18062. https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/d1dc3a8270a6f9394f88847d7f0050cf-Paper.pdf Jack Parker-Holder, Aldo Pacchiano, Krzysztof M Choromanski, and Stephen J Roberts. 2020. Effective Diversity in Population Based Reinforcement Learning. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 18050--18062. https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/d1dc3a8270a6f9394f88847d7f0050cf-Paper.pdf"},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8460528"},{"key":"e_1_3_2_2_48_1","volume-title":"International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=BkeU5j0ctQ","author":"Sigaud Pourchot","year":"2019","unstructured":"Pourchot and Sigaud . 2019 . CEM-RL: Combining evolutionary and gradient-based methods for policy search . In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=BkeU5j0ctQ Pourchot and Sigaud. 2019. CEM-RL: Combining evolutionary and gradient-based methods for policy search. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=BkeU5j0ctQ"},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.3389\/frobt.2016.00040"},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3449639.3459320"},{"key":"e_1_3_2_2_51_1","unstructured":"Tim Salimans Jonathan Ho Xi Chen Szymon Sidor and Ilya Sutskever. 2017. Evolution Strategies as a Scalable Alternative to Reinforcement Learning. arXiv:1703.03864 [stat.ML]  Tim Salimans Jonathan Ho Xi Chen Szymon Sidor and Ilya Sutskever. 2017. Evolution Strategies as a Scalable Alternative to Reinforcement Learning. arXiv:1703.03864 [stat.ML]"},{"key":"e_1_3_2_2_52_1","volume-title":"Proceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"1320","author":"Schaul Tom","year":"2015","unstructured":"Tom Schaul , Daniel Horgan , Karol Gregor , and David Silver . 2015 . Universal Value Function Approximators . In Proceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research , Vol. 37), Francis Bach and David Blei (Eds.). PMLR, Lille, France, 1312-- 1320 . https:\/\/proceedings.mlr.press\/v37\/schaul15.html Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver. 2015. Universal Value Function Approximators. In Proceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 37), Francis Bach and David Blei (Eds.). PMLR, Lille, France, 1312--1320. https:\/\/proceedings.mlr.press\/v37\/schaul15.html"},{"key":"e_1_3_2_2_53_1","volume-title":"Proceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"1897","author":"Schulman John","year":"2015","unstructured":"John Schulman , Sergey Levine , Pieter Abbeel , Michael Jordan , and Philipp Moritz . 2015 . Trust Region Policy Optimization . In Proceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research , Vol. 37), Francis Bach and David Blei (Eds.). PMLR, Lille, France, 1889-- 1897 . https:\/\/proceedings.mlr.press\/v37\/schulman15.html John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015. Trust Region Policy Optimization. In Proceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 37), Francis Bach and David Blei (Eds.). PMLR, Lille, France, 1889--1897. https:\/\/proceedings.mlr.press\/v37\/schulman15.html"},{"key":"e_1_3_2_2_54_1","volume-title":"Proximal Policy Optimization Algorithms. CoRR abs\/1707.06347","author":"Schulman John","year":"2017","unstructured":"John Schulman , Filip Wolski , Prafulla Dhariwal , Alec Radford , and Oleg Klimov . 2017. Proximal Policy Optimization Algorithms. CoRR abs\/1707.06347 ( 2017 ). arXiv:1707.06347 http:\/\/arxiv.org\/abs\/1707.06347 John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. CoRR abs\/1707.06347 (2017). arXiv:1707.06347 http:\/\/arxiv.org\/abs\/1707.06347"},{"key":"e_1_3_2_2_55_1","volume-title":"Barto","author":"Sutton Richard S.","year":"2018","unstructured":"Richard S. Sutton and Andrew G . Barto . 2018 . Reinforcement Learning : An Introduction (second ed.). The MIT Press . http:\/\/incompleteideas.net\/book\/the-book-2nd.html Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction (second ed.). The MIT Press. http:\/\/incompleteideas.net\/book\/the-book-2nd.html"},{"key":"e_1_3_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.5555\/3463952.3464104"},{"key":"e_1_3_2_2_57_1","unstructured":"Bryon Tjanaka Matthew C. Fontaine Yulun Zhang Sam Sommerer Nathan Dennler and Stefanos Nikolaidis. 2021. pyribs: A bare-bones Python library for quality diversity optimization. https:\/\/github.com\/icaros-usc\/pyribs.  Bryon Tjanaka Matthew C. Fontaine Yulun Zhang Sam Sommerer Nathan Dennler and Stefanos Nikolaidis. 2021. pyribs: A bare-bones Python library for quality diversity optimization. https:\/\/github.com\/icaros-usc\/pyribs."},{"key":"e_1_3_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2017.8202133"},{"key":"e_1_3_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3205455.3205602"},{"key":"e_1_3_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2638566"},{"key":"e_1_3_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/CEC.2008.4631255"}],"event":{"name":"GECCO '22: Genetic and Evolutionary Computation Conference","location":"Boston Massachusetts","acronym":"GECCO '22","sponsor":["SIGEVO ACM Special Interest Group on Genetic and Evolutionary Computation"]},"container-title":["Proceedings of the Genetic and Evolutionary Computation Conference"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3512290.3528705","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3512290.3528705","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3512290.3528705","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:29Z","timestamp":1750186829000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3512290.3528705"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,8]]},"references-count":61,"alternative-id":["10.1145\/3512290.3528705","10.1145\/3512290"],"URL":"https:\/\/doi.org\/10.1145\/3512290.3528705","relation":{},"subject":[],"published":{"date-parts":[[2022,7,8]]},"assertion":[{"value":"2022-07-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}