{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,25]],"date-time":"2026-08-25T17:21:00Z","timestamp":1787678460135,"version":"build-2784847793"},"publisher-location":"New York, NY, USA","reference-count":37,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T00:00:00Z","timestamp":1783641600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,7,13]]},"DOI":"10.1145\/3795095.3805180","type":"proceedings-article","created":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T17:44:29Z","timestamp":1783705469000},"page":"348-355","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Evolutionary Discovery of Reinforcement Learning Algorithms via Large Language Models"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-4357-9533","authenticated-orcid":false,"given":"Alkis","family":"Sygkounas","sequence":"first","affiliation":[{"name":"\u00d6rebro University, \u00d6rebro, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3122-693X","authenticated-orcid":false,"given":"Amy","family":"Loutfi","sequence":"additional","affiliation":[{"name":"\u00d6rebro University, \u00d6rebro, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7649-9109","authenticated-orcid":false,"given":"Andreas","family":"Persson","sequence":"additional","affiliation":[{"name":"\u00d6rebro University, \u00d6rebro, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,10]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control. arXiv preprint","author":"Alzorgan Hazim","year":"2025","unstructured":"Hazim Alzorgan and Abolfazl Razi. 2025. Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control. arXiv preprint (2025). https:\/\/arxiv.org\/abs\/2505.09029v1 Reports HalfCheetah-v4 returns around 12750 for strong baselines."},{"key":"e_1_3_2_2_2_1","volume-title":"Learning to learn by gradient descent by gradient descent. Advances in neural information processing systems 29","author":"Andrychowicz Marcin","year":"2016","unstructured":"Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas. 2016. Learning to learn by gradient descent by gradient descent. Advances in neural information processing systems 29 (2016)."},{"key":"e_1_3_2_2_3_1","unstructured":"Erick Cant\u00fa-Paz et al. 1998. A survey of parallel genetic algorithms. Calculateurs paralleles reseaux et systems repartis 10 2 (1998) 141\u2013171."},{"key":"e_1_3_2_2_4_1","volume-title":"Evolving Reinforcement Learning Algorithms. In 9th International Conference on Learning Representations (ICLR). https:\/\/openreview.net\/forum?id=9XlAMdLMrB","author":"Co-Reyes John D.","year":"2021","unstructured":"John D. Co-Reyes, Xue Bin Peng, Sergey Levine, Pieter Abbeel, and John Schulman. 2021. Evolving Reinforcement Learning Algorithms. In 9th International Conference on Learning Representations (ICLR). https:\/\/openreview.net\/forum?id=9XlAMdLMrB"},{"key":"e_1_3_2_2_5_1","volume-title":"International conference on machine learning. PMLR, 9104\u20139149","author":"Eimer Theresa","year":"2023","unstructured":"Theresa Eimer, Marius Lindauer, and Roberta Raileanu. 2023. Hyperparameters in reinforcement learning and how to tune them. In International conference on machine learning. PMLR, 9104\u20139149."},{"key":"e_1_3_2_2_6_1","volume-title":"Implementation matters in deep policy gradients: A case study on ppo and trpo. arXiv preprint arXiv:2005.12729","author":"Engstrom Logan","year":"2020","unstructured":"Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry. 2020. Implementation matters in deep policy gradients: A case study on ppo and trpo. arXiv preprint arXiv:2005.12729 (2020)."},{"key":"e_1_3_2_2_7_1","volume-title":"Gymnasium: A Standard API for Reinforcement Learning. https:\/\/gymnasium.farama.org\/. Gymnasium documentation.","author":"Foundation Farama","year":"2023","unstructured":"Farama Foundation. 2023. Gymnasium: A Standard API for Reinforcement Learning. https:\/\/gymnasium.farama.org\/. Gymnasium documentation."},{"key":"e_1_3_2_2_8_1","volume-title":"Neural architecture evolution in deep reinforcement learning for continuous control. arXiv preprint arXiv:1910.12824","author":"Franke J\u00f6rg KH","year":"2019","unstructured":"J\u00f6rg KH Franke, Gregor K\u00f6hler, Noor Awad, and Frank Hutter. 2019. Neural architecture evolution in deep reinforcement learning for continuous control. arXiv preprint arXiv:1910.12824 (2019)."},{"key":"e_1_3_2_2_9_1","volume-title":"The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=cJPUpL8mOw","author":"Hazra Rishi","year":"2025","unstructured":"Rishi Hazra, Alkis Sygkounas, Andreas Persson, Amy Loutfi, and Pedro Zuidberg Dos Martires. 2025. REvolve: Reward Evolution with Large Language Models using Human Feedback. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=cJPUpL8mOw"},{"key":"e_1_3_2_2_10_1","volume-title":"Reinforcement learning with automated auxiliary loss search. Advances in neural information processing systems 35","author":"He Tairan","year":"2022","unstructured":"Tairan He, Yuge Zhang, Kan Ren, Minghuan Liu, Che Wang, Weinan Zhang, Yuqing Yang, and Dongsheng Li. 2022. Reinforcement learning with automated auxiliary loss search. Advances in neural information processing systems 35 (2022), 1820\u20131834."},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1007\/s10710-024-09494-2","article-title":"Evolving code with a large language model","volume":"25","author":"Hemberg Erik","year":"2024","unstructured":"Erik Hemberg, Stephen Moskal, and Una-May O'Reilly. 2024. Evolving code with a large language model. Genetic Programming and Evolvable Machines 25, 2 (2024), 21.","journal-title":"Genetic Programming and Evolvable Machines"},{"key":"e_1_3_2_2_12_1","volume-title":"Proceedings of the AAAI conference on artificial intelligence","volume":"32","author":"Henderson Peter","year":"2018","unstructured":"Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. 2018. Deep reinforcement learning that matters. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32."},{"key":"e_1_3_2_2_13_1","unstructured":"Max Jaderberg Valentin Dalibard Simon Osindero Wojciech M Czarnecki Jeff Donahue Ali Razavi Oriol Vinyals Tim Green Iain Dunning Karen Simonyan et al. 2017. Population based training of neural networks. arXiv preprint arXiv:1711.09846 (2017)."},{"key":"e_1_3_2_2_14_1","volume-title":"Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu.","author":"Jaderberg Max","year":"2016","unstructured":"Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu. 2016. Reinforcement learning with unsupervised auxiliary tasks. arXiv preprint arXiv:1611.05397 (2016)."},{"key":"e_1_3_2_2_15_1","volume-title":"Evolution-guided policy gradient in reinforcement learning. Advances in Neural Information Processing Systems 31","author":"Khadka Shauharda","year":"2018","unstructured":"Shauharda Khadka and Kagan Tumer. 2018. Evolution-guided policy gradient in reinforcement learning. Advances in Neural Information Processing Systems 31 (2018)."},{"key":"e_1_3_2_2_16_1","volume-title":"Actor-critic algorithms. Advances in neural information processing systems 12","author":"Konda Vijay","year":"1999","unstructured":"Vijay Konda and John Tsitsiklis. 1999. Actor-critic algorithms. Advances in neural information processing systems 12 (1999)."},{"key":"e_1_3_2_2_17_1","volume-title":"Genetic Programming: On the Programming of Computers by Means of Natural Selection","author":"Koza John R.","year":"1992","unstructured":"John R. Koza. 1992. Genetic Programming: On the Programming of Computers by Means of Natural Selection. MIT Press, Cambridge, MA, USA."},{"key":"e_1_3_2_2_18_1","volume-title":"Soviet physics-doklady","author":"Lcvenshtcin VI","unstructured":"VI Lcvenshtcin. 1966. Binary coors capable or 'correcting deletions, insertions, and reversals. In Soviet physics-doklady, Vol. 10."},{"key":"e_1_3_2_2_19_1","volume-title":"Proceedings of the 13th annual conference on Genetic and evolutionary computation. 211\u2013218","author":"Lehman Joel","year":"2011","unstructured":"Joel Lehman and Kenneth O Stanley. 2011. Evolving a diversity of virtual creatures through novelty search and local competition. In Proceedings of the 13th annual conference on Genetic and evolutionary computation. 211\u2013218."},{"key":"e_1_3_2_2_20_1","volume-title":"The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=IEduRUO55F","author":"Ma Yecheng Jason","year":"2024","unstructured":"Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2024. Eureka: Human-Level Reward Design via Coding Large Language Models. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=IEduRUO55F"},{"key":"e_1_3_2_2_21_1","volume-title":"proceedings of the Genetic and Evolutionary Computation Conference. 1110\u20131118","author":"Nasir Muhammad Umair","year":"2024","unstructured":"Muhammad Umair Nasir, Sam Earle, Julian Togelius, Steven James, and Christopher Cleghorn. 2024. Llmatic: neural architecture search via large language models and quality diversity optimization. In proceedings of the Genetic and Evolutionary Computation Conference. 1110\u20131118."},{"key":"e_1_3_2_2_22_1","first-page":"1060","article-title":"Discovering Reinforcement Learning Algorithms","volume":"33","author":"Oh Junhyuk","year":"2020","unstructured":"Junhyuk Oh, Rishabh Agarwal, Kelvin Xu, Dale Schuurmans, Quoc V. Le, and Mohammad Norouzi. 2020. Discovering Reinforcement Learning Algorithms. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 33. 1060\u20131070. https:\/\/papers.nips.cc\/paper_files\/paper\/2020\/hash\/a322852ce0df73e204b7dfbce9c8cfd0-Abstract.html","journal-title":"Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_3_2_2_23_1","unstructured":"OpenAI Gym. 2016. Leaderboard Solved Definitions (LunarLander-v2). https:\/\/github.com\/openai\/gym\/wiki\/Leaderboard. Lists average reward 200 as solved for LunarLander."},{"key":"e_1_3_2_2_24_1","unstructured":"OpenAI Gym. 2016. MountainCar-v0. https:\/\/github.com\/openai\/gym\/wiki\/mountaincar-v0. OpenAI Gym Wiki."},{"key":"e_1_3_2_2_25_1","first-page":"1","article-title":"Empirical design in reinforcement learning","volume":"25","author":"Patterson Andrew","year":"2024","unstructured":"Andrew Patterson, Samuel Neumann, Martha White, and Adam White. 2024. Empirical design in reinforcement learning. Journal of Machine Learning Research 25, 318 (2024), 1\u201363.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_2_26_1","volume-title":"Fast efficient hyper-parameter tuning for policy gradient methods. Advances in Neural Information Processing Systems 32","author":"Paul Supratik","year":"2019","unstructured":"Supratik Paul, Vitaly Kurin, and Shimon Whiteson. 2019. Fast efficient hyper-parameter tuning for policy gradient methods. Advances in Neural Information Processing Systems 32 (2019)."},{"key":"e_1_3_2_2_27_1","volume-title":"International conference on machine learning. PMLR, 4095\u20134104","author":"Pham Hieu","year":"2018","unstructured":"Hieu Pham, Melody Guan, Barret Zoph, Quoc Le, and Jeff Dean. 2018. Efficient neural architecture search via parameters sharing. In International conference on machine learning. PMLR, 4095\u20134104."},{"key":"e_1_3_2_2_28_1","volume-title":"Proceedings of the aaai conference on artificial intelligence","volume":"33","author":"Real Esteban","year":"2019","unstructured":"Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le. 2019. Regularized evolution for image classifier architecture search. In Proceedings of the aaai conference on artificial intelligence, Vol. 33. 4780\u20134789."},{"key":"e_1_3_2_2_29_1","volume-title":"International conference on machine learning. PMLR, 2902\u20132911","author":"Real Esteban","year":"2017","unstructured":"Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena, Yutaka Leon Suematsu, Jie Tan, Quoc V Le, and Alexey Kurakin. 2017. Large-scale evolution of image classifiers. In International conference on machine learning. PMLR, 2902\u20132911."},{"key":"e_1_3_2_2_30_1","volume-title":"Evolution strategies as a scalable alternative to reinforcement learning. arXiv preprint arXiv:1703.03864","author":"Salimans Tim","year":"2017","unstructured":"Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. 2017. Evolution strategies as a scalable alternative to reinforcement learning. arXiv preprint arXiv:1703.03864 (2017)."},{"key":"e_1_3_2_2_31_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Sato Rei","year":"2021","unstructured":"Rei Sato, Jun Sakuma, and Youhei Akimoto. 2021. Advantagenas: Efficient neural architecture search with credit assignment. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 9489\u20139496."},{"key":"e_1_3_2_2_32_1","volume-title":"Evolving neural networks through augmenting topologies. Evolutionary computation 10, 2","author":"Stanley Kenneth O","year":"2002","unstructured":"Kenneth O Stanley and Risto Miikkulainen. 2002. Evolving neural networks through augmenting topologies. Evolutionary computation 10, 2 (2002), 99\u2013127."},{"key":"e_1_3_2_2_33_1","volume-title":"Policy gradient methods for reinforcement learning with function approximation. Advances in neural information processing systems 12","author":"Sutton Richard S","year":"1999","unstructured":"Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour. 1999. Policy gradient methods for reinforcement learning with function approximation. Advances in neural information processing systems 12 (1999)."},{"key":"e_1_3_2_2_34_1","volume-title":"Paired open-ended trailblazer (poet): Endlessly generating increasingly complex and diverse learning environments and their solutions. arXiv preprint arXiv:1901.01753","author":"Wang Rui","year":"2019","unstructured":"Rui Wang, Joel Lehman, Jeff Clune, and Kenneth O Stanley. 2019. Paired open-ended trailblazer (poet): Endlessly generating increasingly complex and diverse learning environments and their solutions. arXiv preprint arXiv:1901.01753 (2019)."},{"key":"e_1_3_2_2_35_1","volume-title":"Meta-gradient reinforcement learning. Advances in neural information processing systems 31","author":"Xu Zhongwen","year":"2018","unstructured":"Zhongwen Xu, Hado P van Hasselt, and David Silver. 2018. Meta-gradient reinforcement learning. Advances in neural information processing systems 31 (2018)."},{"key":"e_1_3_2_2_36_1","volume-title":"Neural Architecture Search with Reinforcement Learning. In 5th International Conference on Learning Representations (ICLR). https:\/\/openreview.net\/forum?id=rlUe8Hcxg Oral Presentation.","author":"Zoph Barret","unstructured":"Barret Zoph and Quoc V. Le. 2017. Neural Architecture Search with Reinforcement Learning. In 5th International Conference on Learning Representations (ICLR). https:\/\/openreview.net\/forum?id=rlUe8Hcxg Oral Presentation."},{"key":"e_1_3_2_2_37_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence, Vol 35","author":"Zou Haosheng","year":"2021","unstructured":"Haosheng Zou, Tongzheng Ren, Dong Yan, Hang Su, and Jun Zhu. 2021. Learning task-distribution reward shaping with meta-learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol 35. 11210\u201311218."}],"event":{"name":"GECCO '26: Genetic and Evolutionary Computation Conference","location":"Centro Internacional de Convenciones CIC-ANDE San Jose Costa Rica","acronym":"GECCO '26","sponsor":["SIGEVO ACM Special Interest Group on Genetic and Evolutionary Computation"]},"container-title":["Proceedings of the Genetic and Evolutionary Computation Conference"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3795095.3805180","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,25]],"date-time":"2026-08-25T16:40:50Z","timestamp":1787676050000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3795095.3805180"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,10]]},"references-count":37,"alternative-id":["10.1145\/3795095.3805180","10.1145\/3795095"],"URL":"https:\/\/doi.org\/10.1145\/3795095.3805180","relation":{},"subject":[],"published":{"date-parts":[[2026,7,10]]},"assertion":[{"value":"2026-07-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}