{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,5]],"date-time":"2026-03-05T16:19:48Z","timestamp":1772727588858,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":36,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,9,15]],"date-time":"2020-09-15T00:00:00Z","timestamp":1600128000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,9,15]]},"DOI":"10.1145\/3402942.3402944","type":"proceedings-article","created":{"date-parts":[[2020,9,17]],"date-time":"2020-09-17T16:38:30Z","timestamp":1600360710000},"page":"1-10","source":"Crossref","is-referenced-by-count":13,"title":["Strategies for Using Proximal Policy Optimization in Mobile Puzzle Games"],"prefix":"10.1145","author":[{"given":"Jeppe Theiss","family":"Kristensen","sequence":"first","affiliation":[{"name":"IT University of Copenhagen, Denmark"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paolo","family":"Burelli","sequence":"additional","affiliation":[{"name":"IT University of Copenhagen, Denmark"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,9,17]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Zafarali Ahmed Nicolas\u00a0Le Roux Mohammad Norouzi and Dale Schuurmans. 2018. Understanding the impact of entropy on policy optimization. (2018). arXiv:arXiv:1811.11214 Zafarali Ahmed Nicolas\u00a0Le Roux Mohammad Norouzi and Dale Schuurmans. 2018. Understanding the impact of entropy on policy optimization. (2018). arXiv:arXiv:1811.11214"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"crossref","unstructured":"Mohammed Alshiekh Roderick Bloem Ruediger Ehlers Bettina K\u00f6nighofer Scott Niekum and Ufuk Topcu. 2017. Safe Reinforcement Learning via Shielding. (2017). arXiv:arXiv:1708.08611 Mohammed Alshiekh Roderick Bloem Ruediger Ehlers Bettina K\u00f6nighofer Scott Niekum and Ufuk Topcu. 2017. Safe Reinforcement Learning via Shielding. (2017). arXiv:arXiv:1708.08611","DOI":"10.1609\/aaai.v32i1.11797"},{"key":"e_1_3_2_1_3_1","volume-title":"IJCAI International Joint Conference on Artificial Intelligence 2015-Janua(2015)","author":"Bellemare G.","year":"2015"},{"key":"e_1_3_2_1_4_1","unstructured":"Petros Christodoulou. 2019. Soft Actor-Critic for Discrete Action Settings. (2019). arXiv:arXiv:1910.07207 Petros Christodoulou. 2019. Soft Actor-Critic for Discrete Action Settings. (2019). arXiv:arXiv:1910.07207"},{"key":"e_1_3_2_1_5_1","unstructured":"Karl Cobbe Oleg Klimov Chris Hesse Taehoon Kim and John Schulman. 2018. Quantifying Generalization in Reinforcement Learning. arXiv:arXiv:1812.02341 Karl Cobbe Oleg Klimov Chris Hesse Taehoon Kim and John Schulman. 2018. Quantifying Generalization in Reinforcement Learning. arXiv:arXiv:1812.02341"},{"key":"e_1_3_2_1_6_1","unstructured":"Prafulla Dhariwal Christopher Hesse Oleg Klimov Alex Nichol Matthias Plappert Alec Radford John Schulman Szymon Sidor Yuhuai Wu and Peter Zhokhov. 2017. OpenAI Baselines. https:\/\/github.com\/openai\/baselines. Prafulla Dhariwal Christopher Hesse Oleg Klimov Alex Nichol Matthias Plappert Alec Radford John Schulman Szymon Sidor Yuhuai Wu and Peter Zhokhov. 2017. OpenAI Baselines. https:\/\/github.com\/openai\/baselines."},{"key":"e_1_3_2_1_7_1","unstructured":"Jesse Farebrother Marlos\u00a0C. Machado and Michael Bowling. 2018. Generalization and Regularization in DQN. (2018). arXiv:arXiv:1810.00123 Jesse Farebrother Marlos\u00a0C. Machado and Michael Bowling. 2018. Generalization and Regularization in DQN. (2018). arXiv:arXiv:1810.00123"},{"key":"e_1_3_2_1_8_1","volume-title":"Automated Curriculum Learning for Neural Networks. 34th International Conference on Machine Learning, ICML 2017 3","author":"Graves Alex","year":"2017"},{"key":"e_1_3_2_1_9_1","volume-title":"Human-Like Playtesting with Deep Learning. In 2018 IEEE Conference on Computational Intelligence and Games (CIG). 1\u20138.","author":"Gudmundsson Stefan\u00a0Freyr","year":"2018"},{"key":"e_1_3_2_1_10_1","volume-title":"35th International Conference on Machine Learning, ICML 2018 5 (1 2018","author":"Haarnoja Tuomas","year":"2018"},{"key":"e_1_3_2_1_11_1","unstructured":"Perttu H\u00e4m\u00e4l\u00e4inen Amin Babadi Xiaoxiao Ma and Jaakko Lehtinen. 2018. PPO-CMA: Proximal Policy Optimization with Covariance Matrix Adaptation. (2018). arXiv:arXiv:1810.02541 Perttu H\u00e4m\u00e4l\u00e4inen Amin Babadi Xiaoxiao Ma and Jaakko Lehtinen. 2018. PPO-CMA: Proximal Policy Optimization with Covariance Matrix Adaptation. (2018). arXiv:arXiv:1810.02541"},{"key":"e_1_3_2_1_12_1","unstructured":"Nicolas Heess Dhruva TB Srinivasan Sriram Jay Lemmon Josh Merel Greg Wayne Yuval Tassa Tom Erez Ziyu Wang S.\u00a0M.\u00a0Ali Eslami Martin Riedmiller and David Silver. 2017. Emergence of Locomotion Behaviours in Rich Environments. (2017). arXiv:arXiv:1707.02286 Nicolas Heess Dhruva TB Srinivasan Sriram Jay Lemmon Josh Merel Greg Wayne Yuval Tassa Tom Erez Ziyu Wang S.\u00a0M.\u00a0Ali Eslami Martin Riedmiller and David Silver. 2017. Emergence of Locomotion Behaviours in Rich Environments. (2017). arXiv:arXiv:1707.02286"},{"key":"e_1_3_2_1_13_1","volume-title":"32nd AAAI Conference on Artificial Intelligence, AAAI 2018","author":"Hessel Matteo","year":"2018"},{"key":"e_1_3_2_1_14_1","unstructured":"Ashley Hill Antonin Raffin Maximilian Ernestus Adam Gleave Anssi Kanervisto Rene Traore Prafulla Dhariwal Christopher Hesse Oleg Klimov Alex Nichol Matthias Plappert Alec Radford John Schulman Szymon Sidor and Yuhuai Wu. 2018. Stable Baselines. https:\/\/github.com\/hill-a\/stable-baselines. Ashley Hill Antonin Raffin Maximilian Ernestus Adam Gleave Anssi Kanervisto Rene Traore Prafulla Dhariwal Christopher Hesse Oleg Klimov Alex Nichol Matthias Plappert Alec Radford John Schulman Szymon Sidor and Yuhuai Wu. 2018. Stable Baselines. https:\/\/github.com\/hill-a\/stable-baselines."},{"key":"e_1_3_2_1_15_1","volume-title":"Generative adversarial imitation learning. Advances in Neural Information Processing Systems","author":"Ho Jonathan","year":"2016"},{"key":"e_1_3_2_1_16_1","volume-title":"Are Deep Policy Gradient Algorithms Truly Policy Gradient Algorithms? (11","author":"Ilyas Andrew","year":"2018"},{"key":"e_1_3_2_1_17_1","volume-title":"Unity: A General Platform for Intelligent Agents.","author":"Juliani Arthur","year":"2018"},{"key":"e_1_3_2_1_18_1","volume-title":"Deep Learning for Video Game Playing","author":"Justesen Niels","year":"2019"},{"key":"e_1_3_2_1_19_1","volume-title":"Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation. (6","author":"Justesen Niels","year":"2018"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CIG.2019.8848003"},{"key":"e_1_3_2_1_21_1","unstructured":"Zachary Kenton Angelos Filos Owain Evans and Yarin Gal. 2019. Generalizing from a few environments in safety-critical reinforcement learning. (2019) 1\u201316. http:\/\/arxiv.org\/abs\/1907.01475 Zachary Kenton Angelos Filos Owain Evans and Yarin Gal. 2019. Generalizing from a few environments in safety-critical reinforcement learning. (2019) 1\u201316. http:\/\/arxiv.org\/abs\/1907.01475"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1611835114"},{"key":"e_1_3_2_1_23_1","volume-title":"Using Restart Heuristics to Improve Agent Performance in Angry Birds. (5","author":"Liu Tommy","year":"2019"},{"key":"e_1_3_2_1_24_1","volume-title":"Playing Atari with Deep Reinforcement Learning. (12","author":"Mnih Volodymyr","year":"2013"},{"key":"e_1_3_2_1_25_1","volume-title":"Human-level control through deep reinforcement learning. Nature 518, 7540 (2","author":"Mnih Volodymyr","year":"2015"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CIG.2019.8848057"},{"key":"e_1_3_2_1_27_1","unstructured":"Charles Packer Katelyn Gao Jernej Kos Philipp Kr\u00e4henb\u00fchl Vladlen Koltun and Dawn Song. 2018. Assessing Generalization in Deep Reinforcement Learning. (2018). arXiv:arXiv:1810.12282 Charles Packer Katelyn Gao Jernej Kos Philipp Kr\u00e4henb\u00fchl Vladlen Koltun and Dawn Song. 2018. Assessing Generalization in Deep Reinforcement Learning. (2018). arXiv:arXiv:1810.12282"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"crossref","unstructured":"Julian Schrittwieser Ioannis Antonoglou Thomas Hubert Karen Simonyan Laurent Sifre Simon Schmitt Arthur Guez Edward Lockhart Demis Hassabis Thore Graepel Timothy Lillicrap and David Silver. 2019. Mastering Atari Go Chess and Shogi by Planning with a Learned Model. (2019). arXiv:arXiv:1911.08265 Julian Schrittwieser Ioannis Antonoglou Thomas Hubert Karen Simonyan Laurent Sifre Simon Schmitt Arthur Guez Edward Lockhart Demis Hassabis Thore Graepel Timothy Lillicrap and David Silver. 2019. Mastering Atari Go Chess and Shogi by Planning with a Learned Model. (2019). arXiv:arXiv:1911.08265","DOI":"10.1038\/s41586-020-03051-4"},{"key":"e_1_3_2_1_29_1","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. (2017). arXiv:arXiv:1707.06347 John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. (2017). arXiv:arXiv:1707.06347"},{"key":"e_1_3_2_1_30_1","volume-title":"TF-Agents: A library for Reinforcement Learning in TensorFlow. https:\/\/github.com\/tensorflow\/agents. https:\/\/github.com\/tensorflow\/agents [Online","author":"Pablo Castro Ethan Holly Oscar Ramirez","year":"2019"},{"key":"e_1_3_2_1_31_1","volume-title":"I\u2019m afraid I can\u2019t do that","author":"Seurin Mathieu","year":"2019"},{"key":"e_1_3_2_1_32_1","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton R.S.","year":"2018"},{"key":"e_1_3_2_1_33_1","unstructured":"Erving Teng. [n.d.]. Training your agents 7 times faster with ML-Agents - Unity Technologies Blog. https:\/\/blogs.unity3d.com\/2019\/11\/11\/training-your-agents-7-times-faster-with-ml-agents\/ Erving Teng. [n.d.]. Training your agents 7 times faster with ML-Agents - Unity Technologies Blog. https:\/\/blogs.unity3d.com\/2019\/11\/11\/training-your-agents-7-times-faster-with-ml-agents\/"},{"key":"e_1_3_2_1_34_1","volume-title":"NeurIPS (9","author":"Zahavy Tom","year":"2018"},{"key":"e_1_3_2_1_35_1","unstructured":"Amy Zhang Harsh Satija and Joelle Pineau. 2018. Decoupling Dynamics and Reward for Transfer Learning. (2018). http:\/\/arxiv.org\/abs\/1804.10689 Amy Zhang Harsh Satija and Joelle Pineau. 2018. Decoupling Dynamics and Reward for Transfer Learning. (2018). http:\/\/arxiv.org\/abs\/1804.10689"},{"key":"e_1_3_2_1_36_1","volume-title":"John Kolen, Jervis Pinto, Reza Pourabolghasem, Harold Chaput, James Pestrak, Mohsen Sardari, Long Lin, Navid Aghdaie, and Kazi Zaman.","author":"Zhao Yunqi","year":"2019"}],"event":{"name":"FDG '20: International Conference on the Foundations of Digital Games","location":"Bugibba Malta","acronym":"FDG '20"},"container-title":["International Conference on the Foundations of Digital Games"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3402942.3402944","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3402942.3402944","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:41:35Z","timestamp":1750200095000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3402942.3402944"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,15]]},"references-count":36,"alternative-id":["10.1145\/3402942.3402944","10.1145\/3402942"],"URL":"https:\/\/doi.org\/10.1145\/3402942.3402944","relation":{},"subject":[],"published":{"date-parts":[[2020,9,15]]}}}