{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,23]],"date-time":"2026-04-23T14:59:42Z","timestamp":1776956382496,"version":"3.51.4"},"reference-count":90,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2022,7,12]],"date-time":"2022-07-12T00:00:00Z","timestamp":1657584000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"H2020 project PRECRIME"},{"name":"ERC Advanced Grant 2017 Program","award":["787703"],"award-info":[{"award-number":["787703"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2022,10,31]]},"abstract":"<jats:p>The dataset available for pre-release training of a machine-learning based system is often not representative of all possible execution contexts that the system will encounter in the field. Reinforcement Learning (RL) is a prominent approach among those that support continual learning, i.e., learning continually in the field, in the post-release phase. No study has so far investigated any method to test the plasticity of RL-based systems, i.e., their capability to adapt to an execution context that may deviate from the training one.<\/jats:p>\n          <jats:p>We propose an approach to test the plasticity of RL-based systems. The output of our approach is a quantification of the adaptation and anti-regression capabilities of the system, obtained by computing the adaptation frontier of the system in a changed environment. We visualize such frontier as an adaptation\/anti-regression heatmap in two dimensions, or as a clustered projection when more than two dimensions are involved. In this way, we provide developers with information on the amount of changes that can be accommodated by the continual learning component of the system, which is key to decide if online, in-the-field learning can be safely enabled or not.<\/jats:p>","DOI":"10.1145\/3511701","type":"journal-article","created":{"date-parts":[[2022,3,28]],"date-time":"2022-03-28T11:51:53Z","timestamp":1648468313000},"page":"1-46","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["Testing the Plasticity of Reinforcement Learning-based Systems"],"prefix":"10.1145","volume":"31","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7825-3409","authenticated-orcid":false,"given":"Matteo","family":"Biagiola","sequence":"first","affiliation":[{"name":"Universit\u00e0 della Svizzera italiana, Lugano, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paolo","family":"Tonella","sequence":"additional","affiliation":[{"name":"Universit\u00e0 della Svizzera italiana, Lugano, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,7,12]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/2970276.2970311"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3238147.3238192"},{"key":"e_1_3_2_4_2","article-title":"Spinning up in deep reinforcement learning","author":"Achiam Joshua","year":"2018","unstructured":"Joshua Achiam. 2018. Spinning up in deep reinforcement learning. Website Retreived on 29 Dec 2020 https:\/\/spinningup.openai.com\/en\/latest\/.","journal-title":"Website"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1002\/stvr.1486"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.1983.6313077"},{"key":"e_1_3_2_7_2","article-title":"Cartpole-v1 OpenAI Gym","author":"Barto Sutton","year":"2020","unstructured":"Sutton Barto and C. W. Anderson. 2020. Cartpole-v1 OpenAI Gym. Retrieved 29 Dec 2020 from https:\/\/gym.openai.com\/envs\/CartPole-v1\/. (2020).","journal-title":"Retrieved 29 Dec 2020 from https:\/\/gym.openai.com\/envs\/CartPole-v1\/"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.5555\/2566972.2566979"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3180155.3180160"},{"key":"e_1_3_2_10_2","unstructured":"Christopher Berner Greg Brockman Brooke Chan Vicki Cheung Przemys\u0142aw D\u0119biak Christy Dennison David Farhi Quirin Fischer Shariq Hashme Chris Hesse et\u00a0al. 2019. Dota 2 with large scale deep reinforcement learning. CoRR abs\/1912.06680 (2019)."},{"key":"e_1_3_2_11_2","unstructured":"Greg Brockman Vicki Cheung Ludwig Pettersson Jonas Schneider John Schulman Jie Tang and Wojciech Zaremba. 2016. OpenAI Gym. CoRR abs\/1606.01540 (2016)."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1364\/AO.26.004919"},{"key":"e_1_3_2_13_2","first-page":"2048","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Cobbe Karl","year":"2020","unstructured":"Karl Cobbe, Chris Hesse, Jacob Hilton, and John Schulman. 2020. Leveraging procedural generation to benchmark reinforcement learning. In Proceedings of the International Conference on Machine Learning. PMLR, 2048\u20132056."},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","unstructured":"P. Virtanen R. Gommers T. E. Oliphant M. Haberland T. Reddy D. Cournapeau E. Burovski P. Peterson W. Weckesser J. Bright S. J. van der Walt M. Brett J. Wilson K. J. Millman N. Mayorov A. R. J. Nelson E. Jones R. Kern E. Larson C. J. Carey \u0130. Polat Y. Feng E. W. Moore J. VanderPlas D. Laxalde J. Perktold R. Cimrman I. Henriksen E. A. Quintero C. R. Harris A. M. Archibald A. H. Ribeiro F. Pedregosa and P. van Mulbregt. 2020. SciPy 1.0 Contributors. SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python. Nature Methods 17 (2020) 261\u2013272.","DOI":"10.1038\/s41592-020-0772-5"},{"key":"e_1_3_2_15_2","article-title":"OpenAI Baselines","author":"Dhariwal Prafulla","year":"2017","unstructured":"Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov. 2017. OpenAI Baselines. Retrieved 29 Dec 2020 from https:\/\/github.com\/openai\/baselines. (2017).","journal-title":"Retrieved 29 Dec 2020 from https:\/\/github.com\/openai\/baselines"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/MET.2017.2"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3338906.3338954"},{"key":"e_1_3_2_18_2","unstructured":"Richard Everett. 2019. Strategically training and evaluating agents in procedurally generated environments. PhD thesis. University of Oxford."},{"key":"e_1_3_2_19_2","unstructured":"William Fedus Dibya Ghosh John D. Martin Marc G. Bellemare Yoshua Bengio and Hugo Larochelle. 2020. On catastrophic interference in atari 2600 games. CoRR abs\/2002.12499 (2020)."},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1016\/S1364-6613(99)01294-2"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3293882.3330566"},{"key":"e_1_3_2_22_2","unstructured":"Alborz Geramifard Robert H. Klein Christoph Dann William Dabney and Jonathan P. How. 2015. RLPy: A value-function-based reinforcement learning framework for education and research. Journal of Machine Learning Research 16 1 (2015) 1573\u20131578."},{"key":"e_1_3_2_23_2","volume-title":"Studies of Mind and Brain: Neural Principles of Learning, Perception, Development, Cognition, and Motor Control","author":"Grossberg Stephen T.","year":"2012","unstructured":"Stephen T. Grossberg. 2012. Studies of Mind and Brain: Neural Principles of Learning, Perception, Development, Cognition, and Motor Control. Vol. 70. Springer Science & Business Media."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3236024.3264835"},{"key":"e_1_3_2_25_2","unstructured":"Tuomas Haarnoja Aurick Zhou Pieter Abbeel and Sergey Levine. 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning ICML 2018 Stockholmsm\u00e4ssan Stockholm Sweden July 10-15 2018 (2018) J. G. Dy and A. Krause (Eds.). vol. 80 of Proceedings of Machine Learning Research. PMLR 1856\u20131865."},{"key":"e_1_3_2_26_2","article-title":"Automatic Intelligent Parking: Audi at NIPS in Barcelona","author":"Hartmann Christian","year":"2016","unstructured":"Christian Hartmann. 2016. Automatic Intelligent Parking: Audi at NIPS in Barcelona. Retrieved 29 Dec 2020 from https:\/\/www.audi-mediacenter.com\/en\/press-releases\/automatic-intelligent-parking-audi-at-nips-in-barcelona-7139. (2016).","journal-title":"Retrieved 29 Dec 2020 from https:\/\/www.audi-mediacenter.com\/en\/press-releases\/automatic-intelligent-parking-audi-at-nips-in-barcelona-7139"},{"key":"e_1_3_2_27_2","unstructured":"Hado Hasselt. 2010. Double Q-learning. In Advances in Neural Information Processing Systems 23: 24th Annual Conference on Neural Information Processing Systems 2010. Proceedings of a meeting held 6-9 December 2010 Vancouver British Columbia Canada (2010) J. D. Lafferty C. K. I. Williams J. Shawe-Taylor R. S. Zemel and A. Culotta (Eds.). Curran Associates Inc. 2613\u20132621."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11694"},{"key":"e_1_3_2_29_2","article-title":"Stable Baselines","author":"Hill Ashley","year":"2018","unstructured":"Ashley Hill, Antonin Raffin, Maximilian Ernestus, Adam Gleave, Anssi Kanervisto, Rene Traore, Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu. 2018. Stable Baselines. Retrieved 29 Dec 2020 from https:\/\/github.com\/hill-a\/stable-baselines. (2018).","journal-title":"https:\/\/github.com\/hill-a\/stable-baselines"},{"key":"e_1_3_2_30_2","unstructured":"Sandy Huang Nicolas Papernot Ian Goodfellow Yan Duan and Pieter Abbeel. 2017. Adversarial attacks on neural network policies. In 5th International Conference on Learning Representations ICLR 2017 Toulon France April 24-26 2017 Workshop Track Proceedings (2017) . OpenReview.net."},{"key":"e_1_3_2_31_2","article-title":"Deep Reinforcement Learning Doesn\u2019t Work Yet","author":"Irpan Alex","year":"2018","unstructured":"Alex Irpan. 2018. Deep Reinforcement Learning Doesn\u2019t Work Yet. Retrieved from https:\/\/www.alexirpan.com\/2018\/02\/14\/rl-hard.html. (2018).","journal-title":"Retrieved from https:\/\/www.alexirpan.com\/2018\/02\/14\/rl-hard.html"},{"key":"e_1_3_2_32_2","article-title":"Artwork Personalization at Netflix","author":"Amat Justin Basilico J. Ashok Chandrashekar, Fernando","year":"2017","unstructured":"Justin Basilico J. Ashok Chandrashekar, Fernando Amat and Tony Jebara. 2017. Artwork Personalization at Netflix. Retrieved 29 Dec 2020 from https:\/\/netflixtechblog.com\/artwork-personalization-c589f074ad76. (2017).","journal-title":"https:\/\/netflixtechblog.com\/artwork-personalization-c589f074ad76"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICST46399.2020.00018"},{"key":"e_1_3_2_34_2","unstructured":"Jason Gauci Edoardo Conti Yitao Liang Kittipat Virochsiri Zhengxing Chen Yuchen He Zachary Kaden Vivek Narayanan and Xiaohui Ye. 2018. Horizon: Facebook\u2019s opensource applied reinforcement learning platform. CoRR abs\/1811.00260 (2018)."},{"key":"e_1_3_2_35_2","unstructured":"Jason Gauci Edoardo Conti Yitao Liang Kittipat Virochsiri Zhengxing Chen Yuchen He Zachary Kaden Vivek Narayanan and Xiaohui Ye. 2021. A Platform for Reasoning Systems (Reinforcement Learning Contextual Bandits etc.). Retrieved 29 Dec 2020 from https:\/\/github.com\/facebookresearch\/ReAgent. (2021)."},{"key":"e_1_3_2_36_2","article-title":"Policy consolidation for continual reinforcement learning","author":"Kaplanis Christos","year":"2019","unstructured":"Christos Kaplanis, Murray Shanahan, and Claudia Clopath. 2019. Policy consolidation for continual reinforcement learning. arXiv:1902.00255. Retrieved from https:\/\/arxiv.org\/abs\/1902.00255.","journal-title":"arXiv:1902.00255"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019.00108"},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","unstructured":"James Kirkpatrick Razvan Pascanu Neil C. Rabinowitz Joel Veness Guillaume Desjardins Andrei A. Rusu Kieran Milan John Quan Tiago Ramalho Agnieszka Grabska-Barwinska Demis Hassabis Claudia Clopath Dharshan Kumaran and Raia Hadsell. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences 114 13 (2017) 3521\u20133526.","DOI":"10.1073\/pnas.1611835114"},{"key":"e_1_3_2_39_2","article-title":"With Reinforcement Learning, Microsoft Brings a New Class of AI Solutions to Customers","author":"Langston Jennifer","year":"2020","unstructured":"Jennifer Langston. 2020. With Reinforcement Learning, Microsoft Brings a New Class of AI Solutions to Customers. Retrieved 29 Dec 2020 from https:\/\/blogs.microsoft.com\/ai\/reinforcement-learning\/. (2020).","journal-title":"Retrieved 29 Dec 2020 from https:\/\/blogs.microsoft.com\/ai\/reinforcement-learning\/"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-27645-3_5"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2773081"},{"key":"e_1_3_2_42_2","unstructured":"Yen-Chen Lin Zhang-Wei Hong Yuan-Hong Liao Meng-Li Shih Ming-Yu Liu and Min Sun. 2017. Tactics of adversarial attack on deep reinforcement learning agents. In 5th International Conference on Learning Representations ICLR 2017 Toulon France April 24-26 2017 Workshop Track Proceedings (2017) . OpenReview.net."},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3238147.3238202"},{"key":"e_1_3_2_44_2","first-page":"100","volume-title":"Proceedings of the 2018 IEEE 29th International Symposium on Software Reliability Engineering (ISSRE)","year":"2018","unstructured":"Lei Ma, Fuyuan Zhang, Jiyuan Sun, Minhui Xue, Bo Li, Felix Juefei-Xu, Chao Xie, Li Li, Yang Liu, Jianjun Zhao, and Yadong Wang. 2018. Deepmutation: Mutation testing of deep learning systems. In Proceedings of the 2018 IEEE 29th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 100\u2013111."},{"key":"e_1_3_2_45_2","article-title":"AlphaTest: A Tool for Testing the Plasticity of Reinforcement Learning Based Systems","author":"Matteo Biagiola","year":"2021","unstructured":"Biagiola Matteo and Tonella Paolo. 2021. AlphaTest: A Tool for Testing the Plasticity of Reinforcement Learning Based Systems. Retrieved 14 Jan 2022 from https:\/\/github.com\/testingautomated-usi\/rl-plasticity-experiments. (2021).","journal-title":"https:\/\/github.com\/testingautomated-usi\/rl-plasticity-experiments"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0079-7421(08)60536-8"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/JRPROC.1961.287775"},{"key":"e_1_3_2_48_2","doi-asserted-by":"crossref","unstructured":"Volodymyr Mnih Koray Kavukcuoglu David Silver Andrei A. Rusu Joel Veness Marc G. Bellemare Alex Graves Martin A. Riedmiller Andreas Fidjeland Georg Ostrovski Stig Petersen Charles Beattie Amir Sadik Ioannis Antonoglou Helen King Dharshan Kumaran Daan Wierstra Shane Legg and Demis Hassabis. 2015. Human-level control through deep reinforcement learning. Nature 518 7540 (2015) 529\u2013533.","DOI":"10.1038\/nature14236"},{"key":"e_1_3_2_49_2","unstructured":"Andrew William Moore. 1990. Efficient memory-based learning for robot control. Tech. rep."},{"key":"e_1_3_2_50_2","article-title":"MountainCar-v0 OpenAI Gym","author":"Moore Barto","year":"2020","unstructured":"Barto Moore and Sutton. 2020. MountainCar-v0 OpenAI Gym. Retrieved 29 Dec 2020 from https:\/\/gym.openai.com\/envs\/MountainCar-v0. (2020).","journal-title":"https:\/\/gym.openai.com\/envs\/MountainCar-v0"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2017.10.031"},{"key":"e_1_3_2_52_2","article-title":"Pendulum-v0 OpenAI Gym","year":"2020","unstructured":"OpenAI. 2020. Pendulum-v0 OpenAI Gym. Retrieved from https:\/\/gym.openai.com\/envs\/Pendulum-v0\/. (2020).","journal-title":"Retrieved from https:\/\/gym.openai.com\/envs\/Pendulum-v0\/"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2019.01.012"},{"key":"e_1_3_2_54_2","unstructured":"Emilio Parisotto Jimmy Lei Ba and Ruslan Salakhutdinov. 2015. Actor-mimic: Deep multitask and transfer reinforcement learning. In 4th International Conference on Learning Representations ICLR 2016 San Juan Puerto Rico May 2-4 2016 Conference Track Proceedings (2016) Y. Bengio and Y. LeCun (Eds.)."},{"key":"e_1_3_2_55_2","article-title":"Modified Version of the CartPole-v0 OpenAI Environment","author":"Patanjali Aaditya","year":"2017","unstructured":"Aaditya Patanjali. 2017. Modified Version of the CartPole-v0 OpenAI Environment. Retrieved 29 Dec 2020 from https:\/\/github.com\/AadityaPatanjali\/gym-cartpolemod. (2017).","journal-title":"https:\/\/github.com\/AadityaPatanjali\/gym-cartpolemod"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3132747.3132785"},{"key":"e_1_3_2_58_2","volume-title":"Software Testing and Analysis: Process, Principles, and Techniques","author":"Pezz\u00e8 Mauro","year":"2008","unstructured":"Mauro Pezz\u00e8 and Michal Young. 2008. Software Testing and Analysis: Process, Principles, and Techniques. John Wiley & Sons."},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-020-09881-0"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409730"},{"key":"e_1_3_2_61_2","first-page":"350","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Rolnick David","year":"2019","unstructured":"David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy Lillicrap, and Gregory Wayne. 2019. Experience replay for continual learning. In Proceedings of the Advances in Neural Information Processing Systems. 350\u2013360."},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/PRDC.2018.00016"},{"key":"e_1_3_2_63_2","unstructured":"Avraham Ruderman Richard Everett Bristy Sikder Hubert Soyer Jonathan Uesato Ananya Kumar Charlie Beattie and Pushmeet Kohli. 2018. Uncovering surprising behaviors in reinforcement learning via worst-case analysis. Tech. rep."},{"key":"e_1_3_2_64_2","unstructured":"Christian Rupprecht Cyril Ibrahim and Christopher J. Pal. 2019. Finding and visualizing weaknesses of deep reinforcement learning agents. In 8th International Conference on Learning Representations ICLR 2020 Addis Ababa Ethiopia April 26-30 2020 (2020) . OpenReview.net."},{"key":"e_1_3_2_65_2","unstructured":"Andrei A. Rusu Neil C. Rabinowitz Guillaume Desjardins Hubert Soyer James Kirkpatrick Koray Kavukcuoglu Razvan Pascanu and Raia Hadsell. 2016. Progressive neural networks. CoRR abs\/1606.04671 (2016)."},{"key":"e_1_3_2_66_2","article-title":"A Warehouse Robot Learns to Sort Out the Tricky Stuff","author":"Satariano Adam","year":"2020","unstructured":"Adam Satariano and Cade Metz. 2020. A Warehouse Robot Learns to Sort Out the Tricky Stuff. Retrieved 29 Dec 2020 from https:\/\/www.nytimes.com\/2020\/01\/29\/technology\/warehouse-robot.html. (2020).","journal-title":"https:\/\/www.nytimes.com\/2020\/01\/29\/technology\/warehouse-robot.html"},{"key":"e_1_3_2_67_2","unstructured":"Tom Schaul John Quan Ioannis Antonoglou and David Silver. 2015. Prioritized experience replay. In 4th International Conference on Learning Representations ICLR 2016 San Juan Puerto Rico May 2-4 2016 Conference Track Proceedings (2016) Y. Bengio and Y. Le Cun (Eds.)."},{"key":"e_1_3_2_68_2","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal policy optimization algorithms. CoRR abs\/1707.06347 (2017)."},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/QRS-C.2018.00032"},{"key":"e_1_3_2_70_2","first-page":"2990","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Shin Hanul","year":"2017","unstructured":"Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. 2017. Continual learning with deep generative replay. In Proceedings of the Advances in Neural Information Processing Systems. 2990\u20132999."},{"key":"e_1_3_2_71_2","doi-asserted-by":"crossref","unstructured":"David Silver Aja Huang Chris J. Maddison Arthur Guez Laurent Sifre George van den Driessche Julian Schrittwieser Ioannis Antonoglou Vedavyas Panneershelvam Marc Lanctot Sander Dieleman Dominik Grewe John Nham Nal Kalchbrenner Ilya Sutskever Timothy P. Lillicrap Madeleine Leach Koray Kavukcuoglu Thore Graepel and Demis Hassabis. 2016. Mastering the game of Go with deep neural networks and tree search. Nature 529 7587 (2016) 484\u2013489.","DOI":"10.1038\/nature16961"},{"key":"e_1_3_2_72_2","volume-title":"Proceedings of the 2013 AAAI Spring Symposium Series","author":"Silver Daniel L.","year":"2013","unstructured":"Daniel L. Silver, Qiang Yang, and Lianghao Li. 2013. Lifelong machine learning systems: Beyond learning algorithms. In Proceedings of the 2013 AAAI Spring Symposium Series."},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377811.3380353"},{"key":"e_1_3_2_74_2","article-title":"Acrobot-v1 OpenAI Gym","author":"Geramifard Dann, Sutton, and","year":"2020","unstructured":"Dann, Sutton, and Geramifard. 2020. Acrobot-v1 OpenAI Gym. Retrieved 29 Dec 2020 from https:\/\/gym.openai.com\/envs\/Acrobot-v1\/. (2020).","journal-title":"Retrieved 29 Dec 2020 from https:\/\/gym.openai.com\/envs\/Acrobot-v1\/"},{"key":"e_1_3_2_75_2","unstructured":"Richard S. Sutton. 1996. Generalization in reinforcement learning: Successful examples using sparse coarse coding. In Advances in Neural Information Processing Systems 8 NIPS Denver CO USA November 27-30 1995 (1995) D. S. Touretzky M. Mozer and M. E. Hasselmo (Eds.) MIT Press 1038\u20131044."},{"key":"e_1_3_2_76_2","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton Richard S.","year":"2018","unstructured":"Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. MIT press."},{"key":"e_1_3_2_77_2","article-title":"Intriguing properties of neural networks","author":"Szegedy Christian","year":"2013","unstructured":"Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013. Intriguing properties of neural networks. arXiv:1312.6199. Retrieved from https:\/\/arxiv.org\/abs\/1312.6199.","journal-title":"arXiv:1312.6199"},{"key":"e_1_3_2_78_2","article-title":"Transfer learning for reinforcement learning domains: A survey.","author":"Taylor Matthew E.","year":"2009","unstructured":"Matthew E. Taylor and Peter Stone. 2009. Transfer learning for reinforcement learning domains: A survey.Journal of Machine Learning Research 10 (2009), 1633\u20131685.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1145\/3180155.3180220"},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2012.6386109"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/3387940.3391462"},{"key":"e_1_3_2_82_2","first-page":"1075","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Tsitsiklis John N.","year":"1997","unstructured":"John N. Tsitsiklis and Benjamin Van Roy. 1997. Analysis of temporal-diffference learning with function approximation. In Proceedings of the Advances in Neural Information Processing Systems. 1075\u20131081."},{"key":"e_1_3_2_83_2","unstructured":"Cumhur Erkan Tuncali and Georgios Fainekos. 2019. Rapidly-exploring random trees-based test generation for autonomous vehicles. CoRR abs\/1903.10629 (2019)."},{"key":"e_1_3_2_84_2","unstructured":"Jonathan Uesato Ananya Kumar Csaba Szepesvari Tom Erez Avraham Ruderman Keith Anderson Nicolas Heess and Pushmeet Kohli. 2018. Rigorous agent evaluation: An adversarial approach to uncover catastrophic failures. In 7th International Conference on Learning Representations ICLR 2019 New Orleans LA USA May 6-9 2019 (2019) . OpenReview.net."},{"key":"e_1_3_2_85_2","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Maaten Laurens van der","year":"2008","unstructured":"Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, 11 (2008), 2579\u20132605.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_86_2","first-page":"1995","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wang Ziyu","year":"2016","unstructured":"Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas. 2016. Dueling network architectures for deep reinforcement learning. In Proceedings of the International Conference on Machine Learning. PMLR, 1995\u20132003."},{"key":"e_1_3_2_87_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-29044-2"},{"issue":"6","key":"e_1_3_2_88_2","first-page":"2","article-title":"Torcs, the open racing car simulator","volume":"4","author":"Wymann Bernhard","year":"2000","unstructured":"Bernhard Wymann, Eric Espi\u00e9, Christophe Guionneau, Christos Dimitrakakis, R\u00e9mi Coulom, and Andrew Sumner. 2000. Torcs, the open racing car simulator. Software Available at http:\/\/torcs. sourceforge. net 4, 6 (2000), 2.","journal-title":"Software Available at http:\/\/torcs. sourceforge. net"},{"key":"e_1_3_2_89_2","article-title":"Machine learning testing: Survey, landscapes and horizons","author":"Zhang Jie M.","year":"2022","unstructured":"Jie M. Zhang, Mark Harman, Lei Ma, and Yang Liu. 2022. Machine learning testing: Survey, landscapes and horizons. IEEE Transactions on Software Engineering 48, 2 (2022), 1\u201336.","journal-title":"IEEE Transactions on Software Engineering"},{"key":"e_1_3_2_90_2","doi-asserted-by":"publisher","DOI":"10.1145\/3238147.3238187"},{"key":"e_1_3_2_91_2","unstructured":"Zhuangdi Zhu Kaixiang Lin and Jiayu Zhou. 2020. Transfer learning in deep reinforcement learning: A survey. CoRR abs\/2009.07888 (2020)."}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3511701","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3511701","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:11:47Z","timestamp":1750191107000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3511701"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,12]]},"references-count":90,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,10,31]]}},"alternative-id":["10.1145\/3511701"],"URL":"https:\/\/doi.org\/10.1145\/3511701","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,7,12]]},"assertion":[{"value":"2021-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-07-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}