{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T16:54:00Z","timestamp":1783184040119,"version":"3.54.6"},"reference-count":77,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2024,3,14]],"date-time":"2024-03-14T00:00:00Z","timestamp":1710374400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2024,3,31]]},"abstract":"<jats:p>Deep Reinforcement Learning (DRL) has received a lot of attention from the research community in recent years. As the technology moves away from game playing to practical contexts, such as autonomous vehicles and robotics, it is crucial to evaluate the quality of DRL agents.<\/jats:p>\n          <jats:p>\n            In this article, we propose a search-based approach to test such agents. Our approach, implemented in a tool called\n            <jats:sc>Indago<\/jats:sc>\n            , trains a classifier on failure and non-failure environment (i.e., pass) configurations resulting from the DRL training process. The classifier is used at testing time as a surrogate model for the DRL agent execution in the environment, predicting the extent to which a given environment configuration induces a failure of the DRL agent under test. The failure prediction acts as a fitness function, guiding the generation towards failure environment configurations, while saving computation time by deferring the execution of the DRL agent in the environment to those configurations that are more likely to expose failures.\n          <\/jats:p>\n          <jats:p>Experimental results show that our search-based approach finds 50% more failures of the DRL agent than state-of-the-art techniques. Moreover, such failures are, on average, 78% more diverse; similarly, the behaviors of the DRL agent induced by failure configurations are 74% more diverse.<\/jats:p>","DOI":"10.1145\/3631970","type":"journal-article","created":{"date-parts":[[2023,11,11]],"date-time":"2023-11-11T04:37:40Z","timestamp":1699677460000},"page":"1-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":31,"title":["Testing of Deep Reinforcement Learning Agents with Surrogate Models"],"prefix":"10.1145","volume":"33","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7825-3409","authenticated-orcid":false,"given":"Matteo","family":"Biagiola","sequence":"first","affiliation":[{"name":"Universit\u00e0 della Svizzera italiana, Lugano, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3088-0339","authenticated-orcid":false,"given":"Paolo","family":"Tonella","sequence":"additional","affiliation":[{"name":"Universit\u00e0 della Svizzera italiana, Lugano, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,3,14]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/2970276.2970311"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3238147.3238192"},{"key":"e_1_3_1_4_2","unstructured":"Marcin Andrychowicz Dwight Crow Alex Ray Jonas Schneider Rachel Fong Peter Welinder Bob McGrew Josh Tobin Pieter Abbeel and Wojciech Zaremba. 2017. Hindsight experience replay. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017 December 4-9 2017 Long Beach CA USA Isabelle Guyon Ulrike von Luxburg Samy Bengio Hanna M. Wallach Rob Fergus S. V. N. Vishwanathan and Roman Garnett (Eds.). 5048\u20135058. https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/453fadbd8a1a3af50a9df4df899537b5-Abstract.html"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1002\/stvr.1486"},{"key":"e_1_3_1_6_2","volume-title":"k-means++: The Advantages of Careful Seeding","author":"Arthur David","year":"2006","unstructured":"David Arthur and Sergei Vassilvitskii. 2006. k-means++: The Advantages of Careful Seeding. Technical Report. Stanford."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","unstructured":"R. Ben Abdessalem S. Nejati L. C. Briand and T. Stifter. 2018. Testing vision-based control systems using learnable evolutionary algorithms. In IEEE\/ACM 40th International Conference on Software Engineering (ICSE\u201918). 1016\u20131026. 10.1145\/3180155.3180160","DOI":"10.1145\/3180155.3180160"},{"key":"e_1_3_1_8_2","unstructured":"Matteo Biagiola. 2023. Replication package. https:\/\/www.github.com\/matteobiagiola\/drl-testing-experiments. Online; accessed November 2023."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","unstructured":"Matteo Biagiola and Paolo Tonella. 2022. Testing the plasticity of reinforcement learning-based systems. ACM Trans. Softw. Eng. Methodol. 31 4 (2022) 80:1\u201380:46. 10.1145\/3511701","DOI":"10.1145\/3511701"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2020.110542"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","unstructured":"Bruno Brito Achin Agarwal and Javier Alonso-Mora. 2022. Learning interaction-aware guidance for trajectory optimization in dense traffic scenarios. IEEE Transactions on Intelligent Transportation Systems 23 10 (2022) 18808\u201318821. 10.1109\/TITS.2022.3160936","DOI":"10.1109\/TITS.2022.3160936"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","unstructured":"Dong Chen Mohammad R. Hajidavalloo Zhaojian Li Kaian Chen Yongqiang Wang Longsheng Jiang and Yue Wang. 2023. Deep multi-agent reinforcement learning for highway on-ramp merging in mixed traffic. IEEE Trans. Intell. Transp. Syst. 24 11 (2023) 11623\u201311638. 10.1109\/TITS.2023.3285442","DOI":"10.1109\/TITS.2023.3285442"},{"key":"e_1_3_1_13_2","unstructured":"Torch Contributors. 2021. Torch CrossEntropyLoss class documentation. https:\/\/pytorch.org\/docs\/stable\/generated\/torch.nn.CrossEntropyLoss.html. Online; accessed November 2023."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00949"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11791"},{"key":"e_1_3_1_16_2","unstructured":"Richard Everett. 2019. Strategically Training and Evaluating Agents in Procedurally Generated Environments. Ph.D. Dissertation. University of Oxford."},{"key":"e_1_3_1_17_2","unstructured":"Ben Eysenbach Ruslan Salakhutdinov and Sergey Levine. 2021. Robust predictable control. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021 NeurIPS 2021 December 6-14 2021 virtual Marc\u2019Aurelio Ranzato Alina Beygelzimer Yann N. Dauphin Percy Liang and Jennifer Wortman Vaughan (Eds.). 27813\u201327825. https:\/\/proceedings.neurips.cc\/paper\/2021\/hash\/e9f85782949743dcc42079e629332b5f-Abstract.html"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-021-23440-1"},{"key":"e_1_3_1_19_2","unstructured":"Farama Foundation. 2022. Humanoid. https:\/\/www.gymlibrary.dev\/environments\/mujoco\/humanoid\/. Online; accessed November 2023."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3293882.3330566"},{"key":"e_1_3_1_21_2","unstructured":"Tuomas Haarnoja Aurick Zhou Pieter Abbeel and Sergey Levine. 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning ICML 2018 Stockholmsm\u00e4ssan Stockholm Sweden July 10-15 2018 (Proceedings of Machine Learning Research Vol. 80) Jennifer G. Dy and Andreas Krause (Eds.). PMLR 1856\u20131865. http:\/\/proceedings.mlr.press\/v80\/haarnoja18b.html"},{"key":"e_1_3_1_22_2","unstructured":"Hado van Hasselt. 2010. Double Q-learning. In Advances in Neural Information Processing Systems 23: 24th Annual Conference on Neural Information Processing Systems 2010. Proceedings of a Meeting held 6-9 December 2010 Vancouver British Columbia Canada John D. Lafferty Christopher K. I. Williams John Shawe-Taylor Richard S. Zemel and Aron Culotta (Eds.). Curran Associates Inc. 2613\u20132621. https:\/\/proceedings.neurips.cc\/paper\/2010\/hash\/091d584fced301b442654dd8c23b3fc9-Abstract.html"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11796"},{"key":"e_1_3_1_24_2","unstructured":"Sandy H. Huang Nicolas Papernot Ian J. Goodfellow Yan Duan and Pieter Abbeel. 2017. Adversarial attacks on neural network policies. In 5th International Conference on Learning Representations ICLR 2017 Toulon France April 24-26 2017 Workshop Track Proceedings. OpenReview.net. https:\/\/openreview.net\/forum?id=ryvlRyBKl"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377811.3380395"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","unstructured":"Joonho Lee Alexey Dosovitskiy Dario Bellicoso Vassilios Tsounis Vladlen Koltun and Marco Hutter. 2019. Learning agile and dynamic motor skills for legged robots. Sci. Robotics 4 26 (2019). 10.1126\/SCIROBOTICS.AAU5872","DOI":"10.1126\/SCIROBOTICS.AAU5872"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2018.2792536"},{"key":"e_1_3_1_28_2","first-page":"448","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the International Conference on Machine Learning. PMLR, 448\u2013456."},{"key":"e_1_3_1_29_2","unstructured":"Justin Basilico J. Ashok Chandrashekar Fernando Amat and Tony Jebara. 2017. Artwork Personalization at Netflix. https:\/\/netflixtechblog.com\/artwork-personalization-c589f074ad76. Online; accessed November 2023."},{"key":"e_1_3_1_30_2","unstructured":"Kittipat Virochsiri Jason Gauci and Edoardo Conti. 2018. Horizon: The first open source reinforcement learning platform for large-scale products and services. https:\/\/engineering.fb.com\/2018\/11\/01\/ml-applications\/horizon\/. Online; accessed November 2023."},{"key":"e_1_3_1_31_2","unstructured":"Kittipat Virochsiri Jason Gauci and Edoardo Conti. 2021. A platform for Reasoning systems (Reinforcement Learning Contextual Bandits etc.). https:\/\/github.com\/facebookresearch\/ReAgent. Online; accessed November 2023."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00500-003-0328-5"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1093\/oxfordjournals.pan.a004868"},{"key":"e_1_3_1_34_2","unstructured":"Narine Kokhlikyan Vivek Miglani Miguel Martin Edward Wang Bilal Alsallakh Jonathan Reynolds Alexander Melnikov Natalia Kliushkina Carlos Araya Siqi Yan and Orion Reblitz-Richardson. 2020. Captum: A unified and generic model interpretability library for PyTorch. CoRR abs\/2009.07896 (2020). arXiv:2009.07896 https:\/\/arxiv.org\/abs\/2009.07896"},{"key":"e_1_3_1_35_2","first-page":"5556","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Kuznetsov Arsenii","year":"2020","unstructured":"Arsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, and Dmitry Vetrov. 2020. Controlling overestimation bias with truncated mixture of continuous distributional quantile critics. In Proceedings of the International Conference on Machine Learning. PMLR, 5556\u20135566."},{"key":"e_1_3_1_36_2","unstructured":"Jennifer Langston. 2020. With reinforcement learning Microsoft brings a new class of AI solutions to customers. https:\/\/blogs.microsoft.com\/ai\/reinforcement-learning\/. Online; accessed November 2023."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","unstructured":"Timoth\u00e9e Lesort Natalia D\u00edaz Rodr\u00edguez Jean-Fran\u00e7ois Goudou and David Filliat. 2018. State representation learning for control: An overview. Neural Networks 108 (2018) 379\u2013392. 10.1016\/J.NEUNET.2018.07.006","DOI":"10.1016\/J.NEUNET.2018.07.006"},{"key":"e_1_3_1_38_2","unstructured":"Edouard Leurent. 2018. An Environment for Autonomous Driving Decision-Making. https:\/\/github.com\/eleurent\/highway-env. Online; accessed November 2023."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","unstructured":"Yen-Chen Lin Zhang-Wei Hong Yuan-Hong Liao Meng-Li Shih Ming-Yu Liu and Min Sun. 2017. Tactics of adversarial attack on deep reinforcement learning agents. In Proceedings of the 26th International Joint Conference on Artificial Intelligence IJCAI 2017 Melbourne Australia August 19-25 2017 Carles Sierra (Ed.). ijcai.org 3756\u20133762. 10.24963\/IJCAI.2017\/525","DOI":"10.24963\/IJCAI.2017\/525"},{"key":"e_1_3_1_40_2","unstructured":"Viktor Makoviychuk Lukasz Wawrzyniak Yunrong Guo Michelle Lu Kier Storey Miles Macklin David Hoeller Nikita Rudin Arthur Allshire Ankur Handa and Gavriel State. 2021. Isaac Gym: High performance GPU based physics simulation for robot learning. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 NeurIPS Datasets and Benchmarks 2021 December 2021 virtual Joaquin Vanschoren and Sai-Kit Yeung (Eds.). https:\/\/datasets-benchmarks-proceedings.neurips.cc\/paper\/2021\/hash\/28dd2c7955ce926456240b2ff0100bde-Abstract-round2.html"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","unstructured":"Henry B. Mann and Donald R. Whitney. 1947. On a test of whether one of two random variables is stochastically larger than the other. The Annals of Mathematical Statistics 18 1 (1947) 50\u201360. 10.1214\/aoms\/1177730491","DOI":"10.1214\/aoms\/1177730491"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.abk2822"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","unstructured":"Volodymyr Mnih Koray Kavukcuoglu David Silver Andrei A. Rusu Joel Veness Marc G. Bellemare Alex Graves Martin A. Riedmiller Andreas Fidjeland Georg Ostrovski Stig Petersen Charles Beattie Amir Sadik Ioannis Antonoglou Helen King Dharshan Kumaran Daan Wierstra Shane Legg and Demis Hassabis. 2015. Human-level control through deep reinforcement learning. Nat. 518 7540 (2015) 529\u2013533. 10.1038\/NATURE14236","DOI":"10.1038\/NATURE14236"},{"key":"e_1_3_1_44_2","unstructured":"OpenAI and collaborators. 2021. Gym Leaderboard. https:\/\/github.com\/openai\/gym\/wiki\/Leaderboard. Online; accessed November 2023."},{"key":"e_1_3_1_45_2","unstructured":"Adam Paszke Sam Gross Francisco Massa Adam Lerer James Bradbury Gregory Chanan Trevor Killeen Zeming Lin Natalia Gimelshein Luca Antiga Alban Desmaison Andreas K\u00f6pf Edward Z. Yang Zachary DeVito Martin Raison Alykhan Tejani Sasank Chilamkurthy Benoit Steiner Lu Fang Junjie Bai and Soumith Chintala. 2019. PyTorch: An imperative style high-performance deep learning library. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019 NeurIPS 2019 December 8-14 2019 Vancouver BC Canada Hanna M. Wallach Hugo Larochelle Alina Beygelzimer Florence d\u2019Alch\u00e9-Buc Emily B. Fox and Roman Garnett (Eds.). 8024\u20138035. https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/bdbca288fee7f92f2bfa9f7012727740-Abstract.html"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","unstructured":"Fabian Pedregosa Ga\u00ebl Varoquaux Alexandre Gramfort Vincent Michel Bertrand Thirion Olivier Grisel Mathieu Blondel Peter Prettenhofer Ron Weiss Vincent Dubourg Jake VanderPlas Alexandre Passos David Cournapeau Matthieu Brucher Matthieu Perrot and Edouard Duchesnay. 2011. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 12 (2011) 2825\u20132830. 10.5555\/1953048.2078195","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_3_1_47_2","unstructured":"Antonin Raffin. 2020. RL Baselines3 Zoo. https:\/\/github.com\/DLR-RM\/rl-baselines3-zoo. Online; accessed November 2023."},{"issue":"268","key":"e_1_3_1_48_2","first-page":"1","article-title":"Stable-Baselines3: Reliable reinforcement learning implementations","volume":"22","author":"Raffin Antonin","year":"2021","unstructured":"Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann. 2021. Stable-Baselines3: Reliable reinforcement learning implementations. Journal of Machine Learning Research 22, 268 (2021), 1\u20138. Retrieved from http:\/\/jmlr.org\/papers\/v22\/20-1364.html","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_49_2","first-page":"1634","volume-title":"Proceedings of the Conference on Robot Learning","author":"Raffin Antonin","year":"2022","unstructured":"Antonin Raffin, Jens Kober, and Freek Stulp. 2022. Smooth exploration for robotic reinforcement learning. In Proceedings of the Conference on Robot Learning. PMLR, 1634\u20131644."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409730"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCC.2004.843247"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","unstructured":"Peter J. Rousseeuw. 1987. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. J. Comput. Appl. Math. 20 (1987) 53\u201365. 10.1016\/0377-0427(87)90125-7","DOI":"10.1016\/0377-0427(87)90125-7"},{"key":"e_1_3_1_53_2","first-page":"91","volume-title":"Proceedings of the Conference on Robot Learning","author":"Rudin Nikita","year":"2022","unstructured":"Nikita Rudin, David Hoeller, Philipp Reist, and Marco Hutter. 2022. Learning to walk in minutes using massively parallel deep reinforcement learning. In Proceedings of the Conference on Robot Learning. PMLR, 91\u2013100."},{"key":"e_1_3_1_54_2","unstructured":"Tom Schaul John Quan Ioannis Antonoglou and David Silver. 2016. Prioritized experience replay. In 4th International Conference on Learning Representations ICLR 2016 San Juan Puerto Rico May 2-4 2016 Conference Track Proceedings Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1511.05952"},{"key":"e_1_3_1_55_2","unstructured":"John Schulman FilipWolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. CoRR abs\/1707.06347 (2017). arXiv:1707.06347 http:\/\/arxiv.org\/abs\/1707.06347"},{"key":"e_1_3_1_56_2","unstructured":"Karen Simonyan Andrea Vedaldi and Andrew Zisserman. 2014. Deep inside convolutional networks: Visualising image classification models and saliency maps. In 2nd International Conference on Learning Representations ICLR 2014 Banff AB Canada April 14-16 2014 Workshop Track Proceedings Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1312.6034"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2670313"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1177\/03611981211018697"},{"key":"e_1_3_1_59_2","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton Richard S.","year":"2018","unstructured":"Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. MIT Press."},{"key":"e_1_3_1_60_2","unstructured":"Christian Szegedy Wojciech Zaremba Ilya Sutskever Joan Bruna Dumitru Erhan Ian J. Goodfellow and Rob Fergus. 2014. Intriguing properties of neural networks. In 2nd International Conference on Learning Representations ICLR 2014 Banff AB Canada April 14-16 2014 Conference Track Proceedings Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1312.6199"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","unstructured":"Martin Tappler Filip Cano C\u00f3rdoba Bernhard K. Aichernig and Bettina K\u00f6nighofer. 2022. Search-based testing of reinforcement learning. In Proceedings of the 31st International Joint Conference on Artificial Intelligence IJCAI 2022 Vienna Austria 23-29 July 2022 Luc De Raedt (Ed.). ijcai.org 503\u2013510. 10.24963\/IJCAI.2022\/72","DOI":"10.24963\/IJCAI.2022\/72"},{"key":"e_1_3_1_62_2","unstructured":"Maxime Ellerbach Tawn Kramer and contributors. 2021. Self Driving Car Sandbox. https:\/\/github.com\/tawnkramer\/sdsandbox. Online; accessed November 2023."},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2012.6386109"},{"key":"e_1_3_1_64_2","unstructured":"Jonathan Uesato Ananya Kumar Csaba Szepesv\u00e1ri Tom Erez Avraham Ruderman Keith Anderson Krishnamurthy (Dj) Dvijotham Nicolas Heess and Pushmeet Kohli. 2019. Rigorous agent evaluation: An adversarial approach to uncover catastrophic failures. In 7th International Conference on Learning Representations ICLR 2019 New Orleans LA USA May 6-9 2019. OpenReview.net. https:\/\/openreview.net\/forum?id=B1xhQhRcK7"},{"key":"e_1_3_1_65_2","unstructured":"Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9 86 (2008) 2579\u20132605. http:\/\/jmlr.org\/papers\/v9\/vandermaaten08a.html"},{"issue":"2","key":"e_1_3_1_66_2","first-page":"101","article-title":"A critique and improvement of the CL common language effect size statistics of McGraw and Wong","volume":"25","author":"Vargha Andr\u00e1s","year":"2000","unstructured":"Andr\u00e1s Vargha and Harold D. Delaney. 2000. A critique and improvement of the CL common language effect size statistics of McGraw and Wong. Journal of Educational and Behavioral Statistics 25, 2 (2000), 101\u2013132.","journal-title":"Journal of Educational and Behavioral Statistics"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.4271\/2021-01-0248"},{"key":"e_1_3_1_68_2","unstructured":"Ari Viitala Rinu Boney and Juho Kannala. 2020. Learning to drive small scale cars from scratch. CoRR abs\/2008.00715 (2020). arXiv:2008.00715 https:\/\/arxiv.org\/abs\/2008.00715"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","unstructured":"Ari Viitala Rinu Boney Yi Zhao Alexander Ilin and Juho Kannala. 2021. Learning to Drive (L2D) as a Low-Cost Benchmark for Real-World Reinforcement Learning. In 20th International Conference on Advanced Robotics ICAR 2021 Ljubljana Slovenia December 6-10 2021. IEEE 275\u2013281. 10.1109\/ICAR53236.2021.9659342","DOI":"10.1109\/ICAR53236.2021.9659342"},{"key":"e_1_3_1_70_2","unstructured":"Che Wang Yanqiu Wu Quan Vuong and Keith W. Ross. 2019. Towards Simplicity in Deep Reinforcement Learning: Streamlined Off-Policy Learning. CoRR abs\/1910.02208 (2019). arXiv:1910.02208 http:\/\/arxiv.org\/abs\/1910.02208"},{"key":"e_1_3_1_71_2","first-page":"1995","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wang Ziyu","year":"2016","unstructured":"Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas. 2016. Dueling network architectures for deep reinforcement learning. In Proceedings of the International Conference on Machine Learning. PMLR, 1995\u20132003."},{"key":"e_1_3_1_72_2","unstructured":"Automotive World. 2016. Automatic Intelligent Parking: Audi at NIPS in Barcelona. Retrieved July 05 2023 from https:\/\/www.automotiveworld.com\/news-releases\/automatic-intelligent-parking-audi-nips-barcelona\/"},{"key":"e_1_3_1_73_2","unstructured":"Mengdi Xu Peide Huang Fengpei Li Jiacheng Zhu Xuewei Qi Kentaro Oguchi Zhiyuan Huang Henry Lam and Ding Zhao. 2021. Accelerated Policy Evaluation: Learning Adversarial Environments with Adaptive Importance Sampling. CoRR abs\/2106.10566 (2021). arXiv:2106.10566 https:\/\/arxiv.org\/abs\/2106.10566"},{"key":"e_1_3_1_74_2","first-page":"11842","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Yang Yuzhe","year":"2021","unstructured":"Yuzhe Yang, Kaiwen Zha, Yingcong Chen, Hao Wang, and Dina Katabi. 2021. Delving into deep imbalanced regression. In Proceedings of the International Conference on Machine Learning. PMLR, 11842\u201311851."},{"key":"e_1_3_1_75_2","unstructured":"Jie M. Zhang Mark Harman Lei Ma and Yang Liu. 2020. Machine learning testing: Survey landscapes and horizons. IEEE Transactions on Software Engineering (2020)."},{"key":"e_1_3_1_76_2","unstructured":"Qi Zhang and Tao Du. 2019. Self-driving scale car trained by Deep reinforcement Learning. CoRR abs\/1909.03467 (2019). arXiv:1909.03467 http:\/\/arxiv.org\/abs\/1909.03467"},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.1109\/ITSC48978.2021.9564972"},{"key":"e_1_3_1_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR48806.2021.9412011"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3631970","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3631970","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,17]],"date-time":"2025-07-17T18:46:28Z","timestamp":1752777988000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3631970"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,14]]},"references-count":77,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,3,31]]}},"alternative-id":["10.1145\/3631970"],"URL":"https:\/\/doi.org\/10.1145\/3631970","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,14]]},"assertion":[{"value":"2022-09-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-31","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}