{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,4]],"date-time":"2026-04-04T01:04:59Z","timestamp":1775264699845,"version":"3.50.1"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,10,26]],"date-time":"2023-10-26T00:00:00Z","timestamp":1698278400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001659","name":"German Research Foundation","doi-asserted-by":"crossref","award":["389792660"],"award-info":[{"award-number":["389792660"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100008530","name":"European Regional Development Fund","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100008530","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Model. Comput. Simul."],"published-print":{"date-parts":[[2023,10,31]]},"abstract":"<jats:p>\n            Neural networks (NN) are gaining importance in sequential decision-making. Deep reinforcement learning (DRL), in particular, is extremely successful in learning action policies in complex and dynamic environments. Despite this success, however, DRL technology is not without its failures, especially in safety-critical applications: (i) the training objective maximizes\n            <jats:italic>average<\/jats:italic>\n            rewards, which may disregard rare but critical situations and hence lack local robustness; (ii) optimization objectives targeting safety typically yield degenerated reward structures, which, for DRL to work, must be replaced with proxy objectives. Here, we introduce a methodology that can help to address both deficiencies. We incorporate\n            <jats:italic>evaluation stages<\/jats:italic>\n            (ES) into DRL, leveraging recent work on deep statistical model checking (DSMC), which verifies NN policies in Markov decision processes. Our ES apply DSMC at regular intervals to determine state space regions with weak performance. We adapt the subsequent DRL training priorities based on the outcome, (i) focusing DRL on critical situations and (ii) allowing to foster arbitrary objectives.\n          <\/jats:p>\n          <jats:p>We run case studies on two benchmarks. One of them is the Racetrack, an abstraction of autonomous driving that requires navigating a map without crashing into a wall. The other is MiniGrid, a widely used benchmark in the AI community. Our results show that DSMC-based ES can significantly improve both (i) and (ii).<\/jats:p>","DOI":"10.1145\/3607198","type":"journal-article","created":{"date-parts":[[2023,7,12]],"date-time":"2023-07-12T11:45:09Z","timestamp":1689162309000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["DSMC Evaluation Stages: Fostering Robust and Safe Behavior in Deep Reinforcement Learning \u2013 Extended Version"],"prefix":"10.1145","volume":"33","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1100-1952","authenticated-orcid":false,"given":"Timo P.","family":"Gros","sequence":"first","affiliation":[{"name":"Saarland University, Saarland Informatics Campus, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8957-8807","authenticated-orcid":false,"given":"Joschka","family":"Gro\u00df","sequence":"additional","affiliation":[{"name":"Saarland University, Saarland Informatics Campus, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2776-9288","authenticated-orcid":false,"given":"Daniel","family":"H\u00f6ller","sequence":"additional","affiliation":[{"name":"Saarland University, Saarland Informatics Campus, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1590-5876","authenticated-orcid":false,"given":"J\u00f6rg","family":"Hoffmann","sequence":"additional","affiliation":[{"name":"Saarland University and German Research Center for Artificial Intelligence (DFKI), Saarland Informatics Campus Saarbr\u00fccken, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6353-227X","authenticated-orcid":false,"given":"Michaela","family":"Klauck","sequence":"additional","affiliation":[{"name":"Saarland University, Saarland Informatics Campus, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-4974-6607","authenticated-orcid":false,"given":"Hendrik","family":"Meerkamp","sequence":"additional","affiliation":[{"name":"Saarland University, Saarland Informatics Campus, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5932-3395","authenticated-orcid":false,"given":"Nicola J.","family":"M\u00fcller","sequence":"additional","affiliation":[{"name":"Saarland University, Saarland Informatics Campus, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-5937-9448","authenticated-orcid":false,"given":"Lukas","family":"Schaller","sequence":"additional","affiliation":[{"name":"Saarland University, Saarland Informatics Campus, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8460-6007","authenticated-orcid":false,"given":"Verena","family":"Wolf","sequence":"additional","affiliation":[{"name":"Saarland University and German Research Center for Artificial Intelligence (DFKI), Saarland Informatics Campus, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,10,26]]},"reference":[{"key":"e_1_3_4_2_2","doi-asserted-by":"crossref","first-page":"356","DOI":"10.1038\/s42256-019-0070-z","article-title":"Solving the Rubik\u2019s cube with deep reinforcement learning and search","author":"Agostinelli Forest","year":"2019","unstructured":"Forest Agostinelli, Stephen McAleer, Alexander Shmakov, and Pierre Baldi. 2019. Solving the Rubik\u2019s cube with deep reinforcement learning and search. Nat. Mach. Intell. 1, 8 (2019), 356\u2013363.","journal-title":"Nat. Mach. Intell."},{"key":"e_1_3_4_3_2","volume-title":"32nd AAAI Conference on Artificial Intelligence","author":"Alshiekh Mohammed","year":"2018","unstructured":"Mohammed Alshiekh, Roderick Bloem, R\u00fcdiger Ehlers, Bettina K\u00f6nighofer, Scott Niekum, and Ufuk Topcu. 2018. Safe reinforcement learning via shielding. In 32nd AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_4_4_2","first-page":"269","volume-title":"International Conference on Machine Learning","author":"Amit Ron","year":"2020","unstructured":"Ron Amit, Ron Meir, and Kamil Ciosek. 2020. Discount factor as a regularizer in reinforcement learning. In International Conference on Machine Learning. PMLR, 269\u2013278."},{"key":"e_1_3_4_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.11.050"},{"key":"e_1_3_4_6_2","doi-asserted-by":"crossref","first-page":"630","DOI":"10.1007\/978-3-030-25540-4_36","volume-title":"International Conference on Computer Aided Verification","author":"Avni Guy","year":"2019","unstructured":"Guy Avni, Roderick Bloem, Krishnendu Chatterjee, Thomas A. Henzinger, Bettina K\u00f6nighofer, and Stefan Pranger. 2019. Run-time optimization for learned controllers through quantitative games. In International Conference on Computer Aided Verification. Springer, 630\u2013649."},{"key":"e_1_3_4_7_2","first-page":"83","volume-title":"Trustworthy AI\u2013Integrating Learning, Optimization and Reasoning: First International Workshop, TAILOR 2020","author":"Baier Christel","year":"2020","unstructured":"Christel Baier, Maria Christakis, Timo P. Gros, David Gro\u00df, Stefan Gumhold, Holger Hermanns, J\u00f6rg Hoffmann, and Michaela Klauck. 2020. Lab conditions for research on explainable automated decisions. In Trustworthy AI\u2013Integrating Learning, Optimization and Reasoning: First International Workshop, TAILOR 2020. Springer Nature, 83."},{"key":"e_1_3_4_8_2","doi-asserted-by":"publisher","DOI":"10.5555\/1644719.1644720"},{"key":"e_1_3_4_9_2","first-page":"249","volume-title":"International GI\/ITG Conference on Measurement, Modelling, and Evaluation of Computing Systems and Dependability and Fault Tolerance","author":"Bogdoll Jonathan","year":"2012","unstructured":"Jonathan Bogdoll, Arnd Hartmanns, and Holger Hermanns. 2012. Simulation and statistical model checking for modestly nondeterministic models. In International GI\/ITG Conference on Measurement, Modelling, and Evaluation of Computing Systems and Dependability and Fault Tolerance. Springer, 249\u2013252."},{"key":"e_1_3_4_10_2","first-page":"82","volume-title":"IJCAI Workshop on Planning with Uncertainty and Incomplete Information","author":"Bonet Blai","year":"2001","unstructured":"Blai Bonet and Hector Geffner. 2001. GPT: A tool for planning with uncertainty and partial information. In IJCAI Workshop on Planning with Uncertainty and Incomplete Information. 82\u201387."},{"key":"e_1_3_4_11_2","first-page":"12","volume-title":"International Conference on Automated Planning and Scheduling","author":"Bonet Blai","year":"2003","unstructured":"Blai Bonet and Hector Geffner. 2003. Labeled RTDP: Improving the convergence of real-time dynamic programming. In International Conference on Automated Planning and Scheduling. 12\u201321."},{"key":"e_1_3_4_12_2","first-page":"340","volume-title":"International Conference on Tools and Algorithms for the Construction and Analysis of Systems","author":"Budde Carlos E.","year":"2018","unstructured":"Carlos E. Budde, Pedro R. D\u2019Argenio, Arnd Hartmanns, and Sean Sedwards. 2018. A statistical model checker for nondeterminism and rare events. In International Conference on Tools and Algorithms for the Construction and Analysis of Systems. Springer, 340\u2013358."},{"key":"e_1_3_4_13_2","article-title":"Exploration by random network distillation","author":"Burda Yuri","year":"2018","unstructured":"Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov. 2018. Exploration by random network distillation. arXiv preprint arXiv:1810.12894 (2018).","journal-title":"arXiv preprint arXiv:1810.12894"},{"key":"e_1_3_4_14_2","volume-title":"International Conference on Learning Representations","volume":"105","author":"Chevalier-Boisvert Maxime","year":"2019","unstructured":"Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, and Yoshua Bengio. 2019. BabyAI: First steps towards grounded language learning with a human in the loop. In International Conference on Learning Representations, Vol. 105."},{"key":"e_1_3_4_15_2","unstructured":"Maxime Chevalier-Boisvert Bolun Dai Mark Towers Rodrigo de Lazcano Lucas Willems Salem Lahlou Suman Pal Pablo Samuel Castro and Jordan Terry. 2023. Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks. CoRR abs\/2306.13831 (2023)."},{"key":"e_1_3_4_16_2","volume-title":"AAAI Conference on Artificial Intelligence","volume":"31","author":"Ciosek Kamil","year":"2017","unstructured":"Kamil Ciosek and Shimon Whiteson. 2017. Offer: Off-environment reinforcement learning. In AAAI Conference on Artificial Intelligence, Vol. 31."},{"key":"e_1_3_4_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390199"},{"issue":"77","key":"e_1_3_4_18_2","first-page":"1","article-title":"ChainerRL: A deep reinforcement learning library","volume":"22","author":"Fujita Yasuhiro","year":"2021","unstructured":"Yasuhiro Fujita, Prabhat Nagarajan, Toshiki Kataoka, and Takahiro Ishikawa. 2021. ChainerRL: A deep reinforcement learning library. J. Mach. Learn. Res. 22, 77 (2021), 1\u201314.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_4_19_2","article-title":"Reinforcement learning with competitive ensembles of information-constrained primitives","author":"Goyal Anirudh","year":"2019","unstructured":"Anirudh Goyal, Shagun Sodhani, Jonathan Binas, Xue Bin Peng, Sergey Levine, and Yoshua Bengio. 2019. Reinforcement learning with competitive ensembles of information-constrained primitives. arXiv preprint arXiv:1906.10667 (2019).","journal-title":"arXiv preprint arXiv:1906.10667"},{"key":"e_1_3_4_20_2","volume-title":"Tracking the Race: Analyzing Racetrack Agents Trained with Imitation Learning and Deep Reinforcement Learning","author":"Gros Timo P.","year":"2021","unstructured":"Timo P. Gros. 2021. Tracking the Race: Analyzing Racetrack Agents Trained with Imitation Learning and Deep Reinforcement Learning. Master\u2019s thesis. Saarland University, Saarland Informatics Campus, 66123 Saarbr\u00fccken."},{"key":"e_1_3_4_21_2","volume-title":"9th International Symposium on Leveraging Applications of Formal Methods, Verification and Validation. From Verification to Explanation.","author":"Gros Timo P.","year":"2020","unstructured":"Timo P. Gros, David Gro\u00df, Stefan Gumhold, J\u00f6rg Hoffmann, Michaela Klauck, and Marcel Steinmetz. 2020. TraceVis: Towards visualization for deep statistical model checking. In 9th International Symposium on Leveraging Applications of Formal Methods, Verification and Validation. From Verification to Explanation."},{"key":"e_1_3_4_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-50086-3_6"},{"key":"e_1_3_4_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10009-022-00685-9"},{"key":"e_1_3_4_24_2","doi-asserted-by":"crossref","first-page":"197","DOI":"10.1007\/978-3-030-85172-9_11","volume-title":"International Conference on Quantitative Evaluation of Systems","author":"Gros Timo P.","year":"2021","unstructured":"Timo P. Gros, Daniel H\u00f6ller, J\u00f6rg Hoffmann, Michaela Klauck, Hendrik Meerkamp, and Verena Wolf. 2021. DSMC evaluation stages: Fostering robust and safe behavior in deep reinforcement learning. In International Conference on Quantitative Evaluation of Systems. Springer, 197\u2013216."},{"key":"e_1_3_4_25_2","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1007\/978-3-030-59854-9_2","volume-title":"International Conference on Quantitative Evaluation of Systems","author":"Gros Timo P.","year":"2020","unstructured":"Timo P. Gros, Daniel H\u00f6ller, J\u00f6rg Hoffmann, and Verena Wolf. 2020. Tracking the race between deep reinforcement learning and imitation learning. In International Conference on Quantitative Evaluation of Systems. Springer, 11\u201317."},{"key":"e_1_3_4_26_2","first-page":"3389","volume-title":"IEEE International Conference on Robotics and Automation (ICRA\u201917)","author":"Gu Shixiang","year":"2017","unstructured":"Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine. 2017. Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In IEEE International Conference on Robotics and Automation (ICRA\u201917). IEEE, 3389\u20133396."},{"key":"e_1_3_4_27_2","first-page":"1861","volume-title":"International Conference on Machine Learning","author":"Haarnoja Tuomas","year":"2018","unstructured":"Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning. PMLR, 1861\u20131870."},{"key":"e_1_3_4_28_2","article-title":"Dealing with sparse rewards in reinforcement learning","author":"Hare Joshua","year":"2019","unstructured":"Joshua Hare. 2019. Dealing with sparse rewards in reinforcement learning. arXiv preprint arXiv:1910.09281 (2019).","journal-title":"arXiv preprint arXiv:1910.09281"},{"key":"e_1_3_4_29_2","first-page":"593","volume-title":"International Conference on Tools and Algorithms for the Construction and Analysis of Systems (LNCS 8413)","author":"Hartmanns Arnd","year":"2014","unstructured":"Arnd Hartmanns and Holger Hermanns. 2014. The modest toolset: An integrated environment for quantitative modelling and verification. In International Conference on Tools and Algorithms for the Construction and Analysis of Systems (LNCS 8413). 593\u2013598."},{"key":"e_1_3_4_30_2","article-title":"Logically-constrained reinforcement learning","author":"Hasanbeig Mohammadhosein","year":"2018","unstructured":"Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening. 2018. Logically-constrained reinforcement learning. arXiv preprint arXiv:1801.08099 (2018).","journal-title":"arXiv preprint arXiv:1801.08099"},{"key":"e_1_3_4_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-57628-8_1"},{"key":"e_1_3_4_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2012.2205597"},{"key":"e_1_3_4_33_2","unstructured":"Bettina K\u00f6nighofer Roderick Bloem Sebastian Junges Nils Jansen and Alex Serban. 2020. Safe reinforcement learning using probabilistic shields. International Conference on Concurrency Theory: 31st CONCUR ."},{"key":"e_1_3_4_34_2","first-page":"4940","volume-title":"International Conference on Machine Learning","author":"Jiang Minqi","year":"2021","unstructured":"Minqi Jiang, Edward Grefenstette, and Tim Rockt\u00e4schel. 2021. Prioritized level replay. In International Conference on Machine Learning. PMLR, 4940\u20134950."},{"key":"e_1_3_4_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-662-49674-9_8"},{"key":"e_1_3_4_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ROMAN.2012.6343862"},{"key":"e_1_3_4_37_2","first-page":"1097","volume-title":"Neural Information Processing Systems Conference","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet classification with deep convolutional neural networks. In Neural Information Processing Systems Conference. 1097\u20131105."},{"issue":"3","key":"e_1_3_4_38_2","doi-asserted-by":"crossref","first-page":"385","DOI":"10.1109\/TSMC.2014.2358639","article-title":"Multiobjective reinforcement learning: A comprehensive overview","volume":"45","author":"Liu Chunming","year":"2014","unstructured":"Chunming Liu, Xin Xu, and Dewen Hu. 2014. Multiobjective reinforcement learning: A comprehensive overview. IEEE Trans. Syst., Man, Cybern.: Syst. 45, 3 (2014), 385\u2013398.","journal-title":"IEEE Trans. Syst., Man, Cybern.: Syst."},{"key":"e_1_3_4_39_2","first-page":"151","volume-title":"International Conference on Automated Planning and Scheduling","author":"McMahan H. Brendan","year":"2005","unstructured":"H. Brendan McMahan and Geoffrey J. Gordon. 2005. Fast exact planning in Markov decision processes. In International Conference on Automated Planning and Scheduling. 151\u2013160."},{"key":"e_1_3_4_40_2","first-page":"1928","volume-title":"International Conference on Machine Learning","author":"Mnih Volodymyr","year":"2016","unstructured":"Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016. Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning. PMLR, 1928\u20131937."},{"key":"e_1_3_4_41_2","article-title":"Playing Atari with deep reinforcement learning","author":"Mnih Volodymyr","year":"2013","unstructured":"Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013. Playing Atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 (2013).","journal-title":"arXiv preprint arXiv:1312.5602"},{"key":"e_1_3_4_42_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_4_43_2","first-page":"9839","volume-title":"Advances in Neural Information Processing Systems 31","author":"Nazari MohammadReza","year":"2018","unstructured":"MohammadReza Nazari, Afshin Oroojlooy, Lawrence Snyder, and Martin Takac. 2018. Reinforcement learning for solving the vehicle routing problem. In Advances in Neural Information Processing Systems 31, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.). Curran Associates, Inc., 9839\u20139849."},{"key":"e_1_3_4_44_2","first-page":"278","volume-title":"16th International Conference on Machine Learning (ICML\u201999)","author":"Ng Andrew Y.","year":"1999","unstructured":"Andrew Y. Ng, Daishi Harada, and Stuart J. Russell. 1999. Policy invariance under reward transformations: Theory and application to reward shaping. In 16th International Conference on Machine Learning (ICML\u201999). 278\u2013287."},{"key":"e_1_3_4_45_2","first-page":"2350","volume-title":"23rd International Joint Conference on Artificial Intelligence","author":"Pineda Luis Enrique","year":"2013","unstructured":"Luis Enrique Pineda, Yi Lu, Shlomo Zilberstein, and Claudia V. Goldman. 2013. Fault-tolerant planning under uncertainty. In 23rd International Joint Conference on Artificial Intelligence. 2350\u20132356."},{"key":"e_1_3_4_46_2","volume-title":"International Conference on Automated Planning and Scheduling","volume":"24","author":"Pineda Luis Enrique","year":"2014","unstructured":"Luis Enrique Pineda and Shlomo Zilberstein. 2014. Planning under uncertainty using reduced models: Revisiting determinization. In International Conference on Automated Planning and Scheduling, Vol. 24."},{"key":"e_1_3_4_47_2","doi-asserted-by":"publisher","DOI":"10.1002\/9780470316887"},{"key":"e_1_3_4_48_2","article-title":"RIDE: Rewarding Impact-Driven Exploration for procedurally-generated environments","author":"Raileanu Roberta","year":"2020","unstructured":"Roberta Raileanu and Tim Rockt\u00e4schel. 2020. RIDE: Rewarding Impact-Driven Exploration for procedurally-generated environments. arXiv preprint arXiv:2002.12292 (2020).","journal-title":"arXiv preprint arXiv:2002.12292"},{"key":"e_1_3_4_49_2","first-page":"4344","volume-title":"International Conference on Machine Learning","author":"Riedmiller Martin","year":"2018","unstructured":"Martin Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Wiele, Vlad Mnih, Nicolas Heess, and Jost Tobias Springenberg. 2018. Learning by playing solving sparse reward tasks from scratch. In International Conference on Machine Learning. PMLR, 4344\u20134353."},{"key":"e_1_3_4_50_2","doi-asserted-by":"publisher","DOI":"10.2352\/ISSN.2470-1173.2017.19.AVM-023"},{"key":"e_1_3_4_51_2","volume-title":"4th International Conference on Learning Representations","author":"Schaul Tom","year":"2016","unstructured":"Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. 2016. Prioritized experience replay. In 4th International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.)."},{"key":"e_1_3_4_52_2","article-title":"High-dimensional continuous control using generalized advantage estimation","author":"Schulman John","year":"2015","unstructured":"John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. 2015. High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438 (2015).","journal-title":"arXiv preprint arXiv:1506.02438"},{"key":"e_1_3_4_53_2","first-page":"298","volume-title":"10th International Conference on Machine Learning","volume":"298","author":"Schwartz Anton","year":"1993","unstructured":"Anton Schwartz. 1993. A reinforcement learning method for maximizing undiscounted rewards. In 10th International Conference on Machine Learning, Vol. 298. 298\u2013305."},{"key":"e_1_3_4_54_2","first-page":"266","volume-title":"International Conference on Computer Aided Verification","author":"Sen Koushik","year":"2005","unstructured":"Koushik Sen, Mahesh Viswanathan, and Gul Agha. 2005. On statistical model checking of stochastic systems. In International Conference on Computer Aided Verification. 266\u2013280."},{"key":"e_1_3_4_55_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature16961"},{"key":"e_1_3_4_56_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.aar6404"},{"key":"e_1_3_4_57_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature24270"},{"key":"e_1_3_4_58_2","article-title":"rlpyt: A research code base for deep reinforcement learning in PyTorch","author":"Stooke Adam","year":"2019","unstructured":"Adam Stooke and Pieter Abbeel. 2019. rlpyt: A research code base for deep reinforcement learning in PyTorch. arXiv preprint arXiv:1909.01500 (2019).","journal-title":"arXiv preprint arXiv:1909.01500"},{"key":"e_1_3_4_59_2","volume-title":"Reinforcement Learning: An Introduction (2nd ed.)","author":"Sutton Richard S.","year":"2018","unstructured":"Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction (2nd ed.). The MIT Press."},{"key":"e_1_3_4_60_2","article-title":"Grandmaster level in StarCraft II using multi-agent reinforcement learning","volume":"575","author":"Vinyals Oriol","year":"2019","unstructured":"Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Micha\u00ebl Mathieu, Andrew Dudzik, Junyoung Chung, David H. Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, Laurent Sifre, Trevor Cai, John P. Agapiou, Max Jaderberg, Alexander S. Vezhnevets, Remi Leblond, Tobias Pohlen, Valentin Dalibard, David Budden, Yury Sulsky, James Molloy, Tom L. Paine, Caglar Gulcehre, Ziyu Wang, Tobias Pfa, Yuhuai Wu, Roman Ring, Dani Yogatama, Dario W\u00fcnsch, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy Lillicrap, Koray Kavukcuoglu, Demis Hassabis, Chris Apps, and David Silver. 2019. Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature 575 (2019).","journal-title":"Nature"},{"key":"e_1_3_4_61_2","first-page":"20759","article-title":"Safe policy optimization with local generalized linear function approximations","volume":"34","author":"Wachi Akifumi","year":"2021","unstructured":"Akifumi Wachi, Yunyue Wei, and Yanan Sui. 2021. Safe policy optimization with local generalized linear function approximations. Adv. Neural Inf. Process. Syst. 34 (2021), 20759\u201320771.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_4_62_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1511.06581"},{"key":"e_1_3_4_63_2","first-page":"46","volume-title":"International Conference on Tools and Algorithms for the Construction and Analysis of Systems","author":"Younes H\u00e5kan L. S.","year":"2004","unstructured":"H\u00e5kan L. S. Younes, Marta Kwiatkowska, Gethin Norman, and David Parker. 2004. Numerical vs. statistical probabilistic model checking: An empirical study. In International Conference on Tools and Algorithms for the Construction and Analysis of Systems. Springer, 46\u201360."}],"container-title":["ACM Transactions on Modeling and Computer Simulation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3607198","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3607198","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:34Z","timestamp":1750178254000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3607198"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,26]]},"references-count":62,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,10,31]]}},"alternative-id":["10.1145\/3607198"],"URL":"https:\/\/doi.org\/10.1145\/3607198","relation":{},"ISSN":["1049-3301","1558-1195"],"issn-type":[{"value":"1049-3301","type":"print"},{"value":"1558-1195","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,26]]},"assertion":[{"value":"2022-02-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-13","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}