{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T04:57:40Z","timestamp":1760245060205,"version":"3.37.3"},"reference-count":154,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2022,4,1]],"date-time":"2022-04-01T00:00:00Z","timestamp":1648771200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,4,4]],"date-time":"2022-04-04T00:00:00Z","timestamp":1649030400000},"content-version":"vor","delay-in-days":3,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001711","name":"Schweizerischer Nationalfonds zur F\u00f6rderung der Wissenschaftlichen Forschung","doi-asserted-by":"publisher","award":["CSSII5_177179"],"award-info":[{"award-number":["CSSII5_177179"]}],"id":[{"id":"10.13039\/501100001711","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100006389","name":"University of Geneva","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100006389","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2022,4]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Despite the recent success of reinforcement learning in various domains, these approaches remain, for the most part, deterringly sensitive to hyper-parameters and are often riddled with essential engineering feats allowing their success. We consider the case of off-policy generative adversarial imitation learning, and perform an in-depth review, qualitative and quantitative, of the method. We show that forcing the learned reward function to be local Lipschitz-continuous is a<jats:italic>sine qua non<\/jats:italic>condition for the method to perform well. We then study the effects of this necessary condition and provide several theoretical results involving the local Lipschitzness of the state-value function. We complement these guarantees with empirical evidence attesting to the strong positive effect that the consistent satisfaction of the Lipschitzness constraint on the reward has on imitation performance. Finally, we tackle a generic pessimistic reward preconditioning add-on spawning a large class of reward shaping methods, which makes the base method it is plugged into provably more robust, as shown in several additional theoretical guarantees. We then discuss these through a fine-grained lens and share our insights. Crucially, the guarantees derived and reported in this work are valid for<jats:italic>any<\/jats:italic>reward satisfying the Lipschitzness condition, nothing is specific to imitation. As such, these may be of independent interest.<\/jats:p>","DOI":"10.1007\/s10994-022-06144-5","type":"journal-article","created":{"date-parts":[[2022,4,4]],"date-time":"2022-04-04T21:02:46Z","timestamp":1649106166000},"page":"1431-1521","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Lipschitzness is all you need to tame off-policy generative adversarial imitation learning"],"prefix":"10.1007","volume":"111","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1993-3121","authenticated-orcid":false,"given":"Lionel","family":"Blond\u00e9","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pablo","family":"Strasser","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alexandros","family":"Kalousis","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,4,4]]},"reference":[{"key":"6144_CR1","unstructured":"Abbasi-Yadkori, Y., Bartlett, P. L., & Szepesvari, C. (2013). Online learning in Markov decision processes with adversarially chosen transition probability distributions. Preprint retrieved from arXiv:1303.3055"},{"key":"6144_CR2","doi-asserted-by":"publisher","unstructured":"Abbeel, P., & Ng, A. Y. (2004). Apprenticeship learning via inverse reinforcement learning. In It international conference on machine learning (ICML). https:\/\/doi.org\/10.1145\/1015330.1015430","DOI":"10.1145\/1015330.1015430"},{"issue":"1","key":"6144_CR3","doi-asserted-by":"publisher","first-page":"1582","DOI":"10.5555\/2946645.2946691","volume":"17","author":"S Abdallah","year":"2016","unstructured":"Abdallah, S., & Kaisers, M. (2016). Addressing environment non-stationarity by repeating Q-learning updates. Journal of Machine Learning Research (JMLR), 17(1), 1582\u20131612. https:\/\/doi.org\/10.5555\/2946645.2946691","journal-title":"Journal of Machine Learning Research (JMLR)"},{"key":"6144_CR4","unstructured":"Achiam, J., Knight, E., & Abbeel, P. (2019). Towards characterizing divergence in deep q-learning. Preprint retrieved from arXiv:1903.08894"},{"key":"6144_CR5","unstructured":"Anava, O., & Karnin, Z. (2016). Multi-armed Bandits: Competing with optimal sequences. In Neural information processing systems (NeurIPS). https:\/\/papers.nips.cc\/paper\/6341-multi-armed-bandits-competing-with-optimal-sequences.pdf"},{"key":"6144_CR6","unstructured":"Arjovsky, M., & Bottou, L. (2017). Towards principled methods for training generative adversarial networks. In International conference on learning representations (ICLR). Preprint retrieved from arXiv:1701.04862"},{"key":"6144_CR7","unstructured":"Arjovsky, M., Chintala, S., & Bottou, L. (2017). Wasserstein GAN. Preprint retrieved from arXiv:1701.07875"},{"key":"6144_CR8","unstructured":"Atkeson, C. G., & Schaal, S. (1997). Robot learning from demonstration. In International conference on machine learning (ICML) (Vol.\u00a097, pp. 12\u201320). https:\/\/pdfs.semanticscholar.org\/a0e2\/dd2f6b116a12b235caa9b8ef8960f7afc585.pdf"},{"key":"6144_CR9","unstructured":"Auer, P., Cesa-Bianchi, N., Freund, Y., & Schapire, R. E. (1995). The non-stochastic multi-armed bandit problem. Symposium on Foundations of Computer Science. https:\/\/cseweb.ucsd.edu\/~yfreund\/papers\/bandits.pdf"},{"key":"6144_CR10","unstructured":"Auer, P., Gajane, P., & Ortner, R. (2019). Adaptively tracking the best bandit arm with an unknown number of distribution changes. In Conference on learning theory (COLT). http:\/\/proceedings.mlr.press\/v99\/auer19a\/auer19a.pdf"},{"key":"6144_CR11","unstructured":"Ba, J. L., Kiros, J. R., & Hinton, G. E. (2016). Layer normalization. Preprint retrieved from arXiv:1607.06450"},{"key":"6144_CR12","unstructured":"Bagnell, J. A. (2015). An invitation to imitation. Tech. rep., Carnegie Mellon, Robotics Institute. https:\/\/www.ri.cmu.edu\/pub_files\/2015\/3\/InvitationToImitation_3_1415.pdf"},{"key":"6144_CR13","unstructured":"Baram, N., Anschel, O., Caspi, I., & Mannor, S. (2017). End-to-end differentiable adversarial imitation learning. In International conference on machine learning (ICML) (pp. 390\u2013399). http:\/\/proceedings.mlr.press\/v70\/baram17a\/baram17a.pdf"},{"key":"6144_CR14","unstructured":"Besbes, O., Gur, Y., & Zeevi, A. (2014). Stochastic multi-armed-bandit problem with non-stationary rewards. In Neural information processing systems (NeurIPS). http:\/\/papers.nips.cc\/paper\/5378-stochastic-multi-armed-bandit-problem-with-non-stationary-rewards.pdf"},{"key":"6144_CR15","unstructured":"Biewald, L. (2020). Experiment tracking with weights and biases. https:\/\/www.wandb.com\/"},{"key":"6144_CR16","doi-asserted-by":"crossref","unstructured":"Billard, A., Calinon, S., Dillmann, R., & Schaal, S. (2008). Robot programming by demonstration. In Bruno, S., & Oussama, K. (Eds.) Springer handbook of robotics (pp. 1371\u20131394). Springer. http:\/\/link.springer.com\/referenceworkentry\/10.1007\/978-3-540-30301-5_60","DOI":"10.1007\/978-3-540-30301-5_60"},{"issue":"1","key":"6144_CR17","doi-asserted-by":"publisher","first-page":"108","DOI":"10.1162\/neco.1995.7.1.108","volume":"7","author":"CM Bishop","year":"1995","unstructured":"Bishop, C. M. (1995). Training with noise is equivalent to Tikhonov regularization. Neural Computation, 7(1), 108\u2013116. https:\/\/doi.org\/10.1162\/neco.1995.7.1.108","journal-title":"Neural Computation"},{"key":"6144_CR18","unstructured":"Blond\u00e9, L., & Kalousis, A. (2019). Sample-efficient imitation learning via generative adversarial nets. In International conference on artificial intelligence and statistics (AISTATS). Preprint retrieved from arXiv:1809.02064"},{"key":"6144_CR19","unstructured":"Borsa, D., Piot, B., Munos, R., & Pietquin, O. (2017). Observational learning by reinforcement learning. In Neural information processing systems (NeurIPS). Preprint retrieved from arXiv:1706.06617"},{"key":"6144_CR20","unstructured":"Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). OpenAI Gym. Preprint retrieved from arXiv:1606.01540"},{"key":"6144_CR21","unstructured":"Burda, Y., Edwards, H., Storkey, A., & Klimov, O. (2018). Exploration by random network distillation. Preprint retrieved from arXiv:1810.12894"},{"key":"6144_CR22","doi-asserted-by":"crossref","unstructured":"Carpenter, G. A., & Grossberg, S. (1987). A massively parallel architecture for a self-organizing neural pattern recognition machine. Computer Vision, Graphics, and Image Processing, 37(1), 54\u2013115. http:\/\/techlab.bu.edu\/files\/resources\/articles_cns\/CarpenterGrossberg1987.pdf","DOI":"10.1016\/S0734-189X(87)80014-2"},{"key":"6144_CR23","unstructured":"Chen, Y., Lee, C. W., Luo, H., & Wei, C. Y. (2019). A new algorithm for non-stationary contextual bandits: Efficient, optimal, and parameter-free. In Conference on learning theory (COLT). Preprint retrieved from arXiv:1902.00980"},{"key":"6144_CR24","doi-asserted-by":"crossref","unstructured":"Cheung, W. C., Simchi-Levi, D., & Zhu, R. (2019). Learning to optimize under non-stationarity. In International conference on artificial intelligence and statistics (AISTATS). Preprint retrieved from arXiv:1810.03024","DOI":"10.2139\/ssrn.3261050"},{"key":"6144_CR25","doi-asserted-by":"crossref","unstructured":"Cheung, W. C., Simchi-Levi, D., & Zhu, R. (2019). Reinforcement learning under drift. Preprint retrieved from arXiv:1906.02922","DOI":"10.2139\/ssrn.3397818"},{"key":"6144_CR26","unstructured":"Cisse, M., Bojanowski, P., Grave, E., Dauphin, Y., & Usunier, N. (2017). Parseval networks: Improving robustness to adversarial examples. Preprint retrieved from arXiv:1704.08847"},{"key":"6144_CR27","doi-asserted-by":"crossref","unstructured":"Da\u00a0Silva, B. C., Basso, E. W., Bazzan, A. L. C., & Engel, P. M. (2006). Dealing with non-stationary environments using context detection. In International conference on machine learning (ICML). https:\/\/dl.acm.org\/citation.cfm?id=1143872","DOI":"10.1145\/1143844.1143872"},{"key":"6144_CR28","unstructured":"Dick, T., Gyorgy, A., & Szepesvari, C. (2014). Online learning in Markov decision processes with changing cost sequences. In International conference on machine learning (ICML). http:\/\/proceedings.mlr.press\/v32\/dick14.html"},{"key":"6144_CR29","unstructured":"Dinh, L., Pascanu, R., Bengio, S., & Bengio, Y. (2017). Sharp minima can generalize for deep nets. In International conference on machine learning (ICML). Preprint retrieved from arXiv:1703.04933"},{"key":"6144_CR30","doi-asserted-by":"crossref","unstructured":"Doersch, C., Gupta, A., Efros, A. A. (2015). Unsupervised visual representation learning by context prediction. In International conference on computer vision (ICCV). Preprint retrieved from arXiv:1505.05192","DOI":"10.1109\/ICCV.2015.167"},{"key":"6144_CR31","unstructured":"Du, Y., Czarnecki, W. M., Jayakumar, S. M., Pascanu, R., & Lakshminarayanan, B. (2018). Adapting auxiliary losses using gradient similarity. Preprint retrieved from arXiv:1812.02224"},{"key":"6144_CR32","unstructured":"Duan, Y., Andrychowicz, M., Stadie, B. C., Ho, J., Schneider, J., Sutskever, I., Abbeel, P., & Zaremba, W. (2017). One-shot imitation learning. In Neural Information Processing Systems (NeurIPS). Preprint retrieved from arXiv:1703.07326"},{"key":"6144_CR33","unstructured":"Even-dar, E., Kakade, S. M., & Mansour, Y. (2005). Experts in a Markov decision process. In Neural Information Processing Systems (NeurIPS). http:\/\/papers.nips.cc\/paper\/2730-experts-in-a-markov-decision-process.pdf"},{"key":"6144_CR34","doi-asserted-by":"crossref","unstructured":"Everitt, T., Krakovna, V., Orseau, L., Hutter, M., & Legg, S. (2017). Reinforcement learning with a corrupted reward channel. In International joint conference on artificial intelligence (IJCAI). Preprint retrieved from arXiv:1705.08417.","DOI":"10.24963\/ijcai.2017\/656"},{"key":"6144_CR35","unstructured":"Fernando Hernandez-Garcia, J., & Sutton, R. S. (2019). Understanding multi-step deep reinforcement learning: A systematic study of the DQN target. Preprint retrieved from arXiv:1901.07510"},{"key":"6144_CR36","unstructured":"Finlay, C., Calder, J., Abbasi, B., & Oberman, A. (2018). Lipschitz regularized Deep Neural Networks generalize and are adversarially robust. Preprint retrieved from arXiv:1808.09540"},{"key":"6144_CR37","unstructured":"Fonteneau, R., Murphy, S. A., Wehenkel, L., & Ernst, D. (2010). Model-free Monte Carlo\u2013like policy evaluation. In International conference on artificial intelligence and statistics (AISTATS). http:\/\/proceedings.mlr.press\/v9\/fonteneau10a\/fonteneau10a.pdf"},{"key":"6144_CR38","doi-asserted-by":"crossref","unstructured":"Fonteneau, R., Murphy, S. A., Wehenkel, L., & Ernst, D. (2013). Batch mode reinforcement learning based on the synthesis of artificial trajectories. Annals of Operations Research, 208(1), 383\u2013416. http:\/\/dx.doi.org\/10.1007\/s10479-012-1248-5","DOI":"10.1007\/s10479-012-1248-5"},{"key":"6144_CR39","unstructured":"Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C., & Legg, S. (2017). Noisy networks for exploration. Preprint retrieved from arXiv:1706.10295"},{"key":"6144_CR40","unstructured":"Fu, J., Kumar, A., Soh, M., & Levine, S. (2019). Diagnosing bottlenecks in deep Q-learning algorithms. In International conference on machine learning (ICML). Preprint retrieved from arXiv:1902.10250"},{"key":"6144_CR41","unstructured":"Fujimoto, S., van Hoof, H., & Meger, D. (2018). Addressing function approximation error in actor-critic methods. In International conference on machine learning (ICML). Preprint retrieved from arXiv:1802.09477"},{"key":"6144_CR42","unstructured":"Gajane, P., Ortner, R., & Auer, P. (2018). A sliding-window algorithm for Markov decision processes with arbitrarily changing rewards and transitions. Preprint retrieved from arXiv:1805.10066"},{"key":"6144_CR43","doi-asserted-by":"crossref","unstructured":"Gama, J., \u017dliobait\u0117, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys (CSUR). https:\/\/www.win.tue.nl\/~mpechen\/publications\/pubs\/Gama_ACMCS_AdaptationCD_accepted.pdf","DOI":"10.1145\/2523813"},{"key":"6144_CR44","doi-asserted-by":"crossref","unstructured":"Garivier, A., & Moulines, E. (2011). On upper-confidence bound policies for switching bandit problems. In Algorithmic learning theory (ALT) (pp. 174\u2013188). Springer. http:\/\/dx.doi.org\/10.1007\/978-3-642-24412-4_16","DOI":"10.1007\/978-3-642-24412-4_16"},{"key":"6144_CR45","unstructured":"Goodfellow, I. (2017). NIPS 2016 tutorial: Generative adversarial networks. Preprint retrieved from arXiv:1701.00160"},{"key":"6144_CR46","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. In Neural information processing systems (NIPS). http:\/\/papers.nips.cc\/paper\/5423-generative-adversarial-nets.pdf"},{"key":"6144_CR47","doi-asserted-by":"crossref","unstructured":"Gouk, H., Frank, E., Pfahringer, B., & Cree, M. J. (2021). Regularisation of neural networks by enforcing Lipschitz continuity. Machine Learning, 110(2), 393\u2013416. https:\/\/doi.org\/10.1007\/s10994-020-05929-w","DOI":"10.1007\/s10994-020-05929-w"},{"key":"6144_CR48","doi-asserted-by":"crossref","unstructured":"Gu, Z., Li, Z., Di, X., & Shi, R. (2020). An LSTM-based autonomous driving model using Waymo open dataset. Preprint retrieved from arXiv:2002.05878","DOI":"10.3390\/app10062046"},{"key":"6144_CR49","unstructured":"Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., & Courville, A. (2017). Improved training of Wasserstein GANs. In Neural information processing systems (NIPS). Preprint retrieved from arXiv:1704.00028v3"},{"key":"6144_CR50","unstructured":"Ha, S., Xu, P., Tan, Z., Levine, S., & Tan, J. (2020). Learning to walk in the real world with minimal human effort. Preprint retrieved from arXiv:2002.08550"},{"key":"6144_CR51","doi-asserted-by":"crossref","unstructured":"Hafner, R., & Riedmiller, M. (2011). Reinforcement learning in feedback control. Machine Learning, 84(1\u20132), 137\u2013169. https:\/\/link.springer.com\/article\/10.1007\/s10994-011-5235-x","DOI":"10.1007\/s10994-011-5235-x"},{"key":"6144_CR52","doi-asserted-by":"crossref","unstructured":"Hanna, J. P., & Stone, P. (2017). Grounded action transformation for robot learning in simulation. In AAAI conference on artificial intelligence. https:\/\/www.cs.utexas.edu\/~jphanna\/gsl.pdf","DOI":"10.1609\/aaai.v31i1.11044"},{"key":"6144_CR53","unstructured":"Harada, D. (1997). Reinforcement learning with time. In Conference on artificial intelligence (AAAI). https:\/\/www.aaai.org\/Papers\/AAAI\/1997\/AAAI97-090.pdf"},{"key":"6144_CR54","unstructured":"Hardt, M., Recht, B., & Singer, Y. (2015). Train faster, generalize better: Stability of stochastic gradient descent. Preprint retrieved from arXiv:1509.01240"},{"key":"6144_CR59","doi-asserted-by":"crossref","unstructured":"Heckman, J. J. (1979). Sample selection bias as a specification error. Econometrica, 47(1), 153\u2013161. http:\/\/www.jstor.org\/stable\/1912352","DOI":"10.2307\/1912352"},{"key":"6144_CR60","doi-asserted-by":"crossref","unstructured":"Held, D., McCarthy, Z., Zhang, M., Shentu, F., & Abbeel, P. (2017). Probabilistically safe policy transfer. Preprint retrieved from arXiv:1705.05394","DOI":"10.1109\/ICRA.2017.7989680"},{"key":"6144_CR61","unstructured":"Hernandez, D., & Brown, T. B. (2020). Measuring the algorithmic efficiency of neural networks. https:\/\/cdn.openai.com\/papers\/ai_and_efficiency.pdf"},{"key":"6144_CR62","unstructured":"Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., & Silver, D. (2017). Rainbow: Combining improvements in deep reinforcement learning. Preprint retrieved from arXiv:1710.02298"},{"key":"6144_CR63","unstructured":"Ho, J., & Ermon, S. (2016). Generative adversarial imitation learning. In Neural information processing systems (NIPS). Preprint retrieved from arXiv:1606.03476"},{"key":"6144_CR64","unstructured":"Ho, J., Gupta, J. K., & Ermon, S. (2016). Model-free imitation learning with policy optimization. In International conference on machine learning (ICML). Preprint retrieved from arXiv:1605.08478"},{"key":"6144_CR65","doi-asserted-by":"crossref","unstructured":"Hochreiter, S., & Schmidhuber, J. (1997). Flat minima. Neural Computation, 9(1), 1\u201342. https:\/\/www.ncbi.nlm.nih.gov\/pubmed\/9117894","DOI":"10.1162\/neco.1997.9.1.1"},{"key":"6144_CR66","unstructured":"Ioffe, S., & Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Preprint retrieved from arXiv:1502.03167"},{"key":"6144_CR67","unstructured":"Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., & Kavukcuoglu, K. (2016). Reinforcement learning with unsupervised auxiliary tasks. Preprint retrieved from arXiv:1611.05397"},{"key":"6144_CR68","unstructured":"Jaksch, T., Ortner, R., & Auer, P. (2010). Near-optimal regret bounds for reinforcement learning. Journal of Machine Learning Research (JMLR), 11, 1563\u20131600 . http:\/\/www.jmlr.org\/papers\/volume11\/jaksch10a\/jaksch10a.pdf"},{"key":"6144_CR69","unstructured":"Kaelbling, L. P. (1993). Learning to achieve goals. In International joint conference on artificial intelligence (IJCAI). https:\/\/citeseerx.ist.psu.edu\/viewdoc\/download?doi=10.1.1.51.3077&rep=rep1&type=pdf"},{"key":"6144_CR70","doi-asserted-by":"crossref","unstructured":"Kahn, G., Zhang, T., Levine, S., & Abbeel, P. (2016). PLATO: Policy learning using adaptive trajectory optimization. Preprint retrieved from arXiv:1603.00622","DOI":"10.1109\/ICRA.2017.7989379"},{"key":"6144_CR71","doi-asserted-by":"crossref","unstructured":"Karpathy, A., & Van De\u00a0Panne, M. (2012). Curriculum learning for motor skills. In Advances in artificial intelligence (Canadian conference on artificial intelligence) (pp. 325\u2013330). http:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-642-30353-1.pdf#page=339","DOI":"10.1007\/978-3-642-30353-1_31"},{"key":"6144_CR72","unstructured":"Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., & Tang, P. T. P. (2017). On large-batch training for deep learning: Generalization gap and sharp minima. In International conference on learning representations (ICLR). Preprint retrieved from arXiv:1609.04836"},{"key":"6144_CR73","unstructured":"Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. Preprint retrieved from arXiv:1412.6980"},{"key":"6144_CR74","unstructured":"Kodali, N., Abernethy, J., Hays, J., & Kira, Z. (2017). How to train your DRAGAN. Preprint retrieved from arXiv:1705.07215"},{"key":"6144_CR75","unstructured":"Kostrikov, I., Agrawal, K. K., Dwibedi, D., Levine, S., & Tompson, J. (2019). Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning. In International conference on learning representations (ICLR). Preprint retrieved from arXiv:1809.02925"},{"key":"6144_CR76","unstructured":"Kurach, K., Lucic, M., Zhai, X., Michalski, M., & Gelly, S. (2018). The GAN landscape: Losses, architectures, regularization, and normalization. Preprint retrieved from arXiv:1807.04720"},{"key":"6144_CR77","doi-asserted-by":"publisher","unstructured":"Lange, S., Gabel, T., & Riedmiller, M. (2012). Batch reinforcement learning. In Wiering, M., & van Otterlo, M. (Eds.) Reinforcement learning: State-of-the-art (pp. 45\u201373). Springer. https:\/\/doi.org\/10.1007\/978-3-642-27645-3_2","DOI":"10.1007\/978-3-642-27645-3_2"},{"key":"6144_CR78","unstructured":"Lecarpentier, E., & Rachelson, E. (2019). Non-stationary Markov decision processes, a worst-case approach using model-based reinforcement learning. In Neural information processing systems (NeurIPS). Preprint retrieved from arXiv:1904.10090"},{"key":"6144_CR79","unstructured":"Li, H., Xu, Z., Taylor, G., Studer, C., & Goldstein, T. (2018). Visualizing the loss landscape of neural nets. In Neural information processing systems (NeurIPS). Preprint retrieved from arXiv:1712.09913"},{"key":"6144_CR80","unstructured":"Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., & Wierstra, D. (2016). Continuous control with deep reinforcement learning. In International conference on learning representations (ICLR). Preprint retrieved from arXiv:1509.02971"},{"key":"6144_CR81","unstructured":"Lim, S.H., Xu, H., & Mannor, S. (2013). Reinforcement learning in robust Markov decision processes. In Neural information processing systems (NeurIPS). http:\/\/papers.nips.cc\/paper\/5183-reinforcement-learning-in-robust-markov-decision-processes.pdf"},{"issue":"3","key":"6144_CR82","doi-asserted-by":"publisher","first-page":"293","DOI":"10.1007\/BF00992699","volume":"8","author":"LJ Lin","year":"1992","unstructured":"Lin, L. J. (1992). Self-improving reactive agents based on reinforcement learning, planning and teaching. Machine Learning, 8(3), 293\u2013321. https:\/\/doi.org\/10.1007\/BF00992699","journal-title":"Machine Learning"},{"key":"6144_CR83","unstructured":"Loshchilov, I., & Hutter, F. (2017). Fixing weight decay regularization in Adam. Preprint retrieved from arXiv:1711.05101"},{"key":"6144_CR84","unstructured":"Lucic, M., Kurach, K., Michalski, M., Gelly, S., & Bousquet, O. (2017). Are GANs created equal? A large-scale study. Preprint retrieved from arXiv:1711.10337"},{"key":"6144_CR85","unstructured":"Luo, H., Wei, C. Y., Agarwal, A., & Langford, J. (2018). Efficient contextual bandits in non-stationary worlds. In Conference on learning theory (COLT). Preprint retrieved from arXiv:1708.01799"},{"key":"6144_CR86","unstructured":"Maas, A. L., Hannun, A. Y., & Ng, A. Y. (2013). Rectifier nonlinearities improve neural network acoustic models. In International conference on machine learning (ICML). http:\/\/citeseerx.ist.psu.edu\/viewdoc\/download?doi=10.1.1.693.1422&rep=rep1&type=pdf"},{"key":"6144_CR87","unstructured":"Mescheder, L., Geiger, A., & Nowozin, S. (2018). Which training methods for GANs do actually converge? In International conference on machine learning (ICML). Preprint retrieved from arXiv:1801.04406"},{"key":"6144_CR88","unstructured":"Mirowski, P., Pascanu, R., Viola, F., Soyer, H., Ballard, A. J., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, K., Kumaran, D., & Hadsell, R. (2016). Learning to navigate in complex environments. Preprint retrieved from arXiv:1611.03673"},{"key":"6144_CR89","unstructured":"Miyato, T., Kataoka, T., Koyama, M., & Yoshida, Y. (2018). Spectral normalization for generative adversarial networks. In International conference on learning representations (ICLR). https:\/\/openreview.net\/pdf?id=B1QRgziT-"},{"key":"6144_CR90","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., & Riedmiller, M. (2013). Playing Atari with deep reinforcement learning. Preprint retrieved from arXiv:1312.5602"},{"issue":"7540","key":"6144_CR91","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529\u2013533. https:\/\/doi.org\/10.1038\/nature14236","journal-title":"Nature"},{"key":"6144_CR92","unstructured":"M\u00fcller, R., Kornblith, S., & Hinton, G. E. (2019). When does label smoothing help? In Neural information processing systems (NeurIPS). http:\/\/papers.nips.cc\/paper\/8717-when-does-label-smoothing-help.pdf"},{"key":"6144_CR93","unstructured":"Ng, A. Y., Harada, D., & Russell, S. (1999). Policy invariance under reward transformations: Theory and application to reward shaping. In International conference on machine learning (ICML) (pp. 278\u2013287). http:\/\/www.robotics.stanford.edu\/~ang\/papers\/shaping-icml99.pdf"},{"key":"6144_CR94","unstructured":"Ng, A. Y., & Russell, S. J. (2000). Algorithms for inverse reinforcement learning. In International conference on machine learning (ICML) (pp. 663\u2013670). http:\/\/ai.stanford.edu\/~ang\/papers\/icml00-irl.pdf"},{"issue":"5","key":"6144_CR95","doi-asserted-by":"publisher","first-page":"780","DOI":"10.1287\/opre.1050.0216","volume":"53","author":"A Nilim","year":"2005","unstructured":"Nilim, A., & El Ghaoui, L. (2005). Robust control of Markov decision processes with uncertain transition matrices. Operations Research, 53(5), 780\u2013798. https:\/\/doi.org\/10.1287\/opre.1050.0216","journal-title":"Operations Research"},{"key":"6144_CR96","unstructured":"OpenAI: Solving Rubik\u2019s cube with a robot Hand (2019). https:\/\/d4mucfpksywv.cloudfront.net\/papers\/solving-rubiks-cube.pdf"},{"key":"6144_CR97","unstructured":"Orsini, M., Raichuk, A., Hussenot, L., Vincent, D., Dadashi, R., Girgin, S., Geist, M., Bachem, O., Pietquin, O., & Andrychowicz, M. (2021). What matters for adversarial imitation learning? Preprint retrieved from arXiv:2106.00672"},{"key":"6144_CR98","unstructured":"Ortner, R., Gajane, P., & Auer, P. (2019). Variational regret bounds for reinforcement learning. Preprint retrieved from arXiv:1905.05857"},{"key":"6144_CR99","unstructured":"Padakandla, S., Prabuchandran, K. J., & Bhatnagar, S. (2019). Reinforcement learning in non-stationary environments. Preprint retrieved from arXiv:1905.03970"},{"key":"6144_CR100","unstructured":"Pardo, F., Tavakoli, A., Levdik, V., & Kormushev, P. (2018). Time limits in reinforcement learning. In International conference on machine learning (ICML). Preprint retrieved from arXiv:1712.00378"},{"key":"6144_CR101","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., & Chintala, S. (2019). PyTorch: An imperative style, high-performance deep learning library. In Neural information processing systems (NeurIPS). http:\/\/papers.nips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf"},{"key":"6144_CR102","doi-asserted-by":"crossref","unstructured":"Peng, J., & Williams, R. J. (1996). Incremental multi-step Q-learning. Machine Learning, 22(1-3), 283\u2013290. https:\/\/link.springer.com\/article\/10.1023\/A:1018076709321","DOI":"10.1007\/BF00114731"},{"key":"6144_CR103","unstructured":"Peng, X. B., Kanazawa, A., Toyer, S., Abbeel, P., & Levine, S. (2018). Variational discriminator bottleneck: Improving imitation learning, inverse RL, and GANs by constraining information flow. Preprint retrieved from arXiv:1810.00821"},{"key":"6144_CR104","unstructured":"Pfau, D., & Vinyals, O. (2016). Connecting generative adversarial networks and actor-critic methods. Preprint retrieved from arXiv:1610.01945"},{"key":"6144_CR105","unstructured":"Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R.Y., Chen, X., Asfour, T., Abbeel, P., & Andrychowicz, M. (2018). Parameter space noise for exploration. In International conference on learning representations (ICLR). Preprint retrieved from arXiv:1706.01905"},{"key":"6144_CR106","unstructured":"Pomerleau, D. (1989). ALVINN: An autonomous land vehicle in a neural network. In Neural information processing systems (NIPS) (pp. 305\u2013313). http:\/\/papers.nips.cc\/paper\/95-alvinn-an-autonomous-land-vehicle-in-a-neural-network.pdf"},{"key":"6144_CR107","unstructured":"Pomerleau, D. (1990). Rapidly adapting artificial neural networks for autonomous navigation. In Neural information processing systems (NIPS) (pp. 429\u2013435). http:\/\/papers.nips.cc\/paper\/432-rapidly-adapting-artificial-neural-networks-for-autonomous-navigation.pdf"},{"key":"6144_CR108","doi-asserted-by":"crossref","unstructured":"Puterman, M. L. (1994). Markov decision processes: Discrete stochastic dynamic programming. Wiley.","DOI":"10.1002\/9780470316887"},{"key":"6144_CR109","doi-asserted-by":"publisher","unstructured":"Ratliff, N., Bagnell, J. A., Srinivasa, S. S. (2007). Imitation learning for locomotion and manipulation. In IEEE-RAS international conference on humanoid robots (pp. 392\u2013397). https:\/\/doi.org\/10.1109\/ICHR.2007.4813899","DOI":"10.1109\/ICHR.2007.4813899"},{"key":"6144_CR110","unstructured":"Ray, A., Achiam, J., & Amodei, D. (2019). Benchmarking safe exploration in deep reinforcement learning. https:\/\/d4mucfpksywv.cloudfront.net\/safexp-short.pdf"},{"key":"6144_CR111","unstructured":"Reed, S., Aytar, Y., Wang, Z., Paine, T., van den Oord, A., Pfaff, T., Gomez, S., Novikov, A., Budden, D., & Vinyals, O. (2018). Visual imitation with a minimal adversary. Tech. rep., Deepmind."},{"key":"6144_CR112","doi-asserted-by":"crossref","unstructured":"Robbins, H. (1952). Some aspects of the sequential design of experiments. Bulletin of the American Mathematical Society, 58(5), 527\u2013535. https:\/\/projecteuclid.org\/download\/pdf_1\/euclid.bams\/1183517370","DOI":"10.1090\/S0002-9904-1952-09620-8"},{"key":"6144_CR113","unstructured":"Romoff, J., Henderson, P., Pich\u00e9, A., Francois-Lavet, V., & Pineau, J. (2018). Reward estimation for variance reduction in deep reinforcement learning. In Conference on robot learning (CoRL). Preprint retrieved from arXiv:1805.03359"},{"key":"6144_CR114","unstructured":"Rosca, M., Weber, T., Gretton, A., & Mohamed, S. (2020). A case for new neural network smoothness constraints. In NeurIPS workshop \u201cI Can\u2019t Believe It\u2019s Not Better\u201d. http:\/\/proceedings.mlr.press\/v137\/rosca20a\/rosca20a.pdf"},{"key":"6144_CR115","unstructured":"Ross, S., & Bagnell, J. A. (2010). Efficient reductions for imitation learning. In International conference on artificial intelligence and statistics (AISTATS). http:\/\/www.jmlr.org\/proceedings\/papers\/v9\/ross10a\/ross10a.pdf"},{"key":"6144_CR116","unstructured":"Roth, K., Lucchi, A., Nowozin, S., & Hofmann, T. (2017) Stabilizing training of generative adversarial networks through regularization. In Neural information processing systems (NeurIPS). Preprint retrieved from arXiv:1705.09367v2"},{"key":"6144_CR117","unstructured":"Russac, Y., Vernade, C., & Capp\u00e9, O. (2019). Weighted linear bandits for non-stationary environments. In Neural information processing systems (NeurIPS) . Preprint retrieved from arXiv:1909.09146"},{"key":"6144_CR118","unstructured":"Rusu, A. A., Colmenarejo, S. G., Gulcehre, C., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., & Hadsell, R. (2015). Policy distillation. Preprint retrieved from arXiv:1511.06295"},{"key":"6144_CR119","unstructured":"Saxe, A. M., McClelland, J. L., & Ganguli, S. (2013). Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. Preprint retrieved from arXiv:1312.6120"},{"key":"6144_CR120","unstructured":"Schaal, S. (1997)(1997) Learning from demonstration. In Neural information processing systems (NeurIPS). http:\/\/papers.nips.cc\/paper\/1224-learning-from-demonstration.pdf"},{"key":"6144_CR121","doi-asserted-by":"crossref","unstructured":"Schlimmer, J. C., & Granger Jr., R. H. (1986). Incremental learning from noisy data. Machine Learning. https:\/\/link.springer.com\/article\/10.1007\/BF00116895","DOI":"10.1007\/BF00116895"},{"key":"6144_CR122","unstructured":"Schulman, J., Levine, S., Moritz, P., Jordan, M. I., & Abbeel, P. (2015). Trust region policy optimization. In International conference on machine learning (ICML). Preprint retrieved from arXiv:1502.05477"},{"key":"6144_CR123","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Oleg, K. (2017). Proximal policy optimization algorithms. https:\/\/openai-public.s3-us-west-2.amazonaws.com\/blog\/2017-07\/ppo\/ppo-arxiv.pdf"},{"key":"6144_CR124","unstructured":"Shelhamer, E., Mahmoudieh, P., Argus, M., & Darrell, T. (2016). Loss is its own reward: Self-supervision for reinforcement learning. Preprint retrieved from arXiv:1612.07307"},{"issue":"7587","key":"6144_CR125","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., & Hassabis, D. (2016). Mastering the game of go with deep neural networks and tree search. Nature, 529(7587), 484\u2013489. https:\/\/doi.org\/10.1038\/nature16961","journal-title":"Nature"},{"key":"6144_CR126","unstructured":"Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., & Riedmiller, M. (2014). Deterministic policy gradient algorithms. In International conference on machine learning (ICML) (pp. 387\u2013395). http:\/\/proceedings.mlr.press\/v32\/silver14.html"},{"key":"6144_CR127","unstructured":"Singh, S., Lewis, R. L., & Barto, A. G. (2009). Where do rewards come from? http:\/\/www-anw.cs.umass.edu\/pubs\/2009\/singh_l_b_09.pdf"},{"key":"6144_CR128","unstructured":"Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research, 15(56), 1929\u20131958. https:\/\/jmlr.org\/papers\/v15\/srivastava14a.html"},{"key":"6144_CR129","doi-asserted-by":"crossref","unstructured":"Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., Vasudevan, V., Han, W., Ngiam, J., Zhao, H., Timofeev, A., Ettinger, S., Krivokon, M., Gao, A., Joshi, A., Zhao, S., Cheng, S., Zhang, Y., Shlens, J., Chen, Z., & Anguelov, D. (2019). Scalability in perception for autonomous driving: Waymo open dataset. Preprint retrieved from arXiv:1912.04838","DOI":"10.1109\/CVPR42600.2020.00252"},{"key":"6144_CR130","doi-asserted-by":"crossref","unstructured":"Sutton, R. S. (1988). Learning to predict by the methods of temporal differences. Machine Learning, 3(1), 9\u201344. https:\/\/link.springer.com\/article\/10.1007\/BF00115009","DOI":"10.1007\/BF00115009"},{"key":"6144_CR131","doi-asserted-by":"crossref","unstructured":"Sutton, R. S., & Barto, A. G. (1998). Reinforcement learning: An introduction. MIT Press.","DOI":"10.1109\/TNN.1998.712192"},{"key":"6144_CR132","unstructured":"Sutton, R. S., McAllester, D. A., Singh, S. P., & Mansour, Y. (1999). Policy gradient methods for reinforcement learning with function approximation. In Neural information processing systems (NIPS) (pp. 1057\u20131063). http:\/\/papers.nips.cc\/paper\/1713-policy-gradient-methods-for-reinforcement-learning-with-function-approximation.pdf"},{"key":"6144_CR133","doi-asserted-by":"publisher","unstructured":"Syed, U., Bowling, M., & Schapire, R. E. (2008). Apprenticeship learning using linear programming. In International conference on machine learning (ICML) (pp. 1032\u20131039). https:\/\/doi.org\/10.1145\/1390156.1390286","DOI":"10.1145\/1390156.1390286"},{"key":"6144_CR134","unstructured":"Syed, U., & Schapire, R. E. (2008). A game-theoretic approach to apprenticeship learning. In Neural information processing systems (NIPS) (pp. 1449\u20131456). http:\/\/papers.nips.cc\/paper\/3293-a-game-theoretic-approach-to-apprenticeship-learning.pdf"},{"key":"6144_CR135","unstructured":"Thrun, S., & Schwartz, A. (1993). Issues in using function approximation for reinforcement learning. In Proceedings of the 1993 connectionist models summer school. Lawrence Erlbaum. https:\/\/www.ri.cmu.edu\/pub_files\/pub1\/thrun_sebastian_1993_1\/thrun_sebastian_1993_1.pdf"},{"key":"6144_CR136","doi-asserted-by":"publisher","unstructured":"Todorov, E., Erez, T., & Tassa, Y. (2012). MuJoCo: A physics engine for model-based control. In IEEE\/RSJ international conference on intelligent robots and systems (IROS) (pp. 5026\u20135033). https:\/\/doi.org\/10.1109\/IROS.2012.6386109","DOI":"10.1109\/IROS.2012.6386109"},{"issue":"5","key":"6144_CR137","doi-asserted-by":"publisher","first-page":"823","DOI":"10.1103\/PhysRev.36.823","volume":"36","author":"GE Uhlenbeck","year":"1930","unstructured":"Uhlenbeck, G. E., & Ornstein, L. S. (1930). On the theory of the Brownian motion. Physical Reviews, 36(5), 823\u2013841. https:\/\/doi.org\/10.1103\/PhysRev.36.823","journal-title":"Physical Reviews"},{"key":"6144_CR55","unstructured":"van Hasselt, H. (2010). Double Q-learning. In Neural information processing systems (NeurIPS). https:\/\/papers.nips.cc\/paper\/3964-double-q-learning"},{"key":"6144_CR56","unstructured":"van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., & Modayil, J. (2018). Deep reinforcement learning and the deadly triad. Preprint retrieved from arXiv:1812.02648"},{"key":"6144_CR57","unstructured":"van Hasselt, H., Guez, A., Hessel, M., Mnih, V., & Silver, D. (2016). Learning values across many orders of magnitude. In Neural information processing systems (NeurIPS). Preprint retrieved from arXiv:1602.07714"},{"key":"6144_CR58","unstructured":"van Hasselt, H., Guez, A., & Silver, D. (2015). Deep reinforcement learning with double Q-learning. In AAAI conference on artificial intelligence (pp. 2094\u20132100). Preprint retrieved from arXiv:1509.06461"},{"key":"6144_CR138","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-019-1724-z","author":"O Vinyals","year":"2019","unstructured":"Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., & Silver, D. (2019). Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature. https:\/\/doi.org\/10.1038\/s41586-019-1724-z","journal-title":"Nature"},{"key":"6144_CR139","unstructured":"Wang, R., Ciliberto, C., Amadori, P., & Demiris, Y. (2019). Random expert distillation: Imitation learning via expert policy support estimation. In International conference on machine learning (ICML). Preprint retrieved from arXiv:1905.06750"},{"key":"6144_CR140","unstructured":"Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., & de Freitas, N. (2016). Sample efficient actor-critic with experience replay. Preprint retrieved from arXiv:1611.01224"},{"key":"6144_CR141","unstructured":"Wang, Z., Merel, J., Reed, S., Wayne, G., de\u00a0Freitas, N., & Heess, N. (2017). Robust imitation of diverse behaviors. In Neural information processing systems (NIPS). Preprint retrieved from arXiv:1707.02747"},{"key":"6144_CR142","unstructured":"Watkins, C. J. C. H. (1989). Learning from delayed rewards. Ph.D. thesis, King\u2019s College. http:\/\/www.cs.rhul.ac.uk\/~chrisw\/new_thesis.pdf"},{"issue":"3","key":"6144_CR143","doi-asserted-by":"publisher","first-page":"279","DOI":"10.1023\/A:1022676722315","volume":"8","author":"CJCH Watkins","year":"1992","unstructured":"Watkins, C. J. C. H., & Dayan, P. (1992). Technical note: Q-learning. Machine Learning, 8(3), 279\u2013292. https:\/\/doi.org\/10.1023\/A:1022676722315","journal-title":"Machine Learning"},{"issue":"3","key":"6144_CR144","doi-asserted-by":"publisher","first-page":"363","DOI":"10.1109\/72.286908","volume":"5","author":"AR Webb","year":"1994","unstructured":"Webb, A. R. (1994). Functional approximation by feed-forward networks: a least-squares approach to generalization. IEEE Transactions on Neural Networks, 5(3), 363\u2013371. https:\/\/doi.org\/10.1109\/72.286908","journal-title":"IEEE Transactions on Neural Networks"},{"key":"6144_CR145","unstructured":"Xu, D., & Denil, M. (2019). Positive-unlabeled reward learning. Preprint retrieved from arXiv:1911.00459"},{"key":"6144_CR146","doi-asserted-by":"crossref","unstructured":"Xu, H., & Mannor, S. (2007). The robustness-performance tradeoff in Markov decision processes. In Neural information processing systems (NeurIPS). http:\/\/papers.nips.cc\/paper\/3053-the-robustness-performance-tradeoff-in-markov-decision-processes.pdf","DOI":"10.7551\/mitpress\/7503.003.0197"},{"key":"6144_CR147","unstructured":"Yang, Y. Y., Rashtchian, C., Zhang, H., Salakhutdinov, R., & Chaudhuri, K. (2020). Adversarial robustness through local lipschitzness. Preprint retrieved from arXiv:2003.02460"},{"key":"6144_CR148","doi-asserted-by":"publisher","unstructured":"Yu, J. Y., & Mannor, S. (2009). Arbitrarily modulated Markov decision processes. In Conference on decision and control (CDC) (pp. 2946\u20132953). https:\/\/doi.org\/10.1109\/CDC.2009.5400662","DOI":"10.1109\/CDC.2009.5400662"},{"key":"6144_CR149","doi-asserted-by":"publisher","unstructured":"Yu, J. Y., & Mannor, S. (2009). Online learning in Markov decision processes with arbitrarily changing rewards and transitions. In International conference on game theory for networks (pp. 314\u2013322). https:\/\/doi.org\/10.1109\/GAMENETS.2009.5137416","DOI":"10.1109\/GAMENETS.2009.5137416"},{"key":"6144_CR150","unstructured":"Yu, T., & Sra, S. (2019). Efficient policy learning for non-stationary MDPs under adversarial manipulation. Preprint retrieved from arXiv:1907.09350"},{"key":"6144_CR151","doi-asserted-by":"crossref","unstructured":"Zhang, H., Cisse, M., Dauphin, Y. N., & Lopez-Paz, D. (2017). mixup: Beyond empirical risk minimization. Preprint retrieved from arXiv:1710.09412","DOI":"10.1007\/978-1-4899-7687-1_79"},{"key":"6144_CR152","unstructured":"Zhao, J., Mathieu, M., & LeCun, Y. (2017). Energy-based generative adversarial network. In International conference on learning representations (ICLR). Preprint retrieved from arXiv:1609.03126"},{"key":"6144_CR153","unstructured":"Ziebart, B. D., Maas, A. L., Bagnell, J. A., & Dey, A. K. (2008). Maximum entropy inverse reinforcement learning. In AAAI conference on artificial intelligence (pp. 1433\u20131438). http:\/\/www.aaai.org\/Papers\/AAAI\/2008\/AAAI08-227.pdf"},{"key":"6144_CR154","unstructured":"Zolna, K., Reed, S., Novikov, A., Colmenarej, S. G., Budden, D., Cabi, S., Denil, M., de Freitas, N., & Wang, Z. (2019). Task-relevant adversarial imitation learning. Preprint retrieved from arXiv:1910.01077"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06144-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-022-06144-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06144-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,19]],"date-time":"2023-11-19T16:43:56Z","timestamp":1700412236000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-022-06144-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4]]},"references-count":154,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,4]]}},"alternative-id":["6144"],"URL":"https:\/\/doi.org\/10.1007\/s10994-022-06144-5","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"type":"print","value":"0885-6125"},{"type":"electronic","value":"1573-0565"}],"subject":[],"published":{"date-parts":[[2022,4]]},"assertion":[{"value":"1 August 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 January 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 January 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 April 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}