{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,26]],"date-time":"2026-02-26T13:56:50Z","timestamp":1772114210039,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":35,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,3,17]],"date-time":"2023-03-17T00:00:00Z","timestamp":1679011200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,3,17]]},"DOI":"10.1145\/3594315.3594399","type":"proceedings-article","created":{"date-parts":[[2023,8,3]],"date-time":"2023-08-03T00:14:16Z","timestamp":1691021656000},"page":"733-742","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Stable Control Policy and Transferable Reward Function via Inverse Reinforcement Learning"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8572-6016","authenticated-orcid":false,"given":"Keyu","family":"Wu","sequence":"first","affiliation":[{"name":"Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1958-8321","authenticated-orcid":false,"given":"Fengge","family":"Wu","sequence":"additional","affiliation":[{"name":"Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8173-7116","authenticated-orcid":false,"given":"Yijun","family":"Lin","sequence":"additional","affiliation":[{"name":"Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8950-4157","authenticated-orcid":false,"given":"Junsuo","family":"Zhao","sequence":"additional","affiliation":[{"name":"Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,8,2]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Julian Schrittwieser, Ioannis Antonoglou","author":"Silver David","year":"2016","unstructured":"David Silver , Aja Huang , Chris J Maddison , Arthur Guez , Laurent Sifre , GeorgeVan Den Driessche , Julian Schrittwieser, Ioannis Antonoglou , Veda Panneershelvam, Marc Lanctot , 2016 . Mastering the game of go with deep neural networks and tree search. nature, 529(7587):484\u2013489 David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, GeorgeVan Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, 2016. Mastering the game of go with deep neural networks and tree search. nature, 529(7587):484\u2013489"},{"key":"e_1_3_2_1_2_1","volume-title":"Jonathan Raiman, Tim Salimans, Jeremy Schlatter, Jonas Schneider, Szymon Sidor, Ilya Sutskever, Jie Tang, Filip Wolski, and Susan Zhang.","author":"Berner Christopher","year":"2019","unstructured":"Christopher Berner , Greg Brockman , Brooke Chan , Vicki Cheung , Przemyslaw Debiak , Christy Dennison , David Farhi , Quirin Fischer , Shariq Hashme , Christopher Hesse , Rafal J\u00f3zefowicz , Scott Gray , Catherine Olsson , Jakub Pachocki , Michael Petrov , Henrique Pond\u00e9 de Oliveira Pinto , Jonathan Raiman, Tim Salimans, Jeremy Schlatter, Jonas Schneider, Szymon Sidor, Ilya Sutskever, Jie Tang, Filip Wolski, and Susan Zhang. 2019 . Dota 2 with large scale deep reinforcement learning. CoRR , abs\/1912.06680 Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Christopher Hesse, Rafal J\u00f3zefowicz, Scott Gray, Catherine Olsson, Jakub Pachocki,Michael Petrov, Henrique Pond\u00e9 de Oliveira Pinto, Jonathan Raiman, Tim Salimans, Jeremy Schlatter, Jonas Schneider, Szymon Sidor, Ilya Sutskever, Jie Tang, Filip Wolski, and Susan Zhang. 2019. Dota 2 with large scale deep reinforcement learning. CoRR, abs\/1912.06680"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-019-1724-z"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1177\/0278364919887447"},{"key":"e_1_3_2_1_5_1","volume-title":"Michael A Osborne","author":"Nguyen V","year":"2021","unstructured":"V Nguyen , SB Orbell , Dominic T Lennon , Hyungil Moon , Florian Vigneau , Leon C Camenzind , Liuqi Yu , Dominik M Zumb\u00fchl , G Andrew D Briggs , Michael A Osborne , 2021 . Deep reinforcement learning for efficient measurement of quantum devices. npj Quantum Information , 7(1):1\u20139 V Nguyen, SB Orbell, Dominic T Lennon, Hyungil Moon, Florian Vigneau, Leon C Camenzind, Liuqi Yu, Dominik M Zumb\u00fchl, G Andrew D Briggs, Michael A Osborne, 2021. Deep reinforcement learning for efficient measurement of quantum devices. npj Quantum Information, 7(1):1\u20139"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-021-04301-9"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.7763\/IJCTE.2021.V13.1296"},{"key":"e_1_3_2_1_8_1","volume-title":"Reinforcement learning: An introduction","author":"Sutton Richard S","unstructured":"Richard S Sutton and Andrew G Barto . 2018. Reinforcement learning: An introduction . MIT press , London ,England Richard S Sutton and Andrew G Barto. 2018. Reinforcement learning: An introduction. MIT press, London,England"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/REAL.1997.641292"},{"key":"e_1_3_2_1_10_1","first-page":"20","volume-title":"ICML","volume":"97","author":"Atkeson Christopher G","year":"1997","unstructured":"Christopher G Atkeson and Stefan Schaal . 1997 . Robot learning from demonstration . In ICML , volume 97 , pages 12\u2013 20 Christopher G Atkeson and Stefan Schaal. 1997. Robot learning from demonstration. In ICML, volume 97, pages 12\u201320"},{"key":"e_1_3_2_1_11_1","first-page":"2","volume-title":"Icml","volume":"1","author":"Ng Andrew Y","year":"2000","unstructured":"Andrew Y Ng , Stuart Russell , 2000 . Algorithms for inverse reinforcement learning . In Icml , volume 1 , page 2 Andrew Y Ng, Stuart Russell, 2000. Algorithms for inverse reinforcement learning. In Icml, volume 1, page 2"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015330.1015430"},{"key":"e_1_3_2_1_13_1","first-page":"668","volume-title":"Proceedings of the thirteenth international conference on artificial intelligence and statistics","author":"Ross St\u00e9phane","year":"2010","unstructured":"St\u00e9phane Ross and Drew Bagnell . 2010 . Efficient reductions for imitation learning . In Proceedings of the thirteenth international conference on artificial intelligence and statistics , pages 661\u2013 668 . JMLR Workshop and Conference Proceedings St\u00e9phane Ross and Drew Bagnell. 2010. Efficient reductions for imitation learning. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pages 661\u2013668. JMLR Workshop and Conference Proceedings"},{"key":"e_1_3_2_1_14_1","first-page":"1438","volume-title":"Aaai","volume":"8","author":"Ziebart Brian D","year":"2008","unstructured":"Brian D Ziebart , Andrew L Maas , J Andrew Bagnell , Anind K Dey , 2008 . Maximum entropy inverse reinforcement learning . In Aaai , volume 8 , pages 1433\u2013 1438 . Chicago, IL, USA Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, 2008. Maximum entropy inverse reinforcement learning. In Aaai, volume 8, pages 1433\u20131438. Chicago, IL, USA"},{"key":"e_1_3_2_1_15_1","unstructured":"Markus Wulfmeier Peter Ondruska and Ingmar Posner. 2015. Deep inverse reinforcement learning. CoRR abs\/1906.05274  Markus Wulfmeier Peter Ondruska and Ingmar Posner. 2015. Deep inverse reinforcement learning. CoRR abs\/1906.05274"},{"key":"e_1_3_2_1_16_1","first-page":"58","volume-title":"International conference on machine learning","author":"Finn Chelsea","year":"2016","unstructured":"Chelsea Finn , Sergey Levine , and Pieter Abbeel . 2016 . Guided cost learning: Deep inverse optimal control via policy optimization . In International conference on machine learning , pages 49\u2013 58 . PMLR Chelsea Finn, Sergey Levine, and Pieter Abbeel. 2016. Guided cost learning: Deep inverse optimal control via policy optimization. In International conference on machine learning, pages 49\u201358. PMLR"},{"key":"e_1_3_2_1_17_1","unstructured":"Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. Advances in neural information processing systems 29  Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. Advances in neural information processing systems 29"},{"key":"e_1_3_2_1_18_1","first-page":"1277","volume-title":"Conference on Robot Learning","author":"Seyed Ghasemipour Seyed Kamyar","year":"2020","unstructured":"Seyed Kamyar Seyed Ghasemipour , Richard Zemel , and Shixiang Gu . 2020 . A divergence minimization perspective on imitation learning methods . In Conference on Robot Learning , pages 1259\u2013 1277 . PMLR Seyed Kamyar Seyed Ghasemipour, Richard Zemel, and Shixiang Gu. 2020. A divergence minimization perspective on imitation learning methods. In Conference on Robot Learning, pages 1259\u20131277. PMLR"},{"key":"e_1_3_2_1_19_1","unstructured":"Justin Fu Katie Luo and Sergey Levine. 2017. Learning robust rewards with adversarial inverse reinforcement learning. CoRR abs\/1710.11248  Justin Fu Katie Luo and Sergey Levine. 2017. Learning robust rewards with adversarial inverse reinforcement learning. CoRR abs\/1710.11248"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.2517-6161.1966.tb00626.x"},{"key":"e_1_3_2_1_21_1","unstructured":"Tuomas Haarnoja Aurick Zhou Kristian Hartikainen George Tucker Sehoon Ha Jie Tan Vikash Kumar Henry Zhu Abhishek Gupta Pieter Abbeel 2018. Soft actor-critic algorithms and applications. CoRR abs\/1812.05905  Tuomas Haarnoja Aurick Zhou Kristian Hartikainen George Tucker Sehoon Ha Jie Tan Vikash Kumar Henry Zhu Abhishek Gupta Pieter Abbeel 2018. Soft actor-critic algorithms and applications. CoRR abs\/1812.05905"},{"key":"e_1_3_2_1_22_1","first-page":"223","volume-title":"International conference on machine learning","author":"Arjovsky Martin","year":"2017","unstructured":"Martin Arjovsky , Soumith Chintala , and L\u00e9on Bottou . 2017 . Wasserstein generative adversarial networks . In International conference on machine learning , pages 214\u2013 223 . PMLR Martin Arjovsky, Soumith Chintala, and L\u00e9on Bottou. 2017. Wasserstein generative adversarial networks. In International conference on machine learning, pages 214\u2013223. PMLR"},{"key":"e_1_3_2_1_23_1","first-page":"189","volume-title":"Proceedings of the fourteenth international conference on artificial intelligence and statistics","author":"Boularias Abdeslam","year":"2011","unstructured":"Abdeslam Boularias , Jens Kober , and Jan Peters . 2011 . Relative entropy inverse reinforcement learning . In Proceedings of the fourteenth international conference on artificial intelligence and statistics , pages 182\u2013 189 . JMLR Workshop and Conference Proceedings Abdeslam Boularias, Jens Kober, and Jan Peters. 2011. Relative entropy inverse reinforcement learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pages 182\u2013189. JMLR Workshop and Conference Proceedings"},{"key":"e_1_3_2_1_24_1","unstructured":"Lisa Lee Benjamin Eysenbach Emilio Parisotto Eric Xing Sergey Levine and Ruslan Salakhutdinov. 2019. Efficient exploration via state marginal matching. CoRR abs\/1906.05274  Lisa Lee Benjamin Eysenbach Emilio Parisotto Eric Xing Sergey Levine and Ruslan Salakhutdinov. 2019. Efficient exploration via state marginal matching. CoRR abs\/1906.05274"},{"key":"e_1_3_2_1_25_1","first-page":"817","volume-title":"Energy-based imitation learning","author":"Liu Minghuan","unstructured":"Minghuan Liu , Tairan He , Minkai Xu , and Weinan Zhang . 2021. Energy-based imitation learning . pages 809\u2013 817 Minghuan Liu, Tairan He, Minkai Xu, and Weinan Zhang. 2021. Energy-based imitation learning. pages 809\u2013817"},{"key":"e_1_3_2_1_26_1","first-page":"551","volume-title":"Conference on Robot Learning","author":"Ni Tianwei","year":"2021","unstructured":"Tianwei Ni , Harshit Sikchi , Yufei Wang , Tejus Gupta , Lisa Lee , and Ben Eysenbach . 2021 . f-irl: Inverse reinforcement learning via state marginal matching . In Conference on Robot Learning , pages 529\u2013 551 . PMLR Tianwei Ni, Harshit Sikchi, Yufei Wang, Tejus Gupta, Lisa Lee, and Ben Eysenbach. 2021. f-irl: Inverse reinforcement learning via state marginal matching. In Conference on Robot Learning, pages 529\u2013551. PMLR"},{"key":"e_1_3_2_1_27_1","unstructured":"Chelsea Finn Paul Christiano Pieter Abbeel and Sergey Levine. 2016. A connection between generative adversarial networks inverse reinforcement learning and energy-based models. CoRR abs\/1611.03852  Chelsea Finn Paul Christiano Pieter Abbeel and Sergey Levine. 2016. A connection between generative adversarial networks inverse reinforcement learning and energy-based models. CoRR abs\/1611.03852"},{"key":"e_1_3_2_1_28_1","unstructured":"Martin Arjovsky and L\u00e9on Bottou. 2017. Towards principled methods for training generative adversarial networks. CoRR abs\/1701.04862  Martin Arjovsky and L\u00e9on Bottou. 2017. Towards principled methods for training generative adversarial networks. CoRR abs\/1701.04862"},{"key":"e_1_3_2_1_29_1","first-page":"7593","volume-title":"International Conference on Machine Learning","author":"Zhou Zhiming","year":"2019","unstructured":"Zhiming Zhou , Jiadong Liang , Yuxuan Song , Lantao Yu , Hongwei Wang , Weinan Zhang , Yong Yu , and Zhihua Zhang . 2019 . Lipschitz generative adversarial nets . In International Conference on Machine Learning , pages 7584\u2013 7593 . PMLR Zhiming Zhou, Jiadong Liang, Yuxuan Song, Lantao Yu, Hongwei Wang, Weinan Zhang, Yong Yu, and Zhihua Zhang. 2019. Lipschitz generative adversarial nets. In International Conference on Machine Learning, pages 7584\u20137593. PMLR"},{"key":"e_1_3_2_1_30_1","unstructured":"Ishaan Gulrajani Faruk Ahmed Martin Arjovsky Vincent Dumoulin and Aaron C Courville. 2017. Improved training of wasserstein gans. Advances in neural information processing systems 30  Ishaan Gulrajani Faruk Ahmed Martin Arjovsky Vincent Dumoulin and Aaron C Courville. 2017. Improved training of wasserstein gans. Advances in neural information processing systems 30"},{"key":"e_1_3_2_1_31_1","unstructured":"Jonathan Sorg Richard L Lewis and Satinder Singh. 2010. Reward design via online gradient ascent. Advances in Neural Information Processing Systems 23  Jonathan Sorg Richard L Lewis and Satinder Singh. 2010. Reward design via online gradient ascent. Advances in Neural Information Processing Systems 23"},{"key":"e_1_3_2_1_32_1","unstructured":"Xiaoxiao Guo Satinder Singh Richard Lewis and Honglak Lee. 2016. Deep learning for reward design to improve monte carlo tree search in atari games. CoRR abs\/1604.07095  Xiaoxiao Guo Satinder Singh Richard Lewis and Honglak Lee. 2016. Deep learning for reward design to improve monte carlo tree search in atari games. CoRR abs\/1604.07095"},{"key":"e_1_3_2_1_33_1","unstructured":"Zeyu Zheng Junhyuk Oh and Satinder Singh. 2018. On learning intrinsic rewards for policy gradient methods. Advances in Neural Information Processing Systems 31  Zeyu Zheng Junhyuk Oh and Satinder Singh. 2018. On learning intrinsic rewards for policy gradient methods. Advances in Neural Information Processing Systems 31"},{"key":"e_1_3_2_1_34_1","first-page":"1135","volume-title":"International conference on machine learning","author":"Finn Chelsea","year":"2017","unstructured":"Chelsea Finn , Pieter Abbeel , and Sergey Levine . 2017 . Model-agnostic meta-learning for fast adaptation of deep networks . In International conference on machine learning , pages 1126\u2013 1135 . PMLR Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning, pages 1126\u20131135. PMLR"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"crossref","unstructured":"Ricardo Vilalta and Youssef Drissi. 2002. A perspective view and survey of meta-learning. Artificial intelligence review 18(2):77\u201395  Ricardo Vilalta and Youssef Drissi. 2002. A perspective view and survey of meta-learning. Artificial intelligence review 18(2):77\u201395","DOI":"10.1023\/A:1019956318069"}],"event":{"name":"ICCAI 2023: 2023 9th International Conference on Computing and Artificial Intelligence","location":"Tianjin China","acronym":"ICCAI 2023"},"container-title":["Proceedings of the 2023 9th International Conference on Computing and Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3594315.3594399","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3594315.3594399","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:51:16Z","timestamp":1750182676000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3594315.3594399"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,17]]},"references-count":35,"alternative-id":["10.1145\/3594315.3594399","10.1145\/3594315"],"URL":"https:\/\/doi.org\/10.1145\/3594315.3594399","relation":{},"subject":[],"published":{"date-parts":[[2023,3,17]]},"assertion":[{"value":"2023-08-02","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}