{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:13:06Z","timestamp":1750219986619,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":50,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,17]],"date-time":"2022-10-17T00:00:00Z","timestamp":1665964800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Tsinghua University Initiative Scientific Research Program, and Tsinghua Precision Medicine Foundation","award":["10001020109"],"award-info":[{"award-number":["10001020109"]}]},{"name":"Technology and Innovation Major Project of the Ministry of Science and Technology of China","award":["2020AAA0108400"],"award-info":[{"award-number":["2020AAA0108400"]}]},{"name":"Technology and Innovation Major Project of the Ministry of Science and Technology of China","award":["2020AAA0108403"],"award-info":[{"award-number":["2020AAA0108403"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,17]]},"DOI":"10.1145\/3511808.3557357","type":"proceedings-article","created":{"date-parts":[[2022,10,16]],"date-time":"2022-10-16T01:29:57Z","timestamp":1665883797000},"page":"128-137","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Imitation Learning to Outperform Demonstrators by Directly Extrapolating Demonstrations"],"prefix":"10.1145","author":[{"given":"Yuanying","family":"Cai","sequence":"first","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chuheng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Shen","sequence":"additional","affiliation":[{"name":"Baidu Inc., Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaonan","family":"He","sequence":"additional","affiliation":[{"name":"Baidu Inc., Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuyun","family":"Zhang","sequence":"additional","affiliation":[{"name":"Macquarie University, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Longbo","family":"Huang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,17]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Kiant\u00e9 Brantley Wen Sun and Mikael Henaff. 2020. Disagreement-regularized imitation learning. In ICLR.  Kiant\u00e9 Brantley Wen Sun and Mikael Henaff. 2020. Disagreement-regularized imitation learning. In ICLR."},{"key":"e_1_3_2_1_2_1","unstructured":"Daniel Brown Wonjoon Goo Prabhat Nagarajan and Scott Niekum. 2019b. Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations. In ICML. 783--792.  Daniel Brown Wonjoon Goo Prabhat Nagarajan and Scott Niekum. 2019b. Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations. In ICML. 783--792."},{"key":"e_1_3_2_1_3_1","unstructured":"Daniel Brown Wonjoon Goo and Scott Niekum. 2019a. Better-than-demonstrator imitation learning via automatically-ranked demonstrations. In CoRL. PMLR 330--359.  Daniel Brown Wonjoon Goo and Scott Niekum. 2019a. Better-than-demonstrator imitation learning via automatically-ranked demonstrations. In CoRL. PMLR 330--359."},{"key":"e_1_3_2_1_4_1","unstructured":"Letian Chen Rohan Paleja and Matthew Gombolay. 2020. Learning from suboptimal demonstration via self-supervised reward regression. In CoRL.  Letian Chen Rohan Paleja and Matthew Gombolay. 2020. Learning from suboptimal demonstration via self-supervised reward regression. In CoRL."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2019.2891173"},{"key":"e_1_3_2_1_6_1","unstructured":"Paul F Christiano Jan Leike Tom Brown Miljan Martic Shane Legg and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In NeurIPS. 4299--4307.  Paul F Christiano Jan Leike Tom Brown Miljan Martic Shane Legg and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In NeurIPS. 4299--4307."},{"key":"e_1_3_2_1_7_1","unstructured":"Marco Cuturi. 2013. Sinkhorn distances: Lightspeed computation of optimal transport. In NeurIPS. 2292--2300.  Marco Cuturi. 2013. Sinkhorn distances: Lightspeed computation of optimal transport. In NeurIPS. 2292--2300."},{"key":"e_1_3_2_1_8_1","unstructured":"Gwendoline De Bie Gabriel Peyr\u00e9 and Marco Cuturi. 2019. Stochastic deep networks. In ICML. PMLR 1556--1565.  Gwendoline De Bie Gabriel Peyr\u00e9 and Marco Cuturi. 2019. Stochastic deep networks. In ICML. PMLR 1556--1565."},{"key":"e_1_3_2_1_9_1","unstructured":"Yannis Flet-Berliac Johan Ferret Olivier Pietquin etal 2021. Adversarially Guided Actor-Critic. In ICLR.  Yannis Flet-Berliac Johan Ferret Olivier Pietquin et al. 2021. Adversarially Guided Actor-Critic. In ICLR."},{"key":"e_1_3_2_1_10_1","unstructured":"Justin Fu Katie Luo and Sergey Levine. 2018. Learning Robust Rewards with Adverserial Inverse Reinforcement Learning. In ICLR.  Justin Fu Katie Luo and Sergey Levine. 2018. Learning Robust Rewards with Adverserial Inverse Reinforcement Learning. In ICLR."},{"key":"e_1_3_2_1_11_1","unstructured":"Tanmay Gangwani and Jian Peng. 2020. State-only Imitation with Transition Dynamics Mismatch. In ICLR.  Tanmay Gangwani and Jian Peng. 2020. State-only Imitation with Transition Dynamics Mismatch. In ICLR."},{"key":"e_1_3_2_1_12_1","unstructured":"Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza etal 2014. Generative adversarial nets. In NeurIPS.  Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza et al. 2014. Generative adversarial nets. In NeurIPS."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Todd Hester Matej Vecerik Olivier Pietquin etal 2018. Deep Q-learning From Demonstrations. In AAAI.  Todd Hester Matej Vecerik Olivier Pietquin et al. 2018. Deep Q-learning From Demonstrations. In AAAI.","DOI":"10.1609\/aaai.v32i1.11757"},{"key":"e_1_3_2_1_14_1","unstructured":"Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. In NeurIPS. 4565--4573.  Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. In NeurIPS. 4565--4573."},{"key":"e_1_3_2_1_15_1","unstructured":"L\u00e9onard Hussenot Marcin Andrychowicz Damien Vincent etal 2021. Hyperparameter selection for imitation learning. In ICML. PMLR 4511--4522.  L\u00e9onard Hussenot Marcin Andrychowicz Damien Vincent et al. 2021. Hyperparameter selection for imitation learning. In ICML. PMLR 4511--4522."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.aau5872"},{"key":"e_1_3_2_1_17_1","unstructured":"Borja Ibarz Jan Leike Tobias Pohlen Geoffrey Irving Shane Legg and Dario Amodei. 2018. Reward learning from human preferences and demonstrations in atari. In NeurIPS. 8011--8023.  Borja Ibarz Jan Leike Tobias Pohlen Geoffrey Irving Shane Legg and Dario Amodei. 2018. Reward learning from human preferences and demonstrations in atari. In NeurIPS. 8011--8023."},{"key":"e_1_3_2_1_18_1","unstructured":"Alexis Jacq Matthieu Geist Ana Paiva and Olivier Pietquin. 2019. Learning from a Learner. In ICML. 2990--2999.  Alexis Jacq Matthieu Geist Ana Paiva and Olivier Pietquin. 2019. Learning from a Learner. In ICML. 2990--2999."},{"key":"e_1_3_2_1_19_1","volume-title":"Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard.","author":"Jaques Natasha","year":"2019","unstructured":"Natasha Jaques , Asma Ghandeharioun , Judy Hanwen Shen , Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard. 2019 . Way of f-policy batch deep reinforcement learning of implicit human preferences in dialog. arXiv preprint arXiv:1907.00456 (2019). Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard. 2019. Way off-policy batch deep reinforcement learning of implicit human preferences in dialog. arXiv preprint arXiv:1907.00456 (2019)."},{"key":"e_1_3_2_1_20_1","volume-title":"A generalised inverse reinforcement learning framework. arXiv preprint arXiv:2105.11812","author":"Jarboui Firas","year":"2021","unstructured":"Firas Jarboui and Vianney Perchet . 2021. A generalised inverse reinforcement learning framework. arXiv preprint arXiv:2105.11812 ( 2021 ). Firas Jarboui and Vianney Perchet. 2021. A generalised inverse reinforcement learning framework. arXiv preprint arXiv:2105.11812 (2021)."},{"key":"e_1_3_2_1_21_1","volume-title":"Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114","author":"Kingma Diederik P","year":"2013","unstructured":"Diederik P Kingma and Max Welling . 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 ( 2013 ). Diederik P Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013)."},{"key":"e_1_3_2_1_22_1","volume-title":"Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson.","author":"Kostrikov Ilya","year":"2019","unstructured":"Ilya Kostrikov , Kumar Krishna Agrawal , Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson. 2019 . Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning. In ICLR. Ilya Kostrikov, Kumar Krishna Agrawal, Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson. 2019. Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning. In ICLR."},{"key":"e_1_3_2_1_23_1","first-page":"11784","article-title":"Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction","volume":"32","author":"Kumar Aviral","year":"2019","unstructured":"Aviral Kumar , Justin Fu , Matthew Soh , George Tucker , and Sergey Levine . 2019 . Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction . Advances in Neural Information Processing Systems , Vol. 32 (2019), 11784 -- 11794 . Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine. 2019. Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction. Advances in Neural Information Processing Systems, Vol. 32 (2019), 11784--11794.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_24_1","volume-title":"Conservative q-learning for offline reinforcement learning. arXiv preprint arXiv:2006.04779","author":"Kumar Aviral","year":"2020","unstructured":"Aviral Kumar , Aurick Zhou , George Tucker , and Sergey Levine . 2020. Conservative q-learning for offline reinforcement learning. arXiv preprint arXiv:2006.04779 ( 2020 ). Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. 2020. Conservative q-learning for offline reinforcement learning. arXiv preprint arXiv:2006.04779 (2020)."},{"key":"e_1_3_2_1_25_1","first-page":"5639","article-title":"CURL: Contrastive Unsupervised Representations for Reinforcement Learning","volume":"119","author":"Laskin Michael","year":"2020","unstructured":"Michael Laskin , Aravind Srinivas , and Pieter Abbeel . 2020 . CURL: Contrastive Unsupervised Representations for Reinforcement Learning . In ICML , Vol. 119. 5639 -- 5650 . Michael Laskin, Aravind Srinivas, and Pieter Abbeel. 2020. CURL: Contrastive Unsupervised Representations for Reinforcement Learning. In ICML, Vol. 119. 5639--5650.","journal-title":"ICML"},{"key":"e_1_3_2_1_26_1","volume-title":"Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643","author":"Levine Sergey","year":"2020","unstructured":"Sergey Levine , Aviral Kumar , George Tucker , and Justin Fu. 2020. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643 ( 2020 ). Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. 2020. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643 (2020)."},{"key":"e_1_3_2_1_27_1","unstructured":"Fangchen Liu Zhan Ling Tongzhou Mu and Hao Su. 2020. State Alignment-based Imitation Learning. In ICLR.  Fangchen Liu Zhan Ling Tongzhou Mu and Hao Su. 2020. State Alignment-based Imitation Learning. In ICLR."},{"key":"e_1_3_2_1_28_1","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"van der Maaten Laurens","year":"2008","unstructured":"Laurens van der Maaten and Geoffrey Hinton . 2008 . Visualizing data using t-SNE . Journal of machine learning research , Vol. 9 , Nov (2008), 2579 -- 2605 . Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of machine learning research, Vol. 9, Nov (2008), 2579--2605.","journal-title":"Journal of machine learning research"},{"key":"e_1_3_2_1_29_1","volume-title":"Inverse Constrained Reinforcement Learning. In International Conference on Machine Learning. PMLR, 7390--7399","author":"Malik Shehryar","year":"2021","unstructured":"Shehryar Malik , Usman Anwar , Alireza Aghasi , and Ali Ahmed . 2021 . Inverse Constrained Reinforcement Learning. In International Conference on Machine Learning. PMLR, 7390--7399 . Shehryar Malik, Usman Anwar, Alireza Aghasi, and Ali Ahmed. 2021. Inverse Constrained Reinforcement Learning. In International Conference on Machine Learning. PMLR, 7390--7399."},{"key":"e_1_3_2_1_30_1","volume-title":"Nature","volume":"518","author":"Mnih Volodymyr","year":"2015","unstructured":"Volodymyr Mnih , Koray Kavukcuoglu , David Silver , 2015 Human-level control through deep reinforcement learning . Nature , Vol. 518 , 7540 (2015), 529--533. Volodymyr Mnih, Koray Kavukcuoglu, David Silver, et al. 2015Human-level control through deep reinforcement learning. Nature, Vol. 518, 7540 (2015), 529--533."},{"key":"e_1_3_2_1_31_1","first-page":"278","article-title":"Policy invariance under reward transformations: Theory and application to reward shaping","volume":"99","author":"Ng Andrew Y","year":"1999","unstructured":"Andrew Y Ng , Daishi Harada , and Stuart Russell . 1999 . Policy invariance under reward transformations: Theory and application to reward shaping . In ICML , Vol. 99. 278 -- 287 . Andrew Y Ng, Daishi Harada, and Stuart Russell. 1999. Policy invariance under reward transformations: Theory and application to reward shaping. In ICML, Vol. 99. 278--287.","journal-title":"ICML"},{"key":"e_1_3_2_1_32_1","volume-title":"NeuIPS","volume":"33","author":"Ramponi Giorgia","year":"2020","unstructured":"Giorgia Ramponi , Gianluca Drappo , and Marcello Restelli . 2020 . Inverse Reinforcement Learning from a Gradient-based Learner . In NeuIPS , Vol. 33 . Giorgia Ramponi, Gianluca Drappo, and Marcello Restelli. 2020. Inverse Reinforcement Learning from a Gradient-based Learner. In NeuIPS, Vol. 33."},{"key":"e_1_3_2_1_33_1","volume-title":"SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards. In ICLR.","author":"Reddy Siddharth","year":"2020","unstructured":"Siddharth Reddy , Anca D Dragan , and Sergey Levine . 2020 . SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards. In ICLR. Siddharth Reddy, Anca D Dragan, and Sergey Levine. 2020. SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards. In ICLR."},{"key":"e_1_3_2_1_34_1","unstructured":"St\u00e9phane Ross and Drew Bagnell. 2010. Efficient reductions for imitation learning. In AISTATS. 661--668.  St\u00e9phane Ross and Drew Bagnell. 2010. Efficient reductions for imitation learning. In AISTATS. 661--668."},{"key":"e_1_3_2_1_35_1","unstructured":"St\u00e9phane Ross Geoffrey Gordon and Drew Bagnell. 2011. A reduction of imitation learning and structured prediction to no-regret online learning. In AISTATS. 627--635.  St\u00e9phane Ross Geoffrey Gordon and Drew Bagnell. 2011. A reduction of imitation learning and structured prediction to no-regret online learning. In AISTATS. 627--635."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"crossref","unstructured":"Wei Shen Xiaonan He Chuheng Zhang Qiang Ni Wanchun Dou and Yan Wang. 2020. Auxiliary-task Based Deep Reinforcement Learning for Participant Selection Problem in Mobile Crowdsourcing. In CIKM. 1355--1364.  Wei Shen Xiaonan He Chuheng Zhang Qiang Ni Wanchun Dou and Yan Wang. 2020. Auxiliary-task Based Deep Reinforcement Learning for Participant Selection Problem in Mobile Crowdsourcing. In CIKM. 1355--1364.","DOI":"10.1145\/3340531.3411913"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33014902"},{"key":"e_1_3_2_1_38_1","volume-title":"Nature","volume":"529","author":"Silver David","year":"2016","unstructured":"David Silver , Aja Huang , Chris J Maddison , 2016 . Mastering the game of Go with deep neural networks and tree search . Nature , Vol. 529 , 7587 (2016), 484--489. David Silver, Aja Huang, Chris J Maddison, et al. 2016. Mastering the game of Go with deep neural networks and tree search. Nature, Vol. 529, 7587 (2016), 484--489."},{"key":"e_1_3_2_1_39_1","unstructured":"Nazneen N Sultana Hardik Meisheri Vinita Baniwal etal 2020. Reinforcement Learning for Multi-Product Multi-Node Inventory Management in Supply Chains. arXiv preprint arXiv:2006.04037 (2020).  Nazneen N Sultana Hardik Meisheri Vinita Baniwal et al. 2020. Reinforcement Learning for Multi-Product Multi-Node Inventory Management in Supply Chains. arXiv preprint arXiv:2006.04037 (2020)."},{"key":"e_1_3_2_1_40_1","unstructured":"Voot Tangkaratt Nontawat Charoenphakdee and Masashi Sugiyama. 2021. Robust Imitation Learning from Noisy Demonstrations. In AISTATS. PMLR 298--306.  Voot Tangkaratt Nontawat Charoenphakdee and Masashi Sugiyama. 2021. Robust Imitation Learning from Noisy Demonstrations. In AISTATS. PMLR 298--306."},{"key":"e_1_3_2_1_41_1","volume-title":"David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al.","author":"Tassa Yuval","year":"2018","unstructured":"Yuval Tassa , Yotam Doron , Alistair Muldal , Tom Erez , Yazhe Li , Diego de Las Casas , David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al. 2018 . Deepmind control suite. arXiv preprint arXiv:1801.00690 (2018). Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al. 2018. Deepmind control suite. arXiv preprint arXiv:1801.00690 (2018)."},{"key":"e_1_3_2_1_42_1","volume-title":"Mujoco: A physics engine for model-based control","author":"Todorov Emanuel","year":"2012","unstructured":"Emanuel Todorov , Tom Erez , and Yuval Tassa . 2012 . Mujoco: A physics engine for model-based control . In IROS. IEEE , 5026--5033. Emanuel Todorov, Tom Erez, and Yuval Tassa. 2012. Mujoco: A physics engine for model-based control. In IROS. IEEE, 5026--5033."},{"key":"e_1_3_2_1_43_1","volume-title":"ICML Workshop.","author":"Torabi Faraz","year":"2019","unstructured":"Faraz Torabi , Garrett Warnell , and Peter Stone . 2019 a. Generative adversarial imitation from observation . In ICML Workshop. Faraz Torabi, Garrett Warnell, and Peter Stone. 2019a. Generative adversarial imitation from observation. In ICML Workshop."},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"crossref","unstructured":"Faraz Torabi Garrett Warnell and Peter Stone. 2019b. Recent advances in imitation learning from observation. In IJCAI. 6325--6331.  Faraz Torabi Garrett Warnell and Peter Stone. 2019b. Recent advances in imitation learning from observation. In IJCAI. 6325--6331.","DOI":"10.24963\/ijcai.2019\/882"},{"key":"e_1_3_2_1_45_1","volume-title":"International Conference on Machine Learning. PMLR, 10961--10970","author":"Wang Yunke","year":"2021","unstructured":"Yunke Wang , Chang Xu , Bo Du , and Honglak Lee . 2021 . Learning to Weight Imperfect Demonstrations . In International Conference on Machine Learning. PMLR, 10961--10970 . Yunke Wang, Chang Xu, Bo Du, and Honglak Lee. 2021. Learning to Weight Imperfect Demonstrations. In International Conference on Machine Learning. PMLR, 10961--10970."},{"key":"e_1_3_2_1_46_1","volume-title":"Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361","author":"Wu Yifan","year":"2019","unstructured":"Yifan Wu , George Tucker , and Ofir Nachum . 2019b. Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361 ( 2019 ). Yifan Wu, George Tucker, and Ofir Nachum. 2019b. Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361 (2019)."},{"key":"e_1_3_2_1_47_1","unstructured":"Yueh-Hua Wu Nontawat Charoenphakdee Han Bao Voot Tangkaratt and Masashi Sugiyama. 2019a. Imitation Learning from Imperfect Demonstration. In ICML. 6818--6827.  Yueh-Hua Wu Nontawat Charoenphakdee Han Bao Voot Tangkaratt and Masashi Sugiyama. 2019a. Imitation Learning from Imperfect Demonstration. In ICML. 6818--6827."},{"key":"e_1_3_2_1_48_1","unstructured":"Xingrui Yu Yueming Lyu and Ivor Tsang. 2020. Intrinsic reward driven imitation learning via generative model. In ICML. PMLR 10925--10935.  Xingrui Yu Yueming Lyu and Ivor Tsang. 2020. Intrinsic reward driven imitation learning via generative model. In ICML. PMLR 10925--10935."},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"crossref","unstructured":"Jiangchuan Zheng Siyuan Liu and Lionel Ni. 2014. Robust bayesian inverse reinforcement learning with sparse behavior noise. In AAAI. 2198--2205.  Jiangchuan Zheng Siyuan Liu and Lionel Ni. 2014. Robust bayesian inverse reinforcement learning with sparse behavior noise. In AAAI. 2198--2205.","DOI":"10.1609\/aaai.v28i1.8979"},{"key":"e_1_3_2_1_50_1","first-page":"1433","article-title":"Maximum entropy inverse reinforcement learning","volume":"8","author":"Ziebart Brian D","year":"2008","unstructured":"Brian D Ziebart , Andrew L Maas , J Andrew Bagnell , and Anind K Dey . 2008 . Maximum entropy inverse reinforcement learning . In AAAI , Vol. 8. 1433 -- 1438 . Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey. 2008. Maximum entropy inverse reinforcement learning. In AAAI, Vol. 8. 1433--1438.","journal-title":"AAAI"}],"event":{"name":"CIKM '22: The 31st ACM International Conference on Information and Knowledge Management","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web","SIGIR ACM Special Interest Group on Information Retrieval"],"location":"Atlanta GA USA","acronym":"CIKM '22"},"container-title":["Proceedings of the 31st ACM International Conference on Information &amp; Knowledge Management"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3511808.3557357","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3511808.3557357","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:29Z","timestamp":1750182569000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3511808.3557357"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,17]]},"references-count":50,"alternative-id":["10.1145\/3511808.3557357","10.1145\/3511808"],"URL":"https:\/\/doi.org\/10.1145\/3511808.3557357","relation":{},"subject":[],"published":{"date-parts":[[2022,10,17]]},"assertion":[{"value":"2022-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}