{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T20:30:00Z","timestamp":1785357000519,"version":"3.55.0"},"reference-count":191,"publisher":"Springer Science and Business Media LLC","issue":"17","license":[{"start":{"date-parts":[[2025,3,26]],"date-time":"2025-03-26T00:00:00Z","timestamp":1742947200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,3,26]],"date-time":"2025-03-26T00:00:00Z","timestamp":1742947200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100018647","name":"RUDN University","doi-asserted-by":"publisher","award":["202256-2-000"],"award-info":[{"award-number":["202256-2-000"]}],"id":[{"id":"10.13039\/501100018647","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001779","name":"Monash University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001779","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2025,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Reinforcement learning, characterized by trial-and-error learning and delayed rewards, is central to decision-making processes. Its core component, the reward function, is traditionally handcrafted, but designing these functions is often challenging or impossible in real-world scenarios. Inverse reinforcement learning (IRL) addresses this issue by extracting reward functions from expert demonstrations, facilitating optimal policy derivation and offering a deeper understanding of expert behavior. This comprehensive review focuses on three key aspects: the diverse methodologies employed in IRL, its wide-ranging applications across fields such as robotics, autonomous vehicles, and human intent analysis, and the importance of curated datasets in advancing IRL research. A structured analysis of IRL techniques is provided, applications are categorized by domain, and the role of benchmark datasets in evaluating performance and guiding future developments is emphasized. The unique value of IRL in bridging the gap between human and artificial learning is highlighted, demonstrating its potential to unlock advancements in machine learning, decision making, and explainable AI. By summarizing the current state of IRL research and advocating for future directions, this review serves as a valuable resource for researchers and practitioners seeking to explore and advance the field.<\/jats:p>","DOI":"10.1007\/s00521-025-11100-0","type":"journal-article","created":{"date-parts":[[2025,3,29]],"date-time":"2025-03-29T08:59:32Z","timestamp":1743238772000},"page":"11071-11123","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Advances and applications in inverse reinforcement learning: a comprehensive review"],"prefix":"10.1007","volume":"37","author":[{"given":"Saurabh","family":"Deshpande","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rahee","family":"Walambe","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ketan","family":"Kotecha","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7161-2109","authenticated-orcid":false,"given":"Ganeshsree","family":"Selvachandran","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ajith","family":"Abraham","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,3,26]]},"reference":[{"key":"11100_CR1","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2021.103500","volume":"297","author":"S Arora","year":"2021","unstructured":"Arora S, Doshi P (2021) A survey of inverse reinforcement learning: Challenges, methods and progress. Artif Intell 297:103500. https:\/\/doi.org\/10.1016\/j.artint.2021.103500","journal-title":"Artif Intell"},{"issue":"3","key":"11100_CR2","doi-asserted-by":"publisher","first-page":"293","DOI":"10.1108\/17563781211255862","volume":"5","author":"S Zhifei","year":"2012","unstructured":"Zhifei S, Joo EM (2012) A survey of inverse reinforcement learning techniques. Int J Intell Comput Cybernetics 5(3):293\u2013311. https:\/\/doi.org\/10.1108\/17563781211255862","journal-title":"Int J Intell Comput Cybernetics"},{"key":"11100_CR3","doi-asserted-by":"publisher","first-page":"4307","DOI":"10.1007\/s10462-021-10108-x","volume":"55","author":"S Adams","year":"2022","unstructured":"Adams S, Cody T, Beling PA (2022) A survey of inverse reinforcement learning. Artif Intell Rev 55:4307\u20134346. https:\/\/doi.org\/10.1007\/s10462-021-10108-x","journal-title":"Artif Intell Rev"},{"key":"11100_CR4","unstructured":"R.S. Sutton and A.G. Barto, (2018), Reinforcement learning: An introduction, 2nd Ed, MIT press, Cambridge, MA, USA. https:\/\/mitpress.mit.edu\/9780262039246\/reinforcement-learning\/"},{"key":"11100_CR5","doi-asserted-by":"publisher","unstructured":"C. Szepesv\u00e1ri, (2010), \"Algorithms for reinforcement learning,\" Synthesis Lectures on Artificial Intelligence and Machine Learning. https:\/\/doi.org\/10.1007\/978-3-031-01551-9","DOI":"10.1007\/978-3-031-01551-9"},{"key":"11100_CR6","doi-asserted-by":"publisher","unstructured":"L.P. Kaelbling, M.L. Littman and A.W. Moore, (1996), \"Reinforcement learning: A survey,\" Journal of Artificial Intelligence Research, 4, 237\u2013285. https:\/\/doi.org\/10.5555\/1622737.1622748","DOI":"10.5555\/1622737.1622748"},{"key":"11100_CR7","doi-asserted-by":"publisher","unstructured":"S. Russell, (1998), \"Learning agents for uncertain environments,\" in COLT\u2019 98: Proceedings of the Eleventh Annual Conference on Computational Learning Theory, 101\u2013103. https:\/\/doi.org\/10.1145\/279943.279964","DOI":"10.1145\/279943.279964"},{"key":"11100_CR8","doi-asserted-by":"publisher","unstructured":"A.Y. Ng and S.J. Russell, (2000), \"Algorithms for inverse reinforcement learning,\" in ICML '00: Proceedings of the Seventeenth International Conference on Machine Learning, 663 -670. https:\/\/doi.org\/10.5555\/645529.657801","DOI":"10.5555\/645529.657801"},{"key":"11100_CR9","doi-asserted-by":"publisher","unstructured":"S. Levine, Z. Popovic and V. Koltun, (2010), \"Feature construction for inverse reinforcement learning,\" in NIPS'10: Proceedings of the 23rd International Conference on Neural Information Processing Systems, 1, 1342 \u2013 1350. https:\/\/doi.org\/10.5555\/2997189.2997339","DOI":"10.5555\/2997189.2997339"},{"key":"11100_CR10","doi-asserted-by":"publisher","unstructured":"P. Abbeel and A.Y. Ng, (2004), \"Apprenticeship learning via inverse reinforcement learning,\" in ICML '04: Proceedings of the Twenty-First International Conference on Machine learning, 1. https:\/\/doi.org\/10.1145\/1015330.1015430","DOI":"10.1145\/1015330.1015430"},{"key":"11100_CR11","doi-asserted-by":"publisher","unstructured":"D. Ramachandran and E. Amir, (2007), \"Bayesian inverse reinforcement learning,\" in IJCAI'07: Proceedings of the 20th International Joint Conference on Artificial Intelligence, 2586 \u2013 2591. https:\/\/doi.org\/10.5555\/1625275.1625692","DOI":"10.5555\/1625275.1625692"},{"key":"11100_CR12","doi-asserted-by":"publisher","unstructured":"B.D. Ziebart, A.L. Maas, J.A. Bagnell and A.K. Dey, (2008), \"Maximum entropy inverse reinforcement learning,\" in AAAI'08: Proceedings of the 23rd National Conference on Artificial Intelligence, 3, 1433 \u2013 1438. https:\/\/doi.org\/10.5555\/1620270.1620297","DOI":"10.5555\/1620270.1620297"},{"key":"11100_CR13","doi-asserted-by":"publisher","unstructured":"M. Wulfmeier, P. Ondruska and I. Posner, (2015), \"Deep inverse reinforcement learning,\" arXiv:1507.04888. https:\/\/doi.org\/10.48550\/arXiv.1507.04888","DOI":"10.48550\/arXiv.1507.04888"},{"key":"11100_CR14","unstructured":"C. Finn, S. Levine and P. Abbeel, (2016), \"Guided cost learning: Deep inverse optimal control via policy optimization,\" in Proceedings of The 33rd International Conference on Machine Learning, PMLR 48, 49\u201358. https:\/\/proceedings.mlr.press\/v48\/finn16.html"},{"key":"11100_CR15","doi-asserted-by":"publisher","unstructured":"J. Fu, K. Luo and S. Levine, (2017), \"Learning robust rewards with adversarial inverse reinforcement learning,\" arXiv:1710.11248. https:\/\/doi.org\/10.48550\/arXiv.1710.11248","DOI":"10.48550\/arXiv.1710.11248"},{"key":"11100_CR16","unstructured":"A. Boularias, J. Kober and J. Peters, (2011), \"Relative entropy inverse reinforcement learning,\" in Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, PMLR 15, 182\u2013189. https:\/\/proceedings.mlr.press\/v15\/boularias11a.html"},{"key":"11100_CR17","doi-asserted-by":"publisher","unstructured":"J. Peters, K. Mulling and Y. Altun, (2010), \"Relative entropy policy search,\" in AAAI'10: Proceedings of the Twenty-Fourth AAAI Conference on Artificial Intelligence, 1607 \u2013 1612. https:\/\/doi.org\/10.5555\/2898607.2898863","DOI":"10.5555\/2898607.2898863"},{"key":"11100_CR18","unstructured":"I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville and Y. Bengio, (2014), \"Generative adversarial nets,\" in NIPS 2014, Advances in Neural Information Processing Systems, 27, Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, K.Q. Weinberger (eds.), Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2014\/file\/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf"},{"key":"11100_CR19","unstructured":"B. Settles, (2009), \"Active learning literature survey,\" TR1648, 2009. http:\/\/digital.library.wisc.edu\/1793\/60660"},{"key":"11100_CR20","doi-asserted-by":"publisher","unstructured":"M. Lopes, F. Melo and L. Montesano, (2009), \"Active learning for reward estimation in inverse reinforcement learning,\" in: Buntine, W., Grobelnik, M., Mladeni\u0107, D., Shawe-Taylor, J. (eds) Machine Learning and Knowledge Discovery in Databases. ECML PKDD 2009. Lecture Notes in Computer Science, vol 5782. Springer, Berlin, Heidelberg. https:\/\/doi.org\/10.1007\/978-3-642-04174-7_3","DOI":"10.1007\/978-3-642-04174-7_3"},{"key":"11100_CR21","doi-asserted-by":"publisher","unstructured":"P. Odom and S. Natarajan, (2015), \"Active advice seeking for inverse reinforcement learning,\" in AAAI'15: Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, 4186 \u2013 4187. https:\/\/doi.org\/10.5555\/2888116.2888320","DOI":"10.5555\/2888116.2888320"},{"key":"11100_CR22","unstructured":"D.S. Brown, Y. Cui and S. Niekum, (2018), \"Risk-aware active inverse reinforcement learning,\" in Proceedings of The 2nd Conference on Robot Learning, PMLR 87, 362\u2013372. https:\/\/proceedings.mlr.press\/v87\/brown18a.html"},{"key":"11100_CR23","doi-asserted-by":"publisher","first-page":"1713","DOI":"10.1177\/0278364918772017","volume":"37","author":"S Singh","year":"2018","unstructured":"Singh S, Lacotte J, Majumdar A, Pavone M (2018) Risk-sensitive inverse reinforcement learning via semi-and non-parametric methods. The International Journal of Robotics Research 37:1713\u20131740. https:\/\/doi.org\/10.1177\/0278364918772017","journal-title":"The International Journal of Robotics Research"},{"key":"11100_CR24","doi-asserted-by":"publisher","unstructured":"A. Majumdar, S. Singh, A. Mandlekar and M. Pavone, (2017), \"Risk-sensitive inverse reinforcement learning via coherent risk models,\" in Robotics: Science and Systems, 13, N. Amato, S. Srinivasa, N. Ayanian, S. Kuindersma (eds.), MIT Press Journals. https:\/\/doi.org\/10.15607\/rss.2017.xiii.069","DOI":"10.15607\/rss.2017.xiii.069"},{"key":"11100_CR25","doi-asserted-by":"publisher","unstructured":"R. Chen, W. Wang, Z. Zhao and D. Zhao, (2019), \"Active learning for risk-sensitive inverse reinforcement learning,\" arXiv:1909.07843. https:\/\/doi.org\/10.48550\/arXiv.1909.07843","DOI":"10.48550\/arXiv.1909.07843"},{"key":"11100_CR26","doi-asserted-by":"publisher","unstructured":"P. Henderson, W.-D. Chang, P.-L. Bacon, D. Meger, J. Pineau and D. Precup, (2018), \"Optiongan: Learning joint reward-policy options using generative adversarial inverse reinforcement learning,\" in AAAI'18\/IAAI'18\/EAAI'18: Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Artificial Intelligence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence, 391, 3199\u20133206. https:\/\/doi.org\/10.5555\/3504035.3504426","DOI":"10.5555\/3504035.3504426"},{"key":"11100_CR27","doi-asserted-by":"publisher","unstructured":"J. Ho and S. Ermon, (2016), \"Generative adversarial imitation learning,\" in NIPS'16: Proceedings of the 30th International Conference on Neural Information Processing Systems, 4572 \u2013 4580. https:\/\/doi.org\/10.5555\/3157382.3157608","DOI":"10.5555\/3157382.3157608"},{"key":"11100_CR28","doi-asserted-by":"publisher","unstructured":"P.-L. Bacon, J. Harb and D. Precup, (2017), \"The option-critic architecture,\" in AAAI'17: Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, 1726 \u2013 1734. https:\/\/doi.org\/10.5555\/3298483.3298491","DOI":"10.5555\/3298483.3298491"},{"issue":"1\u20132","key":"11100_CR29","doi-asserted-by":"publisher","first-page":"181","DOI":"10.1016\/S0004-3702(99)00052-1","volume":"112","author":"RS Sutton","year":"1999","unstructured":"Sutton RS, Precup D, Singh S (1999) Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning. Artif Intell 112(1\u20132):181\u2013211. https:\/\/doi.org\/10.1016\/S0004-3702(99)00052-1","journal-title":"Artif Intell"},{"key":"11100_CR30","doi-asserted-by":"publisher","unstructured":"A.H. Qureshi, B. Boots and M.C. Yip, (2018), \"Adversarial imitation via variational inverse reinforcement learning,\" arXiv:1809.06404. https:\/\/doi.org\/10.48550\/arXiv.1809.06404","DOI":"10.48550\/arXiv.1809.06404"},{"key":"11100_CR31","doi-asserted-by":"publisher","unstructured":"S. Mohamed and D.J. Rezende, (2015), \"Variational information maximisation for intrinsically motivated reinforcement learning,\" in NIPS'15: Proceedings of the 28th International Conference on Neural Information Processing Systems, 2, 2125 \u2013 2133. https:\/\/doi.org\/10.5555\/2969442.2969477","DOI":"10.5555\/2969442.2969477"},{"key":"11100_CR32","doi-asserted-by":"publisher","unstructured":"D. Venuto, J. Chakravorty, L. Boussioux, J. Wang, G. McCracken and D. Precup, (2020), \"oIRL: Robust adversarial inverse reinforcement learning with temporally extended actions,\" arXiv:2002.09043. https:\/\/doi.org\/10.48550\/arXiv.2002.09043","DOI":"10.48550\/arXiv.2002.09043"},{"key":"11100_CR33","unstructured":"L. Yu, J. Song and S. Ermon, (2019), \"Multi-agent adversarial inverse reinforcement learning,\" in Proceedings of the 36th International Conference on Machine Learning, Long Beach, California, PMLR 97, 2019. https:\/\/proceedings.mlr.press\/v97\/yu19e\/yu19e-supp.pdf"},{"issue":"1","key":"11100_CR34","doi-asserted-by":"publisher","first-page":"6","DOI":"10.1006\/game.1995.1023","volume":"10","author":"RD McKelvey","year":"1995","unstructured":"McKelvey RD, Palfrey TR (1995) Quantal response equilibria for normal form games. Games Econom Behav 10(1):6\u201338. https:\/\/doi.org\/10.1006\/game.1995.1023","journal-title":"Games Econom Behav"},{"key":"11100_CR35","doi-asserted-by":"publisher","first-page":"9","DOI":"10.1023\/A:1009905800005","volume":"1","author":"RD McKelvey","year":"1998","unstructured":"McKelvey RD, Palfrey TR (1998) Quantal response equilibria for extensive form games. Exp Econ 1:9\u201341. https:\/\/doi.org\/10.1023\/A:1009905800005","journal-title":"Exp Econ"},{"key":"11100_CR36","doi-asserted-by":"publisher","unstructured":"W. Jeon, P. Barde, D. Nowrouzezahrai and J. Pineau, (2020), \"Scalable multi-agent inverse reinforcement learning via actor-attention-critic,\" arXiv:2002.10525. https:\/\/doi.org\/10.48550\/arXiv.2002.10525","DOI":"10.48550\/arXiv.2002.10525"},{"key":"11100_CR37","unstructured":"W. Jeon, C.-Y. Su, P. Barde, T. Doan, D. Nowrouzezahrai and J. Pineau, (2021), \"Regularized inverse reinforcement learning,\" International Conference on Learning Representations, 2021. https:\/\/openreview.net\/forum?id=HgLO8yalfwc"},{"key":"11100_CR38","unstructured":"M. Geist, B. Scherrer and O. Pietquin, (2019), \"A theory of regularized Markov decision processes,\" in Proceedings of the 36th International Conference on Machine Learning, Long Beach, California, PMLR 97, 2019. https:\/\/proceedings.mlr.press\/v97\/geist19a\/geist19a.pdf"},{"key":"11100_CR39","doi-asserted-by":"publisher","unstructured":"L. Zhou and K. Small, (2021), \"Inverse reinforcement learning with natural language goals,\" in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI-21), 35(12), 11116\u201311124. https:\/\/doi.org\/10.1609\/aaai.v35i12.17326","DOI":"10.1609\/aaai.v35i12.17326"},{"key":"11100_CR40","unstructured":"J. Fu, A. Korattikara, S. Levine and S. Guadarrama, \"From language to goals: Inverse reinforcement learning for vision-based instruction following,\" International Conference on Learning Representations (ICLR), 2019. https:\/\/openreview.net\/pdf?id=r1lq1hRqYQ"},{"issue":"2","key":"11100_CR41","doi-asserted-by":"publisher","first-page":"1880","DOI":"10.1109\/LRA.2021.3061397","volume":"6","author":"J Sun","year":"2021","unstructured":"Sun J, Yu L, Dong P, Lu B, Zhou B (2021) Adversarial inverse reinforcement learning with self-attention dynamics model. IEEE Robotics and Automation Letters 6(2):1880\u20131886. https:\/\/doi.org\/10.1109\/LRA.2021.3061397","journal-title":"IEEE Robotics and Automation Letters"},{"key":"11100_CR42","unstructured":"M. Yuan, M.-O. Pun, Y. Chen and Q. Cao, (2021), \"Hybrid adversarial inverse reinforcement learning,\" arXiv:2102.02454. https:\/\/arxiv.org\/html\/2402.08848v1"},{"key":"11100_CR43","doi-asserted-by":"publisher","unstructured":"Q. Qiao and P.A. Beling, (2011), \"Inverse reinforcement learning with Gaussian process,\" in Proceedings of the 2011 American Control Conference, San Francisco, CA, USA, 113\u2013118. https:\/\/doi.org\/10.1109\/ACC.2011.5990948","DOI":"10.1109\/ACC.2011.5990948"},{"key":"11100_CR44","doi-asserted-by":"publisher","unstructured":"C. Dimitrakakis and C.A. Rothkopf, (2012), \"Bayesian multitask inverse reinforcement learning,\" in Sanner, S., Hutter, M. (eds) Recent Advances in Reinforcement Learning. EWRL 2011. Lecture Notes in Computer Science, 7188. Springer, Berlin, Heidelberg. https:\/\/doi.org\/10.1007\/978-3-642-29946-9_27","DOI":"10.1007\/978-3-642-29946-9_27"},{"key":"11100_CR45","unstructured":"A.C.Y. Tossou and C. Dimitrakakis, (2013), \"Probabilistic inverse reinforcement learning in unknown environments,\" in Conference on Uncertainty in Artificial Intelligence (UAI), 2013. https:\/\/auai.org\/uai2013\/prints\/papers\/106.pdf"},{"key":"11100_CR46","unstructured":"J.D. Choi and K.-E. Kim, (2011), \"Inverse reinforcement learning in partially observable environments,\" Journal of Machine Learning Research, 12, p. 691\u2013730. https:\/\/www.jmlr.org\/papers\/volume12\/choi11a\/choi11a.pdf"},{"key":"11100_CR47","doi-asserted-by":"publisher","unstructured":"B. Michini and J.P. How, (2012), \"Bayesian nonparametric inverse reinforcement learning,\" in: Flach, P.A., De Bie, T., Cristianini, N. (eds) Machine Learning and Knowledge Discovery in Databases. ECML PKDD 2012. Lecture Notes in Computer Science, 7524. Springer, Berlin, Heidelberg. https:\/\/doi.org\/10.1007\/978-3-642-33486-3_10","DOI":"10.1007\/978-3-642-33486-3_10"},{"issue":"8","key":"11100_CR48","doi-asserted-by":"publisher","first-page":"4125","DOI":"10.1109\/TNNLS.2021.3051012","volume":"33","author":"M Imani","year":"2022","unstructured":"Imani M, Ghoreishi SF (2022) Scalable inverse reinforcement learning through multifidelity Bayesian optimization. IEEE Transactions on Neural Networks and Learning Systems 33(8):4125\u20134132. https:\/\/doi.org\/10.1109\/TNNLS.2021.3051012","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"issue":"1","key":"11100_CR49","doi-asserted-by":"publisher","first-page":"56","DOI":"10.1109\/TCIAIG.2017.2679115","volume":"10","author":"X Lin","year":"2018","unstructured":"Lin X, Beling PA, Cogill R (2018) Multiagent inverse reinforcement learning for two-person zero-sum games. IEEE Transactions on Games 10(1):56\u201368. https:\/\/doi.org\/10.1109\/TCIAIG.2017.2679115","journal-title":"IEEE Transactions on Games"},{"key":"11100_CR50","doi-asserted-by":"publisher","unstructured":"J. Choi and K.-E. Kim, (2013), \"Bayesian nonparametric feature construction for inverse reinforcement learning,\" in IJCAI '13: Proceedings of the Twenty-Third International Joint Conference on Artificial Intelligence, 1287\u20131293. https:\/\/doi.org\/10.5555\/2540128.2540314","DOI":"10.5555\/2540128.2540314"},{"key":"11100_CR51","unstructured":"J. Audiffren, M. Valko, A. Lazaric and M. Ghavamzadeh, (2015), \"Maximum entropy semi-supervised inverse reinforcement learning,\" in International Joint Conference on Artificial Intelligence, 3315\u20133321. https:\/\/www.ijcai.org\/Proceedings\/15\/Papers\/467.pdf"},{"key":"11100_CR52","doi-asserted-by":"publisher","unstructured":"J.A. Mendez, S. Shivkumar and E. Eaton, (2018), \"Lifelong inverse reinforcement learning,\" in NIPS'18: Proceedings of the 32nd International Conference on Neural Information Processing Systems, 4507\u20134518. https:\/\/doi.org\/10.5555\/3327345.3327362","DOI":"10.5555\/3327345.3327362"},{"key":"11100_CR53","doi-asserted-by":"publisher","unstructured":"S. Krishnan, A. Garg, R. Liaw, L. Miller, F.T. Pokorny and K. Goldberg, (2016), \"HIRL: Hierarchical inverse reinforcement learning for long-horizon tasks with delayed rewards,\" arXiv:1604.06508. https:\/\/doi.org\/10.48550\/arXiv.1604.06508","DOI":"10.48550\/arXiv.1604.06508"},{"key":"11100_CR54","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1007\/s10458-020-09485-4","volume":"35","author":"S Arora","year":"2021","unstructured":"Arora S, Doshi P, Banerjee B (2021) I2RL: online inverse reinforcement learning under occlusion. Auton Agent Multi-Agent Syst 35:4. https:\/\/doi.org\/10.1007\/s10458-020-09485-4","journal-title":"Auton Agent Multi-Agent Syst"},{"key":"11100_CR55","doi-asserted-by":"publisher","unstructured":"S. Arora, B. Banerjee and P. Doshi, (2020), \"Maximum entropy multi-task inverse RL,\" arXiv:2004.12873. https:\/\/doi.org\/10.48550\/arXiv.2004.12873","DOI":"10.48550\/arXiv.2004.12873"},{"key":"11100_CR56","doi-asserted-by":"publisher","unstructured":"M. Troussard, E. Pignat, P. Kamalaruban, S. Calinon and V. Cevher, (2020), \"Interaction-limited inverse reinforcement learning,\" arXiv:2007.00425. https:\/\/doi.org\/10.48550\/arXiv.2007.00425","DOI":"10.48550\/arXiv.2007.00425"},{"issue":"4","key":"11100_CR57","doi-asserted-by":"publisher","first-page":"5355","DOI":"10.1109\/LRA.2020.3005126","volume":"5","author":"Z Wu","year":"2020","unstructured":"Wu Z, Sun L, Zhan W, Yang C, Tomizuka M (2020) Efficient sampling-based maximum entropy inverse reinforcement learning with application to autonomous driving. IEEE Robotics and Automation Letters 5(4):5355\u20135362. https:\/\/doi.org\/10.1109\/LRA.2020.3005126","journal-title":"IEEE Robotics and Automation Letters"},{"key":"11100_CR58","doi-asserted-by":"publisher","unstructured":"J. Inga, E. Bischoff, F. K\u00f6pf and S. Hohmann, (2019), \"Inverse dynamic games based on maximum entropy inverse reinforcement learning,\" arXiv:1911.07503. https:\/\/doi.org\/10.48550\/arXiv.1911.07503","DOI":"10.48550\/arXiv.1911.07503"},{"issue":"1","key":"11100_CR59","doi-asserted-by":"publisher","first-page":"14902","DOI":"10.1016\/j.ifacol.2017.08.2537","volume":"50","author":"F K\u00f6pf","year":"2017","unstructured":"K\u00f6pf F, Inga J, Rothfu\u00df S, Flad M, Hohmann S (2017) Inverse reinforcement learning for identification in linear-quadratic dynamic games. IFAC-PapersOnLine 50(1):14902\u201314908. https:\/\/doi.org\/10.1016\/j.ifacol.2017.08.2537","journal-title":"IFAC-PapersOnLine"},{"issue":"4","key":"11100_CR60","doi-asserted-by":"publisher","first-page":"1966","DOI":"10.1109\/TIT.2012.2234824","volume":"59","author":"BD Ziebart","year":"2013","unstructured":"Ziebart BD, Bagnell JA, Dey AK (2013) The principle of maximum causal entropy for estimating interacting processes. IEEE Trans Inf Theory 59(4):1966\u20131980. https:\/\/doi.org\/10.1109\/TIT.2012.2234824","journal-title":"IEEE Trans Inf Theory"},{"key":"11100_CR61","doi-asserted-by":"publisher","unstructured":"B.D. Ziebart, J.A. Bagnell and A.K. Dey, (2010), \"Modeling interaction via the principle of maximum causal entropy,\" in ICML'10: Proceedings of the 27th International Conference on International Conference on Machine Learning, 1255\u20131262. https:\/\/doi.org\/10.5555\/3104322.3104481","DOI":"10.5555\/3104322.3104481"},{"key":"11100_CR62","doi-asserted-by":"publisher","unstructured":"M. Bloem and N. Bambos, (2014), \"Infinite time horizon maximum causal entropy inverse reinforcement learning,\" in 53rd IEEE Conference on Decision and Control, Los Angeles, CA, USA, 4911\u20134916. https:\/\/doi.org\/10.1109\/CDC.2014.7040156","DOI":"10.1109\/CDC.2014.7040156"},{"key":"11100_CR63","unstructured":"M. Herman, T. Gindele, J. Wagner, F. Schmitt and W. Burgard, (2016), \"Inverse reinforcement learning with simultaneous estimation of rewards and dynamics,\" in Proceedings of the 19th International Conference on Artificial Intelligence and Statistics, PMLR 51, 102\u2013110. https:\/\/proceedings.mlr.press\/v51\/herman16.html"},{"key":"11100_CR64","doi-asserted-by":"publisher","unstructured":"K. Shiarlis, J. Messias and S.A. Whiteson, (2016), \"Inverse reinforcement learning from failure,\" in AAMAS '16: Proceedings of the 2016 International Conference on Autonomous Agents & Multiagent Systems, 1060\u20131068. https:\/\/doi.org\/10.5555\/2936924.2937079","DOI":"10.5555\/2936924.2937079"},{"key":"11100_CR65","doi-asserted-by":"publisher","unstructured":"A. Gleave and O. Habryka, (2018), \"Multi-task maximum entropy inverse reinforcement learning,\" arXiv preprint arXiv:1805.08882. https:\/\/doi.org\/10.48550\/arXiv.1805.08882","DOI":"10.48550\/arXiv.1805.08882"},{"key":"11100_CR66","doi-asserted-by":"publisher","unstructured":"P. Kamalaruban, R. Devidze, V. Cevher and A. Singla, (2019), \"Interactive teaching algorithms for inverse reinforcement learning,\" in IJCAI'19: Proceedings of the 28th International Joint Conference on Artificial Intelligence, 2692\u20132700. https:\/\/doi.org\/10.5555\/3367243.3367414","DOI":"10.5555\/3367243.3367414"},{"key":"11100_CR67","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.ins.2020.01.023","volume":"520","author":"T Zhang","year":"2020","unstructured":"Zhang T, Liu Y, Hwang M, Hwang K-S, Ma C, Cheng J (2020) An end-to-end inverse reinforcement learning by a boosting approach with relative entropy. Inf Sci 520:1\u201314. https:\/\/doi.org\/10.1016\/j.ins.2020.01.023","journal-title":"Inf Sci"},{"key":"11100_CR68","doi-asserted-by":"publisher","unstructured":"S. Natarajan, G. Kunapuli, K. Judah, P. Tadepalli, K. Kersting and J. Shavlik, (2010), \"Multi-agent inverse reinforcement learning,\" in 2010 Ninth International Conference on Machine Learning and Applications, Washington, DC, USA, 395\u2013400. https:\/\/doi.org\/10.1109\/ICMLA.2010.65","DOI":"10.1109\/ICMLA.2010.65"},{"key":"11100_CR69","doi-asserted-by":"publisher","unstructured":"T.S. Reddy, V. Gopikrishna, G. Zaruba and M. Huber, (2012), \"Inverse reinforcement learning for decentralized non-cooperative multiagent systems,\" in 2012 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Seoul, Korea (South), 1930\u20131935. https:\/\/doi.org\/10.1109\/ICSMC.2012.6378020","DOI":"10.1109\/ICSMC.2012.6378020"},{"key":"11100_CR70","doi-asserted-by":"publisher","unstructured":"K. Bogert, J.F.-S. Lin, P. Doshi and D. Kulic, (2016), \"Expectation-maximization for inverse reinforcement learning with hidden data,\" in AAMAS '16: Proceedings of the 2016 International Conference on Autonomous Agents & Multiagent Systems, 1034\u20131042. https:\/\/doi.org\/10.5555\/2936924.2937076","DOI":"10.5555\/2936924.2937076"},{"key":"11100_CR71","doi-asserted-by":"publisher","unstructured":"K. Bogert and P. Doshi, (2017), \"Scaling expectation-maximization for inverse reinforcement learning to multiple robots under occlusion,\" in AAMAS '17: Proceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems, 522\u2013529. https:\/\/doi.org\/10.5555\/3091125.3091202","DOI":"10.5555\/3091125.3091202"},{"key":"11100_CR72","doi-asserted-by":"publisher","unstructured":"J. Hahn and A.M. Zoubir, (2015), \"Inverse reinforcement learning using expectation maximization in mixture models,\" in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, QLD, Australia, 3721\u20133725. https:\/\/doi.org\/10.1109\/ICASSP.2015.7178666","DOI":"10.1109\/ICASSP.2015.7178666"},{"key":"11100_CR73","unstructured":"Q.P. Nguyen, B.K.H. Low and P. Jaillet, (2015), \"Inverse reinforcement learning with locally consistent reward functions,\" in Advances in Neural Information Processing Systems 28 (NIPS 2015), Long Beach, CA, NIPS. http:\/\/hdl.handle.net\/1721.1\/113094"},{"key":"11100_CR74","doi-asserted-by":"publisher","unstructured":"S. Levine, Z. Popovic and V. Koltun, (2011), \"Nonlinear inverse reinforcement learning with gaussian processes,\" in NIPS'11: Proceedings of the 24th International Conference on Neural Information Processing Systems, 19\u201327. https:\/\/doi.org\/10.5555\/2986459.2986462","DOI":"10.5555\/2986459.2986462"},{"key":"11100_CR75","doi-asserted-by":"publisher","unstructured":"D.C. Li, Y.Q. He and F. Fu, (2014), \"Nonlinear inverse reinforcement learning with mutual information and Gaussian process,\" in 2014 IEEE International Conference on Robotics and Biomimetics (ROBIO 2014), Bali, Indonesia, 1445\u20131450. https:\/\/doi.org\/10.1109\/ROBIO.2014.7090537","DOI":"10.1109\/ROBIO.2014.7090537"},{"key":"11100_CR76","doi-asserted-by":"publisher","unstructured":"K. Lee, S. Choi and S. Oh, (2016), \"Inverse reinforcement learning with leveraged gaussian processes,\" in 2016 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Daejeon, Korea (South), 3907\u20133912. https:\/\/doi.org\/10.1109\/IROS.2016.7759575","DOI":"10.1109\/IROS.2016.7759575"},{"key":"11100_CR77","unstructured":"M. Jin, A. Damianou, P. Abbeel and C. Spanos, (2017), \"Inverse reinforcement learning via deep Gaussian process,\" in Conference on Uncertainty in Artificial Intelligence (UAI 2017). https:\/\/auai.org\/uai2017\/proceedings\/papers\/48.pdf"},{"key":"11100_CR78","doi-asserted-by":"publisher","unstructured":"M. Pirotta and M. Restelli, (2016), \"Inverse reinforcement learning through policy gradient minimization,\" in Proceedings of the AAAI Conference on Artificial Intelligence, 30(1). https:\/\/doi.org\/10.1609\/aaai.v30i1.10313","DOI":"10.1609\/aaai.v30i1.10313"},{"key":"11100_CR79","doi-asserted-by":"publisher","unstructured":"G. Ramponi, G. Drappo and M. Restelli, (2020), \"Inverse reinforcement learning from a gradient-based learner,\" in NIPS'20: Proceedings of the 34th International Conference on Neural Information Processing Systems, 207, 2458\u20132468. https:\/\/doi.org\/10.5555\/3495724.3495931","DOI":"10.5555\/3495724.3495931"},{"key":"11100_CR80","doi-asserted-by":"publisher","unstructured":"E. Mazumdar, L.J. Ratliff, T. Fiez and S.S. Sastry, (2017), \"Gradient-based inverse risk-sensitive reinforcement learning,\" in 2017 IEEE 56th Annual Conference on Decision and Control (CDC), Melbourne, VIC, Australia, 5796\u20135801. https:\/\/doi.org\/10.1109\/CDC.2017.8264535","DOI":"10.1109\/CDC.2017.8264535"},{"key":"11100_CR81","unstructured":"N. Das, S. Bechtle, T. Davchev, D. Jayaraman, A. Rai and F. Meier, (2021), \"Model-based inverse reinforcement learning from visual demonstrations,\" in Proceedings of the 2020 Conference on Robot Learning, PMLR 155, 1930\u20131942. https:\/\/proceedings.mlr.press\/v155\/das21a.html"},{"key":"11100_CR82","unstructured":"T. Ni, H. Sikchi, Y. Wang, T. Gupta, L. Lee and B. Eysenbach, (2020), \"F-IRL: Inverse reinforcement learning via state marginal matching,\" in 4th Conference on Robot Learning (CoRL 2020), Cambridge MA, USA. https:\/\/proceedings.mlr.press\/v155\/ni21a\/ni21a.pdf"},{"key":"11100_CR83","unstructured":"A.M. Metelli, M. Pirotta and M. Restelli, (2017), \"Compatible reward inverse reinforcement learning,\" in Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA. https:\/\/papers.nips.cc\/paper_files\/paper\/2017\/file\/e6d8545daa42d5ced125a4bf747b3688-Paper.pdf"},{"key":"11100_CR84","doi-asserted-by":"publisher","unstructured":"A. Boularias, O. Kr\u00f6mer and J. Peters, (2012), \"Structured apprenticeship learning,\" in: Flach, P.A., De Bie, T., Cristianini, N. (eds) Machine Learning and Knowledge Discovery in Databases. ECML PKDD 2012. Lecture Notes in Computer Science, 7524, 227\u2013242. Springer, Berlin, Heidelberg. https:\/\/doi.org\/10.1007\/978-3-642-33486-3_15","DOI":"10.1007\/978-3-642-33486-3_15"},{"key":"11100_CR85","doi-asserted-by":"publisher","unstructured":"L. El Asri, B. Piot, M. Geist, R. Laroche and O. Pietquin, (2016), \"Score-based inverse reinforcement learning,\" in AAMAS '16: Proceedings of the 2016 International Conference on Autonomous Agents & Multiagent Systems, 457\u2013465. https:\/\/doi.org\/10.5555\/2936924.2936991","DOI":"10.5555\/2936924.2936991"},{"key":"11100_CR86","doi-asserted-by":"publisher","unstructured":"L. Yu, T. Yu, C. Finn and S. Ermon, (2019), \"Meta-inverse reinforcement learning with probabilistic context variables,\" in Proceedings of the 33rd International Conference on Neural Information Processing Systems, 1054, 11772\u201311783. https:\/\/doi.org\/10.5555\/3454287.3455341","DOI":"10.5555\/3454287.3455341"},{"key":"11100_CR87","doi-asserted-by":"publisher","first-page":"83","DOI":"10.1016\/j.patrec.2022.01.016","volume":"154","author":"R Hwang","year":"2022","unstructured":"Hwang R, Lee H, Hwang HJ (2022) Option compatible reward inverse reinforcement learning. Pattern Recogn Lett 154:83\u201389. https:\/\/doi.org\/10.1016\/j.patrec.2022.01.016","journal-title":"Pattern Recogn Lett"},{"key":"11100_CR88","doi-asserted-by":"publisher","unstructured":"V. Freire Da Silva, A.H. Reali Costa and P. Lima, (2006), \"Inverse reinforcement learning with evaluation,\" in Proceedings 2006 IEEE International Conference on Robotics and Automation, 2006. ICRA 2006, Orlando, FL, USA, 4246\u20134251. https:\/\/doi.org\/10.1109\/ROBOT.2006.1642355","DOI":"10.1109\/ROBOT.2006.1642355"},{"key":"11100_CR89","doi-asserted-by":"publisher","unstructured":"M. Sharma, K.M. Kitani and J. Groeger, (2018), \"Inverse reinforcement learning with conditional choice probabilities,\" in Proceedings of RSS '18 Workshop on Perspectives in Robot Learning: Causality and Imitation. https:\/\/doi.org\/10.48550\/arXiv.1709.07597","DOI":"10.48550\/arXiv.1709.07597"},{"key":"11100_CR90","doi-asserted-by":"publisher","unstructured":"L. Haug, I. Ovinnikon and E. Bykovets, (2020), \"Inverse reinforcement learning via matching of optimality profiles,\" arXiv:2011.09264. https:\/\/doi.org\/10.48550\/arXiv.2011.09264","DOI":"10.48550\/arXiv.2011.09264"},{"key":"11100_CR91","doi-asserted-by":"publisher","unstructured":"E. Uchibe and K. Doya, (2014), \"Inverse reinforcement learning using dynamic policy programming,\" in Proceedings of the 4th International Conference on Development and Learning and on Epigenetic Robotics, Genoa, Italy, 222\u2013228. https:\/\/doi.org\/10.1109\/DEVLRN.2014.6982985","DOI":"10.1109\/DEVLRN.2014.6982985"},{"key":"11100_CR92","unstructured":"E. Klein, M. Geist, B. Piot and O. Pietquin, (2012), \"Inverse reinforcement learning through structured classification,\" in Proceedings of NIPS 2012, Advances in Neural Information Processing Systems, 25 (NIPS 2012). https:\/\/papers.nips.cc\/paper_files\/paper\/2012\/hash\/559cb990c9dffd8675f6bc2186971dc2-Abstract.html"},{"key":"11100_CR93","doi-asserted-by":"publisher","unstructured":"E. Klein, B. Piot, M. Geist and O. Pietquin, (2013), \"A cascaded supervised learning approach to inverse reinforcement learning,\" in: Blockeel, H., Kersting, K., Nijssen, S., \u017delezn\u00fd, F. (eds) Machine Learning and Knowledge Discovery in Databases. ECML PKDD 2013. Lecture Notes in Computer Science, 8188, 1\u201316. Springer, Berlin, Heidelberg. https:\/\/doi.org\/10.1007\/978-3-642-40988-2_1","DOI":"10.1007\/978-3-642-40988-2_1"},{"key":"11100_CR94","doi-asserted-by":"publisher","unstructured":"A. \u0160o\u0161i\u0107, W.R. KhudaBukhsh, A.M. Zoubir and H. Koeppl, (2017), \"Inverse reinforcement learning in swarm systems,\" in AAMAS '17: Proceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems, 1413 \u2013 1421. https:\/\/doi.org\/10.5555\/3091125.3091320","DOI":"10.5555\/3091125.3091320"},{"key":"11100_CR95","doi-asserted-by":"publisher","unstructured":"K. Li and J.W. Burdick, (2017), \"Inverse reinforcement learning in large state spaces via function approximation,\" arXiv:1707.09394, 2017. https:\/\/doi.org\/10.48550\/arXiv.1707.09394","DOI":"10.48550\/arXiv.1707.09394"},{"key":"11100_CR96","doi-asserted-by":"publisher","first-page":"2295","DOI":"10.1007\/s10994-021-05984-x","volume":"110","author":"P Korsunsky","year":"2021","unstructured":"Korsunsky P, Belogolovsky S, Zahavy T, Tessler C, Mannor S (2021) Inverse reinforcement learning in contextual MDPs. Mach Learn 110:2295\u20132334. https:\/\/doi.org\/10.1007\/s10994-021-05984-x","journal-title":"Mach Learn"},{"key":"11100_CR97","doi-asserted-by":"publisher","unstructured":"A. Hallak, D. Di Castro and S. Mannor, (2015), \"Contextual Markov decision processes,\" arXiv:1502.02259. https:\/\/doi.org\/10.48550\/arXiv.1502.02259","DOI":"10.48550\/arXiv.1502.02259"},{"key":"11100_CR98","doi-asserted-by":"publisher","first-page":"1517","DOI":"10.1007\/s10994-018-5730-4","volume":"107","author":"A Kangasr\u00e4\u00e4si\u00f6","year":"2018","unstructured":"Kangasr\u00e4\u00e4si\u00f6 A, Kaski S (2018) Inverse reinforcement learning from summary data. Mach Learn 107:1517\u20131535. https:\/\/doi.org\/10.1007\/s10994-018-5730-4","journal-title":"Mach Learn"},{"key":"11100_CR99","unstructured":"A. Jacq, M. Geist, A. Paiva and O. Pietquin, (2019), \"Learning from a learner,\" in Proceedings of the 36th International Conference on Machine Learning, PMLR 97, 2990\u20132999. https:\/\/proceedings.mlr.press\/v97\/jacq19a.html"},{"key":"11100_CR100","unstructured":"D. Brown, W. Goo, P. Nagarajan and S. Niekum, (2019), \"Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations,\" in Proceedings of the 36th International Conference on Machine Learning, Long Beach, California, PMLR 97. https:\/\/proceedings.mlr.press\/v97\/brown19a\/brown19a.pdf"},{"key":"11100_CR101","doi-asserted-by":"publisher","unstructured":"S. Choi, K. Lee, A. Park and S. Oh, (2016), \"Density matching reward learning,\" arXiv:1608.03694. https:\/\/doi.org\/10.48550\/arXiv.1608.03694","DOI":"10.48550\/arXiv.1608.03694"},{"key":"11100_CR102","doi-asserted-by":"publisher","unstructured":"T. Munzer, B. Piot, M. Geist, O. Pietquin and M. Lopes, (2015), \"Inverse reinforcement learning in relational domains,\" in IJCAI'15: Proceedings of the 24th International Conference on Artificial Intelligence, 3735\u20133741. https:\/\/doi.org\/10.5555\/2832747.2832770","DOI":"10.5555\/2832747.2832770"},{"key":"11100_CR103","doi-asserted-by":"publisher","unstructured":"D. Hadfield-Menell, A. Dragan, P. Abbeel and S. Russell, (2016), \"Cooperative inverse reinforcement learning,\" in NIPS'16: Proceedings of the 30th International Conference on Neural Information Processing Systems, 3916\u20133924. https:\/\/doi.org\/10.5555\/3157382.3157535","DOI":"10.5555\/3157382.3157535"},{"key":"11100_CR104","doi-asserted-by":"publisher","unstructured":"M. Babes-Vroman, V. Marivate, K. Subramanian and M. Littman, (2011), \"Apprenticeship learning about multiple intentions,\" in ICML'11: Proceedings of the 28th International Conference on International Conference on Machine Learning, 897\u2013904. https:\/\/doi.org\/10.5555\/3104482.3104595","DOI":"10.5555\/3104482.3104595"},{"key":"11100_CR105","doi-asserted-by":"publisher","unstructured":"V. Jain, P. Doshi and B. Banerjee, (2019), \"Model-free IRL using maximum likelihood estimation,\" in AAAI'19\/IAAI'19\/EAAI'19: Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, 485, 3951\u20133958. https:\/\/doi.org\/10.1609\/aaai.v33i01.33013951","DOI":"10.1609\/aaai.v33i01.33013951"},{"key":"11100_CR106","doi-asserted-by":"publisher","unstructured":"I. Bica, D. Jarrett, A. H\u00fcy\u00fck and M. van der Schaar, (2020), \"Batch inverse reinforcement learning using counterfactuals for understanding decision making,\" arXiv:2007.13531. https:\/\/doi.org\/10.48550\/arXiv.2007.13531","DOI":"10.48550\/arXiv.2007.13531"},{"key":"11100_CR107","unstructured":"X. Wang and D. Klabjan, (2018), \"Competitive multi-agent inverse reinforcement learning with sub-optimal demonstrations,\" in Proceedings of the 35th International Conference on Machine Learning, PMLR 80, 5143\u20135151. https:\/\/proceedings.mlr.press\/v80\/wang18d.html"},{"key":"11100_CR108","doi-asserted-by":"publisher","unstructured":"R. Self, K. Coleman, H. Bai and R. Kamalapurkar, (2021), \"Online observer-based inverse reinforcement learning,\" in 2021 American Control Conference (ACC), New Orleans, LA, USA, 1959\u20131964. https:\/\/doi.org\/10.23919\/ACC50511.2021.9482906","DOI":"10.23919\/ACC50511.2021.9482906"},{"key":"11100_CR109","doi-asserted-by":"publisher","unstructured":"R. Self, M. Abudia and R. Kamalapurkar, (2020), \"Online inverse reinforcement learning for systems with disturbances,\" in 2020 American Control Conference (ACC), Denver, CO, USA, 1118\u20131123. https:\/\/doi.org\/10.23919\/ACC45564.2020.9147344","DOI":"10.23919\/ACC45564.2020.9147344"},{"key":"11100_CR110","doi-asserted-by":"publisher","unstructured":"E. Klein, M. Geist and O. Pietquin, (2011), \"Batch, off-policy and model-free apprenticeship learning,\" in: Sanner, S., Hutter, M. (eds) Recent Advances in Reinforcement Learning. EWRL 2011. Lecture Notes in Computer Science, 7188, 285\u2013296. Springer, Berlin, Heidelberg. https:\/\/doi.org\/10.1007\/978-3-642-29946-9_28","DOI":"10.1007\/978-3-642-29946-9_28"},{"key":"11100_CR111","doi-asserted-by":"publisher","unstructured":"M. Kuderer, S. Gulati and W. Burgard, (2015), \"Learning driving styles for autonomous vehicles from demonstration,\" in 2015 IEEE International Conference on Robotics and Automation (ICRA), Seattle, WA, USA, 2641\u20132646. https:\/\/doi.org\/10.1109\/ICRA.2015.7139555","DOI":"10.1109\/ICRA.2015.7139555"},{"key":"11100_CR112","doi-asserted-by":"publisher","unstructured":"S. Rosbach, V. James, S. Gro\u00dfjohann, S. Homoceanu and S. Roth, (2019), \"Driving with style: inverse reinforcement learning in general-purpose planning for automated driving,\" in 2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 2658\u20132665. https:\/\/doi.org\/10.1109\/IROS40897.2019.8968205","DOI":"10.1109\/IROS40897.2019.8968205"},{"key":"11100_CR113","doi-asserted-by":"publisher","unstructured":"S. Rosbach, V. James, S. Gro\u00dfjohann, S. Homoceanu, X. Li and S. Roth, (2020), \"Driving style encoder: Situational reward adaptation for general-purpose planning in automated driving,\" in 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 6419\u20136425. https:\/\/doi.org\/10.1109\/ICRA40945.2020.9196778","DOI":"10.1109\/ICRA40945.2020.9196778"},{"key":"11100_CR114","doi-asserted-by":"publisher","unstructured":"S. Rosbach, X. Li, S. Gro\u00dfjohann, S. Homoceanu and S. Roth, (2020), \"Planning on the fast lane: Learning to interact using attention mechanisms in path integral inverse reinforcement learning,\" in 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 5187\u20135193. https:\/\/doi.org\/10.1109\/IROS45743.2020.9340636","DOI":"10.1109\/IROS45743.2020.9340636"},{"issue":"8","key":"11100_CR115","doi-asserted-by":"publisher","first-page":"10239","DOI":"10.1109\/TITS.2021.3088935","volume":"23","author":"Z Huang","year":"2022","unstructured":"Huang Z, Wu J, Lv C (2022) Driving behavior modeling using naturalistic human driving data with inverse reinforcement learning. IEEE Trans Intell Transp Syst 23(8):10239\u201310251. https:\/\/doi.org\/10.1109\/TITS.2021.3088935","journal-title":"IEEE Trans Intell Transp Syst"},{"key":"11100_CR116","doi-asserted-by":"publisher","unstructured":"D. Kishikawa and S. Arai, (2019), \"Comfortable driving by using deep inverse reinforcement learning,\" in 2019 IEEE International Conference on Agents (ICA), Jinan, China, 38\u201343. https:\/\/doi.org\/10.1109\/AGENTS.2019.8929214","DOI":"10.1109\/AGENTS.2019.8929214"},{"issue":"1","key":"11100_CR117","doi-asserted-by":"publisher","first-page":"87","DOI":"10.1109\/MSP.2020.2988287","volume":"38","author":"T Fernando","year":"2021","unstructured":"Fernando T, Denman S, Sridharan S, Fookes C (2021) Deep inverse reinforcement learning for behavior prediction in autonomous driving: accurate forecasts of vehicle motion. IEEE Signal Process Mag 38(1):87\u201396. https:\/\/doi.org\/10.1109\/MSP.2020.2988287","journal-title":"IEEE Signal Process Mag"},{"key":"11100_CR118","doi-asserted-by":"publisher","unstructured":"L. Sun, W. Zhan and M. Tomizuka, (2018), \"Probabilistic prediction of interactive driving behavior via hierarchical inverse reinforcement learning,\" in 2018 21st International Conference on Intelligent Transportation Systems (ITSC), Maui, HI, USA, 2111\u20132117. https:\/\/doi.org\/10.1109\/ITSC.2018.8569453","DOI":"10.1109\/ITSC.2018.8569453"},{"key":"11100_CR119","doi-asserted-by":"publisher","unstructured":"M. Shimosaka, K. Nishi, J. Sato and H. Kataoka, (2015), \"Predicting driving behavior using inverse reinforcement learning with multiple reward functions towards environmental diversity,\" in 2015 IEEE Intelligent Vehicles Symposium (IV), Seoul, Korea (South), 567\u2013572. https:\/\/doi.org\/10.1109\/IVS.2015.7225745","DOI":"10.1109\/IVS.2015.7225745"},{"key":"11100_CR120","doi-asserted-by":"publisher","unstructured":"L. Xin, S.E. Li, P. Wang, W. Cao, B. Nie, C.-Y. Chan and B. Cheng, (2019), \"Accelerated inverse reinforcement learning with randomly pre-sampled policies for autonomous driving reward design,\" in 2019 IEEE Intelligent Transportation Systems Conference (ITSC), Auckland, New Zealand, 2757\u20132764. https:\/\/doi.org\/10.1109\/ITSC.2019.8916952","DOI":"10.1109\/ITSC.2019.8916952"},{"key":"11100_CR121","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.robot.2019.01.003","volume":"114","author":"C You","year":"2019","unstructured":"You C, Lu J, Filev D, Tsiotras P (2019) Advanced planning for autonomous vehicles using reinforcement learning and deep inverse reinforcement learning. Robot Auton Syst 114:1\u201318. https:\/\/doi.org\/10.1016\/j.robot.2019.01.003","journal-title":"Robot Auton Syst"},{"key":"11100_CR122","doi-asserted-by":"publisher","unstructured":"D. Choi, T.-H. An, K. Ahn and J. Choi, (2018), \"Future trajectory prediction via RNN and maximum margin inverse reinforcement learning,\" in 2018 17th IEEE International Conference on Machine Learning and Applications (ICMLA), Orlando, FL, USA, 125\u2013130. https:\/\/doi.org\/10.1109\/ICMLA.2018.00026","DOI":"10.1109\/ICMLA.2018.00026"},{"key":"11100_CR123","doi-asserted-by":"publisher","unstructured":"P. Wang, D. Liu, J. Chen, H. Li and C.-Y. Chan, (2019), \"Human-like decision making for autonomous driving via adversarial inverse reinforcement learning,\" arXiv:1911.08044. https:\/\/doi.org\/10.48550\/arXiv.1911.08044","DOI":"10.48550\/arXiv.1911.08044"},{"key":"11100_CR124","doi-asserted-by":"publisher","unstructured":"K. Sama, Y. Morales, N. Akai, E. Takeuchi and K. Takeda, (2018), \"Learning how to drive in blind intersections from human data,\" in 2018 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Miyazaki, Japan, 317\u2013324. https:\/\/doi.org\/10.1109\/SMC.2018.00064","DOI":"10.1109\/SMC.2018.00064"},{"key":"11100_CR125","doi-asserted-by":"publisher","unstructured":"Z. Zhu, N. Li, R. Sun, D. Xu and H. Zhao, (2020), \"Off-road autonomous vehicles traversability analysis and trajectory planning based on deep inverse reinforcement learning,\" in 2020 IEEE Intelligent Vehicles Symposium (IV), Las Vegas, NV, USA, 971\u2013977. https:\/\/doi.org\/10.1109\/IV47402.2020.9304721","DOI":"10.1109\/IV47402.2020.9304721"},{"key":"11100_CR126","doi-asserted-by":"publisher","unstructured":"M. Shimosaka, T. Kaneko and K. Nishi, (2014), \"Modeling risk anticipation and defensive driving on residential roads with inverse reinforcement learning,\" in 17th International IEEE Conference on Intelligent Transportation Systems (ITSC), Qingdao, China, 1694\u20131700. https:\/\/doi.org\/10.1109\/ITSC.2014.6957937","DOI":"10.1109\/ITSC.2014.6957937"},{"key":"11100_CR127","doi-asserted-by":"publisher","DOI":"10.1016\/j.trc.2024.104572","volume":"161","author":"AR Alozi","year":"2024","unstructured":"Alozi AR, Hussein M (2024) How do active road users act around autonomous vehicles? An inverse reinforcement learning approach. Transportation Research Part C: Emerging Technologies 161:104572. https:\/\/doi.org\/10.1016\/j.trc.2024.104572","journal-title":"Transportation Research Part C: Emerging Technologies"},{"key":"11100_CR128","doi-asserted-by":"publisher","unstructured":"S.J. Lee and Z. Popovi\u0107, (2010), \"Learning behavior styles with inverse reinforcement learning,\" in Proceedings of SIGGRAPH '10: Special Interest Group on Computer Graphics and Interactive Techniques Conference, Los Angeles, California, 122, 1 \u2013 7. https:\/\/doi.org\/10.1145\/1833349.1778859","DOI":"10.1145\/1833349.1778859"},{"key":"11100_CR129","doi-asserted-by":"publisher","unstructured":"S. Liu, M. Araujo, E. Brunskill, R. Rossetti, J. Barros and R. Krishnan, (2013), \"Understanding sequential decisions via inverse reinforcement learning,\" in 2013 IEEE 14th International Conference on Mobile Data Management, Milan, Italy, 177\u2013186. https:\/\/doi.org\/10.1109\/MDM.2013.28","DOI":"10.1109\/MDM.2013.28"},{"key":"11100_CR130","doi-asserted-by":"publisher","unstructured":"S. Das and A. Lavoie, (2014), \"The effects of feedback on human behavior in social media: An inverse reinforcement learning model,\" in AAMAS '14: Proceedings of the 2014 International Conference on Autonomous Agents and Multi-Agent Systems, 653\u2013660. https:\/\/doi.org\/10.5555\/2615731.2615837","DOI":"10.5555\/2615731.2615837"},{"key":"11100_CR131","doi-asserted-by":"publisher","unstructured":"M. Pan, W. Huang, Y. Li, X. Zhou and J. Luo, (2020), \"xGAIL: Explainable generative adversarial imitation learning for explainable human decision analysis,\" in KDD \u201920: Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2020, 1334\u20131343. https:\/\/doi.org\/10.1145\/3394486.3403186","DOI":"10.1145\/3394486.3403186"},{"issue":"1","key":"11100_CR132","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1609\/aiide.v15i1.5244","volume":"15","author":"B Wang","year":"2019","unstructured":"Wang B, Sun T, Zheng XS (2019) Beyond winning and losing: Modeling human motivations and behaviors with vector-valued inverse reinforcement learning. Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment 15(1):195\u2013201. https:\/\/doi.org\/10.1609\/aiide.v15i1.5244","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment"},{"key":"11100_CR133","doi-asserted-by":"publisher","unstructured":"G. Wu, Y. Li, S. Luo, G. Song, Q. Wang, J. He, J. Ye, X. Qie and H. Zhu, (2020), \"A joint inverse reinforcement learning and deep learning model for drivers' behavioral prediction,\" in CIK \u201920: Proceedings of the 29th ACM International Conference on Information & Knowledge Management, 2805\u20132812. https:\/\/doi.org\/10.1145\/3340531.3412682","DOI":"10.1145\/3340531.3412682"},{"issue":"4","key":"11100_CR134","doi-asserted-by":"publisher","first-page":"1029","DOI":"10.1109\/TCSS.2021.3070239","volume":"9","author":"C Wu","year":"2022","unstructured":"Wu C (2022) Connections between relational event model and inverse reinforcement learning for characterizing group interaction sequences. IEEE Transactions on Computational Social Systems 9(4):1029\u20131037. https:\/\/doi.org\/10.1109\/TCSS.2021.3070239","journal-title":"IEEE Transactions on Computational Social Systems"},{"issue":"6","key":"11100_CR135","doi-asserted-by":"publisher","first-page":"1015","DOI":"10.1080\/0952813X.2020.1718773","volume":"32","author":"R Bhattacharyya","year":"2020","unstructured":"Bhattacharyya R, Hazarika SM (2020) A knowledge-driven layered inverse reinforcement learning approach for recognizing human intents. J Exp Theor Artif Intell 32(6):1015\u20131044. https:\/\/doi.org\/10.1080\/0952813X.2020.1718773","journal-title":"J Exp Theor Artif Intell"},{"key":"11100_CR136","doi-asserted-by":"publisher","unstructured":"S. Reddy, A.D. Dragan and S. Levine, (2018), \"Where do you think you're going?: Inferring beliefs about dynamics from behavior,\" in NIPS\u201918: Proceedings of the 32nd International Conference on Neural Information Processing Systems, 1461\u20131472. https:\/\/doi.org\/10.5555\/3326943.3327077","DOI":"10.5555\/3326943.3327077"},{"issue":"1","key":"11100_CR137","doi-asserted-by":"publisher","first-page":"2022340","DOI":"10.1080\/08839514.2021.2022340","volume":"36","author":"K Hantous","year":"2022","unstructured":"Hantous K, Rejeb L, Hellali R (2022) Detecting physiological needs using deep inverse reinforcement learning. Appl Artif Intell 36(1):2022340. https:\/\/doi.org\/10.1080\/08839514.2021.2022340","journal-title":"Appl Artif Intell"},{"issue":"10","key":"11100_CR138","doi-asserted-by":"publisher","first-page":"1364","DOI":"10.3390\/e24101364","volume":"24","author":"KJ Kim","year":"2022","unstructured":"Kim KJ, Santos E Jr, Nguyen H, Pieper S (2022) An application of inverse reinforcement learning to estimate interference in drone swarms. Entropy 24(10):1364. https:\/\/doi.org\/10.3390\/e24101364","journal-title":"Entropy"},{"key":"11100_CR139","doi-asserted-by":"publisher","unstructured":"B.D. Ziebart, A.L. Maas, A.K. Dey and J.A. Bagnell, (2008), \"Navigate like a cabbie: Probabilistic reasoning from observed context-aware behavior,\" in UbiComp \u201908: Proceedings of the 10th International Conference on Ubiquitous Computing, 322\u2013331. https:\/\/doi.org\/10.1145\/1409635.1409678","DOI":"10.1145\/1409635.1409678"},{"key":"11100_CR140","doi-asserted-by":"publisher","unstructured":"M.F. Ozkan and Y. Ma, (2020), \"Inverse reinforcement learning based driver behavior analysis and fuel economy assessment,\" in ASME 2020 Dynamic Systems and Control Conference, DSCC2020\u20133122, V001T02A003; https:\/\/doi.org\/10.1115\/DSCC2020-3122","DOI":"10.1115\/DSCC2020-3122"},{"key":"11100_CR141","doi-asserted-by":"publisher","unstructured":"M. Fahad, Z. Chen and Y. Guo, (2018), \"Learning how pedestrians navigate: A deep inverse reinforcement learning approach,\" in 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 819\u2013826. https:\/\/doi.org\/10.1109\/IROS.2018.8593438","DOI":"10.1109\/IROS.2018.8593438"},{"key":"11100_CR142","doi-asserted-by":"publisher","unstructured":"K. Saleh, M. Hossny and S. Nahavandi, (2018), \"Long-term recurrent predictive model for intent prediction of pedestrians via inverse reinforcement learning,\" in 2018 Digital Image Computing: Techniques and Applications (DICTA), Canberra, ACT, Australia, 1\u20138. https:\/\/doi.org\/10.1109\/DICTA.2018.8615854","DOI":"10.1109\/DICTA.2018.8615854"},{"issue":"18","key":"11100_CR143","doi-asserted-by":"publisher","first-page":"5207","DOI":"10.3390\/s20185207","volume":"20","author":"B Lin","year":"2020","unstructured":"Lin B, Cook DJ (2020) Analyzing sensor-based individual and population behavior patterns via inverse reinforcement learning. Sensors 20(18):5207. https:\/\/doi.org\/10.3390\/s20185207","journal-title":"Sensors"},{"key":"11100_CR144","doi-asserted-by":"publisher","unstructured":"T. Fernando, S. Denman, S. Sridharan and C. Fookes, (2019), \"Neighbourhood context embeddings in deep inverse reinforcement learning for predicting pedestrian motion over long time horizons,\" in 2019 IEEE\/CVF International Conference on Computer Vision Workshop (ICCVW), Seoul, Korea (South), 1179\u20131187. https:\/\/doi.org\/10.1109\/ICCVW.2019.00149","DOI":"10.1109\/ICCVW.2019.00149"},{"issue":"9","key":"11100_CR145","doi-asserted-by":"publisher","first-page":"1479","DOI":"10.3390\/math8091479","volume":"8","author":"F Martinez-Gil","year":"2020","unstructured":"Martinez-Gil F, Lozano M, I. P. Romero, D. Serra and R. Sebasti\u00e1n, (2020) Using inverse reinforcement learning with real trajectories to get more trustworthy pedestrian simulations. Mathematics 8(9):1479. https:\/\/doi.org\/10.3390\/math8091479","journal-title":"Mathematics"},{"key":"11100_CR146","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.122499","volume":"240","author":"B Yang","year":"2024","unstructured":"Yang B, Lu Y, Wan R, Hu H, Yang C, Ni R (2024) Meta-IRLSOT++: A meta-inverse reinforcement learning method for fast adaptation of trajectory prediction networks. Expert Syst Appl 240:122499. https:\/\/doi.org\/10.1016\/j.eswa.2023.122499","journal-title":"Expert Syst Appl"},{"key":"11100_CR147","doi-asserted-by":"publisher","unstructured":"D. Vasquez, B. Okal and K.O. Arras, (2014), \"Inverse reinforcement learning algorithms and features for robot navigation in crowds: An experimental comparison,\" in 2014 IEEE\/RSJ International Conference on Intelligent Robots and Systems, Chicago, IL, USA, 1341\u20131346. https:\/\/doi.org\/10.1109\/IROS.2014.6942731","DOI":"10.1109\/IROS.2014.6942731"},{"key":"11100_CR148","doi-asserted-by":"publisher","unstructured":"M. Herman, V. Fischer, T. Gindele and W. Burgard, (2015), \"Inverse reinforcement learning of behavioral models for online-adapting navigation strategies,\" in 2015 IEEE International Conference on Robotics and Automation (ICRA), Seattle, WA, USA, 3215\u20133222. https:\/\/doi.org\/10.1109\/ICRA.2015.7139642","DOI":"10.1109\/ICRA.2015.7139642"},{"issue":"2","key":"11100_CR149","doi-asserted-by":"publisher","first-page":"651","DOI":"10.1109\/LRA.2020.3048657","volume":"6","author":"A Konar","year":"2021","unstructured":"Konar A, Baghi BH, Dudek G (2021) Learning goal conditioned socially compliant navigation from demonstration using risk-based features. IEEE Robotics and Automation Letters 6(2):651\u2013658. https:\/\/doi.org\/10.1109\/LRA.2020.3048657","journal-title":"IEEE Robotics and Automation Letters"},{"key":"11100_CR150","doi-asserted-by":"publisher","unstructured":"M. Kollmitz, T. Koller, J. Boedecker and W. Burgard, (2020), \"Learning human-aware robot navigation from physical interaction via inverse reinforcement learning,\" in 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 11025\u201311031. https:\/\/doi.org\/10.1109\/IROS45743.2020.9340865","DOI":"10.1109\/IROS45743.2020.9340865"},{"key":"11100_CR151","doi-asserted-by":"crossref","unstructured":"S. Chandramohan, M. Geist, F. Lefevre and O. Pietquin, (2011), \"User simulation in dialogue systems using inverse reinforcement learning,\" in Interspeech 2011, Florence, Italy, 1025\u20131028. https:\/\/centralesupelec.hal.science\/hal-00652446v1","DOI":"10.21437\/Interspeech.2011-302"},{"key":"11100_CR152","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1007\/s10994-009-5110-1","volume":"77","author":"G Neu","year":"2009","unstructured":"Neu G, Szepesv\u00e1ri C (2009) Training parsers by inverse reinforcement learning. Mach Learn 77:303\u2013337. https:\/\/doi.org\/10.1007\/s10994-009-5110-1","journal-title":"Mach Learn"},{"key":"11100_CR153","doi-asserted-by":"publisher","unstructured":"Z. Shi, X. Chen, X. Qiu and X. Huang, (2018), \"Toward diverse text generation with inverse reinforcement learning,\" in IJCAI'18: Proceedings of the 27th International Joint Conference on Artificial Intelligence, 4361\u20134367. https:\/\/doi.org\/10.5555\/3304222.3304376","DOI":"10.5555\/3304222.3304376"},{"key":"11100_CR154","unstructured":"W. Hoiles, V. Krishnamurthy and K. Pattanayak, (2020), \"Rationally inattentive inverse reinforcement learning explains YouTube commenting behavior,\" Journal of Machine Learning Research, 21(1), 170. https:\/\/jmlr.org\/papers\/volume21\/19-872\/19-872.pdf"},{"issue":"1","key":"11100_CR155","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1609\/icwsm.v14i1.7311","volume":"14","author":"L Luceri","year":"2020","unstructured":"Luceri L, Giordano S, Ferrara E (2020) Detecting troll behavior via inverse reinforcement learning: A case study of Russian trolls in the 2016 US election. Proceedings of the International AAAI Conference on Web and Social Media 14(1):417\u2013427. https:\/\/doi.org\/10.1609\/icwsm.v14i1.7311","journal-title":"Proceedings of the International AAAI Conference on Web and Social Media"},{"issue":"04","key":"11100_CR156","doi-asserted-by":"publisher","first-page":"1250","DOI":"10.1109\/TCBB.2018.2830357","volume":"16","author":"M Imani","year":"2019","unstructured":"Imani M, Braga-Neto UM (2019) Control of gene regulatory networks using Bayesian inverse reinforcement learning. IEEE\/ACM Trans Comput Biol Bioinf 16(04):1250\u20131261. https:\/\/doi.org\/10.1109\/TCBB.2018.2830357","journal-title":"IEEE\/ACM Trans Comput Biol Bioinf"},{"issue":"5","key":"11100_CR157","doi-asserted-by":"publisher","first-page":"568","DOI":"10.1177\/0278364920903104","volume":"39","author":"K Li","year":"2020","unstructured":"Li K, Burdick JW (2020) Human motion analysis in medical robotics via high-dimensional inverse reinforcement learning. The International Journal of Robotics Research 39(5):568\u2013585. https:\/\/doi.org\/10.1177\/0278364920903104","journal-title":"The International Journal of Robotics Research"},{"key":"11100_CR158","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1186\/s12911-019-0763-6","volume":"19","author":"C Yu","year":"2019","unstructured":"Yu C, Liu J, Zhao H (2019) Inverse reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units. BMC Med Inform Decis Mak 19:57. https:\/\/doi.org\/10.1186\/s12911-019-0763-6","journal-title":"BMC Med Inform Decis Mak"},{"key":"11100_CR159","doi-asserted-by":"publisher","unstructured":"B. Agyemang, W.-P. Wu, D. Addo, M.Y. Kpiebaareh, E. Nanor and C.R. Haruna, (2021), \"Deep inverse reinforcement learning for structural evolution of small molecules,\" Briefings in Bioinformatics, 22(4), bbaa364. https:\/\/doi.org\/10.1093\/bib\/bbaa364","DOI":"10.1093\/bib\/bbaa364"},{"issue":"02","key":"11100_CR160","doi-asserted-by":"publisher","first-page":"304","DOI":"10.1109\/TPAMI.2018.2873794","volume":"42","author":"N Rhinehart","year":"2020","unstructured":"Rhinehart N, Kitani KM (2020) First-person activity forecasting from video with online inverse reinforcement learning. IEEE Trans Pattern Anal Mach Intell 42(02):304\u2013317. https:\/\/doi.org\/10.1109\/TPAMI.2018.2873794","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11100_CR161","doi-asserted-by":"publisher","unstructured":"X. Chen, L. Yao, A. Sun, X. Wang, X. Xu and L. Zhu, (2021), \"Generative inverse deep reinforcement learning for online recommendation,\" in CIKM '21: Proceedings of the 30th ACM International Conference on Information & Knowledge Management, 201\u2013210. https:\/\/doi.org\/10.1145\/3459637.3482347","DOI":"10.1145\/3459637.3482347"},{"key":"11100_CR162","doi-asserted-by":"publisher","first-page":"7535","DOI":"10.1109\/IROS40897.2019.8968460","volume":"2019","author":"E Tolstaya","year":"2019","unstructured":"Tolstaya E, Ribeiro A, Kumar V, Kapoor A (2019) (2019), \u201cInverse optimal planning for air traffic control,\u201d in. IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) 2019:7535\u20137542. https:\/\/doi.org\/10.1109\/IROS40897.2019.8968460","journal-title":"IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)"},{"key":"11100_CR163","doi-asserted-by":"publisher","unstructured":"Y. Luo, (2020), \"Inverse reinforcement learning for team sports: Valuing actions and players,\" in Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, 3356\u20133363. https:\/\/doi.org\/10.24963\/ijcai.2020\/464","DOI":"10.24963\/ijcai.2020\/464"},{"key":"11100_CR164","doi-asserted-by":"publisher","unstructured":"A. Tucker, A. Gleave and S. Russell, (2018), \"Inverse reinforcement learning for video games,\" arXiv:1810.10593. https:\/\/doi.org\/10.48550\/arXiv.1810.10593","DOI":"10.48550\/arXiv.1810.10593"},{"key":"11100_CR165","unstructured":"R. Pinsler, M. Maag, O. Arenz and G. Neumann, (2018), \"Inverse reinforcement learning of bird flocking behavior,\" in ICRA 2018 Workshop on Swarms: From Biology to Robotics and Back. https:\/\/www.ias.informatik.tu-darmstadt.de\/uploads\/Team\/OlegArenz\/PinslerEtAl_ICRA2018swarms.pdf"},{"key":"11100_CR166","doi-asserted-by":"publisher","unstructured":"M.-h. Oh and G. Iyengar, (2019), \"Sequential anomaly detection using inverse reinforcement learning,\" in KDD '19: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 1480\u20131490. https:\/\/doi.org\/10.1145\/3292500.3330932","DOI":"10.1145\/3292500.3330932"},{"key":"11100_CR167","doi-asserted-by":"publisher","unstructured":"B.G. Doerr, R. Linares and R. Furfaro, (2020), \"Space objects maneuvering prediction via maximum causal entropy inverse reinforcement learning,\" in AIAA 2020\u20130235, AIAA Scitech 2020 Forum. https:\/\/doi.org\/10.2514\/6.2020-0235","DOI":"10.2514\/6.2020-0235"},{"key":"11100_CR168","doi-asserted-by":"publisher","unstructured":"J. Jin, L. Petrich, Z. Zhang, M. Dehghan and M. Jagersand, (2020), \"Visual geometric skill inference by watching human demonstration,\" in 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 8985\u20138991. https:\/\/doi.org\/10.1109\/ICRA40945.2020.9196570","DOI":"10.1109\/ICRA40945.2020.9196570"},{"key":"11100_CR169","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijtst.2024.11.007","author":"S Wang","year":"2024","unstructured":"Wang S, Yang H, Tang Y, Chen J, Zhao C, Du Y (2024) Learning to search for parking like a human: A deep inverse reinforcement learning approach. International Journal of Transportation Science and Technology. https:\/\/doi.org\/10.1016\/j.ijtst.2024.11.007","journal-title":"International Journal of Transportation Science and Technology"},{"key":"11100_CR170","doi-asserted-by":"publisher","unstructured":"G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang and W. Zaremba, (2016), \"OpenAI Gym,\" arXiv:1606.01540. https:\/\/doi.org\/10.48550\/arXiv.1606.01540","DOI":"10.48550\/arXiv.1606.01540"},{"key":"11100_CR171","unstructured":"E. Leurent, (2018), \u201cAn environment for autonomous driving decision-making,\u201d GitHub. https:\/\/github.com\/czh513\/Auto-driving-RL-decision-making"},{"key":"11100_CR172","doi-asserted-by":"publisher","unstructured":"P. Newson and J. Krumm, (2009), \"Hidden Markov map matching through noise and sparseness,\" in GIS '09: Proceedings of the 17th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems, 336\u2013343. https:\/\/doi.org\/10.1145\/1653771.1653818","DOI":"10.1145\/1653771.1653818"},{"key":"11100_CR173","doi-asserted-by":"publisher","unstructured":"M. Piorkowski, N. Sarafijanovic-Djukic and M. Grossglauser, (2009), \"A parsimonious model of mobile partitioned networks with clustering,\" in 2009 First International Communication Systems and Networks and Workshops, Bangalore, India, 1\u201310. https:\/\/doi.org\/10.1109\/COMSNETS.2009.4808865","DOI":"10.1109\/COMSNETS.2009.4808865"},{"key":"11100_CR174","doi-asserted-by":"publisher","unstructured":"J. Choi and K.-E. Kim, (2012), \"Nonparametric Bayesian inverse reinforcement learning for multiple reward functions,\" in NIPS'12: Proceedings of the 25th International Conference on Neural Information Processing Systems, 1, 305\u2013313. https:\/\/doi.org\/10.5555\/2999134.2999169","DOI":"10.5555\/2999134.2999169"},{"key":"11100_CR175","doi-asserted-by":"publisher","unstructured":"W. Zhan, L. Sun, D. Wang, Y. Jin and M. Tomizuka, (2019), \"Constructing a highly interactive vehicle motion dataset,\" in 2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 6415\u20136420. https:\/\/doi.org\/10.1109\/IROS40897.2019.8967724","DOI":"10.1109\/IROS40897.2019.8967724"},{"key":"11100_CR176","doi-asserted-by":"publisher","unstructured":"W. Zhan, L. Sun, D. Wang, H. Shi, A. Clausse, M. Naumann, J. Kummerle, H. Konigshof, C. Stiller, A. de La Fortelle and M. Tomizuka, (2019), \u201cINTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION dataset in interactive driving scenarios with semantic maps,\u201d arXiv:1910.03088. https:\/\/doi.org\/10.48550\/arXiv.1910.03088","DOI":"10.48550\/arXiv.1910.03088"},{"key":"11100_CR177","doi-asserted-by":"publisher","unstructured":"P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. S\u00fcnderhauf, I. Reid, S. Gould and A. Van Den Hengel, (2018), \"Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,\" in 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 3674\u20133683. https:\/\/doi.org\/10.1109\/CVPR.2018.00387","DOI":"10.1109\/CVPR.2018.00387"},{"key":"11100_CR178","doi-asserted-by":"publisher","DOI":"10.1038\/sdata.2016.35","volume":"3","author":"AEW Johnson","year":"2016","unstructured":"Johnson AEW, Pollard TJ, Shen L, Li-Wei HL, Feng M, Ghassemi M, Moody B, Szolovits P, Celi LA, Mark RG (2016) MIMIC-III, a freely accessible critical care database. Scientific Data 3:160035. https:\/\/doi.org\/10.1038\/sdata.2016.35","journal-title":"Scientific Data"},{"key":"11100_CR179","unstructured":"L. El Asri, R. Laroche and O. Pietquin, (2014), \"DINASTI: Dialogues with a negotiating appointment setting interface,\" in Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14), 272\u2013278, Reykjavik, Iceland. http:\/\/www.lrec-conf.org\/proceedings\/lrec2014\/pdf\/576_Paper.pdf"},{"key":"11100_CR180","doi-asserted-by":"publisher","unstructured":"Y.-T. Lee, K.-T. Chen, Y.-M. Cheng and C.-L. Lei, (2011), \"World of Warcraft avatar history dataset,\" in MMSys '11: Proceedings of the Second Annual ACM Conference on Multimedia Systems, 123\u2013128. https:\/\/doi.org\/10.1145\/1943552.1943569","DOI":"10.1145\/1943552.1943569"},{"key":"11100_CR181","doi-asserted-by":"publisher","unstructured":"H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan and O. Beijbom, (2020), \"nuScenes: A multimodal dataset for autonomous driving,\" in 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 11618\u201311628. https:\/\/doi.org\/10.1109\/CVPR42600.2020.01164","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"11100_CR182","doi-asserted-by":"publisher","unstructured":"L. Sun, Z. Wu, H. Ma and M. Tomizuka, (2020), \"Expressing diverse human driving behavior with probabilistic rewards and online inference,\" in 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 2020\u20132026. https:\/\/doi.org\/10.1109\/IROS45743.2020.9341371","DOI":"10.1109\/IROS45743.2020.9341371"},{"issue":"11","key":"11100_CR183","doi-asserted-by":"publisher","first-page":"1231","DOI":"10.1177\/0278364913491297","volume":"32","author":"A Geiger","year":"2013","unstructured":"Geiger A, Lenz P, Stiller C, Urtasun R (2013) Vision meets robotics: The KITTI dataset. The International Journal of Robotics Research 32(11):1231\u20131237. https:\/\/doi.org\/10.1177\/0278364913491297","journal-title":"The International Journal of Robotics Research"},{"key":"11100_CR184","doi-asserted-by":"publisher","unstructured":"A. Robicquet, A. Sadeghian, A. Alahi and S. Savarese, (2016), \"Learning social etiquette: Human trajectory understanding in crowded scenes,\" in Leibe, B., Matas, J., Sebe, N., Welling, M. (eds) Computer Vision \u2013 ECCV 2016. ECCV 2016. Lecture Notes in Computer Science, 9912, 549\u2013565. Springer, Cham. https:\/\/doi.org\/10.1007\/978-3-319-46484-8_33","DOI":"10.1007\/978-3-319-46484-8_33"},{"key":"11100_CR185","doi-asserted-by":"publisher","unstructured":"T. Warnakulasuriya, S. Denman, S. Sridharan and C. Fookes, (2019), \"Pedestrian trajectory prediction with structured memory hierarchies,\" in Berlingerio, M., Bonchi, F., G\u00e4rtner, T., Hurley, N., Ifrim, G. (eds) Machine Learning and Knowledge Discovery in Databases. ECML PKDD 2018. Lecture Notes in Computer Science, 11051, 241\u2013256. Springer, Cham. https:\/\/doi.org\/10.1007\/978-3-030-10925-7_15","DOI":"10.1007\/978-3-030-10925-7_15"},{"key":"11100_CR186","unstructured":"Y. Zheng, X. Xie and W.-Y. Ma, (2010), \"Geolife: A collaborative social networking service among user, location and trajectory.,\" IEEE Data(Base) Engineering Bulletin, 33, 32\u201339. https:\/\/www.microsoft.com\/en-us\/research\/publication\/geolife-a-collaborative-social-networking-service-among-user-location-and-trajectory\/"},{"key":"11100_CR187","doi-asserted-by":"publisher","unstructured":"Y. Zheng, L. Zhang, X. Xie and W.-Y. Ma, (2009), \"Mining interesting locations and travel sequences from GPS trajectories,\" in WWW '09: Proceedings of the 18th International Conference on World Wide Web, 791\u2013800. https:\/\/doi.org\/10.1145\/1526709.1526816","DOI":"10.1145\/1526709.1526816"},{"issue":"3","key":"11100_CR188","doi-asserted-by":"publisher","first-page":"1393","DOI":"10.1109\/TITS.2013.2262376","volume":"14","author":"L Moreira-Matias","year":"2013","unstructured":"Moreira-Matias L, Gama J, Ferreira M, Mendes-Moreira J, Damas L (2013) Predicting taxi passenger demand using streaming data. IEEE Trans Intell Transp Syst 14(3):1393\u20131402. https:\/\/doi.org\/10.1109\/TITS.2013.2262376","journal-title":"IEEE Trans Intell Transp Syst"},{"key":"11100_CR189","doi-asserted-by":"publisher","unstructured":"X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Doll\u00e1r and C.L. Zitnick, (2015), \"Microsoft COCO captions: Data collection and evaluation server,\" arXiv:1504.00325. https:\/\/doi.org\/10.48550\/arXiv.1504.00325","DOI":"10.48550\/arXiv.1504.00325"},{"key":"11100_CR190","doi-asserted-by":"publisher","unstructured":"Q. Diao, M. Qiu, C.-Y. Wu, A.J. Smola, J. Jiang and C. Wang, (2014), \"Jointly modeling aspects, ratings and sentiments for movie recommendation (JMARS),\" in KDD '14: Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 193- 202. https:\/\/doi.org\/10.1145\/2623330.2623758","DOI":"10.1145\/2623330.2623758"},{"key":"11100_CR191","unstructured":"V. Alexiadis, J. Colyar, J. Halkias, R. Hranac and G. McHale, (2004), \"The next generation simulation program,\" Institute of Transportation Engineers. ITE Journal, 74(8), 22\u201326. https:\/\/www.proquest.com\/docview\/224887913\/fulltextPDF\/E65A0D3D1DE4433DPQ\/1?accountid=12528&sourcetype=Scholarly%20Journals"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-025-11100-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-025-11100-0\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-025-11100-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,31]],"date-time":"2025-05-31T09:17:49Z","timestamp":1748683069000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-025-11100-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,26]]},"references-count":191,"journal-issue":{"issue":"17","published-print":{"date-parts":[[2025,6]]}},"alternative-id":["11100"],"URL":"https:\/\/doi.org\/10.1007\/s00521-025-11100-0","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,26]]},"assertion":[{"value":"17 July 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 February 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 March 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial or non-financial interests to disclose.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Authors\u2019 declaration: This manuscript is the authors\u2019 original work, and the content of this paper has not been copied from elsewhere. This manuscript is unpublished and is not currently under review with nor submitted to another journal; all data measurements are genuine results and have not been manipulated. All authors have checked the manuscript and have agreed to this submission. Ethical approval: This article does not contain any studies with human participants or animals performed by any of the authors.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"During the preparation of this work the authors used ChatGPT to improve grammar and sentence structure. After using this tool\/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declaration of generative AI in scientific writing"}}]}}