{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,9]],"date-time":"2026-05-09T16:51:33Z","timestamp":1778345493576,"version":"3.51.4"},"reference-count":71,"publisher":"SAGE Publications","issue":"13-14","license":[{"start":{"date-parts":[[2019,5,22]],"date-time":"2019-05-22T00:00:00Z","timestamp":1558483200000},"content-version":"vor","delay-in-days":365,"URL":"http:\/\/www.sagepub.com\/licence-information-for-chorus"}],"funder":[{"DOI":"10.13039\/100007297","name":"Office of Naval Research Global","doi-asserted-by":"publisher","award":["N00014-15-1-2673"],"award-info":[{"award-number":["N00014-15-1-2673"]}],"id":[{"id":"10.13039\/100007297","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of Robotics Research"],"published-print":{"date-parts":[[2018,12]]},"abstract":"<jats:p>The literature on inverse reinforcement learning (IRL) typically assumes that humans take actions to minimize the expected value of a cost function, i.e., that humans are risk neutral. Yet, in practice, humans are often far from being risk neutral. To fill this gap, the objective of this paper is to devise a framework for risk-sensitive (RS) IRL to explicitly account for a human\u2019s risk sensitivity. To this end, we propose a flexible class of models based on coherent risk measures, which allow us to capture an entire spectrum of risk preferences from risk neutral to worst case. We propose efficient non-parametric algorithms based on linear programming and semi-parametric algorithms based on maximum likelihood for inferring a human\u2019s underlying risk measure and cost function for a rich class of static and dynamic decision-making settings. The resulting approach is demonstrated on a simulated driving game with 10 human participants. Our method is able to infer and mimic a wide range of qualitatively different driving styles from highly risk averse to risk neutral in a data-efficient manner. Moreover, comparisons of the RS-IRL approach with a risk-neutral model show that the RS-IRL framework more accurately captures observed participant behavior both qualitatively and quantitatively, especially in scenarios where catastrophic outcomes such as collisions can occur.<\/jats:p>","DOI":"10.1177\/0278364918772017","type":"journal-article","created":{"date-parts":[[2018,5,23]],"date-time":"2018-05-23T02:33:52Z","timestamp":1527042832000},"page":"1713-1740","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":14,"title":["Risk-sensitive inverse reinforcement learning via semi- and non-parametric methods"],"prefix":"10.1177","volume":"37","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4510-5853","authenticated-orcid":false,"given":"Sumeet","family":"Singh","sequence":"first","affiliation":[{"name":"Department of Aeronautics and Astronautics, Stanford University, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jonathan","family":"Lacotte","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, Stanford University, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anirudha","family":"Majumdar","sequence":"additional","affiliation":[{"name":"Department of Mechanical and Aerospace Engineering, Princeton University, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marco","family":"Pavone","sequence":"additional","affiliation":[{"name":"Department of Aeronautics and Astronautics, Stanford University, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2018,5,22]]},"reference":[{"key":"bibr1-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1145\/1015330.1015430"},{"key":"bibr2-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1145\/1102351.1102352"},{"key":"bibr3-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1016\/S0378-4266(02)00281-9"},{"key":"bibr4-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1016\/S0378-4266(02)00283-2"},{"key":"bibr5-0278364918772017","doi-asserted-by":"publisher","DOI":"10.2307\/1907921"},{"key":"bibr6-0278364918772017","unstructured":"ApS M (2017) MOSEK optimization software. Available at: https:\/\/mosek.com\/."},{"key":"bibr7-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1111\/1467-9965.00068"},{"key":"bibr8-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1109\/CDC.2016.7799166"},{"key":"bibr9-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1257\/jep.27.1.173"},{"key":"bibr10-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1007\/s00186-011-0367-0"},{"key":"bibr11-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1137\/030601296"},{"key":"bibr12-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0167021"},{"key":"bibr13-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-012-0569-0"},{"key":"bibr14-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1109\/ACC.2014.6859437"},{"key":"bibr15-0278364918772017","author":"Chow Y","year":"2015","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr16-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1137\/040605217"},{"key":"bibr17-0278364918772017","doi-asserted-by":"publisher","DOI":"10.2307\/1884324"},{"key":"bibr18-0278364918772017","volume-title":"International Symposium on Robotics Research","author":"Englert P","year":"2015"},{"key":"bibr19-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1287\/moor.14.1.147"},{"key":"bibr20-0278364918772017","volume-title":"International Conference on Machine Learning","author":"Finn C","year":"2016"},{"key":"bibr21-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1666"},{"key":"bibr22-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-20451-2_21"},{"key":"bibr23-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1016\/0304-4068(89)90018-9"},{"key":"bibr24-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-416008-8.00035-8"},{"key":"bibr25-0278364918772017","doi-asserted-by":"publisher","DOI":"10.3982\/ECTA9188"},{"key":"bibr26-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1007\/s11166-010-9102-0"},{"key":"bibr27-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.18.7.356"},{"key":"bibr28-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1126\/science.1115327"},{"key":"bibr29-0278364918772017","doi-asserted-by":"publisher","DOI":"10.2307\/1914185"},{"key":"bibr30-0278364918772017","author":"Kolter JZ","year":"2007","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr31-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1177\/0278364915619772"},{"key":"bibr32-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2015.7139555"},{"key":"bibr33-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1007\/s00211-016-0842-x"},{"key":"bibr34-0278364918772017","volume-title":"International Conference on Machine Learning","author":"Levine S","year":"2012"},{"key":"bibr35-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1109\/CACSD.2004.1393890"},{"key":"bibr36-0278364918772017","volume-title":"International Symposium on Robotics Research","author":"Majumdar A","year":"2017"},{"key":"bibr37-0278364918772017","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2017.XIII.069"},{"key":"bibr38-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1023\/A:1017940631555"},{"key":"bibr39-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1007\/s10514-009-9170-7"},{"key":"bibr40-0278364918772017","volume-title":"International Conference on Machine Learning","author":"Ng A","year":"2000"},{"key":"bibr41-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1287\/opre.1050.0216"},{"key":"bibr42-0278364918772017","author":"Osogami T","year":"2012","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr43-0278364918772017","volume-title":"Robotics: Science and Systems Workshop on Inverse Optimal Control and Robotic Learning from Demonstration","author":"Park T","year":"2013"},{"key":"bibr44-0278364918772017","volume-title":"Proceedings of the Conference on Uncertainty in Artificial Intelligence","author":"Petrik M","year":"2012"},{"key":"bibr45-0278364918772017","volume-title":"International Conference on Machine Learning","author":"Prashanth LA","year":"2016"},{"key":"bibr46-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1016\/0167-2681(82)90008-7"},{"key":"bibr47-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1111\/1468-0262.00158"},{"key":"bibr48-0278364918772017","volume-title":"International Joint Conference on Artificial Intelligence","author":"Ramachandran D","year":"2007"},{"key":"bibr49-0278364918772017","unstructured":"Ratliff LJ, Mazumdar E (2017) Risk-sensitive inverse reinforcement learning via gradient methods. Available at: https:\/\/arxiv.org\/abs\/1703.09842."},{"key":"bibr50-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1287\/educ.1073.0032"},{"key":"bibr51-0278364918772017","doi-asserted-by":"publisher","DOI":"10.21314\/JOR.2000.038"},{"key":"bibr52-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1016\/S0378-4266(02)00271-6"},{"key":"bibr53-0278364918772017","author":"Russell S","year":"1998","journal-title":"Proceedings Computational Learning Theory"},{"key":"bibr54-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-010-0393-3"},{"key":"bibr55-0278364918772017","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2016.XII.029"},{"key":"bibr56-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2016.7759036"},{"key":"bibr57-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1016\/j.orl.2009.02.005"},{"key":"bibr58-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611973433"},{"key":"bibr59-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1162\/NECO_a_00600"},{"key":"bibr60-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.2016.2644871"},{"key":"bibr61-0278364918772017","volume-title":"International Conference on Machine Learning","author":"Tamar A","year":"2012"},{"key":"bibr62-0278364918772017","unstructured":"VIRES Simulationstechnologie GmbH (2017) VTD - Virtual Test Drive. Available at: https:\/\/vires.com\/vtd-vires-virtual-test-drive\/."},{"key":"bibr63-0278364918772017","volume-title":"Theory of Games and Economic Behavior","author":"von Neumann J","year":"1944"},{"key":"bibr64-0278364918772017","volume-title":"International Conference on Machine Learning","author":"Waugh K","year":"2011"},{"key":"bibr65-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1006\/jmaa.1998.6203"},{"key":"bibr66-0278364918772017","unstructured":"Wulfmeier M, Ondruska P, Posner I (2015) Maximum entropy deep inverse reinforcement learning. Available at: https:\/\/arxiv.org\/abs\/1507.04888."},{"key":"bibr67-0278364918772017","author":"Xu H","year":"2010","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr68-0278364918772017","doi-asserted-by":"publisher","DOI":"10.2307\/1911158"},{"key":"bibr69-0278364918772017","volume-title":"Proceedings AAAI Conference on Artificial Intelligence","author":"Ziebart BD","year":"2008"},{"key":"bibr70-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2009.5354147"},{"key":"bibr71-0278364918772017","doi-asserted-by":"publisher","DOI":"10.1109\/ROBOT.2010.5509176"}],"container-title":["The International Journal of Robotics Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/0278364918772017","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/0278364918772017","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/0278364918772017","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/0278364918772017","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T10:15:43Z","timestamp":1777457743000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/0278364918772017"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,5,22]]},"references-count":71,"journal-issue":{"issue":"13-14","published-print":{"date-parts":[[2018,12]]}},"alternative-id":["10.1177\/0278364918772017"],"URL":"https:\/\/doi.org\/10.1177\/0278364918772017","relation":{},"ISSN":["0278-3649","1741-3176"],"issn-type":[{"value":"0278-3649","type":"print"},{"value":"1741-3176","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,5,22]]}}}