{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,20]],"date-time":"2026-01-20T11:19:03Z","timestamp":1768907943222,"version":"3.49.0"},"publisher-location":"New York, NY, USA","reference-count":30,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,3,8]],"date-time":"2021-03-08T00:00:00Z","timestamp":1615161600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"CONIX Research Center"},{"name":"DARPA Assured Autonomy","award":["FA8750-18-C-0101"],"award-info":[{"award-number":["FA8750-18-C-0101"]}]},{"DOI":"10.13039\/100000181","name":"Air Force Office of Scientific Research","doi-asserted-by":"publisher","award":["FA9550-17-1-0308"],"award-info":[{"award-number":["FA9550-17-1-0308"]}],"id":[{"id":"10.13039\/100000181","id-type":"DOI","asserted-by":"publisher"}]},{"name":"German Academic Exchange Service (DAAD)"},{"DOI":"10.13039\/100000006","name":"Office of Naval Research","doi-asserted-by":"publisher","award":["20-1-2736"],"award-info":[{"award-number":["20-1-2736"]}],"id":[{"id":"10.13039\/100000006","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,3,8]]},"DOI":"10.1145\/3434073.3444667","type":"proceedings-article","created":{"date-parts":[[2021,3,5]],"date-time":"2021-03-05T18:02:57Z","timestamp":1614967377000},"page":"216-224","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["Feature Expansive Reward Learning"],"prefix":"10.1145","author":[{"given":"Andreea","family":"Bobu","sequence":"first","affiliation":[{"name":"University of California, Berkeley, Berkeley, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marius","family":"Wiggert","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Claire","family":"Tomlin","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anca D.","family":"Dragan","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,3,8]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Invariant Risk Minimization. ArXi","author":"Arjovsky Mart\u00edn","year":"2019","unstructured":"Mart\u00edn Arjovsky , L\u00e9on Bottou , Ishaan Gulrajani , and David Lopez-Paz . 2019. Invariant Risk Minimization. ArXi , Vol. abs\/ 1907 .02893 ( 2019 ). Mart\u00edn Arjovsky, L\u00e9on Bottou, Ishaan Gulrajani, and David Lopez-Paz. 2019. Invariant Risk Minimization. ArXi, Vol. abs\/1907.02893 (2019)."},{"key":"e_1_3_2_1_2_1","volume-title":"Marcia Kilchenman O'Malley, and Anca D. Dragan","author":"Bajcsy Andrea","year":"2017","unstructured":"Andrea Bajcsy , Dylan P. Losey , Marcia Kilchenman O'Malley, and Anca D. Dragan . 2017 . Learning Robot Objectives from Physical Human Interaction. In CoRL. Andrea Bajcsy, Dylan P. Losey, Marcia Kilchenman O'Malley, and Anca D. Dragan. 2017. Learning Robot Objectives from Physical Human Interaction. In CoRL."},{"key":"e_1_3_2_1_3_1","volume-title":"Proceedings of the 2018 ACM\/IEEE International Conference on Human-Robot Interaction","author":"Bajcsy Andrea","unstructured":"Andrea Bajcsy , Dylan P. Losey , Marcia K. O'Malley , and Anca D. Dragan . 2018. Learning from Physical Human Corrections, One Feature at a Time . In Proceedings of the 2018 ACM\/IEEE International Conference on Human-Robot Interaction ( Chicago, IL, USA) (HRI '18). ACM, New York, NY, USA, 141--149. https:\/\/doi.org\/10.1145\/3171221.3171267 Andrea Bajcsy, Dylan P. Losey, Marcia K. O'Malley, and Anca D. Dragan. 2018. Learning from Physical Human Corrections, One Feature at a Time. In Proceedings of the 2018 ACM\/IEEE International Conference on Human-Robot Interaction (Chicago, IL, USA) (HRI '18). ACM, New York, NY, USA, 141--149. https:\/\/doi.org\/10.1145\/3171221.3171267"},{"key":"e_1_3_2_1_4_1","volume-title":"Goal inference as inverse planning. (01","author":"Baker Chris","year":"2007","unstructured":"Chris Baker , Joshua B Tenenbaum , and Rebecca R Saxe . 2007. Goal inference as inverse planning. (01 2007 ). Chris Baker, Joshua B Tenenbaum, and Rebecca R Saxe. 2007. Goal inference as inverse planning. (01 2007)."},{"key":"e_1_3_2_1_5_1","unstructured":"A. Bobu A. Bajcsy J. F. Fisac S. Deglurkar and A. D. Dragan. 2020. Quantifying Hypothesis Space Misspecification in Learning From Human--Robot Demonstrations and Physical Corrections. IEEE Transactions on Robotics (2020) 1--20. A. Bobu A. Bajcsy J. F. Fisac S. Deglurkar and A. D. Dragan. 2020. Quantifying Hypothesis Space Misspecification in Learning From Human--Robot Demonstrations and Physical Corrections. IEEE Transactions on Robotics (2020) 1--20."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.2307\/2334029"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2021028"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/2540128.2540314"},{"key":"e_1_3_2_1_9_1","volume-title":"Deep reinforcement learning from human preferences. (06","author":"Christiano Paul","year":"2017","unstructured":"Paul Christiano , Jan Leike , Tom B. Brown , Miljan Martic , Shane Legg , and Dario Amodei . 2017. Deep reinforcement learning from human preferences. (06 2017 ). Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. (06 2017)."},{"key":"e_1_3_2_1_10_1","unstructured":"Erwin Coumans and Yunfei Bai. 2016--2019. PyBullet a Python module for physics simulation for games robotics and machine learning. http:\/\/pybullet.org. Erwin Coumans and Yunfei Bai. 2016--2019. PyBullet a Python module for physics simulation for games robotics and machine learning. http:\/\/pybullet.org."},{"key":"e_1_3_2_1_11_1","volume-title":"Proceedings of the 33rd International Conference on International Conference on Machine Learning -","volume":"48","author":"Finn Chelsea","year":"2016","unstructured":"Chelsea Finn , Sergey Levine , and Pieter Abbeel . 2016 . Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization . In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 (New York, NY, USA) (ICML'16). JMLR.org, 49--58. Chelsea Finn, Sergey Levine, and Pieter Abbeel. 2016. Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 (New York, NY, USA) (ICML'16). JMLR.org, 49--58."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2018.XIV.069"},{"key":"e_1_3_2_1_13_1","volume-title":"Tomlin","author":"Fridovich-Keil David","year":"2019","unstructured":"David Fridovich-Keil , Andrea Bajcsy , Jaime F. Fisac , Sylvia L. Herbert , Steven Wang , Anca D. Dragan , and Claire J . Tomlin . 2019 . Confidence-aware motion prediction for real-time collision avoidance. International Journal of Robotics Research ( 2019). David Fridovich-Keil, Andrea Bajcsy, Jaime F. Fisac, Sylvia L. Herbert, Steven Wang, Anca D. Dragan, and Claire J. Tomlin. 2019. Confidence-aware motion prediction for real-time collision avoidance. International Journal of Robotics Research (2019)."},{"key":"e_1_3_2_1_14_1","volume-title":"Learning Robust Rewards with Adverserial Inverse Reinforcement Learning. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=rkHywl-A-","author":"Fu Justin","year":"2018","unstructured":"Justin Fu , Katie Luo , and Sergey Levine . 2018 . Learning Robust Rewards with Adverserial Inverse Reinforcement Learning. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=rkHywl-A- Justin Fu, Katie Luo, and Sergey Levine. 2018. Learning Robust Rewards with Adverserial Inverse Reinforcement Learning. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=rkHywl-A-"},{"key":"e_1_3_2_1_15_1","volume-title":"Staveland","author":"Hart Sandra G.","year":"1988","unstructured":"Sandra G. Hart and Lowell E . Staveland . 1988 . Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research. In Human Mental Workload,, Peter A. Hancock and Najmedin Meshkati (Eds.). Advances in Psychology, Vol. 52 . North-Holland , 139--183. https:\/\/doi.org\/10.1016\/S0166-4115(08)62386-9 Sandra G. Hart and Lowell E. Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research. In Human Mental Workload,, Peter A. Hancock and Najmedin Meshkati (Eds.). Advances in Psychology, Vol. 52. North-Holland, 139--183. https:\/\/doi.org\/10.1016\/S0166-4115(08)62386-9"},{"key":"e_1_3_2_1_16_1","unstructured":"Luis Haug Sebastian Tschiatschek and Adish Singla. 2018. Teaching inverse reinforcement learners via features and demonstrations. In Advances in Neural Information Processing Systems. 8464--8473. Luis Haug Sebastian Tschiatschek and Adish Singla. 2018. Teaching inverse reinforcement learners via features and demonstrations. In Advances in Neural Information Processing Systems. 8464--8473."},{"key":"e_1_3_2_1_17_1","volume-title":"Garnett (Eds.)","volume":"31","author":"Ibarz Borja","year":"2018","unstructured":"Borja Ibarz , Jan Leike , Tobias Pohlen , Geoffrey Irving , Shane Legg , and Dario Amodei . 2018 . Reward learning from human preferences and demonstrations in Atari. In Advances in Neural Information Processing Systems,, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R . Garnett (Eds.) , Vol. 31 . Curran Associates, Inc., 8011--8023. https:\/\/proceedings.neurips.cc\/paper\/ 2018\/file\/8cbe9ce23f42628c98f80fa0fac8b19a-Paper.pdf Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei. 2018. Reward learning from human preferences and demonstrations in Atari. In Advances in Neural Information Processing Systems,, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.), Vol. 31. Curran Associates, Inc., 8011--8023. https:\/\/proceedings.neurips.cc\/paper\/2018\/file\/8cbe9ce23f42628c98f80fa0fac8b19a-Paper.pdf"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1177\/0278364915581193"},{"key":"e_1_3_2_1_19_1","unstructured":"Sergey Levine Zoran Popovic and Vladlen Koltun. 2010. Feature construction for inverse reinforcement learning. In Advances in Neural Information Processing Systems. 1342--1350. Sergey Levine Zoran Popovic and Vladlen Koltun. 2010. Feature construction for inverse reinforcement learning. In Advances in Neural Information Processing Systems. 1342--1350."},{"key":"e_1_3_2_1_20_1","unstructured":"Sergey Levine Zoran Popovic and Vladlen Koltun. 2011. Nonlinear inverse reinforcement learning with gaussian processes. In Advances in Neural Information Processing Systems. 19--27. Sergey Levine Zoran Popovic and Vladlen Koltun. 2011. Nonlinear inverse reinforcement learning with gaussian processes. In Advances in Neural Information Processing Systems. 19--27."},{"key":"e_1_3_2_1_21_1","first-page":"153","article-title":"Individual choice behavior. John Wiley, Oxford","author":"Luce R. Duncan","year":"1959","unstructured":"R. Duncan Luce . 1959 . Individual choice behavior. John Wiley, Oxford , England. xii , 153 -- xii , 153 pages. R. Duncan Luce. 1959. Individual choice behavior. John Wiley, Oxford, England. xii, 153--xii, 153 pages.","journal-title":"England."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Nathan Ratliff David M Bradley Joel Chestnutt and J A Bagnell. 2007. Boosting structured prediction for imitation learning. In Advances in Neural Information Processing Systems. 1153--1160. Nathan Ratliff David M Bradley Joel Chestnutt and J A Bagnell. 2007. Boosting structured prediction for imitation learning. In Advances in Neural Information Processing Systems. 1153--1160.","DOI":"10.7551\/mitpress\/7503.003.0149"},{"key":"e_1_3_2_1_23_1","volume-title":"Proceedings of the 23rd International Conference on Machine Learning","author":"Ratliff Nathan D.","unstructured":"Nathan D. Ratliff , J. Andrew Bagnell , and Martin A. Zinkevich . 2006. Maximum Margin Planning . In Proceedings of the 23rd International Conference on Machine Learning ( Pittsburgh, Pennsylvania, USA) (ICML '06). Association for Computing Machinery, New York, NY, USA, 729--736. https:\/\/doi.org\/10.1145\/1143844.1143936 Nathan D. Ratliff, J. Andrew Bagnell, and Martin A. Zinkevich. 2006. Maximum Margin Planning. In Proceedings of the 23rd International Conference on Machine Learning (Pittsburgh, Pennsylvania, USA) (ICML '06). Association for Computing Machinery, New York, NY, USA, 729--736. https:\/\/doi.org\/10.1145\/1143844.1143936"},{"key":"e_1_3_2_1_24_1","volume-title":"8th International Conference on Learning Representations, ICLR 2020","author":"Reddy Siddharth","year":"2020","unstructured":"Siddharth Reddy , Anca D. Dragan , and Sergey Levine . 2020 . SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards . In 8th International Conference on Learning Representations, ICLR 2020 , Addis Ababa, Ethiopia , April 26-30, 2020. OpenReview.net. https:\/\/openreview.net\/forum?id=S1xKd24twB Siddharth Reddy, Anca D. Dragan, and Sergey Levine. 2020. SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net. https:\/\/openreview.net\/forum?id=S1xKd24twB"},{"key":"e_1_3_2_1_25_1","volume-title":"Learning Human Objectives by Evaluating Hypothetical Behavior. arxiv","author":"Reddy Siddharth","year":"1912","unstructured":"Siddharth Reddy , Anca D. Dragan , Sergey Levine , Shane Legg , and Jan Leike . 2019. Learning Human Objectives by Evaluating Hypothetical Behavior. arxiv : 1912 .05652 [cs.CY] Siddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg, and Jan Leike. 2019. Learning Human Objectives by Evaluating Hypothetical Behavior. arxiv: 1912.05652 [cs.CY]"},{"key":"e_1_3_2_1_26_1","unstructured":"John Schulman Jonathan Ho Alex Lee Ibrahim Awwal Henry Bradlow and Pieter Abbeel. [n.d.]. Finding Locally Optimal Collision-Free Trajectories with Sequential Convex Optimization. John Schulman Jonathan Ho Alex Lee Ibrahim Awwal Henry Bradlow and Pieter Abbeel. [n.d.]. Finding Locally Optimal Collision-Free Trajectories with Sequential Convex Optimization."},{"key":"e_1_3_2_1_27_1","volume-title":"Collision-Free Trajectories with Sequential Convex Optimization.. In Robotics: science and systems","author":"Schulman John","unstructured":"John Schulman , Jonathan Ho , Alex X Lee , Ibrahim Awwal , Henry Bradlow , and Pieter Abbeel . 2013. Finding Locally Optimal , Collision-Free Trajectories with Sequential Convex Optimization.. In Robotics: science and systems , Vol. 9 . Citeseer , 1--10. John Schulman, Jonathan Ho, Alex X Lee, Ibrahim Awwal, Henry Bradlow, and Pieter Abbeel. 2013. Finding Locally Optimal, Collision-Free Trajectories with Sequential Convex Optimization.. In Robotics: science and systems, Vol. 9. Citeseer, 1--10."},{"key":"e_1_3_2_1_28_1","volume-title":"Theory of games and economic behavior","author":"Neumann John Von","unstructured":"John Von Neumann and Oskar Morgenstern . 1945. Theory of games and economic behavior . Princeton University Press Princeton , NJ. John Von Neumann and Oskar Morgenstern. 1945. Theory of games and economic behavior. Princeton University Press Princeton, NJ."},{"key":"e_1_3_2_1_29_1","volume-title":"2016 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS). 2089--2095","author":"Wulfmeier M.","unstructured":"M. Wulfmeier , D. Z. Wang , and I. Posner . 2016. Watch this: Scalable cost-function learning for path planning in urban environments . In 2016 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS). 2089--2095 . M. Wulfmeier, D. Z. Wang, and I. Posner. 2016. Watch this: Scalable cost-function learning for path planning in urban environments. In 2016 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS). 2089--2095."},{"key":"e_1_3_2_1_30_1","volume-title":"Proceedings of the 23rd National Conference on Artificial Intelligence -","volume":"3","author":"Ziebart Brian D.","year":"2027","unstructured":"Brian D. Ziebart , Andrew Maas , J. Andrew Bagnell , and Anind K. Dey . 2008. Maximum Entropy Inverse Reinforcement Learning . In Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 3 (Chicago, Illinois) (AAAI'08). AAAI Press, 1433--1438. http:\/\/dl.acm.org\/citation.cfm?id=16 2027 0.1620297 Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey. 2008. Maximum Entropy Inverse Reinforcement Learning. In Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 3 (Chicago, Illinois) (AAAI'08). AAAI Press, 1433--1438. http:\/\/dl.acm.org\/citation.cfm?id=1620270.1620297"}],"event":{"name":"HRI '21: ACM\/IEEE International Conference on Human-Robot Interaction","location":"Boulder CO USA","acronym":"HRI '21","sponsor":["SIGAI ACM Special Interest Group on Artificial Intelligence","SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["Proceedings of the 2021 ACM\/IEEE International Conference on Human-Robot Interaction"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3434073.3444667","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3434073.3444667","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3434073.3444667","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:03:00Z","timestamp":1750197780000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3434073.3444667"}},"subtitle":["Rethinking Human Input"],"short-title":[],"issued":{"date-parts":[[2021,3,8]]},"references-count":30,"alternative-id":["10.1145\/3434073.3444667","10.1145\/3434073"],"URL":"https:\/\/doi.org\/10.1145\/3434073.3444667","relation":{},"subject":[],"published":{"date-parts":[[2021,3,8]]},"assertion":[{"value":"2021-03-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}