{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:20:40Z","timestamp":1750220440521,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":25,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,3,8]],"date-time":"2021-03-08T00:00:00Z","timestamp":1615161600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,3,8]]},"DOI":"10.1145\/3434074.3447189","type":"proceedings-article","created":{"date-parts":[[2021,3,8]],"date-time":"2021-03-08T01:33:18Z","timestamp":1615167198000},"page":"344-348","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["When Oracles Go Wrong: Using Preferences as a Means to Explore"],"prefix":"10.1145","author":[{"given":"Isaac S.","family":"Sheidlower","sequence":"first","affiliation":[{"name":"Tufts University, Somerville, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elaine Schaertl","family":"Short","sequence":"additional","affiliation":[{"name":"Tufts University, Medford, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,3,8]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Twenty-first international conference on Machine learning - ICML '04","author":"Abbeel Pieter","year":"2020","unstructured":"Pieter Abbeel and Andrew Y. Ng . 2004. Apprenticeship learning via inverse reinforcement learning . In Twenty-first international conference on Machine learning - ICML '04 . event-place: Banff, Alberta, Canada. ACM Press, 1. doi:10.1145\/1015330.1015430. Retrieved 09\/13\/ 2020 from http:\/\/portal.acm.org\/citation.cfm?doid=1015330.1015430. 10.1145\/1015330.1015430 Pieter Abbeel and Andrew Y. Ng. 2004. Apprenticeship learning via inverse reinforcement learning. In Twenty-first international conference on Machine learning - ICML '04. event-place: Banff, Alberta, Canada. ACM Press, 1. doi:10.1145\/1015330.1015430. Retrieved 09\/13\/2020 from http:\/\/portal.acm.org\/citation.cfm?doid=1015330.1015430."},{"volume-title":"Machine Learning and Knowledge Discovery in Databases.","author":"Akrour Riad","key":"e_1_3_2_1_2_1","unstructured":"Riad Akrour , Marc Schoenauer , and Michele Sebag . 2011. Preference-Based Policy Learning. en . In Machine Learning and Knowledge Discovery in Databases. Volume 6911 . Dimitrios Gunopulos, Thomas Hofmann , Donato Malerba, and Michalis Vazirgiannis, editors. Series Title : Lecture Notes in Computer Science. Springer Berlin Heidelberg , Berlin, Heidelberg, 12--27. isbn: 978--3--642--23779--9978--3--642--23780--5. doi: 10.1007\/978--3--642--23780--5_11. Retrieved 12\/01\/2020from http:\/\/link.springer.com\/10.1007\/978--3--642--23780--5_11. 10.1007\/978--3--642--23780--5_11 Riad Akrour, Marc Schoenauer, and Michele Sebag. 2011. Preference-Based Policy Learning. en. In Machine Learning and Knowledge Discovery in Databases. Volume 6911. Dimitrios Gunopulos, Thomas Hofmann, Donato Malerba, and Michalis Vazirgiannis, editors. Series Title: Lecture Notes in Computer Science. Springer Berlin Heidelberg, Berlin, Heidelberg, 12--27. isbn: 978--3--642--23779--9978--3--642--23780--5. doi: 10.1007\/978--3--642--23780--5_11. Retrieved 12\/01\/2020from http:\/\/link.springer.com\/10.1007\/978--3--642--23780--5_11."},{"key":"e_1_3_2_1_3_1","volume-title":"Sophie Saskin, and Michael L. Littman.","author":"Arumugam Dilip","year":"2019","unstructured":"Dilip Arumugam , Jun Ki Lee , Sophie Saskin, and Michael L. Littman. 2019 . Deep Reinforcement Learning from Policy-Dependent Human Feedback. arXiv:1902.04257[cs, stat], (February 2019). arXiv: 1902.04257. Retrieved 10\/05\/2020 from http:\/\/arxiv.org\/abs\/1902.04257. Dilip Arumugam, Jun Ki Lee, Sophie Saskin, and Michael L. Littman. 2019. Deep Reinforcement Learning from Policy-Dependent Human Feedback. arXiv:1902.04257[cs, stat], (February 2019). arXiv: 1902.04257. Retrieved 10\/05\/2020 from http:\/\/arxiv.org\/abs\/1902.04257."},{"key":"e_1_3_2_1_4_1","volume-title":"Proceedings of the 2020 ACM\/IEEE International Conference on Human-Robot Interaction. event-place: Cambridge United Kingdom. ACM,(March 2020","author":"Bobu Andreea","year":"2020","unstructured":"Andreea Bobu , Dexter R. R. Scobee , Jaime F. Fisac , S. Shankar Sastry , and Anca D. Dragan . 2020. LESS is More: Rethinking Probabilistic Models of Human Behavior . In Proceedings of the 2020 ACM\/IEEE International Conference on Human-Robot Interaction. event-place: Cambridge United Kingdom. ACM,(March 2020 ), 429--437. isbn: 978--1--4503--6746--2.doi: 10.1145\/3319502.3374811. Retrieved 09\/13\/ 2020 from https:\/\/dl.acm.org\/doi\/10.1145\/3319502.3374811. 10.1145\/3319502.3374811 Andreea Bobu, Dexter R. R. Scobee, Jaime F. Fisac, S. Shankar Sastry, and Anca D. Dragan. 2020. LESS is More: Rethinking Probabilistic Models of Human Behavior. In Proceedings of the 2020 ACM\/IEEE International Conference on Human-Robot Interaction. event-place: Cambridge United Kingdom. ACM,(March 2020), 429--437. isbn: 978--1--4503--6746--2.doi: 10.1145\/3319502.3374811. Retrieved 09\/13\/2020 from https:\/\/dl.acm.org\/doi\/10.1145\/3319502.3374811."},{"key":"e_1_3_2_1_5_1","volume-title":"Guy Hoffman, and Andrea L Thomaz.","author":"Faulkner Taylor Kessler","year":"2019","unstructured":"Taylor Kessler Faulkner , Reymundo A Gutierrez , Elaine Schaertl Short , Guy Hoffman, and Andrea L Thomaz. 2019 . Active Attention-Modified Policy Shaping , 9. Taylor Kessler Faulkner, Reymundo A Gutierrez, Elaine Schaertl Short, Guy Hoffman, and Andrea L Thomaz. 2019. Active Attention-Modified Policy Shaping, 9."},{"key":"e_1_3_2_1_6_1","volume-title":"andAndrea L Thomaz","author":"Griffith Shane","year":"2013","unstructured":"Shane Griffith , Kaushik Subramanian , Jonathan Scholz , Charles L Isbell , andAndrea L Thomaz . 2013 . Policy Shaping : Integrating Human Feedback with Reinforcement Learning. In Advances in Neural Information Processing Systems 26. C. J. C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Q. Weinberger, editors. Curran Associates, Inc., 2625--2633. Retrieved 10\/11\/2020 from http:\/\/papers.nips.cc\/paper\/5187-policy-shaping-integrating-human-feedback-with-reinforcement-learning.pdf. Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L Isbell, andAndrea L Thomaz. 2013. Policy Shaping: Integrating Human Feedback with Reinforcement Learning. In Advances in Neural Information Processing Systems 26. C. J. C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Q. Weinberger, editors. Curran Associates, Inc., 2625--2633. Retrieved 10\/11\/2020 from http:\/\/papers.nips.cc\/paper\/5187-policy-shaping-integrating-human-feedback-with-reinforcement-learning.pdf."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12369-012-0161-z"},{"key":"e_1_3_2_1_8_1","volume-title":"Advances in Neural Information Processing Systems 29","author":"Hadfield-Menell Dylan","year":"2020","unstructured":"Dylan Hadfield-Menell , Stuart J Russell , Pieter Abbeel , and Anca Dragan . 2016. Cooperative Inverse Reinforcement Learning . In Advances in Neural Information Processing Systems 29 . D. D. Lee, M. Sugiyama, U. V. Luxburg, I. Guyon, and R. Garnett, editors. Curran Associates, Inc. , 3909--3917. Retrieved 09\/13\/ 2020 from http:\/\/papers.nips.cc\/paper\/6420- cooperative- inverse- reinforcement-learning.pdf. Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan. 2016. Cooperative Inverse Reinforcement Learning. In Advances in Neural Information Processing Systems 29. D. D. Lee, M. Sugiyama, U. V. Luxburg, I. Guyon, and R. Garnett, editors. Curran Associates, Inc., 3909--3917. Retrieved 09\/13\/2020 from http:\/\/papers.nips.cc\/paper\/6420- cooperative- inverse- reinforcement-learning.pdf."},{"key":"e_1_3_2_1_9_1","volume-title":"Advances in Neural Information Processing Systems 29","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho and Stefano Ermon . 2016. Generative Adversarial Imitation Learning . In Advances in Neural Information Processing Systems 29 . D. D. Lee, M. Sugiyama, U. V. Luxburg, I. Guyon, and R. Garnett, editors. Curran Associates, Inc. , 4565--4573. Retrieved 10\/19\/ 2020 from http:\/\/papers.nips.cc\/paper\/6391-generative-adversarial-imitation-learning.pdf. Jonathan Ho and Stefano Ermon. 2016. Generative Adversarial Imitation Learning. In Advances in Neural Information Processing Systems 29. D. D. Lee, M. Sugiyama, U. V. Luxburg, I. Guyon, and R. Garnett, editors. Curran Associates, Inc., 4565--4573. Retrieved 10\/19\/2020 from http:\/\/papers.nips.cc\/paper\/6391-generative-adversarial-imitation-learning.pdf."},{"key":"e_1_3_2_1_10_1","unstructured":"Borja Ibarz Jan Leike Tobias Pohlen Geoffrey Irving Shane Legg and Dario Amodei. [n. d.] Reward learning from human preferences and demonstrations in Atari. en 13.  Borja Ibarz Jan Leike Tobias Pohlen Geoffrey Irving Shane Legg and Dario Amodei. [n. d.] Reward learning from human preferences and demonstrations in Atari. en 13."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1597735.1597738"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3277904"},{"key":"e_1_3_2_1_13_1","volume-title":"AC-Teach: A Bayesian Actor-Critic Method for Policy Learning with an Ensemble of Suboptimal Teachers. arXiv:1909.04121 [cs, stat],(December","author":"Kurenkov Andrey","year":"2019","unstructured":"Andrey Kurenkov , Ajay Mandlekar , Roberto Martin-Martin , Silvio Savarese , and Animesh Garg . 2019. AC-Teach: A Bayesian Actor-Critic Method for Policy Learning with an Ensemble of Suboptimal Teachers. arXiv:1909.04121 [cs, stat],(December 2019 ). arXiv: 1909.04121. Retrieved 10\/19\/2020 from http:\/\/arxiv.org\/abs\/1909.04121. Andrey Kurenkov, Ajay Mandlekar, Roberto Martin-Martin, Silvio Savarese, and Animesh Garg. 2019. AC-Teach: A Bayesian Actor-Critic Method for Policy Learning with an Ensemble of Suboptimal Teachers. arXiv:1909.04121 [cs, stat],(December 2019). arXiv: 1909.04121. Retrieved 10\/19\/2020 from http:\/\/arxiv.org\/abs\/1909.04121."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3319502.3374832"},{"key":"e_1_3_2_1_15_1","volume-title":"Littman","author":"MacGlashan James","year":"2017","unstructured":"James MacGlashan , Mark K. Ho , Robert Loftin , Bei Peng , David Roberts , Matthew E. Taylor , and Michael L . Littman . 2017 . Interactive Learning from Policy-Dependent Human Feedback. arXiv:1701.06049 [cs], (January 2017). arXiv: 1701.06049. Retrieved 10\/05\/2020 from http:\/\/arxiv.org\/abs\/1701.06049. James MacGlashan, Mark K. Ho, Robert Loftin, Bei Peng, David Roberts, Matthew E. Taylor, and Michael L. Littman. 2017. Interactive Learning from Policy-Dependent Human Feedback. arXiv:1701.06049 [cs], (January 2017). arXiv: 1701.06049. Retrieved 10\/05\/2020 from http:\/\/arxiv.org\/abs\/1701.06049."},{"key":"e_1_3_2_1_16_1","unstructured":"St\u00e9phane Ross Geoffrey J Gordon and J Andrew Bagnell. [n. d.] A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. en 9.  St\u00e9phane Ross Geoffrey J Gordon and J Andrew Bagnell. [n. d.] A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. en 9."},{"key":"e_1_3_2_1_17_1","unstructured":"Himanshu Sahni Brent Harrison Thomas Cederborg Charles Isbell Kaushik Subramanian and Andrea Thomaz. [n. d.] Policy Shaping in Domains with Multiple Optimal Policies (Extended Abstract) 2.  Himanshu Sahni Brent Harrison Thomas Cederborg Charles Isbell Kaushik Subramanian and Andrea Thomaz. [n. d.] Policy Shaping in Domains with Multiple Optimal Policies (Extended Abstract) 2."},{"key":"e_1_3_2_1_18_1","unstructured":"Matthew E Taylor Halit Bener Suay and Sonia Chernova. [n. d.] Integrating reinforcement learning with human demonstrations of varying ability. en 8.  Matthew E Taylor Halit Bener Suay and Sonia Chernova. [n. d.] Integrating reinforcement learning with human demonstrations of varying ability. en 8."},{"key":"e_1_3_2_1_19_1","unstructured":"Andrea L Thomaz. [n. d.] Reinforcement Learning with Human Teachers: Evidence of Feedback and Guidance with Implications for Learning Performance 6.  Andrea L Thomaz. [n. d.] Reinforcement Learning with Human Teachers: Evidence of Feedback and Guidance with Implications for Learning Performance 6."},{"key":"e_1_3_2_1_20_1","volume-title":"Behavioral Cloning from Observation. arXiv:1805.01954 [cs], (May","author":"Torabi Faraz","year":"2018","unstructured":"Faraz Torabi , Garrett Warnell , and Peter Stone . 2018. Behavioral Cloning from Observation. arXiv:1805.01954 [cs], (May 2018 ). arXiv: 1805.01954. Retrieved 10\/19\/2020 from http:\/\/arxiv.org\/abs\/1805.01954. Faraz Torabi, Garrett Warnell, and Peter Stone. 2018. Behavioral Cloning from Observation. arXiv:1805.01954 [cs], (May 2018). arXiv: 1805.01954. Retrieved 10\/19\/2020 from http:\/\/arxiv.org\/abs\/1805.01954."},{"key":"e_1_3_2_1_21_1","volume-title":"Deep TAMER: Interactive Agent Shaping in High-Dimensional State Spaces.arXiv:1709.10163 [cs], (January","author":"Warnell Garrett","year":"2018","unstructured":"Garrett Warnell , Nicholas Waytowich , Vernon Lawhern , and Peter Stone . 2018. Deep TAMER: Interactive Agent Shaping in High-Dimensional State Spaces.arXiv:1709.10163 [cs], (January 2018 ). arXiv: 1709.10163. Retrieved 09\/29\/2020 from http:\/\/arxiv.org\/abs\/1709.10163. Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone. 2018. Deep TAMER: Interactive Agent Shaping in High-Dimensional State Spaces.arXiv:1709.10163 [cs], (January 2018). arXiv: 1709.10163. Retrieved 09\/29\/2020 from http:\/\/arxiv.org\/abs\/1709.10163."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1177\/0278364920910802"},{"key":"e_1_3_2_1_23_1","unstructured":"Christian Wirth Riad Akrour Gerhard Neumann and Johannes F\u00fcrnkranz. [n.d.] A Survey of Preference-Based Reinforcement Learning Methods. en 46.  Christian Wirth Riad Akrour Gerhard Neumann and Johannes F\u00fcrnkranz. [n.d.] A Survey of Preference-Based Reinforcement Learning Methods. en 46."},{"key":"e_1_3_2_1_24_1","volume-title":"Dongxu wang, and Guangliang Li","author":"Yu Chao","year":"2018","unstructured":"Chao Yu , Tianpei Yang , Wenxuan Zhu , Dongxu wang, and Guangliang Li . 2018 . Learning Shaping Strategies in Human-in-the-loop Interactive Reinforcement Learning. arXiv:1811.04272 [cs], (November 2018). arXiv: 1811.04272. Retrieved 12\/04\/2020 from http:\/\/arxiv.org\/abs\/1811.04272. Chao Yu, Tianpei Yang, Wenxuan Zhu, Dongxu wang, and Guangliang Li. 2018. Learning Shaping Strategies in Human-in-the-loop Interactive Reinforcement Learning. arXiv:1811.04272 [cs], (November 2018). arXiv: 1811.04272. Retrieved12\/04\/2020 from http:\/\/arxiv.org\/abs\/1811.04272."},{"key":"e_1_3_2_1_25_1","volume-title":"Human-guided Robot Behavior Learning: A GAN-assisted Preference-based Reinforcement Learning Approach.arXiv:2010.07467 [cs], (October","author":"Zhan Huixin","year":"2020","unstructured":"Huixin Zhan , Feng Tao , and Yongcan Cao . 2020. Human-guided Robot Behavior Learning: A GAN-assisted Preference-based Reinforcement Learning Approach.arXiv:2010.07467 [cs], (October 2020 ). arXiv: 2010.07467. Retrieved 12\/01\/2020from http:\/\/arxiv.org\/abs\/2010.07467. Huixin Zhan, Feng Tao, and Yongcan Cao. 2020. Human-guided Robot Behavior Learning: A GAN-assisted Preference-based Reinforcement Learning Approach.arXiv:2010.07467 [cs], (October 2020). arXiv: 2010.07467. Retrieved 12\/01\/2020from http:\/\/arxiv.org\/abs\/2010.07467."}],"event":{"name":"HRI '21: ACM\/IEEE International Conference on Human-Robot Interaction","sponsor":["SIGAI ACM Special Interest Group on Artificial Intelligence","SIGCHI ACM Special Interest Group on Computer-Human Interaction"],"location":"Boulder CO USA","acronym":"HRI '21"},"container-title":["Companion of the 2021 ACM\/IEEE International Conference on Human-Robot Interaction"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3434074.3447189","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3434074.3447189","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:47:28Z","timestamp":1750193248000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3434074.3447189"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,8]]},"references-count":25,"alternative-id":["10.1145\/3434074.3447189","10.1145\/3434074"],"URL":"https:\/\/doi.org\/10.1145\/3434074.3447189","relation":{},"subject":[],"published":{"date-parts":[[2021,3,8]]},"assertion":[{"value":"2021-03-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}