{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T03:00:04Z","timestamp":1776394804475,"version":"3.51.2"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2018,7,31]],"date-time":"2018-07-31T00:00:00Z","timestamp":1532995200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100007297","name":"Office of Naval Research","doi-asserted-by":"publisher","award":["N000141410003"],"award-info":[{"award-number":["N000141410003"]}],"id":[{"id":"10.13039\/100007297","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["J. Hum.-Robot Interact."],"published-print":{"date-parts":[[2018,7,31]]},"abstract":"<jats:p>A goal of interactive machine learning (IML) is to enable people with no specialized training to intuitively teach intelligent agents how to perform tasks. Toward achieving that goal, we are studying how the design of the interaction method for a Bayesian Q-Learning algorithm impacts aspects of the human\u2019s experience of teaching the agent using human-centric metrics such as frustration in addition to traditional ML performance metrics. This study investigated two methods of natural language instruction: critique and action advice. We conducted a human-in-the-loop experiment in which people trained two agents with different teaching methods but, unknown to each participant, the same underlying reinforcement learning algorithm. The results show an agent that learns from action advice creates a better user experience compared to an agent that learns from binary critique in terms of frustration, perceived performance, transparency, immediacy, and perceived intelligence. We identified nine main characteristics of an IML algorithm\u2019s design that impact the human\u2019s experience with the agent, including using human instructions about the future, compliance with input, empowerment, transparency, immediacy, a deterministic interaction, the complexity of the instructions, accuracy of the speech recognition software, and the robust and flexible nature of the interaction algorithm.<\/jats:p>","DOI":"10.1145\/3277904","type":"journal-article","created":{"date-parts":[[2019,5,7]],"date-time":"2019-05-07T12:15:57Z","timestamp":1557231357000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":26,"title":["Interaction Algorithm Effect on Human Experience with Reinforcement Learning"],"prefix":"10.1145","volume":"7","author":[{"given":"Samantha","family":"Krening","sequence":"first","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Karen M.","family":"Feigh","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,10,24]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2008.4651020"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12369-009-0013-7"},{"key":"e_1_2_1_3_1","volume-title":"International Joint Conferences on Artificial Intelligence (IJCAI'15)","author":"Cederborg Thomas","year":"2015"},{"key":"e_1_2_1_4_1","unstructured":"Richard Dearden Nir Friedman and Stuart Russell. 1998. Bayesian Q-learning. In AAAI\/IAAI. 761--768.   Richard Dearden Nir Friedman and Stuart Russell. 1998. Bayesian Q-learning. In AAAI\/IAAI. 761--768."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/778712.778756"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0921-8890(02)00374-3"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0921-8890(02)00372-X"},{"key":"e_1_2_1_8_1","unstructured":"Shane Griffith Kaushik Subramanian Jonathan Scholz Charles Isbell and Andrea L. Thomaz. 2013. Policy shaping: Integrating human feedback with reinforcement learning. In Advances in Neural Information Processing Systems. 2625--2633.   Shane Griffith Kaushik Subramanian Jonathan Scholz Charles Isbell and Andrea L. Thomaz. 2013. Policy shaping: Integrating human feedback with reinforcement learning. In Advances in Neural Information Processing Systems. 2625--2633."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1177\/154193120605000909"},{"key":"e_1_2_1_10_1","unstructured":"Jay Heinrichs. 2017. Thank You for Arguing: What Aristotle Lincoln and Homer Simpson Can Teach Us about the Art of Persuasion. Three Rivers Press (CA).  Jay Heinrichs. 2017. Thank You for Arguing: What Aristotle Lincoln and Homer Simpson Can Teach Us about the Art of Persuasion. Three Rivers Press (CA)."},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the 2006 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201906)","volume":"1","author":"Huggins-Daines David"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/375735.376334"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10458-006-0005-z"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICTAI.2012.152"},{"key":"e_1_2_1_15_1","volume-title":"Proceedings of the 9th International Conference on Autonomous Agents and Multiagent Systems.","volume":"1","author":"Bradley Knox W.","year":"2010"},{"key":"e_1_2_1_16_1","unstructured":"Samantha Krening. 2018. Newtonian action advice: Integrating human verbal instruction with reinforcement learning. arXiv arXiv:1804.05821.  Samantha Krening. 2018. Newtonian action advice: Integrating human verbal instruction with reinforcement learning. arXiv arXiv:1804.05821."},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the Human Factors and Ergonomics Society Annual Meeting.","author":"Krening Samantha"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCDS.2016.2628365"},{"key":"e_1_2_1_19_1","volume-title":"The AAAI 2004 Workshop on Supervisory Control of Learning and Adaptive Systems.","author":"Kuhlmann Gregory","year":"2004"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2015.XI.018"},{"key":"e_1_2_1_21_1","unstructured":"Richard Maclin Jude Shavlik Lisa Torrey Trevor Walker and Edward Wild. 2005. Giving advice about preferred actions to reinforcement learners via knowledge-based kernel regression. In Association for the Advancement of Artificial Intelligence (AAAI'05). 819--824.   Richard Maclin Jude Shavlik Lisa Torrey Trevor Walker and Edward Wild. 2005. Giving advice about preferred actions to reinforcement learners via knowledge-based kernel regression. In Association for the Advancement of Artificial Intelligence (AAAI'05). 819--824."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-5010"},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the 2014 International Conference on Autonomous Agents and Multi-Agent Systems. International Foundation for Autonomous Agents and Multiagent Systems, 1069--1076","author":"Meri\u00e7li Cetin","year":"2014"},{"key":"e_1_2_1_24_1","volume-title":"International Conference on Machine Learning (ICML'00)","author":"Andrew"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.3115\/1219840.1219855"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1561\/1500000011"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1518\/001872097778543886"},{"key":"#cr-split#-e_1_2_1_28_1.1","unstructured":"Taylor Phillips. 2010. Put your money where your mouth is: The effects of southern vs. standard accent on perceptions of speakers. Social Sciences (2010) 53--56. https:\/\/scholar.google.com\/scholar?hl&equals;en&as_sdt&equals;&equals;&equals;0%2C6&q&equals;&equals;&equals;Taylor+Phillips.+2010.+Put+your+money+where+your+mouth+is%3A+The+effects+of+southern+vs.+standard+accent+on+perceptions+of+speakers.+S&btnG&equals;&equals;&equals"},{"key":"#cr-split#-e_1_2_1_28_1.2","unstructured":"Taylor Phillips. 2010. Put your money where your mouth is: The effects of southern vs. standard accent on perceptions of speakers. Social Sciences (2010) 53--56. https:\/\/scholar.google.com\/scholar?hl&equals;en&as_sdt&equals;&equals;&equals;0%2C6&q&equals;&equals;&equals;Taylor+Phillips.+2010.+Put+your+money+where+your+mouth+is%3A+The+effects+of+southern+vs.+standard+accent+on+perceptions+of+speakers.+S&btnG&equals;&equals;&equals;"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1037\/a0021522"},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the 2016 International Conference on AAMAS. International Foundation for AAMAS, 1455--1456","author":"Sahni Himanshu","year":"2016"},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the Twenty-Fifth International Florida Artificial Intelligence Research Society Conference (FLAIRS'12)","author":"Sivamurugan Manimaran Sivasamy","year":"2012"},{"key":"e_1_2_1_32_1","unstructured":"Burrhus Frederic Skinner. 1990. The behavior of organisms: An experimental analysis. BF Skinner Foundation.  Burrhus Frederic Skinner. 1990. The behavior of organisms: An experimental analysis. BF Skinner Foundation."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.7244649"},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201913)","volume":"1631","author":"Socher Richard","year":"2013"},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the 2016 International Conference on AAMAS. International Foundation for AAMAS, 447--456","author":"Subramanian Kaushik"},{"key":"e_1_2_1_36_1","unstructured":"Richard S. Sutton and Andrew G. Barto. 1998. Reinforcement Learning: An Introduction. Vol. 1. MIT Press Cambridge.   Richard S. Sutton and Andrew G. Barto. 1998. Reinforcement Learning: An Introduction. Vol. 1. MIT Press Cambridge."},{"key":"e_1_2_1_37_1","doi-asserted-by":"crossref","unstructured":"Stefanie Tellex Thomas Kollar Steven Dickerson Matthew R. Walter Ashis Gopal Banerjee Seth J. Teller and Nicholas Roy. 2011. Understanding natural language commands for robotic navigation and mobile manipulation. In Association for the Advancement of Artificial Intelligence (AAAI'11). Vol. 1. 2.   Stefanie Tellex Thomas Kollar Steven Dickerson Matthew R. Walter Ashis Gopal Banerjee Seth J. Teller and Nicholas Roy. 2011. Understanding natural language commands for robotic navigation and mobile manipulation. In Association for the Advancement of Artificial Intelligence (AAAI'11). Vol. 1. 2.","DOI":"10.1609\/aaai.v25i1.7979"},{"key":"e_1_2_1_38_1","volume-title":"International Joint Conferences on Artificial Intelligence (IJCAI'16)","author":"Thomason Jesse"},{"key":"e_1_2_1_39_1","volume-title":"International Joint Conferences on Artificial Intelligence (IJCAI'15)","author":"Thomason Jesse","year":"2015"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2007.09.009"}],"container-title":["ACM Transactions on Human-Robot Interaction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3277904","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3277904","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:54:09Z","timestamp":1750204449000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3277904"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,7,31]]},"references-count":41,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2018,7,31]]}},"alternative-id":["10.1145\/3277904"],"URL":"https:\/\/doi.org\/10.1145\/3277904","relation":{},"ISSN":["2573-9522"],"issn-type":[{"value":"2573-9522","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,7,31]]},"assertion":[{"value":"2018-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-08-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-10-24","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}