{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,5]],"date-time":"2026-01-05T18:29:22Z","timestamp":1767637762035,"version":"3.48.0"},"reference-count":16,"publisher":"Maximum Academic Press","license":[{"start":{"date-parts":[[2019,11,12]],"date-time":"2019-11-12T00:00:00Z","timestamp":1573516800000},"content-version":"unspecified","delay-in-days":315,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["The Knowledge Engineering Review"],"published-print":{"date-parts":[[2019]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>One way to address this low sample efficiency of reinforcement learning (RL) is to employ human expert demonstrations to speed up the RL process (RL from demonstration or RLfD). The research so far has focused on demonstrations from a single expert. However, little attention has been given to the case where demonstrations are collected from multiple experts, whose expertise may vary on different aspects of the task. In such scenarios, it is likely that the demonstrations will contain conflicting advice in many parts of the state space. We propose a two-level Q-learning algorithm, in which the RL agent not only learns the policy of deciding on the optimal action but also learns to select the most trustworthy expert according to the current state. Thus, our approach removes the traditional assumption that demonstrations come from one single source and are mostly conflict-free. We evaluate our technique on three different domains and the results show that the state-of-the-art RLfD baseline fails to converge or performs similarly to conventional Q-learning. In contrast, the performance level of our novel algorithm increases with more experts being involved in the learning process and the proposed approach has the capability to handle demonstration conflicts well.<\/jats:p>","DOI":"10.1017\/s0269888919000092","type":"journal-article","created":{"date-parts":[[2019,11,12]],"date-time":"2019-11-12T04:59:46Z","timestamp":1573534786000},"source":"Crossref","is-referenced-by-count":4,"title":["Two-level Q-learning: learning from conflict demonstrations"],"prefix":"10.48130","volume":"34","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8500-1165","authenticated-orcid":false,"given":"Mao","family":"Li","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yi","family":"Wei","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daniel","family":"Kudenko","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"27968","published-online":{"date-parts":[[2019,11,12]]},"reference":[{"key":"S0269888919000092_ref15","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"},{"key":"S0269888919000092_ref14","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2017\/422"},{"key":"S0269888919000092_ref13","unstructured":"Taylor, M. E. , Suay, H. B. & Chernova, S. 2011. Integrating reinforcement learning with human demonstrations of varying ability. In The 10th International Conference on Autonomous Agents and Multiagent Systems-Volume 2, 617\u2013624. International Foundation for Autonomous Agents and Multiagent Systems."},{"key":"S0269888919000092_ref4","doi-asserted-by":"publisher","DOI":"10.1613\/jair.639"},{"key":"S0269888919000092_ref1","unstructured":"Brockman, G. , Cheung, V. , Pettersson, L. , Schneider, J. , Schulman, J. , Tang, J. & Zaremba, W. 2016. Openai gym. arXiv preprint arXiv:1606.01540."},{"key":"S0269888919000092_ref5","doi-asserted-by":"publisher","DOI":"10.1145\/1160633.1160762"},{"key":"S0269888919000092_ref11","unstructured":"Schaal, S. 1997. Learning from demonstration. In Proceedings of the 1997 Conference on Neural Information Processing Systems (NIPS97). Denver, CO, pp. 1040\u20131046."},{"key":"S0269888919000092_ref3","unstructured":"Devlin, S. & Kudenko, D. 2012. Dynamic potential-based reward shaping. In Proceedings of the 11th International Conference on Autonomous Agents and Multiagent Systems-Volume 1, 433\u2013440. International Foundation for Autonomous Agents and Multiagent Systems."},{"key":"S0269888919000092_ref6","doi-asserted-by":"crossref","unstructured":"Hester, T. , Vecerik, M. , Pietquin, O. , Lanctot, M. , Schaul, T. , Piot, B. , Horgan, D. , Quan, J. , Sendonaris, A. , Dulac-Arnold, G. , Osband, I. , Agapiou, J. 2018. Deep Q-learning from demonstrations. In Thirty-Second AAAI Conference on Artificial Intelligence 2018 Apr 29.","DOI":"10.1609\/aaai.v32i1.11757"},{"key":"S0269888919000092_ref12","volume-title":"Reinforcement Learning: An Introduction","volume":"1","author":"Sutton","year":"1998"},{"key":"S0269888919000092_ref9","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"S0269888919000092_ref2","unstructured":"Brys, T. , Harutyunyan, A. , Suay, H. B. , Chernova, S. , Taylor, M. E. & Now\u00e9, A. 2015. Reinforcement learning from demonstration through shaping. In Proceedings of the 24th International Conference on Artificial Intelligence, 3352\u20133358. AAAI Press."},{"key":"S0269888919000092_ref7","unstructured":"Kaelbling, L. P. 1993. Hierarchical learning in stochastic domains: Preliminary results. In Proceedings of the Tenth International Conference on Machine Learning, 951, 167\u2013173."},{"key":"S0269888919000092_ref16","unstructured":"Wiewiora, E. , Cottrell, G. W. & Elkan, C. 2003. Principled methods for advising reinforcement learning agents. In Proceedings of the 20th International Conference on Machine Learning (ICML-03), 792\u2013799."},{"key":"S0269888919000092_ref10","unstructured":"Ng, A. Y. , Harada, D. & Russell, S. 1999. Policy invariance under reward transformations: Theory and application to reward shaping. In ICML, 99, 278\u2013287."},{"key":"S0269888919000092_ref8","unstructured":"Kingma, D. P. & Ba, J. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980."}],"container-title":["The Knowledge Engineering Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S0269888919000092","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,5]],"date-time":"2026-01-05T14:42:12Z","timestamp":1767624132000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S0269888919000092\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019]]},"references-count":16,"alternative-id":["S0269888919000092"],"URL":"https:\/\/doi.org\/10.1017\/s0269888919000092","relation":{},"ISSN":["0269-8889","1469-8005"],"issn-type":[{"type":"print","value":"0269-8889"},{"type":"electronic","value":"1469-8005"}],"subject":[],"published":{"date-parts":[[2019]]},"article-number":"e14"}}