{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T07:15:43Z","timestamp":1740122143585,"version":"3.37.3"},"reference-count":86,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2021,8,16]],"date-time":"2021-08-16T00:00:00Z","timestamp":1629072000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,8,16]],"date-time":"2021-08-16T00:00:00Z","timestamp":1629072000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Auton Agent Multi-Agent Syst"],"published-print":{"date-parts":[[2021,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Deep reinforcement learning methods have achieved significant successes in complex decision-making problems. In fact, they traditionally rely on well-designed extrinsic rewards, which limits their applicability to many real-world tasks where rewards are naturally sparse. While cloning behaviors provided by an expert is a promising approach to the exploration problem, learning from a fixed set of demonstrations may be impracticable due to lack of state coverage or distribution mismatch\u2014when the learner\u2019s goal deviates from the demonstrated behaviors. Besides, we are interested in learning how to reach a wide range of goals from the same set of demonstrations. In this work we propose a novel goal-conditioned method that leverages very small sets of goal-driven demonstrations to massively accelerate the learning process. Crucially, we introduce the concept of active goal-driven demonstrations to query the demonstrator only in hard-to-learn and uncertain regions of the state space. We further present a strategy for prioritizing sampling of goals where the disagreement between the expert and the policy is maximized. We evaluate our method on a variety of benchmark environments from the Mujoco domain. Experimental results show that our method outperforms prior imitation learning approaches in most of the tasks in terms of exploration efficiency and average scores.<\/jats:p>","DOI":"10.1007\/s10458-021-09527-5","type":"journal-article","created":{"date-parts":[[2021,8,16]],"date-time":"2021-08-16T08:02:52Z","timestamp":1629100972000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Goal-driven active learning"],"prefix":"10.1007","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9856-0038","authenticated-orcid":false,"given":"Nicolas","family":"Bougie","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ryutaro","family":"Ichise","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,8,16]]},"reference":[{"key":"9527_CR1","doi-asserted-by":"crossref","unstructured":"Abbeel, P., & Ng, A. Y. (2004). Apprenticeship learning via inverse reinforcement learning (pp. 1\u20138).","DOI":"10.1145\/1015330.1015430"},{"key":"9527_CR2","unstructured":"Agarwal, P., de\u00a0Beaucorps, P., & de\u00a0Charette, R. (2021). Sparse curriculum reinforcement learning for end-to-end driving. 2103.09189."},{"key":"9527_CR3","unstructured":"Amir, O., Kamar, E., Kolobov, A., & Grosz, B. J. (2016). Interactive teaching strategies for agent training. In Proceedings of the twenty-fifth international joint conference on artificial intelligence (pp. 804\u2013811)."},{"key":"9527_CR4","unstructured":"Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O.P., & Zaremba, W. (2017). Hindsight experience replay. In Advances in neural information processing systems (pp. 5048\u20135058)."},{"issue":"5","key":"9527_CR5","doi-asserted-by":"publisher","first-page":"469","DOI":"10.1016\/j.robot.2008.10.024","volume":"57","author":"BD Argall","year":"2009","unstructured":"Argall, B. D., Chernova, S., Veloso, M., & Browning, B. (2009). A survey of robot learning from demonstration. Robotics and autonomous systems, 57(5), 469\u2013483.","journal-title":"Robotics and autonomous systems"},{"key":"9527_CR6","unstructured":"Bassich, A., Foglino, F., Leonetti, M., & Kudenko, D. (2020). Curriculum learning with a progression function. arXiv preprint arXiv:200800511."},{"key":"9527_CR7","unstructured":"Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., & Munos, R. (2016). Unifying count-based exploration and intrinsic motivation. In Proceedings of advances in neural information processing systems (pp. 1471\u20131479)."},{"key":"9527_CR8","doi-asserted-by":"crossref","unstructured":"Bengio, Y., Louradour, J., Collobert, R., & Weston, J. (2009). Curriculum learning. In Proceedings of the 26th annual international conference on machine learning (pp. 41\u201348).","DOI":"10.1145\/1553374.1553380"},{"key":"9527_CR9","doi-asserted-by":"crossref","unstructured":"Bougie, N., & Ichise, R. (2019). Skill-based curiosity for intrinsically motivated reinforcement learning. Machine Learning, 493\u2013512.","DOI":"10.1007\/s10994-019-05845-8"},{"issue":"2","key":"9527_CR10","doi-asserted-by":"publisher","first-page":"33","DOI":"10.1145\/3243064.3243067","volume":"18","author":"N Bougie","year":"2018","unstructured":"Bougie, N., Cheng, L. K., & Ichise, R. (2018). Combining deep reinforcement learning with prior knowledge and reasoning. ACM SIGAPP Applied Computing Review, 18(2), 33\u201345.","journal-title":"ACM SIGAPP Applied Computing Review"},{"key":"9527_CR11","unstructured":"Burda, Y., Edwards, H., Storkey, A., & Klimov, O. (2019). Exploration by random network distillation. In Proceedings of the international conference on learning representations."},{"issue":"11","key":"9527_CR12","doi-asserted-by":"publisher","first-page":"3207","DOI":"10.1007\/s00477-018-1573-6","volume":"32","author":"AJ Cannon","year":"2018","unstructured":"Cannon, A. J. (2018). Non-crossing nonlinear regression quantiles by monotone composite quantile regression neural network, with application to rainfall extremes. Stochastic Environmental Research and Risk Assessment, 32(11), 3207\u20133225. https:\/\/doi.org\/10.1007\/s00477-018-1573-6","journal-title":"Stochastic Environmental Research and Risk Assessment"},{"key":"9527_CR13","unstructured":"Chen, M., Wang, Y., Liu, T., Yang, Z., Li, X., Wang, Z., & Zhao, T. (2020). On computation and generalization of generative adversarial imitation learning. In International conference on learning representations"},{"key":"9527_CR14","doi-asserted-by":"crossref","unstructured":"Chen, S. A., Tangkaratt, V., Lin, H. T., & Sugiyama, M. (2019). Active deep q-learning with demonstration. Machine Learning, 1\u201327.","DOI":"10.1007\/s10994-019-05849-4"},{"key":"9527_CR15","unstructured":"Chen, X., Zhou, Z., Wang, Z., Wang, C., Wu, Y., & Ross, K. (2020b). Bail: Best-action imitation learning for batch deep reinforcement learning. Advances in Neural Information Processing Systems 33"},{"key":"9527_CR16","doi-asserted-by":"crossref","unstructured":"Chernova, S., & Veloso, M. (2007). Confidence-based policy learning from demonstration using gaussian mixture models. In Proceedings of the international joint conference on autonomous agents and multiagent systems (pp. 1\u20138).","DOI":"10.1145\/1329125.1329407"},{"key":"9527_CR17","unstructured":"Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. In Advances in Neural information processing systems (pp. 4299\u20134307)."},{"issue":"1","key":"9527_CR18","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s10458-019-09430-0","volume":"34","author":"FL Da Silva","year":"2020","unstructured":"Da Silva, F. L., Warnell, G., Costa, A. H. R., & Stone, P. (2020). Agents teaching agents: Asurvey on inter-agent transfer learning. Autonomous Agents and Multi-Agent Systems, 34(1), 1\u201317.","journal-title":"Autonomous Agents and Multi-Agent Systems"},{"key":"9527_CR19","unstructured":"Ding, Y., Florensa, C., Abbeel, P., & Phielipp, M. (2019). Goal-conditioned imitation learning. In Advances in neural information processing systems (pp. 15298\u201315309)."},{"key":"9527_CR20","unstructured":"Duan, Y., Andrychowicz, M., Stadie, B., Jonathan\u00a0Ho, O., Schneider, J., Sutskever, I., Abbeel, P., & Zaremba, W. (2017). One-shot imitation learning. In Advances in neural information processing systems (Vol.\u00a030)."},{"key":"9527_CR21","unstructured":"Eysenbach, B., Gupta, A., Ibarz, J., & Levine, S. (2019). Diversity is all you need: Learning skills without a reward function. In International conference on learning representations."},{"issue":"1","key":"9527_CR22","doi-asserted-by":"publisher","first-page":"21","DOI":"10.3390\/make1010002","volume":"1","author":"A Fachantidis","year":"2019","unstructured":"Fachantidis, A., Taylor, M. E., & Vlahavas, I. (2019). Learning to teach reinforcement learning agents. Machine Learning and Knowledge Extraction, 1(1), 21\u201342.","journal-title":"Machine Learning and Knowledge Extraction"},{"key":"9527_CR23","unstructured":"Fan, Y., Tian, F., Qin, T., Li, X. Y., & Liu, T. Y. (2018). Learning to teach. In International conference on learning representations."},{"key":"9527_CR24","unstructured":"Finn, C., Abbeel, P., & Levine, S. (2017). Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning (pp. 1126\u20131135)."},{"key":"9527_CR25","unstructured":"Finn, C., Yu, T., Zhang, T., Abbeel, P., & Levine, S. (2017). One-shot visual imitation learning via meta-learning. In Proceedings of the 1st annual conference on robot learning, Proceedings of machine learning research (Vol.\u00a078, pp. 357\u2013368)."},{"key":"9527_CR26","unstructured":"Florensa, C., Held, D., Geng, X., & Abbeel, P. (2017). Automatic goal generation for reinforcement learning agents. arXiv preprint: arXiv:170506366."},{"key":"9527_CR27","unstructured":"Florensa, C., Held, D., Wulfmeier, M., Zhang, M., & Abbeel, P. (2017) Reverse curriculum generation for reinforcement learning. arXiv preprint arXiv:170705300."},{"key":"9527_CR28","unstructured":"Forestier, S., Mollard, Y., & Oudeyer, P.Y. (2017). Intrinsically motivated goal exploration processes with automatic curriculum learning. arXiv preprint: arXiv:170802190."},{"key":"9527_CR29","unstructured":"Fujimoto, S., Conti, E., Ghavamzadeh, M., & Pineau, J. (2019). Benchmarking batch deep reinforcement learning algorithms. arXiv preprint arXiv:191001708"},{"key":"9527_CR30","unstructured":"Fujimoto, S., Meger, D., & Precup, D. (2019). Off-policy deep reinforcement learning without exploration. In International conference on machine learning, PMLR (pp. 2052\u20132062)."},{"key":"9527_CR31","unstructured":"Gal, Y., & Ghahramani, Z. (2016). Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of the international conference on machine learning (pp. 1050\u20131059)."},{"key":"9527_CR32","unstructured":"Gasthaus, J., Benidis, K., Wang, Y., Rangapuram, S. S., Salinas, D., Flunkert, V., & Januschowski, T. (2019). Probabilistic forecasting with spline quantile function rnns. In The 22nd international conference on artificial intelligence and statistics (pp. 1901\u20131910)."},{"key":"9527_CR33","unstructured":"Graves, A., Bellemare, M. G., Menick, J., Munos, R., & Kavukcuoglu, K. (2017). Automated curriculum learning for neural networks. In Proceedings of the 34th international conference on machine learning (pp. 1311\u20131320)."},{"key":"9527_CR34","doi-asserted-by":"crossref","unstructured":"Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., Dulac-Arnold, G., Agapiou, J., Leibo, J. Z., & Gruslys, A. (2018). Deep q-learning from demonstrations. In Proceedings of the annual meeting of the association for the advancement of artificial intelligence.","DOI":"10.1609\/aaai.v32i1.11757"},{"key":"9527_CR35","unstructured":"Ho, J., & Ermon, S. (2016). Generative adversarial imitation learning. In Advances in neural information processing systems (pp. 4565\u20134573)."},{"key":"9527_CR36","unstructured":"Houthooft, R., Chen, X., Chen, X., Duan, Y., Schulman, J., De Turck, F., & Abbeel, P. (2016). Vime: Variational information maximizing exploration. In Proceedings of advances in neural information processing systems (pp 1109\u20131117)."},{"key":"9527_CR37","unstructured":"Hsu, D. (2019). A new framework for query efficient active imitation learning. arXiv preprint arXiv:191213037."},{"key":"9527_CR38","unstructured":"Huang, S., & Onta\u00f1\u00f3n, S. (2020). Action guidance: Getting the best of sparse rewards and shaped rewards for real-time strategy games. 2010.03956."},{"key":"9527_CR39","unstructured":"Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., & Amodei, D. (2018). Reward learning from human preferences and demonstrations in atari. In Advances in neural information processing systems (pp 8011\u20138023)."},{"key":"9527_CR40","unstructured":"James, S., Bloesch, M., & Davison, A. J. (2018). Task-embedded control networks for few-shot imitation learning. In Conference on robot learning (pp 783\u2013795)."},{"issue":"2","key":"9527_CR41","first-page":"86","volume":"3","author":"OO John","year":"2015","unstructured":"John, O. O. (2015). Robustness of quantile regression to outliers. American Journal of Applied Mathematics and Statistics, 3(2), 86\u201388.","journal-title":"American Journal of Applied Mathematics and Statistics"},{"key":"9527_CR42","unstructured":"Judah, K., Fern, A., & Dietterich, T.G. (2012). Active imitation learning via reduction to iid active learning. arXiv preprint arXiv:12104876."},{"issue":"120","key":"9527_CR43","first-page":"4105","volume":"15","author":"K Judah","year":"2014","unstructured":"Judah, K., Fern, A. P., Dietterich, T. G., & Tadepalli, P. (2014). Active imitation learning: Formal and practical reductions to i.i.d. learning. Journal of Machine Learning Research, 15(120), 4105\u20134143.","journal-title":"Journal of Machine Learning Research"},{"key":"9527_CR44","unstructured":"Kaelbling, L. P. (1993). Learning to achieve goals. In Proceedings of the international joint conferences on artificial intelligence (pp 1094\u20131098)."},{"key":"9527_CR45","unstructured":"Kang, B., Jie, Z., & Feng, J. (2018). Policy optimization with demonstrations. In International conference on machine learning (pp. 2469\u20132478)."},{"key":"9527_CR46","unstructured":"Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:14126980."},{"key":"9527_CR47","doi-asserted-by":"crossref","unstructured":"Klyubin, A. S., Polani, D., & Nehaniv, C. L. (2005). Empowerment: A universal agent-centric measure of control. Proceedings of the IEEE Congress on Evolutionary Computation, 1, 128\u2013135.","DOI":"10.1109\/CEC.2005.1554676"},{"issue":"1","key":"9527_CR48","first-page":"1334","volume":"17","author":"S Levine","year":"2016","unstructured":"Levine, S., Finn, C., Darrell, T., & Abbeel, P. (2016). End-to-end training of deep visuomotor policies. The Journal of Machine Learning Research, 17(1), 1334\u20131373.","journal-title":"The Journal of Machine Learning Research"},{"key":"9527_CR49","unstructured":"Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., & Wierstra, . (2015). Continuous control with deep reinforcement learning. arXiv preprint arXiv:150902971."},{"key":"9527_CR50","unstructured":"Machado, M. C., Bellemare, M. G., & Bowling, M. (2018). Count-based exploration with the successor representation. arXiv preprint arXiv:180711622"},{"key":"9527_CR51","doi-asserted-by":"crossref","unstructured":"Martin, J., Sasikumar, S. N., Everitt, T., & Hutter, M. (2017). Count-based exploration in feature space for reinforcement learning. In Proceedings of the international joint conference on artificial intelligence.","DOI":"10.24963\/ijcai.2017\/344"},{"key":"9527_CR52","unstructured":"Mathewson, K.W., Pilarski, P.M. (2017). Actor-critic reinforcement learning with simultaneous human control and feedback. arXiv preprint arXiv:170301274."},{"issue":"7540","key":"9527_CR53","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., & Georg. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529.","journal-title":"Nature"},{"key":"9527_CR54","doi-asserted-by":"crossref","unstructured":"Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., & Abbeel, P. (2018). Overcoming exploration in reinforcement learning with demonstrations. In Proceedings of the IEEE international conference on robotics and automation (pp. 6292\u20136299). IEEE.","DOI":"10.1109\/ICRA.2018.8463162"},{"key":"9527_CR55","unstructured":"Nair, A. V., Pong, V., Dalal, M., Bahl, S., Lin, S., & Levine, S. (2018b). Visual reinforcement learning with imagined goals. In Proceedings of the international conference on machine learning (pp. 9191\u20139200)."},{"key":"9527_CR56","unstructured":"Ng, A. Y., & Russell, S. J. (2000). Algorithms for inverse reinforcement learning. In Proceedings of the international conference on machine learning (pp 663\u2013670)."},{"key":"9527_CR57","unstructured":"Ostrovski, G., Bellemare, M. G., van den Oord, A., & Munos, R. (2017). Count-based exploration with neural density models. In Proceedings of the international conference on machine learning (pp. 2721\u20132730)"},{"issue":"2","key":"9527_CR58","doi-asserted-by":"publisher","first-page":"265","DOI":"10.1109\/TEVC.2006.890271","volume":"11","author":"PY Oudeyer","year":"2007","unstructured":"Oudeyer, P. Y., Kaplan, F., & Hafner, V. V. (2007). Intrinsic motivation systems for autonomous mental development. IEEE transactions on evolutionary computation, 11(2), 265\u2013286.","journal-title":"IEEE transactions on evolutionary computation"},{"key":"9527_CR59","doi-asserted-by":"crossref","unstructured":"Pathak, D., Agrawal, P., Efros, A. A., & Darrell, T. (2017). Curiosity-driven exploration by self-supervised prediction. In Proceedings of the international conference on international conference on machine learning (pp. 2778\u20132787).","DOI":"10.1109\/CVPRW.2017.70"},{"key":"9527_CR60","unstructured":"Pere, A., Forestier, S., Sigaud, O., & Oudeyer, P. Y. (2018). Unsupervised learning of goal spaces for intrinsically motivated goal exploration. In Proceedings of the international conference on learning representations."},{"key":"9527_CR61","unstructured":"Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., Kumar, V., & Zaremba, W. (2018). Multi-goal reinforcement learning: Challenging robotics environments and request for research. arXiv preprint arXiv:180209464."},{"key":"9527_CR62","unstructured":"Pomerleau, D. A. (1988). Alvinn: an autonomous land vehicle in a neural network. In Proceedings of the 1st international conference on neural information processing systems (pp 305\u2013313)."},{"key":"9527_CR63","unstructured":"Pong, V. H., Dalal, M., Lin, S., Nair, A., Bahl, S., & Levine, S. (2019). Skew-fit: State-covering self-supervised reinforcement learning. arXiv preprint:190303698."},{"key":"9527_CR64","unstructured":"Racaniere, S., Lampinen, A., Santoro, A., Reichert, D., Firoiu, V., & Lillicrap, T. (2020). Automated curriculum generation through setter-solver interactions. In International conference on learning representations"},{"key":"9527_CR65","unstructured":"Saunders, W., Sastry, G., Stuhlmueller, A., & Evans, O. (2018). Trial without error: Towards safe reinforcement learning via human intervention. In Proceedings of the international conference on autonomous agents and multiAgent systems (pp. 2067\u20132069)."},{"key":"9527_CR66","unstructured":"Savinov, N., Raichuk, A., Marinier, R., Vincent, D., Pollefeys, M., Lillicrap, T., & Gelly, S. (2019). Episodic curiosity through reachability. In Proceedings of the international conference on learning representations."},{"issue":"6","key":"9527_CR67","doi-asserted-by":"publisher","first-page":"233","DOI":"10.1016\/S1364-6613(99)01327-3","volume":"3","author":"S Schaal","year":"1999","unstructured":"Schaal, S. (1999). Is imitation learning the route to humanoid robots? Trends in cognitive sciences, 3(6), 233\u2013242.","journal-title":"Trends in cognitive sciences"},{"key":"9527_CR68","unstructured":"Schaul, T., Horgan, D., Gregor, K., & Silver, D. (2015). Universal value function approximators. In Proceedings of the international conference on machine learning (pp 1312\u20131320)."},{"key":"9527_CR69","unstructured":"Shon, A. P., Verma, D., & Rao, R. P. (2007). Active imitation learning. In Proceedings of the AAAI conference on artificial intelligence (pp. 756\u2013762)."},{"issue":"7587","key":"9527_CR70","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., & Hassabis, D. (2016). Mastering the game of go with deep neural networks and tree search. Nature, 529(7587), 484.","journal-title":"nature"},{"key":"9527_CR71","unstructured":"Stadie, B. C., Levine, .S, & Abbeel, P. (2015). Incentivizing exploration in reinforcement learning with deep predictive models. arXiv preprint arXiv:150700814."},{"issue":"8","key":"9527_CR72","doi-asserted-by":"publisher","first-page":"1309","DOI":"10.1016\/j.jcss.2007.08.009","volume":"74","author":"AL Strehl","year":"2008","unstructured":"Strehl, A. L., & Littman, M. L. (2008). An analysis of model-based interval estimation for markov decision processes. Journal of Computer and System Sciences, 74(8), 1309\u20131331.","journal-title":"Journal of Computer and System Sciences"},{"key":"9527_CR73","unstructured":"Sun, W., Venkatraman, A., Gordon, G.J., Boots, B., Bagnell, J.A. (2017). Deeply aggrevated: Differentiable imitation learning for sequential prediction. In Proceedings of the 34th international conference on machine learning (Vol. 70, pp. 3309\u20133318)."},{"key":"9527_CR74","first-page":"6417","volume":"32","author":"N Tagasovska","year":"2019","unstructured":"Tagasovska, N., & Lopez-Paz, D. (2019). Single-model uncertainties for deep learning. Advances in Neural Information Processing Systems, 32, 6417\u20136428.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"9527_CR75","unstructured":"Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, X., Duan, Y., Schulman, J., De Turck, F., & Abbeel, P. (2017). # exploration: Astudy of count-based exploration for deep reinforcement learning. In Proceedings of the 31st international conference on neural information processing systems (pp. 2750\u20132759)."},{"issue":"1","key":"9527_CR76","doi-asserted-by":"publisher","first-page":"45","DOI":"10.1080\/09540091.2014.885279","volume":"26","author":"ME Taylor","year":"2014","unstructured":"Taylor, M. E., Carboni, N., Fachantidis, A., Vlahavas, I., & Torrey, L. (2014). Reinforcement learning agents providing advice in complex video games. Connection Science, 26(1), 45\u201363.","journal-title":"Connection Science"},{"key":"9527_CR77","doi-asserted-by":"crossref","unstructured":"Todorov, E., Erez, T., & Tassa, Y. (2012). Mujoco: A physics engine for model-based control. In IEEE\/RSJ international conference on intelligent robots and systems (pp. 5026\u20135033).","DOI":"10.1109\/IROS.2012.6386109"},{"key":"9527_CR78","unstructured":"Vecerik, M., Hester, T., Scholz, J., Wang, F., Pietquin, O., Piot, B., Heess, N., Roth\u00f6rl, T., Lampe, T., & Riedmiller, M. (2017). Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards. arXiv preprint arXiv:170708817."},{"key":"9527_CR79","doi-asserted-by":"crossref","unstructured":"Warnell, G., Waytowich, N., Lawhern, V., & Stone, P. (2018). Deep tamer: Interactive agent shaping in high-dimensional state spaces. In Thirty-second AAAI conference on artificial intelligence (pp 1545\u20131554).","DOI":"10.1609\/aaai.v32i1.11485"},{"key":"9527_CR80","unstructured":"Wilson, A., Fern, A., Tadepalli, P. (2012). A bayesian approach for policy learning from trajectory preference queries. In Advances in neural information processing systems (pp. 1133\u20131141)."},{"issue":"1","key":"9527_CR81","first-page":"4945","volume":"18","author":"C Wirth","year":"2017","unstructured":"Wirth, C., Akrour, R., Neumann, G., & F\u00fcrnkranz, J. (2017). A survey of preference-based reinforcement learning methods. The Journal of Machine Learning Research, 18(1), 4945\u20134990.","journal-title":"The Journal of Machine Learning Research"},{"key":"9527_CR82","unstructured":"Yin, H., Seiler, P., Jin, M., & Arcak, M. (2020). Imitation learning with stability and safety guarantees. arXiv preprint arXiv:201209293."},{"key":"9527_CR83","unstructured":"Zaremba, W., & Sutskever, I. (2015). Learning to execute. 1410.4615."},{"key":"9527_CR84","unstructured":"Zhang, X., & Ma, H. (2018). Pretraining deep actor-critic reinforcement learning algorithms with expert demonstrations. arXiv preprint arXiv:180110459."},{"key":"9527_CR85","unstructured":"Zimmer, M., Viappiani, P., & Weng, P. (2014). Teacher-student framework: Areinforcement learning approach. In AAMAS workshop autonomous robots and multirobot systems."},{"key":"9527_CR86","doi-asserted-by":"crossref","unstructured":"Zuo, G., Zhao, Q., Lu, J., & Li, J. (2020). Efficient hindsight reinforcement learning using demonstrations for robotic tasks with sparse rewards. International Journal of Advanced Robotic Systems, 17.","DOI":"10.1177\/1729881419898342"}],"container-title":["Autonomous Agents and Multi-Agent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10458-021-09527-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10458-021-09527-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10458-021-09527-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,7]],"date-time":"2023-01-07T12:20:20Z","timestamp":1673094020000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10458-021-09527-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,16]]},"references-count":86,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,10]]}},"alternative-id":["9527"],"URL":"https:\/\/doi.org\/10.1007\/s10458-021-09527-5","relation":{},"ISSN":["1387-2532","1573-7454"],"issn-type":[{"type":"print","value":"1387-2532"},{"type":"electronic","value":"1573-7454"}],"subject":[],"published":{"date-parts":[[2021,8,16]]},"assertion":[{"value":"28 July 2021","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 August 2021","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"44"}}