{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T17:43:45Z","timestamp":1780508625595,"version":"3.54.1"},"reference-count":61,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2023,7,10]],"date-time":"2023-07-10T00:00:00Z","timestamp":1688947200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,7,10]],"date-time":"2023-07-10T00:00:00Z","timestamp":1688947200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100012338","name":"Alan Turing Institute","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100012338","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100000288","name":"Royal Society","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100000288","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100000266","name":"Engineering and Physical Sciences Research Council","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Auton Robot"],"published-print":{"date-parts":[[2023,8]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Informative path-planning is a well established approach to visual-servoing and active viewpoint selection in robotics, but typically assumes that a suitable cost function or goal state is known. This work considers the inverse problem, where the goal of the task is unknown, and a reward function needs to be inferred from exploratory example demonstrations provided by a demonstrator, for use in a downstream informative path-planning policy. Unfortunately, many existing reward inference strategies are unsuited to this class of problems, due to the exploratory nature of the demonstrations. In this paper, we propose an alternative approach to cope with the class of problems where these sub-optimal, exploratory demonstrations occur. We hypothesise that, in tasks which require discovery, successive states of any demonstration are progressively more likely to be associated with a higher reward, and use this hypothesis to generate time-based binary comparison outcomes and infer reward functions that support these ranks, under a probabilistic generative model. We formalise this <jats:italic>probabilistic temporal ranking<\/jats:italic> approach and show that it improves upon existing approaches to perform reward inference for autonomous ultrasound scanning, a novel application of learning from demonstration in medical imaging while also being of value across a broad range of goal-oriented learning from demonstration tasks.<\/jats:p>","DOI":"10.1007\/s10514-023-10120-w","type":"journal-article","created":{"date-parts":[[2023,7,10]],"date-time":"2023-07-10T18:02:42Z","timestamp":1689012162000},"page":"733-751","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["Learning rewards from exploratory demonstrations using probabilistic temporal ranking"],"prefix":"10.1007","volume":"47","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7426-1498","authenticated-orcid":false,"given":"Michael","family":"Burke","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Katie","family":"Lu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniel","family":"Angelov","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Art\u016bras","family":"Strai\u017eys","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Craig","family":"Innes","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kartic","family":"Subr","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Subramanian","family":"Ramamoorthy","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,7,10]]},"reference":[{"key":"10120_CR1","doi-asserted-by":"publisher","unstructured":"Abbeel, P., & Ng, AY. (2004). Apprenticeship learning via inverse reinforcement learning. In Proceedings of the twenty-first international conference on Machine learning, ACM, p\u00a01 https:\/\/doi.org\/10.1145\/1015330.1015430","DOI":"10.1145\/1015330.1015430"},{"issue":"1","key":"10120_CR2","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1109\/70.988970","volume":"18","author":"P Abolmaesumi","year":"2002","unstructured":"Abolmaesumi, P., Salcudean, S. E., Zhu, Wen-Hong., Sirouspour, M. R., & DiMaio, S. P. (2002). Image-guided control of a robot for medical ultrasound. IEEE Transactions on Robotics and Automation, 18(1), 11\u201323. https:\/\/doi.org\/10.1109\/70.988970","journal-title":"IEEE Transactions on Robotics and Automation"},{"key":"10120_CR3","doi-asserted-by":"crossref","unstructured":"Angelov, D., Hristov, Y., Burke, M., Ramamoorthy, S. (2020). Composing diverse policies for temporally extended tasks. Robotics and automation letters (RA-L) arXiv:1907.08199.","DOI":"10.1109\/LRA.2020.2972794"},{"key":"10120_CR4","unstructured":"Bagnell, JAD. (2015). An Invitation to Imitation. Tech. Rep. CMU-RI-TR-15-08, Carnegie Mellon University, Pittsburgh, PA https:\/\/www.ri.cmu.edu\/pub_files\/2015\/3\/InvitationToImitation_3_1415.pdf."},{"key":"10120_CR5","doi-asserted-by":"crossref","unstructured":"Barto, AG. (2013). Intrinsic motivation and reinforcement learning. In Intrinsically Motivated Learning in natural and Artificial Systems, Springer, pp. 17\u201347.","DOI":"10.1007\/978-3-642-32375-1_2"},{"key":"10120_CR6","doi-asserted-by":"publisher","unstructured":"Binney, J., & Sukhatme, GS. (2012). Branch and bound for informative path planning. In 2012 IEEE International Conference on Robotics and Automation, pp. 2147\u20132154 https:\/\/doi.org\/10.1109\/ICRA.2012.6224902.","DOI":"10.1109\/ICRA.2012.6224902"},{"key":"10120_CR7","doi-asserted-by":"publisher","unstructured":"Biyik, E., Huynh, N., Kochenderfer, M., & Sadigh, D. (2020). Active Preference-Based Gaussian Process Regression for Reward Learning. In Proceedings of Robotics: Science and Systems. https:\/\/doi.org\/10.15607\/RSS.2020.XVI.041.","DOI":"10.15607\/RSS.2020.XVI.041"},{"key":"10120_CR8","unstructured":"Boularias, A., Kober. J., & Peters, J. (2011). Relative entropy inverse reinforcement learning. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, pp. 182\u2013189 http:\/\/proceedings.mlr.press\/v15\/boularias11a\/boularias11a.pdf."},{"key":"10120_CR9","unstructured":"Braziunas, D., & Boutilier, C. (2006). Preference elicitation and generalized additive utility. AAAI 21(2) https:\/\/www.aaai.org\/Papers\/AAAI\/2006\/AAAI06-253.pdf."},{"key":"10120_CR10","doi-asserted-by":"publisher","unstructured":"Brochu, E., Brochu, T., de\u00a0Freitas, N. (2010). A Bayesian interactive optimization approach to procedural animation design. In Proceedings of the 2010 ACM SIGGRAPH\/Eurographics Symposium on Computer Animation, Eurographics Association, pp. 103\u2013112 https:\/\/doi.org\/10.5555\/1921427.1921443.","DOI":"10.5555\/1921427.1921443"},{"key":"10120_CR11","unstructured":"Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). Openai gym. arXiv:1606.01540."},{"key":"10120_CR12","unstructured":"Brown, D., Goo, W., Nagarajan, P., & Niekum, S. (2019). Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations. In International Conference on Machine Learning, pp. 783\u2013792 arXiv:1904.06387."},{"key":"10120_CR13","unstructured":"Brown, DS., Goo, W., & Niekum, S. (2020). Better-than-demonstrator imitation learning via automatically-ranked demonstrations. In Conference on robot learning, PMLR, pp. 330\u2013359."},{"key":"10120_CR14","unstructured":"Burke, M., Mbonambi, S., Molala, P., & Sefala, R. (2017). Rapid Probabilistic Interest Learning from Domain-Specific Pairwise Image Comparisons. arXiv preprint arXiv:1706.05850"},{"key":"10120_CR15","doi-asserted-by":"publisher","unstructured":"Calandra, R., Seyfarth, A., Peters, J., & Deisenroth, MP. (2014). An experimental comparison of Bayesian optimization for bipedal locomotion. In 2014 IEEE International Conference on Robotics and Automation (ICRA), pp. 1951\u20131958 https:\/\/doi.org\/10.1109\/ICRA.2014.6907117.","DOI":"10.1109\/ICRA.2014.6907117"},{"key":"10120_CR16","doi-asserted-by":"publisher","unstructured":"Chatelain, P., Krupa, A., & Marchal, M. (2013). Real-time needle detection and tracking using a visually servoed 3D ultrasound probe. In 2013 IEEE International Conference on Robotics and Automation, pp. 1676\u20131681 https:\/\/doi.org\/10.1109\/ICRA.2013.6630795.","DOI":"10.1109\/ICRA.2013.6630795"},{"key":"10120_CR17","doi-asserted-by":"crossref","unstructured":"Chatelain, P., Krupa, A., & Navab, N. (2015). Optimization of ultrasound image quality via visual servoing. In 2015 IEEE International Conference on Robotics and Automation (ICRA), IEEE, pp. 5997\u20136002.","DOI":"10.1109\/ICRA.2015.7140040"},{"key":"10120_CR18","doi-asserted-by":"crossref","unstructured":"Cho, DH., Ha, JS., Lee, S., Moon, S., & Choi, HL. (2018). Informative path planning and mapping with multiple UAVs in wind fields. In Distributed Autonomous Robotic Systems, Springer, pp. 269\u2013283. arXiv:1610.01303","DOI":"10.1007\/978-3-319-73008-0_19"},{"key":"10120_CR19","doi-asserted-by":"crossref","unstructured":"Chu, W., & Ghahramani, Z. (2005). Preference learning with gaussian processes. In Proceedings of the 22nd International Conference on Machine Learning, pp. 137\u2013144.","DOI":"10.1145\/1102351.1102369"},{"key":"10120_CR20","doi-asserted-by":"publisher","unstructured":"Deisenroth, M., & Rasmussen, CE. (2011). PILCO: A model-based and data-efficient approach to policy search. In Proceedings of the 28th International Conference on Machine Learning (ICML-11), pp.465\u2013472 https:\/\/doi.org\/10.5555\/3104482.3104541.","DOI":"10.5555\/3104482.3104541"},{"key":"10120_CR21","unstructured":"Finn, C., Christiano, P., Abbeel, P., & Levine, S. (2016). A connection between generative adversarial networks, inverse reinforcement learning, and energy-based MODELS. arXiv preprint arXiv:1611.03852."},{"key":"10120_CR22","unstructured":"Fu, J., Luo, K., & Levine, S. (2018). Learning robust rewards with adversarial inverse reinforcement learning. In International Conference on Learning Representations (ICLR). arXiv:1710.11248."},{"key":"10120_CR23","unstructured":"Ghasemipour, SKS., Zemel, R., & Gu, S. (2019). A Divergence Minimization Perspective on Imitation Learning Methods. In Conference on Robot Learning (CoRL). arXiv:1911.02256."},{"key":"10120_CR24","unstructured":"Gleave, A., Taufeeque, M., Rocamonde, J., Jenner, E., Wang, SH., Toyer, S., Ernestus, M., Belrose, N., Emmons, S., & Russell, S. (2022). imitation: Clean imitation learning implementations. arXiv:2211.11972v1 [cs.LG], https:\/\/arxiv.org\/abs\/2211.11972, 2211.11972."},{"key":"10120_CR25","doi-asserted-by":"crossref","unstructured":"Herbrich, R., Minka, T., & Graepel, T. (2007). TrueSkill$${^{TM}}$$: a Bayesian skill rating system. In Advances in neural information processing systems pp 569\u2013576 https:\/\/papers.nips.cc\/paper\/3079-trueskilltm-a-bayesian-skill-rating-system.pdf","DOI":"10.7551\/mitpress\/7503.003.0076"},{"key":"10120_CR26","unstructured":"Ho, J., & Ermon, S. (2016). Generative adversarial imitation learning. In Advances in neural information processing systems, pp. 4565\u20134573. http:\/\/papers.nips.cc\/paper\/6391-generative-adversarial-imitation-learning.pdf."},{"key":"10120_CR27","doi-asserted-by":"publisher","unstructured":"Kiapour, MH., Yamaguchi, K., Berg. AC., & Berg, TL. (2014). Hipster wars: Discovering elements of fashion styles. In European conference on computer vision, Springer, pp. 472\u2013488. https:\/\/doi.org\/10.1007\/978-3-319-10590-1_31.","DOI":"10.1007\/978-3-319-10590-1_31"},{"key":"10120_CR28","doi-asserted-by":"crossref","unstructured":"Konstantinova, J., Li, M., Althoefer, K., Nanayakkara, T., Dasgupta, P. (2013). Palpation strategies for artificial soft tissue examination. In 3rd Joint Workshop on New technologies for Computer\/Robot Assisted Surgery","DOI":"10.1109\/IROS.2013.6696622"},{"issue":"1","key":"10120_CR29","first-page":"430","volume":"18","author":"A Kucukelbir","year":"2017","unstructured":"Kucukelbir, A., Tran, D., Ranganath, R., Gelman, A., & Blei, D. M. (2017). Automatic differentiation variational inference. Journal of Machine Learning Research, 18(1), 430\u2013474.","journal-title":"Journal of Machine Learning Research"},{"key":"10120_CR30","doi-asserted-by":"crossref","unstructured":"Lee, K., Choi, S., & Oh, S. (2016). Inverse reinforcement learning with leveraged gaussian processes. In 2016 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, pp. 3907\u20133912.","DOI":"10.1109\/IROS.2016.7759575"},{"key":"10120_CR31","unstructured":"Levine, S., Popovic, Z., & Koltun, V. (2011). Nonlinear inverse reinforcement learning with Gaussian processes. In Advances in Neural Information Processing Systems, pp. 19\u201327 https:\/\/papers.nips.cc\/paper\/4420-nonlinear-inverse-reinforcement-learning-with-gaussian-processes.pdf."},{"key":"10120_CR32","doi-asserted-by":"publisher","unstructured":"Li, T., Kermorgant, O., Krupa, A. (2012). Maintaining visibility constraints during tele-echography with ultrasound visual servoing. In 2012 IEEE International Conference on Robotics and Automation, pp. 4856\u20134861 https:\/\/doi.org\/10.1109\/ICRA.2012.6224974.","DOI":"10.1109\/ICRA.2012.6224974"},{"issue":"1","key":"10120_CR33","doi-asserted-by":"publisher","first-page":"173","DOI":"10.1016\/j.ultrasmedbio.2009.08.014","volume":"36","author":"K Liang","year":"2010","unstructured":"Liang, K., Rogers, A. J., Light, E. D., von Allmen, D., & Smith, S. W. (2010). Three-dimensional ultrasound guidance of autonomous robotic breast biopsy: Feasibility study. Ultrasound in Medicine & Biology, 36(1), 173\u2013177. https:\/\/doi.org\/10.1016\/j.ultrasmedbio.2009.08.014","journal-title":"Ultrasound in Medicine & Biology"},{"key":"10120_CR34","doi-asserted-by":"crossref","unstructured":"Ling, CK., Low, KH., & Jaillet, P. (2016). Gaussian process planning with Lipschitz continuous reward functions: Towards unifying Bayesian optimization, active learning, and beyond. In Thirtieth AAAI Conference on Artificial Intelligence. arXiv:1511.06890.","DOI":"10.1609\/aaai.v30i1.10210"},{"key":"10120_CR35","doi-asserted-by":"crossref","unstructured":"Lopes, M., Melo, F., & Montesano, L. (2009). Active learning for reward estimation in inverse reinforcement learning. In W. Buntine, M. Grobelnik, D. Mladeni\u0107, & J. Shawe-Taylor (Eds.), Machine Learning and Knowledge Discovery in Databases pp. 31\u201346. Springer.","DOI":"10.1007\/978-3-642-04174-7_3"},{"key":"10120_CR36","doi-asserted-by":"crossref","unstructured":"Majumdar, A., Singh, S., Mandlekar, A., & Pavone, M. (2017). Risk-sensitive Inverse Reinforcement Learning via Coherent Risk Models. In Proceedings of Robotics: Science and Systems, Cambridge, Massachusetts. http:\/\/www.roboticsproceedings.org\/rss13\/p69.pdf.","DOI":"10.15607\/RSS.2017.XIII.069"},{"key":"10120_CR37","doi-asserted-by":"publisher","unstructured":"Marchant, R., & Ramos, F. (2014). Bayesian Optimisation for informative continuous path planning. In 2014 IEEE International Conference on Robotics and Automation (ICRA), pp. 6136\u20136143. https:\/\/doi.org\/10.1109\/ICRA.2014.6907763.","DOI":"10.1109\/ICRA.2014.6907763"},{"key":"10120_CR38","doi-asserted-by":"publisher","unstructured":"Martinez-Cantin, R. (2017). Bayesian optimization with adaptive kernels for robot control. In 2017 IEEE International Conference on Robotics and Automation (ICRA), pp. 3350\u20133356. https:\/\/doi.org\/10.1109\/ICRA.2017.7989380.","DOI":"10.1109\/ICRA.2017.7989380"},{"issue":"2","key":"10120_CR39","doi-asserted-by":"publisher","first-page":"93","DOI":"10.1007\/s10514-009-9130-2","volume":"27","author":"R Martinez-Cantin","year":"2009","unstructured":"Martinez-Cantin, R., de Freitas, N., Brochu, E., Castellanos, J., & Doucet, A. (2009). A Bayesian exploration-exploitation approach for optimal online sensing and planning with a visually guided mobile robot. Autonomous Robots, 27(2), 93\u2013103. https:\/\/doi.org\/10.1007\/s10514-009-9130-2","journal-title":"Autonomous Robots"},{"key":"10120_CR40","unstructured":"Mnih, V., Badia, AP., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., & Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning, pp. 1928\u20131937."},{"key":"10120_CR41","doi-asserted-by":"publisher","first-page":"34851","DOI":"10.1038\/srep34851","volume":"6","author":"C Murawski","year":"2016","unstructured":"Murawski, C., & Bossaerts, P. (2016). How humans solve complex problems: The case of the knapsack problem. Scientific Reports, 6, 34851.","journal-title":"Scientific Reports"},{"key":"10120_CR42","doi-asserted-by":"publisher","unstructured":"Naik, N., Philipoom, J., Raskar, R., & Hidalgo, C. (2014). Streetscore \u2013 predicting the perceived safety of one million streetscapes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 779\u2013785. https:\/\/doi.org\/10.1109\/CVPRW.2014.121.","DOI":"10.1109\/CVPRW.2014.121"},{"key":"10120_CR43","doi-asserted-by":"crossref","unstructured":"Neal, RM. (1996). Priors for infinite networks. In Bayesian Learning for Neural Networks, Springer. pp. 29\u201353. https:\/\/www.cs.toronto.edu\/~radford\/ftp\/pin.pdf.","DOI":"10.1007\/978-1-4612-0745-0_2"},{"key":"10120_CR44","unstructured":"Ng, AY., & Russell, SJ. (2000). Algorithms for inverse reinforcement learning. In ICML, vol.\u00a01, p.\u00a02."},{"issue":"268","key":"10120_CR45","first-page":"1","volume":"22","author":"A Raffin","year":"2021","unstructured":"Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., & Dormann, N. (2021). Stable-baselines3: Reliable reinforcement learning implementations. Journal of Machine Learning Research, 22(268), 1\u20138.","journal-title":"Journal of Machine Learning Research"},{"key":"10120_CR46","doi-asserted-by":"publisher","DOI":"10.7717\/peerj-cs.55","volume":"2","author":"J Salvatier","year":"2016","unstructured":"Salvatier, J., Wiecki, T. V., & Fonnesbeck, C. (2016). Probabilistic programming in Python using PyMC3. PeerJ Computer Science, 2, e55. https:\/\/doi.org\/10.7717\/peerj-cs.55","journal-title":"PeerJ Computer Science"},{"key":"10120_CR47","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347."},{"key":"10120_CR48","doi-asserted-by":"crossref","unstructured":"Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal S., & Levine, S. (2018). Time-contrastive networks: Self-supervised learning from video. In 2018 IEEE International Conference on Robotics and Automation (ICRA), pp. 1134\u20131141. arXiv:1704.06888.","DOI":"10.1109\/ICRA.2018.8462891"},{"key":"10120_CR49","unstructured":"Shiarlis, K., ao Messias, J., & Whiteson, S. (2016). Inverse reinforcement learning from failure. In AAMAS 2016: Proceedings of the Fifteenth International Joint Conference on Autonomous Agents and Multi-Agent Systems, pp. 1060\u20131068."},{"issue":"11","key":"10120_CR50","doi-asserted-by":"publisher","first-page":"704","DOI":"10.1016\/j.artint.2010.04.022","volume":"174","author":"M Sridharan","year":"2010","unstructured":"Sridharan, M., Wyatt, J., & Dearden, R. (2010). Planning to see: A hierarchical approach to planning visual actions on a robot using pomdps. Artificial Intelligence, 174(11), 704\u2013725.","journal-title":"Artificial Intelligence"},{"key":"10120_CR51","doi-asserted-by":"crossref","unstructured":"Sugiyama, H., Meguro, T., & Minami, Y. (2012). Preference-learning based inverse reinforcement learning for dialog control. In Thirteenth Annual Conference of the International Speech Communication Association. https:\/\/www.isca-speech.org\/archive\/archive_papers\/interspeech_2012\/i12_0222.pdf.","DOI":"10.21437\/Interspeech.2012-72"},{"key":"10120_CR52","unstructured":"Sui, Y., Zhuang, V., Burdick, JW., & Yue, Y. (2017). Multi-dueling bandits with dependent arms. arXiv preprint arXiv:1705.00253."},{"key":"10120_CR53","doi-asserted-by":"crossref","unstructured":"Thurstone, LL. (2017). A law of comparative judgment. In Scaling, Routledge. pp. 81\u201392.","DOI":"10.4324\/9781315128948-7"},{"key":"10120_CR54","doi-asserted-by":"crossref","unstructured":"Tucker, M., Novoseller, E., Kann, C., Sui, Y., Yue, Y., Burdick, JW., & Ames, AD. (2020). Preference-based learning for exoskeleton gait optimization. In 2020 IEEE International Conference on Robotics and Automation (ICRA), IEEE, pp. 2351\u20132357.","DOI":"10.1109\/ICRA40945.2020.9196661"},{"key":"10120_CR55","unstructured":"Valko, M., Ghavamzadeh, M., & Lazaric, A. (2013). Semi-supervised apprenticeship learning. In European Workshop on Reinforcement Learning, pp. 131\u2013142."},{"key":"10120_CR56","doi-asserted-by":"crossref","unstructured":"Williams, CK., & Rasmussen, CE. (2006). Gaussian processes for machine learning, vol\u00a02. MIT press Cambridge, MA http:\/\/www.gaussianprocess.org\/gpml\/chapters\/RW.pdf.","DOI":"10.7551\/mitpress\/3206.001.0001"},{"issue":"1","key":"10120_CR57","first-page":"4945","volume":"18","author":"C Wirth","year":"2017","unstructured":"Wirth, C., Akrour, R., Neumann, G., & F\u00fcrnkranz, J. (2017). A survey of preference-based reinforcement learning methods. The Journal of Machine Learning Research, 18(1), 4945\u20134990.","journal-title":"The Journal of Machine Learning Research"},{"key":"10120_CR58","unstructured":"Wu, YH., Charoenphakdee, N., Bao, H., Tangkaratt, V., Sugiyama, M. (2019). Imitation learning from imperfect demonstration. In International Conference on Machine Learning, pp. 6818\u20136827."},{"key":"10120_CR59","unstructured":"Wulfmeier, M., Ondruska, P., Posner, I. (2015). Maximum entropy deep inverse reinforcement learning. arXiv preprint arXiv:1507.04888."},{"key":"10120_CR60","doi-asserted-by":"publisher","unstructured":"Yi, Z., Calandra, R., Veiga, F., van Hoof, H., Hermans, T., Zhang, Y., & Peters, J. (2016). Active tactile object exploration with Gaussian processes. In 2016 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 4925\u20134930. https:\/\/doi.org\/10.1109\/IROS.2016.7759723.","DOI":"10.1109\/IROS.2016.7759723"},{"key":"10120_CR61","unstructured":"Ziebart, BD., Maas, A., Bagnell, JA., & Dey, AK. (2008). Maximum entropy inverse reinforcement learning. In Proceedings of AAAI\u201908 Proceedings of the 23rd national conference on Artifical intelligence, vol\u00a03, pp. 1433\u20131438. https:\/\/www.aaai.org\/Papers\/AAAI\/2008\/AAAI08-227.pdf."}],"container-title":["Autonomous Robots"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10514-023-10120-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10514-023-10120-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10514-023-10120-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,7,28]],"date-time":"2023-07-28T02:27:53Z","timestamp":1690511273000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10514-023-10120-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,10]]},"references-count":61,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,8]]}},"alternative-id":["10120"],"URL":"https:\/\/doi.org\/10.1007\/s10514-023-10120-w","relation":{},"ISSN":["0929-5593","1573-7527"],"issn-type":[{"value":"0929-5593","type":"print"},{"value":"1573-7527","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,10]]},"assertion":[{"value":"22 February 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 June 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 July 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"S. Ramamoorthy is vice president at five.ai, an autonomous driving company whose focus lies outside the domain of this paper. The remaining authors confirm that no other conflict of interest exists.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of interest"}}]}}