{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:04:06Z","timestamp":1720742646415},"reference-count":30,"publisher":"Walter de Gruyter GmbH","issue":"1","license":[{"start":{"date-parts":[[2018,8,1]],"date-time":"2018-08-01T00:00:00Z","timestamp":1533081600000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2018,8,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Using assistive robots for educational applications requires robots to be able to adapt their behavior specifically for each child with whom they interact.Among relevant signals, non-verbal cues such as the child\u2019s gaze can provide the robot with important information about the child\u2019s current engagement in the task, and whether the robot should continue its current behavior or not. Here we propose a reinforcement learning algorithm extended with active state-specific exploration and show its applicability to child engagement maximization as well as more classical tasks such as maze navigation. We first demonstrate its adaptive nature on a continuous maze problem as an enhancement of the classic grid world. There, parameterized actions enable the agent to learn single moves until the end of a corridor, similarly to \u201coptions\u201d but without explicit hierarchical representations.We then apply the algorithm to a series of simulated scenarios, such as an extended Tower of Hanoi where the robot should find the appropriate speed of movement for the interacting child, and to a pointing task where the robot should find the child-specific appropriate level of expressivity of action. We show that the algorithm enables to cope with both global and local non-stationarities in the state space while preserving a stable behavior in other stationary portions of the state space. Altogether, these results suggest a promising way to enable robot learning based on non-verbal cues and the high degree of non-stationarities that can occur during interaction with children.<\/jats:p>","DOI":"10.1515\/pjbr-2018-0016","type":"journal-article","created":{"date-parts":[[2018,9,13]],"date-time":"2018-09-13T09:02:09Z","timestamp":1536829329000},"page":"235-253","source":"Crossref","is-referenced-by-count":6,"title":["Adaptive reinforcement learning with active state-specific exploration for engagement maximization during simulated child-robot interaction"],"prefix":"10.1515","volume":"9","author":[{"given":"George","family":"Velentzas","sequence":"first","affiliation":[{"name":"School of Electrical and Computer Engineering, National Technical University of Athens, Athens , Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Theodore","family":"Tsitsimis","sequence":"additional","affiliation":[{"name":"School of Electrical and Computer Engineering, National Technical University of Athens, Athens , Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"I\u00f1aki","family":"Ra\u00f1\u00f3","sequence":"additional","affiliation":[{"name":"Intelligent Systems Research Centre, Ulster University, Coleraine , UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Costas","family":"Tzafestas","sequence":"additional","affiliation":[{"name":"School of Electrical and Computer Engineering, National Technical University of Athens, Athens , Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mehdi","family":"Khamassi","sequence":"additional","affiliation":[{"name":"Sorbonne Universit\u00e9,CNRS, Institute of Intelligent Systems and Robotics, Paris , France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"374","published-online":{"date-parts":[[2018,9,1]]},"reference":[{"key":"2022042712092644022_j_pjbr-2018-0016_ref_001_w2aab3b7c16b1b6b1ab1ab1Aa","doi-asserted-by":"crossref","unstructured":"[1] T. Fong, I. Nourbakhsh, K. Dautenhahn, A survey of socially interactive robots, Robotics and Autonomous Systems, 2003, 42, 143-16610.1016\/S0921-8890(02)00372-X","DOI":"10.1016\/S0921-8890(02)00372-X"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_002_w2aab3b7c16b1b6b1ab1ab2Aa","doi-asserted-by":"crossref","unstructured":"[2] T. Kanda, T. Hirano, D. Eaton, H. Ishiguro, Interactive robots as social partners and peer tutors for children: A field trial, Human- Computer Interaction, 2004, 19(1), 61-8410.1207\/s15327051hci1901&2_4","DOI":"10.1207\/s15327051hci1901&2_4"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_003_w2aab3b7c16b1b6b1ab1ab3Aa","doi-asserted-by":"crossref","unstructured":"[3] B. Robins, K. Dautenhahn, R. Te Boekhorst, A. Billard, Robotic assistants in therapy and education of children with autism: Can a small humanoid robot help encourage social interaction skills? Universal Access in the Information Society, 2005, 4(2), 105-12010.1007\/s10209-005-0116-3","DOI":"10.1007\/s10209-005-0116-3"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_004_w2aab3b7c16b1b6b1ab1ab4Aa","doi-asserted-by":"crossref","unstructured":"[4] T. Belpaeme, P. E. Baxter, R. Read, R. Wood, H. Cuay\u00e1huitl, B. Kiefer, et al.,Multimodal child-robot interaction: Building social bonds, Journal of Human-Robot Interaction, 2012, 1(2), 33-5310.5898\/JHRI.1.2.Belpaeme","DOI":"10.5898\/JHRI.1.2.Belpaeme"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_005_w2aab3b7c16b1b6b1ab1ab5Aa","doi-asserted-by":"crossref","unstructured":"[5] K.-Y. Chin, Z.-W. Hong, Y.-L. Chen, Impact of using an educational robot-based learning system on students motivation in elementary education, IEEE Transactions on Learning Technologies, 2014, 7(4), 333-34510.1109\/TLT.2014.2346756","DOI":"10.1109\/TLT.2014.2346756"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_006_w2aab3b7c16b1b6b1ab1ab6Aa","doi-asserted-by":"crossref","unstructured":"[6] C. Rich, B. Ponsler, A. Holroyd, C. L. Sidner, Recognizing engagement in human-robot interaction, In: 5th ACM\/IEEE International Conference on Human-Robot Interaction (HRI), IEEE, 2010, 375- 38210.1109\/HRI.2010.5453163","DOI":"10.1109\/HRI.2010.5453163"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_007_w2aab3b7c16b1b6b1ab1ab7Aa","doi-asserted-by":"crossref","unstructured":"[7] S. Ivaldi, S. Lefort, J. Peters, M. Chetouani, J. Provasi, E. Zibetti, Towards engagement models that consider individual factors in HRI: on the relation of extroversion and negative attitude towards robots to gaze and speech during a human-robot assembly task, International Journal of Social Robotics, 2017, 9(1), 63-8610.1007\/s12369-016-0357-8","DOI":"10.1007\/s12369-016-0357-8"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_008_w2aab3b7c16b1b6b1ab1ab8Aa","doi-asserted-by":"crossref","unstructured":"[8] S. Lemaignan, M.Warnier, E.A. Sisbot, A. Clodic, R. Alami, Artificial cognition for social human-robot interaction: An implementation, Artificial Intelligence, 2017, 247, 45-6910.1016\/j.artint.2016.07.002","DOI":"10.1016\/j.artint.2016.07.002"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_009_w2aab3b7c16b1b6b1ab1ab9Aa","doi-asserted-by":"crossref","unstructured":"[9] C. L. Sidner, C. Lee, C. D. Kidd, N. Lesh, C. Rich, Explorations in engagement for humans and robots, Artificial Intelligence, 2005, 166(1-2), 140-16410.1016\/j.artint.2005.03.005","DOI":"10.1016\/j.artint.2005.03.005"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_010_w2aab3b7c16b1b6b1ab1ac10Aa","doi-asserted-by":"crossref","unstructured":"[10] S. M. Anzalone, S. Boucenna, S. Ivaldi, M. Chetouani, Evaluating the engagement with social robots, International Journal of Social Robotics, 2015, 7(4), 465-47810.1007\/s12369-015-0298-7","DOI":"10.1007\/s12369-015-0298-7"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_011_w2aab3b7c16b1b6b1ab1ac11Aa","doi-asserted-by":"crossref","unstructured":"[11] M. Khamassi, S. Lall\u00e9e, P. Enel, E. Procyk, P. F. Dominey, Robot cognitive control with a neurophysiologically inspired reinforcement learning model, Frontiers in Neurorobotics, 2011, 5, 110.3389\/fnbot.2011.00001","DOI":"10.3389\/fnbot.2011.00001"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_012_w2aab3b7c16b1b6b1ab1ac12Aa","unstructured":"[12] J. Kober, J. Peters, Policy search for motor primitives in robotics, Machine Learning, 2011, 84, 171-20310.1007\/s10994-010-5223-6"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_013_w2aab3b7c16b1b6b1ab1ac13Aa","doi-asserted-by":"crossref","unstructured":"[13] F. Stulp, O. Sigaud, Robot skill learning: From reinforcement learning to evolution strategies, Paladyn Journal of Behavioral Robotics, 2013, 4(1), 49-6110.2478\/pjbr-2013-0003","DOI":"10.2478\/pjbr-2013-0003"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_014_w2aab3b7c16b1b6b1ab1ac14Aa","doi-asserted-by":"crossref","unstructured":"[14] J. Kober, J. A. Bagnell, J. Peters, Reinforcement learning in robotics: A survey, The International Journal of Robotics Research, 2013, 32(11), 1238-127410.1177\/0278364913495721","DOI":"10.1177\/0278364913495721"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_015_w2aab3b7c16b1b6b1ab1ac15Aa","doi-asserted-by":"crossref","unstructured":"[15] M. Khamassi, G. Velentzas, T. Tsitsimis, C. Tzafestas, Active exploration and parameterized reinforcement learning applied to a simulated human-robot interaction task, In: 2017 First IEEE International Conference on Robotic Computing (IRC), Taichung, Taiwan, 2017, 28-3510.1109\/IRC.2017.33","DOI":"10.1109\/IRC.2017.33"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_016_w2aab3b7c16b1b6b1ab1ac16Aa","doi-asserted-by":"crossref","unstructured":"[16] M. Khamassi, G. Velentzas, T. Tsitsimis, C. Tzafestas, Robot fast adaptation to changes in human engagement during simulated dynamic social interaction with active exploration in parameterized reinforcement learning, IEEE Transactions on Cognitive and Developmental Systems, 2018 (in press)10.1109\/TCDS.2018.2843122","DOI":"10.1109\/TCDS.2018.2843122"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_017_w2aab3b7c16b1b6b1ab1ac17Aa","doi-asserted-by":"crossref","unstructured":"[17] W. Masson, P. Ranchod, G. Konidaris, Reinforcement learning with parameterized actions, In: Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (AAAI-16), 2016","DOI":"10.1609\/aaai.v30i1.10226"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_018_w2aab3b7c16b1b6b1ab1ac18Aa","unstructured":"[18] M. Hausknecht, P. Stone, Deep reinforcement learning in parameterized action space, In: International Conference on Learning Representations (ICLR 2016), 2016"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_019_w2aab3b7c16b1b6b1ab1ac19Aa","doi-asserted-by":"crossref","unstructured":"[19] J. Schmidhuber, Developmental robotics, optimal artificial curiosity, creativity, music, and the fine arts, Connection Science, 2006, 18(2), 173-18710.1080\/09540090600768658","DOI":"10.1080\/09540090600768658"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_020_w2aab3b7c16b1b6b1ab1ac20Aa","doi-asserted-by":"crossref","unstructured":"[20] A. Baranes, P.-Y. Oudeyer, Active learning of inverse models with intrinsically motivated goal exploration in robots, Robotics and Autonomous Systems, 2013, 61(1), 49-7310.1016\/j.robot.2012.05.008","DOI":"10.1016\/j.robot.2012.05.008"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_021_w2aab3b7c16b1b6b1ab1ac21Aa","doi-asserted-by":"crossref","unstructured":"[21] C. Moulin-Frier, P.-Y. Oudeyer, Exploration strategies in developmental robotics: a unified probabilistic framework, In: 2013 IEEE Third Joint International Conference on Development and Learning and Epigenetic Robotics (ICDL), IEEE, 2013, 1-610.1109\/DevLrn.2013.6652535","DOI":"10.1109\/DevLrn.2013.6652535"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_022_w2aab3b7c16b1b6b1ab1ac22Aa","doi-asserted-by":"crossref","unstructured":"[22] F. C. Y. Benureau, P.-Y. Oudeyer, Behavioral diversity generation in autonomous exploration through reuse of past experience, Frontiers in Robotics and AI, 2016, 310.3389\/frobt.2016.00008","DOI":"10.3389\/frobt.2016.00008"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_023_w2aab3b7c16b1b6b1ab1ac23Aa","unstructured":"[23] J. X. Wang, Z. Kurth-Nelson, D. Tirumala, H. Soyer, J. Z. Leibo, R. Munos, et al., Learning to reinforcement learn, 2016, arXiv:1611.05763"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_024_w2aab3b7c16b1b6b1ab1ac24Aa","doi-asserted-by":"crossref","unstructured":"[24] N. Schweighofer, K. Doya, Meta-learning in reinforcement learning, Neural Networks, 2003, 16(1), 5-910.1016\/S0893-6080(02)00228-9","DOI":"10.1016\/S0893-6080(02)00228-9"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_025_w2aab3b7c16b1b6b1ab1ac25Aa","doi-asserted-by":"crossref","unstructured":"[25] K. Doya, Metalearning and neuromodulation, Neural Networks, 2002, 15(4-6), 495-50610.1016\/S0893-6080(02)00044-8","DOI":"10.1016\/S0893-6080(02)00044-8"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_026_w2aab3b7c16b1b6b1ab1ac26Aa","doi-asserted-by":"crossref","unstructured":"[26] G. Velentzas, C. Tzafestas, M. Khamassi, Bio-inspired meta learning for active exploration during non-stationary multiarmed bandit tasks, In: IEEE Intelligent Systems Conference 2017, London, UK, 201710.1109\/IntelliSys.2017.8324365","DOI":"10.1109\/IntelliSys.2017.8324365"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_027_w2aab3b7c16b1b6b1ab1ac27Aa","unstructured":"[27] A. Garivier, E.Moulines, On upper-confidence bound policies for non-stationary bandit problems, 2008, arXiv:0805.3415"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_028_w2aab3b7c16b1b6b1ab1ac28Aa","doi-asserted-by":"crossref","unstructured":"[28] H. van Hasselt, M. Wiering, Reinforcement learning in continuous action spaces, In: IEEE Symposium on Approximate Dynamic Programming and Reinforcement Learning, 2007, 272-27910.1109\/ADPRL.2007.368199","DOI":"10.1109\/ADPRL.2007.368199"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_029_w2aab3b7c16b1b6b1ab1ac29Aa","unstructured":"[29] R. S. Sutton, A. G. Barto, Reinforcement Learning: An Introduction, Cambridge, MA: MIT Press, 199810.1109\/TNN.1998.712192"},{"key":"2022042712092644022_j_pjbr-2018-0016_ref_030_w2aab3b7c16b1b6b1ab1ac30Aa","doi-asserted-by":"crossref","unstructured":"[30] L. Schilbach, M. Wilms, S. B. Eickhoff, S. Romanzetti, R. Tepest, G. Bente, N. J. Shah, G. R. Fink, K. Vogeley, Minds made for sharing: Initiating joint attention recruits reward-related neurocircuitry, Journal of Cognitive Neuroscience, 2010, 22(12), 2702- 2715.10.1162\/jocn.2009.21401","DOI":"10.1162\/jocn.2009.21401"}],"container-title":["Paladyn, Journal of Behavioral Robotics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/www.degruyter.com\/view\/j\/pjbr.2018.9.issue-1\/pjbr-2018-0016\/pjbr-2018-0016.xml","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyter.com\/document\/doi\/10.1515\/pjbr-2018-0016\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyter.com\/document\/doi\/10.1515\/pjbr-2018-0016\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,9,1]],"date-time":"2022-09-01T07:23:16Z","timestamp":1662016996000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.degruyter.com\/document\/doi\/10.1515\/pjbr-2018-0016\/html"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,8,1]]},"references-count":30,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2018,7,25]]},"published-print":{"date-parts":[[2018,7,1]]}},"alternative-id":["10.1515\/pjbr-2018-0016"],"URL":"https:\/\/doi.org\/10.1515\/pjbr-2018-0016","relation":{},"ISSN":["2081-4836"],"issn-type":[{"value":"2081-4836","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,8,1]]}}}