{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T15:57:56Z","timestamp":1777910276410,"version":"3.51.4"},"reference-count":27,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2016,9,5]],"date-time":"2016-09-05T00:00:00Z","timestamp":1473033600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"DOI":"10.13039\/501100004731","name":"Zhejiang Provincial Natural Science Foundation","doi-asserted-by":"crossref","award":["LY15F030005"],"award-info":[{"award-number":["LY15F030005"]}],"id":[{"id":"10.13039\/501100004731","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61004066"],"award-info":[{"award-number":["61004066"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Transactions of the Institute of Measurement and Control"],"published-print":{"date-parts":[[2018,1]]},"abstract":"<jats:p>Nowadays, transfer learning (TL) has become a crucial technique to accelerate the slow optimization procedure of reinforcement learning (RL) by re-utilizing knowledge acquired in a previous related task. Nevertheless, most of the current relevant research acquires knowledge through RL training in the source task, which would be too time-consuming. In view of this situation, in this paper, we propose a novel TL framework where the agent extracts knowledge from human-demonstration trajectories of the source task and reuses the knowledge in RL in the target task. As for what to transfer, two forms of knowledge deduced from the demonstration trajectories, which are the k-nearest neighbour of the current state in source samples and visit frequency of homologous states, are adopted. For how to transfer, the two forms of knowledge are respectively used to recommend a preferred action when random exploration is needed and to shape an instantaneous reward for RL. Simulation experiments of balancing Cart-Poles with different difficulties suggest that both the two forms of knowledge accelerate the learning process of RL obviously. What is more, the effect is even more significant when they are used in combination. In this case, the experimental results manifest the positive role of our framework in RL.<\/jats:p>","DOI":"10.1177\/0142331216649655","type":"journal-article","created":{"date-parts":[[2016,9,5]],"date-time":"2016-09-05T20:48:34Z","timestamp":1473108514000},"page":"94-101","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":4,"title":["Transferring knowledge from human-demonstration trajectories to reinforcement learning"],"prefix":"10.1177","volume":"40","author":[{"given":"Guo-fang","family":"Wang","sequence":"first","affiliation":[{"name":"School of Aeronautics and Astronautics, Zhejiang University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhou","family":"Fang","sequence":"additional","affiliation":[{"name":"School of Aeronautics and Astronautics, Zhejiang University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ping","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Control Science and Engineering, Zhejiang University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bo","family":"Li","sequence":"additional","affiliation":[{"name":"School of Aeronautics and Astronautics, Zhejiang University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2016,9,5]]},"reference":[{"key":"bibr1-0142331216649655","volume-title":"Apprenticeship Learning and Reinforcement Learning with Application to Robotic Control","author":"Abbeel P","year":"2008"},{"key":"bibr2-0142331216649655","first-page":"554","volume-title":"Encyclopedia of Machine Learning","author":"Abbeel P","year":"2010"},{"key":"bibr3-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2008.10.024"},{"key":"bibr4-0142331216649655","first-page":"3352","volume-title":"Proceedings of the 24th international joint conference on artificial intelligence (IJCAI)","author":"Brys T","year":"2015"},{"key":"bibr5-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1201\/9781439821091"},{"key":"bibr6-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2012.08.039"},{"key":"bibr7-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1145\/1160633.1160762"},{"key":"bibr8-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1109\/ChiCC.2015.7260106"},{"key":"bibr9-0142331216649655","volume-title":"Machine Learning in Action","author":"Harrington P","year":"2012"},{"key":"bibr10-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913495721"},{"key":"bibr11-0142331216649655","first-page":"1333","volume":"13","author":"Konidaris G","year":"2012","journal-title":"Journal of Machine Learning Research"},{"key":"bibr12-0142331216649655","first-page":"1107","volume":"4","author":"Lagoudakis MG","year":"2003","journal-title":"The Journal of Machine Learning Research"},{"key":"bibr13-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390225"},{"key":"bibr14-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1177\/0142331213509828"},{"key":"bibr15-0142331216649655","first-page":"361","volume-title":"Proceedings of the 18th international conference on machine learning (ICML)","author":"McGovern A","year":"2001"},{"key":"bibr16-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1007\/11552246_35"},{"key":"bibr17-0142331216649655","first-page":"278","volume-title":"Proceedings of the 16th international conference on machine learning (ICML)","volume":"99","author":"Ng AY","year":"1999"},{"key":"bibr18-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2009.191"},{"key":"bibr19-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2007.11.026"},{"key":"bibr20-0142331216649655","volume-title":"Science and Human Behavior","author":"Skinner BF","year":"1953"},{"key":"bibr21-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.1998.712192"},{"key":"bibr22-0142331216649655","first-page":"1633","volume":"10","author":"Taylor ME","year":"2009","journal-title":"Journal of Machine Learning Research"},{"key":"bibr23-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1609\/aimag.v32i1.2329"},{"key":"bibr24-0142331216649655","first-page":"617","volume-title":"Proceedings of the 10th international conference on autonomous agents and multiagent systems (AAMAS)","volume":"2","author":"Taylor ME","year":"2011"},{"key":"bibr25-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1007\/11871842_41"},{"key":"bibr26-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2012.11.007"},{"key":"bibr27-0142331216649655","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2013.08.037"}],"container-title":["Transactions of the Institute of Measurement and Control"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/0142331216649655","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/0142331216649655","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/0142331216649655","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T14:53:11Z","timestamp":1777647191000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/0142331216649655"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,9,5]]},"references-count":27,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2018,1]]}},"alternative-id":["10.1177\/0142331216649655"],"URL":"https:\/\/doi.org\/10.1177\/0142331216649655","relation":{},"ISSN":["0142-3312","1477-0369"],"issn-type":[{"value":"0142-3312","type":"print"},{"value":"1477-0369","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,9,5]]}}}