{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T06:26:31Z","timestamp":1777703191567,"version":"3.51.4"},"reference-count":27,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2018,1,12]],"date-time":"2018-01-12T00:00:00Z","timestamp":1515715200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Journal of Intelligent &amp; Fuzzy Systems: Applications in Engineering and Technology"],"published-print":{"date-parts":[[2018,1,12]]},"abstract":"<jats:p>\n                    Reusing knowledge obtained in other related but different tasks to accelerate the learning procedure of reinforcement learning (RL) has attracted more and more attention and expert knowledge transfer is the root cause of positive effect. Nevertheless, compared with acquiring knowledge by RL training in source tasks, this paper proposes to transfer knowledge contained in human-demonstrations of source tasks. Based on this, three specific forms of knowledge in total are mined from demonstration trajectories to be reused in the target task to shape RL and all of them are closely associated with the similarity between states of different tasks which can be measured by Euclidean distance via human-supplied inter-task mappings. In more detail, the similarity between the target state and the most similar state in source samples, the proportion of different actions among the\n                    <jats:italic toggle=\"yes\">k<\/jats:italic>\n                    -NN of the target state in source samples and the proportion of different actions under a constant similarity with the target state in source samples are respectively selected to initialize the value of state-action function. Simulation experiments of mountain car problems with different difficulties and different dimensions suggest that all the three shaping methods could obviously speed up RL. In comparison, it can also be found that the two latter methods are more robust and efficient to the quality of human demonstrations as it takes more source samples\u2019 information into consideration.\n                  <\/jats:p>","DOI":"10.3233\/jifs-17052","type":"journal-article","created":{"date-parts":[[2018,1,19]],"date-time":"2018-01-19T10:56:49Z","timestamp":1516359409000},"page":"711-720","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":0,"title":["Shaping in reinforcement learning by knowledge transferred from human-demonstrations of a simple similar task"],"prefix":"10.1177","volume":"34","author":[{"given":"Guo-Fang","family":"Wang","sequence":"first","affiliation":[{"name":"Zhejiang University","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhou","family":"Fang","sequence":"additional","affiliation":[{"name":"Zhejiang University","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ping","family":"Li","sequence":"additional","affiliation":[{"name":"Zhejiang University","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2018,1,12]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2013.08.037"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913495721"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature16961"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(99)00052-1"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.conb.2012.05.008"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCC.2012.2218595"},{"issue":"5","key":"e_1_3_2_8_2","first-page":"156","article-title":"Autonomous reinforcement learning with experience replay[J]","volume":"41","author":"Wawrzy\u0144ski P.","year":"2012","unstructured":"Wawrzy\u0144skiP. and TanwaniA.K., Autonomous reinforcement learning with experience replay[J], Neural Networks the Official Journal of the International Neural Network Society 41(5) (2012), 156\u2013167.","journal-title":"Neural Networks the Official Journal of the International Neural Network Society"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2008.01.004"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2015.03.009"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2009.191"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390225"},{"issue":"10","key":"e_1_3_2_13_2","first-page":"1633","article-title":"Transfer learning for reinforcement learning domains: A survey[J]","volume":"10","author":"Taylor M.E.","year":"2009","unstructured":"TaylorM.E. and StoneP., Transfer learning for reinforcement learning domains: A survey[J], Journal of Machine Learning Research 10(10) (2009), 1633\u20131685.","journal-title":"Journal of Machine Learning Research"},{"issue":"1","key":"e_1_3_2_14_2","first-page":"1333","article-title":"Transfer in reinforcement learning via shared features[J]","volume":"13","author":"Konidaris G.","year":"2012","unstructured":"KonidarisG., ScheidwasserI. and BartoA.G., Transfer in reinforcement learning via shared features[J], Journal of Machine Learning Research 13(1) (2012), 1333\u20131371.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/11871842_41"},{"key":"e_1_3_2_16_2","first-page":"3352","volume-title":"Reinforcement learning from demonstration through shaping[C]","author":"Brys T.","year":"2015","unstructured":"BrysT., HarutyunyanA., SuayH.B., et al., Reinforcement learning from demonstration through shaping[C], International Conference on Artificial Intelligence, AAAI Press, 2015, pp. 3352\u20133358."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/1160633.1160762"},{"issue":"1","key":"e_1_3_2_18_2","first-page":"1437","article-title":"A comprehensive survey on safe reinforcement learning[J]","volume":"16","author":"Garc\u00eda J.","year":"2015","unstructured":"Garc\u00edaJ. and Fern\u00e1ndezF., A comprehensive survey on safe reinforcement learning[J], Journal of Machine Learning Research 16(1) (2015), 1437\u20131480.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_19_2","first-page":"3137","article-title":"A generalized path integral control approach to reinforcement learning[J]","author":"Theodorou E.","year":"2010","unstructured":"TheodorouE., BuchliJ., SchaalS., et al., A generalized path integral control approach to reinforcement learning[J], Journal of Machine Learning Research (2010), 3137\u20133181.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_20_2","first-page":"617","volume-title":"Integrating reinforcement learning with human demonstrations of varying ability[C]","author":"Taylor M.E.","year":"2011","unstructured":"TaylorM.E., SuayH.B. and ChernovaS., Integrating reinforcement learning with human demonstrations of varying ability[C], International Conference on Autonomous Agents and Multiagent Systems DBLP, 2011, pp. 617\u2013624."},{"key":"e_1_3_2_21_2","first-page":"3033","volume-title":"Shaping in reinforcement learning via knowledge transferred from human-demonstrations [C]","author":"Wang G.F.","year":"2015","unstructured":"WangG.F., FangZ., LiP., et al., Shaping in reinforcement learning via knowledge transferred from human-demonstrations [C], Chinese Control Conference, 2015, pp. 3033\u20133038."},{"key":"e_1_3_2_22_2","article-title":"Transferring knowledge from human-demonstration trajectories to reinforcement learning[J]","author":"Wang G.F.","year":"2016","unstructured":"WangG.F., FangZ., LiP., et al., Transferring knowledge from human-demonstration trajectories to reinforcement learning[J], Transactions of the Institute of Measurement & Control, 2016.","journal-title":"Transactions of the Institute of Measurement & Control"},{"issue":"7","key":"e_1_3_2_23_2","first-page":"1180","article-title":"Natural actor-critic[J]","volume":"71","author":"Peters J.","year":"2005","unstructured":"PetersJ. and SchaalS., Natural actor-critic[J], Neurocomputing 71(7\/9) (2005), 1180\u20131190.","journal-title":"Neurocomputing"},{"key":"e_1_3_2_24_2","first-page":"278","volume-title":"Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping[C]","author":"Ng A.Y.","year":"1999","unstructured":"NgA.Y., HaradaD. and RussellS.J., Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping[C], Sixteenth International Conference on Machine Learning. Morgan Kaufmann Publishers Inc, 1999, pp. 278\u2013287."},{"issue":"1","key":"e_1_3_2_25_2","first-page":"205","article-title":"Potential-based shaping and Q-value initialization are equivalent[J]","volume":"19","author":"Wiewiora E.","year":"2011","unstructured":"WiewioraE., Potential-based shaping and Q-value initialization are equivalent[J], Journal of Artificial Intelligence Research 19(1) (2011), 205\u2013208.","journal-title":"Journal of Artificial Intelligence Research"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2008.10.024"},{"key":"e_1_3_2_27_2","unstructured":"AbbeelP. Apprenticeship learning and reinforcement learning with application to robotic control[M] Stanford University 2008."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2012.08.039"}],"container-title":["Journal of Intelligent &amp; Fuzzy Systems: Applications in Engineering and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/JIFS-17052","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.3233\/JIFS-17052","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/JIFS-17052","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:38:07Z","timestamp":1777455487000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.3233\/JIFS-17052"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,1,12]]},"references-count":27,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2018,1,12]]}},"alternative-id":["10.3233\/JIFS-17052"],"URL":"https:\/\/doi.org\/10.3233\/jifs-17052","relation":{},"ISSN":["1064-1246","1875-8967"],"issn-type":[{"value":"1064-1246","type":"print"},{"value":"1875-8967","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,1,12]]}}}