{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T21:54:32Z","timestamp":1740174872232,"version":"3.37.3"},"reference-count":17,"publisher":"Wiley","license":[{"start":{"date-parts":[[2021,11,8]],"date-time":"2021-11-08T00:00:00Z","timestamp":1636329600000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Basic Research Program of China","doi-asserted-by":"publisher","award":["2020YFB1708503","2019YFB1705003","2018YFB1308801"],"award-info":[{"award-number":["2020YFB1708503","2019YFB1705003","2018YFB1308801"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Scientific Programming"],"published-print":{"date-parts":[[2021,11,8]]},"abstract":"<jats:p>We propose in this paper a new approach to solve the decision problem of robot-following. Different from the existing single policy model, we propose a multipolicy model, which can change the following policy in time according to the scene. The value of this paper is to obtain a multipolicy robot-following model by the self-learning method, which is used to improve the safety, efficiency, and stability of robot-following in the complex environments. Empirical investigation on a number of datasets reveals that overall, the proposed approach tends to have superior out-of-sample performance when compared to alternative robot-following decision methods. The performance of the model has been improved by about 2 times in situations where there are few obstacles and about 6 times in situations where there are lots of obstacles.<\/jats:p>","DOI":"10.1155\/2021\/5692105","type":"journal-article","created":{"date-parts":[[2021,11,8]],"date-time":"2021-11-08T23:20:33Z","timestamp":1636413633000},"page":"1-8","source":"Crossref","is-referenced-by-count":1,"title":["Multipolicy Robot-Following Model Based on Reinforcement Learning"],"prefix":"10.1155","volume":"2021","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7411-3874","authenticated-orcid":true,"given":"Ning","family":"Yu","sequence":"first","affiliation":[{"name":"Key Laboratory of Networked Control Systems, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Institutes for Robotics and Intelligent Manufacturing, Chinese Academy of Sciences, Shenyang 110169, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6904-3533","authenticated-orcid":true,"given":"Lin","family":"Nan","sequence":"additional","affiliation":[{"name":"Key Laboratory of Networked Control Systems, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Institutes for Robotics and Intelligent Manufacturing, Chinese Academy of Sciences, Shenyang 110169, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7080-9089","authenticated-orcid":true,"given":"Tao","family":"Ku","sequence":"additional","affiliation":[{"name":"Key Laboratory of Networked Control Systems, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Institutes for Robotics and Intelligent Manufacturing, Chinese Academy of Sciences, Shenyang 110169, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","reference":[{"key":"1","doi-asserted-by":"publisher","DOI":"10.23919\/ACC.1992.4792410"},{"key":"2","doi-asserted-by":"publisher","DOI":"10.1109\/IVS.2017.7995721"},{"issue":"6","key":"3","first-page":"667","article-title":"A vehicle adaptive cruise control algorithm based on simulating driver\u2019s multi-objective decision making","volume":"37","author":"Z. Gao","year":"2015","journal-title":"Automotive Engineering"},{"issue":"11","key":"4","first-page":"2031","article-title":"A control method of unmanned car following under time-varying relative distance and angle","volume":"48","author":"R. Li","year":"2018","journal-title":"Acta Automatica Sinica"},{"author":"R. Dechter","key":"5","article-title":"Learning while searching in constraint-satisfaction problems"},{"issue":"1","key":"6","first-page":"86","article-title":"The review of reinforcement learning","volume":"30","author":"Y. Gao","year":"2004","journal-title":"Acta Automatica Sinica"},{"key":"7","doi-asserted-by":"publisher","DOI":"10.1109\/tits.2017.2706963"},{"key":"8","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4302-5990-9_7","article-title":"Deep neural networks","volume-title":"Efficient Learning Machines","author":"M. Awad","year":"2015"},{"issue":"3","key":"9","first-page":"182","article-title":"Survey on sparse reward in deep reinforcement learning","volume":"47","author":"W. Yang","year":"2020","journal-title":"Computer Science"},{"key":"10","doi-asserted-by":"publisher","DOI":"10.1177\/1729881419853185"},{"issue":"6","key":"11","first-page":"1","article-title":"Multi-objective vehicle following decision algorithm based on reinforcement learning","volume":"36","author":"X. Deng","year":"2020","journal-title":"Control and Decision"},{"key":"12","doi-asserted-by":"publisher","DOI":"10.1108\/ir-07-2018-0154"},{"issue":"6","key":"13","first-page":"1","article-title":"Car-following method based on inverse reinforcement learning for autonomous vehicle decision-making","volume":"15","author":"H. Gao","year":"2019","journal-title":"International Journal of Advanced Robotic Systems"},{"author":"R. S. Sutton","key":"14","article-title":"Policy gradient methods for reinforcement learning with function approximation"},{"key":"15","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2007.11.026"},{"key":"16","doi-asserted-by":"publisher","DOI":"10.1109\/icisce.2017.120"},{"key":"17","doi-asserted-by":"publisher","DOI":"10.3390\/s19204373"}],"container-title":["Scientific Programming"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/downloads.hindawi.com\/journals\/sp\/2021\/5692105.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/sp\/2021\/5692105.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/sp\/2021\/5692105.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,11,8]],"date-time":"2021-11-08T23:20:49Z","timestamp":1636413649000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.hindawi.com\/journals\/sp\/2021\/5692105\/"}},"subtitle":[],"editor":[{"given":"Tongguang","family":"Ni","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2021,11,8]]},"references-count":17,"alternative-id":["5692105","5692105"],"URL":"https:\/\/doi.org\/10.1155\/2021\/5692105","relation":{},"ISSN":["1875-919X","1058-9244"],"issn-type":[{"type":"electronic","value":"1875-919X"},{"type":"print","value":"1058-9244"}],"subject":[],"published":{"date-parts":[[2021,11,8]]}}}