{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T18:48:17Z","timestamp":1775069297185,"version":"3.50.1"},"reference-count":21,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2022,12,5]],"date-time":"2022-12-05T00:00:00Z","timestamp":1670198400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neurorobot."],"abstract":"<jats:p>With the application and development of UAV technology and navigation and positioning technology, higher requirements are put forward for UAV maneuvering obstacle avoidance ability and real-time route planning. In this paper, for the problem of real-time UAV route planning in the unknown environment, we combine the ideas of artificial potential field method to modify the state observation and reward function, which solves the problem of sparse rewards of reinforcement learning algorithm, improves the convergence speed of the algorithm, and improves the generalization of the algorithm by step-by-step training based on the ideas of curriculum learning and transfer learning according to the difficulty of the task. The simulation results show that the improved SAC algorithm has fast convergence speed, good timeliness and strong generalization, and can better complete the UAV route planning task.<\/jats:p>","DOI":"10.3389\/fnbot.2022.1025817","type":"journal-article","created":{"date-parts":[[2022,12,5]],"date-time":"2022-12-05T04:52:54Z","timestamp":1670215974000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Real-time route planning of unmanned aerial vehicles based on improved soft actor-critic algorithm"],"prefix":"10.3389","volume":"16","author":[{"given":"Yuxiang","family":"Zhou","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiansheng","family":"Shu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaolong","family":"Zheng","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hui","family":"Hao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huan","family":"Song","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2022,12,5]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","DOI":"10.1145\/1553374.1553380","article-title":"Curriculum Learning","author":"Bengio","year":"2009","journal-title":"Proceedings of the 26th Annual International Conference on Machine Learning - ICML \u201809, 41\u201348. ICML \u201909"},{"key":"B2","doi-asserted-by":"publisher","DOI":"10.1007\/s10846-021-01367-5","article-title":"Soft actor-critic for navigation of mobile robots.","volume":"102","author":"De Jesus","year":"2021","journal-title":"J. Intell. Robot. Syst."},{"key":"B3","doi-asserted-by":"publisher","first-page":"269","DOI":"10.1007\/BF01386390","article-title":"A note on two problems in connexion with graphs.","volume":"1","author":"Dijkstra","year":"1959","journal-title":"Numer. Math."},{"key":"B4","first-page":"1587","article-title":"Addressing Function Approximation Error in Actor-Critic Methods","author":"Fujimoto","year":"2018","journal-title":"Proceedings of the 35th International Conference on Machine Learning"},{"key":"B5","doi-asserted-by":"publisher","DOI":"10.3390\/s20195493","article-title":"Deep reinforcement learning for indoor mobile robot path planning.","volume":"20","author":"Gao","year":"2020","journal-title":"Sensors"},{"key":"B6","doi-asserted-by":"publisher","DOI":"10.1007\/s10846-021-01568-y","article-title":"Double critic deep reinforcement learning for mapless 3d navigation of unmanned aerial vehicles.","volume":"104","author":"Grando","year":"2022","journal-title":"J. Intell. Robot. Syst."},{"key":"B7","doi-asserted-by":"publisher","first-page":"1861","DOI":"10.48550\/arXiv.1801.01290","article-title":"Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor","author":"Haarnoja","year":"2018","journal-title":"International Conference on Machine Learning"},{"key":"B8","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1016\/j.asoc.2017.03.035","article-title":"Mobile robot path planning with surrounding point set and path improvement.","volume":"57","author":"Han","year":"2017","journal-title":"Appl. Soft Comput."},{"key":"B9","doi-asserted-by":"publisher","first-page":"100","DOI":"10.1109\/TSSC.1968.300136","article-title":"A formal basis for the heuristic determination of minimum cost paths.","volume":"4","author":"Hart","year":"1968","journal-title":"IEEE Trans. Syst. Sci. Cybern."},{"key":"B10","doi-asserted-by":"publisher","DOI":"10.1016\/j.ast.2021.107052","article-title":"Explainable deep reinforcement learning for UAV autonomous path planning.","volume":"118","author":"He","year":"2021","journal-title":"Aerosp. Sci. Technol."},{"key":"B11","doi-asserted-by":"publisher","first-page":"846","DOI":"10.1177\/0278364911406761","article-title":"Sampling-based algorithms for optimal motion planning.","volume":"30","author":"Karaman","year":"2011","journal-title":"Int. J. Robot. Res."},{"key":"B12","doi-asserted-by":"publisher","first-page":"500","DOI":"10.1109\/ROBOT.1985.1087247","article-title":"Real-time obstacle avoidance for manipulators and mobile robots.","volume":"2","author":"Khatib","year":"1985","journal-title":"Proc. IEEE Int. Conf. Robot. Autom."},{"key":"B13","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1155\/2018\/5781591","article-title":"Dynamic path planning of unknown environment based on deep reinforcement learning.","volume":"2018","author":"Lei","year":"2018","journal-title":"J. Robot."},{"key":"B14","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-016-9115-2","article-title":"Path planning for mobile robot using self-adaptive learning particle swarm optimization.","volume":"61","author":"Li","year":"2018","journal-title":"Sci. China Inf. Sci."},{"key":"B15","doi-asserted-by":"publisher","first-page":"5829","DOI":"10.1007\/s00500-016-2161-7","article-title":"An improved ant colony algorithm for robot path planning.","volume":"21","author":"Liu","year":"2017","journal-title":"Soft Comput."},{"key":"B16","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1312.5602","article-title":"Playing atari with deep reinforcement learning.","author":"Mnih","year":"2013","journal-title":"arXiv"},{"key":"B17","doi-asserted-by":"publisher","DOI":"10.1016\/j.ast.2020.105882","article-title":"Stereo vision based obstacle collision avoidance for a quadrotor using ellipsoidal bounding box and hierarchical clustering.","volume":"103","author":"Park","year":"2020","journal-title":"Aerosp. Sci. Technol."},{"key":"B18","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1707.06347","article-title":"Proximal policy optimization algorithms.","author":"Schulman","year":"2017","journal-title":"arXiv"},{"key":"B19","doi-asserted-by":"publisher","first-page":"415","DOI":"10.1016\/j.isatra.2019.08.018","article-title":"Efficient path planning for uav formation via comprehensively improved particle swarm optimization.","volume":"97","author":"Shao","year":"2020","journal-title":"ISA Trans."},{"key":"B20","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2010.13166","article-title":"A survey on curriculum learning.","author":"Wang","year":"2021","journal-title":"arXiv"},{"key":"B21","doi-asserted-by":"publisher","DOI":"10.3389\/fnbot.2020.00063","article-title":"The path planning of mobile robot by neural networks and hierarchical reinforcement learning.","volume":"14","author":"Yu","year":"2020","journal-title":"Front. Neurorobot."}],"container-title":["Frontiers in Neurorobotics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2022.1025817\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,5]],"date-time":"2022-12-05T04:52:59Z","timestamp":1670215979000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2022.1025817\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,5]]},"references-count":21,"alternative-id":["10.3389\/fnbot.2022.1025817"],"URL":"https:\/\/doi.org\/10.3389\/fnbot.2022.1025817","relation":{},"ISSN":["1662-5218"],"issn-type":[{"value":"1662-5218","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,5]]},"article-number":"1025817"}}