{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T17:42:19Z","timestamp":1779385339818,"version":"3.53.1"},"reference-count":31,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2023,11,15]],"date-time":"2023-11-15T00:00:00Z","timestamp":1700006400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neurorobot."],"abstract":"<jats:sec><jats:title>Introduction<\/jats:title><jats:p>Deep deterministic policy gradient (DDPG)-based path planning algorithms for intelligent robots struggle to discern the value of experience transitions during training due to their reliance on a random experience replay. This can lead to inappropriate sampling of experience transitions and overemphasis on edge experience transitions. As a result, the algorithm's convergence becomes slower, and the success rate of path planning diminishes.<\/jats:p><\/jats:sec><jats:sec><jats:title>Methods<\/jats:title><jats:p>We comprehensively examines the impacts of immediate reward, temporal-difference error (TD-error), and Actor network loss function on the training process. It calculates experience transition priorities based on these three factors. Subsequently, using information entropy as a weight, the three calculated priorities are merged to determine the final priority of the experience transition. In addition, we introduce a method for adaptively adjusting the priority of positive experience transitions to focus on positive experience transitions and maintain a balanced distribution. Finally, the sampling probability of each experience transition is derived from its respective priority.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>The experimental results showed that the test time of our method is shorter than that of PER algorithm, and the number of collisions with obstacles is less. It indicated that the determined experience transition priority accurately gauges the significance of distinct experience transitions for path planning algorithm training.<\/jats:p><\/jats:sec><jats:sec><jats:title>Discussion<\/jats:title><jats:p>This method enhances the utilization rate of transition conversion and the convergence speed of the algorithm and also improves the success rate of path planning.<\/jats:p><\/jats:sec>","DOI":"10.3389\/fnbot.2023.1281166","type":"journal-article","created":{"date-parts":[[2023,11,15]],"date-time":"2023-11-15T09:08:02Z","timestamp":1700039282000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Prioritized experience replay in path planning via multi-dimensional transition priority fusion"],"prefix":"10.3389","volume":"17","author":[{"given":"Nuo","family":"Cheng","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peng","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guangyuan","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cui","family":"Ni","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Erkin","family":"Nematov","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2023,11,15]]},"reference":[{"key":"B1","first-page":"1510","article-title":"\u201cHigh-value prioritized experience replay for off-policy reinforcement learning,\u201d","volume-title":"Proceedings of the 2019 IEEE 31st International Conference on Tools with Artificial Intelligence (ICTAI)","author":"Cao","year":"2019"},{"key":"B2","doi-asserted-by":"publisher","first-page":"16842","DOI":"10.1109\/TITS.2021.3131473","article-title":"An adaptive clustering-based algorithm for automatic path planning of heterogeneous UAVs","volume":"23","author":"Chen","year":"2021","journal-title":"IEEE Trans. Intell. Transp. Syst"},{"key":"B3","doi-asserted-by":"publisher","DOI":"10.3390\/app9204198","article-title":"Mapless collaborative navigation for a multi-robot system based on the deep reinforcement learning","author":"Chen","year":"2019","journal-title":"Appl. Sci"},{"key":"B4","first-page":"1255","article-title":"\u201cOff-policy correction for deep deterministic policy gradient algorithms via batch prioritized experience replay,\u201d","volume-title":"Proceedings of the 2021 IEEE 33rd International Conference on Tools with Artificial Intelligence (ICTAI)","author":"Cicek","year":"2021"},{"key":"B5","first-page":"52","article-title":"\u201cMobile robot path planning based on improved DDPG reinforcement learning algorithm,\u201d","volume-title":"Proceedings of the 2020 IEEE 11th International Conference on Software Engineering and Service Science (ICSESS)","author":"Dong","year":"2020"},{"key":"B6","unstructured":"An equivalence between loss functions and non-uniform sampling in experience replay1421914230\n            FujimotoS.\n            MegerD.\n            PrecupD.\n          Adv. Neural Inf. Process. Syst.332020"},{"key":"B7","first-page":"4548","article-title":"\u201cCan Q-learning be improved with advice?\u201d","volume-title":"Proceedings of the Conference on Learning Theory","author":"Golowich","year":"2022"},{"key":"B8","doi-asserted-by":"publisher","first-page":"210","DOI":"10.3390\/jmse9020210","article-title":"Path planning of coastal ships based on optimized DQN reward function","volume":"9","author":"Guo","year":"2021","journal-title":"J. Mar. Sci. Eng."},{"key":"B9","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2020.106736","article-title":"Regularly updated deterministic policy gradient algorithm","author":"Han","year":"2021","journal-title":"Knowl. Based Syst"},{"key":"B10","doi-asserted-by":"publisher","DOI":"10.3390\/ijgi10110785","article-title":"Improved A-star algorithm for long-distance off-road path planning using terrain data map","author":"Hong","year":"2021","journal-title":"ISPRS Int. J. Geo. Inf"},{"key":"B11","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2021.103949","article-title":"Enhanced ant colony algorithm with communication mechanism for mobile robot path planning","author":"Hou","year":"2022","journal-title":"Robot. Auton. Syst"},{"key":"B12","doi-asserted-by":"publisher","first-page":"108875","DOI":"10.1016\/j.patcog.2022.108875","article-title":"Clustering experience replay for the effective exploitation in reinforcement learning","volume":"131","author":"Li","year":"2022","journal-title":"Pattern Recognit."},{"key":"B13","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1155\/2021\/5169460","article-title":"Research on dynamic path planning of mobile robot based on improved DDPG algorithm","volume":"2021","author":"Li","year":"2021","journal-title":"Mob. Inf. Syst"},{"key":"B14","doi-asserted-by":"publisher","DOI":"10.1016\/j.compag.2021.106350","article-title":"Collision-free path planning for a guava-harvesting robot based on recurrent deep reinforcement learning","author":"Lin","year":"2021","journal-title":"Comput. Agric"},{"key":"B15","doi-asserted-by":"publisher","DOI":"10.1155\/2021\/8881684","article-title":"Path planning for smart car based on Dijkstra algorithm and dynamic window approach","author":"Liu","year":"2021","journal-title":"Wirel. Commun. Mob, Comput"},{"key":"B16","doi-asserted-by":"publisher","DOI":"10.1016\/j.aei.2021.101360","article-title":"Deep reinforcement learning-based safe interaction for industrial human-robot collaboration using intrinsic reward function","author":"Liu","year":"2021","journal-title":"Adv. Eng. Inform"},{"key":"B17","doi-asserted-by":"publisher","first-page":"1901","DOI":"10.1007\/s10586-021-03235-1","article-title":"A path planning method based on the particle swarm optimization trained fuzzy neural network algorithm","volume":"24","author":"Liu","year":"2021","journal-title":"Clust. Comput"},{"key":"B18","doi-asserted-by":"publisher","first-page":"4408","DOI":"10.1007\/s10489-020-02095-3","article-title":"Aspect-gated graph convolutional networks for aspect-based sentiment analysis","volume":"51","author":"Lu","year":"2021","journal-title":"Appl. Intell"},{"key":"B19","doi-asserted-by":"publisher","DOI":"10.1016\/j.cie.2021.107230","article-title":"Path planning optimization of indoor mobile robot based on adaptive ant colony algorithm","author":"Miao","year":"2021","journal-title":"Comput. Ind. Eng"},{"key":"B20","doi-asserted-by":"publisher","first-page":"247","DOI":"10.1023\/A:1017988514716","article-title":"Continuous-action Q-learning","volume":"49","author":"Mill\u00e1n","year":"2002","journal-title":"Mach. Learn."},{"key":"B21","article-title":"\u201cRemember and forget for experience replay,\u201d","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Novati","year":"2019"},{"key":"B22","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2007.07358","article-title":"Learning to sample with local and global contexts in experience replay buffer","author":"Oh","year":"2007","journal-title":"arXiv"},{"key":"B23","doi-asserted-by":"publisher","DOI":"10.1016\/j.oceaneng.2021.108709","article-title":"The hybrid path planning algorithm based on improved A* and artificial potential field for unmanned surface vehicle formations","author":"Sang","year":"2021","journal-title":"Ocean Eng"},{"key":"B24","article-title":"\u201cExperience replay with likelihood-free importance weights,\u201d","volume-title":"Proceedings of the Learning for Dynamics and Control Conference","author":"Sinha","year":"2022"},{"key":"B25","doi-asserted-by":"publisher","first-page":"9326","DOI":"10.1109\/TCYB.2021.3053414","article-title":"Deep reinforcement learning with quantum-inspired experience replay","volume":"52","author":"Wei","year":"2022","journal-title":"IEEE Trans. Cybern."},{"key":"B26","doi-asserted-by":"publisher","first-page":"2588","DOI":"10.1109\/TIE.2021.3070514","article-title":"Deep deterministic policy gradient-DRL enabled multiphysics-constrained fast charging of lithium-ion battery","volume":"69","author":"Wei","year":"2021","journal-title":"IEEE Trans. Ind. Electron"},{"key":"B27","doi-asserted-by":"crossref","first-page":"7112","DOI":"10.1109\/CAC.2017.8244061","article-title":"\u201cApplication of deep reinforcement learning in mobile robot path planning,\u201d","volume-title":"Proceedings of the 2017 Chinese Automation Congress (CAC)","author":"Xin","year":"2017"},{"key":"B28","doi-asserted-by":"publisher","DOI":"10.1016\/j.compeleceng.2022.108015","article-title":"A deep deterministic policy gradient algorithm based on averaged state-action estimation","author":"Xu","year":"2022","journal-title":"Comput. Electr. Eng"},{"key":"B29","doi-asserted-by":"publisher","first-page":"63","DOI":"10.3389\/fnbot.2020.00063","article-title":"The path planning of mobile robot by neural networks and hierarchical reinforcement learning","volume":"14","author":"Yu","year":"2020","journal-title":"Front. Neurorobot."},{"key":"B30","doi-asserted-by":"publisher","first-page":"1243174","DOI":"10.3389\/fnbot.2023.1243174","article-title":"A meta-learning based decision-making method for collaborative UAV swarm","volume":"17","author":"Zhao","year":"2023","journal-title":"Front. Neurorobot."},{"key":"B31","doi-asserted-by":"publisher","first-page":"103223","DOI":"10.1016\/j.ipm.2022.103223","article-title":"Knowledge-guided multi-granularity GCN for ABSA","volume":"60","author":"Zhu","year":"2023","journal-title":"Inf. Process. Manag."}],"container-title":["Frontiers in Neurorobotics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2023.1281166\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,15]],"date-time":"2023-11-15T09:08:10Z","timestamp":1700039290000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2023.1281166\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,15]]},"references-count":31,"alternative-id":["10.3389\/fnbot.2023.1281166"],"URL":"https:\/\/doi.org\/10.3389\/fnbot.2023.1281166","relation":{},"ISSN":["1662-5218"],"issn-type":[{"value":"1662-5218","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,15]]},"article-number":"1281166"}}