{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T15:59:33Z","timestamp":1784563173869,"version":"3.55.0"},"reference-count":25,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,2,13]],"date-time":"2025-02-13T00:00:00Z","timestamp":1739404800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neurorobot."],"abstract":"<jats:p>Aiming at the problems of slow network convergence, poor reward convergence stability, and low path planning efficiency of traditional deep reinforcement learning algorithms, this paper proposes a BiLSTM-D3QN (Bidirectional Long and Short-Term Memory Dueling Double Deep Q-Network) path planning algorithm based on the DDQN (Double Deep Q-Network) decision model. Firstly, a Bidirectional Long Short-Term Memory network (BiLSTM) is introduced to make the network have memory, increase the stability of decision making and make the reward converge more stably; secondly, Dueling Network is introduced to further solve the problem of overestimating the Q-value of the neural network, which makes the network able to be updated quickly; Adaptive reprioritization based on the frequency penalty function is proposed. Experience Playback, which extracts important and fresh data from the experience pool to accelerate the convergence of the neural network; finally, an adaptive action selection mechanism is introduced to further optimize the action exploration. Simulation experiments show that the BiLSTM-D3QN path planning algorithm outperforms the traditional Deep Reinforcement Learning algorithm in terms of network convergence speed, planning efficiency, stability of reward convergence, and success rate in simple environments; in complex environments, the path length of BiLSTM-D3QN is 20\u202fm shorter than that of the improved ERDDQN (Experience Replay Double Deep Q-Network) algorithm, the number of turning points is 7 fewer, the planning time is 0.54\u202fs shorter, and the success rate is 10.4% higher. The superiority of the BiLSTM-D3QN algorithm in terms of network convergence speed and path planning performance is demonstrated.<\/jats:p>","DOI":"10.3389\/fnbot.2025.1512953","type":"journal-article","created":{"date-parts":[[2025,2,13]],"date-time":"2025-02-13T07:09:46Z","timestamp":1739430586000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["Path planning of mobile robot based on improved double deep Q-network algorithm"],"prefix":"10.3389","volume":"19","author":[{"given":"Zhenggang","family":"Wang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuhong","family":"Song","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shenghui","family":"Cheng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2025,2,13]]},"reference":[{"key":"ref1","doi-asserted-by":"publisher","first-page":"2061","DOI":"10.3390\/s24072061","article-title":"Improved double deep Q-network algorithm applied to multi-dimensional environment path planning of hexapod robots","volume":"24","author":"Chen","year":"2024","journal-title":"Sensors"},{"key":"ref2","doi-asserted-by":"publisher","first-page":"115208","DOI":"10.1016\/j.oceaneng.2023.115208","article-title":"Deep reinforcement learning with dynamic window approach based collision avoidance path planning for maritime autonomous surface ships","volume":"284","author":"Chuanbo","year":"2023","journal-title":"Ocean Eng."},{"key":"ref3","doi-asserted-by":"publisher","first-page":"302","DOI":"10.3390\/drones8070302","article-title":"An integrated geometric obstacle avoidance and genetic algorithm TSP model for UAV path planning","volume":"8","author":"Debnath","year":"2024","journal-title":"Drones"},{"key":"ref4","doi-asserted-by":"publisher","first-page":"1523","DOI":"10.3390\/s24051523","article-title":"Enhancing stability and performance in Mobile robot path planning with PMR-dueling DQN algorithm","volume":"24","author":"Deguale","year":"2024","journal-title":"Sensors"},{"key":"ref5","doi-asserted-by":"publisher","first-page":"426","DOI":"10.3390\/s20020426","article-title":"An autonomous path planning model for unmanned ships based on deep reinforcement learning","volume":"20","author":"Guo","year":"2020","journal-title":"Sensors"},{"key":"ref6","doi-asserted-by":"publisher","DOI":"10.3390\/s23125622","article-title":"Improved robot path planning method based on deep reinforcement learning","volume":"23","author":"Huiyan","year":"2023","journal-title":"Sensors (Basel, Switzerland)"},{"key":"ref7","doi-asserted-by":"publisher","first-page":"1767","DOI":"10.3390\/e24121767","article-title":"Path planning research of a UAV Base station searching for disaster victims\u2019 location information based on deep reinforcement learning","volume":"24","author":"Jinduo","year":"2022","journal-title":"Entropy"},{"key":"ref8","doi-asserted-by":"publisher","first-page":"5493","DOI":"10.3390\/s20195493","article-title":"Deep reinforcement learning for indoor Mobile robot path planning","volume":"20","author":"Junli","year":"2020","journal-title":"Sensors (Basel, Switzerland)"},{"key":"ref9","doi-asserted-by":"publisher","first-page":"171302898","DOI":"10.3389\/fnbot.2023.1302898","article-title":"Multi-UAV simultaneous target assignment and path planning based on deep reinforcement learning in dynamic multiple obstacles environments","volume":"17","author":"Kong","year":"2024","journal-title":"Front. Neurorobot."},{"key":"ref10","doi-asserted-by":"publisher","first-page":"119410","DOI":"10.1016\/j.eswa.2022.119410","article-title":"Modified adaptive ant colony optimization algorithm and its application for solving path planning of mobile robot","volume":"215","author":"Lei","year":"2023","journal-title":"Expert Syst. Appl."},{"key":"ref11","doi-asserted-by":"publisher","first-page":"111","DOI":"10.19678\/j.issn.1000-3428.0066348","article-title":"Robot path planning based on improved DQN algorithm","volume":"49","author":"Li","year":"2023","journal-title":"Comput. Eng."},{"key":"ref12","doi-asserted-by":"publisher","first-page":"232","DOI":"10.1504\/IJVD.2023.131056","article-title":"Improved duelling deep Q-networks based path planning for intelligent agents","volume":"91","author":"Lin","year":"2023","journal-title":"Int. J. Veh. Des."},{"key":"ref13","doi-asserted-by":"publisher","first-page":"8392","DOI":"10.3390\/app12178392","article-title":"An overview of variants and advancements of PSO algorithm","volume":"12","author":"Meetu","year":"2022","journal-title":"Appl. Sci."},{"key":"ref14","doi-asserted-by":"publisher","first-page":"842","DOI":"10.1109\/TNNLS.2023.3332172","article-title":"An information-assisted deep reinforcement learning path planning scheme for dynamic and unknown underwater environment","volume":"36","author":"Meng","year":"2023","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref15","doi-asserted-by":"publisher","first-page":"171268447","DOI":"10.3389\/fnbot.2023.1268447","article-title":"A survey of path planning of industrial robots based on rapidly exploring random trees","volume":"17","author":"Sha","year":"2023","journal-title":"Front. Neurorobot."},{"key":"ref16","doi-asserted-by":"publisher","first-page":"30","DOI":"10.19651\/j.cnki.emt.2211675","article-title":"UAV regional coverage path planning strategy based on DDQN","volume":"46","author":"Shen","year":"2023","journal-title":"Electron. Meas. Technol."},{"key":"ref17","doi-asserted-by":"publisher","first-page":"60","DOI":"10.3390\/drones8020060","article-title":"Dynamic scene path planning of UAVs based on deep reinforcement learning","volume":"8","author":"Tang","year":"2024","journal-title":"Drones"},{"key":"ref18","doi-asserted-by":"publisher","first-page":"200","DOI":"10.3969\/j.issn.0255-8297.2024.02.002","article-title":"UAV path and radio mapping based on deep reinforcement learning","volume":"42","author":"Wang","year":"2024","journal-title":"J. Appl. Sci."},{"key":"ref19","doi-asserted-by":"publisher","first-page":"2036","DOI":"10.3390\/s23042036","article-title":"A Mapless local path planning approach using deep reinforcement learning framework","volume":"23","author":"Yan","year":"2023","journal-title":"Sensors"},{"key":"ref20","doi-asserted-by":"publisher","first-page":"521","DOI":"10.1007\/s11370-024-00536-3","article-title":"A* algorithm based on adaptive expansion convolution for unmanned aerial vehicle path planning","volume":"17","author":"Yu","year":"2024","journal-title":"Intell. Serv. Robot."},{"key":"ref21","doi-asserted-by":"publisher","first-page":"923","DOI":"10.20009\/j.cnki.21-1106\/TP.2021-0713","article-title":"Research on D3QN path planning method of Mobile robot priority sampling","volume":"44","author":"Yuan","year":"2023","journal-title":"J. Chin. Comput. Syst."},{"key":"ref22","doi-asserted-by":"publisher","first-page":"4287","DOI":"10.1007\/S40747-022-00948-7","article-title":"DM-DQN: dueling Munchausen deep Q network for robot path planning","volume":"9","author":"Yuwan","year":"2022","journal-title":"Complex Intell. Syst."},{"key":"ref23","doi-asserted-by":"publisher","first-page":"7918","DOI":"10.3390\/s23187918","article-title":"Research on path planning and path tracking control of autonomous vehicles based on improved APF and SMC","volume":"23","author":"Zhang","year":"2023","journal-title":"Sensors"},{"key":"ref24","doi-asserted-by":"publisher","first-page":"127958","DOI":"10.1016\/j.neucom.2024.127958","article-title":"EPPE: an efficient progressive policy enhancement framework of deep reinforcement learning in path planning","volume":"596","author":"Zhao","year":"2024","journal-title":"Neurocomputing"},{"key":"ref25","doi-asserted-by":"publisher","first-page":"9955","DOI":"10.3390\/app13179955","article-title":"Path planning of rail-mounted logistics robots based on the improved Dijkstra algorithm","volume":"13","author":"Zhou","year":"2023","journal-title":"Appl. Sci."}],"container-title":["Frontiers in Neurorobotics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2025.1512953\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,13]],"date-time":"2025-02-13T07:09:51Z","timestamp":1739430591000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2025.1512953\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,13]]},"references-count":25,"alternative-id":["10.3389\/fnbot.2025.1512953"],"URL":"https:\/\/doi.org\/10.3389\/fnbot.2025.1512953","relation":{},"ISSN":["1662-5218"],"issn-type":[{"value":"1662-5218","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,13]]},"article-number":"1512953"}}