{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,21]],"date-time":"2026-04-21T15:37:32Z","timestamp":1776785852592,"version":"3.51.2"},"reference-count":27,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2025,5,19]],"date-time":"2025-05-19T00:00:00Z","timestamp":1747612800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Central Government Guides Local Science and Technology Special Project-Yangtze River Delta Science and Technology Innovation Community Joint Research Program","award":["2023CSJGG1600"],"award-info":[{"award-number":["2023CSJGG1600"]}]},{"name":"Central Government Guides Local Science and Technology Special Project-Yangtze River Delta Science and Technology Innovation Community Joint Research Program","award":["2208085MF173"],"award-info":[{"award-number":["2208085MF173"]}]},{"name":"Central Government Guides Local Science and Technology Special Project-Yangtze River Delta Science and Technology Innovation Community Joint Research Program","award":["2023ZD01"],"award-info":[{"award-number":["2023ZD01"]}]},{"name":"Central Government Guides Local Science and Technology Special Project-Yangtze River Delta Science and Technology Innovation Community Joint Research Program","award":["2023ZD03"],"award-info":[{"award-number":["2023ZD03"]}]},{"name":"Anhui Provincial Natural Science Foundation","award":["2023CSJGG1600"],"award-info":[{"award-number":["2023CSJGG1600"]}]},{"name":"Anhui Provincial Natural Science Foundation","award":["2208085MF173"],"award-info":[{"award-number":["2208085MF173"]}]},{"name":"Anhui Provincial Natural Science Foundation","award":["2023ZD01"],"award-info":[{"award-number":["2023ZD01"]}]},{"name":"Anhui Provincial Natural Science Foundation","award":["2023ZD03"],"award-info":[{"award-number":["2023ZD03"]}]},{"name":"Wuhu \u201cRed Casting Light\u201d Major Science and Technology Project","award":["2023CSJGG1600"],"award-info":[{"award-number":["2023CSJGG1600"]}]},{"name":"Wuhu \u201cRed Casting Light\u201d Major Science and Technology Project","award":["2208085MF173"],"award-info":[{"award-number":["2208085MF173"]}]},{"name":"Wuhu \u201cRed Casting Light\u201d Major Science and Technology Project","award":["2023ZD01"],"award-info":[{"award-number":["2023ZD01"]}]},{"name":"Wuhu \u201cRed Casting Light\u201d Major Science and Technology Project","award":["2023ZD03"],"award-info":[{"award-number":["2023ZD03"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>The behavior decision control of autonomous vehicles is a critical aspect of advancing autonomous driving technology. However, current behavior decision algorithms based on deep reinforcement learning still face several challenges, such as insufficient safety and sparse reward mechanisms. To solve these problems, this paper proposes a dueling double deep Q-network based on dual-priority experience replay\u2014DPD3QN. Initially, the dueling network is integrated with the double deep Q-network, and the original network\u2019s output layer is restructured to enhance the precision of action value estimation. Subsequently, dual-priority experience replay is incorporated to facilitate the model\u2019s ability to swiftly recognize and leverage critical experiences. Ultimately, the training and evaluation are conducted on the OpenAI Gym simulation platform. The test results show that DPD3QN helps to improve the convergence speed of driverless vehicle behavior decision-making. Compared with the currently popular DQN and DDQN algorithms, this algorithm achieves higher success rates in challenging scenarios. Test scenario I increases by 11.8 and 25.8 percentage points, respectively, while the success rates in test scenarios I and II rise by 8.8 and 22.2 percentage points, respectively, indicating a more secure and efficient autonomous driving decision-making capability.<\/jats:p>","DOI":"10.3390\/a18050291","type":"journal-article","created":{"date-parts":[[2025,5,19]],"date-time":"2025-05-19T05:37:13Z","timestamp":1747633033000},"page":"291","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Dual-Priority Delayed Deep Double Q-Network (DPD3QN): A Dueling Double Deep Q-Network with Dual-Priority Experience Replay for Autonomous Driving Behavior Decision-Making"],"prefix":"10.3390","volume":"18","author":[{"given":"Shuai","family":"Li","sequence":"first","affiliation":[{"name":"School of Mechanical and Automotive Engineering, Anhui Polytechnic University, Wuhu 241000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1533-8154","authenticated-orcid":false,"given":"Peicheng","family":"Shi","sequence":"additional","affiliation":[{"name":"School of Mechanical and Automotive Engineering, Anhui Polytechnic University, Wuhu 241000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2029-0537","authenticated-orcid":false,"given":"Aixi","family":"Yang","sequence":"additional","affiliation":[{"name":"Polytechnic Institute, Zhejiang University, Hangzhou 310015, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Heng","family":"Qi","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-3278-6252","authenticated-orcid":false,"given":"Xinlong","family":"Dong","sequence":"additional","affiliation":[{"name":"School of Mechanical and Automotive Engineering, Anhui Polytechnic University, Wuhu 241000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,5,19]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Dong, X., Shi, P., Tang, Y., Yang, L., Yang, A., and Liang, T. (2024). Vehicle Classification Algorithm Based on Improved Vision Transformer. World Electr. Veh. J., 15.","DOI":"10.3390\/wevj15080344"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"102814","DOI":"10.1016\/j.displa.2024.102814","article-title":"TS-BEV: BEV object detection algorithm based on temporal-spatial feature fusion","volume":"84","author":"Dong","year":"2024","journal-title":"Displays"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"103771","DOI":"10.1016\/j.jvcir.2023.103771","article-title":"MT-Net: Fast video instance lane detection based on space time memory and template matching","volume":"91","author":"Shi","year":"2023","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"187","DOI":"10.1146\/annurev-control-060117-105157","article-title":"Planning and decision-making for autonomous vehicles","volume":"1","author":"Schwarting","year":"2018","journal-title":"Annu. Rev. Control Robot. Auton. Syst."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1016\/j.trc.2016.07.007","article-title":"Influence of connected and autonomous vehicles on traffic flow stability and throughput","volume":"71","author":"Talebpour","year":"2016","journal-title":"Transp. Res. Part C Emerg. Technol."},{"key":"ref_6","unstructured":"Bojarski, M., Del Testa, D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., Jackel, L.D., Monfort, M., Muller, U., and Zhang, J. (2016). End to end learning for self-driving cars. arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1007\/BF00992698","article-title":"Q-learning","volume":"8","author":"Watkins","year":"1992","journal-title":"Mach. Learn."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_9","first-page":"390","article-title":"Blind, greedy, and random: Algorithms for matching and clustering using only ordinal information","volume":"30","author":"Anshelevich","year":"2016","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_10","unstructured":"Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N. (2016, January 19\u201324). Dueling network architectures for deep reinforcement learning. Proceedings of the International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_11","first-page":"110341","article-title":"Reward Machines for Deep RL in Noisy and Uncertain Environments","volume":"37","author":"Li","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"328","DOI":"10.1007\/s42154-021-00151-3","article-title":"End-to-end autonomous driving through dueling double deep Q-network","volume":"4","author":"Peng","year":"2021","journal-title":"Automot. Innov."},{"key":"ref_13","unstructured":"Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2015). Prioritized experience replay. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"293","DOI":"10.1007\/BF00992699","article-title":"Self-improving reactive agents based on reinforcement learning, planning and teaching","volume":"8","author":"Lin","year":"1992","journal-title":"Mach. Learn."},{"key":"ref_15","unstructured":"Puterman, M.L. (2014). Markov Decision Processes: Discrete Stochastic Dynamic Programming, John Wiley & Sons."},{"key":"ref_16","unstructured":"Lazaric, A., Restelli, M., and Bonarini, A. (2007). Reinforcement learning in continuous action spaces through sequential monte carlo methods. Adv. Neural Inf. Process. Syst., 20, Available online: https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2007\/file\/0f840be9b8db4d3fbd5ba2ce59211f55-Paper.pdf."},{"key":"ref_17","unstructured":"Hausknecht, M.J., and Stone, P. (2015, January 12\u201314). Deep Recurrent Q-Learning for Partially Observable MDPs. Proceedings of the AAAI Fall Symposia, Arlington, VA, USA."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Baird, L.C., and Klopf, A.H. (1993). Reinforcement Learning with High-Dimensional Continuous Actions, Wright Laboratory. Technical Report WL-TR-93-1147.","DOI":"10.21236\/ADA280844"},{"key":"ref_19","unstructured":"Fujimoto, S., Hoof, H., and Meger, D. (2018, January 10\u201315). Addressing function approximation error in actor-critic methods. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Rubinstein, R.Y., and Kroese, D.P. (2016). Simulation and the Monte Carlo Method, John Wiley & Sons.","DOI":"10.1002\/9781118631980"},{"key":"ref_21","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Sutton, R.S., and Barto, A.G. (1998). Reinforcement Learning: An Introduction, MIT Press.","DOI":"10.1109\/TNN.1998.712192"},{"key":"ref_23","unstructured":"Leurent, E., and Mercat, J. (2019). Social attention for autonomous decision-making in dense traffic. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"288","DOI":"10.1177\/03611981231205877","article-title":"Research on dueling double deep Q network algorithm based on single-step momentum update","volume":"2678","author":"Shi","year":"2024","journal-title":"Transp. Res. Rec."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1805","DOI":"10.1103\/PhysRevE.62.1805","article-title":"Congested traffic states in empirical observations and microscopic simulations","volume":"62","author":"Treiber","year":"2000","journal-title":"Phys. Rev. E"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"5597","DOI":"10.1103\/PhysRevE.55.5597","article-title":"Metastable states in a microscopic model of traffic flow","volume":"55","author":"Wagner","year":"1997","journal-title":"Phys. Rev. E"},{"key":"ref_27","first-page":"119","article-title":"Integrated control of longitudinal and lateral motion for autonomous vehicle driving system","volume":"23","author":"Jie","year":"2010","journal-title":"China J. Highw. Transp."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/18\/5\/291\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T17:35:01Z","timestamp":1760031301000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/18\/5\/291"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,19]]},"references-count":27,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2025,5]]}},"alternative-id":["a18050291"],"URL":"https:\/\/doi.org\/10.3390\/a18050291","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,19]]}}}