{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T01:57:04Z","timestamp":1782957424993,"version":"3.54.5"},"reference-count":23,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2021,6,11]],"date-time":"2021-06-11T00:00:00Z","timestamp":1623369600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No. 61741303"],"award-info":[{"award-number":["No. 61741303"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"The key laboratory of spatial information and geomatics (Guilin University of Technology)","award":["No.19-185-10-08"],"award-info":[{"award-number":["No.19-185-10-08"]}]},{"name":"The Scientific Research Basic Ability Enhancement Program for Young and Middle-aged Teachers of Guangxi","award":["No. 2021KY0260"],"award-info":[{"award-number":["No. 2021KY0260"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Directing at various problems of the traditional Q-Learning algorithm, such as heavy repetition and disequilibrium of explorations, the reinforcement-exploration strategy was used to replace the decayed \u03b5-greedy strategy in the traditional Q-Learning algorithm, and thus a novel self-adaptive reinforcement-exploration Q-Learning (SARE-Q) algorithm was proposed. First, the concept of behavior utility trace was introduced in the proposed algorithm, and the probability for each action to be chosen was adjusted according to the behavior utility trace, so as to improve the efficiency of exploration. Second, the attenuation process of exploration factor \u03b5 was designed into two phases, where the first phase centered on the exploration and the second one transited the focus from the exploration into utilization, and the exploration rate was dynamically adjusted according to the success rate. Finally, by establishing a list of state access times, the exploration factor of the current state is adaptively adjusted according to the number of times the state is accessed. The symmetric grid map environment was established via OpenAI Gym platform to carry out the symmetrical simulation experiments on the Q-Learning algorithm, self-adaptive Q-Learning (SA-Q) algorithm and SARE-Q algorithm. The experimental results show that the proposed algorithm has obvious advantages over the first two algorithms in the average number of turning times, average inside success rate, and number of times with the shortest planned route.<\/jats:p>","DOI":"10.3390\/sym13061057","type":"journal-article","created":{"date-parts":[[2021,6,14]],"date-time":"2021-06-14T22:26:01Z","timestamp":1623709561000},"page":"1057","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":36,"title":["A Self-Adaptive Reinforcement-Exploration Q-Learning Algorithm"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6976-846X","authenticated-orcid":false,"given":"Lieping","family":"Zhang","sequence":"first","affiliation":[{"name":"College of Mechanical and Control Engineering, Guilin University of Technology, Guilin 541006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Liu","family":"Tang","sequence":"additional","affiliation":[{"name":"College of Mechanical and Control Engineering, Guilin University of Technology, Guilin 541006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shenglan","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Mechanical and Control Engineering, Guilin University of Technology, Guilin 541006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhengzhong","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Mechanical and Control Engineering, Guilin University of Technology, Guilin 541006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xianhao","family":"Shen","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Guilin University of Technology, Guilin 541006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zuqiong","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Guilin University of Technology, Guilin 541006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,6,11]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Zhou, X.M., Bai, T., Gao, Y.B., and Han, Y. (2019). Vision-Based Robot Navigation through Combining Unsupervised Learning and Hierarchical Reinforcement Learning. Sensors, 19.","DOI":"10.3390\/s19071576"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"106372","DOI":"10.1016\/j.ultras.2021.106372","article-title":"Supervised learning strategy for classification and regression tasks applied to aeronautical structural health monitoring problems","volume":"113","author":"Miorelli","year":"2021","journal-title":"Ultrasonics"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"107865","DOI":"10.1016\/j.comnet.2021.107865","article-title":"Towards website domain name classification using graph based semi-supervised learning","volume":"188","author":"Faroughi","year":"2021","journal-title":"Comput. Netw."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Zeng, J.J., Qin, L., Hu, Y., and Yin, Q. (2019). Combining Subgoal Graphs with Reinforcement Learning to Build a Rational Pathfinder. Appl. Sci., 9.","DOI":"10.3390\/app9020323"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"135426","DOI":"10.1109\/ACCESS.2020.3011438","article-title":"A Survey on Visual Navigation for Artificial Agents with Deep Reinforcement Learning","volume":"8","author":"Zeng","year":"2020","journal-title":"IEEE Access"},{"key":"ref_6","first-page":"13","article-title":"Overview on Algorithms and Applications for Reinforcement Learning","volume":"29","author":"Li","year":"2020","journal-title":"Comput. Syst. Appl."},{"key":"ref_7","first-page":"1","article-title":"Hybrid genetic algorithm based smooth global-path planning for a mobile robot","volume":"2021","author":"Luan","year":"2021","journal-title":"Mech. Based Des. Struct. Mach."},{"key":"ref_8","first-page":"91","article-title":"An Improved Q-Learning Algorithm and Its Application in Path Planning","volume":"52","author":"Mao","year":"2021","journal-title":"J. Taiyuan Univ. Technol."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"426","DOI":"10.1016\/j.jmsy.2021.02.014","article-title":"A study on a Q-Learning algorithm application to a manufacturing assembly problem","volume":"59","author":"Neves","year":"2021","journal-title":"J. Manuf. Syst."},{"key":"ref_10","unstructured":"Han, X.C., Yu, S.P., Yuan, Z.M., and Cheng, L.J. (2021). High-speed railway dynamic scheduling based on Q-Learning method. Control Theory Appl., Available online: https:\/\/kns.cnki.net\/kcms\/detail\/44.1240.TP.20210330.1333.042.html."},{"key":"ref_11","first-page":"1747","article-title":"Neural network-based reinforcement learning applied to obstacle avoidance","volume":"48","author":"Qiao","year":"2008","journal-title":"J. Tsinghua Univ. Sci. Technol."},{"key":"ref_12","first-page":"1623","article-title":"Initialization in reinforcement learning for mobile robots path planning","volume":"29","author":"Song","year":"2012","journal-title":"Control Theory Appl."},{"key":"ref_13","unstructured":"Zhao, Y.N. (2017). Research of Path Planning Problem Based on Reinforcement Learning. [Master\u2019s Thesis, Harbin Institute of Technology]."},{"key":"ref_14","first-page":"185","article-title":"Research of path planning based on the supervised reinforcement learning","volume":"35","author":"Zeng","year":"2018","journal-title":"Comput. Appl. Softw."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"da Silva, A.G., dos Santos, D.H., de Negreiros, A.P.F., Silva, J.M., and Gon\u00e7alves, L.M.G. (2020). High-Level Path Planning for an Autonomous Sailboat Robot Using Q-Learning. Sensors, 20.","DOI":"10.3390\/s20061550"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1016\/j.robot.2019.02.013","article-title":"Solving the optimal path planning of a mobile robot using improved Q-learning","volume":"115","author":"Low","year":"2019","journal-title":"Robot. Auton. Syst."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Park, J.H., and Lee, K.H. (2021). Computational Design of Modular Robots Based on Genetic Algorithm and Reinforcement Learning. Symmetry, 13.","DOI":"10.3390\/sym13030471"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"29064","DOI":"10.1109\/ACCESS.2020.2971780","article-title":"Path Planning for UAV Ground Target Tracking via Deep Reinforcement Learning","volume":"8","author":"Li","year":"2020","journal-title":"IEEE Access"},{"key":"ref_19","unstructured":"Yan, J.J., Zhang, Q.S., and Hu, X.P. (2021). Review of Path Planning Techniques Based on Reinforcement Learning. Comput. Eng."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Seo, K., and Yang, J. (2020). Differentially Private Actor and Its Eligibility Trace. Electronics, 9.","DOI":"10.3390\/electronics9091486"},{"key":"ref_21","first-page":"180","article-title":"Overview of Research on Model-free Reinforcement Learning","volume":"48","author":"Qin","year":"2021","journal-title":"Comput. Sci."},{"key":"ref_22","unstructured":"Li, T. (2020). Research of Path Planning Algorithm based on Reinforcement Learning. [Master\u2019s Thesis, Jilin University]."},{"key":"ref_23","unstructured":"Li, T., and Li, Y. (2019, January 16\u201317). A Novel Path Planning Algorithm Based on Q-learning and Adaptive Exploration Strategy. Proceedings of the 2019 Scientific Conference on Network, Power Systems and Computing (NPSC 2019), Guilin, China."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/13\/6\/1057\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:13:24Z","timestamp":1760163204000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/13\/6\/1057"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,6,11]]},"references-count":23,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2021,6]]}},"alternative-id":["sym13061057"],"URL":"https:\/\/doi.org\/10.3390\/sym13061057","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,6,11]]}}}