{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T14:59:46Z","timestamp":1753887586102,"version":"3.41.2"},"reference-count":31,"publisher":"Wiley","issue":"1","license":[{"start":{"date-parts":[[2021,10,8]],"date-time":"2021-10-08T00:00:00Z","timestamp":1633651200000},"content-version":"vor","delay-in-days":280,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61762032","11961018"],"award-info":[{"award-number":["61762032","11961018"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computational Intelligence and Neuroscience"],"published-print":{"date-parts":[[2021,1]]},"abstract":"<jats:p>The deep Q\u2010network (DQN) is one of the most successful reinforcement learning algorithms, but it has some drawbacks such as slow convergence and instability. In contrast, the traditional reinforcement learning algorithms with linear function approximation usually have faster convergence and better stability, although they easily suffer from the curse of dimensionality. In recent years, many improvements to DQN have been made, but they seldom make use of the advantage of traditional algorithms to improve DQN. In this paper, we propose a novel Q\u2010learning algorithm with linear function approximation, called the minibatch recursive least squares Q\u2010learning (MRLS\u2010Q). Different from the traditional Q\u2010learning algorithm with linear function approximation, the learning mechanism and model structure of MRLS\u2010Q are more similar to those of DQNs with only one input layer and one linear output layer. It uses the experience replay and the minibatch training mode and uses the agent\u2019s states rather than the agent\u2019s state\u2010action pairs as the inputs. As a result, it can be used alone for low\u2010dimensional problems and can be seamlessly integrated into DQN as the last layer for high\u2010dimensional problems as well. In addition, MRLS\u2010Q uses our proposed average RLS optimization technique, so that it can achieve better convergence performance whether it is used alone or integrated with DQN. At the end of this paper, we demonstrate the effectiveness of MRLS\u2010Q on the CartPole problem and four Atari games and investigate the influences of its hyperparameters experimentally.<\/jats:p>","DOI":"10.1155\/2021\/5370281","type":"journal-article","created":{"date-parts":[[2021,10,10]],"date-time":"2021-10-10T21:35:09Z","timestamp":1633901709000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Minibatch Recursive Least Squares Q\u2010Learning"],"prefix":"10.1155","volume":"2021","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2785-0510","authenticated-orcid":false,"given":"Chunyuan","family":"Zhang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qi","family":"Song","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zeng","family":"Meng","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2021,10,8]]},"reference":[{"volume-title":"Reinforcement Learning: An Introduction","year":"2018","author":"Sutton R. S.","key":"e_1_2_9_1_2"},{"key":"e_1_2_9_2_2","unstructured":"MnihV. KavukcuogluK. SilverD. GravesA. AntonoglouI. WierstraD. andRiedmillerM. Playing atari with deep reinforcement learning Proceedings of the 27th Conference on Neural Information Processing Systems Deep Learning Workshop 2013 Lake Tahoe USA."},{"key":"e_1_2_9_3_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_2_9_4_2","unstructured":"HaarnojaT. ZhouA. AbbeelP. andLevineS. Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor Proceedings of the 35th International Conference on Machine Learning July 2018 Stockholm Sweden 1856\u20131865."},{"key":"e_1_2_9_5_2","unstructured":"KimD. LiuM. RiemerM. SunC. AbdulhaiM. HabibiG. Lopez-CotS. TesauroG. andHowJ. P. A policy gradient algorithm for learning to learn in multiagent reinforcement learning Proceedings of the 38th International Conference on Machine Learning 2021 5541\u20135550."},{"key":"e_1_2_9_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/tvt.2021.3099129"},{"key":"e_1_2_9_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/tii.2020.3024611"},{"key":"e_1_2_9_8_2","doi-asserted-by":"crossref","unstructured":"WangP.andChanC. Formulation of deep reinforcement learning architecture toward autonomous driving for on-ramp merge Proceedings of the 20th International Conference on Intelligent Transportation Systems 2017.","DOI":"10.1109\/ITSC.2017.8317735"},{"key":"e_1_2_9_9_2","doi-asserted-by":"crossref","unstructured":"YangB.andLiuM. Keeping in touch with collaborative UAVs: a deep reinforcement learning approach Proceedings of the 27th International Joint Conference on Artificial Intelligence 2018 562\u2013568.","DOI":"10.24963\/ijcai.2018\/78"},{"key":"e_1_2_9_10_2","doi-asserted-by":"crossref","unstructured":"LiuL. TianB. ZhaoX. andZongQ. UAV autonomous trajectory planning in target tracking tasks via a DQN approach Proceedings of the 2019 International Conference on Real-time Computing and Robotics August 2019 Irkutsk Russia.","DOI":"10.1109\/RCAR47638.2019.9044134"},{"key":"e_1_2_9_11_2","doi-asserted-by":"crossref","unstructured":"HasseltH. V. GuezA. andSilverD. Deep reinforcement learning with double Q-learning Proceedings of the 30th AAAI Conference on Artificial Intelligence February 2016 Phoenix AZ USA 2094\u20132100.","DOI":"10.1609\/aaai.v30i1.10295"},{"key":"e_1_2_9_12_2","unstructured":"WangZ. SchaulT. HesselM. van HasseltH. LanctotM. andde FreitasN. Dueling network architectures for deep reinforcement learning Proceedings of the 33rd International Conference on Machine Learning June 2016 New York NY USA 1995\u20132003."},{"key":"e_1_2_9_13_2","unstructured":"HausknechtM. J.andStoneP. Deep recurrent Q-learning for partially observable MDPs Proceedings of the 2015 AAAI Fall Symposia November 2015 Arlington County VA USA 29\u201337."},{"key":"e_1_2_9_14_2","doi-asserted-by":"crossref","unstructured":"KimS. AsadiK. LittmanM. andKonidarisG. DeepMellow: removing the need for a target network in deep Q-learning Proceedings of the 28th International Joint Conference on Artificial Intelligence August 2019 Macao China 2733\u20132739.","DOI":"10.24963\/ijcai.2019\/379"},{"key":"e_1_2_9_15_2","unstructured":"AnschelO. BaramN. andShimkinN. Averaged-DQN: variance reduction and stabilization for deep reinforcement learning Proceedings of the 34th International Conference on Machine Learning August 2017 Sydney Australia 176\u2013185."},{"key":"e_1_2_9_16_2","unstructured":"SchaulT. QuanJ. AntonoglouI. andSilverD. Prioritized experience replay Proceedings of the 4th International Conference on Learning Representations May 2016 Juan Puerto Rico."},{"key":"e_1_2_9_17_2","unstructured":"FortunatoM. AzarM. G. PiotB. MenickJ. OsbandI. GravesA. MnihV. MunosR. HassabisD. PietquinO. BlundellC. andLeggS. Noisy networks for exploration Proceedings of the 6th International Conference on Learning Representations 2018 Vancouver Canada."},{"key":"e_1_2_9_18_2","unstructured":"LeeS. Y. ChoiS. andChungS. Sample-efficient deep reinforcement learning via episodic backward update Proceedings of the 33rd Conference on Neural Information Processing Systems 2019 Vancouver Canada 2110\u20132119."},{"key":"e_1_2_9_19_2","unstructured":"MnihV. BadiaA. P. MirzaM. GravesA. LillicrapT. P. HarleyT. SilverD. andKavukcuogluK. Asynchronous methods for deep reinforcement learning Proceedings of the 33nd International Conference on Machine Learning June 2016 New York NY USA 1928\u20131937."},{"key":"e_1_2_9_20_2","unstructured":"LevineN. ZahavyT. MankowitzD. J. TamarA. andMannorS. Shallow updates for deep reinforcement learning Proceedings of the 31st International Conference on Neural Information Processing Systems December 2017 Long Beach CA USA 3135\u20133145."},{"key":"e_1_2_9_21_2","doi-asserted-by":"publisher","DOI":"10.1162\/jmlr.2003.4.6.1107"},{"key":"e_1_2_9_22_2","first-page":"503","article-title":"Tree-based batch mode reinforcement learning","volume":"6","author":"Ernst D.","year":"2005","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_9_23_2","doi-asserted-by":"publisher","DOI":"10.1002\/acs.3282"},{"key":"e_1_2_9_24_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2005.12.126"},{"key":"e_1_2_9_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/tnnls.2017.2716952"},{"key":"e_1_2_9_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/tnnls.2020.3043110"},{"key":"e_1_2_9_27_2","doi-asserted-by":"crossref","unstructured":"ZhangC. LiuC. SongQ. andZhaoJ. Recursive least squares policy control with echo state network Proceedings of the 4th International Conference on Artificial Intelligence and Big Data September 2021 Zibo China 104\u2013108.","DOI":"10.1109\/ICAIBD51990.2021.9458984"},{"volume-title":"Online Learning and Stochastic Approximations","year":"1998","author":"Bottou L.","key":"e_1_2_9_28_2"},{"key":"e_1_2_9_29_2","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177729893"},{"key":"e_1_2_9_30_2","doi-asserted-by":"crossref","unstructured":"Ek\u015fio\u011fluE. M. RLS adaptive filtering with sparsity regularization Proceedings of the 10th International Conference on Information Science Signal Processing and their Applications May 2010 Kuala Lumpur Malaysia 550\u2013553.","DOI":"10.1109\/ISSPA.2010.5605592"},{"key":"e_1_2_9_31_2","first-page":"2825","article-title":"Scikit-learn: machine learning in Python","volume":"12","author":"Pedregosa F.","year":"2011","journal-title":"Journal of Machine Learning Research"}],"container-title":["Computational Intelligence and Neuroscience"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/downloads.hindawi.com\/journals\/cin\/2021\/5370281.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/cin\/2021\/5370281.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1155\/2021\/5370281","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,6]],"date-time":"2024-08-06T11:08:57Z","timestamp":1722942537000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1155\/2021\/5370281"}},"subtitle":[],"editor":[{"given":"Yugen","family":"Yi","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2021,1]]},"references-count":31,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,1]]}},"alternative-id":["10.1155\/2021\/5370281"],"URL":"https:\/\/doi.org\/10.1155\/2021\/5370281","archive":["Portico"],"relation":{},"ISSN":["1687-5265","1687-5273"],"issn-type":[{"type":"print","value":"1687-5265"},{"type":"electronic","value":"1687-5273"}],"subject":[],"published":{"date-parts":[[2021,1]]},"assertion":[{"value":"2021-08-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-09-23","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-10-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"5370281"}}