{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,15]],"date-time":"2026-03-15T23:19:49Z","timestamp":1773616789695,"version":"3.50.1"},"reference-count":37,"publisher":"Wiley","issue":"1","license":[{"start":{"date-parts":[[2021,8,4]],"date-time":"2021-08-04T00:00:00Z","timestamp":1628035200000},"content-version":"vor","delay-in-days":215,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61906198"],"award-info":[{"award-number":["61906198"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004608","name":"Natural Science Foundation of Jiangsu Province","doi-asserted-by":"publisher","award":["BK20190622"],"award-info":[{"award-number":["BK20190622"]}],"id":[{"id":"10.13039\/501100004608","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computational Intelligence and Neuroscience"],"published-print":{"date-parts":[[2021,1]]},"abstract":"<jats:p>The reinforcement learning algorithms based on policy gradient may fall into local optimal due to gradient disappearance during the update process, which in turn affects the exploration ability of the reinforcement learning agent. In order to solve the above problem, in this paper, the cross\u2010entropy method (CEM) in evolution policy, maximum mean difference (MMD), and twin delayed deep deterministic policy gradient algorithm (TD3) are combined to propose a diversity evolutionary policy deep reinforcement learning (DEPRL) algorithm. By using the maximum mean discrepancy as a measure of the distance between different policies, some of the policies in the population maximize the distance between them and the previous generation of policies while maximizing the cumulative return during the gradient update. Furthermore, combining the cumulative returns and the distance between policies as the fitness of the population encourages more diversity in the offspring policies, which in turn can reduce the risk of falling into local optimal due to the disappearance of the gradient. The results in the MuJoCo test environment show that DEPRL has achieved excellent performance on continuous control tasks; especially in the Ant\u2010v2 environment, the return of DEPRL ultimately achieved a nearly 20% improvement compared to TD3.<\/jats:p>","DOI":"10.1155\/2021\/5300189","type":"journal-article","created":{"date-parts":[[2021,8,4]],"date-time":"2021-08-04T18:05:11Z","timestamp":1628100311000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Diversity Evolutionary Policy Deep Reinforcement Learning"],"prefix":"10.1155","volume":"2021","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0470-924X","authenticated-orcid":false,"given":"Jian","family":"Liu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liming","family":"Feng","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2021,8,4]]},"reference":[{"key":"e_1_2_9_1_2","doi-asserted-by":"publisher","DOI":"10.1109\/tnn.2004.842673"},{"key":"e_1_2_9_2_2","first-page":"39","article-title":"Transfer of reinforcement learning: the state of the art","volume":"36","author":"Wang H.","year":"2008","journal-title":"Acta Electronica Sinica"},{"key":"e_1_2_9_3_2","article-title":"Machine-learning research","volume":"18","author":"Dietterich T. G.","year":"1997","journal-title":"AI Magazine"},{"key":"e_1_2_9_4_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.aaa8415"},{"key":"e_1_2_9_5_2","unstructured":"SilverD. LeverG. HeessN. DegrisT. WierstraD. andRiedmillerM. Deterministic policy gradient algorithms Proceedings of the International Conference on Machine Learning June 2014 Beijing China 387\u2013395."},{"key":"e_1_2_9_6_2","unstructured":"SchulmanJ. WolskiF. DhariwalP. RadfordA. andKlimovO. Proximal policy optimization algorithms 2017 https:\/\/arxiv.org\/abs\/1707.06347."},{"key":"e_1_2_9_7_2","first-page":"1889","article-title":"Trust region policy optimization","volume":"3","author":"Schulman J.","year":"2015","journal-title":"Computer Science"},{"key":"e_1_2_9_8_2","unstructured":"TesslerC. TennenholtzG. andMannorS. Distributional policy optimization: an alternative approach for continuous control 2019 https:\/\/arxiv.org\/abs\/1905.09855."},{"key":"e_1_2_9_9_2","unstructured":"LillicrapT. P. HuntJ. J. PritzelA.et al. Continuous control with deep reinforcement learning 2015 https:\/\/arxiv.org\/abs\/1509.02971."},{"key":"e_1_2_9_10_2","unstructured":"FujimotoS. HoofH. andMegerD. Addressing function approximation error in actor-critic methods 2018 https:\/\/arxiv.org\/abs\/1802.09477."},{"key":"e_1_2_9_11_2","unstructured":"HaarnojaT. ZhouA. AbbeelP. andLevineS. Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor 2018 https:\/\/arxiv.org\/abs\/1801.01290."},{"key":"e_1_2_9_12_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_2_9_13_2","doi-asserted-by":"crossref","unstructured":"Van HasseltH. GuezA. andSilverD. Deep reinforcement learning with double Q-learning 2016 https:\/\/arxiv.org\/abs\/1509.06461.","DOI":"10.1609\/aaai.v30i1.10295"},{"key":"e_1_2_9_14_2","unstructured":"FortunatoM. AzarM. G. PiotB.et al. Noisy networks for exploration 2017 https:\/\/arxiv.org\/abs\/1706.10295."},{"key":"e_1_2_9_15_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-019-1724-z"},{"key":"e_1_2_9_16_2","first-page":"1958","article-title":"Application of multi-agent reinforcement learning in robot soccer","volume":"38","author":"Liu C. Y.","year":"2010","journal-title":"Acta Electronica Sinica"},{"key":"e_1_2_9_17_2","doi-asserted-by":"crossref","unstructured":"WilliamsJ. D. AtuiK. A. andZweigG. Hybrid code networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning 2017 https:\/\/arxiv.org\/abs\/1702.03274.","DOI":"10.18653\/v1\/P17-1062"},{"key":"e_1_2_9_18_2","first-page":"949","article-title":"Natural evolution policies","volume":"15","author":"Wierstra D.","year":"2014","journal-title":"The Journal of Machine Learning Research"},{"key":"e_1_2_9_19_2","unstructured":"TimS. HoJ. ChenX. SidorS. andSutskeverI. Evolution policies as a scalable alternative to reinforcement learning 2017 https:\/\/arxiv.org\/abs\/1703.03864."},{"key":"e_1_2_9_20_2","unstructured":"KhadkaS.andTumerK. Evolutionary reinforcement learning 2018 https:\/\/arxiv.org\/abs\/1805.07917."},{"key":"e_1_2_9_21_2","unstructured":"PourchotA.andSigaudO. CEM-RL: combining evolutionary and gradient-based methods for policy search 2018 https:\/\/arxiv.org\/abs\/1810.01222."},{"key":"e_1_2_9_22_2","first-page":"1201","article-title":"Research progress of genetic algorithm","volume":"4","author":"Yun W. X.","year":"2012","journal-title":"Application Research of Computers"},{"key":"e_1_2_9_23_2","volume-title":"Estimation of Distribution Algorithms: A New Tool for Evolutionary Computation","author":"Pedro L.","year":"2001"},{"key":"e_1_2_9_24_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.swevo.2011.08.003"},{"key":"e_1_2_9_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-015-1874-3"},{"key":"e_1_2_9_26_2","doi-asserted-by":"publisher","DOI":"10.1177\/1687814015624832"},{"key":"e_1_2_9_27_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_2_9_28_2","first-page":"901","article-title":"The mode and algorithm for the target threat assessment base on elman-adaboost storing predictor","volume":"40","author":"Wang G. G.","year":"2012","journal-title":"Acta Electronica Sinica"},{"key":"e_1_2_9_29_2","doi-asserted-by":"publisher","DOI":"10.1155\/2013\/632437"},{"key":"e_1_2_9_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/tii.2018.2822680"},{"key":"e_1_2_9_31_2","unstructured":"BrockmanG. CheungV. PetterssonL.et al. Openai gym 2016 https:\/\/arxiv.org\/abs\/1606.01540."},{"key":"e_1_2_9_32_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-015-1923-y"},{"key":"e_1_2_9_33_2","doi-asserted-by":"publisher","DOI":"10.1504\/ijbic.2018.093328"},{"key":"e_1_2_9_34_2","doi-asserted-by":"crossref","unstructured":"WangG. DebS. andCoelhoL. Elephant herding optimization Proceedings of the 2015 3rd International Symposium on Computational And Business Intelligence (ISCBI) December 2015 Bali Indonesia https:\/\/doi.org\/10.1109\/ISCBI.2015.8 2-s2.0-84964829151.","DOI":"10.1109\/ISCBI.2015.8"},{"key":"e_1_2_9_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2020.03.055"},{"key":"e_1_2_9_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12293-016-0212-3"},{"key":"e_1_2_9_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2019.02.028"}],"container-title":["Computational Intelligence and Neuroscience"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/downloads.hindawi.com\/journals\/cin\/2021\/5300189.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/cin\/2021\/5300189.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1155\/2021\/5300189","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,6]],"date-time":"2024-08-06T12:09:18Z","timestamp":1722946158000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1155\/2021\/5300189"}},"subtitle":[],"editor":[{"given":"Nian","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2021,1]]},"references-count":37,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,1]]}},"alternative-id":["10.1155\/2021\/5300189"],"URL":"https:\/\/doi.org\/10.1155\/2021\/5300189","archive":["Portico"],"relation":{},"ISSN":["1687-5265","1687-5273"],"issn-type":[{"value":"1687-5265","type":"print"},{"value":"1687-5273","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,1]]},"assertion":[{"value":"2021-06-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-07-19","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-08-04","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"5300189"}}