{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T21:07:27Z","timestamp":1761599247792,"version":"3.37.3"},"reference-count":28,"publisher":"Springer Science and Business Media LLC","issue":"21","license":[{"start":{"date-parts":[[2023,7,28]],"date-time":"2023-07-28T00:00:00Z","timestamp":1690502400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2023,7,28]],"date-time":"2023-07-28T00:00:00Z","timestamp":1690502400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61873008","62103053"],"award-info":[{"award-number":["61873008","62103053"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Beijing Natural Science Foundation","award":["4192010"],"award-info":[{"award-number":["4192010"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Appl Intell"],"published-print":{"date-parts":[[2023,11]]},"DOI":"10.1007\/s10489-023-04816-w","type":"journal-article","created":{"date-parts":[[2023,7,28]],"date-time":"2023-07-28T14:02:10Z","timestamp":1690552930000},"page":"24792-24803","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["A stable data-augmented reinforcement learning method with ensemble exploration and exploitation"],"prefix":"10.1007","volume":"53","author":[{"given":"Guoyu","family":"Zuo","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhipeng","family":"Tian","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0003-0184","authenticated-orcid":false,"given":"Gao","family":"Huang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,7,28]]},"reference":[{"issue":"7540","key":"4816_CR1","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, Graves A, Riedmiller M, Fidjeland AK, Ostrovski G et al (2015) Human-level control through deep reinforcement learning. Nature 518(7540):529\u2013533","journal-title":"Nature"},{"key":"4816_CR2","unstructured":"Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O (2017) Proximal policy optimization algorithms. arXiv preprint\u00a0https:\/\/arxiv.org\/abs\/1707.06347 arXiv:1707.06347"},{"key":"4816_CR3","first-page":"741","volume":"33","author":"AX Lee","year":"2020","unstructured":"Lee AX, Nagabandi A, Abbeel P, Levine S (2020) Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model. Adv Neural Inf Process Syst 33:741\u2013752","journal-title":"Adv Neural Inf Process Syst"},{"key":"4816_CR4","first-page":"10674","volume":"35","author":"D Yarats","year":"2021","unstructured":"Yarats D, Zhang A, Kostrikov I, Amos B, Pineau J, Fergus R (2021) Improving sample efficiency in model-free reinforcement learning from images. Proceed AAAI Conf Artif Intell 35:10674\u201310681","journal-title":"Proceed AAAI Conf Artif Intell"},{"key":"4816_CR5","doi-asserted-by":"publisher","unstructured":"Dwibedi D, Tompson J, Lynch C, Sermanet P (2018) Learning actionable representations from visual observations. In 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, pp 1577\u20131584. https:\/\/doi.org\/10.1109\/IROS.2018.8593951","DOI":"10.1109\/IROS.2018.8593951"},{"key":"4816_CR6","doi-asserted-by":"publisher","unstructured":"Yarats D, Zhang A, Kostrikov I, Amos B, Pineau J, Fergus R (2019) Improving sample efficiency in model-free reinforcement learning from images. In Proceedings of the AAAI Conference on Artificial Intelligence 35(12):10674\u201310681.\u00a0https:\/\/doi.org\/10.1609\/aaai.v35i12.17276","DOI":"10.1609\/aaai.v35i12.17276"},{"key":"4816_CR7","unstructured":"Igl M, Ciosek K, Li Y, Tschiatschek S, Zhang C, Devlin S, Hofmann K (2019) Generalization in reinforcement learning with selective noise injection and information bottleneck. Adv Neural Inf Proces Syst, 32. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2019\/file\/e2ccf95a7f2e1878fcafc8376649b6e8-Paper.pdf"},{"key":"4816_CR8","unstructured":"Haarnoja T, Zhou A, Abbeel P, Levine S (2018) Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In International conference on machine learning, PMLR, pp 1861\u20131870. https:\/\/proceedings.mlr.press\/v80\/haarnoja18b.html"},{"key":"4816_CR9","unstructured":"Choi J, Guo Y, Moczulski M, Oh J, Wu N, Norouzi M, Lee H (2018) Contingency-Aware Exploration in Reinforcement Learning"},{"key":"4816_CR10","unstructured":"Osband I, Blundell C, Pritzel A, Roy BV (2016) Deep Exploration via Bootstrapped DQN.\u00a0Advances in neural information processing systems, 29. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2016\/file\/8d8818c8e140c64c743113f563cf750f-Paper.pdf"},{"key":"4816_CR11","unstructured":"Chen RY, Sidor S, Abbeel P, Schulman J (2017) UCB Exploration via Q-Ensembles.\u00a0arXiv preprint https:\/\/arxiv.org\/abs\/1706.01502 arXiv:1706.01502"},{"issue":"3","key":"4816_CR12","doi-asserted-by":"publisher","first-page":"21","DOI":"10.1109\/MCAS.2006.1688199","volume":"6","author":"R Polikar","year":"2006","unstructured":"Polikar R (2006) Essemble based systems in decision making. IEEE Circ Syst Mag 6(3):21\u201345. https:\/\/doi.org\/10.1109\/MCAS.2006.1688199","journal-title":"IEEE Circ Syst Mag"},{"issue":"1","key":"4816_CR13","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1109\/MCI.2015.2471235","volume":"11","author":"Y Ren","year":"2016","unstructured":"Ren Y, Zhang L, Suganthan PN (2016) Ensemble classification and regression-recent developments, applications and future directions [review article]. IEEE Comput Intell Mag 11(1):41\u201353","journal-title":"IEEE Comput Intell Mag"},{"issue":"10","key":"4816_CR14","doi-asserted-by":"publisher","first-page":"8986","DOI":"10.1016\/j.eswa.2012.02.025","volume":"39","author":"Y Kim","year":"2012","unstructured":"Kim Y, Sohn SY (2012) Stock fraud detection using peer group analysis. Expert Syst Appl 39(10):8986\u20138992","journal-title":"Expert Syst Appl"},{"issue":"11\u201312","key":"4816_CR15","doi-asserted-by":"publisher","first-page":"4224","DOI":"10.1080\/01431161.2013.774099","volume":"34","author":"T Kavzoglu","year":"2013","unstructured":"Kavzoglu T, Colkesen I (2013) An assessment of the effectiveness of a rotation forest ensemble for land-use and land-cover mapping. Int J Remote Sens 34(11\u201312):4224\u20134241","journal-title":"Int J Remote Sens"},{"issue":"pt.a","key":"4816_CR16","first-page":"65","volume":"149","author":"H Min","year":"2015","unstructured":"Min H, Liu B (2015) Ensemble of extreme learning machine for remote sensing image classification. Neurocomputing 149(pt.a):65\u201370","journal-title":"Neurocomput"},{"key":"4816_CR17","unstructured":"Laskin M, Srinivas A, Abbeel P (2020) Curl: Contrastive unsupervised representations for reinforcement learning. In International Conference on Machine Learning, pages 5639\u20135650. PMLR"},{"key":"4816_CR18","unstructured":"Laskin M, Lee K, Stooke A, Pinto L, Abbeel P, Srinivas A (2020) Reinforcement learning with augmented data. Advances in neural information processing systems 33:19884\u201319895. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2020\/file\/e615c82aba461681ade82da2da38004a-Paper.pdf"},{"key":"4816_CR19","unstructured":"Kostrikov I, Yarats D, Fergus R (2020) Image augmentation is all you need: Regularizing deep reinforcement learning from pixels. arXiv preprint https:\/\/arxiv.org\/abs\/2004.13649 arXiv:2004.13649"},{"key":"4816_CR20","unstructured":"Tassa Y, Doron Y, Muldal A, Erez T, Li Y, de Las Casas D, Budden D, Abdolmaleki A, Merel J, Lefrancq A, et al. (2018) Deepmind control suite.\u00a0arXiv preprint https:\/\/arxiv.org\/abs\/1801.00690 arXiv:1801.00690"},{"key":"4816_CR21","unstructured":"Haarnoja T, Zhou A, Hartikainen K, Tucker G, Ha S, Tan J, Kumar V, Zhu H, Gupta A, Abbeel P, et al. (2018) Soft actor-critic algorithms and applications.\u00a0arXiv preprint https:\/\/arxiv.org\/abs\/1812.05905\u00a0arXiv:1812.05905"},{"key":"4816_CR22","unstructured":"Schwarzer M, Anand A, Goel R, Hjelm RD, Courville A, Bachman P (2020) Data-efficient reinforcement learning with self-predictive representations. arXiv preprint\u00a0https:\/\/arxiv.org\/abs\/2007.05929\u00a0arXiv:2007.05929"},{"key":"4816_CR23","unstructured":"Osband I, Blundell C, Pritzel A, Van Roy B (2016) Deep exploration via bootstrapped\u00a0DQN. Adv Neural Inf Proces Syst, 29. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2016\/file\/8d8818c8e140c64c743113f563cf750f-Paper.pdf"},{"key":"4816_CR24","unstructured":"Lee K, Laskin M, Srinivas A, Abbeel P (2021) Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning. In International Conference on Machine Learning, pages 6131\u20136141. PMLR"},{"key":"4816_CR25","doi-asserted-by":"publisher","unstructured":"Todorov E, Erez T, Tassa Y (2012) Mujoco: A physics engine for model-based control. In 2012 IEEE\/RSJ International Conference on Intelligent Robots and Systems, IEEE, 5026\u20135033.\u00a0https:\/\/doi.org\/10.1109\/IROS.2012.6386109","DOI":"10.1109\/IROS.2012.6386109"},{"key":"4816_CR26","unstructured":"Hafner D, Lillicrap T, Fischer I, Villegas R, Ha D, Lee H, Davidson J (2019) Learning latent dynamics for planning from pixels. In International Conference on Machine Learning,\u00a0PMLR pp. 2555\u20132565.\u00a0https:\/\/proceedings.mlr.press\/v97\/hafner19a.html"},{"key":"4816_CR27","unstructured":"Hafner D, Lillicrap T, Ba J, Norouzi M (2019) Dream to control: Learning behaviors by latent imagination.\u00a0arXiv preprint https:\/\/arxiv.org\/abs\/1912.01603 arXiv:1912.01603"},{"key":"4816_CR28","unstructured":"Lan Q, Pan Y, Fyshe A, White M (2019) Maxmin q-learning: Controlling the estimation bias of q-learning.\u00a0arXiv preprint https:\/\/arxiv.org\/abs\/2002.06487 arXiv:2002.06487"}],"container-title":["Applied Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-023-04816-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10489-023-04816-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-023-04816-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,10,23]],"date-time":"2023-10-23T14:07:36Z","timestamp":1698070056000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10489-023-04816-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,28]]},"references-count":28,"journal-issue":{"issue":"21","published-print":{"date-parts":[[2023,11]]}},"alternative-id":["4816"],"URL":"https:\/\/doi.org\/10.1007\/s10489-023-04816-w","relation":{},"ISSN":["0924-669X","1573-7497"],"issn-type":[{"type":"print","value":"0924-669X"},{"type":"electronic","value":"1573-7497"}],"subject":[],"published":{"date-parts":[[2023,7,28]]},"assertion":[{"value":"19 June 2023","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 July 2023","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest. This paper has not been previously published, it is published with the permission of the authors\u2019 institution, and all authors of this paper are responsible for the authenticity of the data in the paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical Approval"}},{"value":"All authors of this paper have been informed of the revision and publication of the paper, have checked all data, figures and tables in the manuscript, and are responsible for their truthfulness and accuracy. Names of all contributing authors: Guoyu Zuo; Zhipeng Tian; Gao Huang.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to Participate"}},{"value":"The publication has been approved by all co-authors.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"All authors of this paper declare no conflict of interest in this paper and agree to submit this manuscript to the journal of Applied Intelligence.","order":6,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing Interests"}}]}}