{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,17]],"date-time":"2026-08-17T15:25:09Z","timestamp":1786980309248,"version":"3.56.0"},"reference-count":39,"publisher":"Springer Science and Business Media LLC","issue":"30","license":[{"start":{"date-parts":[[2023,8,14]],"date-time":"2023-08-14T00:00:00Z","timestamp":1691971200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,8,14]],"date-time":"2023-08-14T00:00:00Z","timestamp":1691971200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"RESEARCH-CREATE-INNOVATE","award":["T2EDK-02743"],"award-info":[{"award-number":["T2EDK-02743"]}]},{"name":"Centre for Research & Technology Hellas"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2023,10]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    While reinforcement learning (RL) algorithms have generated impressive strategies for a wide range of tasks, the performance improvements in continuous-domain, real-world problems do not follow the same trend. Poor exploration and quick convergence to locally optimal solutions play a dominant role. Advanced RL algorithms attempt to mitigate this issue by introducing exploration signals during the training procedure. This successful integration has paved the way to introduce signals from the intrinsic exploration branch. ACRE algorithm is a framework that concretely describes the conditions for such an integration, avoiding transforming the Markov decision process into time varying, and as a result, making the whole optimization scheme brittle and susceptible to instability. The key distinction of ACRE lies in the way of handling and storing both extrinsic and intrinsic rewards. ACRE is an off-policy, actor-critic style RL algorithm that separately approximates the forward novelty return. ACRE is shipped with a Gaussian mixture model to calculate the instantaneous novelty; however, different options could also be integrated. Using such an effective early exploration, ACRE results in substantial improvements over alternative RL methods, in a range of continuous control RL environments, such as learning from policy-misleading reward signals. Open-source implementation is available here:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/athakapo\/ACRE\">https:\/\/github.com\/athakapo\/ACRE<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s00521-023-08845-x","type":"journal-article","created":{"date-parts":[[2023,8,14]],"date-time":"2023-08-14T07:02:06Z","timestamp":1691996526000},"page":"22563-22576","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["ACRE: Actor-Critic with Reward-Preserving Exploration"],"prefix":"10.1007","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1688-036X","authenticated-orcid":false,"given":"Athanasios Ch.","family":"Kapoutsis","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dimitrios I.","family":"Koutras","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christos D.","family":"Korkas","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Elias B.","family":"Kosmatopoulos","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,8,14]]},"reference":[{"key":"8845_CR1","unstructured":"Badia AP, Sprechmann P, Vitvitskyi A et\u00a0al. (2019) Never give up: learning directed exploration strategies. In: International conference on learning representations"},{"key":"8845_CR2","unstructured":"Badia AP, Piot B, Kapturowski S et\u00a0al. (2020) Agent57: outperforming the atari human benchmark. In: International conference on machine learning. PMLR, pp 507\u2013517"},{"key":"8845_CR3","unstructured":"Barth-Maron G, Hoffman MW, Budden D et\u00a0al. (2018) Distributed distributional deterministic policy gradients. In: International conference on learning representations"},{"key":"8845_CR4","doi-asserted-by":"publisher","first-page":"253","DOI":"10.1613\/jair.3912","volume":"47","author":"MG Bellemare","year":"2013","unstructured":"Bellemare MG, Naddaf Y, Veness J et al (2013) The arcade learning environment: an evaluation platform for general agents. J Artif Intell Res 47:253\u2013279","journal-title":"J Artif Intell Res"},{"key":"8845_CR5","unstructured":"Brockman G, Cheung V, Pettersson L et\u00a0al. (2016) Openai gym. arXiv:1606.01540"},{"key":"8845_CR6","unstructured":"Burda Y, Edwards H, Pathak D et\u00a0al. (2018) Large-scale study of curiosity-driven learning. In: International conference on learning representations"},{"key":"8845_CR7","unstructured":"Burda Y, Edwards H, Storkey A et\u00a0al. (2019) Exploration by random network distillation. In: Seventh international conference on learning representations, pp 1\u201317"},{"key":"8845_CR8","doi-asserted-by":"publisher","first-page":"2419","DOI":"10.1007\/s10994-021-05961-4","volume":"110","author":"G Dulac-Arnold","year":"2021","unstructured":"Dulac-Arnold G, Levine N, Mankowitz DJ et al (2021) Challenges of real-world reinforcement learning: definitions, benchmarks and analysis. Mach Learn 110:2419\u20132468","journal-title":"Mach Learn"},{"key":"8845_CR9","unstructured":"Fujimoto S, Hoof H, Meger D (2018) Addressing function approximation error in actor-critic methods. In: International conference on machine learning. PMLR, pp 1587\u20131596"},{"key":"8845_CR10","doi-asserted-by":"crossref","unstructured":"Gu S, Holly E, Lillicrap T et\u00a0al. (2017) Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In: 2017 IEEE international conference on robotics and automation (ICRA). IEEE, pp 3389\u20133396","DOI":"10.1109\/ICRA.2017.7989385"},{"key":"8845_CR11","unstructured":"Haarnoja T, Tang H, Abbeel P et\u00a0al. (2017) Reinforcement learning with deep energy-based policies. In: International conference on machine learning. PMLR, pp 1352\u20131361"},{"key":"8845_CR12","unstructured":"Haarnoja T, Zhou A, Abbeel P et\u00a0al. (2018) Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor. In: International conference on machine learning. PMLR, pp 1861\u20131870"},{"issue":"2","key":"8845_CR13","doi-asserted-by":"publisher","first-page":"1312","DOI":"10.1109\/LRA.2021.3057023","volume":"6","author":"G Kahn","year":"2021","unstructured":"Kahn G, Abbeel P, Levine S (2021) Badgr: an autonomous self-supervised learning-based navigation system. IEEE Robot Autom Lett 6(2):1312\u20131319","journal-title":"IEEE Robot Autom Lett"},{"issue":"4","key":"8845_CR14","doi-asserted-by":"publisher","first-page":"411","DOI":"10.3233\/ICA-220690","volume":"29","author":"GD Karatzinis","year":"2022","unstructured":"Karatzinis GD, Michailidis P, Michailidis IT et\u00a0al (2022) Coordinating heterogeneous mobile sensing platforms for effectively monitoring a dispersed gas plume. Integr Comput-Aided Eng 29(4):411\u2013429. https:\/\/doi.org\/10.3233\/ICA-220690","journal-title":"Integr Comput-Aided Eng"},{"key":"8845_CR15","doi-asserted-by":"crossref","unstructured":"Khorasgani H, Wang H, Gupta C et\u00a0al. (2021) Deep reinforcement learning with adjustments. In: 2021 IEEE 19th international conference on industrial informatics (INDIN). IEEE, pp 1\u20138","DOI":"10.1109\/INDIN45523.2021.9557543"},{"key":"8845_CR16","unstructured":"Kingma DP, Welling M (2014) Auto-encoding variational bayes. In: International conference on learning representations"},{"key":"8845_CR17","unstructured":"Lillicrap TP, Hunt JJ, Pritzel A et\u00a0al. (2016) Continuous control with deep reinforcement learning. In: ICLR (Poster)"},{"key":"8845_CR18","first-page":"2579","volume":"9","author":"L Van der Maaten","year":"2008","unstructured":"Van\u00a0der Maaten L, Hinton G (2008) Visualizing data using t-sne. J Mach Learn Res 9:2579\u20132605","journal-title":"J Mach Learn Res"},{"key":"8845_CR19","unstructured":"Mnih V, Kavukcuoglu K, Silver D et\u00a0al. (2013) Playing atari with deep reinforcement learning. In: NIPS deep learning workshop"},{"issue":"7540","key":"8845_CR20","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih V, Kavukcuoglu K, Silver D et al (2015) Human-level control through deep reinforcement learning. Nature 518(7540):529\u2013533","journal-title":"Nature"},{"key":"8845_CR21","unstructured":"Mnih V, Badia AP, Mirza M et\u00a0al. (2016) Asynchronous methods for deep reinforcement learning. In: International conference on machine learning. PMLR, pp 1928\u20131937"},{"key":"8845_CR22","unstructured":"Nikishin E, Schwarzer M, D\u2019Oro P et\u00a0al. (2022) The primacy bias in deep reinforcement learning. In: International conference on machine learning. PMLR, pp 16,828\u201316,847"},{"key":"8845_CR23","first-page":"35","volume":"22","author":"M Ornik","year":"2021","unstructured":"Ornik M, Topcu U (2021) Learning and planning for time-varying mdps using maximum likelihood estimation. J Mach Learn Res 22:35\u20131","journal-title":"J Mach Learn Res"},{"key":"8845_CR24","unstructured":"Ostrovski G, Bellemare MG, Oord A et\u00a0al. (2017) Count-based exploration with neural density models. In: International conference on machine learning. PMLR, pp 2721\u20132730"},{"key":"8845_CR25","first-page":"27730","volume":"35","author":"L Ouyang","year":"2022","unstructured":"Ouyang L, Wu J, Jiang X et al (2022) Training language models to follow instructions with human feedback. Adv Neural Inf Process Syst 35:27730\u201327744","journal-title":"Adv Neural Inf Process Syst"},{"key":"8845_CR26","unstructured":"Pardo F (2020) Tonic: a deep reinforcement learning library for fast prototyping and benchmarking. arXiv:2011.07537"},{"key":"8845_CR27","doi-asserted-by":"crossref","unstructured":"Pathak D, Agrawal P, Efros AA et\u00a0al. (2017) Curiosity-driven exploration by self-supervised prediction. In: International conference on machine learning. PMLR, pp 2778\u20132787","DOI":"10.1109\/CVPRW.2017.70"},{"key":"8845_CR28","unstructured":"Pathak D, Gandhi D, Gupta A (2019) Self-supervised exploration via disagreement. In: International conference on machine learning. PMLR, pp 5062\u20135071"},{"key":"8845_CR29","unstructured":"Schulman J, Levine S, Abbeel P et\u00a0al. (2015) Trust region policy optimization. In: International conference on machine learning. PMLR, pp 1889\u20131897"},{"key":"8845_CR30","unstructured":"Schulman J, Wolski F, Dhariwal P et\u00a0al. (2017) Proximal policy optimization algorithms. arXiv:1707.06347"},{"issue":"7587","key":"8845_CR31","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver D, Huang A, Maddison CJ et al (2016) Mastering the game of go with deep neural networks and tree search. Nature 529(7587):484\u2013489","journal-title":"Nature"},{"key":"8845_CR32","doi-asserted-by":"crossref","unstructured":"Smith L, Kew JC, Peng XB et\u00a0al. (2021) Legged robots that keep on learning: fine-tuning locomotion policies in the real world. arXiv:2110.05457","DOI":"10.1109\/ICRA46639.2022.9812166"},{"key":"8845_CR33","doi-asserted-by":"crossref","unstructured":"Smith L, Kostrikov I, Levine S (2022) A walk in the park: learning to walk in 20 minutes with model-free reinforcement learning. arXiv:2208.07860","DOI":"10.15607\/RSS.2023.XIX.056"},{"key":"8845_CR34","volume-title":"Reinforcement learning: an introduction","author":"RS Sutton","year":"2018","unstructured":"Sutton RS, Barto AG (2018) Reinforcement learning: an introduction. MIT Press, Cambridge"},{"key":"8845_CR35","unstructured":"Tang H, Houthooft R, Foote D et\u00a0al. (2017) # exploration: a study of count-based exploration for deep reinforcement learning. In: 31st conference on neural information processing systems (NIPS), pp 1\u201318"},{"key":"8845_CR36","unstructured":"Tassa Y, Doron Y, Muldal A et\u00a0al. (2018) Deepmind control suite. arXiv:1801.00690"},{"key":"8845_CR37","unstructured":"Thrun S, Schwartz A (1993) Issues in using function approximation for reinforcement learning. In: Proceedings of the fourth connectionist models summer school, Hillsdale, NJ, pp 255\u2013263"},{"key":"8845_CR38","doi-asserted-by":"crossref","unstructured":"Todorov E, Erez T, Tassa Y (2012) Mujoco: a physics engine for model-based control. In: 2012 IEEE\/RSJ international conference on intelligent robots and systems. IEEE, pp 5026\u20135033","DOI":"10.1109\/IROS.2012.6386109"},{"key":"8845_CR39","unstructured":"Wu P, Escontrela A, Hafner D et\u00a0al. (2023) Daydreamer: world models for physical robot learning. In: Conference on robot learning. PMLR, pp 2226\u20132240"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-023-08845-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-023-08845-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-023-08845-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,16]],"date-time":"2023-09-16T11:09:58Z","timestamp":1694862598000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-023-08845-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,14]]},"references-count":39,"journal-issue":{"issue":"30","published-print":{"date-parts":[[2023,10]]}},"alternative-id":["8845"],"URL":"https:\/\/doi.org\/10.1007\/s00521-023-08845-x","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,14]]},"assertion":[{"value":"27 January 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 June 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 August 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"On behalf of all authors, the corresponding author states that there is no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}