{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,26]],"date-time":"2026-04-26T05:40:39Z","timestamp":1777182039613,"version":"3.51.4"},"reference-count":49,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2024,7,24]],"date-time":"2024-07-24T00:00:00Z","timestamp":1721779200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,7,24]],"date-time":"2024-07-24T00:00:00Z","timestamp":1721779200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003475","name":"Hasler Stiftung","doi-asserted-by":"publisher","award":["Robolutionary"],"award-info":[{"award-number":["Robolutionary"]}],"id":[{"id":"10.13039\/501100003475","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Intell Robot Syst"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Deep Reinforcement Learning applications are growing due to their capability of teaching the agent any task autonomously and generalizing the learning. However, this comes at the cost of a large number of samples and interactions with the environment. Moreover, the robustness of learned policies is usually achieved by a tedious tuning of hyper-parameters and reward functions. In order to address this issue, this paper proposes an evolutionary RL algorithm for the adaptive optimization of hyper-parameters. The policy is trained using an on-policy algorithm, Proximal Policy Optimization (PPO), coupled with an evolutionary algorithm. The achieved results demonstrate an improvement in the sample efficiency of the RL training on a robotic grasping task. In particular, the learning is improved with respect to the baseline case of a non-evolutionary agent. The evolutionary agent needs <jats:inline-formula><jats:alternatives><jats:tex-math>$$60$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mrow>\n                    <mml:mn>60<\/mml:mn>\n                  <\/mml:mrow>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>% fewer samples to completely learn the grasping task, enabled by the adaptive transfer of knowledge between the agents through the evolutionary algorithm. The proposed approach also demonstrates the possibility of updating reward parameters during training, potentially providing a general approach to creating reward functions.<\/jats:p>","DOI":"10.1007\/s10846-024-02138-8","type":"journal-article","created":{"date-parts":[[2024,7,24]],"date-time":"2024-07-24T14:03:11Z","timestamp":1721829791000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Adaptive Optimization of Hyper-Parameters for Robotic Manipulation through Evolutionary Reinforcement Learning"],"prefix":"10.1007","volume":"110","author":[{"given":"Giulio","family":"Onori","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Asad Ali","family":"Shahid","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Francesco","family":"Braghin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4427-536X","authenticated-orcid":false,"given":"Loris","family":"Roveda","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,7,24]]},"reference":[{"key":"2138_CR1","doi-asserted-by":"publisher","first-page":"181855","DOI":"10.1109\/ACCESS.2020.3028740","volume":"8","author":"Q Bai","year":"2020","unstructured":"Bai, Q., Li, S., Yang, J., Song, Q., Li, Z., Zhang, X.: Object detection recognition and robot grasping based on machine learning: A survey. IEEE Access 8, 181855\u2013181879 (2020)","journal-title":"IEEE Access"},{"key":"2138_CR2","doi-asserted-by":"crossref","unstructured":"Semeraro, F., Griffiths, A., Cangelosi, A.: Human\u2013robot collaboration and machine learning: A systematic review of recent research. Robot. Comput.-Integrated Manufac. 79, 102432 (2023)","DOI":"10.1016\/j.rcim.2022.102432"},{"issue":"21","key":"2138_CR3","doi-asserted-by":"publisher","first-page":"15429","DOI":"10.1007\/s00521-023-08361-y","volume":"35","author":"X Song","year":"2023","unstructured":"Song, X., Sun, P., Song, S., Stojanovic, V.: Quantized neural adaptive finite-time preassigned performance control for interconnected nonlinear systems. Neural Comput. Appl. 35(21), 15429\u201315446 (2023)","journal-title":"Neural Comput. Appl."},{"key":"2138_CR4","doi-asserted-by":"publisher","DOI":"10.1016\/j.jprocont.2023.103112","volume":"132","author":"H Tao","year":"2023","unstructured":"Tao, H., Zheng, J., Wei, J., Paszke, W., Rogers, E., Stojanovic, V.: Repetitive process based indirect-type iterative learning control for batch processes with model uncertainty and input delay. J. Process Control 132, 103112 (2023)","journal-title":"J. Process Control"},{"key":"2138_CR5","doi-asserted-by":"crossref","unstructured":"Billard, A.G., Calinon, S., Dillmann, R.: Learning from humans. Springer handbook of robotics, 1995\u20132014 (2016)","DOI":"10.1007\/978-3-319-32552-1_74"},{"key":"2138_CR6","doi-asserted-by":"publisher","DOI":"10.1016\/j.oceaneng.2023.113937","volume":"273","author":"R Deraj","year":"2023","unstructured":"Deraj, R., Kumar, R.S., Alam, M.S., Somayajula, A.: Deep reinforcement learning based controller for ship navigation. Ocean Eng. 273, 113937 (2023)","journal-title":"Ocean Eng."},{"issue":"11","key":"2138_CR7","doi-asserted-by":"publisher","first-page":"1238","DOI":"10.1177\/0278364913495721","volume":"32","author":"J Kober","year":"2013","unstructured":"Kober, J., Bagnell, J.A., Peters, J.: Reinforcement learning in robotics: A survey. Int. J. Robot. Res. 32(11), 1238\u20131274 (2013)","journal-title":"Int. J. Robot. Res."},{"key":"2138_CR8","doi-asserted-by":"crossref","unstructured":"Yarats, D., Zhang, A., Kostrikov, I., Amos, B., Pineau, J., Fergus, R.: Improving sample efficiency in model-free reinforcement learning from images. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 10674\u201310681 (2021)","DOI":"10.1609\/aaai.v35i12.17276"},{"key":"2138_CR9","unstructured":"Grze\u015b, M.: Reward shaping in episodic reinforcement learning. In: Proceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems, pp. 565\u2013573 (2017)"},{"key":"2138_CR10","doi-asserted-by":"crossref","unstructured":"Shahid, A.A., Roveda, L., Piga, D., Braghin, F.: Learning continuous control actions for robotic grasping with reinforcement learning. In: 2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC), pp. 4066\u20134072 (2020)","DOI":"10.1109\/SMC42975.2020.9282951"},{"issue":"02","key":"2138_CR11","first-page":"57","volume":"2","author":"MA Wiering","year":"2010","unstructured":"Wiering, M.A., et al.: Self-play and using an expert to learn to play backgammon with temporal difference learning. J. Intell. Learn. Syst. Appl. 2(02), 57 (2010)","journal-title":"J. Intell. Learn. Syst. Appl."},{"issue":"7587","key":"2138_CR12","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver, D., Huang, A., Maddison, C.J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al.: Mastering the game of go with deep neural networks and tree search. Nature 529(7587), 484\u2013489 (2016)","journal-title":"Nature"},{"issue":"7540","key":"2138_CR13","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., Graves, A., Riedmiller, M., Fidjeland, A.K., Ostrovski, G., et al.: Human-level control through deep reinforcement learning. Nature 518(7540), 529\u2013533 (2015)","journal-title":"Nature"},{"key":"2138_CR14","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1016\/j.isatra.2023.07.043","volume":"142","author":"R Wang","year":"2023","unstructured":"Wang, R., Zhuang, Z., Tao, H., Paszke, W., Stojanovic, V.: Q-learning based fault estimation and fault tolerant iterative learning control for mimo systems. ISA Trans. 142, 123\u2013135 (2023)","journal-title":"ISA Trans."},{"issue":"8","key":"2138_CR15","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3467477","volume":"54","author":"A Telikani","year":"2021","unstructured":"Telikani, A., Tahmassebi, A., Banzhaf, W., Gandomi, A.H.: Evolutionary machine learning: A survey. ACM Comput. Surv. (CSUR) 54(8), 1\u201335 (2021)","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"2138_CR16","doi-asserted-by":"crossref","unstructured":"Liang, J., Meyerson, E., Hodjat, B., Fink, D., Mutch, K., Miikkulainen, R.: Evolutionary neural automl for deep learning. In: Proceedings of the Genetic and Evolutionary Computation Conference, pp. 401\u2013409 (2019)","DOI":"10.1145\/3321707.3321721"},{"key":"2138_CR17","doi-asserted-by":"publisher","first-page":"4","DOI":"10.3389\/frobt.2015.00004","volume":"2","author":"S Doncieux","year":"2015","unstructured":"Doncieux, S., Bredeche, N., Mouret, J.-B., Eiben, A.E.: Evolutionary robotics: what, why, and where to. Front. Robot. AI 2, 4 (2015)","journal-title":"Front. Robot. AI"},{"key":"2138_CR18","doi-asserted-by":"crossref","unstructured":"Paul, C., Bongard, J.C.: The road less travelled: Morphology in the optimization of biped robot locomotion. In: Proceedings 2001 IEEE\/RSJ International Conference on Intelligent Robots and Systems., vol. 1, pp. 226\u2013232 (2001)","DOI":"10.1109\/IROS.2001.973363"},{"key":"2138_CR19","unstructured":"Zykov, V., Bongard, J., Lipson, H.: Evolving dynamic gaits on a physical robot. In: Proceedings of Genetic and Evolutionary Computation Conference (2004)"},{"key":"2138_CR20","doi-asserted-by":"crossref","unstructured":"Lund, H.H.: Co-evolving control and morphology with lego robots. In: Morpho-functional Machines: the New Species: Designing Embodied Intelligence, pp. 59\u201379 (2003)","DOI":"10.1007\/978-4-431-67869-4_4"},{"issue":"4","key":"2138_CR21","doi-asserted-by":"publisher","first-page":"353","DOI":"10.1162\/artl.1994.1.4.353","volume":"1","author":"K Sims","year":"1994","unstructured":"Sims, K.: Evolving 3d morphology and behavior by competition. Artif. Life 1(4), 353\u2013372 (1994)","journal-title":"Artif. Life"},{"key":"2138_CR22","doi-asserted-by":"crossref","unstructured":"Fang, Z., Liang, X.: Intelligent obstacle avoidance path planning method for picking manipulator combined with artificial potential field method. Indust. Robot: Int. J. Robot. Res. Appl. 49(5), 835\u2013850 (2022)","DOI":"10.1108\/IR-09-2021-0194"},{"key":"2138_CR23","doi-asserted-by":"crossref","unstructured":"Larsen, L., Kim, J.: Path planning of cooperating industrial robots using evolutionary algorithms. Robot. Comput.-Integrated Manufac. 67, 102053 (2021)","DOI":"10.1016\/j.rcim.2020.102053"},{"key":"2138_CR24","unstructured":"Lin, H.-S., Xiao, J., Michalewicz, Z.: Evolutionary algorithm for path planning in mobile robot environment. In: Proceedings of the First IEEE Conference on Evolutionary Computation, pp. 211\u2013216 (1994)"},{"issue":"3","key":"2138_CR25","doi-asserted-by":"publisher","first-page":"1155","DOI":"10.1109\/TCDS.2021.3098229","volume":"14","author":"R Wu","year":"2021","unstructured":"Wu, R., Chao, F., Zhou, C., Huang, Y., Yang, L., Lin, C.-M., Chang, X., Shen, Q., Shang, C.: A developmental evolutionary learning framework for robotic chinese stroke writing. IEEE Trans. Cognit. Develop. Syst. 14(3), 1155\u20131169 (2021)","journal-title":"IEEE Trans. Cognit. Develop. Syst."},{"key":"2138_CR26","doi-asserted-by":"publisher","unstructured":"Mouret, J., Clune, J.: Illuminating search spaces by mapping elites. CoRR (2015) https:\/\/doi.org\/10.48550\/ARXIV.1504.04909","DOI":"10.48550\/ARXIV.1504.04909"},{"key":"2138_CR27","doi-asserted-by":"crossref","unstructured":"Lehman, J., Stanley, K.O.: Evolving a diversity of virtual creatures through novelty search and local competition. In: Proceedings of the 13th Annual Conference on Genetic and Evolutionary Computation, pp. 211\u2013218 (2011)","DOI":"10.1145\/2001576.2001606"},{"key":"2138_CR28","doi-asserted-by":"publisher","first-page":"151","DOI":"10.3389\/frobt.2019.00151","volume":"6","author":"R Kaushik","year":"2020","unstructured":"Kaushik, R., Desreumaux, P., Mouret, J.-B.: Adaptive prior selection for repertoire-based online adaptation in robotics. Front. Robot. AI 6, 151 (2020)","journal-title":"Front. Robot. AI"},{"key":"2138_CR29","doi-asserted-by":"crossref","unstructured":"Nordmoen, J., Ellefsen, K.O., Glette, K.: Combining map-elites and incremental evolution to generate gaits for a mammalian quadruped robot. In: Applications of Evolutionary Computation: 21st International Conference, EvoApplications 2018, Parma, Italy, April 4-6, 2018, Proceedings 21, pp. 719\u2013733 (2018)","DOI":"10.1007\/978-3-319-77538-8_48"},{"key":"2138_CR30","doi-asserted-by":"crossref","unstructured":"Bossens, D.M., Mouret, J.-B., Tarapore, D.: Learning behaviour-performance maps with meta-evolution. In: Proceedings of the 2020 Genetic and Evolutionary Computation Conference, pp. 49\u201357 (2020)","DOI":"10.1145\/3377930.3390181"},{"key":"2138_CR31","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2020.103710","volume":"136","author":"S Kim","year":"2021","unstructured":"Kim, S., Coninx, A., Doncieux, S.: From exploration to control: learning object manipulation skills through novelty search and local adaptation. Robot. Autonomous Syst. 136, 103710 (2021)","journal-title":"Robot. Autonomous Syst."},{"key":"2138_CR32","doi-asserted-by":"crossref","unstructured":"Morel, A., Kunimoto, Y., Coninx, A., Doncieux, S.: Automatic acquisition of a repertoire of diverse grasping trajectories through behavior shaping and novelty search. In: 2022 International Conference on Robotics and Automation (ICRA), pp. 755\u2013761 (2022)","DOI":"10.1109\/ICRA46639.2022.9811837"},{"key":"2138_CR33","unstructured":"Khadka, S., Majumdar, S., Nassar, T., Dwiel, Z., Tumer, E., Miret, S., Liu, Y., Tumer, K.: Collaborative evolutionary reinforcement learning. In: International Conference on Machine Learning, pp. 3341\u20133350 (2019)"},{"key":"2138_CR34","doi-asserted-by":"crossref","unstructured":"Ma, Y., Liu, T., Wei, B., Liu, Y., Xu, K., Li, W.: Evolutionary action selection for gradient-based policy learning. In: International Conference on Neural Information Processing, pp. 579\u2013590 (2022)","DOI":"10.1007\/978-3-031-30111-7_49"},{"key":"2138_CR35","unstructured":"Marchesini, E., Corsi, D., Farinelli, A.: Genetic soft updates for policy evolution in deep reinforcement learning. In: International Conference on Learning Representations (2021)"},{"key":"2138_CR36","doi-asserted-by":"crossref","unstructured":"Bodnar, C., Day, B., Li\u00f3, P.: Proximal distilled evolutionary reinforcement learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, pp. 3283\u20133290 (2020)","DOI":"10.1609\/aaai.v34i04.5728"},{"key":"2138_CR37","unstructured":"Parker-Holder, J., Nguyen, V., Roberts, S.J.: Provably efficient online hyperparameter optimization with population-based bandits. Adv Neural Inf Process. Syst. 33, 17200\u201317211 (2020)"},{"key":"2138_CR38","doi-asserted-by":"publisher","unstructured":"Afshar, R.R., Zhang, Y., Vanschoren, J., Kaymak, U.: Automated reinforcement learning: An overview. CoRR (2022) https:\/\/doi.org\/10.48550\/ARXIV.2201.05000","DOI":"10.48550\/ARXIV.2201.05000"},{"key":"2138_CR39","doi-asserted-by":"publisher","unstructured":"Sehgal, A., Ward, N., La, H.M., Louis, S.J.: Automatic parameter optimization using genetic algorithm in deep reinforcement learning for robotic manipulation tasks. CoRR (2022) https:\/\/doi.org\/10.48550\/ARXIV.2204.03656","DOI":"10.48550\/ARXIV.2204.03656"},{"key":"2138_CR40","doi-asserted-by":"crossref","unstructured":"Shahid, A.A., Piga, D., Braghin, F., Roveda, L.: Continuous control actions learning and adaptation for robotic manipulation through reinforcement learning. Autonomous Robot. 46(3), 483\u2013498 (2022)","DOI":"10.1007\/s10514-022-10034-z"},{"key":"2138_CR41","doi-asserted-by":"publisher","unstructured":"Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W.M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., Fernando, C., Kavukcuoglu, K.: Population based training of neural networks. CoRR (2017) https:\/\/doi.org\/10.48550\/ARXIV.1711.09846","DOI":"10.48550\/ARXIV.1711.09846"},{"key":"2138_CR42","doi-asserted-by":"crossref","unstructured":"Todorov, E., Erez, T., Tassa, Y.: Mujoco: A physics engine for model-based control. In: 2012 IEEE\/RSJ International Conference on Intelligent Robots and Systems, pp. 5026\u20135033 (2012)","DOI":"10.1109\/IROS.2012.6386109"},{"key":"2138_CR43","unstructured":"Fan, L., Zhu, Y., Zhu, J., Liu, Z., Zeng, O., Gupta, A., Creus-Costa, J., Savarese, S., Fei-Fei, L.: Surreal: Open-source reinforcement learning framework and robot manipulation benchmark. In: Conference on Robot Learning, pp. 767\u2013782 (2018)"},{"issue":"268","key":"2138_CR44","first-page":"1","volume":"22","author":"A Raffin","year":"2021","unstructured":"Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., Dormann, N.: Stable-baselines3: Reliable reinforcement learning implementations. J. Mach. Learn. Res. 22(268), 1\u20138 (2021)","journal-title":"J. Mach. Learn. Res."},{"key":"2138_CR45","doi-asserted-by":"publisher","unstructured":"Shahid, A.A., Narang, Y.S., Petrone, V., Ferrentino, E., Handa, A., Fox, D., Pavone, M., Roveda, L.: Scaling population-based reinforcement learning with GPU accelerated simulation. CoRR (2024) https:\/\/doi.org\/10.48550\/ARXIV.2404.03336","DOI":"10.48550\/ARXIV.2404.03336"},{"key":"2138_CR46","doi-asserted-by":"publisher","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., Klimov, O.: Proximal policy optimization algorithms. CoRR (2017) https:\/\/doi.org\/10.48550\/ARXIV.1707.06347","DOI":"10.48550\/ARXIV.1707.06347"},{"key":"2138_CR47","unstructured":"Schulman, J., Levine, S., Abbeel, P., Jordan, M., Moritz, P.: Trust region policy optimization. In: International Conference on Machine Learning, pp. 1889\u20131897 (2015)"},{"key":"2138_CR48","unstructured":"Schulman, J., Moritz, P., Levine, S., Jordan, M.I., Abbeel, P.: High-dimensional continuous control using generalized advantage estimation. In: Bengio, Y., LeCun, Y. (eds.) 4th International Conference on Learning Representations, ICLR (2016)"},{"key":"2138_CR49","doi-asserted-by":"crossref","unstructured":"Eiben, A.E., Schoenauer, M.: Evolutionary computing. Inf. Process. Lett. 82(1), 1\u20136 (2002)","DOI":"10.1016\/S0020-0190(02)00204-1"}],"container-title":["Journal of Intelligent &amp; Robotic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10846-024-02138-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10846-024-02138-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10846-024-02138-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,26]],"date-time":"2024-09-26T13:12:15Z","timestamp":1727356335000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10846-024-02138-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,24]]},"references-count":49,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2024,9]]}},"alternative-id":["2138"],"URL":"https:\/\/doi.org\/10.1007\/s10846-024-02138-8","relation":{},"ISSN":["1573-0409"],"issn-type":[{"value":"1573-0409","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,24]]},"assertion":[{"value":"7 November 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 July 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 July 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"On behalf of all authors, the corresponding author states that there is no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest\/Competing interests"}},{"value":"No ethical approval is required for this research.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval"}},{"value":"Not applicable","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"This work does not require any consent for publication","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}],"article-number":"108"}}