{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T15:22:15Z","timestamp":1784301735264,"version":"3.55.0"},"reference-count":45,"publisher":"MDPI AG","issue":"15","license":[{"start":{"date-parts":[[2024,8,1]],"date-time":"2024-08-01T00:00:00Z","timestamp":1722470400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"European Union","award":["FSE-REACT-EU"],"award-info":[{"award-number":["FSE-REACT-EU"]}]},{"name":"European Union","award":["2014-2020 DM1062\/2021"],"award-info":[{"award-number":["2014-2020 DM1062\/2021"]}]},{"name":"PON Research and Innovation","award":["FSE-REACT-EU"],"award-info":[{"award-number":["FSE-REACT-EU"]}]},{"name":"PON Research and Innovation","award":["2014-2020 DM1062\/2021"],"award-info":[{"award-number":["2014-2020 DM1062\/2021"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>We consider a complex control problem: making a monopod accurately reach a target with a single jump. The monopod can jump in any direction at different elevations of the terrain. This is a paradigm for a much larger class of problems, which are extremely challenging and computationally expensive to solve using standard optimization-based techniques. Reinforcement learning (RL) is an interesting alternative, but an end-to-end approach in which the controller must learn everything from scratch can be non-trivial with a sparse-reward task like jumping. Our solution is to guide the learning process within an RL framework leveraging nature-inspired heuristic knowledge. This expedient brings widespread benefits, such as a drastic reduction of learning time, and the ability to learn and compensate for possible errors in the low-level execution of the motion. Our simulation results reveal a clear advantage of our solution against both optimization-based and end-to-end RL approaches.<\/jats:p>","DOI":"10.3390\/s24154981","type":"journal-article","created":{"date-parts":[[2024,8,1]],"date-time":"2024-08-01T11:34:20Z","timestamp":1722512060000},"page":"4981","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Efficient Reinforcement Learning for 3D Jumping Monopods"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-8181-1815","authenticated-orcid":false,"given":"Riccardo","family":"Bussola","sequence":"first","affiliation":[{"name":"Dipartimento di Ingegneria and Scienza Dell\u2019Informazione (DISI), University of Trento, 38123 Trento, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4888-5595","authenticated-orcid":false,"given":"Michele","family":"Focchi","sequence":"additional","affiliation":[{"name":"Dipartimento di Ingegneria and Scienza Dell\u2019Informazione (DISI), University of Trento, 38123 Trento, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1275-2851","authenticated-orcid":false,"given":"Andrea","family":"Del Prete","sequence":"additional","affiliation":[{"name":"Dipartimento di Ingegneria Industriale (DII), University of Trento, 38123 Trento, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5486-9989","authenticated-orcid":false,"given":"Daniele","family":"Fontanelli","sequence":"additional","affiliation":[{"name":"Dipartimento di Ingegneria Industriale (DII), University of Trento, 38123 Trento, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2393-479X","authenticated-orcid":false,"given":"Luigi","family":"Palopoli","sequence":"additional","affiliation":[{"name":"Dipartimento di Ingegneria and Scienza Dell\u2019Informazione (DISI), University of Trento, 38123 Trento, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,8,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"3395","DOI":"10.1109\/TRO.2022.3186804","article-title":"TAMOLS: Terrain-Aware Motion Optimization for Legged Systems","volume":"38","author":"Jenelten","year":"2022","journal-title":"IEEE Trans. Robot."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"7210","DOI":"10.1109\/LRA.2023.3313919","article-title":"Reactive Landing Controller for Quadruped Robots","volume":"8","author":"Roscia","year":"2023","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1177\/0278364917694244","article-title":"High-speed bounding with the MIT Cheetah 2: Control design and experiments","volume":"36","author":"Park","year":"2017","journal-title":"Int. J. Robot. Res."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"3422","DOI":"10.1109\/LRA.2020.2976597","article-title":"Precision Robotic Leaping and Landing Using Stance-Phase Balance","volume":"5","author":"Yim","year":"2020","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Nguyen, C., and Nguyen, Q. (2022, January 23\u201327). Contact-timing and Trajectory Optimization for 3D Jumping on Quadruped Robots. Proceedings of the 2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Kyoto, Japan.","DOI":"10.1109\/IROS47612.2022.9981284"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Chignoli, M., and Kim, S. (June, January 30). Online Trajectory Optimization for Dynamic Aerial Motions of a Quadruped Robot. Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi\u2019an, China.","DOI":"10.1109\/ICRA48506.2021.9560855"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Garc\u00eda, G., Griffin, R., and Pratt, J. (June, January 30). Time-Varying Model Predictive Control for Highly Dynamic Motions of Quadrupedal Robots. Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi\u2019an, China.","DOI":"10.1109\/ICRA48506.2021.9561913"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Chignoli, M., Morozov, S., and Kim, S. (2022, January 23\u201327). Rapid and Reliable Quadruped Motion Planning with Omnidirectional Jumping. Proceedings of the 2022 International Conference on Robotics and Automation (ICRA), Philadelphia, PA, USA.","DOI":"10.1109\/ICRA46639.2022.9812088"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Mastalli, C., Merkt, W., Xin, G., Shim, J., Mistry, M., Havoutis, I., and Vijayakumar, S. (2022). Agile Maneuvers in Legged Robots:a Predictive Control Approach. arXiv.","DOI":"10.21203\/rs.3.rs-1870369\/v1"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Li, H., and Wensing, P.M. (2024). Cafe-Mpc: A Cascaded-Fidelity Model Predictive Control Framework with Tuning-Free Whole-Body Control. arXiv.","DOI":"10.1109\/TRO.2024.3504132"},{"key":"ref_11","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N.M.O., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1109\/MRA.2015.2505910","article-title":"Practice Makes Perfect: An Optimization-Based Approach to Controlling Agile Motions for a Quadruped Robot","volume":"23","author":"Gehring","year":"2016","journal-title":"IEEE Robot. Autom. Mag."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"eaau5872","DOI":"10.1126\/scirobotics.aau5872","article-title":"Learning agile and dynamic motor skills for legged robots","volume":"4","author":"Hwangbo","year":"2019","journal-title":"Sci. Robot."},{"key":"ref_14","unstructured":"Peng, X., Coumans, E., Zhang, T., Lee, T.W., Tan, J., and Levine, S. (2020, January 12\u201316). Learning Agile Robotic Locomotion Skills by Imitating Animals. Proceedings of the Robotics: Science and Systems 2020, Corvalis, OR, USA."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"4630","DOI":"10.1109\/LRA.2022.3151396","article-title":"Concurrent Training of a Control Policy and a State Estimator for Dynamic and Robust Legged Locomotion","volume":"7","author":"Ji","year":"2022","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_16","unstructured":"Rudin, N., Hoeller, D., Reist, P., and Hutter, M. (2022, January 14\u201318). Learning to walk in minutes using massively parallel deep reinforcement learning. Proceedings of the Conference on Robot Learning. PMLR, Auckland, New Zealand."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Fankhauser, P., Hutter, M., Gehring, C., Bloesch, M., Hoepflinger, M.A., and Siegwart, R. (2013, January 3\u20137). Reinforcement learning of single legged locomotion. Proceedings of the 2013 IEEE\/RSJ International Conference on Intelligent Robots and Systems, Tokyo, Japan.","DOI":"10.1109\/IROS.2013.6696352"},{"key":"ref_18","unstructured":"OpenAI (2023, February 26). Benchmarks for Spinning Up Implementations. Available online: https:\/\/spinningup.openai.com\/en\/latest\/spinningup\/bench.html#benchmarks-for-spinning-up-implementations."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Bogdanovic, M., Khadiv, M., and Righetti, L. (2022). Model-free reinforcement learning for robust locomotion using demonstrations from trajectory optimization. Front. Robot. AI, 9.","DOI":"10.3389\/frobt.2022.854212"},{"key":"ref_20","unstructured":"Bellegarda, G., Nguyen, C., and Nguyen, Q. (2023). Robust Quadruped Jumping via Deep Reinforcement Learning. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"3318","DOI":"10.1109\/LRA.2023.3266985","article-title":"CACTO: Continuous Actor-Critic with Trajectory Optimization-Towards Global Optimality","volume":"8","author":"Grandesso","year":"2023","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Peng, X.B., and van de Panne, M. (2017, January 28\u201330). Learning Locomotion Skills Using DeepRL: Does the Choice of Action Space Matter?. Proceedings of the ACM SIGGRAPH\/Eurographics Symposium on Computer Animation, Los Angeles, CA, USA.","DOI":"10.1145\/3099564.3099567"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Bellegarda, G., and Byl, K. (2019, January 3\u20138). Training in Task Space to Speed Up and Guide Reinforcement Learning. Proceedings of the 2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China.","DOI":"10.1109\/IROS40897.2019.8967995"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Chen, S., Zhang, B., Mueller, M.W., Rai, A., and Sreenath, K. (2022). Learning Torque Control for Quadrupedal Locomotion. arXiv.","DOI":"10.1109\/Humanoids57100.2023.10375154"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Aractingi, M., L\u00e9ziart, P.A., Flayols, T., Perez, J., Silander, T., and Sou\u00e8res, P. (2023). Controlling the Solo12 Quadruped Robot with Deep Reinforcement Learning. Sci. Rep., 13.","DOI":"10.1038\/s41598-023-38259-7"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Majid, A.Y., Saaybi, S., van Rietbergen, T., Fran\u00e7ois-Lavet, V., Prasad, R.V., and Verhoeven, C. (2021). Deep Reinforcement Learning Versus Evolution Strategies: A Comparative Survey. arXiv.","DOI":"10.36227\/techrxiv.14679504.v2"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Atanassov, V., Ding, J., Kober, J., Havoutis, I., and Santina, C.D. (2024). Curriculum-Based Reinforcement Learning for Quadrupedal Jumping: A Reference-free Design. arXiv.","DOI":"10.1109\/MRA.2024.3487325"},{"key":"ref_28","unstructured":"Yang, Y., Meng, X., Yu, W., Zhang, T., Tan, J., and Boots, B. (2023, January 15\u201316). Continuous Versatile Jumping Using Learned Action Residuals. Proceedings of the Machine Learning Research PMLR, Philadelphia, PA, USA."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Vezzi, F., Ding, J., Raffin, A., Kober, J., and Della Santina, C. (2024, January 13\u201317). Two-Stage Learning of Highly Dynamic Motions with Rigid and Articulated Soft Quadrupeds. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan.","DOI":"10.1109\/ICRA57147.2024.10610561"},{"key":"ref_30","first-page":"10039","article-title":"Towards the systematic reporting of the energy and carbon footprints of machine learning","volume":"21","author":"Henderson","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref_31","first-page":"36","article-title":"A comparison of ppo, td3 and sac reinforcement algorithms for quadruped walking gait generation","volume":"15","author":"Mock","year":"2023","journal-title":"J. Intell. Learn. Syst. Appl."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Shafiee, M., Bellegarda, G., and Ijspeert, A. (2024). ManyQuadrupeds: Learning a Single Locomotion Policy for Diverse Quadruped Robots. arXiv.","DOI":"10.1109\/ICRA57147.2024.10610155"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"3770","DOI":"10.1038\/s41467-019-11786-6","article-title":"A critique of pure learning and what artificial neural networks can learn from animal brains","volume":"10","author":"Zador","year":"2019","journal-title":"Nat. Commun."},{"key":"ref_34","first-page":"66","article-title":"Learning Fast Quadruped Robot Gaits with the RL PoWER Spline Parameterization","volume":"12","author":"Shen","year":"2013","journal-title":"Cybern. Inf. Technol."},{"key":"ref_35","unstructured":"Kim, T., and Lee, S.H. (2021). Quadruped Locomotion on Non-Rigid Terrain using Reinforcement Learning. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Ji, Y., Li, Z., Sun, Y., Peng, X.B., Levine, S., Berseth, G., and Sreenath, K. (2022, January 23\u201327). Hierarchical Reinforcement Learning for Precise Soccer Shooting Skills using a Quadrupedal Robot. Proceedings of the 2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Kyoto, Japan.","DOI":"10.1109\/IROS47612.2022.9981984"},{"key":"ref_37","unstructured":"Grzes, M. (2017). Reward Shaping in Episodic Reinforcement Learning, ACM."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Focchi, M., Roscia, F., and Semini, C. (2024). Locosim: An Open-Source Cross-PlatformRobotics Framework. Synergetic Cooperation between Robots and Humans, Proceedings of the CLAWAR 2023, Florianopolis, Brazil, 2\u20134 October 2023, Springer. Lecture Notes in Networks and Systems.","DOI":"10.1007\/978-3-031-47272-5_33"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Budhiraja, R., Carpentier, J., Mastalli, C., and Mansard, N. (2018, January 6\u20139). Differential Dynamic Programming for Multi-Phase Rigid Contact Dynamics. Proceedings of the IEEE International Conference on Humanoid Robots, Beijing, China.","DOI":"10.1109\/HUMANOIDS.2018.8624925"},{"key":"ref_40","unstructured":"Mastalli, C., Budhiraja, R., Merkt, W., Saurel, G., Hammoud, B., Naveau, M., Carpentier, J., Righetti, L., Vijayakumar, S., and Mansard, N. (August, January 31). Crocoddyl: An Efficient and Versatile Framework for Multi-Contact Optimal Control. Proceedings of the IEEE International Conference on Robotics and Automation, Paris, France."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Carpentier, J., Saurel, G., Buondonno, G., Mirabel, J., Lamiraux, F., Stasse, O., and Mansard, N. (2019, January 14\u201316). The Pinocchio C++ library\u2014A fast and flexible implementation of rigid body dynamics algorithms and their analytical derivatives. Proceedings of the IEEE International Symposium on System Integrations (SII), Paris, France.","DOI":"10.1109\/SII.2019.8700380"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Gangapurwala, S., Campanaro, L., and Havoutis, I. (2022). Learning Low-Frequency Motion Control for Robust and Dynamic Robot Locomotion. arXiv.","DOI":"10.1109\/ICRA48891.2023.10160357"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Zhao, T.Z., Kumar, V., Levine, S., and Finn, C. (2023). Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv.","DOI":"10.15607\/RSS.2023.XIX.016"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Jeon, S.H., Heim, S., Khazoom, C., and Kim, S. (June, January 29). Benchmarking Potential Based Rewards for Learning Humanoid Locomotion. Proceedings of the 2023 IEEE International Conference on Robotics and Automation (ICRA), London, UK.","DOI":"10.1109\/ICRA48891.2023.10160885"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009, January 14\u201318). Curriculum learning. Proceedings of the 26th Annual International Conference on Machine Learning, Montreal, QC, Canada.","DOI":"10.1145\/1553374.1553380"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/15\/4981\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:28:00Z","timestamp":1760110080000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/15\/4981"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,1]]},"references-count":45,"journal-issue":{"issue":"15","published-online":{"date-parts":[[2024,8]]}},"alternative-id":["s24154981"],"URL":"https:\/\/doi.org\/10.3390\/s24154981","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,1]]}}}