{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,25]],"date-time":"2026-07-25T19:13:14Z","timestamp":1785006794672,"version":"3.55.0"},"reference-count":214,"publisher":"Annual Reviews","issue":"1","license":[{"start":{"date-parts":[[2025,5,5]],"date-time":"2025-05-05T00:00:00Z","timestamp":1746403200000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,5,5]]},"abstract":"<jats:p>Reinforcement learning (RL), particularly its combination with deep neural networks, referred to as deep RL (DRL), has shown tremendous promise across a wide range of applications, suggesting its potential for enabling the development of sophisticated robotic behaviors. Robotics problems, however, pose fundamental difficulties for the application of RL, stemming from the complexity and cost of interacting with the physical world. This article provides a modern survey of DRL for robotics, with a particular focus on evaluating the real-world successes achieved with DRL in realizing several key robotic competencies. Our analysis aims to identify the key factors underlying those exciting successes, reveal underexplored areas, and provide an overall characterization of the status of DRL in robotics. We highlight several important avenues for future work, emphasizing the need for stable and sample-efficient real-world RL paradigms; holistic approaches for discovering and integrating various competencies to tackle complex long-horizon, open-world tasks; and principled development and evaluation procedures. This survey is designed to offer insights for both RL practitioners and roboticists toward harnessing RL's power to create generally capable real-world robotic systems.<\/jats:p>","DOI":"10.1146\/annurev-control-030323-022510","type":"journal-article","created":{"date-parts":[[2024,11,26]],"date-time":"2024-11-26T14:40:51Z","timestamp":1732632051000},"page":"153-188","source":"Crossref","is-referenced-by-count":140,"title":["Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes"],"prefix":"10.1146","volume":"8","author":[{"given":"Chen","family":"Tang","sequence":"first","affiliation":[{"name":"1Department of Computer Science, The University of Texas at Austin, Austin, Texas, USA; email: chen.tang@utexas.edu, abba@cs.utexas.edu, jiahengh@utexas.edu, robertomm@cs.utexas.edu, pstone@utexas.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ben","family":"Abbatematteo","sequence":"additional","affiliation":[{"name":"1Department of Computer Science, The University of Texas at Austin, Austin, Texas, USA; email: chen.tang@utexas.edu, abba@cs.utexas.edu, jiahengh@utexas.edu, robertomm@cs.utexas.edu, pstone@utexas.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiaheng","family":"Hu","sequence":"additional","affiliation":[{"name":"1Department of Computer Science, The University of Texas at Austin, Austin, Texas, USA; email: chen.tang@utexas.edu, abba@cs.utexas.edu, jiahengh@utexas.edu, robertomm@cs.utexas.edu, pstone@utexas.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rohan","family":"Chandra","sequence":"additional","affiliation":[{"name":"2Department of Computer Science, The University of Virginia, Charlottesville, Virginia, USA; email: rohanchandra@virginia.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Roberto","family":"Mart\u00edn-Mart\u00edn","sequence":"additional","affiliation":[{"name":"1Department of Computer Science, The University of Texas at Austin, Austin, Texas, USA; email: chen.tang@utexas.edu, abba@cs.utexas.edu, jiahengh@utexas.edu, robertomm@cs.utexas.edu, pstone@utexas.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peter","family":"Stone","sequence":"additional","affiliation":[{"name":"3Sony AI, New York, NY, USA"},{"name":"1Department of Computer Science, The University of Texas at Austin, Austin, Texas, USA; email: chen.tang@utexas.edu, abba@cs.utexas.edu, jiahengh@utexas.edu, robertomm@cs.utexas.edu, pstone@utexas.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"22","reference":[{"key":"B1","volume-title":"Reinforcement Learning: An Introduction","year":"2018"},{"issue":"3\u20134","key":"B2","first-page":"219","article-title":"An introduction to deep reinforcement learning","volume":"11","year":"2018","journal-title":"Found. Trends Mach. Learn."},{"issue":"7839","key":"B3","doi-asserted-by":"crossref","first-page":"604","DOI":"10.1038\/s41586-020-03051-4","article-title":"Mastering Atari, Go, chess and shogi by planning with a learned model","volume":"588","year":"2020","journal-title":"Nature"},{"issue":"7896","key":"B4","doi-asserted-by":"crossref","first-page":"223","DOI":"10.1038\/s41586-021-04357-7","article-title":"Outracing champion Gran Turismo drivers with deep reinforcement learning","volume":"602","year":"2022","journal-title":"Nature"},{"issue":"1","key":"B5","first-page":"5","article-title":"Reinforcement learning in healthcare: a survey","volume":"55","year":"2021","journal-title":"ACM Comput. Surv."},{"issue":"7","key":"B6","first-page":"145","article-title":"Reinforcement learning based recommender systems: a survey","volume":"55","year":"2022","journal-title":"ACM Comput. Surv."},{"issue":"7976","key":"B7","doi-asserted-by":"crossref","first-page":"982","DOI":"10.1038\/s41586-023-06419-4","article-title":"Champion-level drone racing using deep reinforcement learning","volume":"620","year":"2023","journal-title":"Nature"},{"key":"B8","article-title":"Superior robot mobility\u2014where AI meets the real world","year":"2023","journal-title":"ANYbotics"},{"key":"B9","article-title":"Starting on the right foot with reinforcement learning","year":"2024","journal-title":"Boston Dynamics"},{"issue":"6","key":"B10","first-page":"4909","article-title":"Deep reinforcement learning for autonomous driving: a survey","volume":"23","year":"2021","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"issue":"9","key":"B11","doi-asserted-by":"crossref","first-page":"2419","DOI":"10.1007\/s10994-021-05961-4","article-title":"Challenges of real-world reinforcement learning: definitions, benchmarks, and analysis","volume":"110","year":"2021","journal-title":"Mach. Learn."},{"issue":"4\u20135","key":"B12","first-page":"698","article-title":"How to train your robot with deep reinforcement learning: lessons we have learned","volume":"40","year":"2021","journal-title":"Int. J. Robot. Res."},{"issue":"30","key":"B13","first-page":"1","article-title":"A review of robot learning for manipulation: challenges, representations, and algorithms","volume":"22","year":"2021","journal-title":"J. Mach. Learn. Res."},{"issue":"5","key":"B14","doi-asserted-by":"crossref","first-page":"569","DOI":"10.1007\/s10514-022-10039-8","article-title":"Motion planning and control for mobile robot navigation using machine learning: a survey","volume":"46","year":"2022","journal-title":"Auton. Robots"},{"issue":"1\u20132","key":"B15","first-page":"1","article-title":"A survey on policy search for robotics","volume":"2","year":"2011","journal-title":"Found. Trends Robot."},{"key":"B16","doi-asserted-by":"crossref","first-page":"411","DOI":"10.1146\/annurev-control-042920-020211","article-title":"Safe learning in robotics: from learning-based control to safe reinforcement learning","volume":"5","year":"2022","journal-title":"Annu. Rev. Control Robot. Auton. Syst."},{"issue":"11","key":"B17","doi-asserted-by":"crossref","first-page":"1238","DOI":"10.1177\/0278364913495721","article-title":"Reinforcement learning in robotics: a survey","volume":"32","year":"2013","journal-title":"Int. J. Robot. Res."},{"issue":"4\u20135","key":"B18","first-page":"405","article-title":"The limits and potentials of deep learning for robotics","volume":"37","year":"2018","journal-title":"Int. J. Robot. Res."},{"key":"B19","volume-title":"Mechanics of Robotic Manipulation","year":"2001"},{"key":"B20","volume-title":"Springer Handbook of Robotics","year":"2008"},{"key":"B21","first-page":"1","article-title":"Toward robotic manipulation","volume":"1","year":"2018","journal-title":"Annu. Rev. Control Robot. Auton. Syst."},{"key":"B22","first-page":"2497","article-title":"Advanced skills by learning locomotion and local navigation end-to-end","volume-title":"2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems","year":"2022"},{"issue":"82","key":"B23","doi-asserted-by":"crossref","first-page":"eadg1462","DOI":"10.1126\/scirobotics.adg1462","article-title":"Reaching the limit in autonomous racing: optimal control versus reinforcement learning","volume":"8","year":"2023","journal-title":"Sci. Robot."},{"key":"B24","volume-title":"Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles","year":"2018"},{"issue":"1","key":"B25","doi-asserted-by":"crossref","first-page":"6039","DOI":"10.1038\/s41467-022-33128-9","article-title":"Technology readiness levels for machine learning systems","volume":"13","year":"2022","journal-title":"Nat. Commun."},{"key":"B26","first-page":"2619","article-title":"Policy gradient reinforcement learning for fast quadrupedal locomotion","volume-title":"2004 IEEE International Conference on Robotics and Automation","volume":"3","year":"2004"},{"key":"B27","first-page":"1615","article-title":"Autonomous helicopter control using reinforcement learning policy search methods","volume-title":"2001 IEEE International Conference on Robotics and Automation","volume":"2","year":"2001"},{"key":"B28","first-page":"1","article-title":"An application of reinforcement learning to aerobatic helicopter flight","volume-title":"Advances in Neural Information Processing Systems","volume":"19","year":"2006"},{"key":"B29","first-page":"1161","article-title":"Adapting rapid motor adaptation for bipedal robots","volume-title":"2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems","year":"2022"},{"key":"B30","article-title":"Sim-to-real: learning agile locomotion for quadruped robots","volume-title":"Robotics: Science and Systems XIV","volume":"10","year":"2018"},{"issue":"26","key":"B31","doi-asserted-by":"crossref","first-page":"eaau5872","DOI":"10.1126\/scirobotics.aau5872","article-title":"Learning agile and dynamic motor skills for legged robots","volume":"4","year":"2019","journal-title":"Sci. Robot."},{"key":"B32","first-page":"1893","article-title":"GenLoco: generalized locomotion controllers for quadrupedal robots","volume-title":"Proceedings of the 6th Conference on Robot Learning","year":"2023"},{"key":"B33","article-title":"Robust recovery controller for a quadrupedal robot using deep reinforcement learning","year":"2019"},{"issue":"49","key":"B34","doi-asserted-by":"crossref","first-page":"eabb2174","DOI":"10.1126\/scirobotics.abb2174","article-title":"Multi-expert learning of adaptive legged locomotion","volume":"5","year":"2020","journal-title":"Sci. Robot."},{"key":"B35","article-title":"RMA: rapid motor adaptation for legged robots","volume-title":"Robotics: Science and Systems XVII","volume":"11","year":"2021"},{"issue":"47","key":"B36","doi-asserted-by":"crossref","first-page":"eabc5986","DOI":"10.1126\/scirobotics.abc5986","article-title":"Learning quadrupedal locomotion over challenging terrain","volume":"5","year":"2020","journal-title":"Sci. Robot."},{"issue":"62","key":"B37","doi-asserted-by":"crossref","first-page":"eabk2822","DOI":"10.1126\/scirobotics.abk2822","article-title":"Learning robust perceptive locomotion for quadrupedal robots in the wild","volume":"7","year":"2022","journal-title":"Sci. Robot."},{"issue":"5","key":"B38","doi-asserted-by":"crossref","first-page":"2908","DOI":"10.1109\/TRO.2022.3172469","article-title":"RLOC: terrain-aware legged locomotion using reinforcement learning and optimal control","volume":"38","year":"2022","journal-title":"IEEE Trans. Robot."},{"issue":"74","key":"B39","doi-asserted-by":"crossref","first-page":"eade2256","DOI":"10.1126\/scirobotics.ade2256","article-title":"Learning quadrupedal locomotion on deformable terrain","volume":"8","year":"2023","journal-title":"Sci. Robot."},{"key":"B40","first-page":"5078","article-title":"DreamWaQ: learning robust quadrupedal locomotion with implicit terrain imagination via deep reinforcement learning","year":"2023","journal-title":"2023 IEEE International Conference on Robotics and Automation"},{"key":"B41","first-page":"25","article-title":"Adversarial motion priors make good substitutes for complex reward functions","year":"2022","journal-title":"2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems"},{"key":"B42","first-page":"12149","article-title":"Learning arm-assisted fall damage reduction and recovery for legged mobile manipulators","year":"2023","journal-title":"2023 IEEE International Conference on Robotics and Automation"},{"key":"B43","first-page":"928","article-title":"Minimizing energy consumption leads to the emergence of gaits in legged robots","year":"2022","journal-title":"Proceedings of the 5th Conference on Robot Learning"},{"key":"B44","first-page":"7295","article-title":"Learning visual locomotion with cross-modal supervision","year":"2023","journal-title":"2023 IEEE International Conference on Robotics and Automation"},{"key":"B45","first-page":"403","article-title":"Legged locomotion in challenging terrains using egocentric vision","year":"2023","journal-title":"Proceedings of the 6th Conference on Robot Learning"},{"key":"B46","first-page":"1430","article-title":"Neural volumetric memory for visual locomotion control","year":"2023","journal-title":"2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"issue":"86","key":"B47","doi-asserted-by":"crossref","first-page":"eadh5401","DOI":"10.1126\/scirobotics.adh5401","article-title":"DTC: deep tracking control","volume":"9","year":"2024","journal-title":"Sci. Robot."},{"key":"B48","first-page":"2791","article-title":"CAJun: continuous adaptive jumping using a learned centroidal controller","year":"2023","journal-title":"Proceedings of the 7th Conference on Robot Learning"},{"key":"B49","first-page":"1593","article-title":"Legged robots that keep on learning: fine-tuning locomotion policies in the real world","year":"2022","journal-title":"2022 International Conference on Robotics and Automation"},{"key":"B50","first-page":"11443","article-title":"Extreme parkour with legged robots","year":"2024","journal-title":"2024 IEEE International Conference on Robotics and Automation"},{"key":"B51","first-page":"73","article-title":"Robot parkour learning","year":"2023","journal-title":"Proceedings of the 7th Conference on Robot Learning"},{"key":"B52","first-page":"5120","article-title":"Advanced skills through multiple adversarial motion priors in reinforcement learning","year":"2023","journal-title":"2023 IEEE International Conference on Robotics and Automation"},{"key":"B53","first-page":"22","article-title":"Walk these ways: tuning robot control for generalization with multiplicity of behavior","year":"2023","journal-title":"Proceedings of the 6th Conference on Robot Learning"},{"key":"B54","article-title":"Demonstrating a walk in the park: learning to walk in 20 minutes with model-free reinforcement learning","volume-title":"Robotics: Science and Systems XIX","volume":"56","year":"2023"},{"key":"B55","first-page":"2226","article-title":"DayDreamer: world models for physical robot learning","year":"2023","journal-title":"Proceedings of the 6th Conference on Robot Learning"},{"key":"B56","article-title":"Learning memory-based control for human-scale bipedal locomotion","volume-title":"Robotics: Science and Systems XVI","volume":"31","year":"2020"},{"issue":"9","key":"B57","doi-asserted-by":"crossref","first-page":"2469","DOI":"10.1007\/s10994-021-05982-z","article-title":"Grounded action transformation for sim-to-real reinforcement learning","volume":"110","year":"2021","journal-title":"Mach. Learn."},{"key":"B58","first-page":"7309","article-title":"Sim-to-real learning of all common bipedal gaits via periodic reward composition","year":"2021","journal-title":"2021 IEEE International Conference on Robotics and Automation"},{"key":"B59","first-page":"2811","article-title":"Reinforcement learning for robust parameterized locomotion control of bipedal robots","year":"2021","journal-title":"2021 IEEE International Conference on Robotics and Automation"},{"key":"B60","article-title":"Blind bipedal stair traversal via sim-to-real reinforcement learning","volume-title":"Robotics: Science and Systems XVII","volume":"61","year":"2021"},{"key":"B61","doi-asserted-by":"crossref","first-page":"20135","DOI":"10.1109\/ACCESS.2022.3151771","article-title":"Reinforcement learning-based cascade motion policy design for robust 3D bipedal locomotion","volume":"10","year":"2022","journal-title":"IEEE Access"},{"key":"B62","first-page":"56","article-title":"Learning vision-based bipedal locomotion for challenging terrain","year":"2023","journal-title":"2024 IEEE International Conference on Robotics and Automation"},{"issue":"89","key":"B63","first-page":"eadi9579","article-title":"Real-world humanoid locomotion with reinforcement learning","volume":"9","year":"2023","journal-title":"Sci. Robot."},{"key":"B64","article-title":"Reinforcement learning for versatile, dynamic, and robust bipedal locomotion control","year":"2024"},{"issue":"4","key":"B65","doi-asserted-by":"crossref","first-page":"2096","DOI":"10.1109\/LRA.2017.2720851","article-title":"Control of a quadrotor with reinforcement learning","volume":"2","year":"2017","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B66","first-page":"59","article-title":"Sim-to-(multi)-real: transfer of low-level robust control policies to multiple quadrotors","year":"2019","journal-title":"2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems"},{"key":"B67","first-page":"10504","article-title":"A benchmark comparison of learned control policies for agile quadrotor flight","year":"2022","journal-title":"2022 International Conference on Robotics and Automation"},{"key":"B68","first-page":"1263","article-title":"Learning a single near-hover position controller for vastly different quadcopters","year":"2023","journal-title":"2023 IEEE International Conference on Robotics and Automation"},{"issue":"7","key":"B69","doi-asserted-by":"crossref","first-page":"6336","DOI":"10.1109\/LRA.2024.3396025","article-title":"Learning to fly in seconds","volume":"9","year":"2024","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B70","article-title":"Asymmetric actor critic for image-based robot learning","volume-title":"Robotics: Science and Systems XIV","volume":"8","year":"2018"},{"key":"B71","article-title":"Learning vision-guided quadrupedal locomotion end-to-end with cross-modal transformers","year":"2022","journal-title":"The Tenth International Conference on Learning Representations"},{"key":"B72","article-title":"Proximal policy optimization algorithms","year":"2017"},{"key":"B73","first-page":"2030","article-title":"MABEL, a new robotic bipedal walker and runner","year":"2009","journal-title":"2009 American Control Conference"},{"key":"B74","article-title":"Unitree H1 the world's first full-size motor drive humanoid robot flips on ground","year":"2024","journal-title":"YouTube"},{"key":"B75","article-title":"Atlas \u2223 partners in parkour","year":"2021","journal-title":"YouTube"},{"issue":"59","key":"B76","doi-asserted-by":"crossref","first-page":"eabg5810","DOI":"10.1126\/scirobotics.abg5810","article-title":"Learning high-speed flight in the wild","volume":"6","year":"2021","journal-title":"Sci. Robot."},{"key":"B77","first-page":"31","article-title":"Virtual-to-real deep reinforcement learning: continuous control of mobile robots for mapless navigation","year":"2017","journal-title":"2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems"},{"key":"B78","first-page":"9224","article-title":"Benchmarking reinforcement learning techniques for autonomous navigation","year":"2023","journal-title":"2023 IEEE International Conference on Robotics and Automation"},{"issue":"2","key":"B79","doi-asserted-by":"crossref","first-page":"2007","DOI":"10.1109\/LRA.2019.2899918","article-title":"Learning navigation behaviors end-to-end with autoRL","volume":"4","year":"2019","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B80","first-page":"213","article-title":"Learning over subgoals for efficient navigation of structured, unknown environments","year":"2018","journal-title":"Proceedings of the 2nd Conference on Robot Learning"},{"key":"B81","first-page":"3357","article-title":"Target-driven visual navigation in indoor scenes using deep reinforcement learning","year":"2017","journal-title":"2017 IEEE International Conference on Robotics and Automation"},{"key":"B82","first-page":"4247","article-title":"Object goal navigation using goal-oriented semantic exploration","volume-title":"Advances in Neural Information Processing Systems 33","year":"2020"},{"issue":"79","key":"B83","doi-asserted-by":"crossref","first-page":"eadf6991","DOI":"10.1126\/scirobotics.adf6991","article-title":"Navigating to objects in the real world","volume":"8","year":"2023","journal-title":"Sci. Robot."},{"issue":"4","key":"B84","doi-asserted-by":"crossref","first-page":"6670","DOI":"10.1109\/LRA.2020.3013848","article-title":"Sim2real predictivity: Does evaluation in simulation predict real-world performance?","volume":"5","year":"2020","journal-title":"IEEE Robot. Autom. Lett."},{"issue":"2","key":"B85","doi-asserted-by":"crossref","first-page":"1312","DOI":"10.1109\/LRA.2021.3057023","article-title":"BADGR: an autonomous self-supervised learning-based navigation system","volume":"6","year":"2021","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B86","first-page":"44","article-title":"Offline reinforcement learning for visual navigation","year":"2023","journal-title":"Proceedings of the 6th Conference on Robot Learning"},{"key":"B87","first-page":"1714","article-title":"Information theoretic MPC for model-based reinforcement learning","year":"2017","journal-title":"2017 IEEE International Conference on Robotics and Automation"},{"key":"B88","first-page":"3100","article-title":"FastRLAP: a system for learning high-speed driving via deep RL and autonomous practicing","year":"2023","journal-title":"Proceedings of the 7th Conference on Robot Learning"},{"key":"B89","first-page":"8248","article-title":"Learning to drive in a day","year":"2019","journal-title":"2019 IEEE International Conference on Robotics and Automation"},{"key":"B90","article-title":"Reinforcement learning based oscillation dampening: scaling up single-agent RL algorithms to a 100 AV highway field operational test","year":"2024"},{"issue":"3","key":"B91","doi-asserted-by":"crossref","first-page":"5081","DOI":"10.1109\/LRA.2021.3068639","article-title":"Learning a state representation and navigation in cluttered and dynamic environments","volume":"6","year":"2021","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B92","first-page":"859","article-title":"Rethinking sim2real: lower fidelity simulation leads to higher sim2real transfer in navigation","year":"2023","journal-title":"Proceedings of the 6th Conference on Robot Learning"},{"issue":"5","key":"B93","doi-asserted-by":"crossref","first-page":"4798","DOI":"10.1109\/LRA.2024.3385611","article-title":"IndoorSim-to-OutdoorReal: learning to navigate outdoors without any outdoor experience","volume":"9","year":"2024","journal-title":"IEEE Robot. Autom. Lett."},{"issue":"2","key":"B94","doi-asserted-by":"crossref","first-page":"3906","DOI":"10.1109\/LRA.2022.3145947","article-title":"Learning to navigate sidewalks in outdoor environments","volume":"7","year":"2022","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B95","first-page":"34","article-title":"Resilient legged local navigation: learning to traverse with compromised perception end-to-end","year":"2024","journal-title":"2024 IEEE International Conference on Robotics and Automation"},{"issue":"88","key":"B96","doi-asserted-by":"crossref","first-page":"eadi7566","DOI":"10.1126\/scirobotics.adi7566","article-title":"ANYmal parkour: learning agile navigation for quadrupedal robots","volume":"9","year":"2024","journal-title":"Sci. Robot."},{"issue":"89","key":"B97","doi-asserted-by":"crossref","first-page":"eadi9641","DOI":"10.1126\/scirobotics.adi9641","article-title":"Learning robust autonomous navigation and locomotion for wheeled-legged robots","volume":"9","year":"2024","journal-title":"Sci. Robot."},{"key":"B98","first-page":"8649","article-title":"Learning to walk in confined spaces using 3D representation","year":"2024","journal-title":"2024 IEEE International Conference on Robotics and Automation"},{"key":"B99","first-page":"11474","article-title":"Dexterous legged locomotion in confined 3D spaces with reinforcement learning","year":"2024","journal-title":"2024 IEEE International Conference on Robotics and Automation"},{"key":"B100","article-title":"Agile but safe: learning collision-free high-speed legged locomotion","year":"2024"},{"key":"B101","article-title":"CAD2RL: real single-image flight without a single real image","volume-title":"Robotics: Science and Systems XIII","volume":"34","year":"2017"},{"key":"B102","first-page":"6008","article-title":"Generalization through simulation: integrating simulated and real data into deep reinforcement learning for vision-based autonomous flight","year":"2019","journal-title":"2019 IEEE International Conference on Robotics and Automation"},{"key":"B103","first-page":"14777","article-title":"Actor-critic model predictive control","year":"2024","journal-title":"2024 IEEE International Conference on Robotics and Automation"},{"key":"B104","first-page":"3674","article-title":"Vision-and-language navigation: interpreting visually-grounded navigation instructions in real environments","year":"2018","journal-title":"2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"B105","first-page":"5129","article-title":"Self-supervised deep reinforcement learning with generalized computation graphs for robot navigation","year":"2018","journal-title":"2018 IEEE International Conference on Robotics and Automation"},{"key":"B106","article-title":"DD-PPO: learning near-perfect PointGoal navigators from 2.5 billion frames","year":"2019"},{"key":"B107","first-page":"9339","article-title":"Habitat: a platform for embodied AI research","year":"2019","journal-title":"2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"B108","article-title":"Towards monocular vision based obstacle avoidance through deep reinforcement learning","year":"2017"},{"key":"B109","article-title":"Emergence of maps in the memories of blind navigation agents","year":"2023","journal-title":"The Eleventh International Conference on Learning Representations"},{"key":"B110","first-page":"3437","article-title":"NeRF-SLAM: real-time dense monocular slam with neural radiance fields","year":"2023","journal-title":"2023 IEEE\/RSJ International Conference on Intelligent Robots and Systems"},{"issue":"26","key":"B111","doi-asserted-by":"crossref","first-page":"eaau4984","DOI":"10.1126\/scirobotics.aau4984","article-title":"Learning ambidextrous robot grasping policies","volume":"4","year":"2019","journal-title":"Sci. Robot."},{"key":"B112","first-page":"4238","article-title":"Learning synergies between pushing and grasping with self-supervised deep reinforcement learning","year":"2018","journal-title":"2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems"},{"key":"B113","first-page":"651","article-title":"Scalable deep reinforcement learning for vision-based robotic manipulation","year":"2018","journal-title":"Proceedings of the 2nd Conference on Robot Learning"},{"key":"B114","first-page":"12627","article-title":"Sim-to-real via sim-to-sim: data-efficient robotic grasping via randomized-to-canonical adaptation networks","year":"2019","journal-title":"2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"B115","first-page":"1345","article-title":"On-robot learning with equivariant models","year":"2023","journal-title":"Proceedings of the 6th Conference on Robot Learning"},{"issue":"1","key":"B116","first-page":"1334","article-title":"End-to-end training of deep visuomotor policies","volume":"17","year":"2016","journal-title":"J. Mach. Learn. Res."},{"key":"B117","first-page":"557","article-title":"Scaling up multi-task robotic reinforcement learning","year":"2022","journal-title":"Proceedings of the 5th Conference on Robot Learning"},{"key":"B118","first-page":"1518","article-title":"Actionable models: unsupervised offline reinforcement learning of robotic skills","year":"2021","journal-title":"Proceedings of the 38th International Conference on Machine Learning"},{"key":"B119","first-page":"1089","article-title":"Beyond pick-and-place: tackling robotic stacking of diverse shapes","year":"2021","journal-title":"Proceedings of the 5th Conference on Robot Learning"},{"key":"B120","first-page":"1652","article-title":"Don't start from scratch: leveraging prior data to automate robotic reinforcement learning","year":"2023","journal-title":"Proceedings of the 6th Conference on Robot Learning"},{"key":"B121","article-title":"Visual foresight: model-based deep reinforcement learning for vision-based robotic control","year":"2018"},{"key":"B122","first-page":"4344","article-title":"Learning by playing solving sparse reward tasks from scratch","year":"2018","journal-title":"Proceedings of the 35th International Conference on Machine Learning"},{"key":"B123","article-title":"The ingredients of real world robotic reinforcement learning","year":"2020","journal-title":"The Eighth International Conference on Learning Representations"},{"key":"B124","article-title":"VIP: towards universal visual reward and representation via value-implicit pre-training","year":"2023","journal-title":"The Eleventh International Conference on Learning Representations"},{"key":"B125","article-title":"AWAC: accelerating online reinforcement learning with offline datasets","year":"2020"},{"key":"B126","first-page":"7477","article-title":"Augmenting reinforcement learning with behavior primitives for diverse manipulation tasks","year":"2022","journal-title":"2022 IEEE International Conference on Robotics and Automation"},{"key":"B127","first-page":"3909","article-title":"Q-Transformer: scalable offline reinforcement learning via autoregressive Q-functions","year":"2023","journal-title":"Proceedings of the 7th Conference on Robot Learning"},{"key":"B128","first-page":"9191","article-title":"Visual reinforcement learning with imagined goals","volume-title":"Advances in Neural Information Processing Systems 31","year":"2018"},{"key":"B129","first-page":"6023","article-title":"Residual reinforcement learning for robot control","year":"2019","journal-title":"2019 International Conference on Robotics and Automation"},{"key":"B130","article-title":"Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards","year":"2017"},{"key":"B131","article-title":"Robust multi-modal policies for industrial assembly via reinforcement learning and demonstrations: a large-scale study","volume-title":"Robotics: Science and Systems XVII","volume":"88","year":"2021"},{"key":"B132","first-page":"6386","article-title":"Offline meta-reinforcement learning for industrial insertion","year":"2022","journal-title":"2022 IEEE International Conference on Robotics and Automation"},{"key":"B133","article-title":"IndustReal: transferring contact-rich assembly tasks from simulation to reality","volume-title":"Robotics: Science and Systems XIX","volume":"39","year":"2023"},{"key":"B134","first-page":"8973","article-title":"Closing the sim-to-real loop: adapting simulation randomization with real world experience","year":"2019","journal-title":"2019 IEEE International Conference on Robotics and Automation"},{"key":"B135","first-page":"7522","article-title":"Composable interaction primitives: a structured policy class for efficiently learning sustained-contact manipulation skills","year":"2024","journal-title":"2024 IEEE International Conference on Robotics and Automation"},{"key":"B136","article-title":"VAT-Mart: learning visual action trajectory proposals for manipulating 3D articulated objects","year":"2022","journal-title":"The Tenth International Conference on Learning Representations"},{"key":"B137","first-page":"734","article-title":"Sim-to-real reinforcement learning for deformable object manipulation","year":"2018","journal-title":"Proceedings of the 2nd Conference on Robot Learning"},{"key":"B138","article-title":"Learning to manipulate deformable objects without demonstrations","volume-title":"Robotics: Science and Systems XVI","volume":"65","year":"2020"},{"key":"B139","first-page":"1","article-title":"SpeedFolding: learning efficient bimanual folding of garments","year":"2022","journal-title":"2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems"},{"key":"B140","article-title":"One policy to dress them all: learning to dress people with diverse poses and garments","volume-title":"Robotics: Science and Systems XIX","volume":"8","year":"2023"},{"issue":"1","key":"B141","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1177\/0278364919887447","article-title":"Learning dexterous in-hand manipulation","volume":"39","year":"2020","journal-title":"Int. J. Robot. Res."},{"key":"B142","first-page":"5977","article-title":"DeXtreme: transfer of agile in-hand manipulation from simulation to reality","year":"2023","journal-title":"2023 IEEE International Conference on Robotics and Automation"},{"key":"B143","first-page":"1101","article-title":"Deep dynamics models for learning dexterous manipulation","year":"2020","journal-title":"Proceedings of the Conference on Robot Learning"},{"key":"B144","first-page":"2549","article-title":"General in-hand object rotation with vision and touch","year":"2023","journal-title":"Proceedings of the 7th Conference on Robot Learning"},{"issue":"84","key":"B145","doi-asserted-by":"crossref","first-page":"eadc9244","DOI":"10.1126\/scirobotics.adc9244","article-title":"Visual dexterity: in-hand reorientation of novel and complex object shapes","volume":"8","year":"2023","journal-title":"Sci. Robot."},{"key":"B146","first-page":"150","article-title":"Learning to grasp the ungraspable with emergent extrinsic dexterity","year":"2023","journal-title":"Proceedings of the 6th Conference on Robot Learning"},{"key":"B147","first-page":"241","article-title":"HACMan: learning hybrid actor-critic maps for 6D non-prehensile manipulation","year":"2023","journal-title":"Proceedings of the 7th Conference on Robot Learning"},{"key":"B148","article-title":"CORN: contact-based object representation for nonprehensile manipulation of general unseen objects","year":"2024","journal-title":"The Twelfth International Conference on Learning Representations"},{"key":"B149","article-title":"Born out of research","year":"2024","journal-title":"Covariant"},{"key":"B150","first-page":"2745","article-title":"Learning purely tactile in-hand manipulation with a torque-controlled hand","year":"2022","journal-title":"2022 International Conference on Robotics and Automation"},{"key":"B151","first-page":"12112","article-title":"Learning a shape-conditioned agent for purely tactile in-hand manipulation of various objects","year":"2024","journal-title":"2024 IEEE\/RSJ International Conference on Intelligent Robots and Systems"},{"key":"B152","article-title":"SAM-RL: sensing-aware model-based reinforcement learning via differentiable physics-based simulation and rendering","volume-title":"Robotics: Science and Systems XIX","volume":"40","year":"2023"},{"key":"B153","article-title":"Geometric fabrics: a safe guiding medium for policy learning","year":"2024"},{"key":"B154","first-page":"1149","article-title":"Efficient bimanual manipulation using learned task schemas","year":"2020","journal-title":"2020 IEEE International Conference on Robotics and Automation"},{"issue":"6","key":"B155","doi-asserted-by":"crossref","first-page":"3850","DOI":"10.1109\/TRO.2022.3176207","article-title":"Learning to play table tennis from scratch using muscular robots","volume":"38","year":"2022","journal-title":"IEEE Trans. Robot."},{"issue":"10","key":"B156","doi-asserted-by":"crossref","first-page":"6451","DOI":"10.1109\/LRA.2023.3308061","article-title":"LEAGUE: guided skill learning and abstraction for long-horizon manipulation","volume":"8","year":"2023","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B157","first-page":"1401","article-title":"Learn2Assemble with structured representations and search for robotic architectural construction","year":"2022","journal-title":"Proceedings of the 5th Conference on Robot Learning"},{"issue":"2","key":"B158","doi-asserted-by":"crossref","first-page":"2377","DOI":"10.1109\/LRA.2022.3143567","article-title":"Combining learning-based locomotion policy with model-based manipulation for legged mobile manipulators","volume":"7","year":"2022","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B159","first-page":"138","article-title":"Deep whole-body control: learning a unified policy for manipulation and locomotion","year":"2023","journal-title":"Proceedings of the 6th Conference on Robot Learning"},{"issue":"3","key":"B160","doi-asserted-by":"crossref","first-page":"939","DOI":"10.3390\/s20030939","article-title":"Learning mobile manipulation through deep reinforcement learning","volume":"20","year":"2020","journal-title":"Sensors"},{"key":"B161","article-title":"HumanPlus: humanoid shadowing and imitation from humans","year":"2024"},{"key":"B162","article-title":"Causal policy gradient for whole-body mobile manipulation","volume-title":"Robotics: Science and Systems XIX","volume":"49","year":"2023"},{"key":"B163","article-title":"Harmonic mobile manipulation","year":"2023"},{"key":"B164","first-page":"5106","article-title":"Legs as manipulator: pushing quadrupedal agility beyond locomotion","year":"2023","journal-title":"2023 IEEE International Conference on Robotics and Automation"},{"key":"B165","first-page":"1479","article-title":"Hierarchical reinforcement learning for precise soccer shooting skills using a quadrupedal robot","year":"2022","journal-title":"2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems"},{"key":"B166","first-page":"5155","article-title":"Dribblebot: dynamic legged manipulation in the wild","year":"2023","journal-title":"2023 IEEE International Conference on Robotics and Automation"},{"issue":"5","key":"B167","doi-asserted-by":"crossref","first-page":"3601","DOI":"10.1109\/TRO.2023.3284346","article-title":"N2M2 : learning navigation for arbitrary mobile manipulation motions in unseen and dynamic environments","volume":"39","year":"2023","journal-title":"IEEE Trans. Robot."},{"key":"B168","first-page":"308","article-title":"Fully autonomous real-world reinforcement learning with applications to mobile manipulation","year":"2022","journal-title":"Proceedings of the 5th Conference on Robot Learning"},{"issue":"3","key":"B169","doi-asserted-by":"crossref","first-page":"8399","DOI":"10.1109\/LRA.2022.3188109","article-title":"Robot learning of mobile manipulation with reachability behavior priors","volume":"7","year":"2022","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B170","article-title":"Adaptive mobile manipulation for articulated objects in the open world","year":"2024"},{"key":"B171","first-page":"18133","article-title":"SPIN: simultaneous perception interaction and navigation","year":"2024","journal-title":"2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"B172","article-title":"Visual whole-body control for legged loco-manipulation","year":"2024"},{"issue":"8","key":"B173","doi-asserted-by":"crossref","first-page":"4601","DOI":"10.1109\/LRA.2023.3286171","article-title":"Cascaded compositional residual learning for complex interactive behaviors","volume":"8","year":"2023","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B174","first-page":"11690","article-title":"M-EMBER: tackling long-horizon mobile manipulation via factorized domain transfer","year":"2023","journal-title":"2023 IEEE International Conference on Robotics and Automation"},{"issue":"1","key":"B175","first-page":"779","article-title":"Adaptive skill coordination for robotic mobile manipulation","volume":"9","year":"2023","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B176","article-title":"Deep RL at scale: sorting waste in office buildings with a fleet of mobile manipulators","year":"2023"},{"key":"B177","first-page":"2641","article-title":"A whole-body control framework for humanoids operating in human environments","year":"2006","journal-title":"2006 IEEE International Conference on Robotics and Automation"},{"issue":"2","key":"B178","first-page":"566","article-title":"Human-centered collaborative robots with deep reinforcement learning","volume":"6","year":"2020","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B179","first-page":"3168","article-title":"SynH2R: synthesizing hand-object motions for learning human-to-robot handovers","year":"2023","journal-title":"2023 IEEE International Conference on Robotics and Automation"},{"key":"B180","first-page":"9654","article-title":"Learning human-to-robot handovers from point clouds","year":"2023","journal-title":"2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"B181","first-page":"1011","article-title":"Reinforcement learning of variable admittance control for human-robot co-manipulation","year":"2015","journal-title":"2015 IEEE\/RSJ International Conference on Intelligent Robots and Systems"},{"key":"B182","first-page":"1343","article-title":"Socially aware motion planning with deep reinforcement learning","year":"2017","journal-title":"2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems"},{"key":"B183","doi-asserted-by":"crossref","first-page":"10357","DOI":"10.1109\/ACCESS.2021.3050338","article-title":"Collision avoidance in pedestrian-rich environments with deep reinforcement learning","volume":"9","year":"2021","journal-title":"IEEE Access"},{"key":"B184","first-page":"4221","article-title":"Crowd-steer: realtime smooth and collision-free robot navigation in densely crowded scenarios trained using high-fidelity simulation","year":"2021","journal-title":"Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence"},{"key":"B185","article-title":"SELFI: autonomous self-improvement with reinforcement learning for social navigation","year":"2024"},{"key":"B186","first-page":"9449","article-title":"Safe reinforcement learning of dynamic high-dimensional robotic tasks: navigation, manipulation, interaction","year":"2023","journal-title":"2023 IEEE International Conference on Robotics and Automation"},{"key":"B187","first-page":"1303","article-title":"Learning language-conditioned robot behavior from offline data and crowd-sourced annotation","year":"2022","journal-title":"Proceedings of the 5th Conference on Robot Learning"},{"key":"B188","article-title":"Shared autonomy via deep reinforcement learning","volume-title":"Robotics: Science and Systems XIV","year":"2018"},{"key":"B189","article-title":"Residual policy learning for shared autonomy","volume-title":"Robotics: Science and Systems XVI","volume":"72","year":"2020"},{"key":"B190","first-page":"285","article-title":"Decentralized non-communicating multiagent collision avoidance with deep reinforcement learning","year":"2017","journal-title":"2017 IEEE International Conference on Robotics and Automation"},{"key":"B191","first-page":"3052","article-title":"Motion planning among dynamic, decision-making agents with deep reinforcement learning","year":"2018","journal-title":"2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems"},{"issue":"7","key":"B192","doi-asserted-by":"crossref","first-page":"856","DOI":"10.1177\/0278364920916531","article-title":"Distributed multi-robot collision avoidance via deep reinforcement learning for navigation in complex scenarios","volume":"39","year":"2020","journal-title":"Int. J. Robot. Res."},{"issue":"3","key":"B193","doi-asserted-by":"crossref","first-page":"5896","DOI":"10.1109\/LRA.2022.3161699","article-title":"Reinforcement learned distributed multi-robot navigation with reciprocal velocity obstacle shaped rewards","volume":"7","year":"2022","journal-title":"IEEE Robot. Autom. Lett."},{"issue":"3","key":"B194","doi-asserted-by":"crossref","first-page":"2378","DOI":"10.1109\/LRA.2019.2903261","article-title":"PRIMAL: pathfinding via reinforcement and imitation multi-agent learning","volume":"4","year":"2019","journal-title":"IEEE Robot. Autom. Lett."},{"key":"B195","first-page":"110","article-title":"Multi-agent manipulation via locomotion using hierarchical sim2real","year":"2019","journal-title":"Proceedings of the Conference on Robot Learning"},{"issue":"89","key":"B196","doi-asserted-by":"crossref","first-page":"eadi8022","DOI":"10.1126\/scirobotics.adi8022","article-title":"Learning agile soccer skills for a bipedal robot with deep reinforcement learning","volume":"9","year":"2024","journal-title":"Sci. Robot."},{"key":"B197","first-page":"3","article-title":"Reciprocal n-body collision avoidance","year":"2011","journal-title":"Robotics Research: The 14th International Symposium ISRR"},{"key":"B198","first-page":"34556","article-title":"Jump-start reinforcement learning","year":"2023","journal-title":"Proceedings of the 40th International Conference on Machine Learning"},{"key":"B199","doi-asserted-by":"crossref","first-page":"61857","DOI":"10.52202\/075280-2704","article-title":"Residual Q-learning: offline and online policy customization without value","volume-title":"Advances in Neural Information Processing Systems 36","year":"2023"},{"key":"B200","article-title":"TD-MPC2: scalable, robust world models for continuous control","year":"2023","journal-title":"The Eleventh International Conference on Learning Representations"},{"key":"B201","article-title":"BaRiFlex: a robotic gripper with versatility and collision robustness for robot learning","year":"2023"},{"key":"B202","article-title":"Diversity is all you need: learning skills without a reward function","year":"2019","journal-title":"The Seventh International Conference on Learning Representations"},{"key":"B203","first-page":"2594","article-title":"Curiosity-driven learning of joint locomotion and manipulation tasks","year":"2023","journal-title":"Proceedings of the 7th Conference on Robot Learning"},{"key":"B204","first-page":"4955","article-title":"Dynamics randomization revisited: a case study for quadrupedal locomotion","year":"2021","journal-title":"2021 IEEE International Conference on Robotics and Automation"},{"key":"B205","first-page":"1010","article-title":"Variable impedance control in end-effector space: an action space for reinforcement learning in contact-rich tasks","year":"2019","journal-title":"2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems"},{"key":"B206","first-page":"4583","article-title":"ReLMoGen: integrating motion generation in reinforcement learning for mobile manipulation","year":"2021","journal-title":"2021 IEEE International Conference on Robotics and Automation"},{"key":"B207","article-title":"FMB: a functional manipulation benchmark for generalizable robotic learning","year":"2024"},{"key":"B208","article-title":"FurnitureBench: reproducible real-world benchmark for long-horizon complex manipulation","volume-title":"Robotics: Science and Systems XIX","volume":"23","year":"2023"},{"key":"B209","first-page":"214","article-title":"RoboCup@Home: adaptive benchmarking of robot bodies and minds","year":"2011","journal-title":"Social Robotics: Third International Conference on Social Robotics, ICSR 2011"},{"key":"B210","article-title":"Evaluating real-world robot manipulation policies in simulation","year":"2024"},{"key":"B211","article-title":"Foundation models in robotics: applications, challenges, and the future","year":"2023"},{"key":"B212","article-title":"Toward general-purpose robots via foundation models: a survey and meta-analysis","year":"2023"},{"key":"B213","article-title":"Foundation models for decision making: problems, methods, and opportunities","year":"2023"},{"key":"B214","article-title":"DrEureka: language model guided sim-to-real transfer","year":"2024"}],"container-title":["Annual Review of Control, Robotics, and Autonomous Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.annualreviews.org\/content\/journals\/10.1146\/annurev-control-030323-022510?crawler=true&mimetype=application\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T04:18:52Z","timestamp":1773461932000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.annualreviews.org\/content\/journals\/10.1146\/annurev-control-030323-022510"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,5]]},"references-count":214,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,5,5]]}},"URL":"https:\/\/doi.org\/10.1146\/annurev-control-030323-022510","relation":{},"ISSN":["2573-5144"],"issn-type":[{"value":"2573-5144","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,5]]}}}