{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T19:34:55Z","timestamp":1785353695096,"version":"3.55.0"},"reference-count":106,"publisher":"American Association for the Advancement of Science (AAAS)","issue":"105","license":[{"start":{"date-parts":[[2026,8,20]],"date-time":"2026-08-20T00:00:00Z","timestamp":1787184000000},"content-version":"vor","delay-in-days":365,"URL":"https:\/\/www.science.org\/content\/page\/science-licenses-journal-article-reuse"}],"content-domain":{"domain":["www.science.org"],"crossmark-restriction":true},"short-container-title":["Sci. Robot."],"published-print":{"date-parts":[[2025,8,20]]},"abstract":"<jats:p>Robotic manipulation remains one of the most difficult challenges in robotics, with approaches ranging from classical model-based control to modern imitation learning. Although these methods have enabled substantial progress, they often require extensive manual design, struggle with performance, and demand large-scale data collection. These limitations hinder their real-world deployment at scale, where reliability, speed, and robustness are essential. Reinforcement learning (RL) offers a powerful alternative by enabling robots to autonomously acquire complex manipulation skills through interaction. However, realizing the full potential of RL in the real world remains challenging because of issues of sample efficiency and safety. We present a human-in-the-loop, vision-based RL system that achieved strong performance on a wide range of dexterous manipulation tasks, including precise assembly, dynamic manipulation, and dual-arm coordination. These tasks reflect realistic industrial tolerances, with small but critical variations in initial object placements that demand sophisticated reactive control. Our method integrates demonstrations, human corrections, sample-efficient RL algorithms, and system-level design to directly learn RL policies in the real world. Within 1 to 2.5 hours of real-world training, our approach outperformed other baselines by improving task success by 2\u00d7, achieving near-perfect success rates, and executing 1.8\u00d7 faster on average. Through extensive experiments and analysis, our results suggest that RL can learn a wide range of complex vision-based manipulation policies directly in the real world within practical training times. We hope that this work will inspire a new generation of learned robotic manipulation techniques, benefiting both industrial applications and research advancements.<\/jats:p>","DOI":"10.1126\/scirobotics.ads5033","type":"journal-article","created":{"date-parts":[[2025,8,20]],"date-time":"2025-08-20T17:58:23Z","timestamp":1755712703000},"update-policy":"https:\/\/doi.org\/10.34133\/aaas_crossmark","source":"Crossref","is-referenced-by-count":55,"title":["Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning"],"prefix":"10.1126","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-8029-7794","authenticated-orcid":true,"given":"Jianlan","family":"Luo","sequence":"first","affiliation":[{"name":"Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, CA 94720, USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9126-3680","authenticated-orcid":true,"given":"Charles","family":"Xu","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, CA 94720, USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jeffrey","family":"Wu","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, CA 94720, USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6764-2743","authenticated-orcid":true,"given":"Sergey","family":"Levine","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, CA 94720, USA."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"221","reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.abd9461"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.aau5872"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.abc5986"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.adc9244"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.abg5810"},{"key":"e_1_3_2_7_2","unstructured":"OpenAI I. Akkaya M. Andrychowicz M. Chociej M. Litwin B. McGrew A. Petron Paino M. Plappert G. Powell R. Ribas J. Schneider N. Tezak J. Tworek P. Welinder L. Weng Q. Yuan W. Zaremba L. Zhang Solving Rubik\u2019s cube with a robot hand. arXiv:1910.07113 [cs.LG] (2019)."},{"key":"e_1_3_2_8_2","unstructured":"D. Kalashnikov A. Irpan P. Pastor J. Ibarz A. Herzog E. Jang D. Quillen E. Holly M. Kalakrishnan V. Vanhoucke S. Levine QT-Opt: Scalable deep reinforcement learning for vision-based robotic manipulation. arXiv:1806.10293 [cs.LG] (2018)."},{"key":"e_1_3_2_9_2","unstructured":"D. Kalashnikov J. Varley Y. Chebotar B. Swanson R. Jonschkowski C. Finn S. Levine K. Hausman MT-Opt: Continuous multi-task robotic reinforcement learning at scale. arXiv:2104.08212 [cs.RO] (2021)."},{"key":"e_1_3_2_10_2","first-page":"3137","article-title":"A generalized path integral control approach to reinforcement learning","volume":"11","author":"Theodorou E.","year":"2010","unstructured":"E. Theodorou, J. Buchli, S. Schaal, A generalized path integral control approach to reinforcement learning. J. Mach. Learn. Res. 11, 3137\u20133181 (2010).","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Y. Chebotar M. Kalakrishnan A. Yahya A. Li S. Schaal S. Levine Path integral guided policy search. arXiv:1610.00529 [cs.RO] (2016).","DOI":"10.1109\/ICRA.2017.7989384"},{"key":"e_1_3_2_12_2","unstructured":"P. J. Ball L. Smith I. Kostrikov S. Levine Efficient online reinforcement learning with offline data. arXiv:2302.02948 [cs.LG] (2023)."},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","unstructured":"T. Tang H.-C. Lin Y. Zhao W. Chen M. Tomizuka \u201cAutonomous alignment of peg and hole by force\/torque measurement for robotic assembly\u201d in 2016 IEEE International Conference on Automation Science and Engineering (CASE) (IEEE 2016) pp. 162\u2013167.","DOI":"10.1109\/COASE.2016.7743375"},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","unstructured":"S. Jin X. Zhu C. Wang M. Tomizuka \u201cContact pose identification for peg-in-hole assembly under uncertainties\u201d in 2021 American Control Conference (ACC) (IEEE 2021) pp. 48\u201353.","DOI":"10.23919\/ACC50511.2021.9482981"},{"key":"e_1_3_2_15_2","doi-asserted-by":"crossref","unstructured":"A. S. Morgan B. Wen J. Liang A. Boularias A. M. Dollar K. Bekris \u201cVision-driven compliant manipulation for reliable high-precision assembly tasks\u201d in Proceedings of Robotics: Science and Systems (MIT Press Journals 2021) 10.15607\/RSS.2021.XVII.070.","DOI":"10.15607\/RSS.2021.XVII.070"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIE.2021.3108710"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"B. Tang M. A. Lin I. Akinola A. Handa G. S. Sukhatme F. Ramos D. Fox Y. Narang IndustReal: Transferring contact-rich assembly tasks from simulation to reality. arXiv:2305.17110 [cs.RO] (2023).","DOI":"10.15607\/RSS.2023.XIX.039"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2025.3551637"},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","unstructured":"Y. Narang K. Storey I. Akinola M. Macklin P. Reist L. Wawrzyniak Y. Guo A. Moravanszky G. State M. Lu A. Handa D. Fox Factory: Fast contact for robotic assembly. arXiv:2205.03532 [cs.RO] (2022).","DOI":"10.15607\/RSS.2022.XVIII.035"},{"key":"e_1_3_2_20_2","unstructured":"Y. Guo B. Tang I. Akinola D. Fox A. Gupta Y. Narang SRSA: Skill retrieval and adaptation for robotic assembly tasks. arXiv:2503.04538 [cs.RO] (2025)."},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","unstructured":"B. Tang I. Akinola J. Xu B.Wen A. Handa K. V.Wyk D. Fox G. S. Sukhatme F. Ramos Y. Narang AutoMate: Specialist and generalist assembly policies over diverse geometries. arXiv:2407.08028 [cs.RO] (2024).","DOI":"10.15607\/RSS.2024.XX.064"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"O. Spector V. Tchuiev D. D. Castro InsertionNet 2.0: Minimal contact multi-step insertion using multimodal multiview sensory input. arXiv:2203.01153 [cs.RO] (2022).","DOI":"10.1109\/ICRA46639.2022.9811798"},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","unstructured":"H. Chang A. Boularias S. Jain \u201cInsert-one: One-shot robust visual-force servoing for novel object insertion with 6-DoF tracking\u201d in 2024 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS 2024) (IEEE 2024) pp. 2935\u20132942.","DOI":"10.1109\/IROS58592.2024.10801884"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"H.-C. Song M.-C. Kim J.-B. Song \u201cUSB assembly strategy based on visual servoing and impedance control\u201d in 2015 12th International Conference on Ubiquitous Robots and Ambient Intelligence (URAI) (IEEE 2015) pp. 114\u2013117.","DOI":"10.1109\/URAI.2015.7358873"},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","unstructured":"K. J. Astrom R. M. Murray Feedback Systems: An Introduction for Scientists and Engineers (Princeton Univ. Press 2008).","DOI":"10.1515\/9781400828739"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"M. Mason K. Lynch \u201cDynamic manipulation\u201d in Proceedings of 1993 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS \u201893) (IEEE 1993) pp. 152\u2013159.","DOI":"10.1109\/IROS.1993.583093"},{"key":"e_1_3_2_27_2","doi-asserted-by":"crossref","unstructured":"P. Kormushev S. Calinon D. G. Caldwell \u201cRobot motor skill coordination with EM-based reinforcement learning\u201d in 2010 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IEEE 2010) pp. 3232\u20133237.","DOI":"10.1109\/IROS.2010.5649089"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1162\/NECO_a_00393"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.aav3123"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2024.3353075"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","unstructured":"S. Jin C. Wang M. Tomizuka \u201cRobust deformation model approximation for robotic cable manipulation\u201d in 2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE 2019) pp. 6586\u20136593.","DOI":"10.1109\/IROS40897.2019.8968157"},{"key":"e_1_3_2_32_2","unstructured":"V. Viswanath K. Shivakumar J. Ajmera M. Parulekar J. Kerr J. Ichnowski R. Cheng T. Kollar K. Goldberg HANDLOOM: Learned tracing of one-dimensional objects for inspection and manipulation. arXiv:2303.08975 [cs.RO] (2023)."},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","unstructured":"K. Shivakumar V. Viswanath A. Gu Y. Avigal J. Kerr J. Ichnowski R. Cheng T. Kollar K. Goldberg \u201cSGTM 2.0: Autonomously untangling long cables using interactive perception\u201d in 2023 IEEE International Conference on Robotics and Automation (ICRA) (IEEE 2023) pp. 5837\u20135843.","DOI":"10.1109\/ICRA48891.2023.10160574"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","unstructured":"V. Viswanath K. Shivakumar J. Kerr B. Thananjeyan E. Novoseller J. Ichnowski A. Escontrela M. Laskey J. E. Gonzalez K. Goldberg Autonomously untangling long cables. arXiv:2207.07813 [cs.RO] (2022).","DOI":"10.15607\/RSS.2022.XVIII.034"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10514-009-9120-4"},{"key":"e_1_3_2_36_2","first-page":"1","article-title":"End-to-end training of deep visuomotor policies","volume":"17","author":"Levine S.","year":"2016","unstructured":"S. Levine, C. Finn, T. Darrell, P. Abbeel, End-to-end training of deep visuomotor policies. J. Mach. Learn. Res. 17, 1\u201340 (2016).","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","unstructured":"J. Luo O. Sushkov R. Pevceviciute W. Lian C. Su M. Vecerik N. Ye S. Schaal J. Scholz \u201cRobust multi-modal policies for industrial assembly via reinforcement learning and demonstrations: A large-scale study\u201d in Proceedings of Robotics: Science and Systems (MIT Press Journals 2021) 10.15607\/RSS.2021.XVII.088.","DOI":"10.15607\/RSS.2021.XVII.088"},{"key":"e_1_3_2_38_2","unstructured":"Y. Yang K. Caluwaerts A. Iscen T. Zhang J. Tan V. Sindhwani \u201cData efficient reinforcement learning for legged robots\u201d in Conference on Robot Learning (PMLR) (MLResearchPress 2020) pp. 1\u201310."},{"key":"e_1_3_2_39_2","unstructured":"A. Zhan R. Zhao L. Pinto P. Abbeel M. Laskin \u201cA framework for efficient robotic manipulation\u201d in DEEP RL Workshop NeurIPS 2021 (2021) pp. 1\u201315."},{"key":"e_1_3_2_40_2","doi-asserted-by":"crossref","unstructured":"J. Tebbe L. Krauch Y. Gao A. Zell \u201cSample-efficient reinforcement learning in robotic table tennis\u201d in 2021 IEEE International Conference on Robotics and Automation (ICRA) (IEEE 2021) pp. 4171\u20134178.","DOI":"10.1109\/ICRA48506.2021.9560764"},{"key":"e_1_3_2_41_2","unstructured":"I. Popov N. Heess T. Lillicrap R. Hafner G. Barth-Maron M. Vecerik T. Lampe Y. Tassa T. Erez M. Riedmiller Data-efficient deep reinforcement learning for dexterous manipulation. arXiv:1704.03073 [cs.LG] (2017)."},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","unstructured":"J. Luo E. Solowjow C. Wen J. A. Ojea A. M. Agogino A. Tamar P. Abbeel \u201cReinforcement learning on variable impedance controller for high-precision robotic assembly\u201d in 2019 International Conference on Robotics and Automation (ICRA) (IEEE 2019) pp. 3080\u20133087.","DOI":"10.1109\/ICRA.2019.8793506"},{"key":"e_1_3_2_43_2","doi-asserted-by":"crossref","unstructured":"T. Z. Zhao J. Luo O. Sushkov R. Pevceviciute N. Heess J. Scholz S. Schaal S. Levine \u201cOffline meta-reinforcement learning for industrial insertion\u201d in 2022 International Conference on Robotics and Automation (ICRA) (IEEE 2022) pp. 6386\u20136393.","DOI":"10.1109\/ICRA46639.2022.9812312"},{"key":"e_1_3_2_44_2","unstructured":"Z. Hu A. Rovinsky J. Luo V. Kumar A. Gupta S. Levine REBOOT: Reuse data for bootstrapping efficient real-world dexterous manipulation. arXiv:2309.03322 [cs.LG] (2024)."},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","unstructured":"T. Johannink S. Bahl A. Nair J. Luo A. Kumar M. Loskyll J. A. Ojea E. Solowjow S. Levine \u201cResidual reinforcement learning for robot control\u201d in 2019 International Conference on Robotics and Automation (ICRA) (IEEE 2019) pp. 6023\u20136029.","DOI":"10.1109\/ICRA.2019.8794127"},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","unstructured":"H. Hu S. Mirchandani D. Sadigh Imitation bootstrapped reinforcement learning. arXiv:2311.02198 [cs.LG] (2024).","DOI":"10.15607\/RSS.2024.XX.056"},{"key":"e_1_3_2_47_2","doi-asserted-by":"crossref","unstructured":"A. Rajeswaran V. Kumar A. Gupta G. Vezzani J. Schulman E. Todorov S. Levine \u201cLearning complex dexterous manipulation with deep reinforcement learning and demonstrations\u201d in Proceedings of Robotics: Science and Systems (MIT Press Journals 2018) 10.15607\/RSS.2018.XIV.049.","DOI":"10.15607\/RSS.2018.XIV.049"},{"key":"e_1_3_2_48_2","doi-asserted-by":"crossref","unstructured":"G. Schoettler A. Nair J. Luo S. Bahl J. Aparicio Ojea E. Solowjow S. Levine \u201cDeep reinforcement learning for industrial insertion tasks with visual inputs and natural rewards\u201d in 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE 2020) pp. 5548\u20135555.","DOI":"10.1109\/IROS45743.2020.9341714"},{"key":"e_1_3_2_49_2","doi-asserted-by":"crossref","unstructured":"J. Luo Z. Hu C. Xu Y. L. Tan J. Berg A. Sharma S. Schaal C. Finn A. Gupta S. Levine \u201cSERL: A software suite for sample-efficient robotic reinforcement learning\u201d in 2024 IEEE International Conference on Robotics and Automation (ICRA) (IEEE 2024) pp. 16961\u201316969.","DOI":"10.1109\/ICRA57147.2024.10610040"},{"key":"e_1_3_2_50_2","unstructured":"I. Kostrikov L. M. Smith S. Levine \u201cDemonstrating a walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning\u201d in Proceedings of Robotics: Science and Systems (MIT Press Journals 2023) 10.15607\/RSS.2023.XIX.056."},{"key":"e_1_3_2_51_2","unstructured":"J. Luo P. Dong Y. Zhai Y. Ma S. Levine RLIF: Interactive imitation learning as reinforcement learning. arXiv:2311.12996 [cs.AI] (2023)."},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-012-5322-7"},{"key":"e_1_3_2_53_2","unstructured":"P. Wu A. Escontrela D. Hafner P. Abbeel K. Goldberg \u201cDayDreamer: World models for physical robot learning\u201d in Proceedings of the 6th Conference on Robot Learning (MLResearchPress 2023) pp. 2226\u20132240."},{"key":"e_1_3_2_54_2","unstructured":"A. Nagabandi K. Konolige S. Levine V. Kumar \u201cDeep dynamics models for learning dexterous manipulation\u201d in Proceedings of the Conference on Robot Learning (MLResearchPress 2020) pp. 1101\u20131112."},{"key":"e_1_3_2_55_2","unstructured":"R. Rafailov T. Yu A. Rajeswaran C. Finn \u201cOffline reinforcement learning from images with latent space models\u201d in Proceedings of the 3rd Conference on Learning for Dynamics and Control (MLResearchPress 2021) pp. 1154\u20131168."},{"key":"e_1_3_2_56_2","doi-asserted-by":"crossref","unstructured":"J. Luo E. Solowjow C. Wen J. A. Ojea A. M. Agogino \u201cDeep reinforcement learning for robotic assembly of mixed deformable and rigid objects\u201d in 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE 2018) pp. 2062\u20132069.","DOI":"10.1109\/IROS.2018.8594353"},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","unstructured":"H. Zhu A. Gupta A. Rajeswaran S. Levine V. Kumar \u201cDexterous manipulation with deep reinforcement learning: Efficient general and low-cost\u201d in International Conference on Robotics and Automation (ICRA) (IEEE 2019) pp. 3651\u20133657.","DOI":"10.1109\/ICRA.2019.8794102"},{"key":"e_1_3_2_58_2","unstructured":"J. Fu A. Singh D. Ghosh L. Yang S. Levine \u201cVariational inverse control with events: A general framework for data-driven reward definition\u201d in Advances in Neural Information Processing Systems 31 (NeurIPS 2018) (Curran Associates 2018) pp. 8538\u20138547."},{"key":"e_1_3_2_59_2","unstructured":"K. Li A. Gupta A. Reddy V. H. Pong A. Zhou J. Yu S. Levine \u201cMURAL: Meta-learning uncertainty-aware rewards for outcome-driven reinforcement learning\u201d in Proceedings of the 38th International Conference on Machine Learning ICML 2021 (MLResearchPress 2021) pp. 6346\u20136356."},{"key":"e_1_3_2_60_2","unstructured":"Y. Du K. Konyushkova M. Denil A. Raju J. Landon F. Hill N. de Freitas S. Cabi Vision-language models as success detectors. arXiv:2303.077280 [cs.LG] (2023)."},{"key":"e_1_3_2_61_2","unstructured":"P. Mahmoudieh D. Pathak T. Darrell \u201cZero-shot reward specification via grounded natural language\u201d in International Conference on Machine Learning ICML 2022 (MLResearchPress 2022) pp. 14743\u201314752."},{"key":"e_1_3_2_62_2","unstructured":"L. Fan G. Wang Y. Jiang A. Mandlekar Y. Yang H. Zhu A. Tang D. Huang Y. Zhu A. Anandkumar \u201cMineDojo: Building open-ended embodied agents with internet-scale knowledge\u201d in Advances in Neural Information Processing Systems 35 (Curran Associates 2022) pp. 18343\u201318362."},{"key":"e_1_3_2_63_2","unstructured":"Y. J. Ma S. Sodhani D. Jayaraman O. Bastani V. Kumar A. Zhang \u201cVIP: Towards universal visual reward and representation via value-implicit pre-training\u201d in The Eleventh International Conference on Learning Representations ICLR 2023 (ICLR 2023) pp. 1\u201335."},{"key":"e_1_3_2_64_2","unstructured":"Y. J. Ma V. Kumar A. Zhang O. Bastani D. Jayaraman \u201cLIV: Language-image representations and rewards for robotic control\u201d in International Conference on Machine Learning ICML 2023 (MLResearchPress 2023) pp. 23301\u201323320."},{"key":"e_1_3_2_65_2","doi-asserted-by":"crossref","unstructured":"A. Gupta J. Yu T. Z. Zhao V. Kumar A. Rovinsky K. Xu T. Devlin S. Levine \u201cReset-free reinforcement learning via multi-task learning: Learning dexterous manipulation behaviors without human intervention\u201d in IEEE International Conference on Robotics and Automation ICRA 2021 (IEEE 2021) pp. 6664\u20136671.","DOI":"10.1109\/ICRA48506.2021.9561384"},{"key":"e_1_3_2_66_2","unstructured":"A. Sharma K. Xu N. Sardana A. Gupta K. Hausman S. Levine C. Finn Autonomous reinforcement learning: Benchmarking and formalism. arXiv:2112.09605 [cs.LG] (2021)."},{"key":"e_1_3_2_67_2","unstructured":"H. Zhu J. Yu A. Gupta D. Shah K. Hartikainen A. Singh V. Kumar S. Levine \u201cThe ingredients of real world robotic reinforcement learning\u201d in 8th International Conference on Learning Representations ICLR 2020 (ICLR 2020) pp. 1\u201321."},{"key":"e_1_3_2_68_2","unstructured":"A. Xie F. Tajwar A. Sharma C. Finn \u201cWhen to ask for help: Proactive interventions in autonomous reinforcement learning\u201d in Advances in Neural Information Processing Systems 35 (NeurIPS 2022) (Curran Associates Inc. 2022) pp. 16918\u201316930."},{"key":"e_1_3_2_69_2","unstructured":"A. Sharma A. M. Ahmed R. Ahmad C. Finn Self-improving robots: End-to-end autonomous visuomotor reinforcement learning. arXiv:2303.01488 [cs.RO] (2023)."},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2020.2965869"},{"key":"e_1_3_2_71_2","unstructured":"S. Ross G. Gordon D. Bagnell \u201cA reduction of imitation learning and structured prediction to no-regret online learning\u201d in Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (2011) pp. 627\u2013635."},{"key":"e_1_3_2_72_2","doi-asserted-by":"crossref","unstructured":"M. Kelly C. Sidrane K. Driggs-Campbell M. J. Kochenderfer \u201cHG-DAgger: Interactive imitation learning with human experts\u201d in 2019 International Conference on Robotics and Automation (ICRA) (IEEE 2019) pp. 8077\u20138083.","DOI":"10.1109\/ICRA.2019.8793698"},{"key":"e_1_3_2_73_2","doi-asserted-by":"crossref","unstructured":"C. Chi Z. Xu S. Feng E. Cousineau Y. Du B. Burchfiel R. Tedrake S. Song Diffusion policy: Visuomotor policy learning via action diffusion. arXiv:2303.04137 [cs.RO] (2024).","DOI":"10.1177\/02783649241273668"},{"key":"e_1_3_2_74_2","unstructured":"V. A. Papavassiliou S. Russell \u201cConvergence of reinforcement learning with general function approximators\u201d in Proceedings of the 16th International Joint Conference on Artificial Intelligence (Morgan Kaufmann Publishers Inc. 1999) pp. 748\u2013755."},{"key":"e_1_3_2_75_2","unstructured":"J. Bhandari D. Russo R. Singal \u201cA finite time analysis of temporal di!erence learning with linear function approximation\u201d in Proceedings of the 31st Conference On Learning Theory (MLResearchPress 2018) pp. 1691\u20131692."},{"key":"e_1_3_2_76_2","unstructured":"C. Jin Z. Yang Z. Wang M. I. Jordan \u201cProvably efficient reinforcement learning with linear function approximation\u201d in Proceedings of Thirty Third Conference on Learning Theory (MLResearchPress 2020) pp. 2137\u20132143."},{"key":"e_1_3_2_77_2","unstructured":"L. F. Yang M. Wang Sample-optimal parametric Q-learning using linearly additive features. arXiv:1902.04779 [cs.LG] (2019)."},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.1177\/02783649922066385"},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364910369189"},{"key":"e_1_3_2_80_2","doi-asserted-by":"crossref","unstructured":"T. Marcucci R. Deits M. Gabiccini A. Bicchi R. Tedrake \u201cApproximate hybrid model predictive control for multi-contact push recovery in complex environments\u201d in 2017 IEEERAS 17th International Conference on Humanoid Robotics (Humanoids) (2017) pp. 31\u201338.","DOI":"10.1109\/HUMANOIDS.2017.8239534"},{"key":"e_1_3_2_81_2","unstructured":"F. R. Hogan A. Rodriguez Feedback control of the pusher-slider system: A story of hybrid and underactuated contact dynamics. arXiv:1611.08268 [cs.RO] (2016)."},{"key":"e_1_3_2_82_2","doi-asserted-by":"crossref","unstructured":"B. Aceituno-Cabezas A. Rodriguez \u201cA global quasi-dynamic model for contact-trajectory optimization\u201d in Proceedings of Robotics: Science and Systems (RSS Foundation 2020) 10.15607\/RSS.2020.XVI.047.","DOI":"10.15607\/RSS.2020.XVI.047"},{"key":"e_1_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1108\/09576059710159655"},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0272-6963(02)00108-0"},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.3390\/app13031687"},{"key":"e_1_3_2_86_2","doi-asserted-by":"crossref","unstructured":"A. Brohan N. Brown J. Carbajal Y. Chebotar J. Dabis C. Finn K. Gopalakrishnan K. Hausman A. Herzog J. Hsu J. Ibarz B. Ichter A. Irpan T. Jackson S. Jesmonth N. J. Joshi R. Julian D. Kalashnikov Y. Kuang I. Leal K.-H. Lee S. Levine Y. Lu U. Malla D. Manjunath I. Mordatch O. Nachum C. Parada J. Peralta E. Perez K. Pertsch J. Quiambao K. Rao M. Ryoo G. Salazar P. Sanketi K. Sayed J. Singh S. Sontakke A. Stone C. Tan H. Tran V. Vanhoucke S. Vega Q. Vuong F. Xia T. Xiao P. Xu S. Xu T. Yu B. Zitkovich RT-1: Robotics transformer for real-world control at scale. arXiv:2212.06817 (2023).","DOI":"10.15607\/RSS.2023.XIX.025"},{"key":"e_1_3_2_87_2","unstructured":"A. Brohan N. Brown J. Carbajal Y. Chebotar X. Chen K. Choromanski T. Ding D. Driess A. Dubey C. Finn P. Florence C. Fu M. G. Arenas K. Gopalakrishnan K. Han K. Hausman A. Herzog J. Hsu B. Ichter A. Irpan N. Joshi R. Julian D. Kalashnikov Y. Kuang I. Leal L. Lee T.-W. E. Lee S. Levine Y. Lu H. Michalewski I. Mordatch K. Pertsch K. Rao K. Reymann M. Ryoo G. Salazar P. Sanketi P. Sermanet J. Singh A. Singh R. Soricut H. Tran V. Vanhoucke Q. Vuong A. Wahid S. Welker P. Wohlhart J. Wu F. Xia T. Xiao P. Xu S. Xu T. Yu B. Zitkovich RT-2: Vision-language-action models transfer web knowledge to robotic control. arXiv:2307.15818 [cs.RO] (2023)."},{"key":"e_1_3_2_88_2","unstructured":"Open X-Embodiment Collaboration A. Padalkar A. Pooley A. Jain A. Bewley A. Herzog A. Irpan A. Khazatsky A. Rai A. Singh A. Brohan A. Raffin A.Wahid B. Burgess-Limerick B. Kim B. Sch\u00a8olkopf B. Ichter C. Lu C.Xu C. Finn C.Xu C. Chi C. Huang C. Chan C. Pan C. Fu C. Devin D. Driess D. Pathak D. Shah D. B\u00a8uchler D. Kalashnikov D. Sadigh E. Johns F. Ceola F. Xia F. Stulp G. Zhou G. S. Sukhatme G. Salhotra G. Yan G. Schiavi H. Su H.-S. Fang H. Shi H. B. Amor H. I. Christensen H. Furuta H.Walke H. Fang I. Mordatch I. Radosavovic I. Leal J. Liang J. Kim J. Schneider J. Hsu J. Bohg J. Bingham J. Wu J. Wu J. Luo J. Gu J. Tan J. Oh J. Malik J. Tompson J. Yang J. J. Lim J. Silverio J. Han K. Rao K. Pertsch K. Hausman K. Go K. Gopalakrishnan K. Goldberg K. Byrne K. Oslund K. Kawaharazuka K. Zhang K.Majd K. Rana K. Srinivasan L. Y. Chen L. Pinto L. Tan L. Ott L. Lee M. Tomizuka M. Du M. Ahn M. Zhang M. Ding M. K. Srirama M. Sharma M. J. Kim N. Kanazawa N. Hansen N. Heess N. J. Joshi N. Suenderhauf N. D. Palo N. M. M. Shafiullah O. Mees O. Kroemer P. R. Sanketi P. Wohlhart P. Xu P. Sermanet P. Sundaresan Q. Vuong R. Rafailov R. Tian R. Doshi R. Mart\u00edn-Mart\u00edn R. Mendonca R. Shah R. Hoque R. Julian S. Bustamante S. Kirmani S. Levine S. Moore S. Bahl S. Dass S. Song S. Xu S. Haldar S. Adebola S. Guist S. Nasiriany S. Schaal S. Welker S. Tian S. Dasari S. Belkhale T. Osa T. Harada T. Matsushima T. Xiao T. Yu T. Ding T. Davchev T. Z. Zhao T. Armstrong T. Darrell V. Jain V. Vanhoucke W. Zhan W. Zhou W. Burgard X. Chen X. Wang X. Zhu X. Li Y. Lu Y. Chebotar Y. Zhou Y. Zhu Y. Xu Y. Wang Y. Bisk Y. Cho Y. Lee Y. Cui Y. Wu Y. Tang Y. Zhu Y. Li Y. Iwasawa Y. Matsuo Z. Xu Z. J. Cui \u201cOpen X-embodiment: Robotic learning datasets and RT-X models\u201d in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) (IEEE 2024) pp. 6892\u20136903."},{"key":"e_1_3_2_89_2","unstructured":"Octo Model Team D. Ghosh H. Walke K. Pertsch K. Black O. Mees S. Dasari J. Hejna T. Kreiman C. Xu J. Luo Y. L. Tan L. Y. Chen P. Sanketi Q. Vuong T. Xiao D. Sadigh C. Finn S. Levine Octo: An open-source generalist robot policy. arXiv:2405.12213 [cs.RO] (2024)."},{"key":"e_1_3_2_90_2","unstructured":"M. J. Kim K. Pertsch S. Karamcheti T. Xiao A. Balakrishna S. Nair R. Rafailov E. Foster G. Lam P. Sanketi Q. Vuong T. Kollar B. Burchfiel R. Tedrake D. Sadigh S. Levine P. Liang C. Finn OpenVLA: An open-source vision-language-action model. arXiv:2406.09246 [cs.RO] (2024)."},{"key":"e_1_3_2_91_2","unstructured":"Y. Song Y. Zhou A. Sekhari D. Bagnell A. Krishnamurthy W. Sun \u201cHybrid RL: Using both offine and online data can make RL efficient \u201d poster presented at the 11th International Conference on Learning Representations Kigali Rwanda 1 May to 5 May 2023."},{"key":"e_1_3_2_92_2","unstructured":"V. Mnih K. Kavukcuoglu D. Silver A. Graves I. Antonoglou D. Wierstra M. Riedmiller Playing Atari with deep reinforcement learning. arXiv:1312.5602 [cs.LG] (2013)."},{"key":"e_1_3_2_93_2","unstructured":"T. Haarnoja A. Zhou P. Abbeel S. Levine \u201cSoft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor\u201d in Proceedings of the 35th International Conference on Machine Learning (MLResearchPress 2018) pp. 1861\u20131870."},{"key":"e_1_3_2_94_2","unstructured":"T. Haarnoja A. Zhou P. Abbeel S. Levine Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. arXiv:1801.01290 [cs.LG] (2018)."},{"key":"e_1_3_2_95_2","unstructured":"A. Radford J. W. Kim C. Hallacy A. Ramesh G. Goh S. Agarwal G. Sastry A. Askell P. Mishkin J. Clark G. Krueger I. Sutskever Learning transferable visual models from natural language supervision. arXiv:2103.00020 [cs.CV] (2021)."},{"key":"e_1_3_2_96_2","unstructured":"A. Dosovitskiy L. Beyer A. Kolesnikov D. Weissenborn X. Zhai T. Unterthiner M. Dehghani M. Minderer G. Heigold S. Gelly J. Uszkoreit N. Houlsby An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929 [cs.CV] (2021)."},{"key":"e_1_3_2_97_2","doi-asserted-by":"crossref","unstructured":"A. Kolesnikov L. Beyer X. Zhai J. Puigcerver J. Yung S. Gelly N. Houlsby Big transfer (BiT): General visual representation learning. arXiv:1912.11370 [cs.CV] (2020).","DOI":"10.1007\/978-3-030-58558-7_29"},{"key":"e_1_3_2_98_2","unstructured":"S. S. Du S. M. Kakade R.Wang L. F. Yang Is a good representation sufficient for sample efficient reinforcement learning? arXiv:1910.03016 [cs.LG] (2020)."},{"key":"e_1_3_2_99_2","unstructured":"K. He X. Zhang S. Ren J. Sun Deep residual learning for image recognition. arXiv:1512.03385 [cs.CV] (2023)."},{"key":"e_1_3_2_100_2","doi-asserted-by":"crossref","unstructured":"J. Deng W. Dong R. Socher L.-J. Li K. Li L. Fei-Fei \u201cImageNet: A large-scale hierarchical image database\u201d in 2009 IEEE Conference on Computer Vision and Pattern Recognition (IEEE 2009) pp. 248\u2013255.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_101_2","unstructured":"A. Y. Ng D. Harada S. J. Russell \u201cPolicy invariance under reward transformations: Theory and application reward shaping\u201d in Proceedings of the Sixteenth International Conference on Machine Learning ICML \u201899 (Morgan Kaufmann Publishers Inc. 1999) p. 278\u2013287."},{"key":"e_1_3_2_102_2","unstructured":"C. Florensa D. Held X. Geng P. Abbeel Automatic goal generation for reinforcement learning agents. arXiv:1705.06366 [cs.LG] (2018)."},{"key":"e_1_3_2_103_2","unstructured":"C. Florensa D. Held M.Wulfmeier M. Zhang P. Abbeel \u201cReverse curriculum generation for reinforcement learning\u201d in Proceedings of the 1st Annual Conference on Robot Learning (MLResearchPress 2017) pp. 482\u2013495."},{"key":"e_1_3_2_104_2","doi-asserted-by":"crossref","unstructured":"H. van Hasselt A. Guez D. Silver Deep reinforcement learning with double Q-learning. arXiv:1509.06461 [cs.LG] (2015).","DOI":"10.1609\/aaai.v30i1.10295"},{"key":"e_1_3_2_105_2","unstructured":"C. Jin Z. Allen-Zhu S. Bubeck M. Jordan \u201cIs Q-learning provably efficient?\u201d in Advances in Neural Information Processing Systems 31 (NeurIPS 2018) (Curran Associates 2018) pp. 4863\u20134873."},{"key":"e_1_3_2_106_2","unstructured":"M. G. Azar R. Munos B. Kappen On the sample complexity of reinforcement learning with a generative model. arXiv:1206.6461 [cs.LG] (2012)."},{"key":"e_1_3_2_107_2","unstructured":"M. J. Kearns S. P. Singh \u201cNear-optimal reinforcement learning in polynominal time\u201d in Proceedings of the Fifteenth International Conference on Machine Learning (Morgan Kaufmann Publishers Inc. 1998) pp. 260\u2013268."}],"container-title":["Science Robotics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.science.org\/doi\/pdf\/10.1126\/scirobotics.ads5033","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.science.org\/doi\/pdf\/10.1126\/scirobotics.ads5033","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,20]],"date-time":"2025-08-20T17:58:43Z","timestamp":1755712723000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.science.org\/doi\/10.1126\/scirobotics.ads5033"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,20]]},"references-count":106,"journal-issue":{"issue":"105","published-print":{"date-parts":[[2025,8,20]]}},"alternative-id":["10.1126\/scirobotics.ads5033"],"URL":"https:\/\/doi.org\/10.1126\/scirobotics.ads5033","relation":{},"ISSN":["2470-9476"],"issn-type":[{"value":"2470-9476","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,20]]},"assertion":[{"value":"2024-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-23","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"eads5033"}}