{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,9,8]],"date-time":"2026-09-08T15:05:19Z","timestamp":1788879919250,"version":"build-2803163510"},"reference-count":61,"publisher":"American Association for the Advancement of Science (AAAS)","issue":"89","content-domain":{"domain":["www.science.org"],"crossmark-restriction":true},"short-container-title":["Sci. Robot."],"published-print":{"date-parts":[[2024,4,17]]},"abstract":"<jats:p>Humanoid robots that can autonomously operate in diverse environments have the potential to help address labor shortages in factories, assist elderly at home, and colonize new planets. Although classical controllers for humanoid robots have shown impressive results in a number of settings, they are challenging to generalize and adapt to new environments. Here, we present a fully learning-based approach for real-world humanoid locomotion. Our controller is a causal transformer that takes the history of proprioceptive observations and actions as input and predicts the next action. We hypothesized that the observation-action history contains useful information about the world that a powerful transformer model can use to adapt its behavior in context, without updating its weights. We trained our model with large-scale model-free reinforcement learning on an ensemble of randomized environments in simulation and deployed it to the real-world zero-shot. Our controller could walk over various outdoor terrains, was robust to external disturbances, and could adapt in context.<\/jats:p>","DOI":"10.1126\/scirobotics.adi9579","type":"journal-article","created":{"date-parts":[[2024,4,17]],"date-time":"2024-04-17T13:59:52Z","timestamp":1713362392000},"update-policy":"https:\/\/doi.org\/10.34133\/aaas_crossmark","source":"Crossref","is-referenced-by-count":212,"title":["Real-world humanoid locomotion with reinforcement learning"],"prefix":"10.1126","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-9200-5980","authenticated-orcid":true,"given":"Ilija","family":"Radosavovic","sequence":"first","affiliation":[{"name":"University of California, Berkeley CA, USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6440-2393","authenticated-orcid":true,"given":"Tete","family":"Xiao","sequence":"additional","affiliation":[{"name":"University of California, Berkeley CA, USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2652-4318","authenticated-orcid":true,"given":"Bike","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of California, Berkeley CA, USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Trevor","family":"Darrell","sequence":"additional","affiliation":[{"name":"University of California, Berkeley CA, USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3695-1580","authenticated-orcid":true,"given":"Jitendra","family":"Malik","sequence":"additional","affiliation":[{"name":"University of California, Berkeley CA, USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5346-3637","authenticated-orcid":true,"given":"Koushil","family":"Sreenath","sequence":"additional","affiliation":[{"name":"University of California, Berkeley CA, USA."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"221","reference":[{"key":"e_1_3_2_2_2","unstructured":"I. Kato Development of WABOT 1 in Biomechanism (University of Tokyo Press 1973)."},{"key":"e_1_3_2_3_2","doi-asserted-by":"crossref","unstructured":"K. Hirai M. Hirose Y. Haikawa T. Takenaka The development of Honda humanoid robot in IEEE International Conference on Robotics and Automation (ICRA) (IEEE 1998) vol. 2 pp. 1321\u20131326.","DOI":"10.1109\/ROBOT.1998.677288"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.7210\/jrsj.30.372"},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","unstructured":"O. Stasse T. Flayols R. Budhiraja K. Giraud-Esclasse J. Carpentier J. Mirabel A. Del Prete P. Sou\u00e8res N. Mansard F. Lamiraux J. P. Laumond L. Marchionni H. Tome F. Ferro TALOS: A new humanoid research platform targeted for industrial applications in IEEE-RAS 17th International Conference on Humanoid Robotics (Humanoids) (IEEE 2017) pp. 689\u2013695.","DOI":"10.1109\/HUMANOIDS.2017.8246947"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","unstructured":"M. Chignoli D. Kim E. Stanger-Jones S. Kim The MIT humanoid robot: Design motion planning and control for acrobatic behaviors in IEEE-RAS 20th International Conference on Humanoid Robots (Humanoids) (IEEE 2021) pp. 1\u20138.","DOI":"10.1109\/HUMANOIDS47582.2021.9555782"},{"key":"e_1_3_2_7_2","doi-asserted-by":"crossref","unstructured":"M. H. Raibert Legged Robots That Balance (MIT Press 1986).","DOI":"10.1109\/MEX.1986.4307016"},{"key":"e_1_3_2_8_2","unstructured":"S. Kajita F. Kanehiro K. Kaneko K. Yokoi H. Hirukawa The 3D linear inverted pendulum mode: A simple modeling for a biped walking pattern generation in IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE 2001)."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.2002.806653"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.1107799"},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Y. Tassa T. Erez E. Todorov Synthesis and stabilization of complex behaviors through online trajectory optimization in IEEE\/RSJ International Conference on Intelligent Robots and Systems (IEEE 2012) pp. 4906\u20134913.","DOI":"10.1109\/IROS.2012.6386025"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10514-015-9479-3"},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","unstructured":"J. Di Carlo P. M. Wensing B. Katz G. Bledt S. Kim Dynamic locomotion in the MIT Cheetah 3 through convex model-predictive control in IEEE\/RSJ International Conference on Intelligent Robots And Systems (IROS) (IEEE 2018) pp. 1\u20139.","DOI":"10.1109\/IROS.2018.8594448"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364919887447"},{"key":"e_1_3_2_15_2","unstructured":"OpenAI I. Akkaya M. Andrychowicz M. Chociej M. Litwin B. McGrew A. Petron A. Paino M. Plappert G. Powell R. Ribas J. Schneider N. Tezak J. Tworek P. Welinder L. Weng Q. Yuan W. Zaremba L. Zhang Solving Rubik\u2019s cube with a robot hand. arXiv:1910.07113 (2019)."},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","unstructured":"A. Handa A. Allshire V. Makoviychuk A. Petrenko R. Singh J. Liu D. Makoviichuk K. Van Wyk A. Zhurkevich B. Sundaralingam Y. Narang DeXtreme: Transfer of agile in-hand manipulation from simulation to reality in Proceedings of the 2023 IEEE International Conference on Robotics and Automation (ICRA) (IEEE 2023) pp. 5977\u20135984.","DOI":"10.1109\/ICRA48891.2023.10160216"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.aau5872"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.abc5986"},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","unstructured":"A. Kumar Z. Fu D. Pathak J. Malik RMA: Rapid motor adaptation for legged robots Proceedings of the Robotics: Science and Systems (RSS); Virtual Event 12 to 16 July 2021.","DOI":"10.15607\/RSS.2021.XVII.011"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0921-8890(97)00043-2"},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","unstructured":"R. Tedrake T. W. Zhang H. S. Seung Stochastic policy gradient reinforcement learning on a simple 3D biped in 2004 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE 2004) vol. 3 pp. 2849\u20132854.","DOI":"10.1109\/IROS.2004.1389841"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"Z. Xie G. Berseth P. Clary J. Hurst M. van de Panne Feedback control for cassie with deep reinforcement learning in IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE 2018) pp. 1241\u20131246.","DOI":"10.1109\/IROS.2018.8593722"},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","unstructured":"J. Siekmann Y. Godse A. Fern J. Hurst Sim-to-real learning of all common bipedal gaits via periodic reward composition in IEEE International Conference on Robotics and Automation (ICRA) (IEEE 2021) pp. 7309\u20137315.","DOI":"10.1109\/ICRA48506.2021.9561814"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"J. Siekmann K. Green J. Warila A. Fern J. Hurst Blind bipedal stair traversal via sim-to-real reinforcement learning in Proceedings of the Robotics: Science and Systems (RSS) (RSS 2021).","DOI":"10.15607\/RSS.2021.XVII.061"},{"key":"e_1_3_2_25_2","unstructured":"S. Iida S. Kato K. Kuwayama T. Kunitachi M. Kanoh H. Itoh Humanoid robot control based on reinforcement learning in Micro-Nanomechatronics and Human Science 2004 and The Fourth Symposium Micro-Nanomechatronics for Information-Based Society 2004 (IEEE 2004) pp. 353\u2013358."},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"D. Rodriguez S. Behnke Deepwalk: Omnidirectional bipedal gait by deep reinforcement learning in 2021 IEEE International Conference on Robotics and Automation (ICRA) (IEEE 2021) pp. 3033\u20133039.","DOI":"10.1109\/ICRA48506.2021.9561717"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2022.3151771"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2022.3143227"},{"key":"e_1_3_2_29_2","unstructured":"R. Antonova S. Cruciani C. Smith D. Kragic Reinforcement learning for pivoting task arXiv:1703.00472 (2017)."},{"key":"e_1_3_2_30_2","doi-asserted-by":"crossref","unstructured":"F. Sadeghi S. Levine Cad2rl: Real single-image flight without a single real image in Proceedings of the Robotics: Science and Systems (RSS) (RSS 2016).","DOI":"10.15607\/RSS.2017.XIII.034"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","unstructured":"J. Tobin R. Fong A. Ray J. Schneider W. Zaremba P. Abbeel Domain randomization for transferring deep neural networks from simulation to the real world in IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE 2017) pp. 23\u201330.","DOI":"10.1109\/IROS.2017.8202133"},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","unstructured":"X. B. Peng M. Andrychowicz W. Zaremba P. Abbeel Sim-to-real transfer of robotic control with dynamics randomization in IEEE International Conference on Robotics and Automation (ICRA) (IEEE 2018) pp. 3803\u20133810.","DOI":"10.1109\/ICRA.2018.8460528"},{"key":"e_1_3_2_33_2","unstructured":"T. Brown B. Mann N. Ryder M. Subbiah J. D. Kaplan P. Dhariwal A. Neelakantan P. Shyam G. Sastry A. Askell S. Agarwal A. Herbert-Voss G. Krueger T. Henighan R. Child A. Ramesh D. M. Ziegler J. Wu C. Winter C. Hesse M. Chen E. Sigler M. Litwin S. Gray B. Chess J. Clark C. Berner S. M. Candlish A. Radford I. Sutskever D. Amodei Language models are few-shot learners in Proceedings of the 34th International Conference on Neural Information Processing Systems (NeurIPS 2020) pp. 1877\u20131901."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2022.3158231"},{"key":"e_1_3_2_35_2","doi-asserted-by":"crossref","unstructured":"G. A. Castillo B. Weng W. Zhang A. Hereid Robust feedback motion policy design using reinforcement learning on a 3D digit bipedal robot in IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE 2021) pp. 5136\u20135143.","DOI":"10.1109\/IROS51168.2021.9636467"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","unstructured":"Y. Gao Y. Gong V. Paredes A. Hereid Y. Gu Time-varying alip model and robust foot-placement control for underactuated bipedal robotic walking on a swaying rigid surface in 2023 American Control Conference (ACC) (IEEE 2023) pp. 3282\u20133287.","DOI":"10.23919\/ACC55779.2023.10156254"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMECH.2022.3176015"},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","unstructured":"A. Adu-Bredu N. Devraj O. C. Jenkins Optimal constrained task planning as mixed integer programming in IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE 2022) pp. 12029\u201312036.","DOI":"10.1109\/IROS47612.2022.9981237"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2022.3204367"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2023.3299524"},{"key":"e_1_3_2_41_2","unstructured":"D. J. Morton D. D. Fuller Human Locomotion and Body Form: A Study of Gravity and Man (Williams & Wilkins 1952)."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1242\/jeb.008573"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1098\/rspb.2009.0664"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbiomech.2008.06.039"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbiomech.2008.05.024"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1093\/ptj\/47.4.272"},{"key":"e_1_3_2_47_2","unstructured":"J. Kaplan S. McCandlish T. Henighan T. B. Brown B. Chess R. Child S. Gray A. Radford J.Wu D. Amodei Scaling laws for neural language models. arXiv:2001.08361 (2020)."},{"key":"e_1_3_2_48_2","first-page":"23716","article-title":"Flamingo: A visual language model for few-shot learning","volume":"35","author":"Alayrac J.-B.","year":"2022","unstructured":"J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Bi\u0144kowski, R. Barreira, O. Vinyals, A. Zisserman, K. Simonyan, Flamingo: A visual language model for few-shot learning. Adv. Neural Inf. Process. Syst. 35, 23716 (2022).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_49_2","unstructured":"A. Dosovitskiy L. Beyer A. Kolesnikov D. Weissenborn X. Zhai T. Unterthiner M. Dehghani M. Minderer G. Heigold S. Gelly J. Uszkoreit N. Houlsby An image is worth 16x16 words: Transformers for image recognition at scale in International Conference on Learning Representations (2021)."},{"key":"e_1_3_2_50_2","unstructured":"A. Radford K. Narasimhan T. Salimans I. Sutskever Improving language understanding by generative pre-training (2018)."},{"key":"e_1_3_2_51_2","unstructured":"A. Vaswani N. Shazeer N. Parmar J. Uszkoreit L. Jones A. N. Gomez \u0141. Kaiser I. Polosukhin Attention is all you need Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS 2017) pp. 6000\u20136010."},{"key":"e_1_3_2_52_2","unstructured":"J. Devlin M.-W. Chang K. Lee K. Toutanova BERT: Pre-training of deep bidirectional transformers for language understanding arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_53_2","doi-asserted-by":"crossref","unstructured":"L. Dong S. Xu B. Xu Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition in IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP) (IEEE 2018) pp. 5884\u20135888.","DOI":"10.1109\/ICASSP.2018.8462506"},{"key":"e_1_3_2_54_2","doi-asserted-by":"crossref","unstructured":"N. Carion F. Massa G. Synnaeve N. Usunier A. Kirillov S. Zagoruyko End-to-end object detection with transformers in ECCV 2020: 16th European Conference vol. 12346 of Lecture Notes in Computer Science (Springer 2020) pp. 213\u2013229.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_2_55_2","unstructured":"J. Schulman F. Wolski P. Dhariwal A. Radford O. Klimov Proximal policy optimization algorithms arXiv:1707.06347 (2017)."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2016.2582731"},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","unstructured":"Y. Gong R. Hartley X. Da A. Hereid O. Harib J.-K. Huang J. Grizzle Feedback control of a cassie bipedal robot: Walking standing and riding a Segway in 2019 American Control Conference (ACC) (IEEE 2019) pp. 4559\u20134566.","DOI":"10.23919\/ACC.2019.8814833"},{"key":"e_1_3_2_58_2","unstructured":"V. Makoviychuk L. Wawrzyniak Y. Guo M. Lu K. Storey M. Macklin D. Hoeller N. Rudin A. Allshire A. Handa G. State Isaac Gym: High performance GPU-based physics simulation for robot learning arXiv:2108.10470 (2021)."},{"key":"e_1_3_2_59_2","unstructured":"N. Rudin D. Hoeller P. Reist M. Hutter Learning to walk in minutes using massively parallel deep reinforcement learning in Conference on Robot Learning (MLResearchPress 2022)."},{"key":"e_1_3_2_60_2","unstructured":"S. Bai J. Z. Kolter V. Koltun An empirical evaluation of generic convolutional and recurrent networks for sequence modeling arXiv:1803.01271 (2018)."},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_62_2","doi-asserted-by":"crossref","unstructured":"J. Tan T. Zhang E. Coumans A. Iscen Y. Bai D. Hafner S. Bohez and V. Vanhoucke Sim-to-real: Learning agile locomotion for quadruped robots in Proceedings of the Robotics: Science and Systems (RSS) (RSS 2018).","DOI":"10.15607\/RSS.2018.XIV.010"}],"container-title":["Science Robotics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.science.org\/doi\/pdf\/10.1126\/scirobotics.adi9579","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,16]],"date-time":"2024-11-16T08:00:31Z","timestamp":1731744031000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.science.org\/doi\/10.1126\/scirobotics.adi9579"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,17]]},"references-count":61,"journal-issue":{"issue":"89","published-print":{"date-parts":[[2024,4,17]]}},"alternative-id":["10.1126\/scirobotics.adi9579"],"URL":"https:\/\/doi.org\/10.1126\/scirobotics.adi9579","relation":{},"ISSN":["2470-9476"],"issn-type":[{"value":"2470-9476","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,17]]},"assertion":[{"value":"2023-05-31","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-26","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"eadi9579"}}