{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T09:56:13Z","timestamp":1777715773761,"version":"3.51.4"},"reference-count":48,"publisher":"SAGE Publications","issue":"10-11","license":[{"start":{"date-parts":[[2021,8,21]],"date-time":"2021-08-21T00:00:00Z","timestamp":1629504000000},"content-version":"vor","delay-in-days":365,"URL":"http:\/\/www.sagepub.com\/licence-information-for-chorus"}],"funder":[{"DOI":"10.13039\/100000006","name":"Office of Naval Research","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000006","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["EECS-0926052"],"award-info":[{"award-number":["EECS-0926052"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["IIS-1017134"],"award-info":[{"award-number":["IIS-1017134"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["IIS-1205249"],"award-info":[{"award-number":["IIS-1205249"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004189","name":"Max-Planck-Gesellschaft","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100004189","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004399","name":"Okawa Foundation for Information and Telecommunications","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100004399","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of Robotics Research"],"published-print":{"date-parts":[[2020,9]]},"abstract":"<jats:p>We present a strategy for simulation-to-real transfer, which builds on recent advances in robot skill decomposition. Rather than focusing on minimizing the simulation\u2013reality gap, we propose a method for increasing the sample efficiency and robustness of existing simulation-to-real approaches which exploits hierarchy and online adaptation. Instead of learning a unique policy for each desired robotic task, we learn a diverse set of skills and their variations, and embed those skill variations in a continuously parameterized space. We then interpolate, search, and plan in this space to find a transferable policy which solves more complex, high-level tasks by combining low-level skills and their variations. In this work, we first characterize the behavior of this learned skill space, by experimenting with several techniques for composing pre-learned latent skills. We then discuss an algorithm which allows our method to perform long-horizon tasks never seen in simulation, by intelligently sequencing short-horizon latent skills. Our algorithm adapts to unseen tasks online by repeatedly choosing new skills from the latent space, using live sensor data and simulation to predict which latent skill will perform best next in the real world. Importantly, our method learns to control a real robot in joint-space to achieve these high-level tasks with little or no on-robot time, despite the fact that the low-level policies may not be perfectly transferable from simulation to real, and that the low-level skills were not trained on any examples of high-level tasks. In addition to our results indicating a lower sample complexity for families of tasks, we believe that our method provides a promising template for combining learning-based methods with proven classical robotics algorithms such as model-predictive control.<\/jats:p>","DOI":"10.1177\/0278364920944474","type":"journal-article","created":{"date-parts":[[2020,8,21]],"date-time":"2020-08-21T08:12:11Z","timestamp":1597997531000},"page":"1259-1278","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":1,"title":["Scaling simulation-to-real transfer by learning a latent space of robot skills"],"prefix":"10.1177","volume":"39","author":[{"given":"Ryan C","family":"Julian","sequence":"first","affiliation":[{"name":"University of Southern California, Los Angeles, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eric","family":"Heiden","sequence":"additional","affiliation":[{"name":"University of Southern California, Los Angeles, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhanpeng","family":"He","sequence":"additional","affiliation":[{"name":"Columbia University, New York, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hejia","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Southern California, Los Angeles, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stefan","family":"Schaal","sequence":"additional","affiliation":[{"name":"X, Mountain View, CA USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joseph J","family":"Lim","sequence":"additional","affiliation":[{"name":"University of Southern California, Los Angeles, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gaurav S","family":"Sukhatme","sequence":"additional","affiliation":[{"name":"University of Southern California, Los Angeles, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Karol","family":"Hausman","sequence":"additional","affiliation":[{"name":"Google Brain, Robotics at Google Team, Mountain View, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2020,8,21]]},"reference":[{"key":"bibr1-0278364920944474","first-page":"5048","author":"Andrychowicz M","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr2-0278364920944474","author":"Barrett S","year":"2010","journal-title":"Ninth International Conference on Autonomous Agents and Multiagent Systems - Adaptive Learning Agents Workshop (AAMAS - ALA)"},{"key":"bibr3-0278364920944474","first-page":"703","author":"Chebotar Y","year":"2017","journal-title":"Proceedings of the 34th International Conference on Machine Learning"},{"key":"bibr4-0278364920944474","author":"Clavera I","year":"2019","journal-title":"International Conference on Learning Representations"},{"key":"bibr5-0278364920944474","author":"Co-Reyes JD","year":"2018","journal-title":"Proceedings of the International Conference on Machine Learning"},{"issue":"5","key":"bibr6-0278364920944474","first-page":"12","volume":"33","author":"Cooper PA","year":"1993","journal-title":"Educational Technology"},{"key":"bibr7-0278364920944474","doi-asserted-by":"publisher","DOI":"10.7551\/mitpress\/4378.001.0001"},{"key":"bibr8-0278364920944474","first-page":"1329","author":"Duan Y","year":"2016","journal-title":"International Conference on Machine Learning"},{"key":"bibr9-0278364920944474","author":"Duan Y","year":"2016","journal-title":"CoRR"},{"key":"bibr10-0278364920944474","author":"Eysenbach B","year":"2019","journal-title":"International Conference on Learning Representations"},{"key":"bibr11-0278364920944474","first-page":"1126","author":"Finn C","year":"2017","journal-title":"Proceedings of the 34th International Conference on Machine Learning"},{"key":"bibr12-0278364920944474","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2017.7989385"},{"key":"bibr13-0278364920944474","first-page":"5302","author":"Gupta A","year":"2018","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr14-0278364920944474","first-page":"1846","author":"Haarnoja T","year":"2018","journal-title":"International Conference on Machine Learning"},{"key":"bibr15-0278364920944474","author":"Hausman K","year":"2018","journal-title":"International Conference on Learning Representations"},{"key":"bibr16-0278364920944474","author":"Heess N","year":"2016","journal-title":"CoRR"},{"key":"bibr17-0278364920944474","author":"James S","year":"2017","journal-title":"Conference on Robot Learning (CoRL)"},{"key":"bibr18-0278364920944474","author":"Julian RC","year":"2018","journal-title":"International Symposium on Experimental Robotics"},{"key":"bibr19-0278364920944474","first-page":"1701","author":"Kamthe S","year":"2018","journal-title":"International Conference on Artificial Intelligence and Statistics"},{"key":"bibr20-0278364920944474","author":"Kroemer O","year":"2016","journal-title":"CoRR"},{"key":"bibr21-0278364920944474","author":"Lee Y","year":"2019","journal-title":"International Conference on Learning Representations"},{"issue":"1","key":"bibr22-0278364920944474","first-page":"1334","volume":"17","author":"Levine S","year":"2016","journal-title":"The Journal of Machine Learning Research"},{"key":"bibr23-0278364920944474","doi-asserted-by":"publisher","DOI":"10.1177\/0278364917710318"},{"key":"bibr24-0278364920944474","author":"Lillicrap TP","year":"2015","journal-title":"CoRR"},{"key":"bibr25-0278364920944474","author":"Mishra N","year":"2017","journal-title":"CoRR"},{"key":"bibr26-0278364920944474","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"bibr27-0278364920944474","doi-asserted-by":"publisher","DOI":"10.1109\/HUMANOIDS.2012.6651537"},{"key":"bibr28-0278364920944474","author":"Peng XB","year":"2017","journal-title":"CoRR"},{"key":"bibr29-0278364920944474","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8460528"},{"key":"bibr30-0278364920944474","author":"Preiss JA","year":"2018","journal-title":"NeurIPS Workshop on Reinforcement Learning under Partial Observability"},{"key":"bibr31-0278364920944474","author":"Rajendran J","year":"2015","journal-title":"CoRR"},{"key":"bibr32-0278364920944474","volume-title":"Conference Track Proceedings of the 5th International Conference on Learning Representations (ICLR 2017)","author":"Rajeswaran A","year":"2017"},{"key":"bibr33-0278364920944474","author":"Rakelly K","year":"2019","journal-title":"arXiv preprint arXiv:1903.08254"},{"key":"bibr34-0278364920944474","unstructured":"Reinforcement Learning Working Group Contributors (2019) Garage: A toolkit for reproducible reinforcement learning. Available at: https:\/\/github.com\/rlworkgroup\/garage."},{"key":"bibr35-0278364920944474","author":"Rothfuss J","year":"2019","journal-title":"International Conference on Learning Representations"},{"key":"bibr36-0278364920944474","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2015.7139390"},{"key":"bibr37-0278364920944474","volume-title":"Artificial Intelligence: A Modern Approach","author":"Russell SJ","year":"2016"},{"key":"bibr38-0278364920944474","author":"Rusu AA","year":"2016","journal-title":"CoRR"},{"key":"bibr39-0278364920944474","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2017.XIII.034"},{"key":"bibr40-0278364920944474","first-page":"642","volume-title":"Proceedings of the Thirty-Fourth Conference on Uncertainty in Artificial Intelligence (UAI 2018)","author":"S\u00e6mundsson S","year":"2018"},{"key":"bibr41-0278364920944474","author":"Schulman J","year":"2017","journal-title":"CoRR"},{"key":"bibr42-0278364920944474","volume-title":"Process dynamics and control","author":"Seborg DE","year":"2010"},{"key":"bibr43-0278364920944474","first-page":"23","author":"Tobin J","year":"2017","journal-title":"International Conference on Intelligent Robots and Systems (IROS)"},{"key":"bibr44-0278364920944474","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2012.6386109"},{"key":"bibr45-0278364920944474","author":"Tzeng E","year":"2015","journal-title":"CoRR"},{"key":"bibr46-0278364920944474","author":"Visser A","year":"2011","journal-title":"The International Micro Air Vehicles Conference"},{"key":"bibr47-0278364920944474","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2017.XIII.048"},{"key":"bibr48-0278364920944474","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2016.7487175"}],"container-title":["The International Journal of Robotics Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/0278364920944474","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/0278364920944474","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/0278364920944474","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/0278364920944474","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T10:16:19Z","timestamp":1777457779000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/0278364920944474"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,8,21]]},"references-count":48,"journal-issue":{"issue":"10-11","published-print":{"date-parts":[[2020,9]]}},"alternative-id":["10.1177\/0278364920944474"],"URL":"https:\/\/doi.org\/10.1177\/0278364920944474","relation":{},"ISSN":["0278-3649","1741-3176"],"issn-type":[{"value":"0278-3649","type":"print"},{"value":"1741-3176","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,8,21]]}}}