{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T16:20:30Z","timestamp":1781194830465,"version":"3.54.1"},"reference-count":64,"publisher":"SAGE Publications","issue":"12-14","license":[{"start":{"date-parts":[[2021,10,11]],"date-time":"2021-10-11T00:00:00Z","timestamp":1633910400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"name":"2020 HAI-AWS Cloud Credits Grants"},{"DOI":"10.13039\/100016680","name":"toyota research institute, north america","doi-asserted-by":"publisher","award":["138695"],"award-info":[{"award-number":["138695"]}],"id":[{"id":"10.13039\/100016680","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of Robotics Research"],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:p>We aim to endow a robot with the ability to learn manipulation concepts that link natural language instructions to motor skills. Our goal is to learn a single multi-task policy that takes as input a natural language instruction and an image of the initial scene and outputs a robot motion trajectory to achieve the specified task. This policy has to generalize over different instructions and environments. Our insight is that we can approach this problem through learning from demonstration by leveraging large-scale video datasets of humans performing manipulation actions. Thereby, we avoid more time-consuming processes such as teleoperation or kinesthetic teaching. We also avoid having to manually design task-specific rewards. We propose a two-stage learning process where we first learn single-task policies through reinforcement learning. The reward is provided by scoring how well the robot visually appears to perform the task. This score is given by a video-based action classifier trained on a large-scale human activity dataset. In the second stage, we train a multi-task policy through imitation learning to imitate all the single-task policies. In extensive simulation experiments, we show that the multi-task policy learns to perform a large percentage of the 78 different manipulation tasks on which it was trained. The tasks are of greater variety and complexity than previously considered robot manipulation tasks. We show that the policy generalizes over variations of the environment. We also show examples of successful generalization over novel but similar instructions.<\/jats:p>","DOI":"10.1177\/02783649211046285","type":"journal-article","created":{"date-parts":[[2021,10,11]],"date-time":"2021-10-11T03:16:59Z","timestamp":1633922219000},"page":"1419-1434","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":68,"title":["Concept2Robot: Learning manipulation concepts from instructions and human demonstrations"],"prefix":"10.1177","volume":"40","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3412-5755","authenticated-orcid":false,"given":"Lin","family":"Shao","sequence":"first","affiliation":[{"name":"Stanford AI Lab, Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Toki","family":"Migimatsu","sequence":"additional","affiliation":[{"name":"Stanford AI Lab, Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qiang","family":"Zhang","sequence":"additional","affiliation":[{"name":"Zhiyuan College, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Karen","family":"Yang","sequence":"additional","affiliation":[{"name":"Stanford AI Lab, Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jeannette","family":"Bohg","sequence":"additional","affiliation":[{"name":"Stanford AI Lab, Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2021,10,11]]},"reference":[{"key":"bibr1-02783649211046285","first-page":"3674","author":"Anderson P","year":"2018","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"bibr2-02783649211046285","first-page":"166","volume":"70","author":"Andreas J","year":"2017","journal-title":"Proceedings of the 34th International Conference on Machine Learning"},{"key":"bibr3-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2008.10.024"},{"key":"bibr4-02783649211046285","author":"Bahdanau D","year":"2018","journal-title":"arXiv preprint arXiv:1806.01946"},{"key":"bibr5-02783649211046285","unstructured":"Coumans E, Bai Y (2016\u20132019) Pybullet, a Python module for physics simulation for games, robotics and machine learning. http:\/\/pybullet.org."},{"key":"bibr6-02783649211046285","author":"Das N","year":"2020","journal-title":"Conference on Robot Learning (CoRL)"},{"key":"bibr7-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1007\/s10479-005-5724-z"},{"key":"bibr8-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2017.7989250"},{"key":"bibr9-02783649211046285","author":"Devlin J","year":"2018","journal-title":"arXiv preprint arXiv:1810.04805"},{"key":"bibr10-02783649211046285","first-page":"2625","author":"Donahue J","year":"2015","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"bibr11-02783649211046285","first-page":"1473","author":"Fang H","year":"2015","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"bibr12-02783649211046285","first-page":"8538","author":"Fu J","year":"2018","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr13-02783649211046285","author":"Fujimoto S","year":"2018","journal-title":"arXiv preprint arXiv:1802.09477"},{"key":"bibr14-02783649211046285","first-page":"1","volume-title":"2018 IEEE International Conference on Robotics and Automation (ICRA)","author":"Gams A","year":"2018"},{"key":"bibr15-02783649211046285","author":"Goyal R","year":"2017","journal-title":"CoRR"},{"key":"bibr16-02783649211046285","author":"Gupta A","year":"2019","journal-title":"Conference on Robot Learning (CoRL)"},{"key":"bibr17-02783649211046285","first-page":"770","author":"He K","year":"2016","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"bibr18-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1109\/ROBOT.2002.1014739"},{"key":"bibr19-02783649211046285","author":"Jiang Y","year":"2019","journal-title":"arXiv preprint arXiv:1906.07343"},{"key":"bibr20-02783649211046285","first-page":"9414","author":"Jiang Y","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr21-02783649211046285","volume-title":"Proceedings of the 1st AAAI Conference on Bridging the Gap Between Task and Motion Planning (AAAIWS\u201910-01)","author":"Kaelbling LP","year":"2010"},{"key":"bibr22-02783649211046285","author":"Kalashnikov D","year":"2018","journal-title":"arXiv preprint arXiv:1806.10293"},{"key":"bibr23-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2011.2159412"},{"key":"bibr24-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2011.2163863"},{"key":"bibr25-02783649211046285","first-page":"817","volume-title":"Proceedings of the 20th International Conference on Neural Information Processing Systems (NIPS\u201907)","author":"Langford J","year":"2007"},{"key":"bibr26-02783649211046285","author":"L\u00e1zaro-Gredilla M","year":"2018","journal-title":"arXiv preprint arXiv:1812.02788"},{"key":"bibr27-02783649211046285","author":"Lillicrap TP","year":"2015","journal-title":"arXiv preprint arXiv:1509.02971"},{"key":"bibr28-02783649211046285","author":"Ma CY","year":"2019","journal-title":"Proceedings of the International Conference on Learning Representations (ICLR)"},{"key":"bibr29-02783649211046285","volume-title":"Concepts: Core Readings","author":"Margolis E","year":"1999"},{"key":"bibr30-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2010.5651049"},{"key":"bibr31-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2011.6094676"},{"key":"bibr32-02783649211046285","author":"Misra D","year":"2018","journal-title":"arXiv preprint arXiv:1809.00786"},{"key":"bibr33-02783649211046285","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2020.XVI.080"},{"key":"bibr34-02783649211046285","first-page":"2616","author":"Paraschos A","year":"2013","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr35-02783649211046285","author":"Parisotto E","year":"2015","journal-title":"arXiv preprint arXiv:1511.06342"},{"key":"bibr36-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1109\/ROBOT.2009.5152385"},{"key":"bibr37-02783649211046285","author":"Pourchot | Sigaud","year":"2019","journal-title":"International Conference on Learning Representations"},{"key":"bibr38-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4757-4321-0"},{"key":"bibr39-02783649211046285","author":"Rusu AA","year":"2015","journal-title":"arXiv preprint arXiv:1511.06295"},{"key":"bibr40-02783649211046285","first-page":"527","author":"Sener O","year":"2018","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr41-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8462891"},{"key":"bibr42-02783649211046285","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2020.XVI.082"},{"key":"bibr43-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2018.2856525"},{"key":"bibr44-02783649211046285","author":"Shao L","year":"2020","journal-title":"arXiv preprint arXiv:2009.08973"},{"key":"bibr45-02783649211046285","first-page":"10740","author":"Shridhar M","year":"2020","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"bibr46-02783649211046285","author":"Shu T","year":"2017","journal-title":"arXiv preprint arXiv:1712.07294"},{"key":"bibr47-02783649211046285","author":"Simonyan K","year":"2014","journal-title":"arXiv preprint arXiv:1409.1556"},{"key":"bibr48-02783649211046285","author":"Singh A","year":"2019","journal-title":"arXiv preprint arXiv:1904.07854"},{"key":"bibr49-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1111\/j.1551-6709.2010.01129.x"},{"key":"bibr50-02783649211046285","author":"Smith L","year":"2020","journal-title":"arXiv preprint arXiv:1912.04443"},{"key":"bibr51-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1177\/105971230501300102"},{"key":"bibr52-02783649211046285","first-page":"4496","author":"Teh Y","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr53-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.10744"},{"key":"bibr54-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2017.8202133"},{"key":"bibr55-02783649211046285","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2018.XIV.044"},{"key":"bibr56-02783649211046285","first-page":"4489","author":"Tran D","year":"2015","journal-title":"Proceedings of the IEEE International Conference on Computer Vision"},{"key":"bibr57-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2010.2065430"},{"key":"bibr58-02783649211046285","first-page":"3156","author":"Vinyals O","year":"2015","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"bibr59-02783649211046285","first-page":"6629","author":"Wang X","year":"2019","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"bibr60-02783649211046285","first-page":"37","author":"Wang X","year":"2018","journal-title":"Proceedings of the European Conference on Computer Vision (ECCV)"},{"key":"bibr61-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8794024"},{"key":"bibr62-02783649211046285","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_5"},{"key":"bibr63-02783649211046285","author":"Yu T","year":"2019","journal-title":"arXiv preprint arXiv:1910.10897"},{"key":"bibr64-02783649211046285","first-page":"2223","author":"Zhu JY","year":"2017","journal-title":"Proceedings of the IEEE International Conference on Computer Vision"}],"container-title":["The International Journal of Robotics Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649211046285","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/02783649211046285","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649211046285","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T10:16:45Z","timestamp":1777457805000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/02783649211046285"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,11]]},"references-count":64,"journal-issue":{"issue":"12-14","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["10.1177\/02783649211046285"],"URL":"https:\/\/doi.org\/10.1177\/02783649211046285","relation":{},"ISSN":["0278-3649","1741-3176"],"issn-type":[{"value":"0278-3649","type":"print"},{"value":"1741-3176","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,10,11]]}}}