{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T10:01:55Z","timestamp":1777716115622,"version":"3.51.4"},"reference-count":96,"publisher":"SAGE Publications","issue":"10-11","license":[{"start":{"date-parts":[[2024,11,1]],"date-time":"2024-11-01T00:00:00Z","timestamp":1730419200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of Robotics Research"],"published-print":{"date-parts":[[2025,9]]},"abstract":"<jats:p>Generalization to unseen real-world scenarios for robot manipulation requires exposure to diverse datasets during training. However, collecting large real-world datasets is intractable due to high operational costs. For robot learning to generalize despite these challenges, it is essential to leverage sources of data or priors beyond the robot\u2019s direct experience. In this work, we posit that image-text generative models, which are pre-trained on large corpora of web-scraped data, can serve as such a data source. These generative models encompass a broad range of real-world scenarios beyond a robot\u2019s direct experience and can synthesize novel synthetic experiences that expose robotic agents to additional world priors aiding real-world generalization at no extra cost. In particular, our approach leverages pre-trained generative models as an effective tool for data augmentation. We propose a generative augmentation framework for semantically controllable augmentations and rapidly multiplying robot datasets while inducing rich variations that enable real-world generalization. Based on diverse augmentations of robot data, we show how scalable robot manipulation policies can be trained and deployed both in simulation and in unseen real-world environments such as kitchens and table-tops. By demonstrating the effectiveness of image-text generative models in diverse real-world robotic applications, our generative augmentation framework provides a scalable and efficient path for boosting generalization in robot learning at no extra human cost.<\/jats:p>","DOI":"10.1177\/02783649241273686","type":"journal-article","created":{"date-parts":[[2024,11,1]],"date-time":"2024-11-01T09:10:59Z","timestamp":1730452259000},"page":"1705-1726","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":5,"title":["Semantically controllable augmentations for generalizable robot learning"],"prefix":"10.1177","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-4790-9107","authenticated-orcid":false,"given":"Zoey","family":"Chen","sequence":"first","affiliation":[{"name":"University of Washington"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-5877-920X","authenticated-orcid":false,"given":"Zhao","family":"Mandi","sequence":"additional","affiliation":[{"name":"Stanford University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Homanga","family":"Bharadhwaj","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"},{"name":"FAIR"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mohit","family":"Sharma","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuran","family":"Song","sequence":"additional","affiliation":[{"name":"Stanford University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Abhishek","family":"Gupta","sequence":"additional","affiliation":[{"name":"University of Washington"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vikash","family":"Kumar","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2024,11]]},"reference":[{"key":"e_1_3_4_2_1","doi-asserted-by":"crossref","unstructured":"Bahl S Gupta A Pathak D (2022) Human-to-robot Imitation in the Wild. arXiv preprint arXiv:2207.09450.","DOI":"10.15607\/RSS.2022.XVIII.026"},{"key":"e_1_3_4_3_1","doi-asserted-by":"crossref","unstructured":"Bahl S Mendonca R Chen L et al. (2023) Affordances from Human Videos as a Versatile Representation for Robotics. arXiv preprint arXiv:2304.08488.","DOI":"10.1109\/CVPR52729.2023.01324"},{"key":"e_1_3_4_4_1","first-page":"17605","article-title":"Learning invariances in neural networks from training data","volume":"33","author":"Benton G","year":"2020","unstructured":"Benton G, Finzi M, Izmailov P, et al. (2020) Learning invariances in neural networks from training data. Advances in Neural Information Processing Systems 33: 17605\u201317616.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_5_1","doi-asserted-by":"crossref","unstructured":"Berscheid L R\u00fchr T Kr\u00f6ger T (2019) Improving data efficiency of self-supervised learning for robotic grasping. In: 2019 International Conference on Robotics and Automation (ICRA). Montreal QC 20-24 May 2019 2125\u20132131.","DOI":"10.1109\/ICRA.2019.8793952"},{"key":"e_1_3_4_6_1","doi-asserted-by":"crossref","unstructured":"Bharadhwaj H Gupta A Tulsiani S (2023a) Visual affordance prediction for guiding robot exploration. In: 2023 IEEE International Conference on Robotics and Automation (ICRA). London 2023.","DOI":"10.1109\/ICRA48891.2023.10161288"},{"key":"e_1_3_4_7_1","unstructured":"Bharadhwaj H Gupta A Tulsiani S et al. (2023b) Zero-shot robot manipulation from passive human videos. arXiv preprint arXiv:2302.02011."},{"key":"e_1_3_4_8_1","doi-asserted-by":"crossref","unstructured":"Bharadhwaj H Vakil J Sharma M et al. (2023c) Roboagent: generalization and efficiency in robot manipulation via semantic augmentations and action chunking. arXiv preprint arXiv:2309.01918.","DOI":"10.1109\/ICRA57147.2024.10611293"},{"key":"e_1_3_4_9_1","unstructured":"Bousmalis K Vezzani G Rao D et al. (2023) Robocat: a self-improving foundation agent for robotic manipulation. arXiv preprint arXiv:2306.11706."},{"key":"e_1_3_4_10_1","unstructured":"Brohan A Brown N Carbajal J et al. (2022) Rt-1: Robotics Transformer for Real-World Control at Scale. arXiv preprint arXiv:2212.06817."},{"key":"e_1_3_4_11_1","first-page":"287","volume-title":"Conference on Robot Learning","author":"Brohan A","year":"2023","unstructured":"Brohan A, Chebotar Y, Finn C, et al. (2023) Do as i can, not as i say: grounding language in robotic affordances. In: Conference on Robot Learning. PMLR, 287\u2013318."},{"key":"e_1_3_4_12_1","doi-asserted-by":"crossref","unstructured":"Chen Z Kiami S Gupta A et al. (2023) Genaug: retargeting behaviors to unseen situations via generative augmentation. arXiv preprint arXiv:2302.06671.","DOI":"10.15607\/RSS.2023.XIX.010"},{"key":"e_1_3_4_13_1","doi-asserted-by":"crossref","unstructured":"Cubuk ED Zoph B Mane D et al. (2018) Autoaugment: learning augmentation policies from data. arXiv preprint arXiv:1805.09501.","DOI":"10.1109\/CVPR.2019.00020"},{"key":"e_1_3_4_14_1","doi-asserted-by":"publisher","unstructured":"Deng J Dong W Socher R et al. (2009a) Imagenet: a large-scale hierarchical image database. In: 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009) Miami FL 20-25 June 2009. IEEE Computer Society 248\u2013255. DOI: 10.1109\/CVPR.2009.5206848.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_4_15_1","doi-asserted-by":"crossref","unstructured":"Deng J Dong W Socher R et al. (2009b) Imagenet: a large-scale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition. Miami FL 2009. 248\u2013255.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_4_16_1","doi-asserted-by":"crossref","unstructured":"Deng C Litany O Duan Y et al. (2021) Vector neurons: a general framework for so (3)-equivariant networks. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision 2021 12200\u201312209.","DOI":"10.1109\/ICCV48922.2021.01198"},{"key":"e_1_3_4_17_1","doi-asserted-by":"crossref","unstructured":"Downs L Francis A Koenig N et al. (2022) Google scanned objects: a high-quality dataset of 3d scanned household items. arXiv preprint arXiv:2204.11918.","DOI":"10.1109\/ICRA46639.2022.9811809"},{"key":"e_1_3_4_18_1","unstructured":"Espeholt L Soyer H Munos R et al. (2018) Impala: scalable distributed deep-rl with importance weighted actor-learner architectures. arXiv preprint arXiv:1802.01561."},{"key":"e_1_3_4_19_1","unstructured":"Finn C Yu T Zhang T et al. (2017) One-shot visual imitation learning via meta-learning. In: Conference on Robot Learning. arXiv preprint PMLR 357\u2013368."},{"key":"e_1_3_4_20_1","volume-title":"Free3d","author":"Free3D","unstructured":"Free3D. Free3d. https:\/\/free3d.com\/"},{"key":"e_1_3_4_21_1","unstructured":"Gadre SY Wortsman M Ilharco G et al. (2022) Clip on wheels: zero-shot object navigation as object localization and exploration. arXiv preprint arXiv:2203.10421."},{"key":"e_1_3_4_22_1","unstructured":"Grauman K Westbury A Byrne E et al. (2022) Ego4d: around the world in 3 000 hours of egocentric video. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2022 18995\u201319012."},{"key":"e_1_3_4_23_1","doi-asserted-by":"crossref","unstructured":"Gupta A Dollar P Girshick R (2019) LVIS: a dataset for large vocabulary instance segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2019.","DOI":"10.1109\/CVPR.2019.00550"},{"key":"e_1_3_4_24_1","doi-asserted-by":"crossref","unstructured":"Gupta A Yu J Zhao TZ et al. (2021) Reset-free reinforcement learning via multi-task learning: learning dexterous manipulation behaviors without human intervention. In: 2021 IEEE International Conference on Robotics and Automation (ICRA). Xi'an 2021 6664\u20136671.","DOI":"10.1109\/ICRA48506.2021.9561384"},{"key":"e_1_3_4_25_1","unstructured":"Ha D Schmidhuber J (2018) World Models. arXiv preprint arXiv:1803.10122."},{"key":"e_1_3_4_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/MRA.2021.3138382"},{"key":"e_1_3_4_27_1","unstructured":"Hafner D Lillicrap T Ba J et al. (2019) Dream to control: learning behaviors by latent imagination. arXiv preprint arXiv:1912.01603."},{"key":"e_1_3_4_28_1","doi-asserted-by":"crossref","unstructured":"Handa A Allshire A Makoviychuk V et al. (2023) Dextreme: transfer of agile in-hand manipulation from simulation to reality. In: 2023 IEEE International Conference on Robotics and Automation (ICRA). London 2023 5977\u20135984.","DOI":"10.1109\/ICRA48891.2023.10160216"},{"key":"e_1_3_4_29_1","unstructured":"Hansen N Yuan Z Ze Y et al. (2022) On pre-training for visuo-motor control: revisiting a learning-from-scratch baseline. arXiv preprint arXiv:2212.05749."},{"key":"e_1_3_4_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2020.2974707"},{"key":"e_1_3_4_31_1","unstructured":"Jiang Y Gupta A Zhang Z et al. (2022) Vima: general robot manipulation with multimodal prompts. arXiv preprint arXiv:2210.03094."},{"key":"e_1_3_4_32_1","unstructured":"Kaiser L Babaeizadeh M Milos P et al. (2019) Model-based reinforcement learning for Atari. arXiv preprint arXiv:1903.00374."},{"key":"e_1_3_4_33_1","article-title":"A natural policy gradient","volume":"14","author":"Kakade SM","year":"2001","unstructured":"Kakade SM (2001) A natural policy gradient. Advances in neural information processing systems 14.","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_4_34_1","doi-asserted-by":"crossref","unstructured":"Kapelyukh I Vosylius V Johns E (2022) Dall-e-bot: introducing web-scale diffusion models to robotics. arXiv preprint arXiv:2210.02438.","DOI":"10.1109\/LRA.2023.3272516"},{"key":"e_1_3_4_35_1","unstructured":"Kingma DP Welling M (2013) Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114."},{"key":"e_1_3_4_36_1","doi-asserted-by":"crossref","unstructured":"Kirillov A Mintun E Ravi N et al. (2023) Segment anything. arXiv preprint arXiv:2304.02643.","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"e_1_3_4_37_1","unstructured":"Kostrikov I Yarats D Fergus R (2020) Image augmentation is all you need: regularizing deep reinforcement learning from pixels. arXiv preprint arXiv:2004.13649."},{"key":"e_1_3_4_38_1","doi-asserted-by":"crossref","unstructured":"Kumar V Todorov E (2015) Mujoco haptix: a virtual reality system for hand manipulation. In: 2015 IEEE-RAS 15th International Conference on Humanoid Robots (Humanoids). Seoul Korea 2015 657\u2013663.","DOI":"10.1109\/HUMANOIDS.2015.7363441"},{"key":"e_1_3_4_39_1","article-title":"End-to-end training of deep visuomotor policies","author":"Levine S","year":"2015","unstructured":"Levine S, Finn C, Darrell T, et al. (2015) End-to-end training of deep visuomotor policies. CoRR abs\/1504: 00702, URL https:\/\/arxiv.org\/abs\/1504.00702","journal-title":"CoRR abs\/1504: 00702"},{"key":"e_1_3_4_40_1","doi-asserted-by":"crossref","unstructured":"Lynch C Sermanet P (2020) Language conditioned imitation learning over unstructured data. arXiv preprint arXiv:2005.07648.","DOI":"10.15607\/RSS.2021.XVII.047"},{"key":"e_1_3_4_41_1","unstructured":"Lynch C Khansari M Xiao T et al. (2020) Learning latent plans from play. In: Conference on Robot Learning. arXiv preprint PMLR 1113\u20131132."},{"key":"e_1_3_4_42_1","unstructured":"Ma YJ Sodhani S Jayaraman D et al. (2022) Vip: towards universal visual reward and representation via value-implicit pre-training. arXiv preprint arXiv:2210.00030."},{"key":"e_1_3_4_43_1","unstructured":"Majumdar A Yadav K Arnaud S et al. (2023) Where are we in the search for an artificial visual cortex for embodied intelligence? arXiv preprint arXiv:2303.18240."},{"key":"e_1_3_4_44_1","unstructured":"Mandi Z Bharadhwaj H Moens V et al. (2022) Cacti: a framework for scalable multi-task multi-scene visual imitation learning. arXiv preprint arXiv:2212.05711."},{"key":"e_1_3_4_45_1","unstructured":"Mandlekar A Zhu Y Garg A et al. (2018) Roboturk: a crowdsourcing platform for robotic skill learning through imitation. In: Conference on Robot Learning. arXiv preprint PMLR 879\u2013893."},{"key":"e_1_3_4_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2023.3270034"},{"key":"e_1_3_4_47_1","doi-asserted-by":"crossref","unstructured":"Momeni L Caron M Nagrani A et al. (2023) Verbs in action: improving verb understanding in video-language models. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision 2023 15579\u201315591.","DOI":"10.1109\/ICCV51070.2023.01428"},{"key":"e_1_3_4_48_1","volume-title":"CoRL","author":"Nagabandi A","year":"2019","unstructured":"Nagabandi A, Konoglie K, Levine S, et al. (2019) Deep dynamics models for learning dexterous manipulation. In: CoRL."},{"key":"e_1_3_4_49_1","article-title":"Visual reinforcement learning with imagined goals","volume":"31","author":"Nair AV","year":"2018","unstructured":"Nair AV, Pong V, Dalal M, et al. (2018) Visual reinforcement learning with imagined goals. Advances in neural information processing systems 31.","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_4_50_1","unstructured":"Nair S Rajeswaran A Kumar V et al. (2022) R3M: a universal visual representation for robot manipulation. arXiv preprint arXiv:2203.12601."},{"key":"e_1_3_4_51_1","doi-asserted-by":"crossref","unstructured":"Nguyen A Kanoulas D Muratore L et al. (2018) Translating videos to commands for robotic manipulation with deep recurrent neural networks. In: 2018 IEEE International Conference on Robotics and Automation (ICRA). Brisbane QLD 2018 3782\u20133788.","DOI":"10.1109\/ICRA.2018.8460857"},{"key":"e_1_3_4_52_1","unstructured":"Parisi S Rajeswaran A Purushwalkam S et al. (2022) The unsurprising effectiveness of pre-trained vision models for control. arXiv preprint arXiv:2203.03580."},{"key":"e_1_3_4_53_1","unstructured":"Perez L Wang J (2017) The effectiveness of data augmentation in image classification using deep learning. arXiv preprint arXiv:1712.04621."},{"key":"e_1_3_4_54_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11671"},{"key":"e_1_3_4_55_1","doi-asserted-by":"crossref","unstructured":"Pinto L Gupta A (2017) Learning to push by grasping: using multiple tasks for effective learning. In: 2017 IEEE International Conference on Robotics and Automation (ICRA). Singapore 2017 2161\u20132168.","DOI":"10.1109\/ICRA.2017.7989249"},{"key":"e_1_3_4_56_1","doi-asserted-by":"crossref","unstructured":"Pinto L Andrychowicz M Welinder P et al. (2017) Asymmetric actor critic for image-based robot learning. arXiv preprint arXiv:1710.06542.","DOI":"10.15607\/RSS.2018.XIV.008"},{"key":"e_1_3_4_57_1","article-title":"Motion planning networks","author":"Qureshi AH","year":"2018","unstructured":"Qureshi AH, Bency MJ, Yip MC (2018) Motion planning networks. CoRR abs\/1806: 05767, https:\/\/arxiv.org\/abs\/1806.05767","journal-title":"CoRR abs\/1806: 05767"},{"key":"e_1_3_4_58_1","unstructured":"Radosavovic I Xiao T James S et al. (2023) Real-world robot learning with masked visual pre-training. In: Conference on Robot Learning. arXiv preprint PMLR 416\u2013426."},{"key":"e_1_3_4_59_1","unstructured":"Ramesh A Dhariwal P Nichol A et al. (2022) Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125."},{"key":"e_1_3_4_60_1","doi-asserted-by":"crossref","unstructured":"Rao K Harris C Irpan A et al. (2020) Rl-cyclegan: reinforcement learning aware simulation-to-real. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2020 11157\u201311166.","DOI":"10.1109\/CVPR42600.2020.01117"},{"key":"e_1_3_4_61_1","volume-title":"A Generalist Agent","author":"Reed S","year":"2022","unstructured":"Reed S, Zolna K, Parisotto E, et al. (2022) A Generalist Agent. arXiv preprint arXiv:2205.06175."},{"key":"e_1_3_4_62_1","volume-title":"High-resolution Image Synthesis with Latent Diffusion Models","author":"Rombach R","year":"2021","unstructured":"Rombach R, Blattmann A, Lorenz D, et al. (2021) High-resolution Image Synthesis with Latent Diffusion Models."},{"key":"e_1_3_4_63_1","doi-asserted-by":"crossref","unstructured":"Rombach R Blattmann A Lorenz D et al. (2022) High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2022 10684\u201310695.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_4_64_1","unstructured":"Saharia C Chan W Saxena S et al. (2022) Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487."},{"key":"e_1_3_4_65_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-020-03051-4"},{"key":"e_1_3_4_66_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2210.08402"},{"key":"e_1_3_4_67_1","first-page":"25278","article-title":"Laion-5b: an open large-scale dataset for training next generation image-text models","volume":"35","author":"Schuhmann C","year":"2022","unstructured":"Schuhmann C, Beaumont R, Vencu R, et al. (2022b) Laion-5b: an open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems 35: 25278\u201325294.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_68_1","unstructured":"Shafiullah NMM Cui ZJ Altanzaya A et al. (2022) Behavior transformers: cloning K modes with one stone. arXiv preprint arXiv:2206.11251."},{"key":"e_1_3_4_69_1","unstructured":"Shah R Kumar V (2021) Rrl: Resnet as Representation for Reinforcement Learning. arXiv preprint arXiv:2107.03380."},{"key":"e_1_3_4_70_1","doi-asserted-by":"publisher","DOI":"10.1177\/02783649211046285"},{"key":"e_1_3_4_71_1","unstructured":"Sharma M Fantacci C Zhou Y et al. (2023) Lossless adaptation of pretrained vision models for robotic manipulation. arXiv preprint arXiv:2304.06600."},{"key":"e_1_3_4_72_1","unstructured":"Shaw K Bahl S Pathak D (2023) Videodex: learning dexterity from internet videos. In: Conference on Robot Learning. arXiv preprint PMLR 654\u2013665."},{"key":"e_1_3_4_73_1","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-019-0197-0"},{"key":"e_1_3_4_74_1","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-019-0197-0"},{"key":"e_1_3_4_75_1","unstructured":"Shridhar M Manuelli L Fox D (2021) Cliport: what and where pathways for robotic manipulation. In: Proceedings of the 5th Conference on Robot Learning (CoRL). arXiv preprint."},{"key":"e_1_3_4_76_1","unstructured":"Shridhar M Manuelli L Fox D (2022) Cliport: what and where pathways for robotic manipulation. In: Conference on Robot Learning. arXiv preprint. PMLR 894\u2013906."},{"key":"e_1_3_4_77_1","unstructured":"Singer U Polyak A Hayes T et al. (2022) Make-a-video: text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792."},{"key":"e_1_3_4_78_1","unstructured":"Sodhani S Zhang A Pineau J (2021) Multi-task reinforcement learning with context-based representations. In: International Conference on Machine Learning. PMLR arXiv preprint 9767\u20139779."},{"key":"e_1_3_4_79_1","unstructured":"Song HF Abdolmaleki A Springenberg JT et al. (2019) V-MPO: on-policy maximum a posteriori policy optimization for discrete and continuous control. arXiv preprint arXiv:1909.12238."},{"key":"e_1_3_4_80_1","first-page":"13139","article-title":"Language-conditioned imitation learning for robot manipulation tasks","volume":"33","author":"Stepputtis S","year":"2020","unstructured":"Stepputtis S, Campbell J, Phielipp M, et al. (2020) Language-conditioned imitation learning for robot manipulation tasks. Advances in Neural Information Processing Systems 33: 13139\u201313150.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_81_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v25i1.7979"},{"key":"e_1_3_4_82_1","doi-asserted-by":"crossref","unstructured":"Tobin J Fong R Ray A et al. (2017) Domain randomization for transferring deep neural networks from simulation to the real world. In: 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS). Vancouver BC 2017 23\u201330.","DOI":"10.1109\/IROS.2017.8202133"},{"key":"e_1_3_4_83_1","article-title":"Attention is all you need","volume":"30","author":"Vaswani A","year":"2017","unstructured":"Vaswani A, Shazeer N, Parmar N, et al. (2017) Attention is all you need. Advances in neural information processing systems 30.","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_4_84_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2203.04923"},{"key":"e_1_3_4_85_1","unstructured":"Yang J Gao M Li Z et al. (2023a) Track anything: segment anything meets videos. arXiv preprint arXiv:2304.11968."},{"key":"e_1_3_4_86_1","unstructured":"Yang M Du Y Ghasemipour K et al. (2023b) Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114."},{"key":"e_1_3_4_87_1","unstructured":"Young S Gandhi D Tulsiani S et al. (2020) Visual imitation made easy. In: Conference on Robot Learning (CoRL). arXiv preprint."},{"key":"e_1_3_4_88_1","first-page":"5824","article-title":"Gradient surgery for multi-task learning","volume":"33","author":"Yu T","year":"2020","unstructured":"Yu T, Kumar S, Gupta A, et al. (2020a) Gradient surgery for multi-task learning. Advances in Neural Information Processing Systems 33: 5824\u20135836.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_89_1","unstructured":"Yu T Quillen D He Z et al. (2020b) Meta-world: a benchmark and evaluation for multi-task and meta reinforcement learning. In: Conference on Robot Learning. arXiv preprint. PMLR 1094\u20131100."},{"key":"e_1_3_4_90_1","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2023.XIX.027"},{"key":"e_1_3_4_91_1","doi-asserted-by":"publisher","DOI":"10.3390\/app12178643"},{"key":"e_1_3_4_92_1","unstructured":"Zeng A Florence P Tompson J et al. (2020) Transporter networks: rearranging the visual world for robotic manipulation. arXiv preprint arXiv:2010.14406."},{"key":"e_1_3_4_93_1","unstructured":"Zhao Y Misra I Kr\u00e4henb\u00fchl P et al. (2022) Learning video representations from large language models. arXiv preprint arXiv:2212: 04501."},{"key":"e_1_3_4_94_1","doi-asserted-by":"crossref","unstructured":"Zhao TZ Kumar V Levine S et al. (2023) Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705.","DOI":"10.15607\/RSS.2023.XIX.016"},{"key":"e_1_3_4_95_1","doi-asserted-by":"crossref","unstructured":"Zhou Y Aytar Y Bousmalis K (2021) Manipulator-independent representations for visual imitation. arXiv preprint arXiv:2103.09016.","DOI":"10.15607\/RSS.2021.XVII.002"},{"key":"e_1_3_4_96_1","unstructured":"Zhu Y Wong J Mandlekar A et al. (2020) robosuite: a modular simulation framework and benchmark for robot learning. arXiv preprint arXiv:2009.12293."},{"key":"e_1_3_4_97_1","unstructured":"Zitkovich B Yu T Xu S et al. (2023) Rt-2: vision-language-action models transfer web knowledge to robotic control. In: Conference on Robot Learning. arXiv preprint PMLR 2165\u20132183."}],"container-title":["The International Journal of Robotics Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649241273686","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/02783649241273686","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649241273686","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T10:17:23Z","timestamp":1777457843000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/02783649241273686"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11]]},"references-count":96,"journal-issue":{"issue":"10-11","published-print":{"date-parts":[[2025,9]]}},"alternative-id":["10.1177\/02783649241273686"],"URL":"https:\/\/doi.org\/10.1177\/02783649241273686","relation":{},"ISSN":["0278-3649","1741-3176"],"issn-type":[{"value":"0278-3649","type":"print"},{"value":"1741-3176","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11]]}}}