{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T09:29:54Z","timestamp":1781342994505,"version":"3.54.1"},"reference-count":405,"publisher":"SAGE Publications","issue":"7","license":[{"start":{"date-parts":[[2025,11,20]],"date-time":"2025-11-20T00:00:00Z","timestamp":1763596800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62536001"],"award-info":[{"award-number":["62536001"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62173197"],"award-info":[{"award-number":["62173197"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Natural Science Foundation of China for Key International Collaboration","award":["62120106005"],"award-info":[{"award-number":["62120106005"]}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of Robotics Research"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:p>The realization of universal robots is an ultimate goal of researchers. However, a key hurdle in achieving this goal lies in the robots\u2019 ability to manipulate objects in their unstructured environments according to different tasks. The learning-based approach is considered an effective way to address generalization. The impressive performance of foundation models in the fields of computer vision and natural language suggests the potential of embedding foundation models into manipulation tasks as a viable path toward achieving general manipulation capability. However, we believe achieving general manipulation capability requires an overarching framework akin to auto driving. This framework should encompass multiple functional modules, with different foundation models assuming distinct roles in facilitating general manipulation capability. This survey focuses on the contributions of foundation models to robot learning for manipulation. We propose a comprehensive framework and detail how foundation models can address challenges in each module of the framework. What\u2019s more, we examine current approaches, outline challenges, suggest future research directions, and identify potential risks associated with integrating foundation models into this domain.<\/jats:p>","DOI":"10.1177\/02783649251390579","type":"journal-article","created":{"date-parts":[[2025,11,20]],"date-time":"2025-11-20T15:05:09Z","timestamp":1763651109000},"page":"1091-1142","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":6,"title":["What foundation models can bring for robot learning in manipulation: A survey"],"prefix":"10.1177","volume":"45","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8854-9887","authenticated-orcid":false,"given":"Dingzhe","family":"Li","sequence":"first","affiliation":[{"name":"Advanced Technology Team","place":["China"]},{"name":",","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6286-278X","authenticated-orcid":false,"given":"Yixiang","family":"Jin","sequence":"additional","affiliation":[{"name":"Advanced Technology Team","place":["China"]},{"name":",","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"YuHao","family":"Sun","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence","place":["China"]},{"name":",","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-8709-9027","authenticated-orcid":false,"given":"Yong","family":"A","sequence":"additional","affiliation":[{"name":"Advanced Technology Team","place":["China"]},{"name":",","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongze","family":"Yu","sequence":"additional","affiliation":[{"name":"Advanced Technology Team","place":["China"]},{"name":",","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-3770-6663","authenticated-orcid":false,"given":"Jun","family":"Shi","sequence":"additional","affiliation":[{"name":"Advanced Technology Team","place":["China"]},{"name":",","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-4209-6695","authenticated-orcid":false,"given":"Xiaoshuai","family":"Hao","sequence":"additional","affiliation":[{"name":"Advanced Technology Team","place":["China"]},{"name":",","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6448-3062","authenticated-orcid":false,"given":"Peng","family":"Hao","sequence":"additional","affiliation":[{"name":"Advanced Technology Team","place":["China"]},{"name":",","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huaping","family":"Liu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Technology","place":["China"]},{"name":",","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiang","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Automation, Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinde","family":"Li","sequence":"additional","affiliation":[{"name":"School of Automation","place":["China"]},{"name":", Southeast University","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fuchun","family":"Sun","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Technology","place":["China"]},{"name":",","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianwei","family":"Zhang","sequence":"additional","affiliation":[{"name":"Faculty of Mathematics","place":["Germany"]},{"name":"Universit","place":["Germany"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9149-7336","authenticated-orcid":false,"given":"Bin","family":"Fang","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence","place":["China"]},{"name":",","place":["China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2025,11,20]]},"reference":[{"key":"e_1_3_4_2_1","doi-asserted-by":"crossref","unstructured":"Aceituno-Cabezas B Rodriguez A (2020) A global quasi-dynamic model for contact-trajectory optimization in manipulation.","DOI":"10.15607\/RSS.2020.XVI.047"},{"key":"e_1_3_4_3_1","article-title":"GPT-4 technical report","author":"Achiam J","year":"2023","unstructured":"Achiam J, Adler S, Agarwal S, et al. (2023) GPT-4 technical report. arXiv preprint arXiv:2303.08774.","journal-title":"arXiv preprint arXiv:2303.08774"},{"key":"e_1_3_4_4_1","article-title":"Do as I can, not as I say: grounding language in robotic affordances","author":"Ahn M","year":"2022","unstructured":"Ahn M, Brohan A, Brown N, et al. (2022) Do as I can, not as I say: grounding language in robotic affordances. arXiv preprint arXiv:2204.01691.","journal-title":"arXiv preprint arXiv:2204.01691"},{"key":"e_1_3_4_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-005-0590-7"},{"key":"e_1_3_4_6_1","doi-asserted-by":"crossref","unstructured":"Ausserlechner P Haberger D Thalhammer S et al. (2024) Zs6D: zero-shot 6D object pose estimation using vision transformers. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). Yokohama Japan 13\u201317 May 2024 IEEE pp. 463\u2013469.","DOI":"10.1109\/ICRA57147.2024.10611464"},{"key":"e_1_3_4_7_1","article-title":"Program synthesis with large language models","author":"Austin J","year":"2021","unstructured":"Austin J, Odena A, Nye M, et al. (2021) Program synthesis with large language models. arXiv preprint arXiv:2108.07732.","journal-title":"arXiv preprint arXiv:2108.07732"},{"key":"e_1_3_4_8_1","doi-asserted-by":"crossref","unstructured":"Bahl S Mendonca R Chen L et al. (2023) Affordances from human videos as a versatile representation for robotics. In: 2023 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Vancouver BC Canada 17\u201324 June 2023. pp. 13778\u201313790.","DOI":"10.1109\/CVPR52729.2023.01324"},{"key":"e_1_3_4_9_1","article-title":"Survey on fundamental deep learning 3D reconstruction techniques","author":"Bai Y","year":"2024","unstructured":"Bai Y, Wong L, Twan T (2024) Survey on fundamental deep learning 3D reconstruction techniques. arXiv preprint arXiv:2407.08137.","journal-title":"arXiv preprint arXiv:2407.08137"},{"key":"e_1_3_4_10_1","doi-asserted-by":"crossref","unstructured":"Bain M Nagrani A Varol G et al. (2021) Frozen in time: a joint video and image encoder for end-to-end retrieval. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision Montreal QC Canada 10\u201317 October 2021. pp. 1728\u20131738.","DOI":"10.1109\/ICCV48922.2021.00175"},{"key":"e_1_3_4_11_1","article-title":"RT-H: action hierarchies using language","author":"Belkhale S","year":"2024","unstructured":"Belkhale S, Ding T, Xiao T, et al. (2024) RT-H: action hierarchies using language. arXiv preprint arXiv:2403.01823.","journal-title":"arXiv preprint arXiv:2403.01823"},{"key":"e_1_3_4_12_1","article-title":"Roboagent: generalization and efficiency in robot manipulation via semantic augmentations and action chunking","author":"Bharadhwaj H","year":"2023","unstructured":"Bharadhwaj H, Vakil J, Sharma M, et al. (2023) Roboagent: generalization and efficiency in robot manipulation via semantic augmentations and action chunking. arXiv preprint arXiv:2309.01918.","journal-title":"arXiv preprint arXiv:2309.01918"},{"key":"e_1_3_4_13_1","article-title":"Robotic offline RL from internet videos via value-function pre-training","author":"Bhateja C","year":"2023","unstructured":"Bhateja C, Guo D, Ghosh D, et al. (2023) Robotic offline RL from internet videos via value-function pre-training. arXiv preprint arXiv:2309.13041.","journal-title":"arXiv preprint arXiv:2309.13041"},{"key":"e_1_3_4_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/70.897777"},{"key":"e_1_3_4_15_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.aat8414"},{"key":"e_1_3_4_16_1","article-title":"GR00t N1: an open foundation model for generalist humanoid robots","author":"Bjorck J","year":"2025","unstructured":"Bjorck J, Casta\u00f1eda F, Cherniadev N, et al. (2025) GR00t N1: an open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734.","journal-title":"arXiv preprint arXiv:2503.14734"},{"key":"e_1_3_4_17_1","article-title":"Zero-shot robotic manipulation with pretrained image-editing diffusion models","author":"Black K","year":"2023","unstructured":"Black K, Nakamoto M, Atreya P, et al. (2023) Zero-shot robotic manipulation with pretrained image-editing diffusion models. arXiv preprint arXiv:2310.10639.","journal-title":"arXiv preprint arXiv:2310.10639"},{"key":"e_1_3_4_18_1","article-title":"\u03c0_0: a vision-language-action flow model for general robot control","author":"Black K","year":"2024","unstructured":"Black K, Brown N, Driess D, et al. (2024) \u03c0_0: a vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164.","journal-title":"arXiv preprint arXiv:2410.24164"},{"key":"e_1_3_4_19_1","article-title":"Stable video diffusion: scaling latent video diffusion models to large datasets","author":"Blattmann A","year":"2023","unstructured":"Blattmann A, Dockhorn T, Kulal S, et al. (2023) Stable video diffusion: scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127.","journal-title":"arXiv preprint arXiv:2311.15127"},{"key":"e_1_3_4_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2017.2721939"},{"key":"e_1_3_4_21_1","first-page":"343","article-title":"Domain separation networks","volume":"29","author":"Bousmalis K","year":"2016","unstructured":"Bousmalis K, Trigeorgis G, Silberman N, et al. (2016) Domain separation networks. Advances in Neural Information Processing Systems 29: 343\u2013351.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_22_1","article-title":"Robocat: a self-improving foundation agent for robotic manipulation","author":"Bousmalis K","year":"2023","unstructured":"Bousmalis K, Vezzani G, Rao D, et al. (2023) Robocat: a self-improving foundation agent for robotic manipulation. arXiv preprint arXiv:2306.11706.","journal-title":"arXiv preprint arXiv:2306.11706"},{"key":"e_1_3_4_23_1","article-title":"RT-1: robotics transformer for real-world control at scale","author":"Brohan A","year":"2022","unstructured":"Brohan A, Brown N, Carbajal J, et al. (2022) RT-1: robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817.","journal-title":"arXiv preprint arXiv:2212.06817"},{"key":"e_1_3_4_24_1","article-title":"RT-2: vision-language-action models transfer web knowledge to robotic control","author":"Brohan A","year":"2023","unstructured":"Brohan A, Brown N, Carbajal J, et al. (2023) RT-2: vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818.","journal-title":"arXiv preprint arXiv:2307.15818"},{"key":"e_1_3_4_25_1","doi-asserted-by":"crossref","unstructured":"Brooks T Holynski A Efros AA (2023) InstructPix2Pix: learning to follow image editing instructions. In: 2023 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Vancouver BC Canada 17\u201324 June 2023. pp. 18392\u201318402.","DOI":"10.1109\/CVPR52729.2023.01764"},{"key":"e_1_3_4_26_1","unstructured":"Brooks T Peebles B Holmes C et al. (2024) Video generation models as world simulators. https:\/\/openai.com\/research\/video-generation-models-as-world-simulators"},{"key":"e_1_3_4_27_1","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown T","year":"2020","unstructured":"Brown T, Mann B, Ryder N, et al. (2020) Language models are few-shot learners. Advances in Neural Information Processing Systems 33: 1877\u20131901.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_28_1","article-title":"AgiBot world colosseo: a large-scale manipulation platform for scalable and intelligent embodied systems","author":"Bu Q","year":"2025","unstructured":"Bu Q, Cai J, Chen L, et al. (2025) AgiBot world colosseo: a large-scale manipulation platform for scalable and intelligent embodied systems. arXiv preprint arXiv:2503.06669.","journal-title":"arXiv preprint arXiv:2503.06669"},{"key":"e_1_3_4_29_1","doi-asserted-by":"crossref","unstructured":"Bucker A Figueredo L Haddadin S et al. (2023) Latte: language trajectory transformer. In: 2023 IEEE International Conference on Robotics and Automation (ICRA) London United Kingdom 29 May\u201302 June 2023 IEEE pp. 7287\u20137294.","DOI":"10.1109\/ICRA48891.2023.10161068"},{"key":"e_1_3_4_30_1","article-title":"Comparative evaluation of 3D reconstruction methods for object pose estimation","author":"Burde V","year":"2024","unstructured":"Burde V, Benbihi A, Burget P, et al. (2024) Comparative evaluation of 3D reconstruction methods for object pose estimation. arXiv preprint arXiv:2408.08234.","journal-title":"arXiv preprint arXiv:2408.08234"},{"key":"e_1_3_4_31_1","first-page":"457","volume-title":"Proceedings of the First International Conference on Computer Vision Theory and Applications","author":"Butime J","year":"2006","unstructured":"Butime J, Gutierrez I, Corzo LG, et al. (2006) 3D reconstruction methods, a survey. In: Proceedings of the First International Conference on Computer Vision Theory and Applications. INSTICC, pp. 457\u2013463."},{"key":"e_1_3_4_32_1","article-title":"OV9D: open-vocabulary category-level 9D object pose and size estimation","author":"Cai J","year":"2024","unstructured":"Cai J, He Y, Yuan W, et al. (2024) OV9D: open-vocabulary category-level 9D object pose and size estimation. arXiv preprint arXiv:2403.12396.","journal-title":"arXiv preprint arXiv:2403.12396"},{"key":"e_1_3_4_33_1","doi-asserted-by":"publisher","DOI":"10.1177\/0278364917700714"},{"key":"e_1_3_4_34_1","first-page":"414","volume-title":"European Conference on Computer Vision","author":"Caraffa A","year":"2024","unstructured":"Caraffa A, Boscaini D, Hamza A, et al. (2024) Freeze: training-free zero-shot 6D pose estimation with geometric and vision foundation models. In: European Conference on Computer Vision. Springer, pp. 414\u2013431."},{"key":"e_1_3_4_35_1","doi-asserted-by":"crossref","unstructured":"Caron M Touvron H Misra I et al. (2021) Emerging properties in self-supervised vision transformers. In: 2021 Proceedings of the IEEE\/CVF International Conference on Computer Vision Montreal QC Canada 10\u201317 October 2021. pp. 9650\u20139660.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"e_1_3_4_36_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2006.11.001"},{"key":"e_1_3_4_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58526-6_36"},{"key":"e_1_3_4_38_1","article-title":"Hourvideo: 1-hour video-language understanding","author":"Chandrasegaran K","year":"2024","unstructured":"Chandrasegaran K, Gupta A, Hadzic LM, et al. (2024) Hourvideo: 1-hour video-language understanding. arXiv preprint arXiv:2411.04998.","journal-title":"arXiv preprint arXiv:2411.04998"},{"key":"e_1_3_4_39_1","article-title":"Shapenet: an information-rich 3D model repository","author":"Chang AX","year":"2015","unstructured":"Chang AX, Funkhouser T, Guibas L, et al. (2015) Shapenet: an information-rich 3D model repository. arXiv preprint arXiv:1512.03012.","journal-title":"arXiv preprint arXiv:1512.03012"},{"key":"e_1_3_4_40_1","article-title":"Gr-2: a generative video-language-action model with web-scale knowledge for robot manipulation","author":"Cheang CL","year":"2024","unstructured":"Cheang CL, Chen G, Jing Y, et al. (2024) Gr-2: a generative video-language-action model with web-scale knowledge for robot manipulation. arXiv preprint arXiv:2410.06158.","journal-title":"arXiv preprint arXiv:2410.06158"},{"key":"e_1_3_4_41_1","first-page":"3909","volume-title":"Conference on Robot Learning","author":"Chebotar Y","year":"2023","unstructured":"Chebotar Y, Vuong Q, Hausman K, et al. (2023) Q-transformer: scalable offline reinforcement learning via autoregressive Q-functions. In: Conference on Robot Learning. PMLR, pp. 3909\u20133928."},{"key":"e_1_3_4_42_1","article-title":"Forgetful large language models: lessons learned from using LLMs in robot programming","author":"Chen JT","year":"2023","unstructured":"Chen JT, Huang CM (2023) Forgetful large language models: lessons learned from using LLMs in robot programming. arXiv preprint arXiv:2310.06646.","journal-title":"arXiv preprint arXiv:2310.06646"},{"key":"e_1_3_4_43_1","doi-asserted-by":"crossref","unstructured":"Chen C Culbertson P Lepert M et al. (2021a) Trajectotree: trajectory optimization meets tree search for planning multi-contact dexterous manipulation. In: 2021 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) Prague Czech Republic 27 September\u201301 October 2021 IEEE pp. 8262\u20138268.","DOI":"10.1109\/IROS51168.2021.9636346"},{"key":"e_1_3_4_44_1","article-title":"Evaluating large language models trained on code","author":"Chen M","year":"2021","unstructured":"Chen M, Tworek J, Jun H, et al. (2021b) Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.","journal-title":"arXiv preprint arXiv:2107.03374"},{"key":"e_1_3_4_45_1","doi-asserted-by":"crossref","unstructured":"Chen B Xia F Ichter B et al. (2023a) Open-vocabulary queryable scene representations for real world planning. In:2023 IEEE International Conference on Robotics and Automation (ICRA) London United Kingdom 29 May\u201302 June 2023 IEEE pp. 11509\u201311522.","DOI":"10.1109\/ICRA48891.2023.10161534"},{"key":"e_1_3_4_46_1","article-title":"PaLI-X: on scaling up a multilingual vision and language model","author":"Chen X","year":"2023","unstructured":"Chen X, Djolonga J, Padlewski P, et al. (2023b) PaLI-X: on scaling up a multilingual vision and language model. arXiv preprint arXiv:2305.18565.","journal-title":"arXiv preprint arXiv:2305.18565"},{"key":"e_1_3_4_47_1","article-title":"GenAug: retargeting behaviors to unseen situations via generative augmentation","author":"Chen Z","year":"2023","unstructured":"Chen Z, Kiami S, Gupta A, et al. (2023c) GenAug: retargeting behaviors to unseen situations via generative augmentation. arXiv preprint arXiv:2302.06671.","journal-title":"arXiv preprint arXiv:2302.06671"},{"key":"e_1_3_4_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2023.3237042"},{"key":"e_1_3_4_49_1","doi-asserted-by":"publisher","DOI":"10.1089\/soro.2023.0062"},{"key":"e_1_3_4_50_1","doi-asserted-by":"crossref","unstructured":"Chen B Xu Z Kirmani S et al. (2024b) SpatialVLM: endowing vision-language models with spatial reasoning capabilities. In: 2024 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 16\u201322 June 2024 pp. 14455\u201314465.","DOI":"10.1109\/CVPR52733.2024.01370"},{"key":"e_1_3_4_51_1","article-title":"URDformer: a pipeline for constructing articulated simulation environments from real-world images","author":"Chen Z","year":"2024","unstructured":"Chen Z, Walsman A, Memmel M, et al. (2024c) URDformer: a pipeline for constructing articulated simulation environments from real-world images. arXiv preprint arXiv:2405.11656.","journal-title":"arXiv preprint arXiv:2405.11656"},{"key":"e_1_3_4_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19815-1_37"},{"key":"e_1_3_4_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2023.3333699"},{"key":"e_1_3_4_54_1","article-title":"Open-television: teleoperation with immersive active visual feedback","author":"Cheng X","year":"2024","unstructured":"Cheng X, Li J, Yang S, et al. (2024) Open-television: teleoperation with immersive active visual feedback. arXiv preprint arXiv:2407.01512.","journal-title":"arXiv preprint arXiv:2407.01512"},{"key":"e_1_3_4_55_1","article-title":"Diffusion policy: visuomotor policy learning via action diffusion","author":"Chi C","year":"2023","unstructured":"Chi C, Feng S, Du Y, et al. (2023) Diffusion policy: visuomotor policy learning via action diffusion. arXiv preprint arXiv:2303.04137.","journal-title":"arXiv preprint arXiv:2303.04137"},{"key":"e_1_3_4_56_1","article-title":"Universal manipulation interface: in-the-wild robot teaching without in-the-wild robots","author":"Chi C","year":"2024","unstructured":"Chi C, Xu Z, Pan C, et al. (2024) Universal manipulation interface: in-the-wild robot teaching without in-the-wild robots. arXiv preprint arXiv:2402.10329.","journal-title":"arXiv preprint arXiv:2402.10329"},{"key":"e_1_3_4_57_1","article-title":"Pybullet, a python module for physics simulation for games","author":"Coumans E","year":"2016","unstructured":"Coumans E, Bai Y (2016) Pybullet, a python module for physics simulation for games. Robotics and Machine Learning.","journal-title":"Robotics and Machine Learning"},{"key":"e_1_3_4_58_1","doi-asserted-by":"crossref","unstructured":"Cruciani S Smith C Kragic D et al. (2018) Dexterous manipulation graphs. In: 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) Madrid Spain 01\u201305 October 2018 IEEE pp. 2040\u20132047.","DOI":"10.1109\/IROS.2018.8594303"},{"key":"e_1_3_4_59_1","first-page":"893","volume-title":"Learning for Dynamics and Control Conference","author":"Cui Y","year":"2022","unstructured":"Cui Y, Niekum S, Gupta A, et al. (2022) Can foundation models perform zero-shot task specification for robot manipulation? In: Learning for Dynamics and Control Conference. PMLR, pp. 893\u2013905."},{"key":"e_1_3_4_60_1","doi-asserted-by":"crossref","unstructured":"Cui Y Karamcheti S Palleti R et al. (2023) No to the right: online language corrections for robotic manipulation via shared autonomy. In: Proceedings of the 2023 ACM\/IEEE International Conference on Human-Robot Interaction Stockholm Sweden 13-16 March 2023 pp. 93\u2013101.","DOI":"10.1145\/3568162.3578623"},{"key":"e_1_3_4_61_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19842-7_22"},{"key":"e_1_3_4_62_1","doi-asserted-by":"crossref","unstructured":"Dai Q Zhu Y Geng Y et al. (2023) GraspNeRF: multiview-based 6-DoF grasp detection for transparent and specular objects using generalizable NeRF. In:2023 IEEE International Conference on Robotics and Automation (ICRA) London United Kingdom 29 May\u201302 June 2023 IEEE pp. 1757\u20131763.","DOI":"10.1109\/ICRA48891.2023.10160842"},{"key":"e_1_3_4_63_1","article-title":"Automated creation of digital cousins for robust policy learning","author":"Dai T","year":"2024","unstructured":"Dai T, Wong J, Jiang Y, et al. (2024) Automated creation of digital cousins for robust policy learning. arXiv preprint arXiv:2410.07408.","journal-title":"arXiv preprint arXiv:2410.07408"},{"key":"e_1_3_4_64_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01531-2"},{"key":"e_1_3_4_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2023.3236952"},{"key":"e_1_3_4_66_1","article-title":"Robonet: large-scale multi-robot learning","author":"Dasari S","year":"2019","unstructured":"Dasari S, Ebert F, Tian S, et al. (2019) Robonet: large-scale multi-robot learning. arXiv preprint arXiv:1910.11215.","journal-title":"arXiv preprint arXiv:1910.11215"},{"key":"e_1_3_4_67_1","doi-asserted-by":"crossref","unstructured":"De Pace F Gorjup G Bai H et al. (2021) Leveraging enhanced virtual reality methods and environments for efficient intuitive and immersive teleoperation of robots. In: 2021 IEEE International Conference on Robotics and Automation (ICRA) Xi'an China 30 May\u201305 June 2021 IEEE pp. 12967\u201312973.","DOI":"10.1109\/ICRA48506.2021.9560757"},{"key":"e_1_3_4_68_1","first-page":"7480","volume-title":"International Conference on Machine Learning","author":"Dehghani M","year":"2023","unstructured":"Dehghani M, Djolonga J, Mustafa B, et al. (2023) Scaling vision transformers to 22 billion parameters. In: International Conference on Machine Learning. PMLR, pp. 7480\u20137512."},{"key":"e_1_3_4_69_1","doi-asserted-by":"publisher","DOI":"10.52202\/068431-0433"},{"key":"e_1_3_4_70_1","doi-asserted-by":"crossref","unstructured":"Deitke M Schwenk D Salvador J et al. (2023) Objaverse: a universe of annotated 3D objects. In: 2023 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Vancouver BC Canada 17\u201324 June 2023. pp. 13142\u201313153.","DOI":"10.1109\/CVPR52729.2023.01263"},{"key":"e_1_3_4_71_1","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1003"},{"key":"e_1_3_4_72_1","doi-asserted-by":"crossref","unstructured":"Deng J Dong W Socher R et al. (2009) ImageNet: a large-scale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition Miami FL USA 20\u201325 June 2009 IEEE pp. 248\u2013255.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_4_73_1","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2021.3056043"},{"key":"e_1_3_4_74_1","article-title":"BERT: pre-training of deep bidirectional transformers for language understanding","author":"Devlin J","year":"2018","unstructured":"Devlin J, Chang MW, Lee K, et al. (2018) BERT: pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.","journal-title":"arXiv preprint arXiv:1810.04805"},{"key":"e_1_3_4_75_1","article-title":"Towards a unified agent with foundation models","author":"Di Palo N","year":"2023","unstructured":"Di Palo N, Byravan A, Hasenclever L, et al. (2023) Towards a unified agent with foundation models. arXiv preprint arXiv:2307.09668.","journal-title":"arXiv preprint arXiv:2307.09668"},{"key":"e_1_3_4_76_1","article-title":"Robot task planning and situation handling in open worlds","author":"Ding Y","year":"2022","unstructured":"Ding Y, Zhang X, Amiri S, et al. (2022) Robot task planning and situation handling in open worlds. arXiv preprint arXiv:2210.01287.","journal-title":"arXiv preprint arXiv:2210.01287"},{"key":"e_1_3_4_77_1","article-title":"Task and motion planning with large language models for object rearrangement","author":"Ding Y","year":"2023","unstructured":"Ding Y, Zhang X, Paxton C, et al. (2023) Task and motion planning with large language models for object rearrangement. arXiv preprint arXiv:2303.06247.","journal-title":"arXiv preprint arXiv:2303.06247"},{"key":"e_1_3_4_78_1","article-title":"PaLM-E: an embodied multimodal language model","author":"Driess D","year":"2023","unstructured":"Driess D, Xia F, Sajjadi MS, et al. (2023) PaLM-E: an embodied multimodal language model. arXiv preprint arXiv:2303.03378.","journal-title":"arXiv preprint arXiv:2303.03378"},{"key":"e_1_3_4_79_1","first-page":"1","article-title":"Learning universal policies via text-guided video generation","volume":"36","author":"Du Y","year":"2024","unstructured":"Du Y, Yang S, Dai B, et al. (2024) Learning universal policies via text-guided video generation. Advances in Neural Information Processing Systems 36: 1\u201317.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_80_1","article-title":"Aha: a vision-language-model for detecting and reasoning over failures in robotic manipulation","author":"Duan J","year":"2024","unstructured":"Duan J, Pumacay W, Kumar N, et al. (2024) Aha: a vision-language-model for detecting and reasoning over failures in robotic manipulation. arXiv preprint arXiv:2410.00371.","journal-title":"arXiv preprint arXiv:2410.00371"},{"issue":"2251","key":"e_1_3_4_81_1","first-page":"20220050","article-title":"Dreamcoder: growing generalizable, interpretable knowledge with wake\u2013sleep Bayesian program learning","volume":"381","author":"Ellis K","year":"2023","unstructured":"Ellis K, Wong L, Nye M, et al. (2023) Dreamcoder: growing generalizable, interpretable knowledge with wake\u2013sleep Bayesian program learning. Philosophical Transactions. Series A, Mathematical, Physical, and Engineering Sciences 381(2251): 20220050.","journal-title":"Philosophical Transactions. Series A, Mathematical, Physical, and Engineering Sciences"},{"key":"e_1_3_4_82_1","article-title":"Off-dynamics reinforcement learning: training for transfer with domain classifiers","author":"Eysenbach B","year":"2020","unstructured":"Eysenbach B, Asawa S, Chaudhari S, et al. (2020) Off-dynamics reinforcement learning: training for transfer with domain classifiers. arXiv preprint arXiv:2006.13916.","journal-title":"arXiv preprint arXiv:2006.13916"},{"key":"e_1_3_4_83_1","article-title":"Learning by watching: a review of video-based learning approaches for robot manipulation","author":"Eze C","year":"2024","unstructured":"Eze C, Crick C (2024) Learning by watching: a review of video-based learning approaches for robot manipulation. arXiv preprint arXiv:2402.07127.","journal-title":"arXiv preprint arXiv:2402.07127"},{"key":"e_1_3_4_84_1","doi-asserted-by":"publisher","DOI":"10.1108\/IR-07-2016-0179"},{"key":"e_1_3_4_85_1","doi-asserted-by":"publisher","DOI":"10.1177\/1729881417717057"},{"key":"e_1_3_4_86_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2020.103592"},{"key":"e_1_3_4_87_1","doi-asserted-by":"crossref","unstructured":"Fang HS Wang C Gou M et al. (2020b) Graspnet-1billion: a large-scale benchmark for general object grasping. In: 2020 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 13\u201319 June 2020 pp. 11444\u201311453.","DOI":"10.1109\/CVPR42600.2020.01146"},{"key":"e_1_3_4_88_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2022.3183256"},{"key":"e_1_3_4_89_1","doi-asserted-by":"crossref","unstructured":"Fang H Fang HS Wang Y et al. (2024) Airexo: low-cost exoskeletons for learning whole-arm manipulation in the wild. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). Yokohama Japan 13\u201317 May 2024 IEEE pp. 15031\u201315038.","DOI":"10.1109\/ICRA57147.2024.10610799"},{"key":"e_1_3_4_90_1","first-page":"6","article-title":"Planning optimal grasps","volume":"3","author":"Ferrari C","year":"1992","unstructured":"Ferrari C, Canny JF (1992) Planning optimal grasps. ICRA 3: 6.","journal-title":"ICRA"},{"key":"e_1_3_4_91_1","unstructured":"Figureai (2025) Helix. https:\/\/www.figure.ai\/news\/helix?ref=fixthenews.com (Accessed 20 February 2025)."},{"key":"e_1_3_4_92_1","doi-asserted-by":"publisher","DOI":"10.1177\/02783649241281508"},{"key":"e_1_3_4_93_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-020-09722-8"},{"key":"e_1_3_4_94_1","article-title":"Humanplus: humanoid shadowing and imitation from humans","author":"Fu Z","year":"2024","unstructured":"Fu Z, Zhao Q, Wu Q, et al. (2024) Humanplus: humanoid shadowing and imitation from humans. arXiv preprint arXiv:2406.10454.","journal-title":"arXiv preprint arXiv:2406.10454"},{"key":"e_1_3_4_95_1","doi-asserted-by":"crossref","unstructured":"Gao J Sarkar B Xia F et al. (2023) Physically grounded vision-language models for robotic manipulation.","DOI":"10.1109\/ICRA57147.2024.10610090"},{"key":"e_1_3_4_96_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-023-01891-x"},{"key":"e_1_3_4_97_1","doi-asserted-by":"crossref","unstructured":"Ge Y Macaluso A Li LE et al. (2023) Policy adaptation from foundation model feedback. In: 2023 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Vancouver BC Canada 17\u201324 June 2023. pp. 19059\u201319069.","DOI":"10.1109\/CVPR52729.2023.01827"},{"key":"e_1_3_4_98_1","doi-asserted-by":"crossref","unstructured":"Geng H Xu H Zhao C et al. (2023a) GAPartNet: cross-category domain-generalizable object perception and manipulation via generalizable and actionable parts. In: 2023 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Vancouver BC Canada 17\u201324 June 2023. pp. 7081\u20137091.","DOI":"10.1109\/CVPR52729.2023.00684"},{"key":"e_1_3_4_99_1","doi-asserted-by":"crossref","unstructured":"Geng Y An B Geng H et al. (2023b) RLAfford: end-to-end affordance learning for robotic manipulation. In:2023 IEEE International Conference on Robotics and Automation (ICRA) London United Kingdom 29 May\u201302 June 2023 IEEE pp. 5880\u20135886.","DOI":"10.1109\/ICRA48891.2023.10161571"},{"key":"e_1_3_4_100_1","unstructured":"Gervet T Xian Z Gkanatsios N et al. (2023) Act3D: 3D feature field transformers for multi-task robotic manipulation. In: 7th Annual Conference on Robot Learning Atlanta USA 2023."},{"key":"e_1_3_4_101_1","doi-asserted-by":"publisher","DOI":"10.4324\/9781315740218"},{"key":"e_1_3_4_102_1","unstructured":"Grauman K Westbury A Byrne E et al. (2022) Ego4D: around the world in 3 000 hours of egocentric video. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18\u201324 June 2022. pp. 18995\u201319012."},{"key":"e_1_3_4_103_1","unstructured":"Grauman K Westbury A Torresani L et al. (2024) Ego-Exo4D: understanding skilled human activity from first-and third-person perspectives. In: 2024 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 16\u201322 June 2024 pp. 19383\u201319400."},{"key":"e_1_3_4_104_1","article-title":"Open-vocabulary object detection via vision and language knowledge distillation","author":"Gu X","year":"2021","unstructured":"Gu X, Lin TY, Kuo W, et al. (2021) Open-vocabulary object detection via vision and language knowledge distillation. arXiv preprint arXiv:2104.13921.","journal-title":"arXiv preprint arXiv:2104.13921"},{"key":"e_1_3_4_105_1","article-title":"Maniskill2: a unified benchmark for generalizable manipulation skills","author":"Gu J","year":"2023","unstructured":"Gu J, Xiang F, Li X, et al. (2023) Maniskill2: a unified benchmark for generalizable manipulation skills. arXiv preprint arXiv:2302.04659.","journal-title":"arXiv preprint arXiv:2302.04659"},{"key":"e_1_3_4_106_1","doi-asserted-by":"publisher","DOI":"10.3390\/s24041076"},{"key":"e_1_3_4_107_1","doi-asserted-by":"publisher","DOI":"10.1145\/3583136"},{"key":"e_1_3_4_108_1","doi-asserted-by":"crossref","unstructured":"Gupta A Yu J Zhao TZ et al. (2021) Reset-free reinforcement learning via multi-task learning: learning dexterous manipulation behaviors without human intervention. In: 2021 IEEE International Conference on Robotics and Automation (ICRA) Xi'an China 30 May\u201305 June 2021 IEEE pp. 6664\u20136671.","DOI":"10.1109\/ICRA48506.2021.9561384"},{"key":"e_1_3_4_109_1","article-title":"Bridging the human to robot dexterity gap through object-oriented rewards","author":"Guzey I","year":"2024","unstructured":"Guzey I, Dai Y, Savva G, et al. (2024) Bridging the human to robot dexterity gap through object-oriented rewards. arXiv preprint arXiv:2410.23289.","journal-title":"arXiv preprint arXiv:2410.23289"},{"key":"e_1_3_4_110_1","article-title":"Semantic abstraction: open-world 3D scene understanding from 2D vision-language models","author":"Ha H","year":"2022","unstructured":"Ha H, Song S (2022) Semantic abstraction: open-world 3D scene understanding from 2D vision-language models. arXiv preprint arXiv:2207.11514.","journal-title":"arXiv preprint arXiv:2207.11514"},{"key":"e_1_3_4_111_1","first-page":"3766","volume-title":"Conference on Robot Learning","author":"Ha H","year":"2023","unstructured":"Ha H, Florence P, Song S (2023) Scaling up and distilling down: language-guided robot skill acquisition. In: Conference on Robot Learning. PMLR, pp. 3766\u20133777."},{"key":"e_1_3_4_112_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2022.3152247"},{"key":"e_1_3_4_113_1","doi-asserted-by":"crossref","unstructured":"Haque A Tancik M Efros AA et al. (2023) Instruct-NeRF2NeRF: editing 3D scenes with instructions. In: 2023 Proceedings of the IEEE\/CVF International Conference on Computer Vision Paris France 1\u20136 October 2023 pp. 19740\u201319750.","DOI":"10.1109\/ICCV51070.2023.01808"},{"key":"e_1_3_4_114_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2021.3056060"},{"key":"e_1_3_4_115_1","doi-asserted-by":"crossref","unstructured":"He K Chen X Xie S et al. (2022) Masked autoencoders are scalable vision learners. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18\u201324 June 2022. pp. 16000\u201316009.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"e_1_3_4_116_1","article-title":"Robot data curation with mutual information estimators","author":"Hejna J","year":"2025","unstructured":"Hejna J, Mirchandani S, Balakrishna A, et al. (2025) Robot data curation with mutual information estimators. arXiv preprint arXiv:2502.08623.","journal-title":"arXiv preprint arXiv:2502.08623"},{"key":"e_1_3_4_117_1","doi-asserted-by":"crossref","unstructured":"Herzog A Rao K Hausman K et al. (2023) Deep rl at scale: sorting waste in office buildings with a fleet of mobile manipulators.","DOI":"10.15607\/RSS.2023.XIX.022"},{"key":"e_1_3_4_118_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-43089-4_51"},{"key":"e_1_3_4_119_1","first-page":"20482","article-title":"3d-llm: injecting the 3d world into large language models","volume":"36","author":"Hong Y","year":"2023","unstructured":"Hong Y, Zhen H, Chen P, et al. (2023) 3d-llm: injecting the 3d world into large language models. Advances in Neural Information Processing Systems 36: 20482\u201320494.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_120_1","article-title":"Predictive sampling: real-time behaviour synthesis with mujoco","author":"Howell T","year":"2022","unstructured":"Howell T, Gileadi N, Tunyasuvunakool S, et al. (2022) Predictive sampling: real-time behaviour synthesis with mujoco. arXiv preprint arXiv:2212.00541.","journal-title":"arXiv preprint arXiv:2212.00541"},{"key":"e_1_3_4_121_1","article-title":"Look before you leap: unveiling the power of gpt-4v in robotic vision-language planning","author":"Hu Y","year":"2023","unstructured":"Hu Y, Lin F, Zhang T, et al. (2023a) Look before you leap: unveiling the power of gpt-4v in robotic vision-language planning. arXiv preprint arXiv:2311.17842.","journal-title":"arXiv preprint arXiv:2311.17842"},{"key":"e_1_3_4_122_1","article-title":"Toward general-purpose robots via foundation models: a survey and meta-analysis","author":"Hu Y","year":"2023","unstructured":"Hu Y, Xie Q, Jain V, et al. (2023b) Toward general-purpose robots via foundation models: a survey and meta-analysis. arXiv preprint arXiv:2312.08782.","journal-title":"arXiv preprint arXiv:2312.08782"},{"key":"e_1_3_4_123_1","doi-asserted-by":"crossref","unstructured":"Hu Y Yang J Chen L et al. (2023c) Planning-oriented autonomous driving. In: 2023 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Vancouver BC Canada 17\u201324 June 2023. pp. 17853\u201317862.","DOI":"10.1109\/CVPR52729.2023.01712"},{"key":"e_1_3_4_124_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-66723-8_29"},{"key":"e_1_3_4_125_1","first-page":"9118","volume-title":"International Conference on Machine Learning","author":"Huang W","year":"2022","unstructured":"Huang W, Abbeel P, Pathak D, et al. (2022) Language models as zero-shot planners: extracting actionable knowledge for embodied agents. In: International Conference on Machine Learning. PMLR, pp. 9118\u20139147."},{"key":"e_1_3_4_126_1","article-title":"An embodied generalist agent in 3D world","author":"Huang J","year":"2023","unstructured":"Huang J, Yong S, Ma X, et al. (2023a) An embodied generalist agent in 3D world. arXiv preprint arXiv:2311.12871.","journal-title":"arXiv preprint arXiv:2311.12871"},{"key":"e_1_3_4_127_1","article-title":"Instruct2ACT: mapping multi-modality instructions to robotic actions with large language model","author":"Huang S","year":"2023","unstructured":"Huang S, Jiang Z, Dong H, et al. (2023b) Instruct2ACT: mapping multi-modality instructions to robotic actions with large language model. arXiv preprint arXiv:2305.11176.","journal-title":"arXiv preprint arXiv:2305.11176"},{"key":"e_1_3_4_128_1","article-title":"VoxPoser: composable 3D value maps for robotic manipulation with language models","author":"Huang W","year":"2023","unstructured":"Huang W, Wang C, Zhang R, et al. (2023c) VoxPoser: composable 3D value maps for robotic manipulation with language models. arXiv preprint arXiv:2307.05973.","journal-title":"arXiv preprint arXiv:2307.05973"},{"key":"e_1_3_4_129_1","article-title":"Grounded decoding: guiding text generation with grounded models for robot control","author":"Huang W","year":"2023","unstructured":"Huang W, Xia F, Shah D, et al. (2023d) Grounded decoding: guiding text generation with grounded models for robot control. arXiv preprint arXiv:2303.00855.","journal-title":"arXiv preprint arXiv:2303.00855"},{"key":"e_1_3_4_130_1","article-title":"ReKep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation","author":"Huang W","year":"2024","unstructured":"Huang W, Wang C, Li Y, et al. (2024a) ReKep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation. arXiv preprint arXiv:2409.01652.","journal-title":"arXiv preprint arXiv:2409.01652"},{"key":"e_1_3_4_131_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3429209"},{"key":"e_1_3_4_132_1","doi-asserted-by":"crossref","unstructured":"Issac J W\u00fcthrich M Cifuentes CG et al. (2016) Depth-based object tracking using a robust gaussian filter. In: 2016 IEEE International Conference on Robotics and Automation (ICRA) Stockholm Sweden 16\u201321 May 2016 IEEE 608\u2013615.","DOI":"10.1109\/ICRA.2016.7487184"},{"key":"e_1_3_4_133_1","article-title":"Open teach: a versatile teleoperation system for robotic manipulation","author":"Iyer A","year":"2024","unstructured":"Iyer A, Peng Z, Dai Y, et al. (2024) Open teach: a versatile teleoperation system for robotic manipulation. arXiv preprint arXiv:2403.07870.","journal-title":"arXiv preprint arXiv:2403.07870"},{"key":"e_1_3_4_134_1","first-page":"4651","volume-title":"International Conference on Machine Learning","author":"Jaegle A","year":"2021","unstructured":"Jaegle A, Gimeno F, Brock A, et al. (2021) Perceiver: general perception with iterative attention. In: International Conference on Machine Learning. PMLR, pp. 4651\u20134664."},{"key":"e_1_3_4_135_1","article-title":"Vid2Robot: end-to-end video-conditioned policy learning with cross-attention transformers","author":"Jain V","year":"2024","unstructured":"Jain V, Attarian M, Joshi NJ, et al. (2024) Vid2Robot: end-to-end video-conditioned policy learning with cross-attention transformers. arXiv preprint arXiv:2403.12943.","journal-title":"arXiv preprint arXiv:2403.12943"},{"key":"e_1_3_4_136_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2020.2974707"},{"key":"e_1_3_4_137_1","first-page":"991","volume-title":"Conference on Robot Learning","author":"Jang E","year":"2022","unstructured":"Jang E, Irpan A, Khansari M, et al. (2022) BC-Z: zero-shot task generalization with robotic imitation learning. In: Conference on Robot Learning. PMLR, pp. 991\u20131002."},{"key":"e_1_3_4_138_1","article-title":"Visually-grounded planning without vision: language models infer detailed plans from high-level instructions","author":"Jansen PA","year":"2020","unstructured":"Jansen PA (2020) Visually-grounded planning without vision: language models infer detailed plans from high-level instructions. arXiv preprint arXiv:2009.14259.","journal-title":"arXiv preprint arXiv:2009.14259"},{"key":"e_1_3_4_139_1","article-title":"Conceptfusion: open-set multimodal 3D mapping","author":"Jatavallabhula KM","year":"2023","unstructured":"Jatavallabhula KM, Kuwajerwala A, Gu Q, et al. (2023) Conceptfusion: open-set multimodal 3D mapping. arXiv preprint arXiv:2302.07241.","journal-title":"arXiv preprint arXiv:2302.07241"},{"key":"e_1_3_4_140_1","article-title":"Synergies between affordance and geometry: 6-DoF grasp detection via implicit representations","author":"Jiang Z","year":"2021","unstructured":"Jiang Z, Zhu Y, Svetlik M, et al. (2021) Synergies between affordance and geometry: 6-DoF grasp detection via implicit representations. arXiv preprint arXiv:2104.01542.","journal-title":"arXiv preprint arXiv:2104.01542"},{"key":"e_1_3_4_141_1","doi-asserted-by":"crossref","unstructured":"Jiang Z Hsu CC Zhu Y (2022) Ditto: building digital twins of articulated objects from interaction. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18-24 June 2022. pp. 5616\u20135626.","DOI":"10.1109\/CVPR52688.2022.00553"},{"key":"e_1_3_4_142_1","unstructured":"Jiang Y Gupta A Zhang Z et al. (2023) VIMA: robot manipulation with multimodal prompts."},{"key":"e_1_3_4_143_1","article-title":"TRANSIC: sim-to-real policy transfer by learning from online correction","author":"Jiang Y","year":"2024","unstructured":"Jiang Y, Wang C, Zhang R, et al. (2024a) TRANSIC: sim-to-real policy transfer by learning from online correction. arXiv preprint arXiv:2405.10315.","journal-title":"arXiv preprint arXiv:2405.10315"},{"key":"e_1_3_4_144_1","doi-asserted-by":"crossref","unstructured":"Jiang Y Yu M Zhu X et al. (2024b) Contact-implicit model predictive control for dexterous in-hand manipulation: a long-horizon and robust approach. In: 2024 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) Abu Dhabi United Arab Emirates 14\u201318 October 2024 IEEE pp. 5260\u20135266.","DOI":"10.1109\/IROS58592.2024.10801751"},{"key":"e_1_3_4_145_1","article-title":"DexMimicGen: automated data generation for bimanual dexterous manipulation via imitation learning","author":"Jiang Z","year":"2024","unstructured":"Jiang Z, Xie Y, Lin K, et al. (2024c) DexMimicGen: automated data generation for bimanual dexterous manipulation via imitation learning. arXiv preprint arXiv:2410.24185.","journal-title":"arXiv preprint arXiv:2410.24185"},{"key":"e_1_3_4_146_1","article-title":"Complementarity-free multi-contact modeling and optimization for dexterous manipulation","author":"Jin W","year":"2024","unstructured":"Jin W (2024) Complementarity-free multi-contact modeling and optimization for dexterous manipulation. arXiv preprint arXiv:2408.07855.","journal-title":"arXiv preprint arXiv:2408.07855"},{"key":"e_1_3_4_147_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2024.3357432"},{"key":"e_1_3_4_148_1","article-title":"Exploring visual pre-training for robot manipulation: datasets, models and methods","author":"Jing Y","year":"2023","unstructured":"Jing Y, Zhu X, Liu X, et al. (2023) Exploring visual pre-training for robot manipulation: datasets, models and methods. arXiv preprint arXiv:2308.03620.","journal-title":"arXiv preprint arXiv:2308.03620"},{"key":"e_1_3_4_149_1","article-title":"Copal: corrective planning of robot actions with large language models","author":"Joublin F","year":"2023","unstructured":"Joublin F, Ceravola A, Smirnov P, et al. (2023) Copal: corrective planning of robot actions with large language models. arXiv preprint arXiv:2310.07263.","journal-title":"arXiv preprint arXiv:2310.07263"},{"key":"e_1_3_4_150_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-021-03819-2"},{"key":"e_1_3_4_151_1","article-title":"Smart-LLM: smart multi-agent robot task planning using large language models","author":"Kannan SS","year":"2023","unstructured":"Kannan SS, Venkatesh VL, Min BC (2023) Smart-LLM: smart multi-agent robot task planning using large language models. arXiv preprint arXiv:2309.10062.","journal-title":"arXiv preprint arXiv:2309.10062"},{"key":"e_1_3_4_152_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2023.3272516"},{"key":"e_1_3_4_153_1","article-title":"Language-driven representation learning for robotics","author":"Karamcheti S","year":"2023","unstructured":"Karamcheti S, Nair S, Chen AS, et al. (2023) Language-driven representation learning for robotics. arXiv preprint arXiv:2302.12766.","journal-title":"arXiv preprint arXiv:2302.12766"},{"key":"e_1_3_4_154_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.celrep.2021.109730"},{"key":"e_1_3_4_155_1","article-title":"Egomimic: scaling imitation learning via egocentric video","author":"Kareer S","year":"2024","unstructured":"Kareer S, Patel D, Punamiya R, et al. (2024) Egomimic: scaling imitation learning via egocentric video. arXiv preprint arXiv:2410.24221.","journal-title":"arXiv preprint arXiv:2410.24221"},{"key":"e_1_3_4_156_1","article-title":"3d diffuser actor: policy diffusion with 3d scene representations","author":"Ke TW","year":"2024","unstructured":"Ke TW, Gkanatsios N, Fragkiadaki K (2024) 3d diffuser actor: policy diffusion with 3d scene representations. arXiv preprint arXiv:2402.10885.","journal-title":"arXiv preprint arXiv:2402.10885"},{"key":"e_1_3_4_157_1","doi-asserted-by":"crossref","unstructured":"Kerbl B Kopanas G Leimk\u00fchler T et al. (2023) 3D gaussian splatting for real-time radiance field rendering. https:\/\/arxiv.org\/abs\/2308.04079","DOI":"10.1145\/3592433"},{"key":"e_1_3_4_158_1","doi-asserted-by":"crossref","unstructured":"Kerr J Kim CM Goldberg K et al. (2023) LERF: language embedded radiance fields. In: 2023 Proceedings of the IEEE\/CVF International Conference on Computer Vision Paris France 1\u20136 October 2023 pp. 19729\u201319739.","DOI":"10.1109\/ICCV51070.2023.01807"},{"key":"e_1_3_4_159_1","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2023.3280028"},{"key":"e_1_3_4_160_1","doi-asserted-by":"publisher","DOI":"10.1145\/3505244"},{"key":"e_1_3_4_161_1","article-title":"Natural language robot programming: NLP integrated with autonomous robotic grasping","author":"Khan MA","year":"2023","unstructured":"Khan MA, Kenney M, Painter J, et al. (2023) Natural language robot programming: NLP integrated with autonomous robotic grasping. arXiv preprint arXiv:2304.02993.","journal-title":"arXiv preprint arXiv:2304.02993"},{"key":"e_1_3_4_162_1","doi-asserted-by":"crossref","unstructured":"Khandelwal A Weihs L Mottaghi R et al. (2022) Simple but effective: clip embeddings for embodied ai. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18\u201324 June 2022. pp. 14829\u201314838.","DOI":"10.1109\/CVPR52688.2022.01441"},{"key":"e_1_3_4_163_1","doi-asserted-by":"crossref","unstructured":"Khanna M Mao Y Jiang H et al. (2024) Habitat synthetic scenes dataset (HSSD-200): an analysis of 3D scene scale and realism tradeoffs for objectgoal navigation. In: 2024 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 16\u201322 June 2024 pp. 16384\u201316393.","DOI":"10.1109\/CVPR52733.2024.01550"},{"key":"e_1_3_4_164_1","article-title":"DROID: a large-scale in-the-wild robot manipulation dataset","author":"Khazatsky A","year":"2024","unstructured":"Khazatsky A, Pertsch K, Nair S, et al. (2024) DROID: a large-scale in-the-wild robot manipulation dataset. arXiv preprint arXiv:2403.12945.","journal-title":"arXiv preprint arXiv:2403.12945"},{"key":"e_1_3_4_165_1","first-page":"29","article-title":"Contact-implicit MPS: controlling diverse quadruped motions without pre-planned contact modes or trajectories","author":"Kim G","year":"2023","unstructured":"Kim G, Kang D, Kim JH, et al. (2023) Contact-implicit MPS: controlling diverse quadruped motions without pre-planned contact modes or trajectories. arXiv preprint arXiv:2312.08961: 29\u201384.","journal-title":"arXiv preprint arXiv:2312.08961"},{"key":"e_1_3_4_166_1","article-title":"Openvla: an open-source vision-language-action model","author":"Kim MJ","year":"2024","unstructured":"Kim MJ, Pertsch K, Karamcheti S, et al. (2024) Openvla: an open-source vision-language-action model. arXiv preprint arXiv:2406.09246.","journal-title":"arXiv preprint arXiv:2406.09246"},{"key":"e_1_3_4_167_1","article-title":"Fine-tuning vision-language-action models: optimizing speed and success","author":"Kim MJ","year":"2025","unstructured":"Kim MJ, Finn C, Liang P (2025) Fine-tuning vision-language-action models: optimizing speed and success. arXiv preprint arXiv:2502.19645.","journal-title":"arXiv preprint arXiv:2502.19645"},{"key":"e_1_3_4_168_1","article-title":"Segment anything","author":"Kirillov A","year":"2023","unstructured":"Kirillov A, Mintun E, Ravi N, et al. (2023) Segment anything. arXiv preprint arXiv:2304.02643.","journal-title":"arXiv preprint arXiv:2304.02643"},{"key":"e_1_3_4_169_1","doi-asserted-by":"publisher","DOI":"10.1007\/s43154-020-00021-6"},{"key":"e_1_3_4_170_1","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913495721"},{"key":"e_1_3_4_171_1","doi-asserted-by":"crossref","unstructured":"Kokic M Stork JA Haustein JA et al. (2017) Affordance detection for task-specific grasping using deep learning. In: 2017 IEEE-RAS 17th International Conference on Humanoid Robotics (Humanoids) Birmingham UK 15\u201317 November 2017 IEEE pp. 91\u201398.","DOI":"10.1109\/HUMANOIDS.2017.8239542"},{"key":"e_1_3_4_172_1","doi-asserted-by":"publisher","DOI":"10.1109\/21.179842"},{"issue":"1","key":"e_1_3_4_173_1","first-page":"1395","article-title":"A review of robot learning for manipulation: challenges, representations, and algorithms","volume":"22","author":"Kroemer O","year":"2021","unstructured":"Kroemer O, Niekum S, Konidaris G (2021) A review of robot learning for manipulation: challenges, representations, and algorithms. Journal of Machine Learning Research 22(1): 1395\u20131476.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_4_174_1","article-title":"NeRFbaselines: consistent and reproducible evaluation of novel view synthesis methods","author":"Kulhanek J","year":"2024","unstructured":"Kulhanek J, Sattler T (2024) NeRFbaselines: consistent and reproducible evaluation of novel view synthesis methods. arXiv preprint arXiv:2406.17345.","journal-title":"arXiv preprint arXiv:2406.17345"},{"key":"e_1_3_4_175_1","unstructured":"Kumar NJ (2023) Will scaling solve robotics? Perspectives from corl 2023. https:\/\/nishanthjkumar.com\/Will-Scaling-Solve-Robotics-Perspectives%-from-CoRL-2023\/ (Accessed 25 November 2023)."},{"key":"e_1_3_4_176_1","article-title":"MegaPose: 6D pose estimation of novel objects via render & compare","author":"Labb\u00e9 Y","year":"2022","unstructured":"Labb\u00e9 Y, Manuelli L, Mousavian A, et al. (2022) MegaPose: 6D pose estimation of novel objects via render & compare. arXiv preprint arXiv:2212.06870.","journal-title":"arXiv preprint arXiv:2212.06870"},{"key":"e_1_3_4_177_1","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2024.3351554"},{"key":"e_1_3_4_178_1","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2024.3434208"},{"key":"e_1_3_4_179_1","first-page":"1696","volume-title":"Conference on Robot Learning","author":"Lee KH","year":"2023","unstructured":"Lee KH, Xiao T, Li A, et al. (2023) PI-QT-Opt: predictive information improves multi-task robotic reinforcement learning at scale. In: Conference on Robot Learning. PMLR, pp. 1696\u20131707."},{"key":"e_1_3_4_180_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.displa.2024.102810"},{"key":"e_1_3_4_181_1","doi-asserted-by":"crossref","unstructured":"Li Y Wang G Ji X et al. (2018) Deepim: deep iterative matching for 6D pose estimation. In: Proceedings of the European Conference on Computer Vision (ECCV) Munich Germany 8-14 September 2018 pp. 683\u2013698.","DOI":"10.1007\/978-3-030-01231-1_42"},{"key":"e_1_3_4_182_1","doi-asserted-by":"crossref","unstructured":"Li S Ma X Liang H et al. (2019) Vision-based teleoperation of shadow dexterous hand using end-to-end deep neural network. In: 2019 International Conference on Robotics and Automation (ICRA) Montreal QC Canada 20\u201324 May 2019 IEEE pp. 416\u2013422.","DOI":"10.1109\/ICRA.2019.8794277"},{"key":"e_1_3_4_183_1","doi-asserted-by":"crossref","unstructured":"Li S Jiang J Ruppel P et al. (2020) A mobile robot hand-arm teleoperation system by vision and IMU. In: 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) Las Vegas NV USA 24 October 2020\u201324 January 2021 IEEE pp. 10900\u201310906.","DOI":"10.1109\/IROS45743.2020.9340738"},{"key":"e_1_3_4_184_1","article-title":"iGibson 2.0: object-centric simulation for robot learning of everyday household tasks","author":"Li C","year":"2021","unstructured":"Li C, Xia F, Mart\u00edn-Mart\u00edn R, et al. (2021) iGibson 2.0: object-centric simulation for robot learning of everyday household tasks. arXiv preprint arXiv:2108.03272.","journal-title":"arXiv preprint arXiv:2108.03272"},{"key":"e_1_3_4_185_1","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-021-0250-8"},{"key":"e_1_3_4_186_1","first-page":"80","volume-title":"Conference on Robot Learning","author":"Li C","year":"2023","unstructured":"Li C, Zhang R, Wong J, et al. (2023a) Behavior-1K: a benchmark for embodied ai with 1,000 everyday activities and realistic simulation. In: Conference on Robot Learning. PMLR, pp. 80\u201393."},{"key":"e_1_3_4_187_1","doi-asserted-by":"crossref","unstructured":"Li F Vutukur SR Yu H et al. (2023b) NeRF-pose: a first-reconstruct-then-regress approach for weakly-supervised 6D object pose estimation. In: 2023 Proceedings of the IEEE\/CVF International Conference on Computer Vision Paris France 1\u20136 October 2023 pp. 2123\u20132133.","DOI":"10.1109\/ICCVW60793.2023.00226"},{"key":"e_1_3_4_188_1","article-title":"Mastering robot manipulation with multimodal prompts through pretraining and multi-task fine-tuning","author":"Li J","year":"2023","unstructured":"Li J, Gao Q, Johnston M, et al. (2023c) Mastering robot manipulation with multimodal prompts through pretraining and multi-task fine-tuning. arXiv preprint arXiv:2310.09676.","journal-title":"arXiv preprint arXiv:2310.09676"},{"key":"e_1_3_4_189_1","article-title":"Vision-language foundation models as effective robot imitators","author":"Li X","year":"2023","unstructured":"Li X, Liu M, Zhang H, et al. (2023d) Vision-language foundation models as effective robot imitators. arXiv preprint arXiv:2311.01378.","journal-title":"arXiv preprint arXiv:2311.01378"},{"key":"e_1_3_4_190_1","doi-asserted-by":"crossref","unstructured":"Liang J Huang W Xia F et al. (2023) Code as policies: language model programs for embodied control. In:2023 IEEE International Conference on Robotics and Automation (ICRA) London United Kingdom 29 May\u201302 June 2023 IEEE pp. 9493\u20139500.","DOI":"10.1109\/ICRA48891.2023.10160591"},{"key":"e_1_3_4_191_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3439737"},{"key":"e_1_3_4_192_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA46639.2022.9811720"},{"key":"e_1_3_4_193_1","article-title":"Text2Motion: from natural language instructions to feasible plans","author":"Lin K","year":"2023","unstructured":"Lin K, Agia C, Migimatsu T, et al. (2023) Text2Motion: from natural language instructions to feasible plans. arXiv preprint arXiv:2303.12153.","journal-title":"arXiv preprint arXiv:2303.12153"},{"key":"e_1_3_4_194_1","article-title":"Data scaling laws in imitation learning for robotic manipulation","author":"Lin F","year":"2024","unstructured":"Lin F, Hu Y, Sheng P, et al. (2024a) Data scaling laws in imitation learning for robotic manipulation. arXiv preprint arXiv:2410.18647.","journal-title":"arXiv preprint arXiv:2410.18647"},{"key":"e_1_3_4_195_1","doi-asserted-by":"crossref","unstructured":"Lin J Liu L Lu D et al. (2024b) SAM-6D: segment anything model meets zero-shot 6D object pose estimation. In: 2024 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 16\u201322 June 2024 pp. 27906\u201327916.","DOI":"10.1109\/CVPR52733.2024.02636"},{"key":"e_1_3_4_196_1","unstructured":"Liu S Zhang X Zhang Z et al. (2021) Editing conditional radiance fields. In: 2023 Proceedings of the IEEE\/CVF International Conference on Computer Vision Paris France 1\u20136 October 2023 pp. 5773\u20135783."},{"key":"e_1_3_4_197_1","doi-asserted-by":"publisher","DOI":"10.1177\/02783649241273901"},{"key":"e_1_3_4_198_1","doi-asserted-by":"crossref","unstructured":"Liu L Xu W Fu H et al. (2022b) AKB-48: a real-world articulated object knowledge base. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18\u201324 June 2022. pp. 14809\u201314818.","DOI":"10.1109\/CVPR52688.2022.01439"},{"key":"e_1_3_4_199_1","first-page":"1","article-title":"Learning world models with identifiable factorization","volume":"36","author":"Liu Y","year":"2024","unstructured":"Liu Y, Huang B, Zhu Z, et al. (2024e) Learning world models with identifiable factorization. Advances in Neural Information Processing Systems 36: 1\u201334.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_200_1","article-title":"LLM+ P: empowering large language models with optimal planning proficiency","author":"Liu B","year":"2023","unstructured":"Liu B, Jiang Y, Zhang X, et al. (2023a) LLM+ P: empowering large language models with optimal planning proficiency. arXiv preprint arXiv:2304.11477.","journal-title":"arXiv preprint arXiv:2304.11477"},{"key":"e_1_3_4_201_1","doi-asserted-by":"crossref","unstructured":"Liu J Mahdavi-Amiri A Savva M (2023b) Paris: part-level reconstruction and motion analysis for articulated objects. In: 2023 Proceedings of the IEEE\/CVF International Conference on Computer Vision Paris France 1\u20136 October 2023 pp. 352\u2013363.","DOI":"10.1109\/ICCV51070.2023.00039"},{"key":"e_1_3_4_202_1","first-page":"53433","article-title":"Weakly supervised 3D open-vocabulary segmentation","volume":"36","author":"Liu K","year":"2023","unstructured":"Liu K, Zhan F, Zhang J, et al. (2023c) Weakly supervised 3D open-vocabulary segmentation. Advances in Neural Information Processing Systems 36: 53433\u201353456.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_203_1","doi-asserted-by":"crossref","unstructured":"Liu M Zhu Y Cai H et al. (2023d) PartSLIP: low-shot part segmentation for 3D point clouds via pretrained image-language models. In: 2023 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Vancouver BC Canada 17\u201324 June 2023. pp. 21736\u201321746.","DOI":"10.1109\/CVPR52729.2023.02082"},{"key":"e_1_3_4_204_1","doi-asserted-by":"crossref","unstructured":"Liu R Wu R Van Hoorick B et al. (2023e) Zero-1-to-3: zero-shot one image to 3D object. In: 2023 Proceedings of the IEEE\/CVF International Conference on Computer Vision Paris France 1\u20136 October 2023 pp. 9298\u20139309.","DOI":"10.1109\/ICCV51070.2023.00853"},{"key":"e_1_3_4_205_1","article-title":"Grounding DINO: marrying DINO with grounded pre-training for open-set object detection","author":"Liu S","year":"2023","unstructured":"Liu S, Zeng Z, Ren T, et al. (2023f) Grounding DINO: marrying DINO with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499.","journal-title":"arXiv preprint arXiv:2303.05499"},{"key":"e_1_3_4_206_1","article-title":"VoxAct-B: voxel-based acting and stabilizing policy for bimanual manipulation","author":"Liu I","year":"2024","unstructured":"Liu I, Arthur C, He S, et al. (2024b) VoxAct-B: voxel-based acting and stabilizing policy for bimanual manipulation. arXiv preprint arXiv:2407.04152.","journal-title":"arXiv preprint arXiv:2407.04152"},{"key":"e_1_3_4_207_1","article-title":"Deep learning-based object pose estimation: a comprehensive survey","author":"Liu J","year":"2024","unstructured":"Liu J, Sun W, Yang H, et al. (2024c) Deep learning-based object pose estimation: a comprehensive survey. arXiv preprint arXiv:2405.07801.","journal-title":"arXiv preprint arXiv:2405.07801"},{"key":"e_1_3_4_208_1","article-title":"RDT-1B: a diffusion foundation model for bimanual manipulation","author":"Liu S","year":"2024","unstructured":"Liu S, Wu L, Li B, et al. (2024d) RDT-1B: a diffusion foundation model for bimanual manipulation. arXiv preprint arXiv:2410.07864.","journal-title":"arXiv preprint arXiv:2410.07864"},{"key":"e_1_3_4_209_1","first-page":"1","article-title":"Libero: benchmarking knowledge transfer for lifelong robot learning","volume":"36","author":"Liu B","year":"2024","unstructured":"Liu B, Zhu Y, Gao C, et al. (2024a) Libero: benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems 36: 1\u201316.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_210_1","first-page":"349","volume-title":"European Conference on Computer Vision","author":"Lu G","year":"2024","unstructured":"Lu G, Zhang S, Wang Z, et al. (2024) Manigaussian: dynamic gaussian splatting for multi-task robotic manipulation. In: European Conference on Computer Vision. Springer, pp. 349\u2013366."},{"key":"e_1_3_4_211_1","article-title":"SERL: a software suite for sample-efficient robotic reinforcement learning","author":"Luo J","year":"2024","unstructured":"Luo J, Hu Z, Xu C, et al. (2024) SERL: a software suite for sample-efficient robotic reinforcement learning. arXiv preprint arXiv:2401.16013.","journal-title":"arXiv preprint arXiv:2401.16013"},{"key":"e_1_3_4_212_1","article-title":"Interactive language: talking to robots in real time","author":"Lynch C","year":"2023","unstructured":"Lynch C, Wahid A, Tompson J, et al. (2023) Interactive language: talking to robots in real time. IEEE Robotics and Automation Letters.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"e_1_3_4_213_1","article-title":"Cross-domain policy adaptation by capturing representation mismatch","author":"Lyu J","year":"2024","unstructured":"Lyu J, Bai C, Yang J, et al. (2024) Cross-domain policy adaptation by capturing representation mismatch. arXiv preprint arXiv:2405.15369.","journal-title":"arXiv preprint arXiv:2405.15369"},{"key":"e_1_3_4_214_1","article-title":"VIP: towards universal visual reward and representation via value-implicit pre-training","author":"Ma YJ","year":"2022","unstructured":"Ma YJ, Sodhani S, Jayaraman D, et al. (2022) VIP: towards universal visual reward and representation via value-implicit pre-training. arXiv preprint arXiv:2210.00030.","journal-title":"arXiv preprint arXiv:2210.00030"},{"key":"e_1_3_4_215_1","first-page":"23301","volume-title":"International Conference on Machine Learning","author":"Ma YJ","year":"2023","unstructured":"Ma YJ, Kumar V, Zhang A, et al. (2023a) LIV: language-image representations and rewards for robotic control. In: International Conference on Machine Learning. PMLR, 23301\u201323320."},{"key":"e_1_3_4_216_1","article-title":"Eureka: human-level reward design via coding large language models","author":"Ma YJ","year":"2023","unstructured":"Ma YJ, Liang W, Wang G, et al. (2023b) Eureka: human-level reward design via coding large language models. arXiv preprint arXiv:2310.12931.","journal-title":"arXiv preprint arXiv:2310.12931"},{"key":"e_1_3_4_217_1","article-title":"Dreureka: language model guided sim-to-real transfer","author":"Ma YJ","year":"2024","unstructured":"Ma YJ, Liang W, Wang HJ, et al. (2024) Dreureka: language model guided sim-to-real transfer. arXiv preprint arXiv:2406.01967.","journal-title":"arXiv preprint arXiv:2406.01967"},{"key":"e_1_3_4_218_1","article-title":"Where are we in the search for an artificial visual cortex for embodied intelligence?","author":"Majumdar A","year":"2023","unstructured":"Majumdar A, Yadav K, Arnaud S, et al. (2023) Where are we in the search for an artificial visual cortex for embodied intelligence? arXiv preprint arXiv:2303.18240.","journal-title":"arXiv preprint arXiv:2303.18240"},{"key":"e_1_3_4_219_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11831-022-09790-z"},{"key":"e_1_3_4_220_1","article-title":"Cacti: a framework for scalable multi-task multi-scene visual imitation learning","author":"Mandi Z","year":"2022","unstructured":"Mandi Z, Bharadhwaj H, Moens V, et al. (2022) Cacti: a framework for scalable multi-task multi-scene visual imitation learning. arXiv preprint arXiv:2212.05711.","journal-title":"arXiv preprint arXiv:2212.05711"},{"key":"e_1_3_4_221_1","article-title":"Real2code: reconstruct articulated objects via code generation","author":"Mandi Z","year":"2024","unstructured":"Mandi Z, Weng Y, Bauer D, et al. (2024) Real2code: reconstruct articulated objects via code generation. arXiv preprint arXiv:2406.08474.","journal-title":"arXiv preprint arXiv:2406.08474"},{"key":"e_1_3_4_222_1","article-title":"Mimicgen: a data generation system for scalable robot learning using human demonstrations","author":"Mandlekar A","year":"2023","unstructured":"Mandlekar A, Nasiriany S, Wen B, et al. (2023) Mimicgen: a data generation system for scalable robot learning using human demonstrations. arXiv preprint arXiv:2310.17596.","journal-title":"arXiv preprint arXiv:2310.17596"},{"key":"e_1_3_4_223_1","first-page":"734","volume-title":"Conference on Robot Learning","author":"Matas J","year":"2018","unstructured":"Matas J, James S, Davison AJ (2018) Sim-to-real reinforcement learning for deformable object manipulation Conference on Robot Learning. PMLR, 734\u2013743."},{"key":"e_1_3_4_224_1","article-title":"Towards generalist robot learning from internet video: a survey","author":"McCarthy R","year":"2024","unstructured":"McCarthy R, Tan DC, Schmidt D, et al. (2024) Towards generalist robot learning from internet video: a survey. arXiv preprint arXiv:2404.19664.","journal-title":"arXiv preprint arXiv:2404.19664"},{"key":"e_1_3_4_225_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2022.3180108"},{"key":"e_1_3_4_226_1","article-title":"Structured world models from human videos","author":"Mendonca R","year":"2023","unstructured":"Mendonca R, Bahl S, Pathak D (2023) Structured world models from human videos. arXiv preprint arXiv:2308.10901.","journal-title":"arXiv preprint arXiv:2308.10901"},{"key":"e_1_3_4_227_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00272"},{"key":"e_1_3_4_228_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503250"},{"key":"e_1_3_4_229_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20080-9_42"},{"key":"e_1_3_4_230_1","article-title":"Lan-grasp: using large language models for semantic object grasping","author":"Mirjalili R","year":"2023","unstructured":"Mirjalili R, Krawez M, Silenzi S, et al. (2023) Lan-grasp: using large language models for semantic object grasping. arXiv preprint arXiv:2310.05239.","journal-title":"arXiv preprint arXiv:2310.05239"},{"key":"e_1_3_4_231_1","doi-asserted-by":"crossref","unstructured":"Mo Y Zhang H Kong T (2023) Towards open-world interactive disambiguation for robotic grasping. In: 2023 IEEE International Conference on Robotics and Automation (ICRA) London United Kingdom 29 May\u201302 June 2023 IEEE pp. 8061\u20138067.","DOI":"10.1109\/ICRA48891.2023.10161333"},{"key":"e_1_3_4_232_1","doi-asserted-by":"crossref","unstructured":"Moon S Son H Hur D et al. (2024) Genflow: generalizable recurrent flow for 6d pose refinement of novel objects. In: 2024 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 16\u201322 June 2024 pp. 10039\u201310049.","DOI":"10.1109\/CVPR52733.2024.00957"},{"key":"e_1_3_4_233_1","article-title":"gradslam: dense slam meets automatic differentiation","author":"Murthy Jatavallabhula K","year":"2019","unstructured":"Murthy Jatavallabhula K, Iyer G, Paull L (2019) gradslam: dense slam meets automatic differentiation. arXiv preprint arXiv:1910.10672.","journal-title":"arXiv preprint arXiv:1910.10672"},{"key":"e_1_3_4_234_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19781-9_24"},{"key":"e_1_3_4_235_1","article-title":"R3M: a universal visual representation for robot manipulation","author":"Nair S","year":"2022","unstructured":"Nair S, Rajeswaran A, Kumar V, et al. (2022) R3M: a universal visual representation for robot manipulation. arXiv preprint arXiv:2203.12601.","journal-title":"arXiv preprint arXiv:2203.12601"},{"key":"e_1_3_4_236_1","doi-asserted-by":"publisher","DOI":"10.1080\/01691864.2020.1813623"},{"key":"e_1_3_4_237_1","article-title":"OpenVid-1M: a large-scale high-quality dataset for text-to-video generation","author":"Nan K","year":"2024","unstructured":"Nan K, Xie R, Zhou P, et al. (2024) OpenVid-1M: a large-scale high-quality dataset for text-to-video generation. arXiv preprint arXiv:2407.02371.","journal-title":"arXiv preprint arXiv:2407.02371"},{"key":"e_1_3_4_238_1","article-title":"Robocasa: large-scale simulation of everyday tasks for generalist robots","author":"Nasiriany S","year":"2024","unstructured":"Nasiriany S, Maddukuri A, Zhang L, et al. (2024) Robocasa: large-scale simulation of everyday tasks for generalist robots. arXiv preprint arXiv:2406.02523.","journal-title":"arXiv preprint arXiv:2406.02523"},{"key":"e_1_3_4_239_1","unstructured":"NDI (2024) Polaris vega XT. https:\/\/www.ndigital.com\/optical-navigation-technology\/polaris-vega-xt\/# (Accessed 2024)."},{"key":"e_1_3_4_240_1","article-title":"Language-conditioned affordance-pose detection in 3D point clouds","author":"Nguyen T","year":"2023","unstructured":"Nguyen T, Vu MN, Huang B, et al. (2023) Language-conditioned affordance-pose detection in 3D point clouds. arXiv preprint arXiv:2309.10911.","journal-title":"arXiv preprint arXiv:2309.10911"},{"key":"e_1_3_4_241_1","doi-asserted-by":"crossref","unstructured":"Nguyen VN Groueix T Salzmann M et al. (2024) GigaPose: fast and robust novel object pose estimation via one correspondence. In: 2024 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 16\u201322 June 2024 pp. 9903\u20139913.","DOI":"10.1109\/CVPR52733.2024.00945"},{"key":"e_1_3_4_242_1","doi-asserted-by":"crossref","unstructured":"Okamura AM Smaby N Cutkosky MR (2000) An overview of dexterous manipulation. In: Proceedings 2000 ICRA. Millennium Conference. IEEE International Conference on Robotics and Automation. Symposia Proceedings (Cat. No. 00CH37065) San Francisco CA USA 24\u201328 April 2000 IEEE Vol. 1 255\u2013262.","DOI":"10.1109\/ROBOT.2000.844067"},{"key":"e_1_3_4_243_1","doi-asserted-by":"crossref","unstructured":"\u00d6nol A\u00d6 Long P Pad\u0131r T (2019) Contact-implicit trajectory optimization based on a variable smooth contact model and successive convexification. In: 2019 International Conference on Robotics and Automation (ICRA) Montreal QC Canada 20\u201324 May 2019 IEEE pp. 2447\u20132453.","DOI":"10.1109\/ICRA.2019.8794250"},{"key":"e_1_3_4_244_1","article-title":"Open X-embodiment: robotic learning datasets and RT-X models","author":"Padalkar A","year":"2023","unstructured":"Padalkar A, Pooley A, Jain A, et al. (2023) Open X-embodiment: robotic learning datasets and RT-X models. arXiv preprint arXiv:2310.08864.","journal-title":"arXiv preprint arXiv:2310.08864"},{"key":"e_1_3_4_245_1","volume-title":"Planning, Sensing, and Control for Contact-Rich Robotic Manipulation With Quasi-Static Contact Models","author":"Pang T","year":"2023","unstructured":"Pang T (2023) Planning, Sensing, and Control for Contact-Rich Robotic Manipulation With Quasi-Static Contact Models. Massachusetts Institute of Technology."},{"key":"e_1_3_4_246_1","doi-asserted-by":"crossref","unstructured":"Pang T Tedrake R (2021) A convex quasistatic time-stepping scheme for rigid multibody systems with contact and friction. In: 2021 IEEE International Conference on Robotics and Automation (ICRA) Xi'an China 30 May\u201305 June 2021 IEEE pp. 6614\u20136620.","DOI":"10.1109\/ICRA48506.2021.9560941"},{"key":"e_1_3_4_247_1","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2023.3300230"},{"key":"e_1_3_4_248_1","article-title":"Pretrained language models as visual planners for human assistance","author":"Patel D","year":"2023","unstructured":"Patel D, Eghbalzadeh H, Kamra N, et al. (2023) Pretrained language models as visual planners for human assistance. arXiv preprint arXiv:2304.09179.","journal-title":"arXiv preprint arXiv:2304.09179"},{"key":"e_1_3_4_249_1","doi-asserted-by":"crossref","unstructured":"Pavlakos G Shan D Radosavovic I et al. (2024) Reconstructing hands in 3D with transformers. In: 2024 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 16\u201322 June 2024 pp. 9826\u20139836.","DOI":"10.1109\/CVPR52733.2024.00938"},{"key":"e_1_3_4_250_1","article-title":"Fast: efficient action tokenization for vision-language-action models","author":"Pertsch K","year":"2025","unstructured":"Pertsch K, Stachowicz K, Ichter B, et al. (2025) Fast: efficient action tokenization for vision-language-action models. arXiv preprint arXiv:2501.09747.","journal-title":"arXiv preprint arXiv:2501.09747"},{"issue":"3","key":"e_1_3_4_251_1","first-page":"3979","article-title":"Learning general and distinctive 3D local deep descriptors for point cloud registration","volume":"45","author":"Poiesi F","year":"2022","unstructured":"Poiesi F, Boscaini D (2022) Learning general and distinctive 3D local deep descriptors for point cloud registration. IEEE Transactions on Pattern Analysis and Machine Intelligence 45(3): 3979\u20133985.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_4_252_1","doi-asserted-by":"crossref","unstructured":"Puig X Ra K Boben M et al. (2018) Virtualhome: simulating household activities via programs. In: 2018 IEEE Conference on Computer Vision and Pattern Recognition CVPR 2018 Salt Lake City UT USA June 18-22 2018 pp. 8494\u20138502.","DOI":"10.1109\/CVPR.2018.00886"},{"key":"e_1_3_4_253_1","first-page":"5099","article-title":"Pointnet++: deep hierarchical feature learning on point sets in a metric space","volume":"30","author":"Qi CR","year":"2017","unstructured":"Qi CR, Yi L, Su H, et al. (2017) Pointnet++: deep hierarchical feature learning on point sets in a metric space. Advances in Neural Information Processing Systems 30: 5099\u20135108.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_254_1","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1685"},{"key":"e_1_3_4_255_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19842-7_33"},{"key":"e_1_3_4_256_1","first-page":"8748","volume-title":"International Conference on Machine Learning","author":"Radford A","year":"2021","unstructured":"Radford A, Kim JW, Hallacy C, et al. (2021) Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning. PMLR, pp. 8748\u20138763."},{"key":"e_1_3_4_257_1","first-page":"416","volume-title":"Conference on Robot Learning","author":"Radosavovic I","year":"2023","unstructured":"Radosavovic I, Xiao T, James S, et al. (2023) Real-world robot learning with masked visual pre-training. In: Conference on Robot Learning. PMLR, pp. 416\u2013426."},{"key":"e_1_3_4_258_1","first-page":"8821","volume-title":"International Conference on Machine Learning","author":"Ramesh A","year":"2021","unstructured":"Ramesh A, Pavlov M, Goh G, et al. (2021) Zero-shot text-to-image generation. In: International Conference on Machine Learning. PMLR, pp. 8821\u20138831."},{"key":"e_1_3_4_259_1","article-title":"Bayessim: adaptive domain randomization via probabilistic inference for robotics simulators","author":"Ramos F","year":"2019","unstructured":"Ramos F, Possas RC, Fox D (2019) Bayessim: adaptive domain randomization via probabilistic inference for robotics simulators. arXiv preprint arXiv:1906.01728.","journal-title":"arXiv preprint arXiv:1906.01728"},{"key":"e_1_3_4_260_1","article-title":"IMLE policy: fast and sample efficient visuomotor policy learning via implicit maximum likelihood estimation","author":"Rana K","year":"2025","unstructured":"Rana K, Lee R, Pershouse D, et al. (2025) IMLE policy: fast and sample efficient visuomotor policy learning via implicit maximum likelihood estimation. arXiv preprint arXiv:2502.12371.","journal-title":"arXiv preprint arXiv:2502.12371"},{"key":"e_1_3_4_261_1","article-title":"A generalist agent","author":"Reed S","year":"2022","unstructured":"Reed S, Zolna K, Parisotto E, et al. (2022) A generalist agent. arXiv preprint arXiv:2205.06175.","journal-title":"arXiv preprint arXiv:2205.06175"},{"key":"e_1_3_4_262_1","article-title":"Robots that ask for help: uncertainty alignment for large language model planners","author":"Ren AZ","year":"2023","unstructured":"Ren AZ, Dixit A, Bodrova A, et al. (2023a) Robots that ask for help: uncertainty alignment for large language model planners. arXiv preprint arXiv:2307.01928.","journal-title":"arXiv preprint arXiv:2307.01928"},{"key":"e_1_3_4_263_1","first-page":"1531","volume-title":"Conference on Robot Learning","author":"Ren AZ","year":"2023","unstructured":"Ren AZ, Govil B, Yang TY, et al. (2023b) Leveraging language for accelerated learning of tool manipulation. In: Conference on Robot Learning. PMLR, pp. 1531\u20131541."},{"key":"e_1_3_4_264_1","first-page":"1321","volume-title":"V-REP: a versatile and scalable robot simulation framework","author":"Rohmer E","year":"2013","unstructured":"Rohmer E, Singh SP, Freese M (2013) V-REP: a versatile and scalable robot simulation framework. In: Proceedings of the 2013 IEEE\/RSJ International Conference on Intelligent Robots and Systems, Tokyo, Japan, November 2013, IEEE, pp. 1321\u20131326."},{"key":"e_1_3_4_265_1","doi-asserted-by":"crossref","unstructured":"Rombach R Blattmann A Lorenz D et al. (2022) High-resolution image synthesis with latent diffusion models. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18\u201324 June 2022. pp. 10684\u201310695.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_4_266_1","doi-asserted-by":"crossref","unstructured":"Ruiz N Li Y Jampani V et al. (2023) Dreambooth: fine tuning text-to-image diffusion models for subject-driven generation. In: 2023 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Vancouver BC Canada 17\u201324 June 2023. pp. 22500\u201322510.","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"e_1_3_4_267_1","first-page":"262","volume-title":"Conference on Robot Learning","author":"Rusu AA","year":"2017","unstructured":"Rusu AA, Ve\u010der\u00edk M, Roth\u00f6rl T, et al. (2017) Sim-to-real robot learning from pixels with progressive nets. In: Conference on Robot Learning. PMLR, pp. 262\u2013270."},{"key":"e_1_3_4_268_1","first-page":"3634","volume-title":"Clear Grasp: 3D shape estimation of transparent objects for manipulation","author":"Sajjan S","year":"2020","unstructured":"Sajjan S, Moore M, Pan M, et al(2020) Clear Grasp: 3D shape estimation of transparent objects for manipulation. In: 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, May 31 - June 15 2020, IEEE, 3634\u20133642."},{"key":"e_1_3_4_269_1","first-page":"20154","article-title":"Graf: generative radiance fields for 3d-aware image synthesis","volume":"33","author":"Schwarz K","year":"2020","unstructured":"Schwarz K, Liao Y, Niemeyer M, et al. (2020) Graf: generative radiance fields for 3d-aware image synthesis. Advances in Neural Information Processing Systems 33: 20154\u201320166.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_270_1","doi-asserted-by":"crossref","unstructured":"Seker MY Kroemer O (2024) Estimating material properties of interacting objects using sum-gp-ucb. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). Yokohama Japan 13\u201317 May 2024 IEEE pp. 16684\u201316690.","DOI":"10.1109\/ICRA57147.2024.10610129"},{"key":"e_1_3_4_271_1","article-title":"RoboCQA: multimodal long-horizon reasoning for robotics","author":"Sermanet P","year":"2023","unstructured":"Sermanet P, Ding T, Zhao J, et al. (2023) RoboCQA: multimodal long-horizon reasoning for robotics. arXiv preprint arXiv:2311.00899.","journal-title":"arXiv preprint arXiv:2311.00899"},{"key":"e_1_3_4_272_1","doi-asserted-by":"crossref","unstructured":"Sermanet P Ding T Zhao J et al. (2024) Robovqa: multimodal long-horizon reasoning for robotics. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). Yokohama Japan 13-17 May 2024 IEEE pp. 645\u2013652.","DOI":"10.1109\/ICRA57147.2024.10610216"},{"key":"e_1_3_4_273_1","article-title":"Clip-fields: weakly supervised semantic fields for robotic memory","author":"Shafiullah NMM","year":"2022","unstructured":"Shafiullah NMM, Paxton C, Pinto L, et al. (2022) Clip-fields: weakly supervised semantic fields for robotic memory. arXiv preprint arXiv:2210.05663.","journal-title":"arXiv preprint arXiv:2210.05663"},{"key":"e_1_3_4_274_1","article-title":"Mutex: learning unified policies from multimodal task specifications","author":"Shah R","year":"2023","unstructured":"Shah R, Mart\u00edn-Mart\u00edn R, Zhu Y (2023) Mutex: learning unified policies from multimodal task specifications. arXiv preprint arXiv:2309.14320.","journal-title":"arXiv preprint arXiv:2309.14320"},{"key":"e_1_3_4_275_1","doi-asserted-by":"crossref","unstructured":"Shan D Geng J Shu M et al. (2020) Understanding human hands in contact at internet scale. In: 2020 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 13\u201319 June 2020 pp. 9869\u20139878.","DOI":"10.1109\/CVPR42600.2020.00989"},{"key":"e_1_3_4_276_1","first-page":"654","volume-title":"Conference on Robot Learning","author":"Shaw K","year":"2023","unstructured":"Shaw K, Bahl S, Pathak D (2023) Videodex: learning dexterity from internet videos. In: Conference on Robot Learning. PMLR, pp. 654\u2013665."},{"key":"e_1_3_4_277_1","article-title":"Distilled feature fields enable few-shot language-guided manipulation","author":"Shen W","year":"2023","unstructured":"Shen W, Yang G, Yu A, et al. (2023) Distilled feature fields enable few-shot language-guided manipulation. arXiv preprint arXiv:2308.07931.","journal-title":"arXiv preprint arXiv:2308.07931"},{"key":"e_1_3_4_278_1","article-title":"ASGrasp: generalizable transparent object reconstruction and grasping from rgb-d active stereo camera","author":"Shi J","year":"2024","unstructured":"Shi J, Jin Y, Li D, et al. (2024) ASGrasp: generalizable transparent object reconstruction and grasping from rgb-d active stereo camera. arXiv preprint arXiv:2405.05648.","journal-title":"arXiv preprint arXiv:2405.05648"},{"key":"e_1_3_4_279_1","article-title":"Is linear feedback on smoothed dynamics sufficient for stabilizing contact-rich plans?","author":"Shirai Y","year":"2024","unstructured":"Shirai Y, Zhao T, Suh H, et al. (2024) Is linear feedback on smoothed dynamics sufficient for stabilizing contact-rich plans? arXiv preprint arXiv:2411.06542.","journal-title":"arXiv preprint arXiv:2411.06542"},{"key":"e_1_3_4_280_1","unstructured":"Shridhar M Manuelli L Fox D (2021) Cliport: what and where pathways for robotic manipulation."},{"key":"e_1_3_4_281_1","first-page":"785","volume-title":"Conference on Robot Learning","author":"Shridhar M","year":"2023","unstructured":"Shridhar M, Manuelli L, Fox D (2023) Perceiver-actor: a multi-task transformer for robotic manipulation. In: Conference on Robot Learning. PMLR, pp. 785\u2013799."},{"key":"e_1_3_4_282_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-72992-8_1"},{"key":"e_1_3_4_283_1","doi-asserted-by":"crossref","unstructured":"Silva AJ Ramirez OAD Vega VP et al. (2009) Phantom omni haptic device: kinematic and manipulability. In: Proceedings of the 2009 Electronics Robotics and Automotive Mechanics Conference (CERMA) Cuernavaca Mexico 25 September 2009 IEEE pp. 193\u2013198.","DOI":"10.1109\/CERMA.2009.55"},{"key":"e_1_3_4_284_1","unstructured":"Silver T Hariprasad V Shuttleworth RS et al. (2022) PDDL planning with pretrained large language models. In: NeurIPS 2022 Foundation Models for Decision Making Workshop New Orleans LA USA 3 December 2022."},{"key":"e_1_3_4_285_1","doi-asserted-by":"crossref","unstructured":"Singh I Blukis V Mousavian A et al. (2023) ProgPrompt: generating situated robot task plans using large language models. In: 2023 IEEE International Conference on Robotics and Automation (ICRA) London United Kingdom 29 May\u201302 June 2023 IEEE pp. 11523\u201311530.","DOI":"10.1109\/ICRA48891.2023.10161317"},{"key":"e_1_3_4_286_1","article-title":"AVIDL: learning multi-stage tasks via pixel-level translation of human videos","author":"Smith L","year":"2019","unstructured":"Smith L, Dhawan N, Zhang M, et al. (2019) AVIDL: learning multi-stage tasks via pixel-level translation of human videos. arXiv preprint arXiv:1912.04443.","journal-title":"arXiv preprint arXiv:1912.04443"},{"key":"e_1_3_4_287_1","doi-asserted-by":"crossref","unstructured":"Song CH Wu J Washington C et al. (2023) LLM-planner: few-shot grounded planning for embodied agents with large language models. In: 2023 Proceedings of the IEEE\/CVF International Conference on Computer Vision Paris France 1\u20136 October 2023 pp. 2998\u20133009.","DOI":"10.1109\/ICCV51070.2023.00280"},{"key":"e_1_3_4_288_1","doi-asserted-by":"publisher","DOI":"10.1177\/027836499101000402"},{"key":"e_1_3_4_289_1","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-023-00669-7"},{"key":"e_1_3_4_290_1","doi-asserted-by":"crossref","unstructured":"Stoiber M Sundermeyer M Triebel R (2022) Iterative corresponding geometry: fusing region and depth for highly efficient 3D tracking of textureless objects. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18\u201324 June 2022. pp. 6855\u20136865.","DOI":"10.1109\/CVPR52688.2022.00673"},{"key":"e_1_3_4_291_1","article-title":"Open-world object manipulation using pre-trained vision-language models","author":"Stone A","year":"2023","unstructured":"Stone A, Xiao T, Lu Y, et al. (2023) Open-world object manipulation using pre-trained vision-language models. arXiv preprint arXiv:2303.00905.","journal-title":"arXiv preprint arXiv:2303.00905"},{"key":"e_1_3_4_292_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2022.3146931"},{"key":"e_1_3_4_293_1","unstructured":"Suh HT Simchowitz M Pang T et al. (2023) How does noising data affect learned contact dynamics?"},{"key":"e_1_3_4_294_1","doi-asserted-by":"crossref","unstructured":"Sun C Sun M Chen HT (2022) Direct voxel grid optimization: super-fast convergence for radiance fields reconstruction. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18\u201324 June 2022. pp. 5459\u20135469.","DOI":"10.1109\/CVPR52688.2022.00538"},{"key":"e_1_3_4_295_1","first-page":"10828","volume-title":"DiffCloud: real-to-sim from point clouds with differentiable simulation and rendering of deformable objects","author":"Sundaresan P","year":"2022","unstructured":"Sundaresan P, Antonova R, Bohgl J (2022) DiffCloud: real-to-sim from point clouds with differentiable simulation and rendering of deformable objects. In: 2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Kyoto, Japan, 23-27 October, 2022, IEEE, pp. 10828\u201310835."},{"key":"e_1_3_4_296_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01424-7_27"},{"key":"e_1_3_4_297_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2023.3320012"},{"key":"e_1_3_4_298_1","article-title":"UAD: unsupervised affordance distillation for generalization in robotic manipulation","author":"Tang Y","year":"2025","unstructured":"Tang Y, Huang W, Wang Y, et al. (2025) UAD: unsupervised affordance distillation for generalization in robotic manipulation. arXiv preprint arXiv:2506.09284.","journal-title":"arXiv preprint arXiv:2506.09284"},{"key":"e_1_3_4_299_1","article-title":"MOSAIC: learning unified multi-sensory object property representations for robot perception","author":"Tatiya G","year":"2023","unstructured":"Tatiya G, Francis J, Wu HH, et al. (2023) MOSAIC: learning unified multi-sensory object property representations for robot perception. arXiv preprint arXiv:2309.08508.","journal-title":"arXiv preprint arXiv:2309.08508"},{"key":"e_1_3_4_300_1","article-title":"Octo: an open-source generalist robot policy","author":"Team OM","year":"2024","unstructured":"Team OM, Ghosh D, Walke H, et al. (2024) Octo: an open-source generalist robot policy. arXiv preprint arXiv:2405.12213.","journal-title":"arXiv preprint arXiv:2405.12213"},{"key":"e_1_3_4_301_1","unstructured":"Tedrake R (2023) Underactuated robotics. https:\/\/underactuated.csail.mit.edu5"},{"key":"e_1_3_4_302_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2014.6907751"},{"key":"e_1_3_4_303_1","first-page":"5026","volume-title":"Mujoco: a physics engine for model-based control","author":"Todorov E","year":"2012","unstructured":"Todorov E, Erez T, Tassa Y (2012) Mujoco: a physics engine for model-based control. In: 2012 IEEE\/RSJ International Conference on Intelligent Robots and Systems, Vilamoura-Algarve, Portugal, 07-12 October 2012, IEEE, pp. 5026\u20135033."},{"key":"e_1_3_4_304_1","article-title":"Reconciling reality through simulation: a real-to-sim-to-real approach for robust manipulation","author":"Torne M","year":"2024","unstructured":"Torne M, Simeonov A, Li Z, et al. (2024) Reconciling reality through simulation: a real-to-sim-to-real approach for robust manipulation. arXiv preprint arXiv:2403.03949.","journal-title":"arXiv preprint arXiv:2403.03949"},{"key":"e_1_3_4_305_1","first-page":"25146","article-title":"Learning to synthesize programs as interpretable and generalizable policies","volume":"34","author":"Trivedi D","year":"2021","unstructured":"Trivedi D, Zhang J, Sun SH, et al. (2021) Learning to synthesize programs as interpretable and generalizable policies. Advances in Neural Information Processing Systems 34: 25146\u201325163.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_306_1","unstructured":"Valmeekam K Marquez M Olmo A et al. (2023) Planbench: an extensible benchmark for evaluating large language models on planning and reasoning about change. In: Thirty-Seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track New Orleans LA USA 10-16 December 2023."},{"key":"e_1_3_4_307_1","first-page":"6306","article-title":"Neural discrete representation learning","volume":"30","author":"Van Den Oord A","year":"2017","unstructured":"Van Den Oord A, Vinyals O, Kavukcuoglu K (2017) Neural discrete representation learning. Advances in Neural Information Processing Systems 30: 6306\u20136315.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_308_1","first-page":"20","article-title":"ChatGPT for robotics: design principles and model abilities","volume":"2","author":"Vemprala S","year":"2023","unstructured":"Vemprala S, Bonatti R, Bucker A, et al. (2023) ChatGPT for robotics: design principles and model abilities. Microsoft Autonomous Systems and Robotics Research 2: 20.","journal-title":"Microsoft Autonomous Systems and Robotics Research"},{"key":"e_1_3_4_309_1","unstructured":"Wang H Sridhar S Huang J et al. (2019) Normalized object coordinate space for category-level 6D object pose and size estimation. In: 2020 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 13\u201319 June 2020 pp. 2642\u20132651."},{"key":"e_1_3_4_310_1","first-page":"9075","article-title":"Is long horizon RL more difficult than short horizon RL?","volume":"33","author":"Wang R","year":"2020","unstructured":"Wang R, Du SS, Yang L, et al. (2020b) Is long horizon RL more difficult than short horizon RL? Advances in Neural Information Processing Systems 33: 9075\u20139085.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_311_1","doi-asserted-by":"crossref","unstructured":"Wang C Chai M He M et al. (2022a) Clip-NeRF: text-and-image driven manipulation of neural radiance fields. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18\u201324 June 2022. pp. 3835\u20133844.","DOI":"10.1109\/CVPR52688.2022.00381"},{"key":"e_1_3_4_312_1","article-title":"Imperative learning: A self-supervised neuro-symbolic learning framework for robot autonomy","volume":"027836492513531","author":"Wang C.","year":"2024","unstructured":"Wang C., Ji K., Geng J., et al. (2024) Imperative learning: A self-supervised neuro-symbolic learning framework for robot autonomy. The International Journal of Robotics Research 02783649251353181.","journal-title":"The International Journal of Robotics Research"},{"key":"e_1_3_4_313_1","first-page":"335","volume-title":"The International Symposium of Robotics Research","author":"Wang D","year":"2022","unstructured":"Wang D, Kohler C, Zhu X, et al. (2022b) BulletArm: an open-source robotic manipulation benchmark and learning framework. In: The International Symposium of Robotics Research. Springer, pp. 335\u2013350."},{"key":"e_1_3_4_314_1","article-title":"Mimicplay: long-horizon imitation learning by watching human play","author":"Wang C","year":"2023","unstructured":"Wang C, Fan L, Sun J, et al. (2023a) Mimicplay: long-horizon imitation learning by watching human play. arXiv preprint arXiv:2302.12422.","journal-title":"arXiv preprint arXiv:2302.12422"},{"key":"e_1_3_4_315_1","first-page":"70497","article-title":"Masked space-time hash encoding for efficient dynamic scene reconstruction","volume":"36","author":"Wang F","year":"2023","unstructured":"Wang F, Chen Z, Wang G, et al. (2023b) Masked space-time hash encoding for efficient dynamic scene reconstruction. Advances in Neural Information Processing Systems 36: 70497\u201370510.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_316_1","article-title":"GenSim: generating robotic simulation tasks via large language models","author":"Wang L","year":"2023","unstructured":"Wang L, Ling Y, Yuan Z, et al. (2023c) GenSim: generating robotic simulation tasks via large language models. arXiv preprint arXiv:2310.01361.","journal-title":"arXiv preprint arXiv:2310.01361"},{"key":"e_1_3_4_317_1","article-title":"InternVid: a large-scale video-text dataset for multimodal understanding and generation","author":"Wang Y","year":"2023","unstructured":"Wang Y, He Y, Li Y, et al. (2023d) InternVid: a large-scale video-text dataset for multimodal understanding and generation. arXiv preprint arXiv:2307.06942.","journal-title":"arXiv preprint arXiv:2307.06942"},{"key":"e_1_3_4_318_1","article-title":"Robogen: towards unleashing infinite data for automated robot learning via generative simulation","author":"Wang Y","year":"2023","unstructured":"Wang Y, Xian Z, Chen F, et al. (2023e) Robogen: towards unleashing infinite data for automated robot learning via generative simulation. arXiv preprint arXiv:2311.01455.","journal-title":"arXiv preprint arXiv:2311.01455"},{"key":"e_1_3_4_319_1","article-title":"VLM see, robot do: human demo video to robot action plan via vision language model","author":"Wang B","year":"2024","unstructured":"Wang B, Zhang J, Dong S, et al. (2024a) VLM see, robot do: human demo video to robot action plan via vision language model. arXiv preprint arXiv:2410.08792.","journal-title":"arXiv preprint arXiv:2410.08792"},{"key":"e_1_3_4_320_1","first-page":"10059","volume-title":"6-PACK: category-level 6D pose tracker with anchor-based keypoints","author":"Wang C","year":"2020","unstructured":"Wang C, Mart\u00edn-Mart\u00edn R, Xu D, et al. (2020a) 6-PACK: category-level 6D pose tracker with anchor-based keypoints. In: 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May - 31 August 2020, IEEE, pp. 10059\u201310066."},{"key":"e_1_3_4_321_1","article-title":"DexCap: scalable and portable mocap data collection system for dexterous manipulation","author":"Wang C","year":"2024","unstructured":"Wang C, Shi H, Wang W, et al. (2024b) DexCap: scalable and portable mocap data collection system for dexterous manipulation. arXiv preprint arXiv:2403.07788.","journal-title":"arXiv preprint arXiv:2403.07788"},{"key":"e_1_3_4_322_1","article-title":"Scaling proprioceptive-visual learning with heterogeneous pre-trained transformers","author":"Wang L","year":"2024","unstructured":"Wang L, Chen X, Zhao J, et al. (2024c) Scaling proprioceptive-visual learning with heterogeneous pre-trained transformers. arXiv preprint arXiv:2409.20537.","journal-title":"arXiv preprint arXiv:2409.20537"},{"key":"e_1_3_4_323_1","article-title":"LLaMA-mesh: unifying 3D mesh generation with language models","author":"Wang Z","year":"2024","unstructured":"Wang Z, Lorraine J, Wang Y, et al. (2024d) LLaMA-mesh: unifying 3D mesh generation with language models. arXiv preprint arXiv:2411.09595.","journal-title":"arXiv preprint arXiv:2411.09595"},{"key":"e_1_3_4_324_1","doi-asserted-by":"crossref","unstructured":"Wen B Bekris K (2021) Bundletrack: 6D pose tracking for novel objects without instance or category-level 3d models. In: 2021 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) Prague Czech Republic 27 September\u201301 October 2021 IEEE pp. 8067\u20138074.","DOI":"10.1109\/IROS51168.2021.9635991"},{"key":"e_1_3_4_325_1","doi-asserted-by":"crossref","unstructured":"Wen C Zhang Y Li Z et al. (2019) Pixel2Mesh++: multi-view 3D mesh generation via deformation. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision Seoul South Korea Oct. 27 - Nov. 2 2019 pp. 1042\u20131051.","DOI":"10.1109\/ICCV.2019.00113"},{"key":"e_1_3_4_326_1","doi-asserted-by":"crossref","unstructured":"Wen B Mitash C Ren B et al. (2020) se (3)-tracknet: data-driven 6d pose tracking by calibrating image residuals in synthetic domains. In: 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) Las Vegas NV USA 24 October 2020\u201324 January 2021 IEEE pp. 10367\u201310373.","DOI":"10.1109\/IROS45743.2020.9341314"},{"key":"e_1_3_4_327_1","doi-asserted-by":"crossref","unstructured":"Wen B Tremblay J Blukis V et al. (2023a) BundleSDF: neural 6-DoF tracking and 3D reconstruction of unknown objects. In: 2023 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Vancouver BC Canada 17\u201324 June 2023. pp. 606\u2013617.","DOI":"10.1109\/CVPR52729.2023.00066"},{"key":"e_1_3_4_328_1","article-title":"Any-point trajectory modeling for policy learning","author":"Wen C","year":"2023","unstructured":"Wen C, Lin X, So J, et al. (2023b) Any-point trajectory modeling for policy learning. arXiv preprint arXiv:2401.00025.","journal-title":"arXiv preprint arXiv:2401.00025"},{"key":"e_1_3_4_329_1","doi-asserted-by":"crossref","unstructured":"Wen B Yang W Kautz J et al. (2024a) Foundationpose: unified 6D pose estimation and tracking of novel objects. In: 2024 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 16\u201322 June 2024 pp. 17868\u201317879.","DOI":"10.1109\/CVPR52733.2024.01692"},{"key":"e_1_3_4_330_1","article-title":"TinyVLA: towards fast, data-efficient vision-language-action models for robotic manipulation","author":"Wen J","year":"2024","unstructured":"Wen J, Zhu Y, Li J, et al. (2024b) TinyVLA: towards fast, data-efficient vision-language-action models for robotic manipulation. arXiv preprint arXiv:2409.12514.","journal-title":"arXiv preprint arXiv:2409.12514"},{"key":"e_1_3_4_331_1","volume-title":"3D reconstruction","author":"Wikipedia Contributors","year":"2025","unstructured":"Wikipedia Contributors (2025) 3D reconstruction. https:\/\/en.wikipedia.org\/wiki\/3D_reconstruction (Accessed 30 January 2025)."},{"key":"e_1_3_4_332_1","article-title":"Vat-mart: learning visual action trajectory proposals for manipulating 3d articulated objects","author":"Wu R","year":"2021","unstructured":"Wu R, Zhao Y, Mo K, et al. (2021) Vat-mart: learning visual action trajectory proposals for manipulating 3d articulated objects. arXiv preprint arXiv:2106.14440.","journal-title":"arXiv preprint arXiv:2106.14440"},{"key":"e_1_3_4_333_1","article-title":"Unleashing large-scale video generative pre-training for visual robot manipulation","author":"Wu H","year":"2023","unstructured":"Wu H, Jing Y, Cheang C, et al. (2023a) Unleashing large-scale video generative pre-training for visual robot manipulation. arXiv preprint arXiv:2312.13139.","journal-title":"arXiv preprint arXiv:2312.13139"},{"key":"e_1_3_4_334_1","article-title":"TidyBot: personalized robot assistance with large language models","author":"Wu J","year":"2023","unstructured":"Wu J, Antonova R, Kan A, et al. (2023b) TidyBot: personalized robot assistance with large language models. arXiv preprint arXiv:2305.05658.","journal-title":"arXiv preprint arXiv:2305.05658"},{"key":"e_1_3_4_335_1","article-title":"GELLO: a general, low-cost, and intuitive teleoperation framework for robot manipulators","author":"Wu P","year":"2023","unstructured":"Wu P, Shentu Y, Yi Z, et al. (2023c) GELLO: a general, low-cost, and intuitive teleoperation framework for robot manipulators. arXiv preprint arXiv:2309.13037.","journal-title":"arXiv preprint arXiv:2309.13037"},{"key":"e_1_3_4_336_1","unstructured":"Xian Z Gkanatsios N Gervet T et al. (2023) ChainedDiffuser: unifying trajectory diffusion and keypose prediction for robotic manipulation. In: 7th Annual Conference on Robot Learning Atlanta Georgia USA 6-9 November 2023."},{"key":"e_1_3_4_337_1","doi-asserted-by":"crossref","unstructured":"Xiang F Qin Y Mo K et al. (2020) SAPIEN: a simulated part-based interactive environment. In: 2020 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 13\u201319 June 2020 pp. 11097\u201311107.","DOI":"10.1109\/CVPR42600.2020.01111"},{"key":"e_1_3_4_338_1","article-title":"Structured 3D latents for scalable and versatile 3D generation","author":"Xiang J","year":"2024","unstructured":"Xiang J, Lv Z, Xu S, et al. (2024) Structured 3D latents for scalable and versatile 3D generation. arXiv preprint arXiv:2412.01506.","journal-title":"arXiv preprint arXiv:2412.01506"},{"key":"e_1_3_4_339_1","article-title":"Robotic skill acquisition via instruction augmentation with vision-language models","author":"Xiao T","year":"2022","unstructured":"Xiao T, Chan H, Sermanet P, et al. (2022a) Robotic skill acquisition via instruction augmentation with vision-language models. arXiv preprint arXiv:2211.11736.","journal-title":"arXiv preprint arXiv:2211.11736"},{"key":"e_1_3_4_340_1","article-title":"Masked visual pre-training for motor control","author":"Xiao T","year":"2022","unstructured":"Xiao T, Radosavovic I, Darrell T, et al. (2022b) Masked visual pre-training for motor control. arXiv preprint arXiv:2203.06173.","journal-title":"arXiv preprint arXiv:2203.06173"},{"key":"e_1_3_4_341_1","article-title":"Robot learning in the era of foundation models: a survey","author":"Xiao X","year":"2023","unstructured":"Xiao X, Liu J, Wang Z, et al. (2023) Robot learning in the era of foundation models: a survey. arXiv preprint arXiv:2311.14379.","journal-title":"arXiv preprint arXiv:2311.14379"},{"key":"e_1_3_4_342_1","article-title":"Text2Reward: automated dense reward function generation for reinforcement learning","author":"Xie T","year":"2023","unstructured":"Xie T, Zhao S, Wu CH, et al. (2023a) Text2Reward: automated dense reward function generation for reinforcement learning. arXiv preprint arXiv:2309.11489.","journal-title":"arXiv preprint arXiv:2309.11489"},{"key":"e_1_3_4_343_1","article-title":"Translating natural language to planning goals with large-language models","author":"Xie Y","year":"2023","unstructured":"Xie Y, Yu C, Zhu T, et al. (2023b) Translating natural language to planning goals with large-language models. arXiv preprint arXiv:2302.05128.","journal-title":"arXiv preprint arXiv:2302.05128"},{"key":"e_1_3_4_344_1","doi-asserted-by":"crossref","unstructured":"Xiong H Li Q Chen YC et al. (2021) Learning by watching: physical imitation of manipulation skills from human videos. In: 2021 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS) Prague Czech Republic 27 September\u201301 October 2021 IEEE pp. 7827\u20137834.","DOI":"10.1109\/IROS51168.2021.9636080"},{"key":"e_1_3_4_345_1","article-title":"DensePhysnet: learning dense physical object representations via multi-step dynamic interactions","author":"Xu Z","year":"2019","unstructured":"Xu Z, Wu J, Zeng A, et al. (2019) DensePhysnet: learning dense physical object representations via multi-step dynamic interactions. arXiv preprint arXiv:1906.03853.","journal-title":"arXiv preprint arXiv:1906.03853"},{"key":"e_1_3_4_346_1","article-title":"A joint modeling of vision-language-action for target-oriented grasping in clutter","author":"Xu K","year":"2023","unstructured":"Xu K, Zhao S, Zhou Z, et al. (2023a) A joint modeling of vision-language-action for target-oriented grasping in clutter. arXiv preprint arXiv:2302.12610.","journal-title":"arXiv preprint arXiv:2302.12610"},{"key":"e_1_3_4_347_1","article-title":"Creative robot tool use with large language models","author":"Xu M","year":"2023","unstructured":"Xu M, Huang P, Yu W, et al. (2023b) Creative robot tool use with large language models. arXiv preprint arXiv:2310.13065.","journal-title":"arXiv preprint arXiv:2310.13065"},{"key":"e_1_3_4_348_1","article-title":"Manifoundation model for general-purpose robotic manipulation of contact synthesis with arbitrary objects and robots","author":"Xu Z","year":"2024","unstructured":"Xu Z, Gao C, Liu Z, et al. (2024) Manifoundation model for general-purpose robotic manipulation of contact synthesis with arbitrary objects and robots. arXiv preprint arXiv:2405.06964.","journal-title":"arXiv preprint arXiv:2405.06964"},{"key":"e_1_3_4_349_1","doi-asserted-by":"crossref","unstructured":"Xue H Hang T Zeng Y et al. (2022) Advancing high-resolution video-language representation with large-scale video transcriptions. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18\u201324 June 2022. pp. 5036\u20135045.","DOI":"10.1109\/CVPR52688.2022.00498"},{"key":"e_1_3_4_350_1","doi-asserted-by":"crossref","unstructured":"Xue L Gao M Xing C et al. (2023a) ULIP: learning a unified representation of language images and point clouds for 3D understanding. In: 2023 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Vancouver BC Canada 17\u201324 June 2023. pp. 1179\u20131189.","DOI":"10.1109\/CVPR52729.2023.00120"},{"key":"e_1_3_4_351_1","article-title":"ULIP-2: towards scalable multimodal pre-training for 3D understanding","author":"Xue L","year":"2023","unstructured":"Xue L, Yu N, Zhang S, et al. (2023b) ULIP-2: towards scalable multimodal pre-training for 3D understanding. arXiv preprint arXiv:2305.08275.","journal-title":"arXiv preprint arXiv:2305.08275"},{"key":"e_1_3_4_352_1","article-title":"VideoCoCa: video-text modeling with zero-shot transfer from contrastive captioners","author":"Yan S","year":"2022","unstructured":"Yan S, Zhu T, Wang Z, et al. (2022) VideoCoCa: video-text modeling with zero-shot transfer from contrastive captioners. arXiv preprint arXiv:2212.04979.","journal-title":"arXiv preprint arXiv:2212.04979"},{"key":"e_1_3_4_353_1","article-title":"DNAct: diffusion guided multi-task 3D policy learning","author":"Yan G","year":"2024","unstructured":"Yan G, Wu YH, Wang X (2024) DNAct: diffusion guided multi-task 3D policy learning. arXiv preprint arXiv:2403.04115.","journal-title":"arXiv preprint arXiv:2403.04115"},{"key":"e_1_3_4_354_1","article-title":"Track anything: segment anything meets videos","author":"Yang J","year":"2023","unstructured":"Yang J, Gao M, Li Z, et al. (2023a) Track anything: segment anything meets videos. arXiv preprint arXiv:2304.11968.","journal-title":"arXiv preprint arXiv:2304.11968"},{"key":"e_1_3_4_355_1","article-title":"Learning interactive real-world simulators","author":"Yang M","year":"2023","unstructured":"Yang M, Du Y, Ghasemipour K, et al. (2023b) Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114.","journal-title":"arXiv preprint arXiv:2310.06114"},{"key":"e_1_3_4_356_1","doi-asserted-by":"crossref","unstructured":"Yang T Jing Y Wu H et al. (2023c) Moma-force: visual-force imitation for real-world mobile manipulation. In: 2023 IEEE\/RSJ international conference on intelligent robots and systems (IROS) Detroit MI USA 01\u201305 October 2023 IEEE pp. 6847\u20136852.","DOI":"10.1109\/IROS55552.2023.10342371"},{"key":"e_1_3_4_357_1","article-title":"Plug in the safety chip: enforcing constraints for LLM-driven robot agents","author":"Yang Z","year":"2023","unstructured":"Yang Z, Raman SS, Shah A, et al. (2023d) Plug in the safety chip: enforcing constraints for LLM-driven robot agents. arXiv preprint arXiv:2309.09919.","journal-title":"arXiv preprint arXiv:2309.09919"},{"key":"e_1_3_4_358_1","doi-asserted-by":"crossref","unstructured":"Yang J Mark MS Vu B et al. (2024a) Robot fine-tuning made easy: pre-training rewards and policies for autonomous real-world reinforcement learning. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). Yokohama Japan 13\u201317 May 2024 IEEE pp. 4804\u20134811.","DOI":"10.1109\/ICRA57147.2024.10610421"},{"key":"e_1_3_4_359_1","first-page":"21875","article-title":"Depth anything v2","volume":"37","author":"Yang L","year":"2024","unstructured":"Yang L, Kang B, Huang Z, et al. (2024b) Depth anything v2. Advances in Neural Information Processing Systems 37: 21875\u201321911.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_360_1","doi-asserted-by":"crossref","unstructured":"Yang Y Jia B Zhi P et al. (2024c) PhyScene: physically interactable 3D scene synthesis for embodied ai. In: 2024 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA 16\u201322 June 2024 pp. 16262\u201316272.","DOI":"10.1109\/CVPR52733.2024.01539"},{"key":"e_1_3_4_361_1","first-page":"9125","article-title":"DetCLIP: dictionary-enriched visual-concept paralleled pre-training for open-world detection","volume":"35","author":"Yao L","year":"2022","unstructured":"Yao L, Han J, Wen Y, et al. (2022a) DetCLIP: dictionary-enriched visual-concept paralleled pre-training for open-world detection. Advances in Neural Information Processing Systems 35: 9125\u20139138.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_362_1","article-title":"React: synergizing reasoning and acting in language models","author":"Yao S","year":"2022","unstructured":"Yao S, Zhao J, Yu D, et al. (2022b) React: synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.","journal-title":"arXiv preprint arXiv:2210.03629"},{"key":"e_1_3_4_363_1","article-title":"Foundation reinforcement learning: towards embodied generalist agents with foundation prior assistance","author":"Ye W","year":"2023","unstructured":"Ye W, Zhang Y, Wang M, et al. (2023a) Foundation reinforcement learning: towards embodied generalist agents with foundation prior assistance. arXiv preprint arXiv:2310.02635.","journal-title":"arXiv preprint arXiv:2310.02635"},{"key":"e_1_3_4_364_1","doi-asserted-by":"crossref","unstructured":"Ye Y Li X Gupta A et al. (2023b) Affordance diffusion: synthesizing hand-object interactions. In: 2023 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Vancouver BC Canada 17\u201324 June 2023. pp. 22479\u201322489.","DOI":"10.1109\/CVPR52729.2023.02153"},{"key":"e_1_3_4_365_1","article-title":"Latent action pretraining from videos","author":"Ye S","year":"2024","unstructured":"Ye S, Jang J, Jeon B, et al. (2024) Latent action pretraining from videos. arXiv preprint arXiv:2410.11758.","journal-title":"arXiv preprint arXiv:2410.11758"},{"key":"e_1_3_4_366_1","doi-asserted-by":"publisher","DOI":"10.3389\/frobt.2025.1581110"},{"key":"e_1_3_4_367_1","doi-asserted-by":"publisher","DOI":"10.3389\/fnbot.2022.861825"},{"key":"e_1_3_4_368_1","doi-asserted-by":"crossref","unstructured":"Yu X Tang L Rao Y et al. (2022) Point-BERT: pre-training 3D point cloud transformers with masked point modeling. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18\u201324 June 2022. pp. 19313\u201319322.","DOI":"10.1109\/CVPR52688.2022.01871"},{"key":"e_1_3_4_369_1","article-title":"Scaling robot learning with semantically imagined experience","author":"Yu T","year":"2023","unstructured":"Yu T, Xiao T, Stone A, et al. (2023) Scaling robot learning with semantically imagined experience. arXiv preprint arXiv:2302.11550.","journal-title":"arXiv preprint arXiv:2302.11550"},{"key":"e_1_3_4_370_1","article-title":"General flow as foundation affordance for scalable robot learning","author":"Yuan C","year":"2024","unstructured":"Yuan C, Wen C, Zhang T, et al. (2024) General flow as foundation affordance for scalable robot learning. arXiv preprint arXiv:2401.11439.","journal-title":"arXiv preprint arXiv:2401.11439"},{"key":"e_1_3_4_371_1","doi-asserted-by":"crossref","unstructured":"Zareian A Rosa KD Hu DH et al. (2021) Open-vocabulary object detection using captions. In: 2021 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Nashville TN USA 20\u201325 June 2021 pp. 14393\u201314402.","DOI":"10.1109\/CVPR46437.2021.01416"},{"key":"e_1_3_4_372_1","doi-asserted-by":"crossref","unstructured":"Zarrin RS Jitosho R Yamane K (2023) Hybrid learning-and model-based planning and control of in-hand manipulation. In: 2023 IEEE\/RSJ international conference on intelligent robots and systems (IROS) Detroit MI USA 01\u201305 October 2023 IEEE pp. 8720\u20138726.","DOI":"10.1109\/IROS55552.2023.10342153"},{"key":"e_1_3_4_373_1","first-page":"1","article-title":"H-index: visual reinforcement learning with hand-informed representations for dexterous manipulation","volume":"36","author":"Ze Y","year":"2024","unstructured":"Ze Y, Liu Y, Shi R, et al. (2024a) H-index: visual reinforcement learning with hand-informed representations for dexterous manipulation. Advances in Neural Information Processing Systems 36: 1\u201316.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_374_1","first-page":"284","volume-title":"Conference on Robot Learning","author":"Ze Y","year":"2023","unstructured":"Ze Y, Yan G, Wu YH, et al. (2023) GNFactor: multi-task real robot learning with generalizable neural feature fields. In: Conference on Robot Learning. PMLR, pp. 284\u2013301."},{"key":"e_1_3_4_375_1","doi-asserted-by":"crossref","unstructured":"Ze Y Zhang G Zhang K et al. (2024b) 3D diffusion policy: generalizable visuomotor policy learning via simple 3D representations. In: ICRA 2024 Workshop on 3D Visual Representations for Robot Manipulation Yokohama Japan 17th May 2024.","DOI":"10.15607\/RSS.2024.XX.067"},{"key":"e_1_3_4_376_1","first-page":"23634","article-title":"MERLOT: multimodal neural script knowledge models","volume":"34","author":"Zellers R","year":"2021","unstructured":"Zellers R, Lu X, Hessel J, et al. (2021) MERLOT: multimodal neural script knowledge models. Advances in Neural Information Processing Systems 34: 23634\u201323651.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_4_377_1","doi-asserted-by":"crossref","unstructured":"Zhai X Kolesnikov A Houlsby N et al. (2022) Scaling vision transformers. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18\u201324 June 2022. pp. 12104\u201312113.","DOI":"10.1109\/CVPR52688.2022.01179"},{"key":"e_1_3_4_378_1","doi-asserted-by":"publisher","DOI":"10.5220\/0011314000003271"},{"key":"e_1_3_4_379_1","doi-asserted-by":"crossref","unstructured":"Zhang R Guo Z Zhang W et al. (2022b) PointCLIP: point cloud understanding by clip. In: 2022 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition New Orleans LA USA 18\u201324 June 2022 pp. 8552\u20138562.","DOI":"10.1109\/CVPR52688.2022.00836"},{"key":"e_1_3_4_380_1","article-title":"Faster segment anything: towards lightweight sam for mobile applications","author":"Zhang C","year":"2023","unstructured":"Zhang C, Han D, Qiao Y, et al. (2023a) Faster segment anything: towards lightweight sam for mobile applications. arXiv preprint arXiv:2306.14289.","journal-title":"arXiv preprint arXiv:2306.14289"},{"key":"e_1_3_4_381_1","article-title":"Building cooperative embodied agents modularly with large language models","author":"Zhang H","year":"2023","unstructured":"Zhang H, Du W, Shan J, et al. (2023b) Building cooperative embodied agents modularly with large language models. arXiv preprint arXiv:2307.02485.","journal-title":"arXiv preprint arXiv:2307.02485"},{"key":"e_1_3_4_382_1","article-title":"Bootstrap your own skills: learning to solve new tasks with large language model guidance","author":"Zhang J","year":"2023","unstructured":"Zhang J, Zhang J, Pertsch K, et al. (2023c) Bootstrap your own skills: learning to solve new tasks with large language model guidance. arXiv preprint arXiv:2310.10021.","journal-title":"arXiv preprint arXiv:2310.10021"},{"key":"e_1_3_4_383_1","doi-asserted-by":"crossref","unstructured":"Zhang X Kundu A Funkhouser T et al. (2023d) Nerflets: local radiance fields for efficient structure-aware 3D scene representation from 2D supervision. In: 2023 Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Vancouver BC Canada 17\u201324 June 2023 pp. 8274\u20138284.","DOI":"10.1109\/CVPR52729.2023.00800"},{"key":"e_1_3_4_384_1","doi-asserted-by":"crossref","unstructured":"Zhang Z Zhang L Wang Z et al. (2023e) Part-level scene reconstruction affords robot interaction. In: 2023 IEEE\/RSJ international conference on intelligent robots and systems (IROS) Detroit MI USA 01\u201305 October 2023 IEEE pp. 11178\u201311185.","DOI":"10.1109\/IROS55552.2023.10342208"},{"key":"e_1_3_4_385_1","article-title":"HiRT: enhancing robotic control with hierarchical robot transformers","author":"Zhang J","year":"2024","unstructured":"Zhang J, Guo Y, Chen X, et al. (2024a) HiRT: enhancing robotic control with hierarchical robot transformers. arXiv preprint arXiv:2410.05273.","journal-title":"arXiv preprint arXiv:2410.05273"},{"key":"e_1_3_4_386_1","article-title":"Vision-language models for vision tasks: a survey","author":"Zhang J","year":"2024","unstructured":"Zhang J, Huang J, Jin S, et al. (2024b) Vision-language models for vision tasks: a survey. IEEE Transactions on Pattern Analysis and Machine Intelligence.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_4_387_1","article-title":"Debiased offline representation learning for fast online adaptation in non-stationary dynamics","author":"Zhang X","year":"2024","unstructured":"Zhang X, Qiu W, Li YC, et al. (2024c) Debiased offline representation learning for fast online adaptation in non-stationary dynamics. arXiv preprint arXiv:2402.11317.","journal-title":"arXiv preprint arXiv:2402.11317"},{"key":"e_1_3_4_388_1","doi-asserted-by":"crossref","unstructured":"Zhao W Queralta JP Westerlund T (2020) Sim-to-real transfer in deep reinforcement learning for robotics: a survey. In: 2020 IEEE symposium series on computational intelligence (SSCI) Canberra ACT Australia 01\u201304 December 2020 IEEE pp. 737\u2013744.","DOI":"10.1109\/SSCI47803.2020.9308468"},{"key":"e_1_3_4_389_1","article-title":"Learning fine-grained bimanual manipulation with low-cost hardware","author":"Zhao TZ","year":"2023","unstructured":"Zhao TZ, Kumar V, Levine S, et al. (2023a) Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705.","journal-title":"arXiv preprint arXiv:2304.13705"},{"key":"e_1_3_4_390_1","article-title":"Fast segment anything","author":"Zhao X","year":"2023","unstructured":"Zhao X, Ding W, An Y, et al. (2023b) Fast segment anything. arXiv preprint arXiv:2306.12156.","journal-title":"arXiv preprint arXiv:2306.12156"},{"key":"e_1_3_4_391_1","article-title":"Chat with the environment: interactive multimodal perception using large language models","author":"Zhao X","year":"2023","unstructured":"Zhao X, Li M, Weber C, et al. (2023c) Chat with the environment: interactive multimodal perception using large language models. arXiv preprint arXiv:2303.08268.","journal-title":"arXiv preprint arXiv:2303.08268"},{"key":"e_1_3_4_392_1","volume-title":"A Survey of Optimization-based Task and Motion Planning: From Classical to Learning Approaches","author":"Zhao Z","year":"2024","unstructured":"Zhao Z, Cheng S, Ding Y, et al. (2024) A Survey of Optimization-based Task and Motion Planning: From Classical to Learning Approaches. IEEE\/ASME Transactions on Mechatronics."},{"key":"e_1_3_4_393_1","article-title":"3D-VLA: a 3D vision-language-action generative world model","author":"Zhen H","year":"2024","unstructured":"Zhen H, Qiu X, Chen P, et al. (2024) 3D-VLA: a 3D vision-language-action generative world model. arXiv preprint arXiv:2403.09631.","journal-title":"arXiv preprint arXiv:2403.09631"},{"key":"e_1_3_4_394_1","doi-asserted-by":"crossref","unstructured":"Zhi S Laidlow T Leutenegger S et al. (2021) In-place scene labelling and understanding with implicit scene representation. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision Montreal QC Canada 10\u201317 October 2021 pp. 15838\u201315847.","DOI":"10.1109\/ICCV48922.2021.01554"},{"key":"e_1_3_4_395_1","article-title":"iBOT: image BERT pre-training with online tokenizer","author":"Zhou J","year":"2021","unstructured":"Zhou J, Wei C, Wang H, et al. (2021) iBOT: image BERT pre-training with online tokenizer. arXiv preprint arXiv:2111.07832.","journal-title":"arXiv preprint arXiv:2111.07832"},{"key":"e_1_3_4_396_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-022-01653-1"},{"key":"e_1_3_4_397_1","article-title":"Language-conditioned learning for robotic manipulation: a survey","author":"Zhou H","year":"2023","unstructured":"Zhou H, Yao X, Meng Y, et al. (2023) Language-conditioned learning for robotic manipulation: a survey. arXiv preprint arXiv:2312.10807.","journal-title":"arXiv preprint arXiv:2312.10807"},{"key":"e_1_3_4_398_1","doi-asserted-by":"publisher","DOI":"10.3390\/s24072314"},{"key":"e_1_3_4_399_1","article-title":"Autonomous improvement of instruction following skills via foundation models","author":"Zhou Z","year":"2024","unstructured":"Zhou Z, Atreya P, Lee A, et al. (2024b) Autonomous improvement of instruction following skills via foundation models. arXiv preprint arXiv:2407.20635.","journal-title":"arXiv preprint arXiv:2407.20635"},{"key":"e_1_3_4_400_1","article-title":"Robosuite: a modular simulation framework and benchmark for robot learning","author":"Zhu Y","year":"2020","unstructured":"Zhu Y, Wong J, Mandlekar A, et al. (2020) Robosuite: a modular simulation framework and benchmark for robot learning. arXiv preprint arXiv:2009.12293.","journal-title":"arXiv preprint arXiv:2009.12293"},{"key":"e_1_3_4_401_1","doi-asserted-by":"crossref","unstructured":"Zhu H Meduri A Righetti L (2023a) Efficient object manipulation planning with monte carlo tree search. In: 2023 IEEE\/RSJ international conference on intelligent robots and systems (IROS) Detroit MI USA 01\u201305 October 2023 IEEE pp. 10628\u201310635.","DOI":"10.1109\/IROS55552.2023.10341813"},{"key":"e_1_3_4_402_1","first-page":"1199","volume-title":"Conference on Robot Learning","author":"Zhu Y","year":"2023","unstructured":"Zhu Y, Joshi A, Stone P, et al. (2023b) VIOLA: imitation learning for vision-based manipulation with object proposal priors. In: Conference on Robot Learning. PMLR, pp. 1199\u20131210."},{"key":"e_1_3_4_403_1","article-title":"Should we learn contact-rich manipulation policies from sampling-based planners?","author":"Zhu H","year":"2024","unstructured":"Zhu H, Zhao T, Ni X, et al. (2024) Should we learn contact-rich manipulation policies from sampling-based planners? arXiv preprint arXiv:2412.09743.","journal-title":"arXiv preprint arXiv:2412.09743"},{"key":"e_1_3_4_404_1","doi-asserted-by":"crossref","unstructured":"Zhuang J Wang C Lin L et al. (2023) Dreameditor: text-driven 3D scene editing with neural fields. In: SIGGRAPH Asia 2023 Conference Papers Sydney NSW Australia 12-15 December 2023 pp. 1\u201310.","DOI":"10.1145\/3610548.3618190"},{"key":"e_1_3_4_405_1","article-title":"GRS: generating robotic simulation tasks from real-world images","author":"Zook A","year":"2024","unstructured":"Zook A, Sun FY, Spjut J, et al. (2024) GRS: generating robotic simulation tasks from real-world images. arXiv preprint arXiv:2410.15536.","journal-title":"arXiv preprint arXiv:2410.15536"},{"key":"e_1_3_4_406_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-024-02183-8"}],"container-title":["The International Journal of Robotics Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649251390579","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/02783649251390579","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649251390579","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T05:59:59Z","timestamp":1780639199000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/02783649251390579"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,20]]},"references-count":405,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["10.1177\/02783649251390579"],"URL":"https:\/\/doi.org\/10.1177\/02783649251390579","relation":{},"ISSN":["0278-3649","1741-3176"],"issn-type":[{"value":"0278-3649","type":"print"},{"value":"1741-3176","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,20]]}}}