{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,4]],"date-time":"2026-08-04T21:02:11Z","timestamp":1785877331260,"version":"3.56.0"},"reference-count":262,"publisher":"SAGE Publications","issue":"5","license":[{"start":{"date-parts":[[2024,9,25]],"date-time":"2024-09-25T00:00:00Z","timestamp":1727222400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"DOI":"10.13039\/100000084","name":"Directorate for Engineering","doi-asserted-by":"publisher","award":["2044149"],"award-info":[{"award-number":["2044149"]}],"id":[{"id":"10.13039\/100000084","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000185","name":"Defense Advanced Research Projects Agency","doi-asserted-by":"publisher","award":["HR001120C0107"],"award-info":[{"award-number":["HR001120C0107"]}],"id":[{"id":"10.13039\/100000185","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000006","name":"Office of Naval Research","doi-asserted-by":"publisher","award":["N00014-23-1-2148"],"award-info":[{"award-number":["N00014-23-1-2148"]}],"id":[{"id":"10.13039\/100000006","id-type":"DOI","asserted-by":"publisher"}]},{"name":"ASEE e-Fellows"},{"name":"NSF Graduate Research Fellowship"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of Robotics Research"],"published-print":{"date-parts":[[2025,4]]},"abstract":"<jats:p>\n                    We survey applications of pretrained foundation models in robotics. Traditional deep learning models in robotics are trained on small datasets tailored for specific tasks, which limits their adaptability across diverse applications. In contrast, foundation models pretrained on internet-scale data appear to have superior generalization capabilities, and in some instances display an emergent ability to find zero-shot solutions to problems that are not present in the training data. Foundation models may hold the potential to enhance various components of the robot autonomy stack, from perception to decision-making and control. For example, large language models can generate code or provide common sense reasoning, while vision-language models enable open-vocabulary visual recognition. However, significant open research challenges remain, particularly around the scarcity of robot-relevant training data, safety guarantees and uncertainty quantification, and real-time execution. In this survey, we study recent papers that have used or built foundation models to solve robotics problems. We explore how foundation models contribute to improving robot capabilities in the domains of perception, decision-making, and control. We discuss the challenges hindering the adoption of foundation models in robot autonomy and provide opportunities and potential pathways for future advancements. The GitHub project corresponding to this paper can be found here:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/robotics-survey\/Awesome-Robotics-Foundation-Models\">https:\/\/github.com\/robotics-survey\/Awesome-Robotics-Foundation-Models<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1177\/02783649241281508","type":"journal-article","created":{"date-parts":[[2024,9,26]],"date-time":"2024-09-26T14:16:10Z","timestamp":1727360170000},"page":"701-739","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":179,"title":["Foundation models in robotics: Applications, challenges, and the future"],"prefix":"10.1177","volume":"44","author":[{"given":"Roya","family":"Firoozi","sequence":"first","affiliation":[{"name":"Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Johnathan","family":"Tucker","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stephen","family":"Tian","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anirudha","family":"Majumdar","sequence":"additional","affiliation":[{"name":"Princeton University, Princeton, NJ, USA"},{"name":"Google DeepMind, Mountain View, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiankai","family":"Sun","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weiyu","family":"Liu","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuke","family":"Zhu","sequence":"additional","affiliation":[{"name":"UT Austin, Austin, TX, USA"},{"name":"NVIDIA, Santa Clara, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuran","family":"Song","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ashish","family":"Kapoor","sequence":"additional","affiliation":[{"name":"Scaled Foundations, Kirkland, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Karol","family":"Hausman","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, CA, USA"},{"name":"Google DeepMind, Mountain View, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Brian","family":"Ichter","sequence":"additional","affiliation":[{"name":"Google DeepMind, Mountain View, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Danny","family":"Driess","sequence":"additional","affiliation":[{"name":"Google DeepMind, Mountain View, CA, USA"},{"name":"TU Berlin, Berlin, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiajun","family":"Wu","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cewu","family":"Lu","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mac","family":"Schwager","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2024,9,25]]},"reference":[{"key":"e_1_3_2_2_1","doi-asserted-by":"crossref","unstructured":"Anderson P Wu Q Teney D et al. (2018) Vision-and-language navigation: interpreting visually-grounded navigation instructions in real environments. In: CVPR pp. 3674\u20133683. IEEE.","DOI":"10.1109\/CVPR.2018.00387"},{"key":"e_1_3_2_3_1","doi-asserted-by":"crossref","unstructured":"Arandjelovic R Gronat P Torii A et al. (2016) NetVLAD: CNN architecture for weakly supervised place recognition. In: CVPR pp. 5297\u20135307. IEEE.","DOI":"10.1109\/CVPR.2016.572"},{"key":"e_1_3_2_4_1","doi-asserted-by":"crossref","unstructured":"Bahl S Mendonca R Chen L et al. (2023) Affordances from human videos as a versatile representation for robotics. In: CVPR. IEEE.","DOI":"10.1109\/CVPR52729.2023.01324"},{"key":"e_1_3_2_5_1","unstructured":"Baker B Akkaya I Zhokov P et al. (2022) Video PreTraining (VPT): Learning to act by watching unlabeled online videos. In: NeurIPS. IEEE."},{"key":"e_1_3_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2798607"},{"key":"e_1_3_2_7_1","article-title":"LL3MVN: leveraging large language models for visual target navigation","author":"Bangguo Y","year":"2023","unstructured":"Bangguo Y, Hamidreza K, Ming C (2023) LL3MVN: leveraging large language models for visual target navigation. arXiv preprint arXiv:2304.05501.","journal-title":"arXiv preprint arXiv:2304.05501"},{"key":"e_1_3_2_8_1","article-title":"On the opportunities and risks of foundation models","author":"Bommasani R","year":"2021","unstructured":"Bommasani R, Hudson DA, Adeli E, et al. (2021) On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.","journal-title":"arXiv preprint arXiv:2108.07258"},{"key":"e_1_3_2_9_1","doi-asserted-by":"crossref","unstructured":"Bonatti R Vemprala S Ma S et al. (2023) PACT: perception-action causal transformer for autoregressive robotics pretraining. In: IROS. IEEE.","DOI":"10.1109\/IROS55552.2023.10342381"},{"key":"e_1_3_2_10_1","article-title":"RoboCat: a self-improving foundation agent for robotic manipulation","author":"Bousmalis K","year":"2023","unstructured":"Bousmalis K, Vezzani G, Rao D, et al. (2023) RoboCat: a self-improving foundation agent for robotic manipulation. arXiv preprint arXiv:2306.11706.","journal-title":"arXiv preprint arXiv:2306.11706"},{"key":"e_1_3_2_11_1","volume-title":"Time Series Analysis: Forecasting and Control","author":"Box GE","year":"2015","unstructured":"Box GE, Jenkins GM, Reinsel GC, et al. (2015) Time Series Analysis: Forecasting and Control. Hoboken, New Jersey: John Wiley & Sons."},{"key":"e_1_3_2_12_1","article-title":"Rt-1: robotics transformer for real-world control at scale","author":"Brohan A","year":"2022","unstructured":"Brohan A, Brown N, Carbajal J, et al. (2022) Rt-1: robotics transformer for real-world control at scale arXiv Preprint arXiv:2212.06817.","journal-title":"arXiv Preprint arXiv:2212.06817"},{"key":"e_1_3_2_13_1","doi-asserted-by":"crossref","unstructured":"Brohan A Brown N Carbajal J et al. (2023a) RT-1: robotics transformer for real-world control at scale. In: RSS.","DOI":"10.15607\/RSS.2023.XIX.025"},{"key":"e_1_3_2_14_1","unstructured":"Brohan A Chebotar Y Finn C et al. (2023b) Do as I can not as I say: grounding language in robotic affordances. In: CoRL pp. 287\u2013318. PMLR."},{"key":"e_1_3_2_15_1","unstructured":"Brooks T Peebles B Holmes C et al. (2024) Video generation models as world simulators. URL: https:\/\/Openai.Com\/Research\/Video-Generation-Models-As-World-Simulators"},{"key":"e_1_3_2_16_1","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown T","year":"2020","unstructured":"Brown T, Mann B, Ryder N, et al. (2020) Language models are few-shot learners. NeurIPS 33: 1877\u20131901.","journal-title":"NeurIPS"},{"key":"e_1_3_2_17_1","doi-asserted-by":"crossref","unstructured":"Bucker A Figueredo L Haddadin S et al. (2023) LATTE: LAnguage trajectory TransformEr. In: ICRA pp. 7287\u20137294. IEEE.","DOI":"10.1109\/ICRA48891.2023.10161068"},{"key":"e_1_3_2_18_1","doi-asserted-by":"crossref","unstructured":"Cai F Koutsoukos X (2020) Real-time out-of-distribution detection in learning-enabled cyber-physical systems. In: ICCPS. IEEE.","DOI":"10.1109\/ICCPS48487.2020.00024"},{"key":"e_1_3_2_19_1","doi-asserted-by":"crossref","unstructured":"Caron M Touvron H Misra I et al. (2021) Emerging properties in self-supervised vision transformers. In: ICCV. IEEE.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"e_1_3_2_20_1","article-title":"ShapeNet: an information-rich 3D model repository","author":"Chang AX","year":"2015","unstructured":"Chang AX, Funkhouser T, Guibas L, et al. (2015) ShapeNet: an information-rich 3D model repository. arXiv preprint arXiv:1512.03012.","journal-title":"arXiv preprint arXiv:1512.03012"},{"key":"e_1_3_2_21_1","doi-asserted-by":"publisher","unstructured":"Chang A Dai A Funkhouser T et al. (2017) Matterport3D: learning from RGB-D data in indoor environments. In: 3DV pp. 667\u2013676. DOI: 10.1109\/3DV.2017.00081.","DOI":"10.1109\/3DV.2017.00081"},{"key":"e_1_3_2_22_1","unstructured":"Chen T Kornblith S Norouzi M et al. (2020) A simple framework for contrastive learning of visual representations. In: ICML."},{"key":"e_1_3_2_23_1","article-title":"Evaluating large language models trained on code","author":"Chen M","year":"2021","unstructured":"Chen M, Tworek J, Jun H, et al. (2021) Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.","journal-title":"arXiv preprint arXiv:2107.03374"},{"key":"e_1_3_2_24_1","article-title":"Leveraging large language models for robot 3D scene understanding","author":"Chen W","year":"2022","unstructured":"Chen W, Hu S, Talak R, et al. (2022a) Leveraging large language models for robot 3D scene understanding. arXiv preprint arXiv:2209.05629.","journal-title":"arXiv preprint arXiv:2209.05629"},{"key":"e_1_3_2_25_1","unstructured":"Chen X Wang X Changpinyo S et al. (2022b) PaLI: a jointly-scaled multilingual language-image model. In NeurIPS. IEEE."},{"key":"e_1_3_2_26_1","article-title":"How to not train your dragon: training-free embodied object goal navigation with semantic frontiers","author":"Chen J","year":"2023","unstructured":"Chen J, Li G, Kumar S, et al. (2023a) How to not train your dragon: training-free embodied object goal navigation with semantic frontiers. arXiv preprint arXiv:2305.16925.","journal-title":"arXiv preprint arXiv:2305.16925"},{"key":"e_1_3_2_27_1","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.adc9244"},{"key":"e_1_3_2_28_1","article-title":"PaLI-X: on scaling up a multilingual vision and language model","author":"Chen X","year":"2023","unstructured":"Chen X, Djolonga J, Padlewski P, et al. (2023c) PaLI-X: on scaling up a multilingual vision and language model. arXiv preprint arXiv:2305.18565.","journal-title":"arXiv preprint arXiv:2305.18565"},{"key":"e_1_3_2_29_1","article-title":"AutoTAMP: autoregressive task and motion planning with llms as translators and checkers","author":"Chen Y","year":"2023","unstructured":"Chen Y, Arkin J, Zhang Y, et al. (2023d) AutoTAMP: autoregressive task and motion planning with llms as translators and checkers. arXiv preprint arXiv:2306.06531.","journal-title":"arXiv preprint arXiv:2306.06531"},{"key":"e_1_3_2_30_1","article-title":"NL2TL: transforming natural languages to temporal logics using large language models","author":"Chen Y","year":"2023","unstructured":"Chen Y, Gandhi R, Zhang Y, et al. (2023e) NL2TL: transforming natural languages to temporal logics using large language models. arXiv preprint arXiv:2305.07766.","journal-title":"arXiv preprint arXiv:2305.07766"},{"key":"e_1_3_2_31_1","doi-asserted-by":"crossref","unstructured":"Chen Z Kiami S Gupta A et al. (2023f) GenAug: retargeting behaviors to unseen situations via generative augmentation. In: RSS.","DOI":"10.15607\/RSS.2023.XIX.010"},{"key":"e_1_3_2_32_1","doi-asserted-by":"crossref","unstructured":"Chen B Xu Z Kirmani S et al. (2024) Spatialvlm: endowing vision-language models with spatial reasoning capabilities. In CVPR. IEEE.","DOI":"10.1109\/CVPR52733.2024.01370"},{"key":"e_1_3_2_33_1","doi-asserted-by":"crossref","unstructured":"Cheng HK Schwing AG (2022) XMem: long-term video object segmentation with an atkinson-shiffrin memory model. In: ECCV pp. 640\u2013658. Springer.","DOI":"10.1007\/978-3-031-19815-1_37"},{"key":"e_1_3_2_34_1","article-title":"PaLM: scaling language modeling with pathways","author":"Chowdhery A","year":"2022","unstructured":"Chowdhery A, Narang S, Devlin J, et al. (2022) PaLM: scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.","journal-title":"arXiv preprint arXiv:2204.02311"},{"key":"e_1_3_2_35_1","volume-title":"Robot Learning for Manipulation of Granular Materials Using Vision and Sound","author":"Clarke S","year":"2019","unstructured":"Clarke S (2019) Robot Learning for Manipulation of Granular Materials Using Vision and Sound. Master\u2019s Thesis. Carnegie Mellon University."},{"key":"e_1_3_2_36_1","doi-asserted-by":"crossref","unstructured":"Dai Z Yang Z Yang Y et al. (2019) Transformer-XL: attentive language models beyond a fixed-length context. In: ACL.","DOI":"10.18653\/v1\/P19-1285"},{"key":"e_1_3_2_37_1","doi-asserted-by":"crossref","unstructured":"Damen D Doughty H Farinella GM et al. (2018) Scaling egocentric vision: the EPIC-KITCHENS dataset. In: ECCV. Springer.","DOI":"10.1007\/978-3-030-01225-0_44"},{"key":"e_1_3_2_38_1","unstructured":"Dasari S Ebert F Tian S et al. (2019) RoboNet: large-scale multi-robot learning. In: CoRL."},{"key":"e_1_3_2_39_1","unstructured":"Dasgupta I Kaeser-Chen C Marino K et al. (2022) Collaborating with language models for embodied reasoning. In: Second Workshop on Language and Reinforcement Learning."},{"key":"e_1_3_2_40_1","unstructured":"Dehghani M Djolonga J Mustafa B et al. (2023) Scaling vision transformers to 22 billion parameters. In: ICML."},{"key":"e_1_3_2_41_1","doi-asserted-by":"crossref","unstructured":"Deitke M Han W Herrasti A et al. (2020) RoboTHOR: an open simulation-to-real embodied AI platform. In: CVPR. IEEE.","DOI":"10.1109\/CVPR42600.2020.00323"},{"key":"e_1_3_2_42_1","doi-asserted-by":"crossref","unstructured":"DeTone D Malisiewicz T Rabinovich A (2018) Superpoint: self-supervised interest point detection and description. In: CVPR deep learning for visual SLAM workshop. IEEE.","DOI":"10.1109\/CVPRW.2018.00060"},{"key":"e_1_3_2_43_1","article-title":"BERT: pre-training of deep bidirectional transformers for language understanding","author":"Devlin J","year":"2018","unstructured":"Devlin J, Chang MW, Lee K, et al. (2018) BERT: pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.","journal-title":"arXiv preprint arXiv:1810.04805"},{"key":"e_1_3_2_44_1","first-page":"1","article-title":"A survey on safety-critical driving scenario generation\u2014a methodological perspective","volume":"99","author":"Ding W","year":"2023","unstructured":"Ding W, Xu C, Arief M, et al. (2023) A survey on safety-critical driving scenario generation\u2014a methodological perspective. IEEE ITSC 99: 1\u201319.","journal-title":"IEEE ITSC"},{"key":"e_1_3_2_45_1","article-title":"A survey for in-context learning","author":"Dong Q","year":"2022","unstructured":"Dong Q, Li L, Dai D, et al. (2022) A survey for in-context learning. arXiv preprint arXiv:2301.00234.","journal-title":"arXiv preprint arXiv:2301.00234"},{"key":"e_1_3_2_46_1","unstructured":"Dosovitskiy A Beyer L Kolesnikov A et al. (2021) An image is worth 16x16 words: transformers for image recognition at scale. In: ICLR."},{"key":"e_1_3_2_47_1","article-title":"PaLM-E: an embodied multimodal language model","author":"Driess D","year":"2023","unstructured":"Driess D, Xia F, Sajjadi MSM, et al. (2023) PaLM-E: an embodied multimodal language model. arXiv preprint arXiv:2303.03378.","journal-title":"arXiv preprint arXiv:2303.03378"},{"key":"e_1_3_2_48_1","article-title":"Zero-shot visual question answering with language model feedback","author":"Du Y","year":"2023","unstructured":"Du Y, Li J, Tang T, et al. (2023a) Zero-shot visual question answering with language model feedback. arXiv preprint arXiv:2305.17006.","journal-title":"arXiv preprint arXiv:2305.17006"},{"key":"e_1_3_2_49_1","unstructured":"Du Y Watkins O Wang Z et al. (2023b) Guiding pretraining in reinforcement learning with large language models. In: ICML pp. 8657\u20138677. PMLR."},{"key":"e_1_3_2_50_1","unstructured":"Du Y Yang M Dai B et al. (2023c) Learning universal policies via text-guided video generation. In: NeurIPS."},{"key":"e_1_3_2_51_1","article-title":"Video language planning","author":"Du Y","year":"2023","unstructured":"Du Y, Yang M, Florence P, et al. (2023d) Video language planning. arXiv preprint arXiv:2310.10625.","journal-title":"arXiv preprint arXiv:2310.10625"},{"key":"e_1_3_2_52_1","article-title":"RL\u02c62: fast reinforcement learning via slow reinforcement learning","author":"Duan Y","year":"2016","unstructured":"Duan Y, Schulman J, Chen X, et al. (2016) RL\u02c62: fast reinforcement learning via slow reinforcement learning. arXiv preprint arXiv:1611.02779.","journal-title":"arXiv preprint arXiv:1611.02779"},{"key":"e_1_3_2_53_1","unstructured":"Dugas D (2023) The gpt-3 architecture on a napkin. URL: https:\/\/dugas.ch\/artificial_curiosity\/GPT_architecture.html ([Online; accessed 28-November-2023])."},{"key":"e_1_3_2_54_1","doi-asserted-by":"crossref","unstructured":"Ehsani K Gupta T Hendrix R et al. (2024) Spoc: imitating shortest paths in simulation enables effective navigation and manipulation in the real world. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition pp. 16238\u201316250. IEEE.","DOI":"10.1109\/CVPR52733.2024.01537"},{"key":"e_1_3_2_55_1","doi-asserted-by":"crossref","unstructured":"Engelbrecht HA Schiele G (2014) Transforming minecraft into a research platform. In: IEEE CCNC. IEEE.","DOI":"10.1109\/CCNC.2014.6866580"},{"key":"e_1_3_2_56_1","unstructured":"Fan L Wang G Jiang Y et al. (2022) MineDojo: building open-ended embodied agents with internet-scale knowledge. In: NeurIPS Datasets and Benchmarks Track."},{"key":"e_1_3_2_57_1","doi-asserted-by":"crossref","unstructured":"Fang HS Wang C Gou M et al. (2020) Graspnet-1billion: a large-scale benchmark for general object grasping. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition pp. 11444\u201311453. IEEE.","DOI":"10.1109\/CVPR42600.2020.01146"},{"key":"e_1_3_2_58_1","unstructured":"Farid A Veer S Majumdar A (2021) Task-driven out-of-distribution detection with statistical guarantees for robot learning. In: CoRL."},{"key":"e_1_3_2_59_1","article-title":"Failure prediction with statistical guarantees for vision-based robot control","author":"Farid A","year":"2022","unstructured":"Farid A, Snyder D, Ren AZ, et al. (2022a) Failure prediction with statistical guarantees for vision-based robot control. arXiv preprint arXiv:2202.05894.","journal-title":"arXiv preprint arXiv:2202.05894"},{"key":"e_1_3_2_60_1","unstructured":"Farid A Veer S Ivanovic B et al. (2022b) Task-relevant failure detection for trajectory predictors in autonomous vehicles. In: CoRL."},{"key":"e_1_3_2_61_1","article-title":"ChessGPT: bridging policy learning and language modeling","author":"Feng X","year":"2023","unstructured":"Feng X, Luo Y, Wang Z, et al. (2023) ChessGPT: bridging policy learning and language modeling. arXiv preprint arXiv:2306.09200.","journal-title":"arXiv preprint arXiv:2306.09200"},{"key":"e_1_3_2_62_1","article-title":"A touch, vision, and language dataset for multimodal alignment","author":"Fu L","year":"2024","unstructured":"Fu L, Datta G, Huang H, et al. (2024) A touch, vision, and language dataset for multimodal alignment. arXiv preprint arXiv:2402.13232.","journal-title":"arXiv preprint arXiv:2402.13232"},{"key":"e_1_3_2_63_1","doi-asserted-by":"crossref","unstructured":"Gadre SY Wortsman M Ilharco G et al. (2023) CoWs on pasture: baselines and benchmarks for language-driven zero-shot object navigation. In: CVPR pp. 23171\u201323181. IEEE.","DOI":"10.1109\/CVPR52729.2023.02219"},{"key":"e_1_3_2_64_1","article-title":"Red teaming language models to reduce harms: methods, scaling behaviors, and lessons learned","author":"Ganguli D","year":"2022","unstructured":"Ganguli D, Lovitt L, Kernion J, et al. (2022) Red teaming language models to reduce harms: methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858.","journal-title":"arXiv preprint arXiv:2209.07858"},{"key":"e_1_3_2_65_1","doi-asserted-by":"crossref","unstructured":"Gao R Si Z Chang YY et al. (2022) Objectfolder 2.0: a multisensory object dataset for sim2real transfer. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition pp. 10598\u201310608. IEEE.","DOI":"10.1109\/CVPR52688.2022.01034"},{"key":"e_1_3_2_66_1","doi-asserted-by":"crossref","unstructured":"Gao J Sarkar B Xia F et al. (2023) Physically grounded vision-language models for robotic manipulation. In: CoRL.","DOI":"10.1109\/ICRA57147.2024.10610090"},{"key":"e_1_3_2_67_1","doi-asserted-by":"publisher","DOI":"10.1146\/annurev-control-091420-084139"},{"key":"e_1_3_2_68_1","article-title":"Navigating to objects in the real world","author":"Gervet T","year":"2023","unstructured":"Gervet T, Chintala S, Batra D, et al. (2023) Navigating to objects in the real world. arXiv preprint arXiv:2212.00922.","journal-title":"arXiv preprint arXiv:2212.00922"},{"key":"e_1_3_2_69_1","unstructured":"Goodwin W Havoutis I Posner I (2022a) You only look at one: category-level object representations for pose estimation from a single example. In: CoRL."},{"key":"e_1_3_2_70_1","doi-asserted-by":"crossref","unstructured":"Goodwin W Vaze S Havoutis I et al. (2022b) Zero-shot category-level object pose estimation. In: ECCV.","DOI":"10.1007\/978-3-031-19842-7_30"},{"key":"e_1_3_2_71_1","unstructured":"Greenberg I Mannor S (2021) Detecting rewards deterioration in episodic reinforcement learning. In: ICML."},{"key":"e_1_3_2_72_1","doi-asserted-by":"crossref","unstructured":"Guzhov A Raue F Hees J et al. (2022) AudioCLIP: extending clip to image text and audio. In: ICASSP pp. 976\u2013980. IEEE.","DOI":"10.1109\/ICASSP43922.2022.9747631"},{"key":"e_1_3_2_73_1","unstructured":"Ha H Song S (2022) Semantic abstraction: open-world 3D scene understanding from 2D vision-language models. In: CoRL."},{"key":"e_1_3_2_74_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2022.3215150"},{"key":"e_1_3_2_75_1","article-title":"Deep residual learning for image recognition","author":"He K","year":"2015","unstructured":"He K, Zhang X, Ren S, et al. (2015) Deep residual learning for image recognition. arXiv preprint arXiv:1512.03385.","journal-title":"arXiv preprint arXiv:1512.03385"},{"key":"e_1_3_2_76_1","doi-asserted-by":"publisher","unstructured":"He K Gkioxari G Doll\u00e1r P et al. (2017) Mask R-CNN. In: ICCV pp. 2980\u20132988. DOI: 10.1109\/ICCV.2017.322.","DOI":"10.1109\/ICCV.2017.322"},{"key":"e_1_3_2_77_1","doi-asserted-by":"crossref","unstructured":"He K Chen X Xie S et al. (2022) Masked autoencoders are scalable vision learners. In: CVPR pp. 16000\u201316009.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"e_1_3_2_78_1","article-title":"AnnoLLM: making large language models to be better crowdsourced annotators","author":"He X","year":"2023","unstructured":"He X, Lin Z, Gong Y, et al. (2023) AnnoLLM: making large language models to be better crowdsourced annotators. arXiv preprint arXiv:2303.16854.","journal-title":"arXiv preprint arXiv:2303.16854"},{"key":"e_1_3_2_79_1","unstructured":"Ho J Jain A Abbeel P (2020) Denoising diffusion probabilistic models. In: NeurIPS."},{"key":"e_1_3_2_80_1","article-title":"3D-LLM: injecting the 3D world into large language models","author":"Hong Y","year":"2023","unstructured":"Hong Y, Zhen H, Chen P, et al. (2023) 3D-LLM: injecting the 3D world into large language models. arXiv preprint arXiv:2307.12981.","journal-title":"arXiv preprint arXiv:2307.12981"},{"key":"e_1_3_2_81_1","article-title":"The safety filter: a unified view of safety-critical control in autonomous systems","author":"Hsu KC","year":"2023","unstructured":"Hsu KC, Hu H, Fisac JF (2023a) The safety filter: a unified view of safety-critical control in autonomous systems. arXiv preprint arXiv:2309.05837.","journal-title":"arXiv preprint arXiv:2309.05837"},{"key":"e_1_3_2_82_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2022.103811"},{"key":"e_1_3_2_83_1","article-title":"GAIA-1: a generative world model for autonomous driving","author":"Hu A","year":"2023","unstructured":"Hu A, Russell L, Yeo H, et al. (2023a) GAIA-1: a generative world model for autonomous driving. arXiv preprint arXiv:2309.17080.","journal-title":"arXiv preprint arXiv:2309.17080"},{"key":"e_1_3_2_84_1","article-title":"Toward general-purpose robots via foundation models: a survey and meta-analysis","author":"Hu Y","year":"2023","unstructured":"Hu Y, Xie Q, Jain V, et al. (2023b) Toward general-purpose robots via foundation models: a survey and meta-analysis. arXiv preprint arXiv:2312.08782.","journal-title":"arXiv preprint arXiv:2312.08782"},{"key":"e_1_3_2_85_1","article-title":"Semantic anything in 3d Gaussians","author":"Hu X","year":"2024","unstructured":"Hu X, Wang Y, Fan L, et al. (2024) Semantic anything in 3d Gaussians. arXiv preprint arXiv:2401.17857.","journal-title":"arXiv preprint arXiv:2401.17857"},{"key":"e_1_3_2_86_1","unstructured":"Huang J Xie S Sun J et al. (2021) Learning a decision module by imitating driver\u2019s control behaviors. In: CoRL pp. 1\u201310. PMLR."},{"key":"e_1_3_2_87_1","unstructured":"Huang W Abbeel P Pathak D et al. (2022a) Language models as zero-shot planners: extracting actionable knowledge for embodied agents. In: ICML."},{"key":"e_1_3_2_88_1","article-title":"Inner monologue: embodied reasoning through planning with language models","author":"Huang W","year":"2022","unstructured":"Huang W, Xia F, Xiao T, et al. (2022b) Inner monologue: embodied reasoning through planning with language models. arXiv preprint arXiv:2207.05608.","journal-title":"arXiv preprint arXiv:2207.05608"},{"key":"e_1_3_2_89_1","article-title":"Audio visual language maps for robot navigation","author":"Huang C","year":"2023","unstructured":"Huang C, Mees O, Zeng A, et al. (2023a) Audio visual language maps for robot navigation. arXiv preprint arXiv:2303.07522.","journal-title":"arXiv preprint arXiv:2303.07522"},{"key":"e_1_3_2_90_1","doi-asserted-by":"crossref","unstructured":"Huang C Mees O Zeng A et al. (2023b) Visual language maps for robot navigation. In: 2023 IEEE International Conference on Robotics and Automation (ICRA) pp. 10608\u201310615. IEEE.","DOI":"10.1109\/ICRA48891.2023.10160969"},{"key":"e_1_3_2_91_1","unstructured":"Huang W Wang C Zhang R et al. (2023c) VoxPoser: composable 3D value maps for robotic manipulation with language models. In: CoRL."},{"key":"e_1_3_2_92_1","unstructured":"Janner M Du Y Tenenbaum J et al. (2022) Planning with diffusion for flexible behavior synthesis. In: ICML."},{"key":"e_1_3_2_93_1","article-title":"Chain-of-thought predictive control","author":"Jia Z","year":"2023","unstructured":"Jia Z, Liu F, Thumuluri V, et al. (2023) Chain-of-thought predictive control. arXiv preprint arXiv:2304.00776.","journal-title":"arXiv preprint arXiv:2304.00776"},{"key":"e_1_3_2_94_1","unstructured":"Jiang Y Gupta A Zhang Z et al. (2023) VIMA: general robot manipulation with multimodal prompts. In: ICML."},{"key":"e_1_3_2_95_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2023.3272516"},{"key":"e_1_3_2_96_1","doi-asserted-by":"crossref","unstructured":"Karamcheti S Nair S Chen AS et al. (2023) Language-driven representation learning for robotics. In: RSS.","DOI":"10.15607\/RSS.2023.XIX.032"},{"key":"e_1_3_2_97_1","doi-asserted-by":"publisher","DOI":"10.1145\/3592433"},{"key":"e_1_3_2_98_1","doi-asserted-by":"crossref","unstructured":"Kerr J Kim CM Goldberg K et al. (2023) LERF: language embedded radiance fields. In: ICCV pp. 19729\u201319739.","DOI":"10.1109\/ICCV51070.2023.01807"},{"key":"e_1_3_2_99_1","doi-asserted-by":"publisher","DOI":"10.1145\/3505244"},{"key":"e_1_3_2_100_1","doi-asserted-by":"crossref","unstructured":"Kirillov A Mintun E Ravi N et al. (2023) Segment anything. In: ICCV pp. 4015\u20134026.","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"e_1_3_2_101_1","unstructured":"Kobayashi S Matsumoto E Sitzmann V (2022) Decomposing NeRF for editing via feature field distillation. In: NeurIPS."},{"key":"e_1_3_2_102_1","article-title":"Rma: rapid motor adaptation for legged robots","author":"Kumar A","year":"2021","unstructured":"Kumar A, Fu Z, Pathak D, et al. (2021) Rma: rapid motor adaptation for legged robots. arXiv preprint arXiv:2107.04034.","journal-title":"arXiv preprint arXiv:2107.04034"},{"key":"e_1_3_2_103_1","article-title":"Collision avoidance testing of the waymo automated driving system","author":"Kusano KD","year":"2022","unstructured":"Kusano KD, Beatty K, Schnelle S, et al. (2022) Collision avoidance testing of the waymo automated driving system. arXiv preprint arXiv:2212.08148.","journal-title":"arXiv preprint arXiv:2212.08148"},{"key":"e_1_3_2_104_1","unstructured":"Kwon M Xie SM Bullard K et al. (2023) Reward design with language models. In: ICLR."},{"key":"e_1_3_2_105_1","unstructured":"Levesque H Davis E Morgenstern L (2012) The Winograd schema challenge. In: KR."},{"key":"e_1_3_2_106_1","unstructured":"Li C Xia F Mart\u00edn-Mart\u00edn R et al. (2021) iGibson 2.0: object-centric simulation for robot learning of everyday household tasks. In: CoRL."},{"key":"e_1_3_2_107_1","unstructured":"Li B Weinberger KQ Belongie S et al. (2022a) Language-driven semantic segmentation. In: ICLR."},{"key":"e_1_3_2_108_1","unstructured":"Li C Zhang R Wong J et al. (2022b) BEHAVIOR-1K: a benchmark for embodied AI with 1 000 everyday activities and realistic simulation. In: CoRL."},{"key":"e_1_3_2_109_1","unstructured":"Li J Li D Xiong C et al. (2022c) BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation. In: ICML."},{"key":"e_1_3_2_110_1","doi-asserted-by":"crossref","unstructured":"Li LH Zhang P Zhang H et al. (2022d) Grounded language-image pre-training. In: CVPR pp. 10965\u201310975.","DOI":"10.1109\/CVPR52688.2022.01069"},{"key":"e_1_3_2_111_1","doi-asserted-by":"crossref","unstructured":"Li Y Fan H Hu R et al. (2023) Scaling language-image pre-training via masking. In: CVPR.","DOI":"10.1109\/CVPR52729.2023.02240"},{"key":"e_1_3_2_112_1","doi-asserted-by":"crossref","unstructured":"Liang J Huang W Xia F et al. (2023) Code as Policies: language model programs for embodied control. In: ICRA pp. 9493\u20139500. IEEE.","DOI":"10.1109\/ICRA48891.2023.10160591"},{"key":"e_1_3_2_113_1","article-title":"Clip-gs: clip-informed gaussian splatting for real-time and view-consistent 3d semantic understanding","author":"Liao G","year":"2024","unstructured":"Liao G, Li J, Bao Z, et al. (2024) Clip-gs: clip-informed gaussian splatting for real-time and view-consistent 3d semantic understanding. arXiv preprint arXiv:2404.14249.","journal-title":"arXiv preprint arXiv:2404.14249"},{"key":"e_1_3_2_114_1","article-title":"Learning to model the world with language","author":"Lin J","year":"2023","unstructured":"Lin J, Du Y, Watkins O, et al. (2023a) Learning to model the world with language. arXiv preprint arXiv:2308.01399.","journal-title":"arXiv preprint arXiv:2308.01399"},{"key":"e_1_3_2_115_1","article-title":"AWQ: activation-aware weight quantization for llm compression and acceleration","author":"Lin J","year":"2023","unstructured":"Lin J, Tang J, Tang H, et al. (2023b) AWQ: activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978.","journal-title":"arXiv preprint arXiv:2306.00978"},{"key":"e_1_3_2_116_1","doi-asserted-by":"crossref","unstructured":"Lin K Agia C Migimatsu T et al. (2023c) Text2Motion: from natural language instructions to feasible plans. In: Special Issue: Large Language Models in Robotics. Autonomous Robots.","DOI":"10.1007\/s10514-023-10131-7"},{"key":"e_1_3_2_117_1","unstructured":"Liu PJ Saleh M Pot E et al. (2018) Generating Wikipedia by summarizing long sequences. In: ICLR."},{"key":"e_1_3_2_118_1","article-title":"RoBERTa: a robustly optimized BERT pretraining approach","author":"Liu Y","year":"2019","unstructured":"Liu Y, Ott M, Goyal N, et al. (2019) RoBERTa: a robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.","journal-title":"arXiv preprint arXiv:1907.11692"},{"key":"e_1_3_2_119_1","first-page":"21736","article-title":"Partslip: low-shot part segmentation for 3d point clouds via pretrained image-language models","author":"Liu M","year":"2023","unstructured":"Liu M, Zhu Y, Cai H, et al. (2023a) Partslip: low-shot part segmentation for 3d point clouds via pretrained image-language models. CVPR: 21736\u201321746.","journal-title":"CVPR"},{"key":"e_1_3_2_120_1","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_3_2_121_1","article-title":"Grounding DINO: marrying DINO with grounded pre-training for open-set object detection","author":"Liu S","year":"2023","unstructured":"Liu S, Zeng Z, Ren T, et al. (2023c) Grounding DINO: marrying DINO with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499.","journal-title":"arXiv preprint arXiv:2303.05499"},{"key":"e_1_3_2_122_1","doi-asserted-by":"crossref","unstructured":"Liu W Du Y Hermans T et al. (2023d) StructDiffusion: language-guided creation of physically-valid structures using unseen objects. In: RSS.","DOI":"10.15607\/RSS.2023.XIX.031"},{"key":"e_1_3_2_123_1","doi-asserted-by":"crossref","unstructured":"Luo R Zhao S Kuck J et al. (2022) Sample-efficient safety assurances using conformal prediction. In: WAFR.","DOI":"10.1007\/978-3-031-21090-7_10"},{"key":"e_1_3_2_124_1","doi-asserted-by":"crossref","unstructured":"Lynch C Sermanet P (2021) Language conditioned imitation learning over unstructured data. In: Robotics: Science and Systems.","DOI":"10.15607\/RSS.2021.XVII.047"},{"key":"e_1_3_2_125_1","unstructured":"Lynch C Khansari M Xiao T et al. (2020) Learning latent plans from play. In: CoRL pp. 1113\u20131132. PMLR."},{"key":"e_1_3_2_126_1","doi-asserted-by":"crossref","unstructured":"Ma S Vemprala S Wang W et al. (2022) COMPASS: contrastive multimodal pretraining for autonomous systems. In: IROS pp. 1000\u20131007. IEEE.","DOI":"10.1109\/IROS47612.2022.9982241"},{"key":"e_1_3_2_127_1","unstructured":"Ma YJ Kumar V Zhang A et al. (2023a) LIV: language-image representations and rewards for robotic control. In: ICML."},{"key":"e_1_3_2_128_1","unstructured":"Ma YJ Sodhani S Jayaraman D et al. (2023b) VIP: towards universal visual reward and representation via value-implicit pre-training. In:ICLR."},{"key":"e_1_3_2_129_1","unstructured":"Mahmoudieh P Pathak D Darrell T (2022) Zero-shot reward specification via grounded natural language. In ICML pp. 14743\u201314752. PMLR."},{"key":"e_1_3_2_130_1","doi-asserted-by":"crossref","unstructured":"Majumdar A Shrivastava A Lee S et al. (2020) Improving vision-and-language navigation with image-text pairs from the web. In: ECCV pp. 259\u2013274. Springer.","DOI":"10.1007\/978-3-030-58539-6_16"},{"key":"e_1_3_2_131_1","article-title":"CACTI: a framework for scalable multi-task multi-scene visual imitation learning","author":"Mandi Z","year":"2022","unstructured":"Mandi Z, Bharadhwaj H, Moens V, et al. (2022) CACTI: a framework for scalable multi-task multi-scene visual imitation learning. arXiv preprint arXiv:2212.05711.","journal-title":"arXiv preprint arXiv:2212.05711"},{"key":"e_1_3_2_132_1","unstructured":"Margolis GB Agrawal P (2023) Walk these ways: tuning robot control for generalization with multiplicity of behavior. In: Conference on Robot Learning pp. 22\u201331. PMLR."},{"key":"e_1_3_2_133_1","doi-asserted-by":"crossref","unstructured":"Mees O Borja-Diaz J Burgard W (2023) Grounding language with visual affordances over unstructured data. In: ICRA pp. 11576\u201311582. IEEE.","DOI":"10.1109\/ICRA48891.2023.10160396"},{"key":"e_1_3_2_134_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503250"},{"key":"e_1_3_2_135_1","doi-asserted-by":"crossref","unstructured":"Minderer M Gritsenko A Stone A et al. (2022) Simple open-vocabulary object detection with vision transformers. In: ECCV pp. 728\u2013755. Springer.","DOI":"10.1007\/978-3-031-20080-9_42"},{"key":"e_1_3_2_136_1","article-title":"Large language models as general pattern machines","author":"Mirchandani S","year":"2023","unstructured":"Mirchandani S, Xia F, Florence P, et al. (2023) Large language models as general pattern machines. arXiv preprint arXiv:2307.04721.","journal-title":"arXiv preprint arXiv:2307.04721"},{"key":"e_1_3_2_137_1","article-title":"EmbodiedGPT: vision-language pre-training via embodied chain of thought","author":"Mu Y","year":"2023","unstructured":"Mu Y, Zhang Q, Hu M, et al. (2023) EmbodiedGPT: vision-language pre-training via embodied chain of thought. arXiv preprint arXiv:2305.15021.","journal-title":"arXiv preprint arXiv:2305.15021"},{"key":"e_1_3_2_138_1","unstructured":"Nair S Mitchell E Chen K et al. (2022a) Learning language-conditioned robot behavior from offline data and crowd-sourced annotation. In: CoRL pp. 1303\u20131315. PMLR."},{"key":"e_1_3_2_139_1","article-title":"R3M: a universal visual representation for robot manipulation","author":"Nair S","year":"2022","unstructured":"Nair S, Rajeswaran A, Kumar V, et al. (2022b) R3M: a universal visual representation for robot manipulation. arXiv preprint arXiv:2203.12601.","journal-title":"arXiv preprint arXiv:2203.12601"},{"key":"e_1_3_2_140_1","article-title":"Representation learning with contrastive predictive coding","author":"Oord A","year":"2018","unstructured":"Oord A, Li Y, Vinyals O (2018) Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748.","journal-title":"arXiv preprint arXiv:1807.03748"},{"key":"e_1_3_2_141_1","article-title":"GPT-4 technical report","author":"OpenAI","year":"2023","unstructured":"OpenAI (2023) GPT-4 technical report. arXiv preprint arXiv:2303.08774.","journal-title":"arXiv preprint arXiv:2303.08774"},{"key":"e_1_3_2_142_1","article-title":"DINOv2: learning robust visual features without supervision","author":"Oquab M","year":"2023","unstructured":"Oquab M, Darcet T, Moutakanni T, et al. (2023) DINOv2: learning robust visual features without supervision. arXiv preprint arXiv:2304.07193.","journal-title":"arXiv preprint arXiv:2304.07193"},{"key":"e_1_3_2_143_1","article-title":"Open X-Embodiment: robotic learning datasets and RT-X models","author":"Padalkar A","year":"2023","unstructured":"Padalkar A, Pooley A, Jain A, et al. (2023) Open X-Embodiment: robotic learning datasets and RT-X models. arXiv preprint arXiv:2310.08864.","journal-title":"arXiv preprint arXiv:2310.08864"},{"key":"e_1_3_2_144_1","unstructured":"Palo ND Byravan A Hasenclever L et al. (2023) Towards a unified agent with foundation models. In: Workshop on Reincarnating Reinforcement Learning at ICLR 2023."},{"key":"e_1_3_2_145_1","doi-asserted-by":"crossref","unstructured":"Park JS O\u2019Brien JC Cai CJ et al. (2023a) Generative Agents: interactive simulacra of human behavior. In: ACM Symposium on User Interface Software and Technology. ACM.","DOI":"10.1145\/3586183.3606763"},{"key":"e_1_3_2_146_1","article-title":"Representation reliability and its impact on downstream tasks","author":"Park YJ","year":"2023","unstructured":"Park YJ, Wang H, Ardeshir S, et al. (2023b) Representation reliability and its impact on downstream tasks. arXiv preprint arXiv:2306.00206.","journal-title":"arXiv preprint arXiv:2306.00206"},{"key":"e_1_3_2_147_1","doi-asserted-by":"crossref","unstructured":"Perez E Huang S Song F et al. (2022) Red teaming language models with language models. In: EMNLP.","DOI":"10.18653\/v1\/2022.emnlp-main.225"},{"key":"e_1_3_2_148_1","doi-asserted-by":"crossref","unstructured":"Puig X Ra K Boben M et al. (2018) VirtualHome: simulating household activities via programs. In: CVPR pp. 8494\u20138502.","DOI":"10.1109\/CVPR.2018.00886"},{"key":"e_1_3_2_149_1","article-title":"Habitat 3.0: a co-habitat for humans, avatars and robots","author":"Puig X","year":"2023","unstructured":"Puig X, Undersander E, Szot A, et al. (2023) Habitat 3.0: a co-habitat for humans, avatars and robots. arXiv preprint arXiv:2310.13724.","journal-title":"arXiv preprint arXiv:2310.13724"},{"key":"e_1_3_2_150_1","article-title":"Langsplat: 3d language Gaussian splatting","author":"Qin M","year":"2023","unstructured":"Qin M, Li W, Zhou J, et al. (2023) Langsplat: 3d language Gaussian splatting. arXiv preprint arXiv:2312.16084.","journal-title":"arXiv preprint arXiv:2312.16084"},{"key":"e_1_3_2_151_1","doi-asserted-by":"publisher","DOI":"10.1109\/JBHI.2023.3316750"},{"key":"e_1_3_2_152_1","unstructured":"Radford A Narasimhan K Salimans T et al. (2018) Improving language understanding by generative pre-training. https:\/\/openai.com\/research\/language-unsupervised"},{"key":"e_1_3_2_153_1","volume-title":"Language Models Are Unsupervised Multitask Learners","author":"Radford A","year":"2019","unstructured":"Radford A, Wu J, Child R, et al. (2019) Language Models Are Unsupervised Multitask Learners. OpenAI Blog."},{"key":"e_1_3_2_154_1","unstructured":"Radford A Kim JW Hallacy C et al. (2021a) Learning transferable visual models from natural language supervision. In: Meila M Zhang T (eds) ICML Proceedings of Machine Learning Research Vol. 139 8748\u20138763. PMLR."},{"key":"e_1_3_2_155_1","unstructured":"Radford A Kim JW Hallacy C et al. (2021b) Learning transferable visual models from natural language supervision. In: ICML pp. 8748\u20138763."},{"key":"e_1_3_2_156_1","unstructured":"Radosavovic I Xiao T James S et al. (2023) Real-world robot learning with masked visual pre-training. In: CoRL pp. 416\u2013426. PMLR."},{"issue":"1","key":"e_1_3_2_157_1","first-page":"5485","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel C","year":"2020","unstructured":"Raffel C, Shazeer N, Roberts A, et al. (2020) Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21(1): 5485\u20135551.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_158_1","unstructured":"Ramesh A Pavlov M Goh G et al. (2021a) Zero-shot text-to-image generation. In: ICML pp. 8821\u20138831. PMLR."},{"key":"e_1_3_2_159_1","unstructured":"Ramesh A Pavlov M Goh G et al. (2021b) Zero-shot text-to-image generation. In: Meila M Zhang T (eds) Proceedings of the 38th International Conference on Machine Learning Proceedings of Machine Learning Research Vol. 139 pp. 8821\u20138831. PMLR."},{"key":"e_1_3_2_160_1","article-title":"Hierarchical text-conditional image generation with CLIP latents","author":"Ramesh A","year":"2022","unstructured":"Ramesh A, Dhariwal P, Nichol A, et al. (2022) Hierarchical text-conditional image generation with CLIP latents. arXiv preprint arXiv:2204.06125.","journal-title":"arXiv preprint arXiv:2204.06125"},{"key":"e_1_3_2_161_1","doi-asserted-by":"crossref","unstructured":"Ranftl R Bochkovskiy A Koltun V (2021) Vision transformers for dense prediction. In: ICCV pp. 12179\u201312188. PMLR.","DOI":"10.1109\/ICCV48922.2021.01196"},{"key":"e_1_3_2_162_1","unstructured":"Rashid A Sharma S Kim CM et al. (2023) Language embedded radiance fields for zero-shot task-oriented grasping. In: 7th Annual Conference on Robot Learning (CoRL) pp. 178\u2013200. PMLR."},{"key":"e_1_3_2_163_1","article-title":"A generalist agent","author":"Reed S","year":"2022","unstructured":"Reed S, Zolna K, Parisotto E, et al. (2022) A generalist agent. arXiv preprint arXiv:2205.06175.","journal-title":"arXiv preprint arXiv:2205.06175"},{"key":"e_1_3_2_164_1","article-title":"Can Wikipedia help offline reinforcement learning?","author":"Reid M","year":"2022","unstructured":"Reid M, Yamada Y, Gu SS (2022) Can Wikipedia help offline reinforcement learning? arXiv preprint arXiv:2201.12122.","journal-title":"arXiv preprint arXiv:2201.12122"},{"key":"e_1_3_2_165_1","unstructured":"Ren AZ Dixit A Bodrova A et al. (2023) Robots that ask for help: uncertainty alignment for large language model planners. In: CoRL."},{"key":"e_1_3_2_166_1","doi-asserted-by":"crossref","unstructured":"Rombach R Blattmann A Lorenz D et al. (2022) High-resolution image synthesis with latent diffusion models. In: CVPR.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_2_167_1","doi-asserted-by":"crossref","unstructured":"Savva M Kadian A Maksymets O et al. (2019) Habitat: a platform for embodied AI research. In: ICCV.","DOI":"10.1109\/ICCV.2019.00943"},{"key":"e_1_3_2_168_1","doi-asserted-by":"crossref","unstructured":"Sennrich R Haddow B Birch A (2016) Neural machine translation of rare words with subword units. In: ACL.","DOI":"10.18653\/v1\/P16-1162"},{"key":"e_1_3_2_169_1","article-title":"Robovqa: multimodal long-horizon reasoning for robotics","author":"Sermanet P","year":"2023","unstructured":"Sermanet P, Ding T, Zhao J, et al. (2023) Robovqa: multimodal long-horizon reasoning for robotics. arXiv:2311.00899.","journal-title":"arXiv:2311.00899"},{"key":"e_1_3_2_170_1","unstructured":"Shafiullah NMM Paxton C Pinto L et al. (2023) CLIP-fields: weakly supervised semantic fields for robotic memory. In: RSS."},{"key":"e_1_3_2_171_1","unstructured":"Shah R Kumar V (2021) RRL: resnet as representation for reinforcement learning. In: ICML."},{"key":"e_1_3_2_172_1","doi-asserted-by":"crossref","unstructured":"Shah S Dey D Lovett C et al. (2017) AirSim: high-fidelity visual and physical simulation for autonomous vehicles. In: Field and Service Robotics.","DOI":"10.1007\/978-3-319-67361-5_40"},{"key":"e_1_3_2_173_1","unstructured":"Shah D Osi\u0144ski B Levine S et al. (2023a) LM-Nav: robotic navigation with large pre-trained models of language vision and action. In: CoRL pp. 492\u2013504. PMLR."},{"key":"e_1_3_2_174_1","unstructured":"Shah D Sridhar A Dashora N et al. (2023b) ViNT: a foundation model for visual navigation. In: CoRL."},{"key":"e_1_3_2_175_1","unstructured":"Shah R Mart\u00edn-Mart\u00edn R Zhu Y (2023c) MUTEX: learning unified policies from multimodal task specifications. In: CoRL."},{"key":"e_1_3_2_176_1","article-title":"Anything-3D: towards single-view anything reconstruction in the wild","author":"Shen Q","year":"2023","unstructured":"Shen Q, Yang X, Wang X (2023a) Anything-3D: towards single-view anything reconstruction in the wild. arXiv preprint arXiv:2304.10261.","journal-title":"arXiv preprint arXiv:2304.10261"},{"key":"e_1_3_2_177_1","unstructured":"Shen W Yang G Yu A et al. (2023b) Distilled feature fields enable few-shot language-guided manipulation. In: 7th Annual Conference on Robot Learning (CoRL)."},{"key":"e_1_3_2_178_1","unstructured":"Shen W Yang G Yu A et al. (2023c) Distilled feature fields enable few-shot manipulation. In: CoRL."},{"key":"e_1_3_2_179_1","unstructured":"Shorinwa O Tucker J Smith A et al. (2024) Splat-mover: multi-stage open-vocabulary robotic manipulation via editable Gaussian splatting."},{"key":"e_1_3_2_180_1","unstructured":"Shridhar M Manuelli L Fox D (2022) CLIPort: what and where pathways for robotic manipulation. In: CoRL pp. 894\u2013906. PMLR."},{"key":"e_1_3_2_181_1","unstructured":"Shridhar M Manuelli L Fox D (2023) Perceiver-Actor: a multi-task transformer for robotic manipulation. In: CoRL pp. 785\u2013799. PMLR."},{"key":"e_1_3_2_182_1","doi-asserted-by":"crossref","unstructured":"Simeonov A Du Y Tagliasacchi A et al. (2022) Neural Descriptor Fields: SE(3)-equivariant object representations for manipulation. In: ICRA.","DOI":"10.1109\/ICRA46639.2022.9812146"},{"key":"e_1_3_2_183_1","doi-asserted-by":"crossref","unstructured":"Singh I Blukis V Mousavian A et al. (2023) ProgPrompt: generating situated robot task plans using large language models. In: ICRA pp. 11523\u201311530. IEEE.","DOI":"10.1109\/ICRA48891.2023.10161317"},{"key":"e_1_3_2_184_1","article-title":"A system-level view on out-of-distribution data in robotics","author":"Sinha R","year":"2022","unstructured":"Sinha R, Sharma A, Banerjee S, et al. (2022) A system-level view on out-of-distribution data in robotics. arXiv preprint arXiv:2212.14020.","journal-title":"arXiv preprint arXiv:2212.14020"},{"key":"e_1_3_2_185_1","unstructured":"Sohl-Dickstein J Weiss E Maheswaranathan N et al. (2015) Deep unsupervised learning using nonequilibrium thermodynamics. In: ICML."},{"key":"e_1_3_2_186_1","unstructured":"Sohn K (2016) Improved deep metric learning with multi-class n-pair loss objective. In: NeurIPS."},{"key":"e_1_3_2_187_1","unstructured":"Song Y Ermon S (2019) Generative modeling by estimating gradients of the data distribution. In: NeurIPS."},{"key":"e_1_3_2_188_1","article-title":"Open-world object manipulation using pre-trained vision-language model","author":"Stone A","year":"2023","unstructured":"Stone A, Xiao T, Lu Y, et al. (2023) Open-world object manipulation using pre-trained vision-language model. arXiv preprint arXiv:2303.00905.","journal-title":"arXiv preprint arXiv:2303.00905"},{"key":"e_1_3_2_189_1","unstructured":"Sun J Sun H Han T et al. (2021a) Neuro-symbolic program search for autonomous driving decision module design. In: CoRL pp. 21\u201330. PMLR."},{"key":"e_1_3_2_190_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2021.3061397"},{"key":"e_1_3_2_191_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2022.3150855"},{"key":"e_1_3_2_192_1","doi-asserted-by":"crossref","unstructured":"Sun J Kousik S Fridovich-Keil D et al. (2022b) Self-supervised traffic advisors: distributed multi-view traffic prediction for smart cities. In: IEEE ITSC.","DOI":"10.1109\/ITSC55140.2022.9922340"},{"key":"e_1_3_2_193_1","unstructured":"Sun M Yan W Abbeel P et al. (2022c) Quantifying uncertainty in foundation models via ensembles. In: NeurIPS Workshop on Robustness in Sequence Modeling."},{"key":"e_1_3_2_194_1","unstructured":"Sun J Jiang Y Qiu J et al. (2023a) Conformal prediction for uncertainty-aware planning with diffusion dynamics model. In: NeurIPS."},{"key":"e_1_3_2_195_1","article-title":"Connected autonomous vehicle motion planning with video predictions from smart, self-supervised infrastructure","author":"Sun J","year":"2023","unstructured":"Sun J, Kousik S, Fridovich-Keil D, et al. (2023b) Connected autonomous vehicle motion planning with video predictions from smart, self-supervised infrastructure. arXiv preprint arXiv:2309.07504.","journal-title":"arXiv preprint arXiv:2309.07504"},{"key":"e_1_3_2_196_1","article-title":"Aria-NeRF: multimodal egocentric view synthesis","author":"Sun J","year":"2023","unstructured":"Sun J, Qiu J, Zheng C, et al. (2023c) Aria-NeRF: multimodal egocentric view synthesis. arXiv preprint arXiv:2311.06455.","journal-title":"arXiv preprint arXiv:2311.06455"},{"key":"e_1_3_2_197_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2023.3293308"},{"key":"e_1_3_2_198_1","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.10298866"},{"key":"e_1_3_2_199_1","unstructured":"Sun Y Ma S Madaan R et al. (2023f) SMART: self-supervised multi-task pretraining with control transformers. In: ICLR."},{"key":"e_1_3_2_200_1","unstructured":"Szot A Clegg A Undersander E et al. (2021) Habitat 2.0: training home assistants to rearrange their habitat. In: NeurIPS."},{"key":"e_1_3_2_201_1","unstructured":"Tarasov D Kurenkov V Kolesnikov S (2022) Prompts and pre-trained language models for offline reinforcement learning. In: ICLR Workshop on Generalizable Policy Learning in Physical World."},{"key":"e_1_3_2_202_1","article-title":"Open-ended learning leads to generally capable agents","author":"Team OEL","year":"2021","unstructured":"Team OEL, Stooke A, Mahajan A, et al. (2021) Open-ended learning leads to generally capable agents. arXiv preprint arXiv:2107.12808.","journal-title":"arXiv preprint arXiv:2107.12808"},{"key":"e_1_3_2_203_1","unstructured":"Team AA Bauer J Baumli K et al. (2023) Human-timescale adaptation in an open-ended task space. In: ICML."},{"key":"e_1_3_2_204_1","unstructured":"Thickstun J (2023) The transformer model in equations. URL: https:\/\/johnthickstun.com\/docs\/transformers.pdf. [Online; accessed 28-November-2023]."},{"key":"e_1_3_2_205_1","article-title":"Mass-producing failures of multimodal systems with language models","author":"Tong S","year":"2023","unstructured":"Tong S, Jones E, Steinhardt J (2023) Mass-producing failures of multimodal systems with language models. arXiv preprint arXiv:2306.12105.","journal-title":"arXiv preprint arXiv:2306.12105"},{"key":"e_1_3_2_206_1","article-title":"LLaMA: open and efficient foundation language models","author":"Touvron H","year":"2023","unstructured":"Touvron H, Lavril T, Izacard G, et al. (2023a) LLaMA: open and efficient foundation language models. arXiv preprint arXiv:2302.13971.","journal-title":"arXiv preprint arXiv:2302.13971"},{"key":"e_1_3_2_207_1","article-title":"LLaMA: open and efficient foundation language models","author":"Touvron H","year":"2023","unstructured":"Touvron H, Lavril T, Izacard G, et al. (2023b) LLaMA: open and efficient foundation language models. arXiv preprint arXiv:2302.13971.","journal-title":"arXiv preprint arXiv:2302.13971"},{"key":"e_1_3_2_208_1","doi-asserted-by":"crossref","unstructured":"Tschernezki V Laina I Larlus D et al. (2022) Neural Feature Fusion Fields: 3D distillation of self-supervised 2D image representations. In: 3DV pp. 443\u2013453. IEEE.","DOI":"10.1109\/3DV57658.2022.00056"},{"key":"e_1_3_2_209_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.88573"},{"key":"e_1_3_2_210_1","unstructured":"Vaswani A Shazeer N Parmar N et al. (2017) Attention is all you need. In: NeurIPS."},{"key":"e_1_3_2_211_1","unstructured":"Vemprala S Bonatti R Bucker A et al. (2023) ChatGPT for robotics: design principles and model abilities. Technical Report MSR-TR-2023-8 Microsoft."},{"key":"e_1_3_2_212_1","unstructured":"Villegas R Babaeizadeh M Kindermans PJ et al. (2023) Phenaki: variable length video generation from open domain textual description. In: ICLR."},{"key":"e_1_3_2_213_1","article-title":"Guarantees on robot system performance using stochastic simulation rollouts","author":"Vincent JA","year":"2023","unstructured":"Vincent JA, Feldman AO, Schwager M (2023) Guarantees on robot system performance using stochastic simulation rollouts. arXiv preprint arXiv:2309.10874.","journal-title":"arXiv preprint arXiv:2309.10874"},{"key":"e_1_3_2_214_1","article-title":"How generalizable is my behavior cloning policy? a statistical approach to trustworthy performance evaluation","author":"Vincent JA","year":"2024","unstructured":"Vincent JA, Nishimura H, Itkina M, et al. (2024) How generalizable is my behavior cloning policy? a statistical approach to trustworthy performance evaluation. arXiv preprint arXiv:2405.05439.","journal-title":"arXiv preprint arXiv:2405.05439"},{"key":"e_1_3_2_215_1","article-title":"GLUE: a multi-task benchmark and analysis platform for natural language understanding","author":"Wang A","year":"2018","unstructured":"Wang A, Singh A, Michael J, et al. (2018) GLUE: a multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461.","journal-title":"arXiv preprint arXiv:1804.07461"},{"key":"e_1_3_2_216_1","doi-asserted-by":"crossref","unstructured":"Wang W Zhu D Wang X et al. (2020) TartanAir: a dataset to push the limits of visual SLAM. In: IROS.","DOI":"10.1109\/IROS45743.2020.9341801"},{"key":"e_1_3_2_217_1","doi-asserted-by":"crossref","unstructured":"Wang C Chai M He M et al. (2022) CLIP-NeRF: text-and-image driven manipulation of neural radiance fields. In: CVPR pp. 3835\u20133844.","DOI":"10.1109\/CVPR52688.2022.00381"},{"key":"e_1_3_2_218_1","unstructured":"Wang C Fan L Sun J et al. (2023a) MimicPlay: long-horizon imitation learning by watching human play. In: CoRL."},{"key":"e_1_3_2_219_1","article-title":"Voyager: an open-ended embodied agent with large language models","author":"Wang G","year":"2023","unstructured":"Wang G, Xie Y, Jiang Y, et al. (2023b) Voyager: an open-ended embodied agent with large language models. arXiv preprint arXiv: Arxiv-2305.16291.","journal-title":"arXiv preprint arXiv: Arxiv-2305.16291"},{"key":"e_1_3_2_220_1","doi-asserted-by":"crossref","unstructured":"Wang S Saharia C Montgomery C et al. (2023c) Imagen editor and editbench: advancing and evaluating text-guided image inpainting. In: CVPR pp. 18359\u201318369.","DOI":"10.1109\/CVPR52729.2023.01761"},{"key":"e_1_3_2_221_1","article-title":"Waymo\u2019s safety methodologies and safety readiness determinations","author":"Webb N","year":"2020","unstructured":"Webb N, Smith D, Ludwick C, et al. (2020) Waymo\u2019s safety methodologies and safety readiness determinations. arXiv preprint arXiv:2011.00054.","journal-title":"arXiv preprint arXiv:2011.00054"},{"key":"e_1_3_2_222_1","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei J","year":"2022","unstructured":"Wei J, Wang X, Schuurmans D, et al. (2022) Chain-of-thought prompting elicits reasoning in large language models. NeurIPS 35: 24824\u201324837.","journal-title":"NeurIPS"},{"key":"e_1_3_2_223_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-023-2689-5"},{"key":"e_1_3_2_224_1","unstructured":"Wikipedia (2023) GPT-3. URL: https:\/\/en.wikipedia.org\/wiki\/GPT-3. [Online; accessed 28-November-2023]."},{"key":"e_1_3_2_225_1","doi-asserted-by":"crossref","unstructured":"Wu Z Xiong Y Yu SX et al. (2018) Unsupervised feature learning via non-parametric instance discrimination. In: CVPR.","DOI":"10.1109\/CVPR.2018.00393"},{"key":"e_1_3_2_226_1","doi-asserted-by":"crossref","unstructured":"Xia F R Zamir A He ZY et al. (2018) Gibson Env: real-world perception for embodied agents. In: CVPR.","DOI":"10.1109\/CVPR.2018.00945"},{"key":"e_1_3_2_227_1","article-title":"Towards generalist robots: a promising paradigm via generative simulation","author":"Xian Z","year":"2023","unstructured":"Xian Z, Gervet T, Xu Z, et al. (2023) Towards generalist robots: a promising paradigm via generative simulation. arXiv preprint arXiv:2305.10455.","journal-title":"arXiv preprint arXiv:2305.10455"},{"key":"e_1_3_2_228_1","article-title":"Masked visual pre-training for motor control","author":"Xiao T","year":"2022","unstructured":"Xiao T, Radosavovic I, Darrell T, et al. (2022) Masked visual pre-training for motor control. arXiv preprint arXiv:2203.06173.","journal-title":"arXiv preprint arXiv:2203.06173"},{"key":"e_1_3_2_229_1","doi-asserted-by":"crossref","unstructured":"Xiao T Chan H Sermanet P et al. (2023a) Robotic skill acquisition via instruction augmentation with vision-language models. In: RSS.","DOI":"10.15607\/RSS.2023.XIX.029"},{"key":"e_1_3_2_230_1","article-title":"Robot learning in the era of foundation models: a survey","author":"Xiao X","year":"2023","unstructured":"Xiao X, Liu J, Wang Z, et al. (2023b) Robot learning in the era of foundation models: a survey. arXiv preprint arXiv:2311.14379.","journal-title":"arXiv preprint arXiv:2311.14379"},{"key":"e_1_3_2_231_1","article-title":"ULIP: learning unified representation of language, image and point cloud for 3D understanding","author":"Xue L","year":"2022","unstructured":"Xue L, Gao M, Xing C, et al. (2022) ULIP: learning unified representation of language, image and point cloud for 3D understanding. arXiv preprint arXiv:2212.05171.","journal-title":"arXiv preprint arXiv:2212.05171"},{"key":"e_1_3_2_232_1","article-title":"ULIP-2: towards scalable multimodal pre-training for 3d understanding","author":"Xue L","year":"2023","unstructured":"Xue L, Yu N, Zhang S, et al. (2023) ULIP-2: towards scalable multimodal pre-training for 3d understanding. arXiv preprint arXiv:2305.08275.","journal-title":"arXiv preprint arXiv:2305.08275"},{"key":"e_1_3_2_233_1","article-title":"Track anything: segment anything meets videos","author":"Yang J","year":"2023","unstructured":"Yang J, Gao M, Li Z, et al. (2023a) Track anything: segment anything meets videos. arXiv preprint arXiv:2304.11968.","journal-title":"arXiv preprint arXiv:2304.11968"},{"key":"e_1_3_2_234_1","article-title":"Foundation models for decision making: problems, methods, and opportunities","author":"Yang S","year":"2023","unstructured":"Yang S, Nachum O, Du Y, et al. (2023b) Foundation models for decision making: problems, methods, and opportunities. arXiv preprint arXiv:2303.04129.","journal-title":"arXiv preprint arXiv:2303.04129"},{"key":"e_1_3_2_235_1","doi-asserted-by":"crossref","unstructured":"Yang F Feng C Chen Z et al. (2024) Binding touch to everything: learning unified multimodal tactile representations. In: Computer Vision and Pattern Recognition (CVPR). Springer.","DOI":"10.1109\/CVPR52733.2024.02488"},{"key":"e_1_3_2_236_1","unstructured":"Yao L Huang R Hou L et al. (2022) FILIP: fine-grained interactive language-image pre-training. In: ICLR."},{"key":"e_1_3_2_237_1","unstructured":"Yao S Zhao J Yu D et al. (2023) ReAct: synergizing reasoning and acting in language models. In: ICLR."},{"key":"e_1_3_2_238_1","article-title":"FeatureNeRF: learning generalizable nerfs by distilling pre-trained vision foundation models","author":"Ye J","year":"2023","unstructured":"Ye J, Wang N, Wang X (2023a) FeatureNeRF: learning generalizable nerfs by distilling pre-trained vision foundation models. arXiv preprint arXiv:2303.12786.","journal-title":"arXiv preprint arXiv:2303.12786"},{"key":"e_1_3_2_239_1","doi-asserted-by":"crossref","unstructured":"Ye Y Li X Gupta A et al. (2023b) Affordance diffusion: synthesizing hand-object interactions. In: CVPR.","DOI":"10.1109\/CVPR52729.2023.02153"},{"key":"e_1_3_2_240_1","article-title":"Homerobot: open vocabulary mobile manipulation","author":"Yenamandra S","year":"2023","unstructured":"Yenamandra S, Ramachandran A, Yadav K, et al. (2023) Homerobot: open vocabulary mobile manipulation. arXiv preprint arXiv:2306.11565.","journal-title":"arXiv preprint arXiv:2306.11565"},{"key":"e_1_3_2_241_1","article-title":"UL2: unifying language learning paradigms","author":"Yi T","year":"2023","unstructured":"Yi T, Mostafa D, Vinh QT, et al. (2023) UL2: unifying language learning paradigms. arXiv preprint arXiv:2205.05131.","journal-title":"arXiv preprint arXiv:2205.05131"},{"key":"e_1_3_2_242_1","article-title":"Statler: state-maintaining language models for embodied reasoning","author":"Yoneda T","year":"2023","unstructured":"Yoneda T, Fang J, Li P, et al. (2023) Statler: state-maintaining language models for embodied reasoning. arXiv preprint arXiv:2306.17840.","journal-title":"arXiv preprint arXiv:2306.17840"},{"key":"e_1_3_2_243_1","doi-asserted-by":"crossref","unstructured":"Yu X Tang L Rao Y et al. (2022) Point-BERT: pre-training 3D point cloud transformers with masked point modeling. In: CVPR pp. 19313\u201319322. Springer.","DOI":"10.1109\/CVPR52688.2022.01871"},{"key":"e_1_3_2_244_1","article-title":"Scaling robot learning with semantically imagined experience","author":"Yu T","year":"2023","unstructured":"Yu T, Xiao T, Stone A, et al. (2023) Scaling robot learning with semantically imagined experience. arXiv Preprint arXiv:2302.11550.","journal-title":"arXiv Preprint arXiv:2302.11550"},{"key":"e_1_3_2_245_1","unstructured":"Yuying G Annabella M Li EL et al. (2023) Policy adaptation from foundation model feedback. In: CVPR."},{"key":"e_1_3_2_246_1","unstructured":"Zeng A Florence P Tompson J et al. (2020) Transporter networks: rearranging the visual world for robotic manipulation. In: CoRL."},{"key":"e_1_3_2_247_1","article-title":"Socratic Models: composing zero-shot multimodal reasoning with language","author":"Zeng A","year":"2022","unstructured":"Zeng A, Attarian M, Ichter B, et al. (2022) Socratic Models: composing zero-shot multimodal reasoning with language. arXiv.","journal-title":"arXiv"},{"key":"e_1_3_2_248_1","unstructured":"Zeng A Liu X Du Z et al. (2023a) GLM-130B: an open bilingual pre-trained model. In: ICLR."},{"key":"e_1_3_2_249_1","doi-asserted-by":"crossref","unstructured":"Zeng Y Jiang C Mao J et al. (2023b) CLIP2: contrastive language-image-point pretraining from real-world point cloud data. In: CVPR.","DOI":"10.1109\/CVPR52729.2023.01463"},{"key":"e_1_3_2_250_1","doi-asserted-by":"crossref","unstructured":"Zhai X Kolesnikov A Houlsby N et al. (2022) Scaling vision transformers. In: CVPR.","DOI":"10.1109\/CVPR52688.2022.01179"},{"key":"e_1_3_2_251_1","doi-asserted-by":"crossref","unstructured":"Zhang R Guo Z Zhang W et al. (2022a) PointCLIP: point cloud understanding by CLIP. In: CVPR pp. 8552\u20138562.","DOI":"10.1109\/CVPR52688.2022.00836"},{"key":"e_1_3_2_252_1","unstructured":"Zhang Y Jiang H Miura Y et al. (2022b) Contrastive learning of medical visual representations from paired images and text. In: MLHC."},{"key":"e_1_3_2_253_1","article-title":"Can offline reinforcement learning help natural language understanding?","author":"Zhang Z","year":"2022","unstructured":"Zhang Z, Wang Y, Zhang Y, et al. (2022c) Can offline reinforcement learning help natural language understanding? arXiv preprint arXiv:2212.03864.","journal-title":"arXiv preprint arXiv:2212.03864"},{"key":"e_1_3_2_254_1","article-title":"Faster segment anything: towards lightweight sam for mobile applications","author":"Zhang C","year":"2023","unstructured":"Zhang C, Han D, Qiao Y, et al. (2023a) Faster segment anything: towards lightweight sam for mobile applications. arXiv preprint arXiv:2306.14289.","journal-title":"arXiv preprint arXiv:2306.14289"},{"key":"e_1_3_2_255_1","article-title":"SAM3D: zero-shot 3D object detection via segment anything model","author":"Zhang D","year":"2023","unstructured":"Zhang D, Liang D, Yang H, et al. (2023b) SAM3D: zero-shot 3D object detection via segment anything model. arXiv preprint arXiv:2306.02245.","journal-title":"arXiv preprint arXiv:2306.02245"},{"key":"e_1_3_2_256_1","doi-asserted-by":"crossref","unstructured":"Zhang X Kundu A Funkhouser T et al. (2023c) Nerflets: local radiance fields for efficient structure-aware 3d scene representation from 2d supervision. In: CVPR pp. 8274\u20138284.","DOI":"10.1109\/CVPR52729.2023.00800"},{"key":"e_1_3_2_257_1","article-title":"Fast segment anything","author":"Zhao X","year":"2023","unstructured":"Zhao X, Ding W, An Y, et al. (2023) Fast segment anything. arXiv preprint arXiv:2306.12156.","journal-title":"arXiv preprint arXiv:2306.12156"},{"key":"e_1_3_2_258_1","article-title":"A survey of optimization-based task and motion planning: from classical to learning approaches","author":"Zhao Z","year":"2024","unstructured":"Zhao Z, Chen S, Ding Y, et al. (2024) A survey of optimization-based task and motion planning: from classical to learning approaches. arXiv preprint arXiv:2404.02817.","journal-title":"arXiv preprint arXiv:2404.02817"},{"key":"e_1_3_2_259_1","doi-asserted-by":"crossref","unstructured":"Zhou C Loy CC Dai B (2022) Extract free dense labels from CLIP. In: European Conference on Computer Vision (ECCV) pp. 696\u2013712. Springer.","DOI":"10.1007\/978-3-031-19815-1_40"},{"key":"e_1_3_2_260_1","article-title":"Feature 3dgs: supercharging 3d Gaussian splatting to enable distilled feature fields","author":"Zhou S","year":"2023","unstructured":"Zhou S, Chang H, Jiang S, et al. (2023) Feature 3dgs: supercharging 3d Gaussian splatting to enable distilled feature fields. arXiv preprint arXiv:2312.03203.","journal-title":"arXiv preprint arXiv:2312.03203"},{"key":"e_1_3_2_261_1","article-title":"Ghost in the Minecraft: generally capable agents for open-world environments via large language models with text-based knowledge and memory","author":"Zhu X","year":"2023","unstructured":"Zhu X, Chen Y, Tian H, et al. (2023) Ghost in the Minecraft: generally capable agents for open-world environments via large language models with text-based knowledge and memory. arXiv preprint arXiv:2305.17144.","journal-title":"arXiv preprint arXiv:2305.17144"},{"key":"e_1_3_2_262_1","unstructured":"Zitkovich B Yu T Xu S et al. (2023) RT-2: vision-language-action models transfer web knowledge to robotic control. In: CoRL."},{"key":"e_1_3_2_263_1","article-title":"Fmgs: foundation model embedded 3d Gaussian splatting for holistic 3d scene understanding","author":"Zuo X","year":"2024","unstructured":"Zuo X, Samangouei P, Zhou Y, et al. (2024) Fmgs: foundation model embedded 3d Gaussian splatting for holistic 3d scene understanding. arXiv preprint arXiv:2401.01970.","journal-title":"arXiv preprint arXiv:2401.01970"}],"container-title":["The International Journal of Robotics Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649241281508","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/02783649241281508","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649241281508","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649241281508","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T10:17:34Z","timestamp":1777457854000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/02783649241281508"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,25]]},"references-count":262,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2025,4]]}},"alternative-id":["10.1177\/02783649241281508"],"URL":"https:\/\/doi.org\/10.1177\/02783649241281508","relation":{},"ISSN":["0278-3649","1741-3176"],"issn-type":[{"value":"0278-3649","type":"print"},{"value":"1741-3176","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,9,25]]}}}