{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,20]],"date-time":"2026-03-20T17:16:39Z","timestamp":1774026999015,"version":"3.50.1"},"reference-count":78,"publisher":"Association for Computing Machinery (ACM)","issue":"3","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Navigation is a fundamental task in the research of Embodied AI, and recent advances in machine learning algorithms have garnered growing interest in developing versatile Embodied AI systems. However, current research in this domain reveals opportunities for improvement. First, the direct application of RNNs and Transformers often overlooks the distinct characteristics of navigation tasks compared to traditional sequential data modeling. These methods are inherently designed to capture long-term dependencies, which are relatively weak in navigation scenarios, potentially limiting their performance in such tasks. Second, the reliance on task-specific configurations, such as pre-trained modules and dataset-specific logic, compromises the generalizability of these methods. We address these constraints by initially exploring the unique differences between Navigation tasks and other sequential data tasks through the lens of Causality, presenting a causal framework to elucidate the inadequacies of conventional sequential methods for Navigation. By leveraging this causal perspective, we propose Causality-Aware Transformer (CAT) Networks for Navigation, featuring a Causal Understanding Module to enhance the model\u2019s Environmental Understanding capability. Meanwhile, our method is devoid of task-specific inductive biases and can be trained in an End-to-End manner, which enhances the method\u2019s generalizability across various contexts. Empirical evaluations demonstrate that our methodology consistently surpasses benchmark performances across a spectrum of settings, tasks, and simulation environments, specifically, in Object Navigation within RoboTHOR, Objective Navigation, Point Navigation in Habitat, and R2R Navigation. Extensive ablation studies reveal that the performance gains can be attributed to the Causal Understanding Module, which demonstrates effectiveness and efficiency in both Reinforcement Learning and Supervised Learning settings. Additionally, further analysis highlights the robustness of our method, demonstrating its capacity to consistently perform well across diverse experimental settings and varying conditions. This robustness underscores the adaptability and generalizability of our approach, reinforcing its potential for application across a wide range of tasks.<\/jats:p>","DOI":"10.1145\/3748659","type":"journal-article","created":{"date-parts":[[2025,7,16]],"date-time":"2025-07-16T13:41:38Z","timestamp":1752673298000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Learning Causality-Aware Exploration with Transformers for Goal-Oriented Navigation"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-6617-0369","authenticated-orcid":false,"given":"Ruoyu","family":"Wang","sequence":"first","affiliation":[{"name":"Computer Science and Engineering, University of New South Wales, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5991-2050","authenticated-orcid":false,"given":"Tong","family":"Yu","sequence":"additional","affiliation":[{"name":"Adobe Inc., San Jose, California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6096-9858","authenticated-orcid":false,"given":"Mingjie","family":"Li","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3153-6732","authenticated-orcid":false,"given":"Yuanjiang","family":"Cao","sequence":"additional","affiliation":[{"name":"Macquarie University, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-0335-1694","authenticated-orcid":false,"given":"Yao","family":"Liu","sequence":"additional","affiliation":[{"name":"Macquarie University, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4149-839X","authenticated-orcid":false,"given":"Lina","family":"Yao","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, University of New South Wales, Sydney, Australia and CSIRO Data61 Business Unit, Eveleigh, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,3,20]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"J. Achiam S. Adler S. Agarwal L. Ahmad I. Akkaya F. L. Aleman D. Almeida J. Altenschmidt S. Altman S. Anadkat et al. 2023. GPT-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475282"},{"key":"e_1_3_1_4_2","first-page":"3674","article-title":"Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments","author":"Anderson P.","year":"2018","unstructured":"P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. S\u00fcnderhauf, I. Reid, S. Gould, and A. Van Den Hengel. 2018. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3674\u20133683.","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"e_1_3_1_5_2","unstructured":"L. Annabi. 2022. Intrinsically motivated learning of causal world models. arXiv:2208.04892. Retrieved from https:\/\/arxiv.org\/abs\/2208.04892"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1016\/0022-247X(65)90154-X"},{"key":"e_1_3_1_7_2","unstructured":"D. Batra A. X. Chang S. Chernova A. J. Davison J. Deng V. Koltun S. Levine J. Malik I. Mordatch R. Mottaghi et al. 2020. Rearrangement: A challenge for embodied AI. arXiv:2011.01975. Retrieved from https:\/\/arxiv.org\/abs\/2011.01975"},{"key":"e_1_3_1_8_2","unstructured":"D. Batra A. Gokaslan A. Kembhavi O. Maksymets R. Mottaghi M. Savva A. Toshev and E. Wijmans. 2020. ObjectNav revisited: On evaluation of embodied agents navigating to objects. arXiv:2006.13171. Retrieved from https:\/\/arxiv.org\/abs\/2006.13171"},{"key":"e_1_3_1_9_2","first-page":"706","volume-title":"Proceedings of the Conference on Robot Learning","author":"Blukis V.","year":"2022","unstructured":"V. Blukis, C. Paxton, D. Fox, A. Garg, and Y. Artzi. 2022. A persistent spatial semantic representation for high-level natural language instruction execution. In Proceedings of the Conference on Robot Learning. PMLR, 706\u2013717."},{"key":"e_1_3_1_10_2","doi-asserted-by":"crossref","unstructured":"A. Chang A. Dai T. Funkhouser M. Halber M. Niessner M. Savva S. Song A. Zeng and Y. Zhang. 2017. Matterport3D: Learning from RGB-D data in indoor environments. arXiv:1709.06158. Retrieved from https:\/\/arxiv.org\/abs\/1709.06158","DOI":"10.1109\/3DV.2017.00081"},{"key":"e_1_3_1_11_2","first-page":"309","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"Chaplot D. S.","year":"2020","unstructured":"D. S. Chaplot, H. Jiang, S. Gupta, and A. Gupta. 2020. Semantic curiosity for active visual learning. In: Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part VI. Springer, 309\u2013326."},{"key":"e_1_3_1_12_2","unstructured":"J. Chen B. Lin R. Xu Z. Chai X. Liang and K. Y. K. Wong. 2024. MapGPT: Map-guided prompting for unified vision-and-language navigation. arXiv:2401.07314. Retrieved from https:\/\/arxiv.org\/abs\/2401.07314"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01112"},{"key":"e_1_3_1_14_2","first-page":"5834","article-title":"History aware multimodal transformer for vision-and-language navigation","volume":"34","author":"Chen S.","year":"2021","unstructured":"S. Chen, P. L. Guhur, C. Schmid, and I. Laptev. 2021 History aware multimodal transformer for vision-and-language navigation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 34, 5834\u20135847.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_15_2","first-page":"1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Das A.","year":"2018","unstructured":"A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra. 2018. Embodied question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1\u201310."},{"key":"e_1_3_1_16_2","article-title":"Causal confusion in imitation learning","volume":"32","author":"De Haan P.","year":"2019","unstructured":"P. De Haan, D. Jayaraman, and S. Levine. 2019. Causal confusion in imitation learning. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 32.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00323"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","first-page":"5982","DOI":"10.52202\/068431-0433","article-title":"ProcTHOR: Large-scale embodied AI using procedural generation","volume":"35","author":"Deitke M.","year":"2022","unstructured":"M. Deitke, E. VanderBilt, A. Herrasti, L. Weihs, K. Ehsani, J. Salvador, W. Han, E. Kolve, A. Kembhavi, and R. Mottaghi. 2022. ProcTHOR: Large-scale embodied AI using procedural generation. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vol. 35, 5982\u20135994.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_3_1_19_2","unstructured":"Z. Deng J. Jiang G. Long and C. Zhang. 2023. Causal reinforcement learning: A survey. arXiv:2307.01452. Retrieved from https:\/\/arxiv.org\/abs\/2307.01452"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TETCI.2022.3141105"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00447"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1.13646"},{"key":"e_1_3_1_23_2","article-title":"Speaker-follower models for vision-and-language navigation","volume":"31","author":"Fried D.","year":"2018","unstructured":"D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L. P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell. 2018 Speaker-follower models for vision-and-language navigation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 31.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_24_2","unstructured":"T. Gupta W. Gong C. Ma N. Pawlowski A. Hilmkil M. Scetbon A. Famoti A. J. Llorens J. Gao S. Bauer et al. 2024. The essential role of causality in foundation world models for embodied AI. arXiv:2402.06665. Retrieved from https:\/\/arxiv.org\/abs\/2402.06665"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"issue":"8","key":"e_1_3_1_26_2","first-page":"2","article-title":"Neural networks for machine learning lecture 6a overview of mini-batch gradient descent","volume":"14","author":"Hinton G.","year":"2012","unstructured":"G. Hinton, N. Srivastava, and K. Swersky. 2012. Neural networks for machine learning lecture 6a overview of mini-batch gradient descent. Cited on 14, 8 (2012), 2.","journal-title":"Cited on"},{"key":"e_1_3_1_27_2","unstructured":"Y. Hong Q. Wu Y. Qi C. Rodriguez-Opazo and S. Gould. 2020. A recurrent vision-and-language BERT for navigation. arXiv:2011.13922. Retrieved from https:\/\/arxiv.org\/abs\/2011.13922"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48891.2023.10160969"},{"key":"e_1_3_1_29_2","unstructured":"Y. Inoue and H. Ohashi. 2022. Prompter: Utilizing large language model prompting for a data efficient embodied instruction following. arXiv:2211.03267. Retrieved from https:\/\/arxiv.org\/abs\/2211.03267"},{"key":"e_1_3_1_30_2","first-page":"6741","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ke L.","year":"2019","unstructured":"L. Ke, X. Li, Y. Bisk, A. Holtzman, Z. Gan, J. Liu, J. Gao, Y. Choi, and S. Srinivasa. 2019. Tactical rewind: Self-correction via backtracking in vision-and-language navigation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6741\u20136749."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01441"},{"key":"e_1_3_1_32_2","unstructured":"D. P. Kingma and J. Ba. 2014. Adam: A method for stochastic optimization. arXiv:1412.6980. Retrieved from https:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_1_33_2","unstructured":"M. Kocaoglu C. Snyder A. G. Dimakis and S. Vishwanath. 2017. CausalGAN: Learning causal implicit generative models with adversarial training. arXiv:1709.02023. Retrieved from https:\/\/arxiv.org\/abs\/1709.02023"},{"key":"e_1_3_1_34_2","unstructured":"E. Kolve R. Mottaghi W. Han E. VanderBilt L. Weihs A. Herrasti M. Deitke K. Ehsani D. Gordon Y. Zhu et al. 2017. AI2-THOR: An interactive 3D environment for visual AI. arXiv:1712.05474. Retrieved from https:\/\/arxiv.org\/abs\/1712.05474"},{"key":"e_1_3_1_35_2","doi-asserted-by":"crossref","unstructured":"A. Ku P. Anderson R. Patel E. Ie and J. Baldridge. 2020. Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. arXiv:2010.07954. Retrieved from https:\/\/arxiv.org\/abs\/2010.07954","DOI":"10.18653\/v1\/2020.emnlp-main.356"},{"key":"e_1_3_1_36_2","unstructured":"M. Li M. Yang F. Liu X. Chen Z. Chen and J. Wang. 2020. Causal world models by unsupervised deconfounding of physical dynamics. arXiv:2012.14228. Retrieved from https:\/\/arxiv.org\/abs\/2012.14228"},{"key":"e_1_3_1_37_2","unstructured":"B. Lin Y. Nie Z. Wei J. Chen S. Ma J. Han H. Xu X. Chang and X. Liang. 2024. NavCoT: Boosting LLM-based vision-and-language navigation via learning disentangled reasoning. arXiv:2403.07376. Retrieved from https:\/\/arxiv.org\/abs\/2403.07376"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA57147.2024.10611565"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS47612.2022.9981646"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00689"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01509"},{"key":"e_1_3_1_42_2","unstructured":"S. Y. Min D. S. Chaplot P. Ravikumar Y. Bisk and R. Salakhutdinov. 2021. Film: Following instructions in language with modular methods. arXiv:2110.07342. Retrieved from https:\/\/arxiv.org\/abs\/2110.07342"},{"key":"e_1_3_1_43_2","unstructured":"B. Pan R. Panda S. Jin R. Feris A. Oliva P. Isola and Y. Kim. 2023. LangNav: Language as a perceptual representation for navigation. arXiv:2310.07889. Retrieved from https:\/\/arxiv.org\/abs\/2310.07889"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01564"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0927-0507(05)80172-0"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01000"},{"key":"e_1_3_1_47_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford A.","year":"2021","unstructured":"A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01716"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00511"},{"key":"e_1_3_1_50_2","unstructured":"J. Richens and T. Everitt. 2024. Robust agents learn causal world models. arXiv:2402.10877. Retrieved from https:\/\/arxiv.org\/abs\/2402.10877"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00943"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2021.3058954"},{"key":"e_1_3_1_53_2","unstructured":"J. Schulman F. Wolski P. Dhariwal A. Radford and O. Klimov. 2017. Proximal policy optimization algorithms. arXiv:1707.06347. Retrieved from https:\/\/arxiv.org\/abs\/1707.06347"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS51168.2021.9636667"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01075"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1162\/1064546053278973"},{"key":"e_1_3_1_57_2","unstructured":"A. Suglia Q. Gao J. Thomason G. Thattai and G. Sukhatme. 2021. Embodied BERT: A transformer model for embodied language-guided visual task completion. arXiv:2108.04927. Retrieved from https:\/\/arxiv.org\/abs\/2108.04927"},{"key":"e_1_3_1_58_2","first-page":"251","article-title":"Habitat 2.0: Training home assistants to rearrange their habitat","volume":"34","author":"Szot A.","year":"2021","unstructured":"A. Szot, A. Clegg, E. Undersander, E. Wijmans, Y. Zhao, J. Turner, N. Maestre, M. Mukadam, D. S. Chaplot, O. Maksymets, et al. 2021. Habitat 2.0: Training home assistants to rearrange their habitat. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 34, 251\u2013266.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_59_2","doi-asserted-by":"crossref","unstructured":"H. Tan L. Yu and M. Bansal. 2019. Learning to navigate unseen environments: Back translation with environmental dropout. arXiv:1904.04195. Retrieved from https:\/\/arxiv.org\/abs\/1904.04195","DOI":"10.18653\/v1\/N19-1268"},{"key":"e_1_3_1_60_2","unstructured":"H. Touvron T. Lavril G. Izacard X. Martinet M. A. Lachaux T. Lacroix B. Rozi\u00e8re N. Goyal E. Hambro F. Azhar et al. 2023. LLaMA: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_1_61_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani A.","year":"2017","unstructured":"A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, \u0141. Kaiser, and I. Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_62_2","first-page":"161","volume-title":"Proceedings of the International Conference on Web Information Systems Engineering","author":"Wang R.","year":"2024","unstructured":"R. Wang, X. Li, and L. Yao. 2024. Deconfounded causality-aware parameter-efficient fine-tuning for problem-solving improvement of LLMs. In Proceedings of the International Conference on Web Information Systems Engineering. Springer, 161\u2013176."},{"key":"e_1_3_1_63_2","unstructured":"R. Wang X. Li C. Wang and L. Yao. 2025. Efficient and generalizable environmental understanding for visual navigation. arXiv:2506.15377. Retrieved from https:\/\/arxiv.org\/abs\/2506.15377"},{"key":"e_1_3_1_64_2","unstructured":"R. Wang Y. Liu Y. Cao and L. Yao. 2024. Causality-aware transformer networks for robotic navigation. arXiv:2409.02669. Retrieved from https:\/\/arxiv.org\/abs\/2409.02669"},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","unstructured":"R. Wang T. Yu J. Wu Y. Liu J. McAuley and L. Yao. 2025. Weakly-supervised VLM-guided partial contrastive learning for visual language navigation. arXiv:2506.15757. Retrieved from https:\/\/arxiv.org\/abs\/2506.15757","DOI":"10.1109\/IROS60139.2025.11247068"},{"key":"e_1_3_1_66_2","unstructured":"Z. Wang X. Xiao Z. Xu Y. Zhu and P. Stone. 2022. Causal dynamics learning for task-independent state abstraction. arXiv:2206.13452. Retrieved from https:\/\/arxiv.org\/abs\/2206.13452"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00586"},{"key":"e_1_3_1_68_2","unstructured":"L. Weihs J. Salvador K. Kotar U. Jain K. H. Zeng R. Mottaghi and A. Kembhavi. 2020. AllenAct: A framework for embodied AI research. arXiv:2008.12760. Retrieved from https:\/\/arxiv.org\/abs\/2008.12760"},{"key":"e_1_3_1_69_2","unstructured":"E. Wijmans A. Kadian A. Morcos S. Lee I. Essa D. Parikh M. Savva and D. Batra. 2019. DD-PPO: Learning near-perfect PointGoal navigators from 2.5 billion frames. arXiv:1911.00357. Retrieved from https:\/\/arxiv.org\/abs\/1911.00357"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00945"},{"key":"e_1_3_1_71_2","doi-asserted-by":"crossref","unstructured":"X. Xiao J. Liu Z. Wang Y. Zhou Y. Qi Q. Cheng B. He and S. Jiang. 2023. Robot learning in the era of foundation models: A survey. arXiv:2311.14379. Retrieved from https:\/\/arxiv.org\/abs\/2311.14379","DOI":"10.2139\/ssrn.4706193"},{"key":"e_1_3_1_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01581"},{"key":"e_1_3_1_73_2","first-page":"2734","article-title":"Interventional few-shot learning","volume":"33","author":"Yue Z.","year":"2020","unstructured":"Z. Yue, H. Zhang, Q. Sun, and X. S. Hua. 2020. Interventional few-shot learning. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 33, 2734\u20132746.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_74_2","unstructured":"A. Zhang R. McAllister R. Calandra Y. Gal and S. Levine. 2020. Learning invariant representations for reinforcement learning without reconstruction. arXiv:2006.10742. Retrieved from https:\/\/arxiv.org\/abs\/2006.10742"},{"key":"e_1_3_1_75_2","doi-asserted-by":"crossref","unstructured":"Y. Zhang and J. Chai. 2021. Hierarchical task learning from language instructions with unified transformers and self-monitoring. arXiv:2106.03427. Retrieved from https:\/\/arxiv.org\/abs\/2106.03427","DOI":"10.18653\/v1\/2021.findings-acl.368"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01293"},{"key":"e_1_3_1_77_2","unstructured":"K. Zheng K. Zhou J. Gu Y. Fan J. Wang Z. Di X. He and X. E. Wang. 2022. JARVIS: A neuro-symbolic commonsense reasoning framework for conversational embodied agents. arXiv:2208.13266. Retrieved from https:\/\/arxiv.org\/abs\/2208.13266"},{"key":"e_1_3_1_78_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i7.28597"},{"key":"e_1_3_1_79_2","unstructured":"Z. M. Zhu X. H. Chen H. L. Tian K. Zhang and Y. Yu. 2022. Offline reinforcement learning with causal structured world models. arXiv:2206.01474. Retrieved from https:\/\/arxiv.org\/abs\/2206.01474"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3748659","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,20]],"date-time":"2026-03-20T16:27:22Z","timestamp":1774024042000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3748659"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,20]]},"references-count":78,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3748659"],"URL":"https:\/\/doi.org\/10.1145\/3748659","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,20]]},"assertion":[{"value":"2024-11-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-10","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}