{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T18:43:16Z","timestamp":1776105796512,"version":"3.50.1"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T00:00:00Z","timestamp":1701734400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Tencent AI Lab Rhino-Bird Focused Research Program","award":["RBFR2022013"],"award-info":[{"award-number":["RBFR2022013"]}]},{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["62322207, U21B2023, U2001206"],"award-info":[{"award-number":["62322207, U21B2023, U2001206"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"GD Natural Science Foundation","award":["2021B1515020085"],"award-info":[{"award-number":["2021B1515020085"]}]},{"name":"Shenzhen Science and Technology Program","award":["RCYX20210609103121030"],"award-info":[{"award-number":["RCYX20210609103121030"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2023,12,5]]},"abstract":"<jats:p>\n            We address the problem of scene-aware activity program generation, which requires decomposing a given activity task into instructions that can be sequentially performed within a target scene to complete the activity. While existing methods have shown the ability to generate rational or executable programs, generating programs with both high rationality and executability still remains a challenge. Hence, we propose a novel method where the key idea is to explicitly combine the language rationality of a powerful language model with dynamic perception of the target scene where instructions are executed, to generate programs with high rationality\n            <jats:italic toggle=\"yes\">and<\/jats:italic>\n            executability. Our method iteratively generates instructions for the activity program. Specifically, a two-branch feature encoder operates on a language-based and graph-based representation of the current generation progress to extract language features and scene graph features, respectively. These features are then used by a predictor to generate the next instruction in the program. Subsequently, another module performs the predicted action and updates the scene for perception in the next iteration. Extensive evaluations are conducted on the VirtualHome-Env dataset, showing the advantages of our method over previous work. Key algorithmic designs are validated through ablation studies, and results on other types of inputs are also presented to show the generalizability of our method.\n          <\/jats:p>","DOI":"10.1145\/3618338","type":"journal-article","created":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T10:20:48Z","timestamp":1701771648000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Scene-Aware Activity Program Generation with Language Guidance"],"prefix":"10.1145","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9121-6779","authenticated-orcid":false,"given":"Zejia","family":"Su","sequence":"first","affiliation":[{"name":"Shenzhen University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1249-2826","authenticated-orcid":false,"given":"Qingnan","family":"Fan","sequence":"additional","affiliation":[{"name":"Vivo, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-0158-9469","authenticated-orcid":false,"given":"Xuelin","family":"Chen","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9869-6832","authenticated-orcid":false,"given":"Oliver","family":"Van Kaick","sequence":"additional","affiliation":[{"name":"Carleton University, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3212-0544","authenticated-orcid":false,"given":"Hui","family":"Huang","sequence":"additional","affiliation":[{"name":"Shenzhen University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6798-0336","authenticated-orcid":false,"given":"Ruizhen","family":"Hu","sequence":"additional","affiliation":[{"name":"Shenzhen University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,12,5]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"Michael Ahn Anthony Brohan Noah Brown Yevgen Chebotar Omar Cortes Byron David Chelsea Finn Keerthana Gopalakrishnan Karol Hausman Alex Herzog et al. 2022. Do as i can not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691 (2022)."},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00209"},{"key":"e_1_2_2_3_1","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020) 1877--1901."},{"key":"e_1_2_2_4_1","volume-title":"Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al.","author":"Chen Mark","year":"2021","unstructured":"Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)."},{"key":"e_1_2_2_5_1","volume-title":"Dzmitry Bahdanau, and Yoshua Bengio.","author":"Cho Kyunghyun","year":"2014","unstructured":"Kyunghyun Cho, Bart Van Merri\u00ebnboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014. On the properties of neural machine translation: Encoder-decoder approaches. arXiv preprint arXiv:1409.1259 (2014)."},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.261"},{"key":"e_1_2_2_7_1","volume-title":"Procthor: Large-scale embodied ai using procedural generation. arXiv preprint arXiv:2206.06994","author":"Deitke Matt","year":"2022","unstructured":"Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Jordi Salvador, Kiana Ehsani, Winson Han, Eric Kolve, Ali Farhadi, Aniruddha Kembhavi, et al. 2022. Procthor: Large-scale embodied ai using procedural generation. arXiv preprint arXiv:2206.06994 (2022)."},{"key":"e_1_2_2_8_1","volume-title":"Speaker-follower models for vision-and-language navigation. Advances in Neural Information Processing Systems 31","author":"Fried Daniel","year":"2018","unstructured":"Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. 2018. Speaker-follower models for vision-and-language navigation. Advances in Neural Information Processing Systems 31 (2018)."},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01075"},{"key":"e_1_2_2_10_1","volume-title":"Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. arXiv preprint arXiv:2201.07207","author":"Huang Wenlong","year":"2022","unstructured":"Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. 2022. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. arXiv preprint arXiv:2201.07207 (2022)."},{"key":"e_1_2_2_11_1","volume-title":"Visually-grounded planning without vision: Language models infer detailed plans from high-level instructions. arXiv preprint arXiv:2009.14259","author":"Jansen Peter A","year":"2020","unstructured":"Peter A Jansen. 2020. Visually-grounded planning without vision: Language models infer detailed plans from high-level instructions. arXiv preprint arXiv:2009.14259 (2020)."},{"key":"e_1_2_2_12_1","volume-title":"Ai2-thor: An interactive 3d environment for visual ai. arXiv preprint arXiv:1712.05474","author":"Kolve Eric","year":"2017","unstructured":"Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi. 2017. Ai2-thor: An interactive 3d environment for visual ai. arXiv preprint arXiv:1712.05474 (2017)."},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00725"},{"key":"e_1_2_2_14_1","volume-title":"Conference on Robot Learning. PMLR, 80--93","author":"Li Chengshu","year":"2023","unstructured":"Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Mart\u00edn-Mart\u00edn, Chen Wang, Gabrael Levine, Michael Lingelbach, Jiankai Sun, et al. 2023. Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation. In Conference on Robot Learning. PMLR, 80--93."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00645"},{"key":"e_1_2_2_16_1","volume-title":"Miguel Eckstein, and William Yang Wang.","author":"Lu Yujie","year":"2022","unstructured":"Yujie Lu, Weixi Feng, Wanrong Zhu, Wenda Xu, Xin Eric Wang, Miguel Eckstein, and William Yang Wang. 2022. Neuro-Symbolic Causal Language Planning with Commonsense Prompting. arXiv preprint arXiv:2206.02928 (2022)."},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357943"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467350"},{"key":"e_1_2_2_19_1","unstructured":"Corey Lynch and Pierre Sermanet. 2020a. Grounding language in play. (2020)."},{"key":"e_1_2_2_20_1","volume-title":"Language conditioned imitation learning over unstructured data. arXiv preprint arXiv:2005.07648","author":"Lynch Corey","year":"2020","unstructured":"Corey Lynch and Pierre Sermanet. 2020b. Language conditioned imitation learning over unstructured data. arXiv preprint arXiv:2005.07648 (2020)."},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58539-6_16"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2019.107000"},{"key":"e_1_2_2_23_1","volume-title":"Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 (2013)."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1096"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1177\/0278364915602060"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2012.6385923"},{"key":"e_1_2_2_27_1","unstructured":"OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL]"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00886"},{"key":"e_1_2_2_29_1","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei Ilya Sutskever et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1 8 (2019) 9."},{"key":"e_1_2_2_30_1","unstructured":"Santhosh K Ramakrishnan Aaron Gokaslan Erik Wijmans Oleksandr Maksymets Alex Clegg John Turner Eric Undersander Wojciech Galuba Andrew Westbury Angel X Chang et al. 2021. Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai. arXiv preprint arXiv:2109.08238 (2021)."},{"key":"e_1_2_2_31_1","volume-title":"Temporal graph networks for deep learning on dynamic graphs. arXiv preprint arXiv:2006.10637","author":"Rossi Emanuele","year":"2020","unstructured":"Emanuele Rossi, Ben Chamberlain, Fabrizio Frasca, Davide Eynard, Federico Monti, and Michael Bronstein. 2020. Temporal graph networks for deep learning on dynamic graphs. arXiv preprint arXiv:2006.10637 (2020)."},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3336191.3371845"},{"key":"e_1_2_2_33_1","volume-title":"Progprompt: Generating situated robot task plans using large language models. arXiv preprint arXiv:2209.11302","author":"Singh Ishika","year":"2022","unstructured":"Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. 2022. Progprompt: Generating situated robot task plans using large language models. arXiv preprint arXiv:2209.11302 (2022)."},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.11164"},{"key":"e_1_2_2_35_1","volume-title":"Understanding and executing instructions for everyday manipulation tasks from the world wide web. In 2010 ieee international conference on robotics and automation","author":"Tenorth Moritz","unstructured":"Moritz Tenorth, Daniel Nyga, and Michael Beetz. 2010. Understanding and executing instructions for everyday manipulation tasks from the world wide web. In 2010 ieee international conference on robotics and automation. IEEE, 1486--1491."},{"key":"e_1_2_2_36_1","volume-title":"International conference on learning representations.","author":"Trivedi Rakshit","year":"2019","unstructured":"Rakshit Trivedi, Mehrdad Farajtabar, Prasenjeet Biswal, and Hongyuan Zha. 2019. Dyrep: Learning representations over dynamic graphs. In International conference on learning representations."},{"key":"e_1_2_2_37_1","doi-asserted-by":"crossref","unstructured":"Shreshth Tuli Rajas Bansal Rohan Paul et al. 2021. Tango: commonsense generalization in predicting tool interactions for mobile manipulators. arXiv preprint arXiv:2105.04556 (2021).","DOI":"10.24963\/ijcai.2021\/577"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00679"},{"key":"e_1_2_2_39_1","volume-title":"Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le.","author":"Wei Jason","year":"2021","unstructured":"Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652 (2021)."},{"key":"e_1_2_2_40_1","volume-title":"Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation. arXiv preprint arXiv:2206.02369","author":"Xu Jin","year":"2022","unstructured":"Jin Xu, Xiaojiang Liu, Jianhao Yan, Deng Cai, Huayang Li, and Jian Li. 2022. Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation. arXiv preprint arXiv:2206.02369 (2022)."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/HUMANOIDS.2014.7041483"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v29i1.9671"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2022.3225327"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380076"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3618338","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3618338","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T10:53:28Z","timestamp":1755773608000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3618338"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,5]]},"references-count":44,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,12,5]]}},"alternative-id":["10.1145\/3618338"],"URL":"https:\/\/doi.org\/10.1145\/3618338","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,5]]},"assertion":[{"value":"2023-12-05","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}