{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,5]],"date-time":"2026-04-05T05:35:04Z","timestamp":1775367304266,"version":"3.50.1"},"reference-count":70,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["62325211, 62132021, 62322207, 62522219, 62372457, 62572477"],"award-info":[{"award-number":["62325211, 62132021, 62322207, 62522219, 62372457, 62572477"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Major Program of Xiangjiang Laboratory","award":["23XJ01009"],"award-info":[{"award-number":["23XJ01009"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>The household rearrangement task involves spotting misplaced objects in a scene and accommodate them with proper places. It depends both on common-sense knowledge on the objective side and human user preference on the subjective side. In achieving such a task, we propose to mine object functionality with user preference alignment directly from the scene itself, without relying on human intervention. To do so, we work with scene graph representation and propose LLM-enhanced scene graph learning which transforms the input scene graph into an Affordance Enhanced Graph (AEG) with information-enriched nodes and newly discovered edges (relations). In AEG, the nodes corresponding to the receptacle objects are augmented with context-induced affordance which encodes what kind of carriable objects can be placed on it. New edges are discovered with newly discovered non-local relations. With AEG, we perform task planning for scene rearrangement by detecting misplaced carriables and determining a proper placement for each of them. We implement an end-to-end robot system for autonomous household rearrangement in unseen environments and test our method by implementing a tiding robot in both simulated environments and real-world scenarios, and perform evaluation on a new benchmark we build. Extensive evaluations demonstrate that our method achieves state-of-the-art performance in misplacement detection and rearrangement planning.<\/jats:p>","DOI":"10.1145\/3795693","type":"journal-article","created":{"date-parts":[[2026,2,24]],"date-time":"2026-02-24T11:40:00Z","timestamp":1771933200000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["LLM-enhanced Scene Graph Learning for Household Rearrangement"],"prefix":"10.1145","volume":"45","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-4487-6148","authenticated-orcid":false,"given":"Wenhao","family":"Li","sequence":"first","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1124-3830","authenticated-orcid":false,"given":"Shilong","family":"Zou","sequence":"additional","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-6269-4639","authenticated-orcid":false,"given":"Zhinan","family":"Yu","sequence":"additional","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-2524-9554","authenticated-orcid":false,"given":"Zheng","family":"Zhou","sequence":"additional","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-1048-8807","authenticated-orcid":false,"given":"Wenxuan","family":"Li","sequence":"additional","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2838-8601","authenticated-orcid":false,"given":"Chenyang","family":"Zhu","sequence":"additional","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6798-0336","authenticated-orcid":false,"given":"Ruizhen","family":"Hu","sequence":"additional","affiliation":[{"name":"Shenzhen University","place":["Shenzhen, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9054-0216","authenticated-orcid":false,"given":"Kai","family":"Xu","sequence":"additional","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]},{"name":"Institute of AI for Industries, Chinese Academy of Science","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,3,14]]},"reference":[{"key":"e_1_3_2_2_1","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et\u00a0al. 2023. Gpt-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_3_1","doi-asserted-by":"crossref","first-page":"422","DOI":"10.1007\/978-3-030-58452-8_25","volume-title":"ECCV 2020: Proceedings of the 16th European Conference on Computer Vision, Part I 16","author":"Achlioptas Panos","year":"2020","unstructured":"Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas Guibas. 2020. Referit3D: Neural listeners for fine-grained 3D object identification in real-world scenes. In ECCV 2020: Proceedings of the 16th European Conference on Computer Vision, Part I 16. Springer, 422\u2013440."},{"key":"e_1_3_2_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00453-012-9717-4"},{"key":"e_1_3_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01854"},{"key":"e_1_3_2_6_1","unstructured":"Shuai Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge Sibo Song Kai Dang Peng Wang Shijie Wang Jun Tang et\u00a0al. 2025. Qwen2. 5-vl technical report. arXiv:2502.13923. Retrieved from https:\/\/arxiv.org\/abs\/2502.13923"},{"key":"e_1_3_2_7_1","unstructured":"Dhruv Batra Angel X. Chang Sonia Chernova Andrew J. Davison Jia Deng Vladlen Koltun Sergey Levine Jitendra Malik Igor Mordatch Roozbeh Mottaghi et\u00a0al. 2020. Rearrangement: A challenge for embodied AI. arXiv:2011.01975. Retrieved from https:\/\/arxiv.org\/abs\/2011.01975"},{"issue":"1","key":"e_1_3_2_8_1","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1056\/NEJMoa1411587","article-title":"A randomized trial of intraarterial treatment for acute ischemic stroke","volume":"372","author":"Berkhemer Olvert A.","year":"2015","unstructured":"Olvert A. Berkhemer, Puck S. S. Fransen, Debbie Beumer, Lucie A. Van Den Berg, Hester F. Lingsma, Albert J. Yoo, Wouter J. Schonewille, Jan Albert Vos, Paul J. Nederkoorn, Marieke J. H. Wermer, et\u00a0al. 2015. A randomized trial of intraarterial treatment for acute ischemic stroke. New England Journal of Medicine 372, 1 (2015), 11\u201320.","journal-title":"New England Journal of Medicine"},{"key":"e_1_3_2_9_1","unstructured":"Yihan Cao Jiazhao Zhang Zhinan Yu Shuzhen Liu Zheng Qin Qin Zou Bo Du and Kai Xu. 2024. Cognav: Cognitive process modeling for object goal navigation with LLMs. arXiv:2412.10439. Retrieved from https:\/\/arxiv.org\/abs\/2412.10439"},{"key":"e_1_3_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00182"},{"key":"e_1_3_2_11_1","first-page":"16980","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Duan Yao","year":"2022","unstructured":"Yao Duan, Chenyang Zhu, Yuqing Lan, Renjiao Yi, Xinwang Liu, and Kai Xu. 2022. Disarm: Displacement aware relation module for 3D detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 16980\u201316989."},{"key":"e_1_3_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00447"},{"key":"e_1_3_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2023.3281153"},{"key":"e_1_3_2_14_1","first-page":"2139","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Fang Kuan","year":"2018","unstructured":"Kuan Fang, Te-Lin Wu, Daniel Yang, Silvio Savarese, and Joseph J. Lim. 2018. Demo2vec: Reasoning object affordances from online videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2139\u20132147."},{"key":"e_1_3_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818057"},{"key":"e_1_3_2_16_1","unstructured":"Yunfan Gao Yun Xiong Xinyu Gao Kangxiang Jia Jinliu Pan Yuxi Bi Yi Dai Jiawei Sun and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv:2312.10997. Retrieved from https:\/\/arxiv.org\/abs\/2312.10997"},{"key":"e_1_3_2_17_1","doi-asserted-by":"crossref","unstructured":"Georgios Georgakis Arsalan Mousavian Alexander C. Berg and Jana Kosecka. 2017. Synthesizing training data for object detection in indoor scenes. arXiv:1702.07836. Retrieved from https:\/\/arxiv.org\/abs\/1702.07836","DOI":"10.15607\/RSS.2017.XIII.043"},{"issue":"2","key":"e_1_3_2_18_1","first-page":"67","article-title":"The theory of affordances","volume":"1","author":"Gibson James J.","year":"1977","unstructured":"James J. Gibson. 1977. The theory of affordances. Hilldale, USA 1, 2 (1977), 67\u201382.","journal-title":"Hilldale, USA"},{"key":"e_1_3_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995327"},{"key":"e_1_3_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995448"},{"key":"e_1_3_2_21_1","unstructured":"Dongge Han Trevor McInroe Adam Jelley Stefano V. Albrecht Peter Bell and Amos Storkey. 2024. LLM-personalize: Aligning LLM planners with human preferences via reinforced self-training for housekeeping robots. arXiv:2404.14285. Retrieved from https:\/\/arxiv.org\/abs\/2404.14285"},{"key":"e_1_3_2_22_1","first-page":"20482","article-title":"3D-LLM: Injecting the 3D world into large language models","volume":"36","author":"Hong Yining","year":"2023","unstructured":"Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan. 2023. 3D-LLM: Injecting the 3D world into large language models. Advances in Neural Information Processing Systems 36 (2023), 20482\u201320494.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_23_1","first-page":"603","volume-title":"Computer Graphics Forum","author":"Hu Ruizhen","year":"2018","unstructured":"Ruizhen Hu, Manolis Savva, and Oliver van Kaick. 2018. Functionality representations and applications for shape analysis. In Computer Graphics Forum, Vol. 37. Wiley Online Library, 603\u2013624."},{"key":"e_1_3_2_24_1","unstructured":"Dehao Huang Chao Tang and Hong Zhang. 2023. Efficient object rearrangement via multi-view fusion. arXiv:2309.08994. Retrieved from https:\/\/arxiv.org\/abs\/2309.08994"},{"key":"e_1_3_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/582415.582418"},{"key":"e_1_3_2_26_1","article-title":"ConceptFusion: Open-set multimodal 3D mapping","author":"Jatavallabhula Krishna Murthy","year":"2023","unstructured":"Krishna Murthy Jatavallabhula, Alihusein Kuwajerwala, Qiao Gu, Mohd Omama, Tao Chen, Shuang Li, Ganesh Iyer, Soroush Saryazdi, Nikhil Keetha, Ayush Tewari, et\u00a0al. 2023. ConceptFusion: Open-set multimodal 3D mapping. Robotics: Science and Systems (RSS) (2023).","journal-title":"Robotics: Science and Systems (RSS)"},{"key":"e_1_3_2_27_1","first-page":"355","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Kant Yash","year":"2022","unstructured":"Yash Kant, Arun Ramachandran, Sriram Yenamandra, Igor Gilitschenski, Dhruv Batra, Andrew Szot, and Harsh Agrawal. 2022. Housekeep: Tidying virtual households using commonsense reasoning. In Proceedings of the European Conference on Computer Vision. Springer, 355\u2013373."},{"key":"e_1_3_2_28_1","doi-asserted-by":"crossref","unstructured":"Mukul Khanna Yongsen Mao Hanxiao Jiang Sanjay Haresh Brennan Shacklett Dhruv Batra Alexander Clegg Eric Undersander Angel X. Chang and Manolis Savva. 2023. Habitat synthetic scenes dataset (HSSD-200): An analysis of 3D scene scale and realism tradeoffs for objectgoal navigation. arxiv:2306.11290 [cs.CV]. Retrieved from https:\/\/arxiv.org\/abs\/2306.11290","DOI":"10.1109\/CVPR52733.2024.01550"},{"key":"e_1_3_2_29_1","doi-asserted-by":"crossref","first-page":"831","DOI":"10.1007\/978-3-319-10578-9_54","volume-title":"ECCV 2014: Proceedings of the 13th European Conference on Computer Vision, Part III 13","author":"Koppula Hema S.","year":"2014","unstructured":"Hema S. Koppula and Ashutosh Saxena. 2014. Physically grounded spatio-temporal object affordances. In ECCV 2014: Proceedings of the 13th European Conference on Computer Vision, Part III 13. Springer, 831\u2013847."},{"key":"e_1_3_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01639"},{"key":"e_1_3_2_31_1","first-page":"9459","article-title":"Retrieval-augmented generation for knowledge-intensive NLP tasks","volume":"33","author":"Lewis Patrick","year":"2020","unstructured":"Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K\u00fcttler, Mike Lewis, Wen-tau Yih, Tim Rockt\u00e4schel, et\u00a0al. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems 33 (2020), 9459\u20139474.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3478513.3480478"},{"key":"e_1_3_2_33_1","unstructured":"Gen Li Deqing Sun Laura Sevilla-Lara and Varun Jampani. 2023. One-shot open affordance learning with foundation models. arXiv:2311.17776. Retrieved from https:\/\/arxiv.org\/abs\/2311.17776"},{"key":"e_1_3_2_34_1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"LI QI","year":"2022","unstructured":"QI LI, Kaichun Mo, Yanchao Yang, Hang Zhao, and Leonidas Guibas. 2022. IFR-explore: Learning inter-object functional relationships in 3D indoor scenes. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=OT3mLgR8Wg8"},{"key":"e_1_3_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3680528.3687607"},{"key":"e_1_3_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01265"},{"key":"e_1_3_2_37_1","unstructured":"Aixin Liu Bei Feng Bing Xue Bingxuan Wang Bochao Wu Chengda Lu Chenggang Zhao Chengqi Deng Chenyu Zhang Chong Ruan et\u00a0al. 2024. Deepseek-v3 technical report. arXiv:2412.19437. Retrieved from https:\/\/arxiv.org\/abs\/2412.19437"},{"key":"e_1_3_2_38_1","doi-asserted-by":"crossref","unstructured":"Weiyu Liu Yilun Du Tucker Hermans Sonia Chernova and Chris Paxton. 2022. Structdiffusion: Language-guided creation of physically-valid structures using unseen objects. arXiv preprint arXiv:2211.04604 (2022).","DOI":"10.15607\/RSS.2023.XIX.031"},{"key":"e_1_3_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA46639.2022.9811931"},{"issue":"1","key":"e_1_3_2_40_1","first-page":"486","article-title":"Ocrtoc: A cloud-based competition and benchmark for robotic grasping and manipulation","volume":"7","author":"Liu Ziyuan","year":"2021","unstructured":"Ziyuan Liu, Wei Liu, Yuzhe Qin, Fanbo Xiang, Minghao Gou, Songyan Xin, Maximo A. Roa, Berk Calli, Hao Su, Yu Sun, et\u00a0al. 2021. Ocrtoc: A cloud-based competition and benchmark for robotic grasping and manipulation. IEEE Robotics and Automation Letters 7, 1 (2021), 486\u2013493.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"e_1_3_2_41_1","first-page":"1610","volume-title":"Proceedings of the Conference on Robot Learning","author":"Lu Shiyang","year":"2023","unstructured":"Shiyang Lu, Haonan Chang, Eric Pu Jing, Abdeslam Boularias, and Kostas Bekris. 2023. Ovir-3D: Open-vocabulary 3D instance retrieval without training on 3D data. In Proceedings of the Conference on Robot Learning. PMLR, 1610\u20131620."},{"key":"e_1_3_2_42_1","first-page":"1666","volume-title":"Proceedings of the Conference on Robot Learning","author":"Mo Kaichun","year":"2022","unstructured":"Kaichun Mo, Yuzhe Qin, Fanbo Xiang, Hao Su, and Leonidas Guibas. 2022. O2O-Afford: Annotation-free large-scale object-object affordance learning. In Proceedings of the Conference on Robot Learning. PMLR, 1666\u20131677."},{"key":"e_1_3_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00878"},{"key":"e_1_3_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00024"},{"key":"e_1_3_2_45_1","first-page":"5692","volume-title":"Proceedings of the 2023 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"Nguyen Toan","year":"2023","unstructured":"Toan Nguyen, Minh Nhat Vu, An Vuong, Dzung Nguyen, Thieu Vo, Ngan Le, and Anh Nguyen. 2023. Open-vocabulary affordance detection in 3D point clouds. In Proceedings of the 2023 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 5692\u20135698."},{"key":"e_1_3_2_46_1","unstructured":"Zhe Ni Xiao-Xin Deng Cong Tai Xin-Yue Zhu Xiang Wu Yong-Jin Liu and Long Zeng. 2023. Grid: Scene-graph-based instruction-driven robotic task planning. arXiv:2309.07726. Retrieved from https:\/\/arxiv.org\/abs\/2309.07726"},{"key":"e_1_3_2_47_1","doi-asserted-by":"publisher","DOI":"10.1080\/02693799408901986"},{"key":"e_1_3_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/IWSSIP48289.2020.9145130"},{"key":"e_1_3_2_49_1","first-page":"e14927","volume-title":"Computer Graphics Forum","author":"Patil Akshay Gadi","year":"2024","unstructured":"Akshay Gadi Patil, Supriya Gadi Patil, Manyi Li, Matthew Fisher, Manolis Savva, and Hao Zhang. 2024. Advances in data-driven analysis and synthesis of 3D indoor scenes. In Computer Graphics Forum, Vol. 43. Wiley Online Library, e14927."},{"key":"e_1_3_2_50_1","unstructured":"Xavier Puig Eric Undersander Andrew Szot Mikael Dallaire Cote Tsung-Yen Yang Ruslan Partsey Ruta Desai Alexander William Clegg Michal Hlavac So Yeon Min et\u00a0al. 2023. Habitat 3.0: A co-habitat for humans avatars and robots. arXiv:2310.13724. Retrieved from https:\/\/arxiv.org\/abs\/2310.13724"},{"key":"e_1_3_2_51_1","unstructured":"Zackary Rackauckas. 2024. Rag-fusion: A new take on retrieval-augmented generation. arXiv:2402.03367. Retrieved from https:\/\/arxiv.org\/abs\/2402.03367"},{"key":"e_1_3_2_52_1","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_2_53_1","unstructured":"Abhinav Rajvanshi Karan Sikka Xiao Lin Bhoram Lee Han-Pang Chiu and Alvaro Velasquez. 2023. Saynav: Grounding large language models for dynamic planning to navigation in new environments. arXiv:2309.04077. Retrieved from https:\/\/arxiv.org\/abs\/2309.04077"},{"key":"e_1_3_2_54_1","volume-title":"Proceedings of the 7th Annual Conference on Robot Learning","author":"Rana Krishan","year":"2023","unstructured":"Krishan Rana, Jesse Haviland, Sourav Garg, Jad Abou-Chakra, Ian Reid, and Niko Suenderhauf. 2023. Sayplan: Grounding large language models using 3D scene graphs for scalable robot task planning. In Proceedings of the 7th Annual Conference on Robot Learning."},{"key":"e_1_3_2_55_1","unstructured":"Tianhe Ren Qing Jiang Shilong Liu Zhaoyang Zeng Wenlong Liu Han Gao Hongjie Huang Zhengyu Ma Xiaoke Jiang Yihao Chen et\u00a0al. 2024. Grounding dino 1.5: Advance the \u201cedge\u201d of open-set object detection. arXiv:2405.10300. Retrieved from https:\/\/arxiv.org\/abs\/2405.10300"},{"key":"e_1_3_2_56_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19842-7_28"},{"issue":"6","key":"e_1_3_2_57_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2661229.2661230","article-title":"SceneGrok: Inferring action maps in 3D environments","volume":"33","author":"Savva Manolis","year":"2014","unstructured":"Manolis Savva, Angel X. Chang, Pat Hanrahan, Matthew Fisher, and Matthias Nie\u00dfner. 2014. SceneGrok: Inferring action maps in 3D environments. ACM Transactions on Graphics (TOG) 33, 6 (2014), 1\u201310.","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_3_2_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925867"},{"key":"e_1_3_2_59_1","first-page":"747","volume-title":"Proceedings of the 2020 IEEE 7th International Conference on Data Science and Advanced Analytics (DSAA)","author":"Shahapure Ketan Rajshekhar","year":"2020","unstructured":"Ketan Rajshekhar Shahapure and Charles Nicholas. 2020. Cluster quality analysis using silhouette score. In Proceedings of the 2020 IEEE 7th International Conference on Data Science and Advanced Analytics (DSAA). IEEE, 747\u2013748."},{"key":"e_1_3_2_60_1","doi-asserted-by":"crossref","first-page":"746","DOI":"10.1007\/978-3-642-33715-4_54","volume-title":"Computer Vision\u2013ECCV 2012: Proceedings of the 12th European Conference on Computer Vision, Part V 12","author":"Silberman Nathan","year":"2012","unstructured":"Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. 2012. Indoor segmentation and support inference from RGBD images. In Computer Vision\u2013ECCV 2012: Proceedings of the 12th European Conference on Computer Vision, Part V 12. Springer, 746\u2013760."},{"key":"e_1_3_2_61_1","unstructured":"Chao Tang Jingwen Yu Weinan Chen and Hong Zhang. 2021. Relationship oriented affordance learning through manipulation graph construction. arXiv:2110.14137. Retrieved from https:\/\/arxiv.org\/abs\/2110.14137"},{"key":"e_1_3_2_62_1","doi-asserted-by":"crossref","unstructured":"Tuan Van Vo Minh Nhat Vu Baoru Huang Toan Nguyen Ngan Le Thieu Vo and Anh Nguyen. 2023. Open-vocabulary affordance detection using knowledge distillation and text-point correlation. .arXiv:2309.10932. Retrieved from https:\/\/arxiv.org\/abs\/2309.10932","DOI":"10.1109\/ICRA57147.2024.10610247"},{"key":"e_1_3_2_63_1","doi-asserted-by":"crossref","unstructured":"Zan Wang Yixin Chen Baoxiong Jia Puhao Li Jinlu Zhang Jingze Zhang Tengyu Liu Yixin Zhu Wei Liang and Siyuan Huang. 2024. Move as you say interact as you can: Language-guided human motion generation with scene affordance. arXiv:2403.18036. Retrieved from https:\/\/arxiv.org\/abs\/2403.18036","DOI":"10.1109\/CVPR52733.2024.00049"},{"key":"e_1_3_2_64_1","unstructured":"Zehan Wang Haifeng Huang Yang Zhao Ziang Zhang and Zhou Zhao. 2023. Chat-3d: Data-efficiently tuning large language model for universal dialogue of 3D scenes. arXiv:2308.08769. Retrieved from https:\/\/arxiv.org\/abs\/2308.08769"},{"key":"e_1_3_2_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00586"},{"key":"e_1_3_2_66_1","article-title":"TidyBot: Personalized robot assistance with large language models","author":"Wu Jimmy","year":"2023","unstructured":"Jimmy Wu, Rika Antonova, Adam Kan, Marion Lepert, Andy Zeng, Shuran Song, Jeannette Bohg, Szymon Rusinkiewicz, and Thomas Funkhouser. 2023. TidyBot: Personalized robot assistance with large language models. Autonomous Robots (2023).","journal-title":"Autonomous Robots"},{"key":"e_1_3_2_67_1","unstructured":"Sriram Yenamandra Arun Ramachandran Karmesh Yadav Austin Wang Mukul Khanna Theophile Gervet Tsung-Yen Yang Vidhi Jain Alex William Clegg John Turner et\u00a0al. 2023. HomeRobot: Open Vocab Mobile Manipulation. Retrieved March 2026 from https:\/\/aihabitat.org\/static\/challenge\/home_robot_ovmm_2023\/OVMM.pdf"},{"key":"e_1_3_2_68_1","unstructured":"Ceng Zhang Xin Meng Dongchen Qi and Gregory S. Chirikjian. 2024. RAIL: Robot affordance imagination with large language models. arXiv:2403.19369. Retrieved from https:\/\/arxiv.org\/abs\/2403.19369"},{"key":"e_1_3_2_69_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00645"},{"key":"e_1_3_2_70_1","first-page":"4534","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Jiazhao","year":"2020","unstructured":"Jiazhao Zhang, Chenyang Zhu, Lintao Zheng, and Kai Xu. 2020. Fusion-aware point convolution for online semantic 3D scene segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 4534\u20134543."},{"key":"e_1_3_2_71_1","first-page":"3823","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhu Yixin","year":"2016","unstructured":"Yixin Zhu, Chenfanfu Jiang, Yibiao Zhao, Demetri Terzopoulos, and Song-Chun Zhu. 2016. Inferring forces and learning human utilities from videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3823\u20133833."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3795693","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T10:35:35Z","timestamp":1773484535000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3795693"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,14]]},"references-count":70,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3795693"],"URL":"https:\/\/doi.org\/10.1145\/3795693","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,14]]},"assertion":[{"value":"2025-08-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-26","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}