{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T06:59:35Z","timestamp":1784185175161,"version":"3.55.0"},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,9]]},"abstract":"<jats:p>3D spatial understanding is essential in real-world applications such as robotics, autonomous vehicles, virtual reality, and medical imaging. Recently, Large Language Models (LLMs), having demonstrated remarkable success across various domains, have been leveraged to enhance 3D understanding tasks, showing potential to surpass traditional computer vision methods. In this survey, we present a comprehensive review of methods integrating LLMs with 3D spatial understanding. We propose a taxonomy that categorizes existing methods into three branches: image-based methods deriving 3D understanding from 2D visual data, point cloud-based methods working directly with 3D representations, and hybrid modality-based methods combining multiple data streams. We systematically review representative methods along these categories, covering data representations, architectural modifications, and training strategies that bridge textual and 3D modalities. Finally, we discuss current limitations, including dataset scarcity and computational challenges, while highlighting promising research directions in spatial perception, multi-modal fusion, and real-world applications.<\/jats:p>","DOI":"10.24963\/ijcai.2025\/1200","type":"proceedings-article","created":{"date-parts":[[2025,9,19]],"date-time":"2025-09-19T08:10:40Z","timestamp":1758269440000},"page":"10817-10825","source":"Crossref","is-referenced-by-count":10,"title":["How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM"],"prefix":"10.24963","author":[{"given":"Jirong","family":"Zha","sequence":"first","affiliation":[{"name":"Shenzhen International Graduate School, Tsinghua University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuxuan","family":"Fan","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology (Guang Zhou)"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiao","family":"Yang","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology (Guang Zhou)"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chen","family":"Gao","sequence":"additional","affiliation":[{"name":"BNRist, Tsinghua University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinlei","family":"Chen","sequence":"additional","affiliation":[{"name":"Shenzhen International Graduate School, Tsinghua University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"10584","event":{"name":"Thirty-Fourth International Joint Conference on Artificial Intelligence {IJCAI-25}","theme":"Artificial Intelligence","location":"Montreal, Canada","acronym":"IJCAI-2025","number":"34","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)"],"start":{"date-parts":[[2025,8,16]]},"end":{"date-parts":[[2025,8,22]]}},"container-title":["Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2025,9,23]],"date-time":"2025-09-23T11:36:26Z","timestamp":1758627386000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2025\/1200"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2025,9]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2025\/1200","relation":{},"subject":[],"published":{"date-parts":[[2025,9]]}}}