{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T07:48:16Z","timestamp":1777880896454,"version":"3.51.4"},"reference-count":89,"publisher":"Elsevier BV","license":[{"start":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T00:00:00Z","timestamp":1775001600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.elsevier.com\/tdm\/userlicense\/1.0\/"},{"start":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T00:00:00Z","timestamp":1775001600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.elsevier.com\/legal\/tdmrep-license"},{"start":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T00:00:00Z","timestamp":1775001600000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-017"},{"start":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T00:00:00Z","timestamp":1775001600000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-037"},{"start":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T00:00:00Z","timestamp":1775001600000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-012"},{"start":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T00:00:00Z","timestamp":1775001600000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-029"},{"start":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T00:00:00Z","timestamp":1775001600000},"content-version":"stm-asf","delay-in-days":0,"URL":"https:\/\/doi.org\/10.15223\/policy-004"}],"content-domain":{"domain":["elsevier.com","sciencedirect.com"],"crossmark-restriction":true},"short-container-title":["Engineering Applications of Artificial Intelligence"],"published-print":{"date-parts":[[2026,4]]},"DOI":"10.1016\/j.engappai.2026.114150","type":"journal-article","created":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T08:12:11Z","timestamp":1770883931000},"page":"114150","update-policy":"https:\/\/doi.org\/10.1016\/elsevier_cm_policy","source":"Crossref","is-referenced-by-count":1,"special_numbering":"C","title":["Scene graph-driven reasoning for action planning of humanoid robot"],"prefix":"10.1016","volume":"169","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1407-2633","authenticated-orcid":false,"given":"Dmitry","family":"Yudin","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-5830-9657","authenticated-orcid":false,"given":"Alexander","family":"Lazarev","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6984-188X","authenticated-orcid":false,"given":"Eva","family":"Bakaeva","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Angelika","family":"Kochetkova","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2180-0990","authenticated-orcid":false,"given":"Alexey","family":"Kovalev","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9747-3837","authenticated-orcid":false,"given":"Aleksandr","family":"Panov","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"78","reference":[{"key":"10.1016\/j.engappai.2026.114150_b1","series-title":"Do as I can and not as I say: Grounding language in robotic affordances","author":"Ahn","year":"2022"},{"key":"10.1016\/j.engappai.2026.114150_b2","series-title":"Do as i can, not as i say: Grounding language in robotic affordances","author":"Ahn","year":"2022"},{"key":"10.1016\/j.engappai.2026.114150_b3","series-title":"Relational inductive biases, deep learning, and graph networks","author":"Battaglia","year":"2018"},{"key":"10.1016\/j.engappai.2026.114150_b4","series-title":"International Conference on Learning Representations","article-title":"Depth pro: Sharp monocular metric depth in less than a second","author":"Bochkovskii","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b5","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"10.1016\/j.engappai.2026.114150_b6","series-title":"2023 21st International Conference on Advanced Robotics","first-page":"206","article-title":"Benchmarking the full-order model optimization based imitation in the humanoid robot reinforcement learning walk","author":"Chaikovskaya","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b7","series-title":"Masked-attention mask transformer for universal image segmentation","author":"Cheng","year":"2022"},{"key":"10.1016\/j.engappai.2026.114150_b8","series-title":"Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality","author":"Chiang","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b9","series-title":"Time is on my sight: scene graph filtering for dynamic environment perception in an LLM-driven robot","author":"Colombani","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b10","series-title":"Computer vision annotation tool (CVAT)","author":"CVAT.ai Corporation","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b11","series-title":"IEEE International Conference on Robotics and Automation","article-title":"Optimal scene graph planning with large language model guidance","author":"Dai","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b12","series-title":"PaLM-E: An embodied multimodal language model","author":"Driess","year":"2023"},{"issue":"3\u20134","key":"10.1016\/j.engappai.2026.114150_b13","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1016\/0004-3702(71)90010-5","article-title":"STRIPS: A new approach to the application of theorem proving to problem solving","volume":"2","author":"Fikes","year":"1971","journal-title":"Artificial Intelligence"},{"key":"10.1016\/j.engappai.2026.114150_b14","doi-asserted-by":"crossref","first-page":"955","DOI":"10.1016\/j.robot.2008.08.007","article-title":"Robot task planning using semantic maps","volume":"56","author":"Galindo","year":"2008","journal-title":"Robot. Auton. Syst."},{"key":"10.1016\/j.engappai.2026.114150_b15","doi-asserted-by":"crossref","unstructured":"Galindo, C., Saffiotti, A., Coradeschi, S., Buschka, P., Fern\u00e1ndez-Madrigal, J.-A., Gonz\u00e1lez, J., 2005. Multi-hierarchical semantic maps for mobile robotics. In: Proceedings 2005 IEEE\/RSJ International Conference on Intelligent Robots and Systems. pp. 2278\u20132283.","DOI":"10.1109\/IROS.2005.1545511"},{"key":"10.1016\/j.engappai.2026.114150_b16","doi-asserted-by":"crossref","first-page":"2280","DOI":"10.1016\/j.patcog.2014.01.005","article-title":"Automatic generation and detection of highly reliable fiducial markers under occlusion","volume":"47","author":"Garrido-Jurado","year":"2014","journal-title":"Pattern Recognit."},{"key":"10.1016\/j.engappai.2026.114150_b17","series-title":"2024 IEEE\/RSJ International Conference on Intelligent Robots and Systems","first-page":"13318","article-title":"Commonsense scene graph-based target localization for object search","author":"Ge","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b18","series-title":"PDDL\u2014The Planning Domain Definition Language","author":"Ghallab","year":"1998"},{"key":"10.1016\/j.engappai.2026.114150_b19","doi-asserted-by":"crossref","unstructured":"Giuliari, F., Skenderi, G., Cristani, M., Wang, Y., Del Bue, A., 2022. Spatial commonsense graph for object localisation in partial scenes. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. pp. 19518\u201319527.","DOI":"10.1109\/CVPR52688.2022.01891"},{"key":"10.1016\/j.engappai.2026.114150_b20","series-title":"2023 IEEE\/RSJ International Conference on Intelligent Robots and Systems","first-page":"3568","article-title":"Generating executable action plans with environmentally-aware language models","author":"Gramopadhye","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b21","series-title":"International Conference on Hybrid Artificial Intelligence Systems","first-page":"224","article-title":"Common sense plan verification with large language models","author":"Grigorev","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b22","series-title":"2025 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","first-page":"18489","article-title":"Verifyllm: Llm-based pre-execution task plan verification for robots","author":"Grigorev","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b23","series-title":"Mask R-CNN","author":"He","year":"2018"},{"key":"10.1016\/j.engappai.2026.114150_b24","series-title":"International Conference on Machine Learning","first-page":"9118","article-title":"Language models as zero-shot planners: Extracting actionable knowledge for embodied agents","author":"Huang","year":"2022"},{"key":"10.1016\/j.engappai.2026.114150_b25","series-title":"Inner monologue: Embodied reasoning through planning with language models","author":"Huang","year":"2022"},{"key":"10.1016\/j.engappai.2026.114150_b26","series-title":"Foundations of spatial perception for robotics: Hierarchical representations and real-time systems","author":"Hughes","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b27","doi-asserted-by":"crossref","unstructured":"Ivanova, A., Eva, B., Volovikova, Z., Kovalev, A., Panov, A., 2025. Ambik: Dataset of ambiguous tasks in kitchen environment. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 33216\u201333241.","DOI":"10.18653\/v1\/2025.acl-long.1593"},{"key":"10.1016\/j.engappai.2026.114150_b28","series-title":"OneFormer: One Transformer to Rule Universal Image Segmentation","author":"Jain","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b29","series-title":"Mistral 7B","author":"Jiang","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b30","series-title":"YOLOv8: Ultralytics next-generation object detection model","author":"Jocher","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b31","series-title":"2015 IEEE Conference on Computer Vision and Pattern Recognition","first-page":"3668","article-title":"Image retrieval using scene graphs","author":"Johnson","year":"2015"},{"key":"10.1016\/j.engappai.2026.114150_b32","series-title":"Challenges and applications of large language models","author":"Kaddour","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b33","series-title":"YOLOv11: An Overview of the Key Architectural Enhancements","author":"Khanam","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b34","series-title":"A review of YOLOv12: Attention-based enhancements vs. Previous versions","author":"Khanam","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b35","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1016\/j.robot.2014.12.006","article-title":"Semantic mapping for mobile robotics tasks: A survey","volume":"66","author":"Kostavelis","year":"2015","journal-title":"Robot. Auton. Syst."},{"key":"10.1016\/j.engappai.2026.114150_b36","series-title":"Dokl. Math.","first-page":"S85","article-title":"Application of pretrained large language models in embodied artificial intelligence","author":"Kovalev","year":"2022"},{"key":"10.1016\/j.engappai.2026.114150_b37","doi-asserted-by":"crossref","DOI":"10.1016\/j.cogsys.2021.09.001","article-title":"Vector Semiotic Model for Visual Question Answering","volume":"71","author":"Kovalev","year":"2022","journal-title":"Cogn. Syst. Res."},{"key":"10.1016\/j.engappai.2026.114150_b38","doi-asserted-by":"crossref","unstructured":"Li, Y., Ouyang, W., Zhou, B., Wang, K., Wang, X., 2017. Scene graph generation from objects, phrases and region captions. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 1261\u20131270.","DOI":"10.1109\/ICCV.2017.142"},{"key":"10.1016\/j.engappai.2026.114150_b39","series-title":"Seeing beyond the scene: Enhancing vision-language models with interactional reasoning","author":"Liang","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b40","doi-asserted-by":"crossref","unstructured":"Lin, B.Y., Huang, C., Liu, Q., Gu, W., Sommerer, S., Ren, X., 2023. On grounded planning for embodied tasks with language models. In: Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 37, pp. 13192\u201313200.","DOI":"10.1609\/aaai.v37i11.26549"},{"key":"10.1016\/j.engappai.2026.114150_b41","doi-asserted-by":"crossref","unstructured":"Linok, S., Zemskova, T., Ladanova, S., Titkov, R., Yudin, D., Monastyrny, M., Valenkov, A., 2025. Beyond bare queries: Open-vocabulary object grounding with 3d scene graph. In: 2025 IEEE International Conference on Robotics and Automation (ICRA). pp. 13582\u201313589.","DOI":"10.1109\/ICRA55743.2025.11128059"},{"key":"10.1016\/j.engappai.2026.114150_b42","series-title":"Llm+ p: Empowering large language models with optimal planning proficiency","author":"Liu","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b43","series-title":"Delta: Decomposed efficient long-term robot task planning using large language models","author":"Liu","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b44","series-title":"SG-reg: Generalizable and efficient scene graph registration","author":"Liu","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b45","series-title":"Few-shot subgoal planning with language models","author":"Logeswaran","year":"2022"},{"key":"10.1016\/j.engappai.2026.114150_b46","doi-asserted-by":"crossref","unstructured":"Lucignano, L., Cutugno, F., Rossi, S., Finzi, A., 2013. A dialogue system for multimodal human-robot interaction. In: Proceedings of the 15th ACM on International Conference on Multimodal Interaction. pp. 197\u2013204.","DOI":"10.1145\/2522848.2522873"},{"key":"10.1016\/j.engappai.2026.114150_b47","series-title":"CoRL 2023 Workshop on Learning Effective Abstractions for Planning","article-title":"Obtaining hierarchy from human instructions: an LLMs-based approach","author":"Luo","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b48","series-title":"Large language models: A survey","author":"Minaee","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b49","series-title":"SLC2-SLAM: Semantic-guided loop closure using shared latent code for NeRF SLAM","author":"Ming","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b50","doi-asserted-by":"crossref","unstructured":"Moravec, H.P., Elfes, A., 1985. High resolution maps from wide angle sonar. In: Proceedings. 1985 IEEE International Conference on Robotics and Automation. pp. 116\u2013121.","DOI":"10.1109\/ROBOT.1985.1087316"},{"key":"10.1016\/j.engappai.2026.114150_b51","series-title":"A comprehensive overview of large language models","author":"Naveed","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b52","series-title":"Toward a graph-theoretic model of belief: Confidence, credibility, and structural coherence","author":"Nikooroo","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b53","series-title":"2025 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","first-page":"19218","article-title":"Lera: Replanning with visual feedback in instruction following","author":"Pchelintsev","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b54","series-title":"Proceedings of the Fifteenth Conference on Uncertainty in Artificial Intelligence","first-page":"531","article-title":"Graphical representations of consensus belief","author":"Pennock","year":"1999"},{"key":"10.1016\/j.engappai.2026.114150_b55","series-title":"Leveraging LLMs, graphs and object hierarchies for task planning in large-scale environments","author":"P\u00e9rez-Dattari","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b56","series-title":"ESGNN: Towards equivariant scene graph neural network for 3D scene understanding","author":"Pham","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b57","series-title":"TESGNN: Temporal equivariant scene graph neural networks for efficient and robust multi-view 3D scene understanding","author":"Pham","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b58","series-title":"Adapt: As-needed decomposition and planning with language models","author":"Prasad","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b59","series-title":"7th Annual Conference on Robot Learning","article-title":"SayPlan: Grounding large language models using 3D scene graphs for scalable task planning","author":"Rana","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b60","doi-asserted-by":"crossref","unstructured":"Roberts, M., Ramapuram, J., Ranjan, A., Kumar, A., Bautista, M.A., Paczan, N., Webb, R., Susskind, J.M., 2021. Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding. In: International Conference on Computer Vision (ICCV) 2021.","DOI":"10.1109\/ICCV48922.2021.01073"},{"key":"10.1016\/j.engappai.2026.114150_b61","series-title":"Artificial General Intelligence","first-page":"222","article-title":"Evaluation of pretrained large language models in embodied planning tasks","author":"Sarkisyan","year":"2023"},{"issue":"1","key":"10.1016\/j.engappai.2026.114150_b62","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1109\/TNN.2008.2005605","article-title":"The graph neural network model","volume":"20","author":"Scarselli","year":"2009","journal-title":"IEEE Trans. Neural Netw."},{"key":"10.1016\/j.engappai.2026.114150_b63","series-title":"NeurIPS 2022 Foundation Models for Decision Making Workshop","article-title":"PDDL planning with pretrained large language models","author":"Silver","year":"2022"},{"key":"10.1016\/j.engappai.2026.114150_b64","series-title":"2023 IEEE International Conference on Robotics and Automation","first-page":"11523","article-title":"ProgPrompt: Generating situated robot task plans using large language models","author":"Singh","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b65","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1146\/annurev-control-101119-071628","article-title":"Robots that use language","volume":"3","author":"Tellex","year":"2020","journal-title":"Annu. Rev. Control. Robot. Auton. Syst."},{"key":"10.1016\/j.engappai.2026.114150_b66","series-title":"Probabilistic Robotics","author":"Thrun","year":"2002"},{"key":"10.1016\/j.engappai.2026.114150_b67","doi-asserted-by":"crossref","unstructured":"Wald, J., Avetisyan, A., Navab, N., Tombari, F., Niessner, M., 2019. RIO: 3D Object Instance Re-Localization in Changing Indoor Environments. In: Proceedings IEEE International Conference on Computer Vision. ICCV.","DOI":"10.1109\/ICCV.2019.00775"},{"key":"10.1016\/j.engappai.2026.114150_b68","doi-asserted-by":"crossref","unstructured":"Wald, J., Dhamo, H., Navab, N., Tombari, F., 2020. Learning 3d semantic scene graphs from 3d indoor reconstructions. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. pp. 3961\u20133970.","DOI":"10.1109\/CVPR42600.2020.00402"},{"key":"10.1016\/j.engappai.2026.114150_b69","doi-asserted-by":"crossref","unstructured":"Wang, Z., Cheng, B., Zhao, L., Xu, D., Tang, Y., Sheng, L., 2024. VL-SAT: Visual-Linguistic Semantics Assisted Training for 3D Semantic Scene Graph Prediction in Point Cloud. In: Proc. IEEE\/CVF Conf. Comput. Vis. Pattern Recognit.. CVPR.","DOI":"10.1109\/CVPR52729.2023.02065"},{"key":"10.1016\/j.engappai.2026.114150_b70","series-title":"YOLOE: Real-time seeing anything","author":"Wang","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b71","doi-asserted-by":"crossref","unstructured":"Wang, R., Xu, S., Dai, C., Xiang, J., Deng, Y., Tong, X., Yang, J., 2025b. Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 5261\u20135271.","DOI":"10.1109\/CVPR52734.2025.00496"},{"key":"10.1016\/j.engappai.2026.114150_b72","series-title":"Moge-2: Accurate monocular geometry with metric scale and sharp details","author":"Wang","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b73","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"10.1016\/j.engappai.2026.114150_b74","series-title":"Incremental 3D semantic scene graph prediction from RGB sequences","author":"Wu","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b75","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1016\/j.procir.2025.03.007","article-title":"Large language model-guided graph convolution network reasoning system for complex human-robot collaboration disassembly operations","volume":"134","author":"Xiao","year":"2025","journal-title":"Procedia CIRP"},{"key":"10.1016\/j.engappai.2026.114150_b76","doi-asserted-by":"crossref","first-page":"937","DOI":"10.1016\/j.jmsy.2025.11.012","article-title":"Intelligent disassembly scenario understanding for human behavior and intention recognition towards self-perception human-robot collaboration system","volume":"83","author":"Xiao","year":"2025","journal-title":"J. Manuf. Syst."},{"key":"10.1016\/j.engappai.2026.114150_b77","series-title":"Translating natural language to planning goals with large-language models","author":"Xie","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b78","doi-asserted-by":"crossref","DOI":"10.1109\/TKDE.2025.3536008","article-title":"Are large language models really good logical reasoners? a comprehensive evaluation and beyond","author":"Xu","year":"2025","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"10.1016\/j.engappai.2026.114150_b79","doi-asserted-by":"crossref","unstructured":"Xu, D., Zhu, Y., Choy, C.B., Fei-Fei, L., 2017. Scene graph generation by iterative message passing. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 5410\u20135419.","DOI":"10.1109\/CVPR.2017.330"},{"key":"10.1016\/j.engappai.2026.114150_b80","series-title":"CVPR","article-title":"Depth anything: Unleashing the power of large-scale unlabeled data","author":"Yang","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b81","series-title":"Depth anything V2","author":"Yang","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b82","series-title":"LLM meets scene graph: Can large language models understand and generate scene graphs? A benchmark and empirical study","first-page":"21335","author":"Yang","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b83","series-title":"SAM3D: Segment anything in 3D scenes","author":"Yang","year":"2023"},{"key":"10.1016\/j.engappai.2026.114150_b84","series-title":"Qwen2. 5 technical report","author":"Yang","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b85","doi-asserted-by":"crossref","first-page":"130169","DOI":"10.1016\/j.neucom.2025.130169","article-title":"SegmATRon: Embodied adaptive semantic segmentation for indoor environment","volume":"638","author":"Zemskova","year":"2025","journal-title":"Neurocomputing"},{"key":"10.1016\/j.engappai.2026.114150_b86","unstructured":"Zemskova, T., Yudin, D., 2025. 3DGraphLLM: Combining semantic graphs and large language models for 3d scene understanding. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision. pp. 8885\u20138895."},{"key":"10.1016\/j.engappai.2026.114150_b87","series-title":"HI-SLAM2: Geometry-aware Gaussian SLAM for fast monocular scene reconstruction","author":"Zhang","year":"2025"},{"key":"10.1016\/j.engappai.2026.114150_b88","series-title":"The Thirty-Eighth Annual Conference on Neural Information Processing Systems","article-title":"Multiview scene graph","author":"Zhang","year":"2024"},{"key":"10.1016\/j.engappai.2026.114150_b89","doi-asserted-by":"crossref","unstructured":"Zong, Z., Song, G., Liu, Y., 2023. Detrs with collaborative hybrid assignments training. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision. pp. 6748\u20136758.","DOI":"10.1109\/ICCV51070.2023.00621"}],"container-title":["Engineering Applications of Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/api.elsevier.com\/content\/article\/PII:S0952197626004318?httpAccept=text\/xml","content-type":"text\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/api.elsevier.com\/content\/article\/PII:S0952197626004318?httpAccept=text\/plain","content-type":"text\/plain","content-version":"vor","intended-application":"text-mining"}],"deposited":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T00:25:56Z","timestamp":1777595156000},"score":1,"resource":{"primary":{"URL":"https:\/\/linkinghub.elsevier.com\/retrieve\/pii\/S0952197626004318"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4]]},"references-count":89,"alternative-id":["S0952197626004318"],"URL":"https:\/\/doi.org\/10.1016\/j.engappai.2026.114150","relation":{},"ISSN":["0952-1976"],"issn-type":[{"value":"0952-1976","type":"print"}],"subject":[],"published":{"date-parts":[[2026,4]]},"assertion":[{"value":"Elsevier","name":"publisher","label":"This article is maintained by"},{"value":"Scene graph-driven reasoning for action planning of humanoid robot","name":"articletitle","label":"Article Title"},{"value":"Engineering Applications of Artificial Intelligence","name":"journaltitle","label":"Journal Title"},{"value":"https:\/\/doi.org\/10.1016\/j.engappai.2026.114150","name":"articlelink","label":"CrossRef DOI link to publisher maintained version"},{"value":"article","name":"content_type","label":"Content Type"},{"value":"\u00a9 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.","name":"copyright","label":"Copyright"}],"article-number":"114150"}}