{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T08:51:32Z","timestamp":1777107092117,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":18,"publisher":"ACM","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,11,14]]},"DOI":"10.1145\/3787279.3787314","type":"proceedings-article","created":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T07:38:47Z","timestamp":1777102727000},"page":"214-219","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["A Vision-Language-Action Framework for End-to-End Robotic Manipulation Using Qwen2-VL-Instruct"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0214-2980","authenticated-orcid":false,"given":"Mariam","family":"Kashkash","sequence":"first","affiliation":[{"name":"Machine Learning Department, Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, Abu Dhabi, United Arab Emirates"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8972-8094","authenticated-orcid":false,"given":"Mohsen","family":"Guizani","sequence":"additional","affiliation":[{"name":"Machine Learning Department, MBZUAI, Abu Dhabi, Abu Dhabi, United Arab Emirates"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,25]]},"reference":[{"key":"e_1_3_3_1_2_2","unstructured":"Michael Ahn Anthony Brohan Noah Brown and et al.2022. Do As I Can Not As I Say: Grounding Language in Robotic Affordances."},{"key":"e_1_3_3_1_3_2","unstructured":"C\u00e9dric Colas Tristan Karch Nicolas Lair Jean-Michel Dussoux Peter\u00a0Ford Dominey and Pierre-Yves Oudeyer. 2020. Language as a Cognitive Tool to Imagine Goals in Curiosity-Driven Exploration. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2002.09253 (2020). https:\/\/arxiv.org\/abs\/2002.09253"},{"key":"e_1_3_3_1_4_2","doi-asserted-by":"publisher","unstructured":"Yao Cong and Hongwei Mo. 2024. An Overview of Robot Embodied Intelligence Based on Multimodal Models: Tasks Models and System Schemes. International Journal of Intelligent Systems 2025 1 (2024) 5124400. 10.1155\/int\/5124400Accessed June 28 2025.","DOI":"10.1155\/int\/5124400"},{"key":"e_1_3_3_1_5_2","unstructured":"Ricardo Garcia Shizhe Chen and Cordelia Schmid. 2024. Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy. arXiv (2024). arxiv:https:\/\/arXiv.org\/abs\/2410.01345\u00a0[cs.RO] https:\/\/arxiv.org\/abs\/2410.01345 Accessed February 22 2025."},{"key":"e_1_3_3_1_6_2","unstructured":"Rishi Hazra Pedro Zuidberg Dos\u00a0Martires and Luc De\u00a0Raedt. 2024. SayCanPay: Heuristic Planning with Large Language Models using Learnable Domain Knowledge. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2308.12682 (2024). https:\/\/arxiv.org\/abs\/2308.12682 Accessed May 24 2025."},{"key":"e_1_3_3_1_7_2","unstructured":"Hengyuan Hu and Dorsa Sadigh. 2023. Language Instructed Reinforcement Learning for Human-AI Coordination. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2304.07297 (2023). https:\/\/arxiv.org\/abs\/2304.07297"},{"key":"e_1_3_3_1_8_2","doi-asserted-by":"publisher","unstructured":"Hyeongyo Jeong Haechan Lee Changwon Kim and Sungtae Shin. 2024. A Survey of Robot Intelligence with Large Language Models. Applied Sciences 14 19 (2024). 10.3390\/app14198868","DOI":"10.3390\/app14198868"},{"key":"e_1_3_3_1_9_2","series-title":"Proceedings of Machine Learning Research","first-page":"557","volume-title":"Proceedings of the 5th Conference on Robot Learning","volume":"164","author":"Kalashnikov Dmitry","year":"2022","unstructured":"Dmitry Kalashnikov, Jake Varley, Yevgen Chebotar, Benjamin Swanson, Rico Jonschkowski, Chelsea Finn, Sergey Levine, and Karol Hausman. 2022. Scaling Up Multi-Task Robotic Reinforcement Learning. In Proceedings of the 5th Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a0164), Aleksandra Faust, David Hsu, and Gerhard Neumann (Eds.). PMLR, 557\u2013575. https:\/\/proceedings.mlr.press\/v164\/kalashnikov22a.html"},{"key":"e_1_3_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICMCR64890.2025.10963286"},{"key":"e_1_3_3_1_11_2","doi-asserted-by":"publisher","unstructured":"Shufei Li Pai Zheng Sichao Liu Zuoxu Wang Xi\u00a0Vincent Wang Lianyu Zheng and Lihui Wang. 2023. Proactive human\u2013robot collaboration: Mutual-cognitive predictable and self-organising perspectives. Robotics and Computer-Integrated Manufacturing 81 (2023) 102510. 10.1016\/j.rcim.2022.102510","DOI":"10.1016\/j.rcim.2022.102510"},{"key":"e_1_3_3_1_12_2","unstructured":"Bo Liu Yuqian Jiang Xiaohan Zhang Qiang Liu Shiqi Zhang Joydeep Biswas and Peter Stone. 2023. LLM+P: Empowering Large Language Models with Optimal Planning Proficiency. arXiv (2023). arxiv:https:\/\/arXiv.org\/abs\/2304.11477\u00a0[cs.AI] https:\/\/arxiv.org\/abs\/2304.11477 Accessed February 22 2025."},{"key":"e_1_3_3_1_13_2","doi-asserted-by":"publisher","unstructured":"Alessandro Palleschi George\u00a0Jose Pollayil Mathew\u00a0Jose Pollayil Manolo Garabini and Lucia Pallottino. 2022. High-Level Planning for Object Manipulation With Multi Heterogeneous Robots in Shared Environments. IEEE Robotics and Automation Letters 7 2 (2022) 3138\u20133145. 10.1109\/LRA.2022.3145987","DOI":"10.1109\/LRA.2022.3145987"},{"key":"e_1_3_3_1_14_2","unstructured":"Mohit Shridhar Lucas Manuelli and Dieter Fox. 2021. CLIPort: What and Where Pathways for Robotic Manipulation. (9 2021). http:\/\/arxiv.org\/abs\/2109.12098"},{"key":"e_1_3_3_1_15_2","unstructured":"Harsh Singh Rocktim\u00a0Jyoti Das Mingfei Han Preslav Nakov and Ivan Laptev. 2024. MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2411.17636 (2024). https:\/\/malmm1.github.io\/"},{"key":"e_1_3_3_1_16_2","unstructured":"Shu Wang Muzhi Han Ziyuan Jiao Zeyu Zhang Ying\u00a0N. Wu Song Zhu and Hangxin Liu. 2024. LLM3: Large Language Model-based Task and Motion Planning with Motion Failure Reasoning. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2403.11552 (2024). https:\/\/arxiv.org\/abs\/2403.11552 Accessed May 24 2025."},{"key":"e_1_3_3_1_17_2","unstructured":"Micha\u0142 Zawalski William Chen Karl Pertsch Oier Mees Chelsea Finn and Sergey Levine. 2024. Robotic Control via Embodied Chain-of-Thought Reasoning. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2407.08693 (2024). https:\/\/arxiv.org\/abs\/2407.08693 Accessed May 24 2025."},{"key":"e_1_3_3_1_18_2","volume-title":"Conference on Robot Learning (CoRL)","author":"Zeng Andy","year":"2020","unstructured":"Andy Zeng, Pete Florence, Jonathan Tompson, Stefan Welker, Jonathan Chien, Maria Attarian, Travis Armstrong, Ivan Krasin, Dan Duong, Vikas Sindhwani, and Johnny Lee. 2020. Transporter Networks: Rearranging the Visual World for Robotic Manipulation. In Conference on Robot Learning (CoRL) (Cambridge, MA, USA). PMLR."},{"key":"e_1_3_3_1_19_2","series-title":"Proceedings of Machine Learning Research","first-page":"2165","volume-title":"Proceedings of The 7th Conference on Robot Learning","volume":"229","author":"Zitkovich Brianna","year":"2023","unstructured":"Brianna Zitkovich, Tianhe Yu, Sichun Xu, and et al. 2023. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. In Proceedings of The 7th Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a0229), Jie Tan, Marc Toussaint, and Kourosh Darvish (Eds.). PMLR, 2165\u20132183. https:\/\/proceedings.mlr.press\/v229\/zitkovich23a.html"}],"event":{"name":"ICAAI 2025: 2025 9th International Conference on Advances in Artificial Intelligence","location":"Manchester United Kingdom","acronym":"ICAAI 2025"},"container-title":["Proceedings of the 2025 9th International Conference on Advances in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3787279.3787314","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T08:23:28Z","timestamp":1777105408000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3787279.3787314"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,14]]},"references-count":18,"alternative-id":["10.1145\/3787279.3787314","10.1145\/3787279"],"URL":"https:\/\/doi.org\/10.1145\/3787279.3787314","relation":{},"subject":[],"published":{"date-parts":[[2025,11,14]]},"assertion":[{"value":"2026-04-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}