{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T18:29:09Z","timestamp":1777487349476,"version":"3.51.4"},"reference-count":34,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/100000006","name":"Office of Naval Research","doi-asserted-by":"crossref","award":["N00014-24-1-2024"],"award-info":[{"award-number":["N00014-24-1-2024"]}],"id":[{"id":"10.13039\/100000006","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000181","name":"Air Force Office of Scientific Research","doi-asserted-by":"crossref","award":["FA9550-18-1-0465"],"award-info":[{"award-number":["FA9550-18-1-0465"]}],"id":[{"id":"10.13039\/100000181","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2026,8,31]]},"abstract":"<jats:p>\n                    LLMs are being increasingly integrated into embodied robotic systems. A useful capability that the LLMs bring to robots is translating noisy spoken human natural language instructions into executable robot actions. However, these integrations are somewhat\n                    <jats:italic toggle=\"yes\">ad hoc<\/jats:italic>\n                    and understudied as they tend to not consider the gamut of syntactic, semantic, as well as pragmatic aspects of embodied human communication. What is missing is a characterization of the different paradigms for integrating LLMs into robotic architectures as well as a set of evaluation metrics that capture whether an LLM-equipped robot can correctly understand these different aspects of human instruction. In this article, we present a suite of evaluation metrics together with data augmentation techniques for evaluating these architectures, using concepts from the cognitive science and human communication literature. To illustrate an application of these metrics and augmentation techniques, we conduct experiments to compare two integration methods: LLMs as pre-processing components that map human instructions into more constrained versions to be processed by the architecture\u2019s natural language understanding (NLU) subsystem, or LLMs as a wholesale replacement for the NLU\u2019s parser. We provide experimental evaluations and a robotic implementation to show the inherent tradeoffs between the methods. Our results suggest that while they offer increased explainability, traditional parsing tools coupled with LLMs do not perform as well as an LLM that replaces a parser entirely. The proposed evaluation metrics together with the characterization of different LLM integration approaches offer the promise of systematically evaluating LLMs as natural language interfaces to robotic systems as well as tackle the important tradeoff between explainability\/verifiability\/interpretability and robustness to noisy input and broad language understanding in an open-world embodied setting.\n                  <\/jats:p>","DOI":"10.1145\/3754340","type":"journal-article","created":{"date-parts":[[2025,8,25]],"date-time":"2025-08-25T13:50:53Z","timestamp":1756129853000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["On Evaluating LLM Integration into Robotic Architectures"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4129-6429","authenticated-orcid":false,"given":"Vasanth","family":"Sarathy","sequence":"first","affiliation":[{"name":"Department of Computer Science, Tufts University, Medford, Massachusetts, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-2142-6665","authenticated-orcid":false,"given":"Marlow","family":"Fawn","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Tufts University, Medford, Massachusetts, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-6319-1384","authenticated-orcid":false,"given":"Matthew","family":"McWilliams","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Tufts University, Medford, Massachusetts, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0064-2789","authenticated-orcid":false,"given":"Matthias","family":"Scheutz","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Tufts University, Medford, Massachusetts, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1053-8248","authenticated-orcid":false,"given":"Bradley","family":"Oosterveld","sequence":"additional","affiliation":[{"name":"Thinking Robots, Inc., Boston, Massachusetts, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,28]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","unstructured":"Divyanshu Aggarwal Ashutosh Sathe and Sunayana Sitaram. 2024. Maple: Multilingual evaluation of parameter efficient finetuning of large language models. arXiv:2401.07598. Retrieved from https:\/\/arxiv.org\/abs\/2401.07598","DOI":"10.18653\/v1\/2024.findings-acl.881"},{"key":"e_1_3_2_3_2","unstructured":"Michael Ahn Anthony Brohan Noah Brown Yevgen Chebotar Omar Cortes Byron David Chelsea Finn Chuyuan Fu Keerthana Gopalakrishnan Karol Hausman et al. 2022. Do As I Can Not As I Say: Grounding Language in Robotic Affordances. arXiv:2204.01691. Retrieved from https:\/\/arxiv.org\/abs\/2204.01691"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1002\/wcs.1234"},{"key":"e_1_3_2_5_2","first-page":"35","volume-title":"Proceedings of the 9th International Conference on Computational Semantics","author":"Baral Chitta","year":"2011","unstructured":"Chitta Baral, Juraj Dzifcak, Marcos Alvarez Gonzalez, and Jiayu Zhou. 2011. Using inverse \\(\\lambda\\) and generalization to translate English to formal languages. In Proceedings of the 9th International Conference on Computational Semantics. Association for Computational Linguistics, 35\u201344."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2018\/1"},{"key":"e_1_3_2_7_2","unstructured":"Shuaichen Chang Jun Wang Mingwen Dong Lin Pan Henghui Zhu Alexander Hanbo Li Wuwei Lan Sheng Zhang Jiarong Jiang Joseph Lilien et al. 2023. Dr.Spider: A diagnostic evaluation bench- mark towards text-to-sql robustness. arXiv:2301.08881. Retrieved from https:\/\/arxiv.org\/abs\/2301.08881"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1162\/coli.2007.33.4.493"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","unstructured":"Tim Dettmers Artidoro Pagnoni Ari Holtzman and Luke Zettlemoyer. 2023. QLoRA: Efficient finetuning of quantized LLMs. arXiv:2305.14314. Retrieved from https:\/\/arxiv.org\/abs\/2305.14314","DOI":"10.52202\/075280-0441"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ROBOT.2009.5152776"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","unstructured":"Catherine Finegan-Dollak Jonathan K. Kummerfeld Li Zhang Karthik Ramanathan Sesh Sadasivam Rui Zhang and Dragomir Radev. 2018. Improving Text-to-SQL evaluation methodology. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics 351\u2013360. DOI: 10.18653\/v1\/P18-1033","DOI":"10.18653\/v1\/P18-1033"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.2307\/416535"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","unstructured":"Zhiqiang Hu Lei Wang Yihuai Lan Wanyu Xu Ee-Peng Lim Lidong Bing Xing Xu Soujanya Poria and Roy Ka-Wei Lee. 2023. LLM-adapters: An adapter family for parameter-efficient fine-tuning of large language models. DOI: 10.48550\/arXiv.2304.01933","DOI":"10.48550\/arXiv.2304.01933"},{"key":"e_1_3_2_14_2","unstructured":"Wenlong Huang Fei Xia Ted Xiao Harris Chan Jacky Liang Pete Florence Andy Zeng Jonathan Tompson Igor Mordatch Yevgen Chebotar et al. 2022. Inner monologue: Embodied reasoning through planning with language models. arXiv:2207.05608. Retrieved from https:\/\/arxiv.org\/abs\/2207.05608"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cogsys.2006.07.004"},{"key":"e_1_3_2_16_2","unstructured":"Bo Liu Yuqian Jiang Xiaohan Zhang Qiang Liu Shiqi Zhan Joydeep Biswas and Peter Stone. 2023. LLM+P: Empowering large language models with optimal planning proficiency. arXiv:2304.11477. Retrieved from https:\/\/arxiv.org\/abs\/2304.11477"},{"key":"e_1_3_2_17_2","unstructured":"Jason Xinyu Liu Ziyi Yang Ifrah Idrees Sam Liang Benjamin Schornstein Stefanie Tellex and Ankit Shah. 2023. Lang2LTL: Translating natural language commands to temporal robot task specification. arXiv:2302.11649. Retrieved from https:\/\/arxiv.org\/abs\/2302.11649"},{"key":"e_1_3_2_18_2","unstructured":"Zeyi Liu Arpit Bahety and Shuran Song. 2023. REFLECT: Summarizing robot experiences for failure explanation and correction. arXiv:2306.15724. Retrieved from https:\/\/arxiv.org\/abs\/2306.15724"},{"key":"e_1_3_2_19_2","volume-title":"Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) Workshop on Open Source Robotics","author":"Quigley Morgan","year":"2009","unstructured":"Morgan Quigley, Brian Gerkey, Ken Conley, Josh Faust, Tully Foote, Jeremy Leibs, Eric Berger, Rob Wheeler, and Andrew Ng. 2009. ROS: An open-source robot operating system. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) Workshop on Open Source Robotics."},{"key":"e_1_3_2_20_2","unstructured":"Baptiste Rozi\u00e8re Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing Ellen Tan Yossi Adi Jingyu Liu Tal Remez J\u00e9r\u00e9my Rapin et al. 2023. Code Llama: Open foundation models for code. arXiv:2308.12950. Retrieved from https:\/\/arxiv.org\/abs\/2308.12950"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","unstructured":"Baptiste Rozi\u00e8re Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing Ellen Tan Yossi Adi Jingyu Liu Tal Remez J\u00e9r\u00e9my Rapin et al. 2023. Code Llama: Open foundation models for code. DOI: 10.48550\/arXiv.2308.12950","DOI":"10.48550\/arXiv.2308.12950"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-97550-4_11"},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","unstructured":"Ishika Singh Valts Blukis Arsalan Mousavian Ankit Goyal Danfei Xu Jonathan Tremblay Dieter Fox Jesse Thomason and Animesh Garg. 2022. ProgPrompt: Generating situated robot task plans using large language models. arXiv:2209.11302. Retrieved from https:\/\/arxiv.org\/abs\/2209.11302","DOI":"10.1109\/ICRA48891.2023.10161317"},{"key":"e_1_3_2_24_2","volume-title":"The Syntactic Process","author":"Steedman Mark","year":"2001","unstructured":"Mark Steedman. 2001. The Syntactic Process. MIT Press."},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","unstructured":"Sai H. Vemprala Rogerio Bonatti Arthur Bucker and Ashish Kapoor. 2024. ChatGPT for Robotics: Design principles and model abilities. IEEE Access 12 (2024) 55682\u201355696.","DOI":"10.1109\/ACCESS.2024.3387941"},{"key":"e_1_3_2_26_2","first-page":"899","volume-title":"Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing","volume":"1","author":"Vo Nguyen","year":"2015","unstructured":"Nguyen Vo, Arindam Mitra, and Chitta Baral. 2015. The NL2KR platform for building natural language translation systems. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing, Vol. 1, 899\u2013908."},{"key":"e_1_3_2_27_2","unstructured":"Bailin Wang Zi Wang Xuezhi Wang Yuan Cao Rif A. Saurous and Yoon Kim. 2023. Grammar prompting for domain-specific language generation with large language models. arXiv:2305.19234. Retrieved from https:\/\/arxiv.org\/abs\/2305.19234"},{"key":"e_1_3_2_28_2","unstructured":"Guanzhi Wang Yuqi Xie Yunfan Jiang Ajay Mandlekar Chaowei Xiao Yuke Zhu Linxi Fan and Anima Anandkumar. 2023. Voyager: An open-ended embodied agent with large language models. arXiv:2305.16291. Retrieved from https:\/\/arxiv.org\/abs\/2305.16291"},{"key":"e_1_3_2_29_2","unstructured":"Lilian Weng. 2023. LLM Powered Autonomous Agents. Retrieved from https:\/\/lilianweng.github.io\/posts\/2023-06-23-agent\/"},{"key":"e_1_3_2_30_2","first-page":"1","volume-title":"Proceedings of the Workshop on Autonomous Mobile Service Robots","author":"Wise Melonee","year":"2016","unstructured":"Melonee Wise, Michael Ferguson, Derek King, Eric Diehr, and David Dymesich. 2016. Fetch and freight: Standard platforms for service robot applications. In Proceedings of the Workshop on Autonomous Mobile Service Robots, 1\u20136."},{"key":"e_1_3_2_31_2","unstructured":"Jimmy Wu Rika Antonova Adam Kan Marion Lepert Andy Zeng Shuran Song Jeannette Bohg Szymon Rusinkiewicz and Thomas Funkhouser. 2023. TidyBot: Personalized robot assistance with large language models. arXiv:2305.05658. Retrieved from https:\/\/arxiv.org\/abs\/2305.05658"},{"key":"e_1_3_2_32_2","unstructured":"Takuma Yoneda Jiading Fang Peng Li Huanyu Zhang Tianchong Jiang Shengjie Lin Ben Picker David Yunis Hongyuan Mei and Matthew R. Walter. 2023. Statler state-maintaining language models for embodied reasoning. arXiv:2306.17840. Retrieved from https:\/\/arxiv.org\/abs\/2306.17840"},{"key":"e_1_3_2_33_2","unstructured":"Tao Yu Rui Zhang Kai Yang Michihiro Yasunaga Dongxu Wang Zifan Li James Ma Irene Li Qingning Yao Shanelle Roman et al. 2019. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. arXiv:1809.08887. Retrieved from https:\/\/arxiv.org\/abs\/1809.08887"},{"key":"e_1_3_2_34_2","unstructured":"Biao Zhang Zhongtao Liu Colin Cherry and Orhan Firat. 2024. When scaling meets LLM finetuning: The effect of data model and finetuning method. arXiv:2402.17193. Retrieved from https:\/\/arxiv.org\/abs\/2402.17193"},{"key":"e_1_3_2_35_2","doi-asserted-by":"crossref","first-page":"3813","DOI":"10.18653\/v1\/2021.findings-acl.334","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (ACL-IJCNLP \u201921)","author":"Zhong Ruiqi","year":"2021","unstructured":"Ruiqi Zhong, Dhruba Ghosh, Dan Klein, and Jacob Steinhardt. 2021. Are larger pretrained language models uniformly better? Comparing performance at the instance level. In Proceedings of the Findings of the Association for Computational Linguistics (ACL-IJCNLP \u201921), 3813\u20133827."}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3754340","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T14:37:45Z","timestamp":1777387065000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3754340"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,28]]},"references-count":34,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,8,31]]}},"alternative-id":["10.1145\/3754340"],"URL":"https:\/\/doi.org\/10.1145\/3754340","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,28]]},"assertion":[{"value":"2024-02-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-19","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}