{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T17:54:34Z","timestamp":1781546074027,"version":"3.54.5"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T00:00:00Z","timestamp":1781481600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62572069"],"award-info":[{"award-number":["62572069"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62232004"],"award-info":[{"award-number":["62232004"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100013314","name":"111 Project","doi-asserted-by":"crossref","award":["B18008"],"award-info":[{"award-number":["B18008"]}],"id":[{"id":"10.13039\/501100013314","id-type":"DOI","asserted-by":"crossref"}]},{"name":"BUPT innovation and entrepreneurship support program","award":["2025-YC T014"],"award-info":[{"award-number":["2025-YC T014"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2026,6,15]]},"abstract":"<jats:p>Smart homes are receiving growing interest in global markets. However, current smart home assistant systems either control appliances solely through users' explicit commands or rely on cloud-based large language models (LLMs) for vague command understanding, which brings drawbacks such as high latency, privacy concern, and high cost. In this work, we propose VCU-LLM, the first system to deploy LLMs on edge devices for local vague command understanding and smart device control plan generation. VCU-LLM introduces a novel vague command knowledge retrieval algorithm that refines device-related information in the input prompt, thereby accelerating the LLM's on-device inference and reducing task complexity. We further construct a dataset for LLM fine-tuning to simulate the use of smart home assistants in controlling devices across different households. During inference, a customized KV-cache technique is applied for further inference acceleration. Our evaluations with both human-based and LLM-based scoring demonstrate that VCU-LLM improves the quality of generated control plans by an average of 43.3% compared with SOTA baselines, while reducing time overhead by an average of 8.44x compared with other on-device baselines. We also implement VCU-LLM through a case study in a real home environment, demonstrating its feasibility in real-world application.<\/jats:p>","DOI":"10.1145\/3810190","type":"journal-article","created":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T17:06:41Z","timestamp":1781543201000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["VCU-LLM: Prompt-efficient On-device Large Language Model for Vague Command Understanding in Smart Homes"],"prefix":"10.1145","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-0262-3323","authenticated-orcid":false,"given":"Zhengyuan","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Computer Science, Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7337-9168","authenticated-orcid":false,"given":"Dong","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Computer Science, Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-3005-7050","authenticated-orcid":false,"given":"Tiancheng","family":"He","sequence":"additional","affiliation":[{"name":"International School, Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-8097-0640","authenticated-orcid":false,"given":"Zilong","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-3456-0316","authenticated-orcid":false,"given":"Xiangyu","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science, Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7199-5047","authenticated-orcid":false,"given":"Huadong","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Computer Science, Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,15]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2025. Alexa Smart Home -Amazon. https:\/\/www.amazon.com\/alexa-smart-home\/."},{"key":"e_1_2_1_2_1","unstructured":"2025. all-MiniLM-L6-v2. https:\/\/huggingface.co\/sentence-transformers\/all-MiniLM-L6-v2."},{"key":"e_1_2_1_3_1","unstructured":"2025. GGUF. https:\/\/huggingface.co\/docs\/hub\/en\/gguf."},{"key":"e_1_2_1_4_1","unstructured":"2025. Home Assistant. https:\/\/www.home-assistant.io\/."},{"key":"e_1_2_1_5_1","unstructured":"2025. Home-Assistant-Requests. https:\/\/huggingface.co\/datasets\/acon96\/Home-Assistant-Requests."},{"key":"e_1_2_1_6_1","unstructured":"2025. HomeLLM. https:\/\/github.com\/acon96\/home-llm."},{"key":"e_1_2_1_7_1","unstructured":"2025. Llama.cpp. https:\/\/github.com\/ggml-org\/llama.cpp."},{"key":"e_1_2_1_8_1","unstructured":"2025. Mi IoT Platform. https:\/\/github.com\/XiaoMi\/ha_xiaomi_home."},{"key":"e_1_2_1_9_1","volume-title":"Number of users of smart homes worldwide from 2019 to","year":"2028","unstructured":"2025. Number of users of smart homes worldwide from 2019 to 2028. https:\/\/www.statista.com\/forecasts\/887613\/number-of-smart-homes-in-the-smart-home-market-in-the-world."},{"key":"e_1_2_1_10_1","unstructured":"2025. OpenAI. https:\/\/openai.com\/."},{"key":"e_1_2_1_11_1","unstructured":"2025. OpenAI API. https:\/\/openai.com\/api\/."},{"key":"e_1_2_1_12_1","unstructured":"2025. Raspberry Pi 5. https:\/\/www.raspberrypi.com\/products\/raspberry-pi-5\/."},{"key":"e_1_2_1_13_1","unstructured":"2025. Smart Home - Worldwide. https:\/\/www.statista.com\/outlook\/cmo\/smart-home\/worldwide."},{"key":"e_1_2_1_14_1","unstructured":"2025. Smart Home | Xiaomi Global. https:\/\/www.mi.com\/global\/smart-home\/."},{"key":"e_1_2_1_15_1","unstructured":"2025. Smart Home Spec. https:\/\/home.miot-spec.com\/."},{"key":"e_1_2_1_16_1","volume-title":"Adoption Report","year":"2018","unstructured":"2025. Smart Speaker Consumer Adoption Report 2018. https:\/\/voicebot.ai\/download-smart-speaker-consumer-adoption-report-2018\/."},{"key":"e_1_2_1_17_1","unstructured":"2025. Welcome to a more helpful home - Google Home. https:\/\/home.google.com\/welcome\/."},{"key":"e_1_2_1_18_1","volume-title":"Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.","author":"Achiam Josh","year":"2023","unstructured":"Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)."},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL). 12562\u201312584","author":"Alizadeh Keivan","year":"2024","unstructured":"Keivan Alizadeh, Seyed Iman Mirzadeh, Dmitry Belenko, S Khatamifard, Minsik Cho, Carlo C Del Mundo, Mohammad Rastegari, and Mehrdad Farajtabar. 2024. Llm in a flash: Efficient large language model inference with limited memory. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL). 12562\u201312584."},{"key":"e_1_2_1_20_1","volume-title":"Jamie Ryan Kiros, and Geoffrey E Hinton","author":"Ba Jimmy Lei","year":"2016","unstructured":"Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. arXiv preprint arXiv: 1607.06450 (2016)."},{"key":"e_1_2_1_21_1","volume-title":"Unified active retrieval for retrieval augmented generation. arXiv preprint arXiv:2406.12534","author":"Cheng Qinyuan","year":"2024","unstructured":"Qinyuan Cheng, Xiaonan Li, Shimin Li, Qin Zhu, Zhangyue Yin, Yunfan Shao, Linyang Li, Tianxiang Sun, Hang Yan, and Xipeng Qiu. 2024. Unified active retrieval for retrieval augmented generation. arXiv preprint arXiv:2406.12534 (2024)."},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the 19th international conference on human-computer interaction with mobile devices and services (MobileHCI). 1\u201312","author":"Cowan Benjamin R","year":"2017","unstructured":"Benjamin R Cowan, Nadia Pantidi, David Coyle, Kellie Morrissey, Peter Clarke, Sara Al-Shehri, David Earley, and Natasha Bandeira. 2017. \u201cWhat can i help you with?\u201d infrequent users' experiences of intelligent personal assistants. In Proceedings of the 19th international conference on human-computer interaction with mobile devices and services (MobileHCI). 1\u201312."},{"key":"e_1_2_1_23_1","volume-title":"Mobile-bench: An evaluation benchmark for llm-based mobile agents. arXiv preprint arXiv:2407.00993","author":"Deng Shihan","year":"2024","unstructured":"Shihan Deng, Weikai Xu, Hongda Sun, Wei Liu, Tao Tan, Jianfeng Liu, Ang Li, Jian Luan, Bin Wang, Rui Yan, et al. 2024. Mobile-bench: An evaluation benchmark for llm-based mobile agents. arXiv preprint arXiv:2407.00993 (2024)."},{"key":"e_1_2_1_24_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. In arXiv preprint arXiv:1810.04805. 1\u201316.","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. In arXiv preprint arXiv:1810.04805. 1\u201316."},{"key":"e_1_2_1_25_1","unstructured":"Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Amy Yang Angela Fan et al. 2024. The llama 3 herd of models. arXiv e-prints (2024) arXiv-2407."},{"key":"e_1_2_1_26_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3678585","article-title":"ChatIoT: Zero-code Generation of Trigger-action Based IoT Programs","volume":"8","author":"Gao Yi","year":"2024","unstructured":"Yi Gao, Kaijie Xiao, Fu Li, Weifeng Xu, Jiaming Huang, and Wei Dong. 2024. ChatIoT: Zero-code Generation of Trigger-action Based IoT Programs. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 3 (2024), 1\u201329.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_1_27_1","volume-title":"Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 2, 1","author":"Gao Yunfan","year":"2023","unstructured":"Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 2, 1 (2023)."},{"key":"e_1_2_1_28_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3381002","article-title":"He is just like me: a study of the long-term use of smart speakers by parents and children","volume":"4","author":"Garg Radhika","year":"2020","unstructured":"Radhika Garg and Subhasree Sengupta. 2020. He is just like me: a study of the long-term use of smart speakers by parents and children. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 4, 1 (2020), 1\u201324.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_1_29_1","unstructured":"Daya Guo Dejian Yang Haowei Zhang Junxiao Song Ruoyu Zhang Runxin Xu Qihao Zhu Shirong Ma Peiyi Wang Xiao Bi et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948 (2025)."},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems","volume":"2","author":"Guo Liwei","year":"2023","unstructured":"Liwei Guo, Wonkyo Choe, and Felix Xiaozhu Lin. 2023. Sti: Turbocharge nlp inference at the edge via elastic pipelining. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. 791\u2013803."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","unstructured":"Matthew Honnibal and Ines Montani. 2017. spaCy 2: Natural language understanding with Bloom embeddings convolutional neural networks and incremental parsing. Zenodo software release. doi:10.5281\/zenodo.1212303","DOI":"10.5281\/zenodo.1212303"},{"key":"e_1_2_1_32_1","first-page":"3","article-title":"Lora: Low-rank adaptation of large language models","volume":"1","author":"Hu Edward J","year":"2022","unstructured":"Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2022. Lora: Low-rank adaptation of large language models. ICLR 1, 2 (2022), 3.","journal-title":"ICLR"},{"key":"e_1_2_1_33_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3703155","article-title":"A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions","volume":"43","author":"Huang Lei","year":"2025","unstructured":"Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2025. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems 43, 2 (2025), 1\u201355.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_2_1_34_1","volume-title":"Understanding the planning of LLM agents: A survey. arXiv preprint arXiv:2402.02716","author":"Huang Xu","year":"2024","unstructured":"Xu Huang, Weiwen Liu, Xiaolong Chen, Xingmei Wang, Hao Wang, Defu Lian, Yasheng Wang, Ruiming Tang, and Enhong Chen. 2024. Understanding the planning of LLM agents: A survey. arXiv preprint arXiv:2402.02716 (2024)."},{"key":"e_1_2_1_35_1","volume-title":"Scaling laws for neural language models. arXiv preprint arXiv:2001.08361","author":"Kaplan Jared","year":"2020","unstructured":"Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361 (2020)."},{"key":"e_1_2_1_36_1","doi-asserted-by":"crossref","first-page":"106914","DOI":"10.1016\/j.chb.2021.106914","article-title":"Exploring older adults' perception and use of smart speaker-based voice assistants: A longitudinal study","volume":"124","author":"Kim Sunyoung","year":"2021","unstructured":"Sunyoung Kim and Abhishek Choudhury. 2021. Exploring older adults' perception and use of smart speaker-based voice assistants: A longitudinal study. Computers in Human Behavior (COHB) 124 (2021), 106914.","journal-title":"Computers in Human Behavior (COHB)"},{"key":"e_1_2_1_37_1","unstructured":"Evan King Haoxiang Yu Sangsu Lee and Christine Julien. 2023. \u201cGet ready for a party\u201d: Exploring smarter smart spaces with help from large language models. arXiv preprint arXiv:2303.14143 (2023)."},{"key":"e_1_2_1_38_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3643505","article-title":"Sasha: creative goal-oriented reasoning in smart homes with large language models","volume":"8","author":"King Evan","year":"2024","unstructured":"Evan King, Haoxiang Yu, Sangsu Lee, and Christine Julien. 2024. Sasha: creative goal-oriented reasoning in smart homes with large language models. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 1 (2024), 1\u201338.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_1_39_1","volume-title":"The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691","author":"Lester Brian","year":"2021","unstructured":"Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691 (2021)."},{"key":"e_1_2_1_40_1","volume-title":"International Conference on Machine Learning (ICML). 19274\u201319286","author":"Leviathan Yaniv","year":"2023","unstructured":"Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning (ICML). 19274\u201319286."},{"key":"e_1_2_1_41_1","unstructured":"Patrick Lewis Ethan Perez Aleksandra Piktus Fabio Petroni Vladimir Karpukhin Naman Goyal Heinrich K\u00fcttler Mike Lewis Wen-tau Yih Tim Rockt\u00e4schel et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems 33 (2020) 9459\u20139474."},{"key":"e_1_2_1_42_1","first-page":"22947","article-title":"Snapkv: Llm knows what you are looking for before generation","volume":"37","author":"Li Yuhong","year":"2024","unstructured":"Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen. 2024. Snapkv: Llm knows what you are looking for before generation. Advances in Neural Information Processing Systems 37 (2024), 22947\u201322970.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_43_1","unstructured":"Yuanchun Li Hao Wen Weijun Wang Xiangyu Li Yizhen Yuan Guohong Liu Jiacheng Liu Wenxing Xu Xiang Wang Yi Sun et al. 2024. Personal llm agents: Insights and survey about the capability efficiency and security. arXiv preprint arXiv:2401.05459 (2024)."},{"key":"e_1_2_1_44_1","volume-title":"Kivi: A tuning-free asymmetric 2bit quantization for kv cache. arXiv preprint arXiv:2402.02750","author":"Liu Zirui","year":"2024","unstructured":"Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu. 2024. Kivi: A tuning-free asymmetric 2bit quantization for kv cache. arXiv preprint arXiv:2402.02750 (2024)."},{"key":"e_1_2_1_45_1","volume-title":"ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 12066\u201312070","author":"Paul Sudipta","year":"2024","unstructured":"Sudipta Paul, Lingyu Zhang, Yilin Shen, and Hongxia Jin. 2024. Enabling device control planning capabilities of small language model. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 12066\u201312070."},{"key":"e_1_2_1_46_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3373759","article-title":"Use of intelligent voice assistants by older adults with low technology use","volume":"27","author":"Pradhan Alisha","year":"2020","unstructured":"Alisha Pradhan, Amanda Lazar, and Leah Findlater. 2020. Use of intelligent voice assistants by older adults with low technology use. ACM Transactions on Computer-Human Interaction (TOCHI) 27, 4 (2020), 1\u201327.","journal-title":"ACM Transactions on Computer-Human Interaction (TOCHI)"},{"key":"e_1_2_1_47_1","volume-title":"Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv","author":"Reimers Nils","year":"2019","unstructured":"Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv: 1908.10084 (2019)."},{"key":"e_1_2_1_48_1","first-page":"34","article-title":"Joint Intent Detection and Slot Filling with Rules","volume":"2242","author":"Ren Shiya","year":"2018","unstructured":"Shiya Ren, Huaming Wang, Dongming Yu, Yuan Li, Zhixing Li, S Hu, and L Zou. 2018. Joint Intent Detection and Slot Filling with Rules. CCKS Tasks 2242 (2018), 34\u201340.","journal-title":"CCKS Tasks"},{"key":"e_1_2_1_49_1","volume-title":"Leveraging Large Language Models for enhanced personalised user experience in Smart Homes. arXiv preprint arXiv:2407.12024","author":"Rey-Jouanchicot Jordan","year":"2024","unstructured":"Jordan Rey-Jouanchicot, Andr\u00e9 Bottaro, Eric Campo, Jean-L\u00e9on Bouraoui, Nadine Vigouroux, and Fr\u00e9d\u00e9ric Vella. 2024. Leveraging Large Language Models for enhanced personalised user experience in Smart Homes. arXiv preprint arXiv:2407.12024 (2024)."},{"key":"e_1_2_1_50_1","volume-title":"AIoT Smart Home via Autonomous LLM Agents","author":"Rivkin Dmitriy","year":"2024","unstructured":"Dmitriy Rivkin, Francois Hogan, Amal Feriani, Abhisek Konar, Adam Sigal, Xue Liu, and Gregory Dudek. 2024. AIoT Smart Home via Autonomous LLM Agents. IEEE Internet of Things Journal (2024)."},{"key":"e_1_2_1_51_1","volume-title":"International Conference on Machine Learning (ICML). 31094\u201331116","author":"Sheng Ying","year":"2023","unstructured":"Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher R\u00e9, Ion Stoica, and Ce Zhang. 2023. Flexgen: High-throughput generative inference of large language models with a single gpu. In International Conference on Machine Learning (ICML). 31094\u201331116."},{"key":"e_1_2_1_52_1","volume-title":"Bridging the gap between natural user expression with complex automation programming in smart homes. arXiv preprint arXiv:2408.12687","author":"Shi Yingtian","year":"2024","unstructured":"Yingtian Shi, Xiaoyi Liu, Chun Yu, Tianao Yang, Cheng Gao, Chen Liang, and Yuanchun Shi. 2024. Bridging the gap between natural user expression with complex automation programming in smart homes. arXiv preprint arXiv:2408.12687 (2024)."},{"key":"e_1_2_1_53_1","unstructured":"Qwen Team. 2024. Qwen2.5: A Party of Foundation Models. https:\/\/qwenlm.github.io\/blog\/qwen2.5\/"},{"key":"e_1_2_1_54_1","volume-title":"A comprehensive survey of hallucination mitigation techniques in large language models. arXiv preprint arXiv:2401.01313 6","author":"Tonmoy SM","year":"2024","unstructured":"SM Tonmoy, SM Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. 2024. A comprehensive survey of hallucination mitigation techniques in large language models. arXiv preprint arXiv:2401.01313 6 (2024)."},{"key":"e_1_2_1_55_1","volume-title":"Llama: Open and efficient foundation language models. In arXiv preprint arXiv:2302.13971. 2556\u20132565.","author":"Touvron Hugo","year":"2023","unstructured":"Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth\u00e9e Lacroix, Baptiste Rozi\u00e8re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. In arXiv preprint arXiv:2302.13971. 2556\u20132565."},{"key":"e_1_2_1_56_1","volume-title":"Proceedings of the 2023 CHI conference on human factors in computing systems (CHI). 1\u201311","author":"Upadhyay Pooja","year":"2023","unstructured":"Pooja Upadhyay, Sharon Heung, Shiri Azenkot, and Robin N Brewer. 2023. Studying exploration & long-term use of voice assistants by older adults. In Proceedings of the 2023 CHI conference on human factors in computing systems (CHI). 1\u201311."},{"key":"e_1_2_1_57_1","volume-title":"AAAI 2025 Workshop on Artificial Intelligence for Wireless Communications and Networking (AI4WCN).","author":"Velaga Krishna Sruthi","year":"2025","unstructured":"Krishna Sruthi Velaga and Yifan Guo. 2025. Optimizing Large Language Models Assisted Smart Home Assistant Systems at the Edge: An Empirical Study. In AAAI 2025 Workshop on Artificial Intelligence for Wireless Communications and Networking (AI4WCN)."},{"key":"e_1_2_1_58_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3547138","article-title":"A survey of joint intent detection and slot filling models in natural language understanding","volume":"55","author":"Weld Henry","year":"2022","unstructured":"Henry Weld, Xiaoqi Huang, Siqu Long, Josiah Poon, and Soyeon Caren Han. 2022. A survey of joint intent detection and slot filling models in natural language understanding. Comput. Surveys 55, 8 (2022), 1\u201338.","journal-title":"Comput. Surveys"},{"key":"e_1_2_1_59_1","doi-asserted-by":"crossref","first-page":"463","DOI":"10.1007\/s00779-014-0813-0","article-title":"Smart homes and their users: a systematic analysis and key challenges","volume":"19","author":"Wilson Charlie","year":"2015","unstructured":"Charlie Wilson, Tom Hargreaves, and Richard Hauxwell-Baldwin. 2015. Smart homes and their users: a systematic analysis and key challenges. Personal and Ubiquitous Computing 19 (2015), 463\u2013476.","journal-title":"Personal and Ubiquitous Computing"},{"key":"e_1_2_1_60_1","volume-title":"EdgeLLM: Fast On-device LLM Inference with Speculative Decoding","author":"Xu Daliang","year":"2024","unstructured":"Daliang Xu, Wangsong Yin, Hao Zhang, Xin Jin, Ying Zhang, Shiyun Wei, Mengwei Xu, and Xuanzhe Liu. 2024. EdgeLLM: Fast On-device LLM Inference with Speculative Decoding. IEEE Transactions on Mobile Computing (TMC) (2024)."},{"key":"e_1_2_1_61_1","unstructured":"An Yang Anfeng Li Baosong Yang Beichen Zhang Binyuan Hui Bo Zheng Bowen Yu Chang Gao Chengen Huang Chenxu Lv et al. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388 (2025)."},{"key":"e_1_2_1_62_1","volume-title":"Kvlink: Accelerating large language models via efficient kv cache reuse. arXiv preprint arXiv:2502.16002","author":"Yang Jingbo","year":"2025","unstructured":"Jingbo Yang, Bairu Hou, Wei Wei, Yujia Bao, and Shiyu Chang. 2025. Kvlink: Accelerating large language models via efficient kv cache reuse. arXiv preprint arXiv:2502.16002 (2025)."},{"key":"e_1_2_1_63_1","volume-title":"TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization. arXiv preprint arXiv:2505.19586","author":"Yao Dingyu","year":"2025","unstructured":"Dingyu Yao, Bowen Shen, Zheng Lin, Wei Liu, Jian Luan, Bin Wang, and Weiping Wang. 2025. TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization. arXiv preprint arXiv:2505.19586 (2025)."},{"key":"e_1_2_1_64_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3411838","article-title":"Trace2tap: Synthesizing trigger-action programs from traces of behavior","volume":"4","author":"Zhang Lefan","year":"2020","unstructured":"Lefan Zhang, Weijia He, Olivia Morkved, Valerie Zhao, Michael L Littman, Shan Lu, and Blase Ur. 2020. Trace2tap: Synthesizing trigger-action programs from traces of behavior. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 4, 3 (2020), 1\u201326.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_1_65_1","volume-title":"Tinyllama: An open-source small language model. arXiv preprint arXiv:2401.02385","author":"Zhang Peiyuan","year":"2024","unstructured":"Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, and Wei Lu. 2024. Tinyllama: An open-source small language model. arXiv preprint arXiv:2401.02385 (2024)."},{"key":"e_1_2_1_66_1","volume-title":"SleepCoT: A Lightweight Personalized Sleep Health Model via Chain-of-Thought Distillation. arXiv preprint arXiv:2410.16924","author":"Zheng Huimin","year":"2024","unstructured":"Huimin Zheng, Xiaofeng Xing, and Xiangmin Xu. 2024. SleepCoT: A Lightweight Personalized Sleep Health Model via Chain-of-Thought Distillation. arXiv preprint arXiv:2410.16924 (2024)."}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3810190","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T17:16:06Z","timestamp":1781543766000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3810190"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,15]]},"references-count":66,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,15]]}},"alternative-id":["10.1145\/3810190"],"URL":"https:\/\/doi.org\/10.1145\/3810190","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,15]]},"assertion":[{"value":"2026-06-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}