{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T10:53:16Z","timestamp":1782298396243,"version":"3.54.5"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"2","funder":[{"DOI":"10.13039\/501100014438","name":"Business Finland 6G Bridge Program","doi-asserted-by":"publisher","award":["8782\/31\/2022"],"award-info":[{"award-number":["8782\/31\/2022"]}],"id":[{"id":"10.13039\/501100014438","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004785","name":"NordForsk","doi-asserted-by":"publisher","award":["168043"],"award-info":[{"award-number":["168043"]}],"id":[{"id":"10.13039\/501100004785","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2026,3,31]]},"abstract":"<jats:p>The rapid rise of Language Models (LMs) has expanded the capabilities of natural language processing, powering applications from text generation to complex decision-making. While state-of-the-art LMs often boast hundreds of billions of parameters and are primarily deployed in data centers, recent trends show a growing focus on compact models\u2014typically under 10 billion parameters\u2013enabled by techniques such as quantization and other model compression techniques. This shift paves the way for LMs on edge devices, offering potential benefits such as enhanced privacy, reduced latency, and improved data sovereignty. However, the inherent complexity of even these smaller models, combined with the limited computing resources of edge hardware, raises critical questions about the practical trade-offs in executing LM inference outside the cloud. To address these challenges, we present a comprehensive evaluation of generative LM inference on representative CPU-based and GPU-accelerated edge devices. Our study measures key performance indicators\u2014including memory usage, inference speed, and energy consumption\u2014across various device configurations. Additionally, we examine throughput-energy trade-offs, cost considerations, and usability, alongside an assessment of qualitative model performance. While quantization helps mitigate memory overhead, it does not fully eliminate resource bottlenecks, especially for larger models. Our findings quantify the memory and energy constraints that must be considered for practical real-world deployments, offering concrete insights into the trade-offs between model size, inference performance, and efficiency. The exploration of LMs at the edge is still in its early stages. We hope this study provides a foundation for future research, guiding the refinement of models, the enhancement of inference efficiency, and the advancement of edge-centric AI systems.<\/jats:p>","DOI":"10.1145\/3788870","type":"journal-article","created":{"date-parts":[[2026,1,12]],"date-time":"2026-01-12T21:09:16Z","timestamp":1768252156000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Sometimes Painful but Promising: Feasibility and Trade-Offs of On-Device Language Model Inference"],"prefix":"10.1145","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-9535-2393","authenticated-orcid":false,"given":"Maximilian","family":"Abstreiter","sequence":"first","affiliation":[{"name":"Department of Computer Science, University of Helsinki","place":["Helsinki, Finland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4220-3650","authenticated-orcid":false,"given":"Sasu","family":"Tarkoma","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Helsinki","place":["Helsinki, Finland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4240-9934","authenticated-orcid":false,"given":"Roberto","family":"Morabito","sequence":"additional","affiliation":[{"name":"Communication Systems Department, EURECOM","place":["Sophia Antipolis, France"]},{"name":"Department of Computer Science, University of Helsinki","place":["Helsinki, Finland"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,3,3]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone","author":"Abdin Marah","year":"2024","unstructured":"Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et\u00a0al. 2024. Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone. arXiv:2404.14219. Retrieved from https:\/\/arxiv.org\/abs\/2404.14219. (2024)."},{"key":"e_1_3_2_3_2","unstructured":"US Energy Information Administration. 2024. Electric Power Monthly November 24. Retrieved December 18 2024 from https:\/\/www.eia.gov\/electricity\/monthly\/current_month\/november2024.pdf"},{"key":"e_1_3_2_4_2","volume-title":"Yi: Open Foundation Models by 01.AI","author":"AI 01","year":"2025","unstructured":"01 AI, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Guoyin Wang, Heng Li, Jiangcheng Zhu, et\u00a0al. 2025. Yi: Open Foundation Models by 01.AI. arXiv:2403.04652. Retrieved from https:\/\/arxiv.org\/abs\/2403.04652. (2025)."},{"key":"e_1_3_2_5_2","volume-title":"Llama 3.2: Revolutionizing edge AI and vision with open, customizable models","author":"AI Meta","year":"2024","unstructured":"Meta AI. 2024. Llama 3.2: Revolutionizing edge AI and vision with open, customizable models. Retrieved January 22, 2025 from https:\/\/ai.meta.com\/blog\/llama-3-2-connect-2024-vision-edge-mobile-devices\/"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3325413.3329793"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-46002-9_23"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","unstructured":"Marc Brysbaert. 2019. How many words do we read per minute? A review and meta-analysis of reading rate. Journal of Memory and Language 109 (Dec.2019) 104047. DOI:10.1016\/j.jml.2019.104047","DOI":"10.1016\/j.jml.2019.104047"},{"key":"e_1_3_2_9_2","unstructured":"Zheng Cai Maosong Cao Haojiong Chen Kai Chen Keyu Chen Xin Chen Xun Chen Zehui Chen Zhi Chen Pei Chu et\u00a0al. 2024. InternLM2 Technical Report. arxiv:2403.17297. Retrieved from https:\/\/arxiv.org\/abs\/2403.17297. (2024)."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1179"},{"key":"e_1_3_2_11_2","unstructured":"Hyung Won Chung Le Hou Shayne Longpre Barret Zoph Yi Tai William Fedus Yunxuan Li Xuezhi Wang Mostafa Dehghani Siddhartha Brahma et\u00a0al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research 25 1 (Jan.2024) 70:3381\u201370:3433."},{"key":"e_1_3_2_12_2","unstructured":"Peter Clark Isaac Cowhey Oren Etzioni Tushar Khot Ashish Sabharwal Carissa Schoenick and Oyvind Tafjord. 2018. Think you have Solved Question Answering? Try ARC the AI2 Reasoning Challenge. arXiv:1803.05457. Retrieved from https:\/\/arxiv.org\/abs\/1803.05457. (2018)."},{"key":"e_1_3_2_13_2","series-title":"ICML\u201923","first-page":"7750","volume-title":"Proceedings of the 40th International Conference on Machine Learning","volume":"202","author":"Dettmers Tim","year":"2023","unstructured":"Tim Dettmers and Luke Zettlemoyer. 2023. The case for 4-Bit precision: K-bit inference scaling laws. In Proceedings of the 40th International Conference on Machine Learning(ICML\u201923, Vol. 202). JMLR.org, Honolulu, Hawaii, USA, 7750\u20137774."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3603287.3651205"},{"key":"e_1_3_2_16_2","unstructured":"Raspberry Pi Documentation. 2024. vcgencmd. Retrieved July 26 2024 from https:\/\/www.raspberrypi.com\/documentation\/computers\/os.html#vcgencmd"},{"key":"e_1_3_2_17_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Frantar Elias","year":"2023","unstructured":"Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2023. GPTQ: Accurate post-training quantization for generative pre-trained transformers. In Proceedings of the 11th International Conference on Learning Representations. OpenReview."},{"key":"e_1_3_2_18_2","unstructured":"Georgi Gerganov. 2023. llama.cpp. Retrieved July 18 2024 from https:\/\/github.com\/ggerganov\/llama.cpp"},{"key":"e_1_3_2_19_2","unstructured":"Google Inc. 2024. Retrieved September 2 2024 from https:\/\/ai.google.dev\/edge\/mediapipe\/solutions\/genai\/llm_inference"},{"key":"e_1_3_2_20_2","volume-title":"The Llama 3 Herd of Models","author":"Grattafiori Aaron","year":"2024","unstructured":"Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et\u00a0al. 2024. The Llama 3 Herd of Models. arXiv:2407.21783. Retrieved from https:\/\/arxiv.org\/abs\/2407.21783. (2024)."},{"key":"e_1_3_2_21_2","unstructured":"Hailo Technologies Ltd. 2025. Hailo-10H M.2 AI Acceleration Module. Retrieved October 25 2025 from https:\/\/hailo.ai\/products\/ai-accelerators\/hailo-10h-m-2-ai-acceleration-module"},{"key":"e_1_3_2_22_2","unstructured":"Huggingface. 2024. Open LLM Leaderboard. Retrieved September 13 2024 from https:\/\/huggingface.co\/spaces\/open-llm-leaderboard\/open_llm_leaderboard"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-11021-5_19"},{"key":"e_1_3_2_24_2","unstructured":"MIT HAN Lab. 2023. Retrieved July 23 2024 from https:\/\/github.com\/mit-han-lab\/TinyChatEngine"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3636534.3690668"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3662006.3662059"},{"key":"e_1_3_2_27_2","unstructured":"Ji Lin Jiaming Tang Haotian Tang Shang Yang Wei-Ming Chen Wei-Chen Wang Guangxuan Xiao Xingyu Dang Chuang Gan and Song Han. 2024. AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration. Proceedings of Machine Learning and Systems 6 (May2024) 87\u2013100."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.229"},{"key":"e_1_3_2_29_2","unstructured":"Jiachen Liu Zhiyu Wu Jae-Won Chung Fan Lai Myungjin Lee and Mosharaf Chowdhury. 2024. Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services. arXiv:2404.16283. Retrieved from https:\/\/arxiv.org\/abs\/2404.16283. (2024)."},{"key":"e_1_3_2_30_2","unstructured":"llama.cpp Team. 2023. k-quants. Retrieved July 18 2024 from https:\/\/github.com\/ggerganov\/llama.cpp\/pull\/1684"},{"key":"e_1_3_2_31_2","unstructured":"llama.cpp Team. 2024. backend cpu: add online flow for aarch64 Q4_0 GEMV\/GEMM kernels. Retrieved February 2 2025 from https:\/\/github.com\/ggerganov\/llama.cpp\/pull\/9921"},{"key":"e_1_3_2_32_2","unstructured":"Stephen Merity Caiming Xiong James Bradbury and Richard Socher. 2017. Pointer sentinel mixture models. In International Conference on Learning Representations. 1851\u20131865. Retrieved from https:\/\/openreview.net\/forum?id=Byj72udxe"},{"key":"e_1_3_2_33_2","unstructured":"MLC team. 2023-2025. MLC-LLM. Retrieved May 8 2025 from https:\/\/github.com\/mlc-ai\/mlc-llm"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","unstructured":"Jeffrey C. Mogul and Anita Borg. 1991. The effect of context switches on cache performance. ACM SIGOPS Operating Systems Review 25 Special Issue (April1991) 75\u201384. DOI:10.1145\/106974.106982","DOI":"10.1145\/106974.106982"},{"key":"e_1_3_2_35_2","unstructured":"Monsoon Solutions Inc. 2025. Monsoon High Voltage Power Monitor. Retrieved October 24 2025 from https:\/\/www.msoon.com\/high-voltage-power-monitor"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","unstructured":"Zeinab Nezami Maryam Hafeez Karim Djemame and Syed Ali Raza Zaidi. 2025. Generative AI on the edge: Architecture and performance evaluation. In ICC 2025 - IEEE International Conference on Communications. 4595\u20134602. DOI:10.1109\/ICC52391.2025.11161569","DOI":"10.1109\/ICC52391.2025.11161569"},{"key":"e_1_3_2_37_2","unstructured":"NVIDIA Corporation. 2022. NVIDIA Jetson AGX Orin Series. Retrieved January 26 2025 from https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/gtcf21\/jetson-orin\/nvidia-jetson-agx-orin-technical-brief.pdf"},{"key":"e_1_3_2_38_2","unstructured":"NVIDIA Corporation. 2024. Jetson Orin Nano Developer Kit Carrier Board Specification. Retrieved October 26 2025 from https:\/\/developer.download.nvidia.com\/assets\/embedded\/secure\/jetson\/orin_nano\/docs\/Jetson-Orin-Nano-DevKit-Carrier-Board-Specification_SP-11324-001_v1.3.pdf"},{"key":"e_1_3_2_39_2","unstructured":"NVIDIA Corporation. 2024. NVIDIA Jetson Orin Nano Series Modules Data Sheet. Retrieved February 26 2025 from https:\/\/developer.download.nvidia.com\/assets\/embedded\/secure\/jetson\/orin_nano\/docs\/Jetson-Orin-Nano-Series-Modules-Datasheet_DS-11105-001_v1.5.pdf"},{"key":"e_1_3_2_40_2","unstructured":"OpenAI. 2024. Gpt-4o system card. arXiv:2410.21276. Retrieved from https:\/\/arxiv.org\/abs\/2410.21276. (2024)."},{"key":"e_1_3_2_41_2","unstructured":"OpenAI Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman et\u00a0al. 2024. GPT-4 Technical Report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774. (2024)."},{"key":"e_1_3_2_42_2","unstructured":"OpenAI Inc.2022. ChatGPT. Retrieved July 29 2024 from https:\/\/chatgpt.com\/"},{"key":"e_1_3_2_43_2","unstructured":"OpenAI Inc.2024. API Pricing. Retrieved December 19 2024 from https:\/\/openai.com\/api\/pricing\/"},{"key":"e_1_3_2_44_2","unstructured":"OpenAI Inc.2024. What are tokens and how to count them? Retrieved August 23 2024 from https:\/\/help.openai.com\/en\/articles\/4936856-what-are-tokens-and-how-to-count-them"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA59077.2024.00019"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","unstructured":"Cheng Peng Xi Yang Aokun Chen Kaleb E. Smith Nima PourNejatian Anthony B. Costa Cheryl Martin Mona G. Flores Ying Zhang Tanja Magoc et\u00a0al. 2023. A study of generative large language model for medical research and healthcare. npj Digital Medicine 6 1 (Nov.2023) 1\u201310. DOI:10.1038\/s41746-023-00958-w","DOI":"10.1038\/s41746-023-00958-w"},{"key":"e_1_3_2_47_2","unstructured":"Inc. Picovoice. 2024. PicoLLM. Retrieved July 23 2024 from https:\/\/github.com\/Picovoice\/picollm"},{"key":"e_1_3_2_48_2","unstructured":"Qualcomm Inc.2024. Unlocking on-device generative AI with an NPU and heterogeneous computing. Retrieved October 10 2025 from https:\/\/www.qualcomm.com\/content\/dam\/qcomm-martech\/dm-assets\/documents\/Unlocking-on-device-generative-AI-with-an-NPU-and-heterogeneous-computing.pdf"},{"key":"e_1_3_2_49_2","unstructured":"Raspberry Pi Ltd. 2025. Raspberry Pi 5. Retrieved February 26 2025 from https:\/\/datasheets.raspberrypi.com\/rpi5\/raspberry-pi-5-product-brief.pdf"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00045"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","unstructured":"Keisuke Sakaguchi Ronan Le Bras Chandra Bhagavatula and Yejin Choi. 2020. WinoGrande: An adversarial winograd schema challenge at Scale. Proceedings of the AAAI Conference on Artificial Intelligence 34 05 (April2020) 8732\u20138740. DOI:10.1609\/aaai.v34i05.6399","DOI":"10.1609\/aaai.v34i05.6399"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3629526.3645054"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/EDGE60047.2023.00034"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3640794.3665582"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","unstructured":"Gemma Team. 2024. Gemma. DOI:10.34740\/KAGGLE\/M\/3301","DOI":"10.34740\/KAGGLE\/M\/3301"},{"key":"e_1_3_2_56_2","first-page":"5998","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141 ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 30. Curran Associates, Inc., Red Hook, NY, USA, 5998\u20136008. Retrieved August 26, 2024 from https:\/\/papers.nips.cc\/paper_files\/paper\/2017\/hash\/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html"},{"key":"e_1_3_2_57_2","unstructured":"Hongyu Wang Shuming Ma Lingxiao Ma Lei Wang Wenhui Wang Li Dong Shaohan Huang Huaijie Wang Jilong Xue Ruiping Wang Yi Wu and Furu Wei. 2025. BitNet: 1-bit pre-training for large language models. Journal of Machine Learning Research 26 125 (2025) 1\u201329. Retrieved from http:\/\/jmlr.org\/papers\/v26\/24-2050.html"},{"key":"e_1_3_2_58_2","first-page":"38087","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Xiao Guangxuan","year":"2023","unstructured":"Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. SmoothQuant: Accurate and efficient post-training quantization for large language models. In Proceedings of the 40th International Conference on Machine Learning. PMLR, 38087\u201338099."},{"key":"e_1_3_2_59_2","unstructured":"An Yang Baosong Yang Binyuan Hui Bo Zheng Bowen Yu Chang Zhou Chengpeng Li Chengyuan Li Dayiheng Liu Fei Huang et\u00a0al. 2024. Qwen2 Technical Report. arxiv:2407.10671. Retrieved from https:\/\/arxiv.org\/abs\/2407.10671. (2024)."},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1472"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.naacl-long.520"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","unstructured":"Xunyu Zhu Jian Li Yong Liu Can Ma and Weiping Wang. 2024. A survey on model compression for large language models. Transactions of the Association for Computational Linguistics 12 (Nov.2024) 1556\u20131577. DOI:10.1162\/tacl_a_00704","DOI":"10.1162\/tacl_a_00704"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3788870","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T08:15:34Z","timestamp":1773389734000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3788870"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,3]]},"references-count":61,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,3,31]]}},"alternative-id":["10.1145\/3788870"],"URL":"https:\/\/doi.org\/10.1145\/3788870","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,3]]},"assertion":[{"value":"2025-05-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-16","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}