{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,29]],"date-time":"2026-08-29T15:58:50Z","timestamp":1788019130278,"version":"build-2784847793"},"reference-count":10,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2025,1,20]],"date-time":"2025-01-20T00:00:00Z","timestamp":1737331200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["GetMobile: Mobile Comp. and Comm."],"published-print":{"date-parts":[[2025,1,20]]},"abstract":"<jats:p>Large language models (LLMs) have transformed numerous AI applications. On-device LLM is becoming increasingly important: running LLMs locally on edge devices can reduce cloud computing costs and protect users' privacy. However, the astronomical model size and the limited hardware resources pose significant deployment challenges. To solve these issues, we propose Activation-aware Weight Quantization (AWQ) and TinyChat, an algorithm-system full-stack solution for efficient on-device LLM deployment. AWQ is a novel quantization method that identifies and protects salient weights based on activation distribution, significantly reducing model size while preserving performance. TinyChat, an optimized inference framework, translates AWQ's theoretical memory savings into practical speedups through techniques such as on-the-fly dequantization, SIMD-aware weight packing, and kernel fusion. Together, they enable 4x model size reduction and 3-4x acceleration across various edge platforms, from high-end desktop GPUs to resource-constrained IoT devices. This solution democratizes on-device LLM deployment, offering privacy-preserving, low-latency AI capabilities across a wide range of applications.<\/jats:p>","DOI":"10.1145\/3714983.3714987","type":"journal-article","created":{"date-parts":[[2025,1,21]],"date-time":"2025-01-21T18:19:02Z","timestamp":1737483542000},"page":"12-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":181,"title":["AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration"],"prefix":"10.1145","volume":"28","author":[{"given":"Ji","family":"Lin","sequence":"first","affiliation":[{"name":"MIT, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiaming","family":"Tang","sequence":"additional","affiliation":[{"name":"MIT, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haotian","family":"Tang","sequence":"additional","affiliation":[{"name":"MIT, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shang","family":"Yang","sequence":"additional","affiliation":[{"name":"MIT, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guangxuan","family":"Xiao","sequence":"additional","affiliation":[{"name":"MIT, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Song","family":"Han","sequence":"additional","affiliation":[{"name":"MIT, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,1,21]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Dettmers T. a. (2022). LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale."},{"key":"e_1_2_1_2_1","unstructured":"Dettmers T. a. (2022). The case for 4-bit precision: k-bit Inference Scaling Laws."},{"key":"e_1_2_1_3_1","volume-title":"Learned step size quantization","author":"Esser S.K.","year":"2019","unstructured":"Esser, S.K. (2019). Learned step size quantization."},{"key":"e_1_2_1_4_1","unstructured":"Frantar E. a. (2022). GPTQ: Accurate Post- Training Quantization for Generative Pre-trained Transformers."},{"key":"e_1_2_1_5_1","volume-title":"SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning","author":"Han H.W.","year":"2020","unstructured":"Han, H.W. (2020). SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning."},{"key":"e_1_2_1_6_1","volume-title":"Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production","author":"Kim Y.J.","year":"2022","unstructured":"Kim, Y.J. (2022). Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production."},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Lin J. a. (2023). Vila: On pre-training for visual language models.","DOI":"10.1109\/CVPR52733.2024.02520"},{"key":"e_1_2_1_8_1","unstructured":"Touvron H. a. (2023). Llama 2: Open foundation and fine-tuned chat models."},{"key":"e_1_2_1_9_1","doi-asserted-by":"crossref","unstructured":"Wei X. a. (2023). Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling.","DOI":"10.18653\/v1\/2023.emnlp-main.102"},{"key":"e_1_2_1_10_1","unstructured":"Xiao G. a. (2022). Smoothquant: Accurate and efficient post-training quantization for large language models."}],"container-title":["GetMobile: Mobile Computing and Communications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3714983.3714987","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3714983.3714987","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T21:17:57Z","timestamp":1750281477000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3714983.3714987"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,20]]},"references-count":10,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,1,20]]}},"alternative-id":["10.1145\/3714983.3714987"],"URL":"https:\/\/doi.org\/10.1145\/3714983.3714987","relation":{},"ISSN":["2375-0529","2375-0537"],"issn-type":[{"value":"2375-0529","type":"print"},{"value":"2375-0537","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,20]]},"assertion":[{"value":"2025-01-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}