{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T16:54:21Z","timestamp":1787504061296,"version":"build-2736575974"},"reference-count":85,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T00:00:00Z","timestamp":1778630400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Army Research Laboratory","award":["W911NF-24-2-0080"],"award-info":[{"award-number":["W911NF-24-2-0080"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>\n                    Large Language Models (LLMs) have become foundational tools in natural language processing, achieving state-of-the-art performance across a variety of tasks. However, their immense size and computational requirements make them impractical for deployment in resource-constrained environments, such as edge devices and embedded systems. In this work, we introduce\n                    <jats:italic toggle=\"yes\">Magnitude and Gradient-Informed Pruning (MaGrIP)<\/jats:italic>\n                    , a novel framework for task-agnostic pruning and compression of LLMs. MaGrIP employs a dual-threshold strategy combining magnitude- and gradient-based saliency measures to efficiently prune redundant neurons while retaining task performance. Our results demonstrate the effectiveness of MaGrIP in compressing state-of-the-art models. The compression reduced the total computational complexity of the FFN layers from\n                    <jats:inline-formula content-type=\"math\/tex\">\n                      <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(\\mathcal {O}(d \\cdot h)\\)<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    to\n                    <jats:inline-formula content-type=\"math\/tex\">\n                      <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(\\mathcal {O}((d - q) \\cdot h)\\)<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    . In terms of model size, our pruning approach significantly reduces both model parameters and storage requirements while maintaining competitive perplexity scores evaluated on WikiText-2. For the Gemma 7B model, our method reduces the total size from 28 GB to 5 GB, while for Gemma 2B, MaGrIP achieves a size reduction from 8 GB to 1.5 GB. MaGrIP furthermore exhibits robust performance across multiple benchmarks, such as BOOLQ, ARC-E, and CSQA. Specifically, the pruned Gemma 7B model at 50% pruning achieved 59.26% accuracy on ARC-E compared to 81.06% for the baseline, and 64.74% accuracy on BoolQ compared to 59.98% for the baseline. Similarly, the pruned Llama 3 8B at 50% pruning achieved 46.76% accuracy on ARC-E compared to 77.57% for the baseline, reflecting the tradeoff between compression and accuracy. LLMs compressed using MaGrIP, when deployed on the Nvidia Jetson Orin Nano, achieved a 2.16\u00d7 improvement in throughput and a 2.3\u00d7 improvement in performance compared to baseline LLMs.\n                  <\/jats:p>","DOI":"10.1145\/3766068","type":"journal-article","created":{"date-parts":[[2025,9,5]],"date-time":"2025-09-05T11:28:25Z","timestamp":1757071705000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["MaGrIP: Magnitude and Gradient-Informed Pruning for Task-Agnostic Large Language Models"],"prefix":"10.1145","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-1150-3903","authenticated-orcid":false,"given":"Uttej","family":"Kallakuri","sequence":"first","affiliation":[{"name":"Electrical and Computer Engineering, Johns Hopkins University","place":["Baltimore, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-3945-0116","authenticated-orcid":false,"given":"Edward","family":"Humes","sequence":"additional","affiliation":[{"name":"Computer Science and Electrical Engineering, University of Maryland Baltimore County","place":["Baltimore, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9983-6929","authenticated-orcid":false,"given":"Hasib-Al","family":"Rashid","sequence":"additional","affiliation":[{"name":"Computer Science and Electrical Engineering, University of Maryland Baltimore County","place":["Baltimore, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5551-2124","authenticated-orcid":false,"given":"Tinoosh","family":"Mohsenin","sequence":"additional","affiliation":[{"name":"Johns Hopkins University","place":["Baltimore, United States"]},{"name":"University of Maryland Baltimore County","place":["Baltimore, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,5,13]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","DOI":"10.1109\/BioCAS67066.2025.00057","article-title":"MedMambaLite: Hardware-aware mamba for medical image classification","author":"Aalishah Romina","year":"2025","unstructured":"Romina Aalishah, Mozhgan Navardi, and Tinoosh Mohsenin. 2025. MedMambaLite: Hardware-aware mamba for medical image classification. In Proceedings of the IEEE Biomedical Circuits and Systems (BioCAS) Conference.","journal-title":"Proceedings of the IEEE Biomedical Circuits and Systems (BioCAS) Conference"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i10.28960"},{"key":"e_1_3_2_4_2","unstructured":"Saleh Ashkboos Maximilian L. Croci Marcelo Gennari do Nascimento Torsten Hoefler and James Hensman. 2024. Slicegpt: Compress large language models by deleting rows and columns. arXiv:2401.15024. Retrieved from https:\/\/arxiv.org\/abs\/2401.15024"},{"key":"e_1_3_2_5_2","unstructured":"Haoli Bai Wei Zhang Lu Hou Lifeng Shang Jing Jin Xin Jiang Qun Liu Michael Lyu and Irwin King. 2020. Binarybert: Pushing the limit of bert quantization. arXiv:2012.15701. Retrieved from https:\/\/arxiv.org\/abs\/2012.15701"},{"key":"e_1_3_2_6_2","unstructured":"Jerry Chee Yaohui Cai Volodymyr Kuleshov and Christopher M. De Sa. 2023. QuIP: 2-bit quantization of large language models with guarantees. In Advances in Neural Information Processing Systems A. Oh T. Naumann A. Globerson K. Saenko M. Hardt and S. Levine (Eds.). Curran Associates Inc. 4396\u20134429. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2023\/file\/0df38cd13520747e1e64e5b123a78ef8-Paper-Conference.pdf"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2021.3061394"},{"key":"e_1_3_2_8_2","unstructured":"Christopher Clark Kenton Lee Ming-Wei Chang Tom Kwiatkowski Michael Collins and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes\/no questions. arXiv:1905.10044. Retrieved from https:\/\/arxiv.org\/abs\/1905.10044"},{"key":"e_1_3_2_9_2","unstructured":"Peter Clark Isaac Cowhey Oren Etzioni Tushar Khot Ashish Sabharwal Carissa Schoenick and Oyvind Tafjord. 2018. Think you have solved question answering? Try ARC the AI2 reasoning challenge. arXiv:1803.05457. Retrieved from https:\/\/arxiv.org\/abs\/1803.05457"},{"key":"e_1_3_2_10_2","unstructured":"NVIDIA Corporation. 2023. Jetson Orin Nano Developer Kit Technical Specifications. Retrieved June 27 2024 from https:\/\/developer.nvidia.com\/embedded\/jetson-orin-nano-developer-kit."},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Tim Dettmers Mike Lewis Younes Belkada and Luke Zettlemoyer. 2022. GPT3.int8(): 8-bit matrix multiplication for transformers at scale. In Advances in Neural Information Processing Systems S. Koyejo S. Mohamed A. Agarwal D. Belgrave K. Cho and A. Oh (Eds.). Curran Associates Inc. 30318\u201330332. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/c3ba4962c05c49636d4c6206a97e9c8a-Paper-Conference.pdf","DOI":"10.52202\/068431-2198"},{"key":"e_1_3_2_12_2","unstructured":"Jacob Devlin. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","unstructured":"Fnu Devvrit Sneha Kudugunta Aditya Kusupati Tim Dettmers Kaifeng Chen Inderjit Dhillon Yulia Tsvetkov Hannaneh Hajishirzi Sham Kakade Ali Farhadi and Prateek Jain. 2024. MatFormer: Nested transformer for elastic inference. In Advances in Neural Information Processing Systems Curran Associates Inc. 140535\u2013140564. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2024\/file\/fe066022bab2a6c6a3c57032a1623c70-Paper-Conference.pdf","DOI":"10.52202\/079017-4461"},{"key":"e_1_3_2_14_2","unstructured":"Dayou Du Yijia Zhang Shijie Cao Jiaqi Guo Ting Cao Xiaowen Chu and Ningyi Xu. 2024. Bitdistiller: Unleashing the potential of sub-4-bit LLMs via self-distillation. arXiv:2402.10631. Retrieved from https:\/\/arxiv.org\/abs\/2402.10631"},{"key":"e_1_3_2_15_2","unstructured":"Cl\u00e9mentine Fourrier Nathan Habib Thomas Wolf and Lewis Tunstall. 2023. LightEval: A lightweight framework for LLM evaluation. Retrieved June 19 2025 from https:\/\/github.com\/huggingface\/lighteval"},{"key":"e_1_3_2_16_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Frantar Elias","year":"2022","unstructured":"Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. OPTQ: Accurate quantization for generative pre-trained transformers. In Proceedings of the 11th International Conference on Learning Representations."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","unstructured":"Amir Gholami Zhewei Yao Sehoon Kim Coleman Hooper Michael W. Mahoney and Kurt Keutzer. 2024. AI and memory wall. IEEE Micro 44 3 (2024) 33\u201339. DOI:10.1109\/MM.2024.3373763","DOI":"10.1109\/MM.2024.3373763"},{"key":"e_1_3_2_18_2","unstructured":"Aaron Grattafiori Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Alex Vaughan et\u00a0al. 2024. The Llama 3 herd of models. arXiv:2407.21783. Retrieved from https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3579371.3589038"},{"key":"e_1_3_2_20_2","unstructured":"Fu-Ming Guo Sijia Liu Finlay S. Mungall Xue Lin and Yanzhi Wang. 2019. Reweighted proximal pruning for large-scale language representation. arXiv:1909.12486. Retrieved from https:\/\/arxiv.org\/abs\/1909.12486"},{"key":"e_1_3_2_21_2","unstructured":"Song Han Huizi Mao and William J. Dally. 2015. Deep compression: Compressing deep neural networks with pruning trained quantization and Huffman coding. arXiv:1510.00149. Retrieved from https:\/\/arxiv.org\/abs\/1510.00149"},{"key":"e_1_3_2_22_2","unstructured":"Coleman Hooper Sehoon Kim Hiva Mohammadzadeh Michael W. Mahoney Yakun Sophia Shao Kurt Keutzer and Amir Gholami. 2024. KVQuant: Towards 10 million context length LLM inference with KV cache quantization. arXiv:2401.18079. Retrieved from https:\/\/arxiv.org\/abs\/2401.18079"},{"key":"e_1_3_2_23_2","volume-title":"Proceedings of the Synthetic Data for Computer Vision Workshop @ CVPR 2025","author":"Humes Edward Steven","year":"2025","unstructured":"Edward Steven Humes, Xiaomin Lin, Utteja Kallakuri, and Tinoosh Mohsenin. 2025. RAFT: Robust augmentation of FeaTures for image segmentation. In Proceedings of the Synthetic Data for Computer Vision Workshop @ CVPR 2025. Retrieved from https:\/\/openreview.net\/forum?id=uHxvIXPT2c"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","first-page":"694","DOI":"10.1145\/3649476.3658699","volume-title":"Proceedings of the Great Lakes Symposium on VLSI 2024","author":"Kallakuri Uttej","year":"2024","unstructured":"Uttej Kallakuri, Edward Humes, and Tinoosh Mohsenin. 2024. Resource-aware saliency-guided differentiable pruning for deep neural networks. In Proceedings of the Great Lakes Symposium on VLSI 2024. 694\u2013699."},{"key":"e_1_3_2_25_2","volume-title":"Proceedings of the 1st Vision and Language for Autonomous Driving and Robotics Workshop","author":"Kallakuri Utteja","year":"2024","unstructured":"Utteja Kallakuri, Bharat Prakash, Arnab Neelim Mazumder, Hasib-Al Rashid, Nicholas R. Waytowich, and Tinoosh Mohsenin. 2024. ATLAS: Adaptive landmark acquisition using LLM-guided navigation. In Proceedings of the 1st Vision and Language for Autonomous Driving and Robotics Workshop. Retrieved from https:\/\/openreview.net\/forum?id=VhpxzSWTWj"},{"key":"e_1_3_2_26_2","article-title":"Enabling on-device medical AI assistants via input-driven saliency adaptation","author":"Kallakurik Uttej","year":"2025","unstructured":"Uttej Kallakurik, Edward Humes, Rithvik Jonna, Xiaomin Lin, and Tinoosh Mohsenin. 2025. Enabling on-device medical AI assistants via input-driven saliency adaptation. In Proceedings of the IEEE Biomedical Circuits and Systems (BioCAS) Conference.","journal-title":"Proceedings of the IEEE Biomedical Circuits and Systems (BioCAS) Conference"},{"key":"e_1_3_2_27_2","unstructured":"Bo-Kyeong Kim Geonmin Kim Tae-Ho Kim Thibault Castells Shinkook Choi Junho Shin and Hyoung-Kyu Song. 2024. Shortened llama: A simple depth pruning for large language models. arXiv:2402.02834. Retrieved from https:\/\/arxiv.org\/abs\/2402.02834"},{"key":"e_1_3_2_28_2","unstructured":"Sehoon Kim Coleman Hooper Amir Gholami Zhen Dong Xiuyu Li Sheng Shen Michael W. Mahoney and Kurt Keutzer. 2023. SqueezeLLM: Dense-and-sparse quantization. arXiv:2306.07629. Retrieved from https:\/\/arxiv.org\/abs\/2306.07629"},{"key":"e_1_3_2_29_2","doi-asserted-by":"crossref","unstructured":"Eldar Kurtic Daniel Campos Tuan Nguyen Elias Frantar Mark Kurtz Benjamin Fineran Michael Goin and Dan Alistarh. 2022. The optimal bert surgeon: Scalable and accurate second-order pruning for large language models. arXiv:2203.07259. Retrieved from https:\/\/arxiv.org\/abs\/2203.07259","DOI":"10.18653\/v1\/2022.emnlp-main.279"},{"key":"e_1_3_2_30_2","unstructured":"Z Lan. 2019. Albert: A lite bert for self-supervised learning of language representations. arXiv:1909.11942. Retrieved from https:\/\/arxiv.org\/abs\/1909.11942"},{"key":"e_1_3_2_31_2","unstructured":"Yann LeCun John Denker and Sara Solla. 1989. Optimal brain damage. In Advances in Neural Information Processing Systems D. Touretzky (Ed.). Morgan-Kaufmann. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/1989\/file\/6c9882bbac1c7093bd25041881277658-Paper.pdf"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btz682"},{"key":"e_1_3_2_33_2","unstructured":"M Lewis. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation translation and comprehension. arXiv:1910.13461. Retrieved from https:\/\/arxiv.org\/abs\/1910.13461"},{"key":"e_1_3_2_34_2","unstructured":"Yun Li Lin Niu Xipeng Zhang Kai Liu Jianchen Zhu and Zhanhui Kang. 2023. E-sparse: Boosting the large language model inference through entropy-based N: M sparsity. arXiv:2310.15929. Retrieved from https:\/\/arxiv.org\/abs\/2310.15929"},{"key":"e_1_3_2_35_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Li Zefan","year":"2021","unstructured":"Zefan Li, Ziwei Lin, Shupeng Liu, Aojun Zhou, Xiaowei Wang, and Yiren Lin. 2021. BRECQ: Pushing the limit of post-training quantization by block reconstruction. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_36_2","unstructured":"Stephanie Lin Jacob Hilton and Owain Evans. 2021. TruthfulQA: Measuring how models mimic human falsehoods. arXiv:2109.07958. Retrieved from https:\/\/arxiv.org\/abs\/2109.07958"},{"key":"e_1_3_2_37_2","unstructured":"Yinhan Liu. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_2_38_2","unstructured":"Yijiang Liu Huanrui Yang Youxin Chen Rongyu Zhang Miao Wang Yuan Du and Li Du. 2024. PAT: Pruning-aware tuning for large language models. arXiv:2408.14721. Retrieved from https:\/\/arxiv.org\/abs\/2408.14721"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-acl.425"},{"key":"e_1_3_2_40_2","unstructured":"Zechun Liu Barlas Oguz Changsheng Zhao Ernie Chang Pierre Stock Yashar Mehdad Yangyang Shi Raghuraman Krishnamoorthi and Vikas Chandra. 2023. LLM-QAT: Data-free quantization aware training for large language models. arXiv:2305.17888. Retrieved from https:\/\/arxiv.org\/abs\/2305.17888"},{"key":"e_1_3_2_41_2","unstructured":"Zirui Liu Jiayi Yuan Hongye Jin Shaochen Zhong Zhaozhuo Xu Vladimir Braverman Beidi Chen and Xia Hu. 2024. KIVI: A tuning-free asymmetric 2bit quantization for KV cache. arXiv:2402.02750. Retrieved from https:\/\/arxiv.org\/abs\/2402.02750"},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","unstructured":"Xinyin Ma Gongfan Fang and Xinchao Wang. 2023. LLM-Pruner: On the structural pruning of large language models. In Advances in Neural Information Processing Systems Curran Associates Inc. 21702\u201321720. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2023\/file\/44956951349095f74492a5471128a7e0-Paper-Conference.pdf","DOI":"10.52202\/075280-0950"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2021.3129415"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3623380"},{"key":"e_1_3_2_45_2","unstructured":"Stephen Merity Caiming Xiong James Bradbury and Richard Socher. 2016. Pointer sentinel mixture models. arXiv:1609.07843. Retrieved from https:\/\/arxiv.org\/abs\/1609.07843"},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","unstructured":"Todor Mihaylov Peter Clark Tushar Khot and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? A new dataset for open book question answering. arXiv:1809.02789. Retrieved from https:\/\/arxiv.org\/abs\/1809.02789","DOI":"10.18653\/v1\/D18-1260"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01152"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.5555\/3524938.3525605"},{"key":"e_1_3_2_49_2","first-page":"180","volume-title":"Proceedings of the AAAI Spring Symposium Series","volume":"5","author":"Navardi Mozhgan","year":"2025","unstructured":"Mozhgan Navardi, Romina Aalishah, Yuzhe Fu, Yueqian Lin, Hai Li, Yiran Chen, and Tinoosh Mohsenin. 2025. GenAI at the edge: Comprehensive survey on empowering edge devices. In Proceedings of the AAAI Spring Symposium Series, Vol. 5. 180\u2013187."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/LES.2024.3446948"},{"key":"e_1_3_2_51_2","unstructured":"Haojie Pan Chengyu Wang Minghui Qiu Yichang Zhang Yaliang Li and Jun Huang. 2020. Meta-KD: A meta knowledge distillation framework for language model compression across domains. arXiv:2012.01266. Retrieved from https:\/\/arxiv.org\/abs\/2012.01266"},{"key":"e_1_3_2_52_2","unstructured":"Gunho Park Baeseong Park Minsub Kim Sungjae Lee Jeonghoon Kim Beomseok Kwon Se Jung Kwon Byeongwook Kim Youngjoo Lee and Dongsoo Lee. 2022. LUT-GEMM: Quantized matrix multiplication based on LUTs for efficient inference in large-scale generative language models. arXiv:2206.09557. Retrieved from https:\/\/arxiv.org\/abs\/2206.09557"},{"issue":"8","key":"e_1_3_2_53_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et\u00a0al. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_2_54_2","doi-asserted-by":"crossref","unstructured":"Hasib-Al Rashid Pretom Roy Ovi Carl Busart Aryya Gangopadhyay and Tinoosh Mohsenin. 2022. Tinym2net: A flexible system algorithm co-designed multimodal learning framework for tiny devices. ArXiv (2022) arXiv\u20132202.","DOI":"10.31219\/osf.io\/e8px7"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3595633"},{"key":"e_1_3_2_56_2","first-page":"1","volume-title":"Proceedings of the 2025 IEEE Conference on Technologies for Sustainability (SusTech)","author":"Rashid Hasib-Al","year":"2025","unstructured":"Hasib-Al Rashid, Eiman Kanjo, and Tinoosh Mohsenin. 2025. HAC-M-DNN: Hardware aware compression of sustainable multimodal deep neural networks for efficient TinyML deployment. In Proceedings of the 2025 IEEE Conference on Technologies for Sustainability (SusTech). IEEE, 1\u20137."},{"key":"e_1_3_2_57_2","volume-title":"Proceedings of the 2nd Workshop on Sustainable AI","author":"Rashid Hasib-Al","year":"2024","unstructured":"Hasib-Al Rashid and Tinoosh Mohsenin. 2024. TinyM \u23032 Net-V3: Memory-aware compressed multimodal deep neural networks for sustainable edge deployment. In Proceedings of the 2nd Workshop on Sustainable AI."},{"key":"e_1_3_2_58_2","unstructured":"Hasib-Al Rashid Argho Sarkar Aryya Gangopadhyay Maryam Rahnemoonfar and Tinoosh Mohsenin. 2024. TinyVQA: Compact multimodal deep neural network for visual question answering on resource-constrained devices. arXiv preprint arXiv:2404.03574 (2024)."},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474381"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","unstructured":"Md Ragib Shaharear Arnab Neelim Mazumder and Tinoosh Mohsenin. 2025. ViT-Reg: Regression-focused hardware-aware fine-tuning for ViT on TinyML platforms. IEEE Design & Test 42 5 (2025) 35\u201344. DOI:10.1109\/MDAT.2024.3521320","DOI":"10.1109\/MDAT.2024.3521320"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP48485.2024.10445737"},{"key":"e_1_3_2_62_2","unstructured":"Wenqi Shao Mengzhao Chen Zhaoyang Zhang Peng Xu Lirui Zhao Zhiqian Li Kaipeng Zhang Peng Gao Yu Qiao and Ping Luo. 2023. Omniquant: Omnidirectionally calibrated quantization for large language models. arXiv:2308.13137. Retrieved from https:\/\/arxiv.org\/abs\/2308.13137"},{"key":"e_1_3_2_63_2","unstructured":"Aarohi Srivastava Abhinav Rastogi Abhishek Rao Abu Awal Md Shoeb Abubakar Abid Adam Fisch Adam R. Brown Adam Santoro Aditya Gupta Adri\u00e0 Garriga-Alonso et\u00a0al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv:2206.04615. Retrieved from https:\/\/arxiv.org\/abs\/2206.04615"},{"key":"e_1_3_2_64_2","unstructured":"Mingjie Sun Zhuang Liu Anna Bair and J. Zico Kolter. 2024. A simple and effective pruning approach for large language models. arxiv:cs.CL\/2306.11695. Retrieved from https:\/\/arxiv.org\/abs\/2306.11695"},{"key":"e_1_3_2_65_2","unstructured":"Siqi Sun Yu Cheng Zhe Gan and Jingjing Liu. 2019. Patient knowledge distillation for bert model compression. arXiv:1908.09355. Retrieved from https:\/\/arxiv.org\/abs\/1908.09355"},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","unstructured":"Siqi Sun Zhe Gan Yu Cheng Yuwei Fang Shuohang Wang and Jingjing Liu. 2020. Contrastive distillation on intermediate representations for language model compression. arXiv:2009.14167. Retrieved from https:\/\/arxiv.org\/abs\/2009.14167","DOI":"10.18653\/v1\/2020.emnlp-main.36"},{"key":"e_1_3_2_67_2","unstructured":"Zhiqing Sun Hongkun Yu Xiaodan Song Renjie Liu Yiming Yang and Denny Zhou. 2020. Mobilebert: A compact task-agnostic bert for resource-limited devices. arXiv:2004.02984. Retrieved from https:\/\/arxiv.org\/abs\/2004.02984"},{"key":"e_1_3_2_68_2","unstructured":"Alon Talmor Jonathan Herzig Nicholas Lourie and Jonathan Berant. 2018. CommonsenseQA: A question answering challenge targeting commonsense knowledge. arXiv:1811.00937. Retrieved from https:\/\/arxiv.org\/abs\/1811.00937"},{"key":"e_1_3_2_69_2","unstructured":"Rohan Taori Ishaan Gulrajani Tianyi Zhang Yann Dubois Xuechen Li Carlos Guestrin Percy Liang and Tatsunori B. Hashimoto. 2023. Stanford Alpaca: An Instruction-following LLaMA model. GitHub Repository."},{"key":"e_1_3_2_70_2","unstructured":"Gemma Team Thomas Mesnard Cassidy Hardin Robert Dadashi Surya Bhupatiraju Shreya Pathak Laurent Sifre Morgane Rivi\u00e8re Mihir Sanjay Kale Juliette Love et\u00a0al. 2024. Gemma: Open models based on Gemini research and technology. arXiv:2403.08295. Retrieved from https:\/\/arxiv.org\/abs\/2403.08295"},{"key":"e_1_3_2_71_2","article-title":"Attention is all you need","author":"Vaswani A.","year":"2017","unstructured":"A. Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_72_2","first-page":"196","volume-title":"Proceedings of the AAAI Spring Symposium Series","volume":"5","author":"Walczak Mikolaj","year":"2025","unstructured":"Mikolaj Walczak, Uttej Kallakuri, and Tinoosh Mohsenin. 2025. ATLASv2: LLM-guided adaptive landmark acquisition and navigation on the edge. In Proceedings of the AAAI Spring Symposium Series, Vol. 5. 196\u2013203."},{"key":"e_1_3_2_73_2","unstructured":"Xiuying Wei Yunchen Zhang Yuhang Li Xiangguo Zhang Ruihao Gong Jinyang Guo and Xianglong Liu. 2023. Outlier suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling. arXiv:2304.09145. Retrieved from https:\/\/arxiv.org\/abs\/2304.09145"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.14778\/3626292.3626303"},{"key":"e_1_3_2_75_2","unstructured":"Mengzhou Xia Tianyu Gao Zhiyuan Zeng and Danqi Chen. 2023. Sheared LLaMA: Accelerating language model pre-training via structured pruning. arXiv:2310.06694. Retrieved from https:\/\/arxiv.org\/abs\/2310.06694"},{"key":"e_1_3_2_76_2","first-page":"38087","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xiao Guangxuan","year":"2023","unstructured":"Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. Smoothquant: Accurate and efficient post-training quantization for large language models. In Proceedings of the International Conference on Machine Learning. PMLR, 38087\u201338099."},{"key":"e_1_3_2_77_2","unstructured":"Dongkuan Xu Ian E. H. Yen Jinxi Zhao and Zhibin Xiao. 2021. Rethinking network pruning\u2013under the pre-train and fine-tune paradigm. arXiv:2104.08682. Retrieved from https:\/\/arxiv.org\/abs\/2104.08682"},{"key":"e_1_3_2_78_2","unstructured":"Yuzhuang Xu Xu Han Zonghan Yang Shuo Wang Qingfu Zhu Zhiyuan Liu Weidong Liu and Wanxiang Che. 2024. OneBit: Towards extremely low-bit large language models. arXiv:2402.11295. Retrieved from https:\/\/arxiv.org\/abs\/2402.11295"},{"key":"e_1_3_2_79_2","doi-asserted-by":"crossref","unstructured":"Zhewei Yao Reza Yazdani Aminabadi Minjia Zhang Xiaoxia Wu Conglong Li and Yuxiong He. 2022. ZeroQuant: Efficient and affordable post-training quantization for large-scale transformers. In Advances in Neural Information Processing Systems Curran Associates Inc. 27168\u201327183. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/adf7fa39d65e2983d724ff7da57f00ac-Paper-Conference.pdf","DOI":"10.52202\/068431-1970"},{"key":"e_1_3_2_80_2","first-page":"27168","article-title":"Zeroquant: Efficient and affordable post-training quantization for large-scale transformers","volume":"35","author":"Yao Zhewei","year":"2022","unstructured":"Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. 2022. Zeroquant: Efficient and affordable post-training quantization for large-scale transformers. Advances in Neural Information Processing Systems 35 (2022), 27168\u201327183.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_81_2","unstructured":"Deming Ye Yankai Lin Yufei Huang and Maosong Sun. 2021. Tr-bert: Dynamic token reduction for accelerating bert inference. arXiv:2105.11618. Retrieved from https:\/\/arxiv.org\/abs\/2105.11618"},{"key":"e_1_3_2_82_2","unstructured":"Zhihang Yuan Lin Niu Jiawei Liu Wenyu Liu Xinggang Wang Yuzhang Shang Guangyu Sun Qiang Wu Jiaxiang Wu and Bingzhe Wu. 2023. RPTQ: Reorder-based post-training quantization for large language models. arXiv:2304.01089. Retrieved from https:\/\/arxiv.org\/abs\/2304.01089"},{"key":"e_1_3_2_83_2","unstructured":"Yuxuan Yue Zhihang Yuan Haojie Duanmu Sifan Zhou Jianlong Wu and Liqiang Nie. 2024. WKVQuant: Quantizing weight and key\/value cache for large language models gains more. arXiv:2402.12065. Retrieved from https:\/\/arxiv.org\/abs\/2402.12065"},{"key":"e_1_3_2_84_2","first-page":"36","volume-title":"Proceedings of the 2019 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition (EMC2-NIPS)","author":"Zafrir Ofir","year":"2019","unstructured":"Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019. Q8bert: Quantized 8bit bert. In Proceedings of the 2019 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition (EMC2-NIPS). IEEE, 36\u201339."},{"key":"e_1_3_2_85_2","doi-asserted-by":"crossref","unstructured":"Rowan Zellers Ari Holtzman Yonatan Bisk Ali Farhadi and Yejin Choi. 2019. HellaSwag: Can a machine really finish your sentence? arXiv:1905.07830. Retrieved from https:\/\/arxiv.org\/abs\/1905.07830","DOI":"10.18653\/v1\/P19-1472"},{"key":"e_1_3_2_86_2","unstructured":"Susan Zhang Stephen Roller Naman Goyal Mikel Artetxe Moya Chen Shuohui Chen Christopher Dewan Mona Diab Xian Li Xi Victoria Lin et\u00a0al. 2022. OPT: Open pre-trained transformer language models. arxiv:cs.CL\/2205.01068. Retrieved from https:\/\/arxiv.org\/abs\/2205.01068"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3766068","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T13:58:11Z","timestamp":1778680691000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3766068"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,13]]},"references-count":85,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3766068"],"URL":"https:\/\/doi.org\/10.1145\/3766068","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,13]]},"assertion":[{"value":"2024-11-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-11","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}