{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T16:10:56Z","timestamp":1784736656471,"version":"3.55.0"},"reference-count":173,"publisher":"Association for Computing Machinery (ACM)","issue":"10","license":[{"start":{"date-parts":[[2025,5,6]],"date-time":"2025-05-06T00:00:00Z","timestamp":1746489600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-sa\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2025,10,31]]},"abstract":"<jats:p>Efficient compression and tuning techniques have become indispensable in addressing the increasing computational and memory demands of large language models (LLMs). While these models have demonstrated exceptional performance across a wide range of natural language processing tasks, their growing size and resource requirements pose significant challenges to accessibility and sustainability. This survey systematically reviews state-of-the-art methods in model compression, including compression techniques such as knowledge distillation, low-rank approximation, parameter pruning, and quantization, as well as tuning techniques such as parameter-efficient fine-tuning and inference optimization. Compression techniques, though well-established in traditional deep learning, require updated methodologies tailored to the scale and dynamics of LLMs. Simultaneously, parameter-efficient fine-tuning, exemplified by techniques like Low-Rank Adaptation (LoRA) and query tuning, emerges as a promising solution for adapting models with minimal resource overhead. This study provides a detailed taxonomy of these methods, examining their practical applications, strengths, and limitations. Critical gaps are identified in scalability, and the integration of compression and tuning strategies, signaling the need for unified frameworks and hybrid approaches to maximize efficiency and performance. By addressing these challenges, this survey aims at guiding researchers toward sustainable, efficient, and accessible LLM development, ensuring their broader applicability across diverse domains while mitigating resource constraints.<\/jats:p>","DOI":"10.1145\/3728636","type":"journal-article","created":{"date-parts":[[2025,4,8]],"date-time":"2025-04-08T11:44:54Z","timestamp":1744112694000},"page":"1-39","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Efficient Compressing and Tuning Methods for Large Language Models: A Systematic Literature Review"],"prefix":"10.1145","volume":"57","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-7126-1245","authenticated-orcid":false,"given":"Gun Il","family":"Kim","sequence":"first","affiliation":[{"name":"Graduate School of Information, Yonsei University, Seoul, Korea (the Republic of)"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1887-7393","authenticated-orcid":false,"given":"Sunga","family":"Hwang","sequence":"additional","affiliation":[{"name":"Graduate School of Information, Yonsei University, Seoul, Korea (the Republic of)"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3911-5935","authenticated-orcid":false,"given":"Beakcheol","family":"Jang","sequence":"additional","affiliation":[{"name":"Graduate School of Information, Yonsei University, Seoul, Korea (the Republic of)"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,5,6]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"1","volume-title":"International Joint Conference on Neural Networks, Rio de Janeiro, Brazil, July 8\u201313, 2018","year":"2018","unstructured":"Md. Zahangir Alom, Adam T. Moody, Naoya Maruyama, Brian C. Van Essen, and Tarek M. Taha. 2018. Effective quantization approaches for recurrent neural networks. In International Joint Conference on Neural Networks, Rio de Janeiro, Brazil, July 8\u201313, 2018. 1\u20138."},{"key":"e_1_3_1_3_2","volume-title":"International Conference for High Performance Computing, Networking, Storage and Analysis, Dallas, TX, USA, November 13\u201318, 2022","year":"2022","unstructured":"Reza Yazdani Aminabadi, Samyam Rajbhandari, Minjia Zhang, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Jeff Rasley, Shaden Smith, Olatunji Ruwase, and Yuxiong He. 2022. DeepSpeed-Inference: Enabling efficient inference of transformer models at unprecedented scale. In International Conference for High Performance Computing, Networking, Storage and Analysis, Dallas, TX, USA, November 13\u201318, 2022."},{"key":"e_1_3_1_4_2","volume-title":"38th AAAI Conference on Artificial Intelligence, 36th Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, 14th Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20\u201327, 2024, Vancouver, Canada","year":"2024","unstructured":"Yongqi An, Xu Zhao, Tao Yu, Ming Tang, and Jinqiao Wang. 2024. Fluctuation-based adaptive structured pruning for large language models. In 38th AAAI Conference on Artificial Intelligence, 36th Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, 14th Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20\u201327, 2024, Vancouver, Canada."},{"key":"e_1_3_1_5_2","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022","year":"2022","unstructured":"Akari Asai, Mohammadreza Salehi, Matthew E. Peters, and Hannaneh Hajishirzi. 2022. ATTEMPT: Parameter-efficient multi-task tuning via attentional mixtures of soft prompts. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022."},{"key":"e_1_3_1_6_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","year":"2024","unstructured":"Saleh Ashkboos, Maximilian L. Croci, Marcelo Gennari do Nascimento, Torsten Hoefler, and James Hensman. 2024. SliceGPT: Compress large language models by deleting rows and columns. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_7_2","volume-title":"Proceedings of the 14th Conference on Machine Translation, WMT 2019, Florence, Italy, 2019 - Volume 2: Shared Task Papers","year":"2019","unstructured":"Loic Barrault, Ondrej Bojar, Marta Ruiz Costa-jussa, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Muller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019. Findings of the 2019 conference on machine translation (WMT19). In Proceedings of the 14th Conference on Machine Translation, WMT 2019, Florence, Italy, 2019 - Volume 2: Shared Task Papers."},{"key":"e_1_3_1_8_2","volume-title":"Proceedings of the 20th Workshop on Biomedical Language Processing","year":"2021","unstructured":"Asma Ben Abacha, Yassine Mrabet, Yuhao Zhang, Chaitanya P. Shivade, Curt P. Langlotz, and Dina Demner-Fushman. 2021. Overview of the MEDIQA 2021 shared task on summarization in the medical domain. In Proceedings of the 20th Workshop on Biomedical Language Processing."},{"key":"e_1_3_1_9_2","unstructured":"Yoshua Bengio Nicholas Leonard and Aaron C. Courville. 2013. Estimating or propagating gradients through stochastic neurons for conditional computation. https:\/\/arxiv.org\/abs\/1308.3432"},{"key":"e_1_3_1_10_2","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, Canada, July 9\u201314, 2023","year":"2023","unstructured":"Chandra Bhagavatula, Jena D. Hwang, Doug Downey, Ronan Le Bras, Ximing Lu, Keisuke Sakaguchi, Swabha Swayamdipta, Peter West, and Yejin Choi. 2023. I2D2: Inductive knowledge distillation with NeuroLogic and self-imitation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, Canada, July 9\u201314, 2023."},{"key":"e_1_3_1_11_2","volume-title":"Proceedings of the International Conference on Machine Learning, 23\u201329 July 2023, Honolulu, Hawaii, USA","year":"2023","unstructured":"Stella Biderman, Hailey Schoelkopf, Quentin G. Anthony, Herbie Bradley, Kyle O\u2019Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. 2023. Pythia: A suite for analyzing large language models across training and scaling. In Proceedings of the International Conference on Machine Learning, 23\u201329 July 2023, Honolulu, Hawaii, USA."},{"key":"e_1_3_1_12_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Ma-teusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020."},{"key":"e_1_3_1_13_2","unstructured":"Arnav Chavan Zhuang Liu Deepak K. Gupta Eric P. Xing and Zhiqiang Shen. 2023. One-for-All: Generalized LoRA for Parameter-Efficient Fine-tuning. https:\/\/arxiv.org\/abs\/2306.07967"},{"key":"e_1_3_1_14_2","volume-title":"Proceedings of the 7th Annual Conference on Machine Learning and Systems, Santa Clara, CA, USA, May 13\u201316, 2024","year":"2024","unstructured":"Lequn Chen, Zihao Ye, Yongji Wu, Danyang Zhuo, Luis Ceze, Arvind Krishnamurthy, and Duke University. 2024. Punica: Multi-Tenant LoRA serving. In Proceedings of the 7th Annual Conference on Machine Learning and Systems, Santa Clara, CA, USA, May 13\u201316, 2024."},{"key":"e_1_3_1_15_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde Jared Kaplan Harrison Edwards Yura Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mo Bavarian Clemens Winter Philippe Tillet Felipe Petroski Such David W. Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William H. Guss Alex Nichol Igor Babuschkin Suchir Balaji Shantanu Jain Andrew Carr Jan Leike Joshua Achiam Vedant Misra Evan Morikawa Alec Radford Matthew M. Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating large language models trained on code. https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_1_16_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, December 6\u201314, 2021, virtual","year":"2021","unstructured":"Patrick H. Chen, Hsiang-Fu Yu, Inderjit S. Dhillon, and Cho-Jui Hsieh. 2021. DRONE: Data-aware low-rank compression for large NLP models. In Proceedings of the Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, December 6\u201314, 2021, virtual."},{"key":"e_1_3_1_17_2","volume-title":"Proceedings of the International Conference on Machine Learning, 23\u201329 July 2023, Honolulu, Hawaii, USA","year":"2023","unstructured":"Wuyang Chen, Yan-Quan Zhou, Nan Du, Yanping Huang, James Laudon, Z. Chen, and Claire Cu. 2023. Lifelong language pretraining with distribution-specialized experts. In Proceedings of the International Conference on Machine Learning, 23\u201329 July 2023, Honolulu, Hawaii, USA."},{"key":"e_1_3_1_18_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","year":"2024","unstructured":"Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia. 2024. LongLoRA: Efficient fine-tuning of long-context large language models. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_19_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\u201311, 2024","author":"Cheng Daixuan","year":"2024","unstructured":"Daixuan Cheng, Shaohan Huang, and Furu Wei. 2024. Adapting large language models via reading comprehension. In Proceedings of the 12th International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_20_2","unstructured":"Karl Cobbe Vineet Kosaraju Mo Bavarian Mark Chen Heewoo Jun Lukasz Kaiser Matthias Plappert Jerry Tworek Jacob Hilton Reiichiro Nakano Christopher Hesse and John Schulman. 2021. Training verifiers to solve math word problems. https:\/\/arxiv.org\/abs\/2110.14168"},{"key":"e_1_3_1_21_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations, San Diego, CA, USA, May 7\u20139, 2015, Workshop Track Proceedings","year":"2015","unstructured":"Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. 2015. Low precision arithmetic for deep learning. In Proceedings of the 3rd International Conference on Learning Representations, San Diego, CA, USA, May 7\u20139, 2015, Workshop Track Proceedings."},{"key":"e_1_3_1_22_2","first-page":"1280","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, August 11\u201316, 2024","year":"2024","unstructured":"Damai Dai, Chengqi Deng, Chenggang Zhao, Runxin Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Yu Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. 2024. DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, August 11\u201316, 2024. Association for Computational Linguistics, 1280\u20131297."},{"key":"e_1_3_1_23_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","author":"Dao Tri","year":"2024","unstructured":"Tri Dao. 2024. FlashAttention-2: Faster attention with better parallelism and work partitioning. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_24_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, New Orleans, LA, USA, November 28\u2013December 9, 2022","author":"Dao Tri","year":"2022","unstructured":"Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher R\u00e9. 2022. FlashAttention: Fast and memory-efficient exact attention with IO-Awareness. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, New Orleans, LA, USA, November 28\u2013December 9, 2022."},{"key":"e_1_3_1_25_2","unstructured":"Rocktim Jyoti Das Liqun Ma and Zhiqiang Shen. 2023. Beyond size: How gradients shape pruning decisions in large language models. https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_1_26_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, New Orleans, LA, USA, November 28\u2013December 9, 2022","year":"2022","unstructured":"Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. LLM.int8(): 8-bit matrix multiplication for transformers at scale. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, New Orleans, LA, USA, November 28\u2013December 9, 2022."},{"key":"e_1_3_1_27_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, New Orleans, LA, USA, December 10\u201316, 2023","year":"2023","unstructured":"Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient finetuning of quantized LLMs. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, New Orleans, LA, USA, December 10\u201316, 2023."},{"key":"e_1_3_1_28_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","year":"2024","unstructured":"Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. 2024. SpQR: A sparse-quantized representation for near-lossless LLM weight compression. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_29_2","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, June 2\u20137, 2019, Volume 1 (Long and Short Papers)","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, June 2\u20137, 2019, Volume 1 (Long and Short Papers)."},{"key":"e_1_3_1_30_2","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, December 6\u201310, 2023","year":"2023","unstructured":"Ning Ding, Xingtai Lv, Qiaosen Wang, Yulin Chen, Bowen Zhou, Zhiyuan Liu, and Maosong Sun. 2023. Sparse low-rank adaptation of pre-trained language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, December 6\u201310, 2023."},{"key":"e_1_3_1_31_2","volume-title":"Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, July 21\u201327, 2024","year":"2024","unstructured":"Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang. 2024. LongRoPE: Extending LLM context window beyond 2 million tokens. In Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, July 21\u201327, 2024."},{"key":"e_1_3_1_32_2","volume-title":"Proceedings of the International Conference on Machine Learning, 17\u201323 July 2022, Baltimore, Maryland, USA","year":"2022","unstructured":"Nan Du, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, Liam Fedus, Maarten Bosma, Zongwei Zhou, Tao Wang, Yu Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathleen S. Meier-Hellstern, Toju Duke, Lucas Dixon, Kun Zhang, Quoc V. Le, Yonghui Wu, Z. Chen, and Claire Cui. 2022. GLaM: Efficient scaling of language models with mixture-of-experts. In Proceedings of the International Conference on Machine Learning, 17\u201323 July 2022, Baltimore, Maryland, USA."},{"key":"e_1_3_1_33_2","unstructured":"Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Amy Yang Angela Fan Anirudh Goyal Anthony S. Hartshorn Aobo Yang Archi Mitra Archie Sravankumar and Artem Korenev. 2024. The Llama 3 herd of models. Retrieved from https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_1_34_2","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Dublin, Ireland, May 22\u201327, 2022","year":"2022","unstructured":"Ali Edalati, Marzieh S. Tahaei, Ahmad Rashid, V. Nia, James J. Clark, and Mehdi Rezagholizadeh. 2022. Kronecker decomposition for GPT compression. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Dublin, Ireland, May 22\u201327, 2022."},{"key":"e_1_3_1_35_2","unstructured":"Elias Frantar Saleh Ashkboos Torsten Hoefler and Dan Alistarh. 2023. GPTQ: Accurate post-training quantization for generative pre-trained transformers. In The Eleventh International Conference on Learning Representations 2023. Kigali Rwanda May 1 - May 5 2023. 1\u201316. https:\/\/arxiv.org\/abs\/2210.17323"},{"key":"e_1_3_1_36_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, New Orleans, LA, USA, November 28\u2013December 9, 2022","author":"Frantar Elias","year":"2022","unstructured":"Elias Frantar and Dan Alistarh. 2022. Optimal brain compression: A framework for accurate post-training quantization and pruning. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, New Orleans, LA, USA, November 28\u2013December 9, 2022."},{"key":"e_1_3_1_37_2","volume-title":"Proceedings of the International Conference on Machine Learning, 23\u201329 July, 2023, Honolulu, Hawaii, USA","author":"Frantar Elias","year":"2023","unstructured":"Elias Frantar and Dan Alistarh. 2023. SparseGPT: Massive language models can be accurately pruned in one-shot. In Proceedings of the International Conference on Machine Learning, 23\u201329 July, 2023, Honolulu, Hawaii, USA."},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","unstructured":"Amir Gholami Sehoon Kim Zhen Dong Zhewei Yao Michael W. Mahoney and Kurt Keutzer. 2021. A survey of quantization methods for efficient neural network inference. https:\/\/arxiv.org\/abs\/2103.13630","DOI":"10.1201\/9781003162810-13"},{"key":"e_1_3_1_39_2","unstructured":"Ian J. Goodfellow Jonathon Shlens and Christian Szegedy. 2015. Explaining and harnessing adversarial examples. In 3rd International Conference on Learning Representations ICLR 2015 San Diego CA USA May 7-9 2015 Conference Track Proceedings. https:\/\/arxiv.org\/abs\/1412.6572"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01453-z"},{"key":"e_1_3_1_41_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 2024","year":"2024","unstructured":"Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2024. MiniLLM: Knowledge distillation of large language models. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 2024."},{"key":"e_1_3_1_42_2","volume-title":"Proceedings of the 50th Annual International Symposium on Computer Architecture, Orlando, FL, USA, June 17\u201321, 2023","year":"2023","unstructured":"Cong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng, Chen Zhang, Fan Yang, Yun-Bo Liu, Minyi Guo, and Yuhao Zhu. 2023. OliVe: Accelerating large language models via hardware-friendly outlier-victim pair quantization. In Proceedings of the 50th Annual International Symposium on Computer Architecture, Orlando, FL, USA, June 17\u201321, 2023."},{"key":"e_1_3_1_43_2","unstructured":"DeepSeek-AI Daya Guo Dejian Yang Haowei Zhang Junxiao Song Ruizhe Zhang Runxin Xu Qihao Zhu Shirong Ma Peiyi Wang Xiao Bi Xiaojin Zhang Xingkai Yu Yuxuan Wu Zhenfeng Wu Zhe Gou Zhihong Shao Zilin Li Zhenda Gao Aixin Liu Bei Xue Bin Wang Bingxuan Wang Bo Liu Chengda Lu Chenggang Zhao Chengqi Deng Chenyu Zhang Chong Ruan Damai Dai Deli Chen Dongjie Ji Erhang Li Fangyun Lin Fucong Dai Fuli Luo Guangbo Hao Guanting Chen Guowei Li Han Bao Hanwei Xu Haocheng Wang Haowei Zhang Honghui Ding Huajian Xin Huazuo Gao Hui Li Hui Qu Jian Liang Jiaqi Ni Jianzhong Guo Jia Li Jiashi Li Jin Chen Jingyang Yuan Junjie Qiu Kai Dong Kaige Gao Kang Guan Lean Wang Lecong Zhang Lei Xu Leyi Xia Liang Zhao Liyue Zhang Meng Li Miaojun Wang Mingchuan Zhang Minghua Zhang Minghui Tang Mingming Li Ning Tian Panpan Huang Peiyi Wang Peng Zhang Qihao Zhu Qinyu Chen Qiushi Du Ruiqi Ge Ruizhe Pan Runxin Xu Ruyi Chen Shanghao Lu Shangyan Zhou Shanhuang Chen Shaoqing Wu Shengfeng Ye Shirong Ma Shiyu Wang Shuang Zhou Shuiping Yu Shunfeng Zhou Size Zheng Tian Pei Tian Yuan Tianyu Sun Wangding Zeng Wei An Wen Liu Wenfeng Liang Wenjun Gao Wentao Zhang Xiangyue Jin Xianzu Wang Xiao Bi Xiaodong Liu Xiaohan Wang Xiaojin Shen Xiaokang Chen Xiaosha Chen Xiaotao Nie Xiaowen Sun Xiaoxiang Wang Xin Liu Xin Xie Xingkai Yu Xinnan Song Xinyi Zhou Xinyu Yang Xuan Lu Xuecheng Su Yanhong Xu Yanping Huang Yao Li Yao Zhao Yaofeng Sun Yaohui Li Yaohui Wang Yi Zheng Yichao Zhang Yiliang Xiong Yilong Zhao Ying He Ying Tang Yishi Piao Yixin Dong Yixuan Tan Yiyuan Liu Yongji Wang Yongqiang Guo Yuchen Zhu Yuduan Wang Yuheng Zou Yukun Zha Yunxian Ma Yuting Yan Yuxiang You Yuxuan Liu Zehui Ren Zhangli Sha Zhe Fu Zhen Huang Zhen Zhang Zhenda Xie Zhewen Hao Zhihong Shao Zhiniu Wen Zhipeng Xu Zhongyu Zhang Zhuoshu Li Zihan Wang Zihui Gu Zilin Li Ziwei Xie and Ziheng Pan. 2025. DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. https:\/\/arxiv.org\/abs\/2501.12948"},{"key":"e_1_3_1_44_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","year":"2024","unstructured":"Han Guo, Philip Greengard, Eric P. Xing, and Yoon Kim. 2024. LQ-LoRA: Low-rank plus quantized matrix decomposition for efficient language model finetuning. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_45_2","volume-title":"Proceedings of the 8th Financial Technology and Natural Language Processing and the 1st Agent AI for Scenario Planning","year":"2024","unstructured":"Lubingzhi Guo, Javier Sanz-Cruzado, and Richard McCreadie. 2024. University of glasgow at the FinLLM challenge task: Adapting Llama for financial news abstractive summarization. In Proceedings of the 8th Financial Technology and Natural Language Processing and the 1st Agent AI for Scenario Planning. 127\u2013132."},{"key":"e_1_3_1_46_2","volume-title":"Proceedings of the Workshop on Efficient Systems for Foundation Models @ ICML2023","year":"2023","unstructured":"Kshitij Gupta, Benjamin Therien, Adam Ibrahim, Mats L. Richter, Quentin G. Anthony, Eugene Belilovsky, Irina Rish, and Timothee Lesort. 2023. Continual pre-training of large language models: How to (re)warm your model?. In Proceedings of the Workshop on Efficient Systems for Foundation Models @ ICML2023."},{"key":"e_1_3_1_47_2","volume-title":"Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2021, December 6\u201314, 2021, Virtual","year":"2021","unstructured":"Habib Hajimolahoseini, Mehdi Rezagholizadeh, Vahid Partovinia, Marzieh S. Tahaei, Omar Mohamed Awad, and Yang Liu. 2021. Compressing pre-trained language models using progressive low rank decomposition. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2021, December 6\u201314, 2021, Virtual."},{"key":"e_1_3_1_48_2","unstructured":"Song Han Huizi Mao and William J. Dally. 2016. Deep compression: Compressing deep neural network with pruning trained quantization and huffman coding. In 4th International Conference on Learning Representations ICLR 2016 San Juan Puerto Rico May 2-4 2016 Conference Track Proceedings. https:\/\/arxiv.org\/abs\/1510.00149"},{"key":"e_1_3_1_49_2","volume-title":"Findings of the Association for Computational Linguistics, Online Event, August 1\u20136, 2021","year":"2021","unstructured":"Tahmid Hasan, Abhik Bhattacharjee, Md Saiful Islam, Kazi Samin Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. XL-Sum: Large-scale multilingual abstractive summarization for 44 languages. In Findings of the Association for Computational Linguistics, Online Event, August 1\u20136, 2021."},{"key":"e_1_3_1_50_2","unstructured":"Qinyao He He Wen Shuchang Zhou Yuxin Wu Cong Yao Xinyu Zhou and Yuheng Zou. 2016. Effective quantization methods for recurrent neural networks. https:\/\/arxiv.org\/abs\/1611.10176"},{"key":"e_1_3_1_51_2","volume-title":"Findings of the Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022","year":"2022","unstructured":"Shwai He, Liang Ding, Daize Dong, Miao Zhang, and Dacheng Tao. 2022. SparseAdapter: An easy approach for improving the parameter-efficiency of adapters. In Findings of the Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022."},{"key":"e_1_3_1_52_2","volume-title":"Proceedings of the 9th International Conference on Learning Representations, Virtual Event, Austria, 2021","year":"2021","unstructured":"Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Xiaodong Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In Proceedings of the 9th International Conference on Learning Representations, Virtual Event, Austria, 2021."},{"key":"e_1_3_1_53_2","unstructured":"Geoffrey E. Hinton Oriol Vinyals and Jeffrey Dean. 2015. Distilling the knowledge in a neural network. https:\/\/arxiv.org\/abs\/1503.02531"},{"key":"e_1_3_1_54_2","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, Canada, July 9\u201314, 2023","year":"2023","unstructured":"Namgyu Ho, Laura Schmid, and Se-Young Yun. 2023. Large language models are reasoning teachers. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, Canada, July 9\u201314, 2023."},{"key":"e_1_3_1_55_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems, Vancouver, BC, Canada, December 10\u201315, 2024","year":"2024","unstructured":"Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. 2024. KVQuant: Towards 10 million context length LLM inference with KV cache quantization. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems, Vancouver, BC, Canada, December 10\u201315, 2024."},{"key":"e_1_3_1_56_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, 2020","year":"2020","unstructured":"Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, and Qun Liu. 2020. DynaBERT: Dynamic BERT with adaptive width and depth. In Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, 2020."},{"key":"e_1_3_1_57_2","volume-title":"Proceedings of the 36th International Conference on Machine Learning, 9\u201315 June 2019, Long Beach, California, USA","year":"2019","unstructured":"Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning, 9\u201315 June 2019, Long Beach, California, USA."},{"key":"e_1_3_1_58_2","volume-title":"Findings of the Association for Computational Linguistics, Toronto, Canada, July 9\u201314, 2023","year":"2023","unstructured":"Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander J. Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023. Distilling step-by-step! Outperforming larger language models with less training data and smaller model sizes. In Findings of the Association for Computational Linguistics, Toronto, Canada, July 9\u201314, 2023."},{"key":"e_1_3_1_59_2","volume-title":"Proceedings of the 10th International Conference on Learning Representations, Virtual Event, April 25\u201329, 2022","year":"2022","unstructured":"Yen-Chang Hsu, Ting Hua, Sung-En Chang, Qiang Lou, Yilin Shen, and Hongxia Jin. 2022. Language model compression with weighted low-rank factorization. In Proceedings of the 10th International Conference on Learning Representations, Virtual Event, April 25\u201329, 2022."},{"key":"e_1_3_1_60_2","volume-title":"Proceedings of the 10th International Conference on Learning Representations, Virtual Event, April 25\u201329, 2022","year":"2022","unstructured":"J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In Proceedings of the 10th International Conference on Learning Representations, Virtual Event, April 25\u201329, 2022."},{"key":"e_1_3_1_61_2","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022","year":"2022","unstructured":"Ting Hua, Yen-Chang Hsu, Felicity Wang, Qiang Lou, Yilin Shen, and Hongxia Jin. 2022. Numerical optimizations for weighted low-rank estimation on language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022."},{"key":"e_1_3_1_62_2","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, August 11\u201316, 2024","year":"2024","unstructured":"Jianheng Huang, Leyang Cui, Ante Wang, Chengyi Yang, Xinting Liao, Linfeng Song, Junfeng Yao, and Jinsong Su. 2024. Mitigating catastrophic forgetting in large language models with self-synthesized rehearsal. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, August 11\u201316, 2024."},{"key":"e_1_3_1_63_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","year":"2024","unstructured":"Ajay Kumar Jaiswal, Zhe Gan, Xianzhi Du, Bowen Zhang, Zhangyang Wang, and Yinfei Yang. 2024. Compressing LLMs: The truth is rarely pure and never simple. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_64_2","unstructured":"Ting Jiang Shaohan Huang Shengyue Luo Zihan Zhang Haizhen Huang Furu Wei Weiwei Deng Feng Sun Qi Zhang Deqing Wang and Fuzhen Zhuang. 2024. Improving domain adaptation through extended-text reading comprehension. https:\/\/arxiv.org\/abs\/2401.07284"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","unstructured":"Di Jin Eileen Pan Nassim Oufattole Wei-Hung Weng Hanyi Fang and Peter Szolovits. 2021. What disease does this patient have? A large-scale open domain question answering dataset from medical exams. Applied Sciences 11 14 (2021) 6421. 10.3390\/app11146421","DOI":"10.3390\/app11146421"},{"key":"e_1_3_1_66_2","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, Hong Kong, China, November 3\u20137, 2019","year":"2019","unstructured":"Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu. 2019. PubMedQA: A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, Hong Kong, China, November 3\u20137, 2019."},{"key":"e_1_3_1_67_2","doi-asserted-by":"crossref","unstructured":"Joel Niklaus Veton Matoshi Matthias St\u00fcrmer Ilias Chalkidis and Daniel E. Ho. 2024. MultiLegalPile: A 689GB multilingual legal corpus. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) L.-W. Ku A. Martins and V. Srikumar (Eds.). 15077\u201315094.","DOI":"10.18653\/v1\/2024.acl-long.805"},{"key":"e_1_3_1_68_2","volume-title":"Proceedings of the International Conference on Machine Learning, 23\u201329 July 2023, Honolulu, Hawaii, USA","year":"2023","unstructured":"Ayush Kaushal, Tejas Vaidhya, and Irina Rish. 2023. LORD: Low rank decomposition of monolingual code LLMs for one-shot compression. In Proceedings of the International Conference on Machine Learning, 23\u201329 July 2023, Honolulu, Hawaii, USA."},{"key":"e_1_3_1_69_2","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022","year":"2022","unstructured":"Minsoo Kim, Sihwa Lee, Suk Joon Hong, Duhyeuk Chang, and Jungwook Choi. 2022. Understanding and improving knowledge distillation for quantization aware training of large transformer encoders. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022."},{"key":"e_1_3_1_70_2","volume-title":"Proceedings of the 42st International Conference on Machine Learning, Vienna, Austria, July 21\u201327, 2024","year":"2024","unstructured":"Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W. Mahoney, and Kurt Keutzer. 2024. SqueezeLLM: Dense-and-sparse quantization. In Proceedings of the 42st International Conference on Machine Learning, Vienna, Austria, July 21\u201327, 2024."},{"key":"e_1_3_1_71_2","unstructured":"D. P. Kingma and J. Ba. 2014. Adam: A Method for stochastic optimization. Proceedings of the 3rd International Conference on Learning Representations (ICLR). https:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_1_72_2","volume-title":"Joint Technical Report TR\/SE0401 and 0400011T.1. Computer Science Department, Keele University and National ICT Australia Ltd.","author":"Kitchenham Barbara Ann","year":"2004","unstructured":"Barbara Ann Kitchenham. 2004. Procedures for performing systematic reviews. In Joint Technical Report TR\/SE0401 and 0400011T.1. Computer Science Department, Keele University and National ICT Australia Ltd."},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2008.09.009"},{"key":"e_1_3_1_74_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","year":"2024","unstructured":"Dawid Jan Kopiczko, Tijmen Blankevoort, and Yuki Markus Asano. 2024. VeRA: Vector-based random matrix adaptation. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_75_2","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022","year":"2022","unstructured":"Eldar Kurtic, Daniel Fernando Campos, Tuan Nguyen, Elias Frantar, Mark Kurtz, Ben Fineran, Michael Goin, and Dan Alistarh. 2022. The optimal BERT surgeon: Scalable and accurate second-order pruning for large language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022."},{"key":"e_1_3_1_76_2","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Virtual Event \/ Punta Cana, Dominican Republic, 7\u201311 November, 2021","year":"2021","unstructured":"Fran\u00e7ois Lagunas, Ella Charlaix, Victor Sanh, and Alexander M. Rush. 2021. Block pruning for faster transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Virtual Event \/ Punta Cana, Dominican Republic, 7\u201311 November, 2021."},{"key":"e_1_3_1_77_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 2, [NIPS Conference, Denver, Colorado, USA, November 27\u201330, 1989]","year":"1989","unstructured":"Yann LeCun, John S. Denker, and Sara A. Solla. 1989. Optimal brain damage. In Proceedings of the Advances in Neural Information Processing Systems 2, [NIPS Conference, Denver, Colorado, USA, November 27\u201330, 1989]."},{"key":"e_1_3_1_78_2","volume-title":"Proceedings of the 9th International Conference on Learning Representations, Virtual Event, Austria, May 3\u20137, 2021","year":"2021","unstructured":"Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam M. Shazeer, and Z. Chen. 2021. GShard: Scaling giant models with conditional computation and automatic sharding. In Proceedings of the 9th International Conference on Learning Representations, Virtual Event, Austria, May 3\u20137, 2021."},{"key":"e_1_3_1_79_2","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Virtual Event \/ Punta Cana, Dominican Republic, 7\u201311 November, 2021","year":"2021","unstructured":"Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Virtual Event \/ Punta Cana, Dominican Republic, 7\u201311 November, 2021."},{"key":"e_1_3_1_80_2","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Toronto, Canada, July 9\u201314, 2023","year":"2023","unstructured":"Jonathan Li, Will Aitken, Rohan Bhambhoria, and Xiao-Dan Zhu. 2023. Prefix propagation: Parameter-efficient tuning for long sequences. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Toronto, Canada, July 9\u201314, 2023."},{"key":"e_1_3_1_81_2","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, Canada, July 9\u201314, 2023","year":"2023","unstructured":"Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, and Yejin Choi. 2023. Symbolic chain-of-thought distillation: Small models can also \u201cThink\u201d step-by-step. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, Canada, July 9\u201314, 2023."},{"key":"e_1_3_1_82_2","unstructured":"Raymond Li Loubna Ben Allal Yangtian Zi Niklas Muennighoff Denis Kocetkov Chenghao Mou Marc Marone Christopher Akiki Jia Li Jenny Chim Qian Liu Evgenii Zheltonozhskii Terry Yue Zhuo Thomas Wang Olivier Dehaene Mishig Davaadorj Joel Lamy-Poirier Jo\u00e3o Monteiro Oleh Shliazhko Nicolas Gontier Nicholas Meade Armel Zebaze Ming-Ho Yee Logesh Kumar Umapathi Jian Zhu Benjamin Lipkin Muhtasham Oblokulov Zhiruo Wang Rudra Murthy V Jason T. Stillerman Siva Sankalp Patel Dmitry Abulkhanov Marco Zocca Manan Dey Zhihan Zhang Nour Fahmy Urvashi Bhattacharyya Wenhao Yu Swayam Singh Sasha Luccioni Paulo Villegas Maxim Kunakov Fedor Zhdanov Manuel Romero Tony Lee Nadav Timor Jennifer Ding Claire Schlesinger Hailey Schoelkopf Jan Ebert Tri Dao Mayank Mishra Alex Gu Jennifer Robinson Carolyn Jane Anderson Brendan Dolan-Gavitt Danish Contractor Siva Reddy Daniel Fried Dzmitry Bahdanau Yacine Jernite Carlos Mu\u00f1oz Ferrandis Sean Hughes Thomas Wolf Arjun Guha Leandro von Werra and Harm de Vries. 2023. StarCoder: May the source be with you! Transactions on Machine Learning Research. 1\u201355."},{"key":"e_1_3_1_83_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.353"},{"key":"e_1_3_1_84_2","unstructured":"Raymond Li Loubna Ben Allal Yangtian Zi Niklas Muennighoff Denis Kocetkov Chenghao Mou Marc Marone Christopher Akiki Jia Li Jenny Chim Qian Liu Evgenii Zheltonozhskii Terry Yue Zhuo Thomas Wang Olivier Dehaene Mishig Davaadorj Joel Lamy-Poirier Jo\u00e3o Monteiro Oleh Shliazhko Nicolas Gontier Nicholas Meade Armel Zebaze Ming-Ho Yee Logesh Kumar Umapathi Jian Zhu Benjamin Lipkin Muhtasham Oblokulov Zhiruo Wang Rudra Murthy V Jason T. Stillerman Siva Sankalp Patel Dmitry Abulkhanov Marco Zocca Manan Dey Zhihan Zhang Nour Fahmy Urvashi Bhattacharyya Wenhao Yu Swayam Singh Sasha Luccioni Paulo Villegas Maxim Kunakov Fedor Zhdanov Manuel Romero Tony Lee Nadav Timor Jennifer Ding Claire Schlesinger Hailey Schoelkopf Jan Ebert Tri Dao Mayank Mishra Alex Gu Jennifer Robinson Carolyn Jane Anderson Brendan Dolan-Gavitt Danish Contractor Siva Reddy Daniel Fried Dzmitry Bahdanau Yacine Jernite Carlos Mu\u00f1oz Ferrandis Sean Hughes Thomas Wolf Arjun Guha Leandro von Werra and Harm de Vries. 2023. StarCoder: May the source be with you! Transactions on Machine Learning Research. 1\u201355."},{"key":"e_1_3_1_85_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","year":"2024","unstructured":"Yixiao Li, Yifan Yu, Chen Liang, Pengcheng He, Nikos Karampatziakis, Weizhu Chen, and Tuo Zhao. 2024. LoftQ: LoRA-fine-tuning-aware quantization for large language models. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_86_2","volume-title":"Proceedings of the 7th Annual Conference on Machine Learning and Systems, Santa Clara, CA, USA, May 13\u201316, 2024","year":"2024","unstructured":"Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. 2024. AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration. In Proceedings of the 7th Annual Conference on Machine Learning and Systems, Santa Clara, CA, USA, May 13\u201316, 2024."},{"key":"e_1_3_1_87_2","unstructured":"DeepSeek-AI Aixin Liu Bei Feng Bin Wang Bingxuan Wang Bo Liu Chenggang Zhao Chengqi Deng Chong Ruan Damai Dai Daya Guo Dejian Yang Deli Chen Dongjie Ji Erhang Li Fangyun Lin Fuli Luo Guangbo Hao Guanting Chen Guowei Li Hao Zhang Hanwei Xu Hao Yang Haowei Zhang Honghui Ding Huajian Xin Huazuo Gao Hui Li Hui Qu J. L. Cai Jian Liang Jianzhong Guo Jiaqi Ni Jiashi Li Jin Chen Jingyang Yuan Junjie Qiu Junxiao Song Kai Dong Kaige Gao Kang Guan Lean Wang Lecong Zhang Lei Xu Leyi Xia Liang Zhao Liyue Zhang Meng Li Miaojun Wang Mingchuan Zhang Minghua Zhang Minghui Tang Mingming Li Ning Tian Panpan Huang Peiyi Wang Peng Zhang Qihao Zhu Qinyu Chen Qiushi Du R. J. Chen R. L. Jin Ruiqi Ge Ruizhe Pan Runxin Xu Ruyi Chen S. S. Li Shanghao Lu Shangyan Zhou Shanhuang Chen Shaoqing Wu Shengfeng Ye Shirong Ma Shiyu Wang Shuang Zhou Shuiping Yu Shunfeng Zhou Size Zheng Tao Wang Tian Pei Tian Yuan Tianyu Sun W. L. Xiao Wangding Zeng Wei An Wen Liu Wenfeng Liang Wenjun Gao Wentao Zhang X. Q. Li Xiangyue Jin Xianzu Wang Xiao Bi Xiaodong Liu Xiaohan Wang Xiaojin Shen Xiaokang Chen Xiaosha Chen Xiaotao Nie and Xiaowen Sun. 2024. DeepSeek-V2: A strong economical and efficient mixture-of-experts language model. https:\/\/arxiv.org\/abs\/2405.04434"},{"key":"e_1_3_1_88_2","unstructured":"DeepSeek-AI Aixin Liu Bei Feng Bing Xue Bingxuan Wang Bochao Wu Chengda Lu Chenggang Zhao Chengqi Deng Chenyu Zhang Chong Ruan Damai Dai Daya Guo Dejian Yang Deli Chen Dongjie Ji Erhang Li Fangyun Lin Fucong Dai Fuli Luo Guangbo Hao Guanting Chen Guowei Li H. Zhang Han Bao Hanwei Xu Haocheng Wang Haowei Zhang Honghui Ding Huajian Xin Huazuo Gao Hui Li Hui Qu J. L. Cai Jian Liang Jianzhong Guo Jiaqi Ni Jiashi Li Jiawei Wang Jin Chen Jingchang Chen Jingyang Yuan Junjie Qiu Junlong Li Junxiao Song Kai Dong Kai Hu Kaige Gao Kang Guan Kexin Huang Kuai Yu Lean Wang Lecong Zhang Lei Xu Leyi Xia Liang Zhao Litong Wang Liyue Zhang Meng Li Miaojun Wang Mingchuan Zhang Minghua Zhang Minghui Tang Mingming Li Ning Tian Panpan Huang Peiyi Wang Peng Zhang Qiancheng Wang Qihao Zhu Qinyu Chen Qiushi Du R. J. Chen R. L. Jin Ruiqi Ge Ruisong Zhang Ruizhe Pan Runji Wang Runxin Xu Ruoyu Zhang Ruyi Chen S. S. Li Shanghao Lu Shangyan Zhou Shanhuang Chen Shaoqing Wu Shengfeng Ye Shengfeng Ye Shirong Ma Shiyu Wang Shuang Zhou Shuiping Yu Shunfeng Zhou Shuting Pan T. Wang Tao Yun Tian Pei Tianyu Sun W. L. Xiao Wangding Zeng Wanjia Zhao Wei An Wen Liu Wenfeng Liang Wenjun Gao Wenqin Yu Wentao Zhang X. Q. Li Xiangyue Jin Xianzu Wang Xiao Bi Xiaodong Liu Xiaohan Wang Xiaojin Shen Xiaokang Chen Xiaokang Zhang Xiaosha Chen Xiaotao Nie Xiaowen Sun Xiaoxiang Wang Xin Cheng Xin Liu Xin Xie Xingchao Liu Xingkai Yu Xinnan Song Xinxia Shan Xinyi Zhou Xinyu Yang Xinyuan Li Xuecheng Su Xuheng Lin Y. K. Li Y. Q. Wang Y. X. Wei Y. X. Zhu Yang Zhang Yanhong Xu Yanhong Xu Yanping Huang Yao Li Yao Zhao Yaofeng Sun Yaohui Li Yaohui Wang Yi Yu Yi Zheng Yichao Zhang Yifan Shi Yiliang Xiong Ying He Ying Tang Yishi Piao Yisong Wang Yixuan Tan Yiyang Ma Yiyuan Liu Yongqiang Guo Yu Wu Yuan Ou Yuchen Zhu Yuduan Wang Yue Gong Yuheng Zou Yujia He Yukun Zha Yunfan Xiong Yunxian Ma Yuting Yan Yuxiang Luo Yuxiang You Yuxuan Liu Yuyang Zhou Z. F. Wu Z. Z. Ren Zehui Ren Zhangli Sha Zhe Fu Zhean Xu Zhen Huang Zhen Zhang Zhenda Xie Zhengyan Zhang Zhewen Hao Zhibin Gou Zhicheng Ma Zhigang Yan Zhihong Shao Zhipeng Xu Zhiyu Wu Zhongyu Zhang Zhuoshu Li Zihui Gu Zijia Zhu Zijun Liu Zilin Li Ziwei Xie Ziyang Song Ziyi Gao and Zizheng Pan. 2024. DeepSeek-V3 Technical Report. https:\/\/arxiv.org\/abs\/2412.19437"},{"key":"e_1_3_1_89_2","volume-title":"Proceedings of the 38th International Conference on Machine Learning, 18\u201324 July 2021, Virtual Event","year":"2021","unstructured":"Shiwei Liu, Decebal Constantin Mocanu, Yulong Pei, and Mykola Pechenizkiy. 2021. Selfish sparse RNN training. In Proceedings of the 38th International Conference on Machine Learning, 18\u201324 July 2021, Virtual Event."},{"key":"e_1_3_1_90_2","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, December 6\u201310, 2023","year":"2023","unstructured":"Shih-Yang Liu, Zechun Liu, Xijie Huang, Pingcheng Dong, and Kwang-Ting Cheng. 2023. LLM-FP4: 4-Bit floating-point quantized transformers. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, December 6\u201310, 2023."},{"key":"e_1_3_1_91_2","volume-title":"Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, July 21\u201327, 2024","year":"2024","unstructured":"Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, and Min-Hung Chen. 2024. DoRA: Weight-decomposed low-rank adaptation. In Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, July 21\u201327, 2024."},{"key":"e_1_3_1_92_2","first-page":"61","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","year":"2022","unstructured":"Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2022. P-Tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 61\u201368."},{"key":"e_1_3_1_93_2","doi-asserted-by":"publisher","unstructured":"Xiao Liu Yanan Zheng Zhengxiao Du Ming Ding Yujie Qian Zhilin Yang and Jie Tang. 2024. GPT understands too. AI Open 5 (2024) 208\u2013215. 10.1016\/j.aiopen.2023.08.012","DOI":"10.1016\/j.aiopen.2023.08.012"},{"key":"e_1_3_1_94_2","volume-title":"Proceedings of the 8th International Conference on Learning Representations, Virtual Event, April 26\u2013May 1, 2020","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. In Proceedings of the 8th International Conference on Learning Representations, Virtual Event, April 26\u2013May 1, 2020."},{"key":"e_1_3_1_95_2","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","year":"2024","unstructured":"Yongkang Liu, Yiqun Zhang, Qian Li, Shi Feng, Daling Wang, Yifei Zhang, and Hinrich Schutze. 2024. HiFT: A hierarchical full parameter fine-tuning strategy. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing."},{"key":"e_1_3_1_96_2","volume-title":"Findings of the Association for Computational Linguistics, Bangkok, Thailand and Virtual Meeting, August 11\u201316, 2024","year":"2024","unstructured":"Zechun Liu, Barlas O?uz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra. 2024. LLM-QAT: Data-free quantization aware training for large language models. In Findings of the Association for Computational Linguistics, Bangkok, Thailand and Virtual Meeting, August 11\u201316, 2024."},{"key":"e_1_3_1_97_2","volume-title":"Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, July 21\u201327, 2024","year":"2024","unstructured":"Zechun Liu, Changsheng Zhao, Forrest N. Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, Liangzhen Lai, and Vikas Chandra. 2024. MobileLLM: Optimizing sub-billion parameter language models for on-device use cases. In Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, July 21\u201327, 2024."},{"key":"e_1_3_1_98_2","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Online, June 6\u201311, 2021","year":"2021","unstructured":"Ximing Lu, Peter West, Rowan Zellers, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. NeuroLogic decoding: (Un)supervised neural text generation with predicate logic constraints. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Online, June 6\u201311, 2021."},{"key":"e_1_3_1_99_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, New Orleans, LA, USA, December 10\u201316, 2023","year":"2023","unstructured":"Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023. LLM-Pruner: On the structural pruning of large language models. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, New Orleans, LA, USA, December 10\u201316, 2023."},{"key":"e_1_3_1_100_2","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Toronto, Canada, July 9\u201314, 2023","year":"2023","unstructured":"Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, and Aliaksei Severyn. 2023. Teaching small language models to reason. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Toronto, Canada, July 9\u201314, 2023."},{"key":"e_1_3_1_101_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, December 6\u201314, 2021, Virtual","year":"2021","unstructured":"Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. 2021. Compacter: Efficient low-rank hypercomplex adapter layers. In Proceedings of the Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, December 6\u201314, 2021, Virtual."},{"key":"e_1_3_1_102_2","volume-title":"Companion of the The Web Conference 2018 on The Web Conference 2018, Lyon , France, April 23\u201327, 2018","year":"2018","unstructured":"Macedo Maia, Siegfried Handschuh, Andre Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018. WWW\u201918 open challenge: Financial opinion mining and question answering. In Companion of the The Web Conference 2018 on The Web Conference 2018, Lyon , France, April 23\u201327, 2018."},{"key":"e_1_3_1_103_2","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, August 11\u201316, 2024","year":"2024","unstructured":"Pratyush Maini, Skyler Seto, Richard He Bai, David Grangier, Yizhe Zhang, and Navdeep Jaitly. 2024. Rephrasing the web: A recipe for compute and data-efficient language modeling. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, August 11\u201316, 2024."},{"key":"e_1_3_1_104_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, New Orleans, LA, USA, December 10\u201316, 2023","year":"2023","unstructured":"Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alexandru Damian, Jason D. Lee, Danqi Chen, and Sanjeev Arora. 2023. Fine-tuning language models with just forward passes. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, New Orleans, LA, USA, December 10\u201316, 2023."},{"key":"e_1_3_1_105_2","doi-asserted-by":"crossref","unstructured":"Michael McCloskey and Neal J. Cohen. 1989. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of Learning and Motivation 24 (1989) 109\u2013165.","DOI":"10.1016\/S0079-7421(08)60536-8"},{"key":"e_1_3_1_106_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, December 8\u201314, 2019, Vancouver, BC, Canada","year":"2019","unstructured":"Paul Michel, Omer Levy, and Graham Neubig. 2019. Are sixteen heads really better than one?. In Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, December 8\u201314, 2019, Vancouver, BC, Canada."},{"key":"e_1_3_1_107_2","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","year":"2021","unstructured":"Linyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, and Nazneen Fatema Rajani. 2021. DART: Open-domain structured data record to text generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies."},{"key":"e_1_3_1_108_2","unstructured":"Sharan Narang Gregory Frederick Diamos Shubho Sengupta and Erich Elsen. 2017. Exploring sparsity in recurrent neural networks. In 5th International Conference on Learning Representations ICLR 2017 Toulon France April 24-26. 110. https:\/\/arxiv.org\/abs\/1704.05119"},{"key":"e_1_3_1_109_2","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","year":"2018","unstructured":"Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don\u2019t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing."},{"key":"e_1_3_1_110_2","volume-title":"Proceedings of the 8th Workshop on Representation Learning for NLP, Toronto, Canada, July 13, 2023","year":"2023","unstructured":"Stephen Obadinma, Hongyu Guo, and Xiao-Dan Zhu. 2023. Effectiveness of data augmentation for parameter efficient tuning with limited data. In Proceedings of the 8th Workshop on Representation Learning for NLP, Toronto, Canada, July 13, 2023."},{"key":"e_1_3_1_111_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, New Orleans, LA, USA, December 10\u201316, 2023","year":"2023","unstructured":"Matteo Pagliardini, Daniele Paliotta, Martin Jaggi, and Fran\u00e7ois Fleuret. 2023. Fast attention over long sequences with dynamic sparse flash attention. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, New Orleans, LA, USA, December 10\u201316, 2023."},{"key":"e_1_3_1_112_2","volume-title":"Proceedings of the Conference on Health, Inference, and Learning, 7\u20138 April 2022, Virtual Event","year":"2022","unstructured":"Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. MedMCQA: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Proceedings of the Conference on Health, Inference, and Learning, 7\u20138 April 2022, Virtual Event."},{"key":"e_1_3_1_113_2","first-page":"688","volume-title":"Proceedings of the 45th ACM\/IEEE Annual International Symposium on Computer Architecture, Los Angeles, CA, USA, June 1\u20136, 2018","year":"2018","unstructured":"Eunhyeok Park, Dongyoung Kim, and Sungjoo Yoo. 2018. Energy-efficient neural network accelerator based on outlier-aware low-precision computation. In Proceedings of the 45th ACM\/IEEE Annual International Symposium on Computer Architecture, Los Angeles, CA, USA, June 1\u20136, 2018. 688\u2013698."},{"key":"e_1_3_1_114_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","year":"2024","unstructured":"Gunho Park, Baeseong Park, Minsub Kim, Sungjae Lee, Jeonghoon Kim, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, and Dongsoo Lee. 2024. LUT-GEMM: Quantized matrix multiplication based on LUTs for efficient inference in large-scale generative language models. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_115_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","year":"2024","unstructured":"Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. 2024. YaRN: Efficient context window extension of large language models. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_116_2","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, EMNLP 2020 - Demos, Online, November 16\u201320, 2020","year":"2020","unstructured":"Jonas Pfeiffer, Andreas Ruckle, Clifton A. Poth, Aishwarya Kamath, Ivan Vulic, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020. AdapterHub: A framework for adapting transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, EMNLP 2020 - Demos, Online, November 16\u201320, 2020."},{"key":"e_1_3_1_117_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.410"},{"key":"e_1_3_1_118_2","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1 8 (2019) 9."},{"key":"e_1_3_1_119_2","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Austin, Texas, USA, November 1\u20134, 2016","year":"2016","unstructured":"Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100, 000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Austin, Texas, USA, November 1\u20134, 2016."},{"key":"e_1_3_1_120_2","volume-title":"Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands, October 11\u201314, 2016","year":"2016","unstructured":"Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. 2016. XNOR-Net: ImageNet classification using binary convolutional neural networks. In Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands, October 11\u201314, 2016."},{"key":"e_1_3_1_121_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, New Orleans, LA, USA, December 10\u201316, 2023","year":"2023","unstructured":"Rajarshi Saha, Varun Srivastava, and Mert Pilanci. 2023. Matrix compression via randomized low rank and low precision factorization. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, New Orleans, LA, USA, December 10\u201316, 2023."},{"key":"e_1_3_1_122_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems, Vancouver, BC, Canada, December 10\u201315, 2024","author":"Sakr Charbel","year":"2024","unstructured":"Charbel Sakr and Brucek Khailany. 2024. ESPACE: Dimensionality reduction of activations for model compression. In Proceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems, Vancouver, BC, Canada, December 10\u201315, 2024."},{"key":"e_1_3_1_123_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, December 6\u201312, 2020, virtual","year":"2020","unstructured":"Victor Sanh, Thomas Wolf, and Alexander M. Rush. 2020. Movement pruning: Adaptive sparsity by fine-tuning. In Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, December 6\u201312, 2020, virtual."},{"key":"e_1_3_1_124_2","unstructured":"Teven Le Scao Angela Fan Christopher Akiki Ellie Pavlick Suzana Ili\u2019c Daniel Hesslow Roman Castagn\u2019e Alexandra Sasha Luccioni Fran\u00e7ois Yvon Matthias Galle Jonathan Tow Alexander M. Rush Stella Biderman Albert Webson and Pawan Sasanka Ammanamanchi. 2022. BLOOM: A 176B-parameter open-access multilingual language model. arXiv:2211.05100. Retrieved from https:\/\/arxiv.org\/abs\/2211.05100"},{"key":"e_1_3_1_125_2","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal policy optimization algorithms. https:\/\/arxiv.org\/abs\/1707.06347"},{"key":"e_1_3_1_126_2","unstructured":"Zihan Shao Peng Wang Qiang Zhu Rui Xu Jiawei Song Ming Zhang Yikang Li Yiming Wu and Daya Guo. 2024. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. https:\/\/arxiv.org\/abs\/2402.03300"},{"key":"e_1_3_1_127_2","volume-title":"Proceedings of the 5th MLSys Conference, Santa Clara, CA, USA, 2024","year":"2024","unstructured":"Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, Joseph Gonzalez, and Ion Stoica. 2024. S-LoRA: Serving thousands of concurrent LoRA adapters. In Proceedings of the 5th MLSys Conference, Santa Clara, CA, USA, 2024."},{"key":"e_1_3_1_128_2","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Online, November 16\u201320, 2020","year":"2020","unstructured":"Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting knowledge from language models with automatically generated prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Online, November 16\u201320, 2020."},{"key":"e_1_3_1_129_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, December 6\u201312, 2020, virtual","author":"Singh Sidak Pal","year":"2020","unstructured":"Sidak Pal Singh and Dan Alistarh. 2020. WoodFisher: Efficient second-order approximations for model compression. In Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, December 6\u201312, 2020, virtual."},{"key":"e_1_3_1_130_2","doi-asserted-by":"crossref","unstructured":"Srinjoy Sridharan John R. Stevens Karthik Roy and Anand Raghunathan. 2023. X-Former: In-memory acceleration of transformers. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 31 (2023) 1223\u20131233.","DOI":"10.1109\/TVLSI.2023.3282046"},{"key":"e_1_3_1_131_2","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020","year":"2020","unstructured":"Haipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama, Eiichiro Sumita, and Tiejun Zhao. 2020. Knowledge distillation for multilingual unsupervised neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020."},{"key":"e_1_3_1_132_2","volume-title":"Proceedings of the 1st Conference on Language Modeling, Philadelphia, PA, USA, October 7\u20139, 2024","year":"2024","unstructured":"Hanshi Sun, Zhuoming Chen, Xinyu Yang, Yuandong Tian, and Beidi Chen. 2024. TriForce: Lossless acceleration of long sequence generation with hierarchical speculative decoding. In Proceedings of the 1st Conference on Language Modeling, Philadelphia, PA, USA, October 7\u20139, 2024."},{"key":"e_1_3_1_133_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","year":"2024","unstructured":"Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter. 2024. A simple and effective pruning approach for large language models. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_134_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, December 6\u201314, 2021, Virtual","year":"2021","unstructured":"Yi-Lin Sung, Varun Nair, and Colin Raffel. 2021. Training neural networks with fixed sparse masks. In Proceedings of the Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, December 6\u201314, 2021, Virtual."},{"key":"e_1_3_1_135_2","unstructured":"Rohan Taori Ishaan Gulrajani Tianyi Zhang Yann Dubois Xuechen Li Carlos Guestrin Percy Liang and Tatsunori Hashimoto. 2023. Stanford alpaca: An instruction-following LLaMA model. GitHub. https:\/\/github.com\/tatsu-lab\/stanford_alpaca. (Accessed on October 16 2024)."},{"key":"e_1_3_1_136_2","unstructured":"Hugo Touvron Louis Martin Kevin R. Stone Peter Albert Amjad Almahairi Yasmine Babaei Niko-lay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale Daniel M. Bikel Lukas Blecher Cristian Canton Ferrer Moya Chen Guillem Cucurull David Esiobu Jude Fernandes Jeremy Fu Wenyin Fu Brian Fuller Cynthia Gao Vedanuj Goswami Naman Goyal Anthony S. Hartshorn Saghar Hosseini Rui Hou Hakan Inan Marcin Kardas Viktor Kerkez Madian Khabsa Isabel M. Kloumann A. V. Korenev Punit Singh Koura Marie-Anne Lachaux Thibaut Lavril Jenya Lee Diana Liskovich Yinghai Lu Yuning Mao Xavier Martinet Todor Mihaylov Pushkar Mishra Igor Molybog Yixin Nie Andrew Poulton Jeremy Reizenstein Rashi Rungta Kalyan Saladi Alan Schelten Ruan Silva Eric Michael Smith R. Subramanian Xia Tan Binh Tang Ross Taylor Adina Williams Jian Xiang Kuan Puxin Xu Zhengxu Yan Iliyan Zarov Yuchen Zhang Angela Fan Melissa Hall Melanie Kambadur Sharan Narang Aur\u2019elien Rodriguez Robert Stojnic Sergey Edunov and Thomas Scialom. 2023. Llama 2: Open foundation and fine-Tuned chat models. https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_1_137_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timothee Lacroix Baptiste Roziere Naman Goyal Eric Hambro Faisal Azhar Aur\u2019elien Rodriguez Armand Joulin Edouard Grave and Guillaume Lample. 2023. LLaMA: open and efficient foundation language models. https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_1_138_2","unstructured":"Zhongwei Wan Xin Wang Che Liu Samiul Alam Yu Zheng Jiachen Liu Zhongnan Qu Shen Yan Yi Zhu Quanlu Zhang Mosharaf Chowdhury and Mi Zhang. 2024. Efficient large language models: A survey. Transactions on Machine Learning Research (2024). 1\u201367."},{"key":"e_1_3_1_139_2","volume-title":"Proceedings of the Workshop: Analyzing and Interpreting Neural Networks for NLP, Brussels, Belgium, 2018","year":"2018","unstructured":"Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the Workshop: Analyzing and Interpreting Neural Networks for NLP, Brussels, Belgium, 2018."},{"key":"e_1_3_1_140_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 32","year":"2019","unstructured":"Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. SuperGLUE: A stickier benchmark for general-purpose language understanding systems. In Proceedings of the Advances in Neural Information Processing Systems 32."},{"key":"e_1_3_1_141_2","unstructured":"Hongyu Wang Shuming Ma Li Dong Shaohan Huang Huaijie Wang Lingxiao Ma Fan Yang Ruiping Wang Yi Wu and Furu Wei. 2023. BitNet: Scaling 1-bit transformers for large language models. https:\/\/arxiv.org\/abs\/2310.11453"},{"key":"e_1_3_1_142_2","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022","year":"2022","unstructured":"Yaqing Wang, Subhabrata Mukherjee, Xiaodong Liu, Jing Gao, and Jianfeng Gao. 2022. AdaMix: Mixture-of-adaptations for parameter-efficient model tuning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022."},{"key":"e_1_3_1_143_2","unstructured":"Yiming Wang Yu Lin Xiaodong Zeng and Guannan Zhang. 2023. MultiLoRA: Democratizing LoRA for better multi-task learning. https:\/\/arxiv.org\/abs\/2311.11501"},{"key":"e_1_3_1_144_2","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Online, November 16\u201320, 2020","year":"2020","unstructured":"Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2020. Structured pruning of large language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Online, November 16\u201320, 2020."},{"key":"e_1_3_1_145_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations, Kigali, Rwanda, May 1\u20135, 2023","year":"2023","unstructured":"Ziqi Wang, Yuexin Wu, Frederick Liu, Daogao Liu, Le Hou, Hongkun Yu, Jing Li, and Heng Ji. 2023. Augmentation with projection: Towards an effective and efficient data augmentation paradigm for distillation. In Proceedings of the 11th International Conference on Learning Representations, Kigali, Rwanda, May 1\u20135, 2023."},{"key":"e_1_3_1_146_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations, Kigali, Rwanda, May 1\u20135, 2023","year":"2023","unstructured":"Zhen Wang, Rameswar Panda, Leonid Karlinsky, Rogerio Schmidt Feris, Huan Sun, and Yoon Kim. 2023. Multitask prompt tuning enables parameter-efficient transfer learning. In Proceedings of the 11th International Conference on Learning Representations, Kigali, Rwanda, May 1\u20135, 2023."},{"key":"e_1_3_1_147_2","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Seattle, WA, United States, July 10\u201315, 2022","year":"2022","unstructured":"Peter West, Chandrasekhar Bhagavatula, Jack Hessel, Jena D. Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi. 2022. Symbolic knowledge distillation: From general language models to commonsense models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Seattle, WA, United States, July 10\u201315, 2022."},{"key":"e_1_3_1_148_2","volume-title":"Proceedings of the 33rd International Joint Conference on Artificial Intelligence, Jeju, South Korea, August 3\u20139, 2024","year":"2024","unstructured":"Herbert Woisetschlager, Alexander Isenko, Shiqiang Wang, Ruben Mayer, and Hans-Arno Jacobsen. 2024. A survey on efficient federated learning methods for foundation model training. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence, Jeju, South Korea, August 3\u20139, 2024."},{"key":"e_1_3_1_149_2","unstructured":"Chuhan Wu Fangzhao Wu Lingjuan Lyu Yongfeng Huang and Xing Xie. 2022. Communication-efficient federated learning via knowledge distillation. Nature Communications (2022)."},{"key":"e_1_3_1_150_2","doi-asserted-by":"crossref","unstructured":"Chaoyi Wu Weixiong Lin Xiaoman Zhang Ya Zhang Weidi Xie and Yanfeng Wang. 2024. PMC-LLaMA: Toward building open-source language models for medicine. Journal of the American Medical Informatics Association 31 9 (2024) 1833\u20131843.","DOI":"10.1093\/jamia\/ocae045"},{"key":"e_1_3_1_151_2","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, May 22\u201327, 2022","year":"2022","unstructured":"Mengzhou Xia, Zexuan Zhong, and Danqi Chen. 2022. Structured pruning learns compact and accurate models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, May 22\u201327, 2022."},{"key":"e_1_3_1_152_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","year":"2024","unstructured":"Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen. 2024. Sheared LLaMA: Accelerating language model pre-training via structured pruning. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_153_2","volume-title":"Proceedings of the International Conference on Machine Learning, 23\u201329 July 2023, Honolulu, Hawaii, USA","year":"2023","unstructured":"Guangxuan Xiao, Ji Lin, Mickael Seznec, Julien Demouth, and Song Han. 2023. SmoothQuant: Accurate and efficient post-training quantization for large language models. In Proceedings of the International Conference on Machine Learning, 23\u201329 July 2023, Honolulu, Hawaii, USA."},{"key":"e_1_3_1_154_2","volume-title":"Findings of the Association for Computational Linguistics, Singapore, December 6\u201310, 2023","year":"2023","unstructured":"Yao Xiao, Lu Xu, Jiaxi Li, Wei Lu, and Xiaoli Li. 2023. Decomposed prompt tuning via low-rank reparameterization. In Findings of the Association for Computational Linguistics, Singapore, December 6\u201310, 2023."},{"key":"e_1_3_1_155_2","volume-title":"Proceedings of the 2024 USENIX Annual Technical Conference, Santa Clara, CA, USA, July 10\u201312, 2024","year":"2024","unstructured":"Mengwei Xu, Dongqi Cai, Yaozong Wu, Xiang Li, and Shangguang Wang. 2024. FwdLLM: Efficient federated finetuning of large language models with perturbed inferences. In Proceedings of the 2024 USENIX Annual Technical Conference, Santa Clara, CA, USA, July 10\u201312, 2024."},{"key":"e_1_3_1_156_2","unstructured":"Mengwei Xu Wangsong Yin Dongqi Cai Rongjie Yi Daliang Xu Qipeng Wang Bingyang Wu Yihao Zhao Chen Yang Shihe Wang Qiyang Zhang Zhenyan Lu Li Zhang Shangguang Wang Yuanchun Li Yunxin Liu Xin Jin and Xuanzhe Liu. 2024. A survey of resource-efficient LLM and multimodal foundation models. https:\/\/arxiv.org\/abs\/2401.08092"},{"key":"e_1_3_1_157_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024","year":"2024","unstructured":"Yuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen, Heng Chang, Hengheng Zhang, Zhensu Chen, Xiaopeng Zhang, and Qi Tian. 2024. QA-LoRA: Quantization-aware low-rank adaptation of large language models. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, May 7\u201311, 2024."},{"key":"e_1_3_1_158_2","unstructured":"An Yang Baosong Yang Beichen Zhang Binyuan Hui Bo Zheng Bowen Yu Chengyuan Li Dayiheng Liu Fei Huang Guanting Dong Haoran Wei Huan Lin Jian Yang Jianhong Tu Jianwei Zhang Jianxin Yang Jiaxin Yang Jingren Zhou Junyang Lin Kai Dang Keming Lu Keqin Bao Kexin Yang Le Yu Mei Li Mingfeng Xue Pei Zhang Qin Zhu Rui Men Runji Lin Tianhao Li Tingyu Xia Xingzhang Ren Xuancheng Ren Yang Fan Yang Su Yi-Chao Zhang Yunyang Wan Yuqi Liu Zeyu Cui Zhenru Zhang Zihan Qiu Shanghaoran Quan and Zekun Wang. 2025. Qwen2.5 Technical Report. ArXiv https:\/\/arxiv.org\/abs\/2412.15115"},{"key":"e_1_3_1_159_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, New Orleans, LA, USA, November 28\u2013December 9, 2022","year":"2022","unstructured":"Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. 2022. ZeroQuant: Efficient and affordable post-training quantization for large-scale transformers. In Proceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, New Orleans, LA, USA, November 28\u2013December 9, 2022."},{"key":"e_1_3_1_160_2","volume-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence","year":"2024","unstructured":"Zhewei Yao, Xiaoxia Wu, Cheng Li, Stephen Youn, and Yuxiong He. 2024. Exploring post-training quantization in LLMs from comprehensive study to low rank compensation. In Proceedings of the 38th AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_1_161_2","volume-title":"Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, July 21\u201327, 2024","year":"2024","unstructured":"Lu Yin, You Wu, Zhenyu (Allen) Zhang, Cheng-Yu Hsieh, Yaqing Wang, Yiling Jia, Mykola Pechenizkiy, Yi Liang, Zhangyang Wang, and Shiwei Liu. 2024. Outlier weighed layerwise sparsity (OWL): A missing secret sauce for pruning LLMs to high sparsity. In Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, July 21\u201327, 2024."},{"key":"e_1_3_1_162_2","unstructured":"Zhihang Yuan Lin Niu Jia-Wen Liu Wenyu Liu Xinggang Wang Yuzhang Shang Guangyu Sun Qiang Wu Jiaxiang Wu and Bingzhe Wu. 2023. RPTQ: Reorder-based Post-training quantization for large language models. https:\/\/arxiv.org\/abs\/2304.01089"},{"key":"e_1_3_1_163_2","first-page":"811","volume-title":"Proceedings of the 53rd Annual IEEE\/ACM International Symposium on Microarchitecture, Athens, Greece, October 17\u201321, 2020","year":"2020","unstructured":"Ali Hadi Zadeh and Andreas Moshovos. 2020. GOBO: Quantizing attention-based NLP models for low latency and energy efficient inference. In Proceedings of the 53rd Annual IEEE\/ACM International Symposium on Microarchitecture, Athens, Greece, October 17\u201321, 2020. 811\u2013824."},{"key":"e_1_3_1_164_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, December 6\u201314, 2021, Workshop","year":"2021","unstructured":"Ofir Zafrir, Ariel Larey, Guy Boudoukh, Haihao Shen, and Moshe Wasserblat. 2021. Prune once for all: Sparse pre-trained language models. In Proceedings of the Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, December 6\u201314, 2021, Workshop."},{"key":"e_1_3_1_165_2","volume-title":"Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, 20\u201325 May, 2024, Torino, Italy","year":"2024","unstructured":"Hongchuan Zeng, Hongshen Xu, Lu Chen, and Kai Yu. 2024. Multilingual brain surgeon: Large language models can be compressed leaving no language behind. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, 20\u201325 May, 2024, Torino, Italy."},{"key":"e_1_3_1_166_2","unstructured":"Hongyi Zhang Moustapha Cisse Yann Dauphin and David Lopez-Paz. 2018. mixup: Beyond empirical Risk Minimization. In 6th International Conference on Learning Representations ICLR 2018 Vancouver BC Canada April 30 - May 3 2018. https:\/\/arxiv.org\/abs\/1710.09412"},{"key":"e_1_3_1_167_2","volume-title":"Findings of the Association for Computational Linguistics, Bangkok, Thailand and Virtual Meeting, August 11\u201316, 2024","year":"2024","unstructured":"Mingyang Zhang, Hao Chen, Chunhua Shen, Zhenyi Yang, Linlin Ou, Xinyi Yu, and Bohan Zhuang. 2024. LoRAPrune: Structured pruning meets low-rank parameter-efficient fine-tuning. In Findings of the Association for Computational Linguistics, Bangkok, Thailand and Virtual Meeting, August 11\u201316, 2024."},{"key":"e_1_3_1_168_2","volume-title":"Findings of the Association for Computational Linguistics, Mexico City, Mexico, June 16\u201321, 2024","year":"2024","unstructured":"Nan Zhang, Yanchi Liu, Xujiang Zhao, Wei Cheng, Runxue Bao, Rui Zhang, Prasenjit Mitra, and Haifeng Chen. 2024. Pruning as a domain-specific LLM extractor. In Findings of the Association for Computational Linguistics, Mexico City, Mexico, June 16\u201321, 2024."},{"key":"e_1_3_1_169_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations, Kigali, Rwanda, May 1\u20135, 2023","year":"2023","unstructured":"Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. 2023. Adaptive budget allocation for parameter-efficient fine-tuning. In Proceedings of the 11th International Conference on Learning Representations, Kigali, Rwanda, May 1\u20135, 2023."},{"key":"e_1_3_1_170_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 2024","year":"2024","unstructured":"Renrui Zhang, Jiaming Han, Chris Liu, Aojun Zhou, Pan Lu, Yu Jiao Qiao, Hongsheng Li, and Peng Gao. 2024. LLaMA-Adapter: Efficient fine-tuning of large language models with zero-initialized attention. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 2024."},{"key":"e_1_3_1_171_2","unstructured":"Susan Zhang Stephen Roller Naman Goyal Mikel Artetxe Moya Chen Shuohui Chen Christopher Dewan Mona T. Diab Xian Li Xi Victoria Lin Todor Mihaylov Myle Ott Sam Shleifer Kurt Shuster Daniel Simig Punit Singh Koura Anjali Sridhar Tianlu Wang and Luke Zettlemoyer. 2022. OPT: Open Pre-trained transformer language models. https:\/\/arxiv.org\/abs\/2205.01068"},{"key":"e_1_3_1_172_2","volume-title":"Proceedings of the 18th International Conference for Artificial Intelligence and Law","year":"2021","unstructured":"Lucia Zheng, Neel Guha, Brandon R. Anderson, Peter Henderson, and Daniel E. Ho. 2021. When does pretraining help?: Assessing self-supervised learning for law and the CaseHOLD dataset of 53, 000+ legal holdings. In Proceedings of the 18th International Conference for Artificial Intelligence and Law."},{"key":"e_1_3_1_173_2","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","year":"2024","unstructured":"Wei Zhu, Aaron Xuxiang Tian, Congrui Yin, Yuan Ni, Xiaoling Wang, and Guo Tong Xie. 2024. IAPT: Instance-aware prompt tuning for large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)."},{"key":"e_1_3_1_174_2","doi-asserted-by":"publisher","unstructured":"Xunyu Zhu Jian Li Yong Liu Can Ma and Weiping Wang. 2023. A survey on model compression for large language models. Transactions of the Association for Computational Linguistics 12 (2023) 1556\u20131577. 10.1162\/tacl_a_00704","DOI":"10.1162\/tacl_a_00704"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728636","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728636","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:36Z","timestamp":1750295916000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728636"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,6]]},"references-count":173,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2025,10,31]]}},"alternative-id":["10.1145\/3728636"],"URL":"https:\/\/doi.org\/10.1145\/3728636","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,6]]},"assertion":[{"value":"2024-03-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-06","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}