{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T14:47:51Z","timestamp":1782571671334,"version":"3.54.5"},"reference-count":107,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T00:00:00Z","timestamp":1782518400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/100031060","name":"European High Performance Computing Joint Undertaking","doi-asserted-by":"crossref","award":["955513 (MAELSTROM) and 101034126 (EU-Pilot)"],"award-info":[{"award-number":["955513 (MAELSTROM) and 101034126 (EU-Pilot)"]}],"id":[{"id":"10.13039\/100031060","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001711","name":"Swiss National Science Foundation","doi-asserted-by":"crossref","award":["185778"],"award-info":[{"award-number":["185778"]}],"id":[{"id":"10.13039\/501100001711","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    As deep learning models grow, sparsity is becoming an increasingly critical component of deep neural networks, enabling improved performance and reduced storage. However, existing frameworks offer poor support for sparsity. Specialized sparsity engines focus exclusively on sparse inference, while general frameworks primarily focus on sparse tensors in classical formats and neglect the broader sparsification pipeline necessary for using sparse models, especially during training. Further, existing frameworks are not easily extensible: adding a new sparse tensor format or operator is challenging and time-consuming. To address this, we propose STen, a sparsity programming model and interface for PyTorch whose key design insight is the decoupling of sparsity layouts, operators, and sparsifiers into composable, first-class abstractions that users can independently define and combine. An automatic dispatch mechanism selects the best available sparse implementation and transparently falls back to dense execution, allowing STen to support virtually all sparsification methods while enabling rapid prototyping without sacrificing performance for optimized paths. We demonstrate the versatility of STen by expressing existing sparsification techniques within its abstraction, achieving a code size reduction of over 2\u00d7. Finally, we develop a novel, high-performance grouped\n                    <jats:italic toggle=\"yes\">n<\/jats:italic>\n                    :\n                    <jats:italic toggle=\"yes\">m<\/jats:italic>\n                    sparsity layout for CPU inference at moderate sparsity, accelerating end-to-end BERT\n                    <jats:sub>\n                      <jats:sc>BASE<\/jats:sc>\n                    <\/jats:sub>\n                    inference by up to 3.2\u00d7. STen brings high performance and ease of use, making sparsity readily accessible for existing PyTorch models.\n                  <\/jats:p>","DOI":"10.1145\/3815424","type":"journal-article","created":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T11:09:05Z","timestamp":1779880145000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["STen: Productive and Efficient Sparsity in PyTorch"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-9487-9990","authenticated-orcid":false,"given":"Andrei","family":"Ivanov","sequence":"first","affiliation":[{"name":"ETH Zurich","place":["Z\u00fcrich, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9965-3647","authenticated-orcid":false,"given":"Nikoli","family":"Dryden","sequence":"additional","affiliation":[{"name":"ETH Zurich","place":["Z\u00fcrich, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3657-6568","authenticated-orcid":false,"given":"Tal","family":"Ben-Nun","sequence":"additional","affiliation":[{"name":"ETH Zurich","place":["Z\u00fcrich, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4884-3934","authenticated-orcid":false,"given":"Timo","family":"Schneider","sequence":"additional","affiliation":[{"name":"ETH Zurich","place":["Z\u00fcrich, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6115-6779","authenticated-orcid":false,"given":"Saleh","family":"Ashkboos","sequence":"additional","affiliation":[{"name":"ETH Zurich","place":["Z\u00fcrich, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1333-9797","authenticated-orcid":false,"given":"Torsten","family":"Hoefler","sequence":"additional","affiliation":[{"name":"ETH Zurich","place":["Z\u00fcrich, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,27]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin et\u00a0al. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. Retrieved June 8 2026 from https:\/\/www.tensorflow.org\/"},{"issue":"2166","key":"e_1_3_2_3_2","first-page":"1","article-title":"Preparing sparse solvers for exascale computing","volume":"378","author":"Anzt Hartwig","year":"2020","unstructured":"Hartwig Anzt, Erik Boman, Rob Falgout, Pieter Ghysels, Michael Heroux, Xiaoye Li, Lois Curfman McInnes, Richard Tran Mills, Sivasankaran Rajamanickam, Karl Rupp, et\u00a0al. 2020. Preparing sparse solvers for exascale computing. Philosophical Transactions of the Royal Society A 378, 2166 (2020), 1\u201317.","journal-title":"Philosophical Transactions of the Royal Society A"},{"key":"e_1_3_2_4_2","unstructured":"Maciej Besta and Torsten Hoefler. 2022. Parallel and distributed graph neural networks: An in-depth concurrency analysis. arXiv:2205.09702. Retrieved from https:\/\/arxiv.org\/abs\/2205.09702"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3544559"},{"key":"e_1_3_2_6_2","article-title":"Language models are few-shot learners","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et\u00a0al. 2020. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NeurIPS).","journal-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_3_2_7_2","unstructured":"Jesse Cai Daniel Haziza and Supriya Rao. 2024. Accelerating Neural Network Training with Semi-Structured (2:4) Sparsity. Retrieved June 8 2026 from https:\/\/pytorch.org\/blog\/accelerating-neural-network-training\/"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2025.3554028"},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et\u00a0al. 2018. TVM: An automated end-to-end optimizing compiler for deep learning. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI)."},{"key":"e_1_3_2_10_2","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems (NeurIPS)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. Learning to optimize tensor programs. In Proceedings of the International Conference on Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_2_11_2","unstructured":"Xuhao Chen. 2018. Escoin: Efficient sparse convolutional neural network inference on GPUs. arXiv:1802.10280. Retrieved from https:\/\/arxiv.org\/abs\/1802.10280"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476182"},{"key":"e_1_3_2_13_2","unstructured":"Rewon Child Scott Gray Alec Radford and Ilya Sutskever. 2019. Generating long sequences with sparse transformers. arXiv:1904.10509. Retrieved from https:\/\/arxiv.org\/abs\/1904.10509"},{"key":"e_1_3_2_14_2","unstructured":"Aakanksha Chowdhery Sharan Narang Jacob Devlin Maarten Bosma Gaurav Mishra Adam Roberts Paul Barham Hyung Won Chung Charles Sutton Sebastian Gehrmann et\u00a0al. 2022. PaLM: Scaling language modeling with pathways. arXiv:2204.02311. Retrieved from https:\/\/arxiv.org\/abs\/2204.02311"},{"key":"e_1_3_2_15_2","volume-title":"Proceedings of the International Conference on Machine Learning (ICML)","author":"Dao Tri","year":"2022","unstructured":"Tri Dao, Beidi Chen, Nimit S. Sohoni, Arjun Desai, Michael Poli, Jessica Grogan, Alexander Liu, Aniruddh Rao, Atri Rudra, and Christopher R\u00e9. 2022. Monarch: Expressive structured matrices for efficient and accurate training. In Proceedings of the International Conference on Machine Learning (ICML)."},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2021.3098483"},{"key":"e_1_3_2_17_2","unstructured":"Tim Dettmers Ruslan Svirschevski Vage Egiazarian Denis Kuznedelev Elias Frantar Saleh Ashkboos Alexander Borzunov Torsten Hoefler and Dan Alistarh. 2023. SpQR: A sparse-quantized representation for near-lossless LLM weight compression. arxiv:2306.03078 [cs.CL]. Retrieved from https:\/\/arxiv.org\/abs\/2306.03078"},{"key":"e_1_3_2_18_2","unstructured":"Tim Dettmers and Luke Zettlemoyer. 2019. Sparse networks from scratch: Faster training without losing performance. arXiv:1907.04840. Retrieved from https:\/\/arxiv.org\/abs\/1907.04840"},{"key":"e_1_3_2_19_2","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL)."},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/MLHPC.2016.004"},{"key":"e_1_3_2_21_2","first-page":"5547","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Du Nan","year":"2022","unstructured":"Nan Du, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et\u00a0al. 2022. Glam: Efficient scaling of language models with mixture-of-experts. In Proceedings of the International Conference on Machine Learning. PMLR, 5547\u20135569."},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01464"},{"key":"e_1_3_2_23_2","first-page":"2943","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Evci Utku","year":"2020","unstructured":"Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen. 2020. Rigging the lottery: Making all tickets winners. In Proceedings of the International Conference on Machine Learning. PMLR, 2943\u20132952."},{"key":"e_1_3_2_24_2","volume-title":"Proceedings of the Identifying and Understanding Deep Learning Phenomena Workshop@ICLR","author":"Evci Utku","year":"2019","unstructured":"Utku Evci, Fabian Pedregosa, Aidan Gomez, and Erich Elsen. 2019. The difficulty of training sparse neural networks. In Proceedings of the Identifying and Understanding Deep Learning Phenomena Workshop@ICLR."},{"key":"e_1_3_2_25_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Evci Utku","year":"2022","unstructured":"Utku Evci, Max Vladymyrov, Thomas Unterthiner, Bart van Merri\u00ebnboer, and Fabian Pedregosa. 2022. GradMax: Growing neural networks using gradient information. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"issue":"120","key":"e_1_3_2_26_2","first-page":"1","article-title":"Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity","volume":"23","author":"Fedus William","year":"2022","unstructured":"William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research 23, 120 (2022), 1\u201339.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_27_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Frankle Jonathan","year":"2019","unstructured":"Jonathan Frankle and Michael Carbin. 2019. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_28_2","first-page":"10323","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Frantar Elias","year":"2023","unstructured":"Elias Frantar and Dan Alistarh. 2023. SparseGPT: Massive language models can be accurately pruned in one-shot. In Proceedings of the International Conference on Machine Learning. PMLR, 10323\u201310337."},{"key":"e_1_3_2_29_2","unstructured":"Elias Frantar Saleh Ashkboos Torsten Hoefler and Dan Alistarh. 2022. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv:2210.17323. Retrieved from https:\/\/arxiv.org\/abs\/2210.17323"},{"key":"e_1_3_2_30_2","doi-asserted-by":"crossref","unstructured":"Elias Frantar Roberto L. Castro Jiale Chen Torsten Hoefler and Dan Alistarh. 2024. MARLIN: Mixed-precision auto-regressive parallel inference on large language models. arXiv:2408.11743. Retrieved from https:\/\/arxiv.org\/abs\/2408.11743","DOI":"10.1145\/3710848.3710871"},{"key":"e_1_3_2_31_2","volume-title":"Sparse-Marlin: Boosting 4-bit inference kernels with 2:4 Sparsity","author":"Frantar Elias","year":"2024","unstructured":"Elias Frantar, Roberto L. Castro, Jiale Chen, Torsten Hoefler, and Dan Alistarh. 2024. Sparse-Marlin: Boosting 4-bit inference kernels with 2:4 Sparsity. Retrieved June 8, 2026 from https:\/\/github.com\/IST-DASLab\/Sparse-Marlin"},{"key":"e_1_3_2_32_2","unstructured":"Joshua Fromm Bing Xu Morgan Funtowicz and Jason Knight. 2020. Leveraging Block Sparsity with Apache TVM to Halve your Cloud Bill for NLP. Retrieved January 31 2023 from https:\/\/octoml.ai\/blog\/leveraging-block-sparsity-with-apache-tvm-to-halve-your-cloud-bill-for-nlp\/"},{"key":"e_1_3_2_33_2","unstructured":"Trevor Gale Erich Elsen and Sara Hooker. 2019. The state of sparsity in deep neural networks. arXiv:1902.09574. Retrieved from https:\/\/arxiv.org\/abs\/1902.09574"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.5555\/3433701.3433723"},{"key":"e_1_3_2_35_2","unstructured":"Yizhao Gao Zhichen Zeng Dayou Du Shijie Cao Peiyuan Zhou Jiaxing Qi Junjie Lai Hayden Kwok-Hay So Ting Cao Fan Yang and Mao Yang. 2025. SeerAttention: Learning intrinsic sparse attention in your LLMs. arxiv:2410.13276 [cs.CL]. Retrieved from https:\/\/arxiv.org\/abs\/2410.13276"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","unstructured":"Amir Gholami Sehoon Kim Zhen Dong Zhewei Yao Michael W. Mahoney and Kurt Keutzer. 2021. A survey of quantization methods for efficient neural network inference. arXiv:2103.13630. Retrieved from https:\/\/arxiv.org\/abs\/2103.13630","DOI":"10.1201\/9781003162810-13"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3410463.3414655"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01453-z"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3020078.3021745"},{"key":"e_1_3_2_40_2","unstructured":"Song Han Huizi Mao and William J. Dally. 2015. Deep compression: Compressing deep neural networks with pruning trained quantization and Huffman coding. arXiv:1510.00149. Retrieved from https:\/\/arxiv.org\/abs\/1510.00149"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123939.3123970"},{"issue":"241","key":"e_1_3_2_43_2","first-page":"1","article-title":"Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks","volume":"22","author":"Hoefler Torsten","year":"2021","unstructured":"Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste. 2021. Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. Journal of Machine Learning Research 22, 241 (2021), 1\u2013124.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_44_2","doi-asserted-by":"crossref","unstructured":"Jordan Hoffmann Sebastian Borgeaud Arthur Mensch Elena Buchatskaya Trevor Cai Eliza Rutherford Diego de Las Casas Lisa Anne Hendricks Johannes Welbl Aidan Clark et\u00a0al. 2022. Training compute-optimal large language models. arXiv:2203.15556. Retrieved from https:\/\/arxiv.org\/abs\/2203.15556","DOI":"10.52202\/068431-2176"},{"issue":"2","key":"e_1_3_2_45_2","first-page":"3","article-title":"Lora: Low-rank adaptation of large language models.","volume":"1","author":"Hu Edward J.","year":"2022","unstructured":"Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. Lora: Low-rank adaptation of large language models. ICLR 1, 2 (2022), 3.","journal-title":"ICLR"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00075"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00076"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1088\/2634-4386\/ac7c8a"},{"key":"e_1_3_2_49_2","volume-title":"Proceedings of Machine Learning and Systems (MLSys)","author":"Ivanov Andrei","year":"2021","unstructured":"Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun, Shigang Li, and Torsten Hoefler. 2021. Data movement is all you need: A case study on optimizing transformers. In Proceedings of Machine Learning and Systems (MLSys)."},{"key":"e_1_3_2_50_2","unstructured":"Jared Kaplan Sam McCandlish Tom Henighan Tom B. Brown Benjamin Chess Rewon Child Scott Gray Alec Radford Jeffrey Wu and Dario Amodei. 2020. Scaling laws for neural language models. arXiv:2001.08361. Retrieved from https:\/\/arxiv.org\/abs\/2001.08361"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3133901"},{"key":"e_1_3_2_52_2","volume-title":"Learning Multiple Layers of Features from Tiny Images","author":"Krizhevsky Alex","year":"2009","unstructured":"Alex Krizhevsky. 2009. Learning Multiple Layers of Features from Tiny Images. Technical Report. Retrieved June 8, 2026 from https:\/\/www.cs.toronto.edu\/kriz\/learning-features-2009-TR.pdf"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2020.2972520"},{"key":"e_1_3_2_54_2","volume-title":"Proceedings of the International Conference on Machine Learning (ICML)","author":"Kurtz Mark","year":"2020","unstructured":"Mark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev, John Carr, Michael Goin, William Leiserson, Sage Moore, Nir Shavit, and Dan Alistarh. 2020. Inducing and exploiting activation sparsity for fast inference on deep neural networks. In Proceedings of the International Conference on Machine Learning (ICML). PMLR."},{"key":"e_1_3_2_55_2","unstructured":"Fan\u00e7ois Lagunas. 2020. Block Sparse Matrices for Smaller and Faster Language Models. Retrieved June 8 2026 from https:\/\/huggingface.co\/blog\/pytorch_block_sparse"},{"key":"e_1_3_2_56_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Lan Zhenzhong","year":"2020","unstructured":"Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A lite BERT for self-supervised learning of language representations. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_57_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Li Hao","year":"2017","unstructured":"Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. 2017. Pruning filters for efficient ConvNets. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503221.3508399"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.5555\/3571885.3571934"},{"key":"e_1_3_2_60_2","unstructured":"Sheng Li Jongsoo Park and Ping Tak Peter Tang. 2017. Enabling sparse Winograd convolution by native pruning. arXiv:1702.08597. Retrieved from https:\/\/arxiv.org\/abs\/1702.08597"},{"key":"e_1_3_2_61_2","unstructured":"Yunlu Li and Artsiom Ablavatski. 2021. Build fast sparse on-device models with the new TF MOT pruning API. Retrieved June 8 2026 from https:\/\/blog.tensorflow.org\/2021\/07\/build-fast-sparse-on-device-models-with-tf-mot-pruning-api.html"},{"key":"e_1_3_2_62_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Lin Yujun","year":"2018","unstructured":"Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J. Dally. 2018. Deep gradient compression: Reducing the communication bandwidth for distributed training. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3649816"},{"key":"e_1_3_2_64_2","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Louizos Christos","year":"2017","unstructured":"Christos Louizos, Karen Ullrich, and Max Welling. 2017. Bayesian compression for deep learning. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 30."},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356156"},{"key":"e_1_3_2_66_2","unstructured":"Sourab Mangrulkar Sylvain Gugger Lysandre Debut Younes Belkada Sayak Paul and Benjamin Bossan. 2022. PEFT: State-of-the-art Parameter-Efficient Fine-Tuning methods. Retrieved June 8 2026 from https:\/\/github.com\/huggingface\/peft"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-018-04316-3"},{"key":"e_1_3_2_68_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICML)","author":"Nair Vinod","year":"2010","unstructured":"Vinod Nair and Geoffrey E. Hinton. 2010. Rectified linear units improve restricted Boltzmann machines. In Proceedings of the International Conference on Learning Representations (ICML)."},{"key":"e_1_3_2_69_2","volume-title":"A fork of the PEFT library, supporting Robust Adaptation (RoSA)","author":"Nikdan Mahdi","year":"2024","unstructured":"Mahdi Nikdan, Soroush Tabesh, Elvir Crn\u010devi\u0107, and Dan Alistarh. 2024. A fork of the PEFT library, supporting Robust Adaptation (RoSA). Retrieved June 8, 2026 from https:\/\/github.com\/IST-DASLab\/peft-rosa"},{"key":"e_1_3_2_70_2","unstructured":"Mahdi Nikdan Soroush Tabesh Elvir Crn\u010devi\u0107 and Dan Alistarh. 2024. RoSA: Accurate parameter-efficient fine-tuning via robust adaptation. arXiv:2401.04679. Retrieved from https:\/\/arxiv.org\/abs\/2401.04679"},{"key":"e_1_3_2_71_2","unstructured":"NVIDIA. 2020. NVIDIA A100 Tensor Core GPU Architecture. Retrieved June 8 2026 from https:\/\/images.nvidia.com\/aem-dam\/en-zz\/Solutions\/data-center\/nvidia-ampere-architecture-whitepaper.pdf"},{"key":"e_1_3_2_72_2","unstructured":"OpenAI. 2017. Block-Sparse GPU Kernels. Retrieved June 8 2026 from https:\/\/openai.com\/blog\/block-sparse-gpu-kernels\/"},{"key":"e_1_3_2_73_2","unstructured":"OpenAI. 2018. AI and Compute. Retrieved June 8 2026 from https:\/\/openai.com\/blog\/ai-and-compute\/"},{"key":"e_1_3_2_74_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Park Jongsoo","year":"2016","unstructured":"Jongsoo Park, Sheng Li, Wei Wen, Ping Tak Peter Tang, Hai Li, Yiran Chen, and Pradeep Dubey. 2016. Faster CNNs with direct sparse convolutions and guided pruning. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_75_2","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems (NeurIPS)","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et\u00a0al. 2019. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of the International Conference on Neural Information Processing Systems (NeurIPS). arxiv:1912.01703 [cs.LG]"},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58526-6_31"},{"key":"e_1_3_2_77_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Plummer Bryan A.","year":"2022","unstructured":"Bryan A. Plummer, Nikoli Dryden, Julius Frost, Torsten Hoefler, and Kate Saenko. 2022. Neural parameter allocation search. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.1145\/356616.356618"},{"key":"e_1_3_2_79_2","first-page":"13316","article-title":"Channel permutations for N: M sparsity","author":"Pool Jeff","year":"2021","unstructured":"Jeff Pool and Chong Yu. 2021. Channel permutations for N: M sparsity. In Proceedings of the 35th International Conference on Neural Information Processing Systems. 13316\u201313327.","journal-title":"Proceedings of the 35th International Conference on Neural Information Processing Systems"},{"key":"e_1_3_2_80_2","volume-title":"Proceedings of Machine Learning and Systems (MLSys)","author":"Pope Reiner","year":"2023","unstructured":"Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Levskaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. 2023. Efficiently scaling transformer inference. In Proceedings of Machine Learning and Systems (MLSys)."},{"key":"e_1_3_2_81_2","unstructured":"PyTorch. 2023. Torchvision. Retrieved June 8 2026 from https:\/\/pytorch.org\/vision\/stable\/index.html"},{"key":"e_1_3_2_82_2","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356222"},{"issue":"140","key":"e_1_3_2_83_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research 21, 140 (2020), 1\u201367.","journal-title":"The Journal of Machine Learning Research"},{"key":"e_1_3_2_84_2","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems (NeurIPS)","author":"Sanh Victor","year":"2020","unstructured":"Victor Sanh, Thomas Wolf, and Alexander Rush. 2020. Movement pruning: Adaptive sparsity by fine-tuning. In Proceedings of the International Conference on Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_2_85_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Savarese Pedro","year":"2019","unstructured":"Pedro Savarese and Michael Maire. 2019. Learning implicitly recurrent CNNs through parameter sharing. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_86_2","doi-asserted-by":"crossref","unstructured":"Jaime Sevilla Lennart Heim Anson Ho Tamay Besiroglu Marius Hobbhahn and Pablo Villalobos. 2022. Compute trends across three eras of machine learning. arXiv:2202.05924. Retrieved from https:\/\/arxiv.org\/abs\/2202.05924","DOI":"10.1109\/IJCNN55064.2022.9891914"},{"key":"e_1_3_2_87_2","unstructured":"Noam Shazeer Azalia Mirhoseini Krzysztof Maziarz Andy Davis Quoc Le Geoffrey Hinton and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arxiv:1701.06538 [cs.LG]. Retrieved from https:\/\/arxiv.org\/abs\/1701.06538"},{"key":"e_1_3_2_88_2","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2670313"},{"key":"e_1_3_2_89_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2015-354"},{"key":"e_1_3_2_90_2","unstructured":"Zhenheng Tang Shaohuai Shi Xiaowen Chu Wei Wang and Bo Li. 2020. Communication-efficient distributed deep learning: A comprehensive survey. arXiv:2003.06307. Retrieved from https:\/\/arxiv.org\/abs\/2003.06307"},{"key":"e_1_3_2_91_2","volume-title":"torchao: PyTorch Native Quantization and Sparsity for Training and Inference","author":"Maintainers Torchao","year":"2024","unstructured":"Torchao Maintainers and Contributors. 2024. torchao: PyTorch Native Quantization and Sparsity for Training and Inference. Retrieved June 8, 2026 from https:\/\/github.com\/pytorch\/ao"},{"key":"e_1_3_2_92_2","unstructured":"Pablo Villalobos Jaime Sevilla Tamay Besiroglu Lennart Heim Anson Ho and Marius Hobbhahn. 2022. Machine learning model sizes and the parameter gap. arXiv:2207.02852. Retrieved from https:\/\/arxiv.org\/abs\/2207.02852"},{"key":"e_1_3_2_93_2","doi-asserted-by":"publisher","DOI":"10.1145\/2503210.2503219"},{"key":"e_1_3_2_94_2","doi-asserted-by":"publisher","DOI":"10.1145\/3410463.3414654"},{"key":"e_1_3_2_95_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"e_1_3_2_96_2","unstructured":"Chai Wah Wu. 2018. ProdSumNet: reducing model parameters in deep neural networks via product-of-sums matrix decompositions. arXiv:1809.02209. Retrieved from https:\/\/arxiv.org\/abs\/1809.02209"},{"key":"e_1_3_2_97_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS.2012.97"},{"key":"e_1_3_2_98_2","unstructured":"Ruyi Xu Guangxuan Xiao Haofeng Huang Junxian Guo and Song Han. 2025. XAttention: Block sparse attention with antidiagonal scoring. arxiv:2503.16428 [cs.CL]. Retrieved from https:\/\/arxiv.org\/abs\/2503.16428"},{"key":"e_1_3_2_99_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33015676"},{"key":"e_1_3_2_100_2","unstructured":"Zihao Ye Ruihang Lai Junru Shao Tianqi Chen and Luis Ceze. 2022. SparseTIR: Composable abstractions for sparse compilation in deep learning. arXiv:2207.04606. Retrieved from https:\/\/arxiv.org\/abs\/2207.04606"},{"key":"e_1_3_2_101_2","doi-asserted-by":"publisher","DOI":"10.1145\/3140659.3080215"},{"key":"e_1_3_2_102_2","unstructured":"Jingyang Yuan Huazuo Gao Damai Dai Junyu Luo Liang Zhao Zhengyan Zhang Zhenda Xie Y. X. Wei Lean Wang Zhiping Xiao et\u00a0al. 2025. Native sparse attention: Hardware-aligned and natively trainable sparse attention. arxiv:2502.11089 [cs.CL]. Retrieved from https:\/\/arxiv.org\/abs\/2502.11089"},{"key":"e_1_3_2_103_2","doi-asserted-by":"publisher","DOI":"10.5244\/C.30.87"},{"key":"e_1_3_2_104_2","first-page":"17283","article-title":"Big bird: Transformers for longer sequences","author":"Zaheer Manzil","year":"2020","unstructured":"Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et\u00a0al. 2020. Big bird: Transformers for longer sequences. In Proceedings of the 34th International Conference on Neural Information Processing Systems. 17283\u201317297.","journal-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems"},{"key":"e_1_3_2_105_2","unstructured":"Susan Zhang Stephen Roller Naman Goyal Mikel Artetxe Moya Chen Shuohui Chen Christopher Dewan Mona Diab Xian Li Xi Victoria Lin et\u00a0al. 2022. Opt: Open pre-trained transformer language models. arXiv:2205.01068. Retrieved from https:\/\/arxiv.org\/abs\/2205.01068"},{"key":"e_1_3_2_106_2","first-page":"213","volume-title":"Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Zheng Ningxin","year":"2022","unstructured":"Ningxin Zheng, Bin Lin, Quanlu Zhang, Lingxiao Ma, Yuqing Yang, Fan Yang, Yang Wang, Mao Yang, and Lidong Zhou. 2022. SparTA: Deep-learning model sparsity via tensor-with-sparsity-attribute. In Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). USENIX Association, Carlsbad, CA, 213\u2013232. Retrieved from https:\/\/www.usenix.org\/conference\/osdi22\/presentation\/zheng-ningxin"},{"key":"e_1_3_2_107_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Zhou Aojun","year":"2021","unstructured":"Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li. 2021. Learning N:M fine-grained structured sparse neural networks from scratch. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_108_2","unstructured":"Michael Zhu and Suyog Gupta. 2017. To prune or not to prune: Exploring the efficacy of pruning for model compression. arXiv:1710.01878. Retrieved from https:\/\/arxiv.org\/abs\/1710.01878"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3815424","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T14:16:43Z","timestamp":1782569803000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3815424"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,27]]},"references-count":107,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3815424"],"URL":"https:\/\/doi.org\/10.1145\/3815424","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,27]]},"assertion":[{"value":"2025-07-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-21","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}