{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T13:44:05Z","timestamp":1782999845710,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":79,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,7,5]],"date-time":"2026-07-05T00:00:00Z","timestamp":1783209600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"National Natural Science Foundation of China","award":["62272122"],"award-info":[{"award-number":["62272122"]}]},{"name":"Guangzhou Municipal Joint Funding Project with Universities and Enterprises","award":["2024A03J0616"],"award-info":[{"award-number":["2024A03J0616"]}]},{"name":"Guangzhou Municipality Big Data Intelligence Key Lab","award":["2023A03J0012"],"award-info":[{"award-number":["2023A03J0012"]}]},{"name":"Hong Kong CRF grants","award":["C7004-22G"],"award-info":[{"award-number":["C7004-22G"]}]},{"name":"Hong Kong CRF grants","award":["C6015-23G"],"award-info":[{"award-number":["C6015-23G"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,7,6]]},"DOI":"10.1145\/3797905.3800548","type":"proceedings-article","created":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T11:50:37Z","timestamp":1782993037000},"page":"1053-1064","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["DynSpAttn: Efficient Attention via Dual-Side Dynamic Sparsity on Sparse Tensor Cores"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-2478-1512","authenticated-orcid":false,"given":"Xiangrui","family":"Yu","sequence":"first","affiliation":[{"name":"Baidu Inc., Beijing, China and The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-7492-2069","authenticated-orcid":false,"given":"Ruibo","family":"Fan","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-4381-0544","authenticated-orcid":false,"given":"Zeyu","family":"Li","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2875-0056","authenticated-orcid":false,"given":"Weile","family":"Luo","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-2269-2358","authenticated-orcid":false,"given":"Gu","family":"Gong","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9745-4372","authenticated-orcid":false,"given":"Xiaowen","family":"Chu","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China and The Hong Kong University of Science and Technology, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,5]]},"reference":[{"key":"e_1_3_3_1_2_2","unstructured":"Hamdy Abdelkhalik Yehia Arafa Nandakishore Santhi and Abdel-Hameed Badawy. 2022. Demystifying the Nvidia Ampere Architecture through Microbenchmarking and Instruction-level Analysis. arxiv:https:\/\/arXiv.org\/abs\/2208.11174\u00a0[cs.AR] https:\/\/arxiv.org\/abs\/2208.11174"},{"key":"e_1_3_3_1_3_2","unstructured":"AI@Meta. 2024. Llama 3 Model Card. (2024). https:\/\/github.com\/meta-llama\/llama3\/blob\/main\/MODEL_CARD.md"},{"key":"e_1_3_3_1_4_2","series-title":"(SC \u201922)","volume-title":"Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis","author":"Aminabadi Reza\u00a0Yazdani","year":"2022","unstructured":"Reza\u00a0Yazdani Aminabadi, Samyam Rajbhandari, Ammar\u00a0Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Olatunji Ruwase, Shaden Smith, Minjia Zhang, Jeff Rasley, and Yuxiong He. 2022. DeepSpeed-inference: enabling efficient inference of transformer models at unprecedented scale. In Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis (Dallas, Texas) (SC \u201922). IEEE Press, Article 46, 15\u00a0pages."},{"key":"e_1_3_3_1_5_2","unstructured":"Jinze Bai Shuai Bai Yunfei Chu Zeyu Cui Kai Dang Xiaodong Deng Yang Fan Wenbin Ge Yu Han Fei Huang Binyuan Hui Luo Ji Mei Li Junyang Lin Runji Lin Dayiheng Liu Gao Liu Chengqiang Lu Keming Lu Jianxin Ma Rui Men Xingzhang Ren Xuancheng Ren Chuanqi Tan Sinan Tan Jianhong Tu Peng Wang Shijie Wang Wei Wang Shengguang Wu Benfeng Xu Jin Xu An Yang Hao Yang Jian Yang Shusheng Yang Yang Yao Bowen Yu Hongyi Yuan Zheng Yuan Jianwei Zhang Xingxuan Zhang Yichang Zhang Zhenru Zhang Chang Zhou Jingren Zhou Xiaohuan Zhou and Tianhang Zhu. 2023. Qwen Technical Report. arxiv:https:\/\/arXiv.org\/abs\/2309.16609\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2309.16609"},{"key":"e_1_3_3_1_6_2","unstructured":"Iz Beltagy Matthew\u00a0E Peters and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2004.05150 (2020)."},{"key":"e_1_3_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6239"},{"key":"e_1_3_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3559009.3569691"},{"key":"e_1_3_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581784.3607087"},{"key":"e_1_3_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476182"},{"key":"e_1_3_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3572848.3577500"},{"key":"e_1_3_3_1_12_2","unstructured":"Wei-Lin Chiang Lianmin Zheng Ying Sheng Anastasios\u00a0Nikolas Angelopoulos Tianle Li Dacheng Li Hao Zhang Banghua Zhu Michael Jordan Joseph\u00a0E. Gonzalez and Ion Stoica. 2024. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arxiv:https:\/\/arXiv.org\/abs\/2403.04132\u00a0[cs.AI] https:\/\/arxiv.org\/abs\/2403.04132"},{"key":"e_1_3_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3489517.3530508"},{"key":"e_1_3_3_1_14_2","unstructured":"Tri Dao. 2023. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning. arxiv:https:\/\/arXiv.org\/abs\/2307.08691\u00a0[cs.LG] https:\/\/arxiv.org\/abs\/2307.08691"},{"key":"e_1_3_3_1_15_2","unstructured":"Tri Dao Daniel\u00a0Y. Fu Stefano Ermon Atri Rudra and Christopher R\u00e9. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. arxiv:https:\/\/arXiv.org\/abs\/2205.14135\u00a0[cs.LG] https:\/\/arxiv.org\/abs\/2205.14135"},{"key":"e_1_3_3_1_16_2","unstructured":"Juechu Dong Boyuan Feng Driss Guessous Yanbo Liang and Horace He. 2024. Flex Attention: A Programming Model for Generating Optimized Attention Kernels. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2412.05496 (2024)."},{"key":"e_1_3_3_1_17_2","unstructured":"Dayou Du Shijie Cao Jianyi Cheng Ting Cao and Mao Yang. 2025. BitDecoding: Unlocking Tensor Cores for Long-Context LLMs Decoding with Low-Bit KV Cache. arxiv:https:\/\/arXiv.org\/abs\/2503.18773\u00a0[cs.AR] https:\/\/arxiv.org\/abs\/2503.18773"},{"key":"e_1_3_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3620666.3651378"},{"key":"e_1_3_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3689031.3717481"},{"key":"e_1_3_3_1_20_2","doi-asserted-by":"crossref","unstructured":"Gongfan Fang Hongxu Yin Saurav Muralidharan Greg Heinrich Jeff Pool Jan Kautz Pavlo Molchanov and Xinchao Wang. 2024. Maskllm: Learnable semi-structured sparsity for large language models. Advances in Neural Information Processing Systems 37 (2024) 7736\u20137758.","DOI":"10.52202\/079017-0248"},{"key":"e_1_3_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Elias Frantar and Dan Alistarh. 2022. Optimal brain compression: A framework for accurate post-training quantization and pruning. Advances in Neural Information Processing Systems 35 (2022) 4475\u20134488.","DOI":"10.52202\/068431-0323"},{"key":"e_1_3_3_1_22_2","unstructured":"Trevor Gale Matei Zaharia Cliff Young and Erich Elsen. 2020. Sparse GPU Kernels for Deep Learning. arxiv:https:\/\/arXiv.org\/abs\/2006.10901\u00a0[cs.LG] https:\/\/arxiv.org\/abs\/2006.10901"},{"key":"e_1_3_3_1_23_2","unstructured":"Yizhao Gao Zhichen Zeng Dayou Du Shijie Cao Hayden Kwok-Hay So Ting Cao Fan Yang and Mao Yang. 2024. SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs. ArXiv. https:\/\/www.microsoft.com\/en-us\/research\/publication\/seerattention-learning-intrinsic-sparse-attention-in-your-llms\/"},{"key":"e_1_3_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3676641.3716011"},{"key":"e_1_3_3_1_25_2","unstructured":"Aaron Grattafiori Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian et\u00a0al. 2024. The Llama 3 Herd of Models. arxiv:https:\/\/arXiv.org\/abs\/2407.21783\u00a0[cs.AI] https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_3_1_26_2","unstructured":"Scott Gray Alec Radford and Diederik\u00a0P Kingma. 2017. Block-sparse gpu kernels."},{"key":"e_1_3_3_1_27_2","unstructured":"Daya Guo Qihao Zhu Dejian Yang Zhenda Xie Kai Dong Wentao Zhang Guanting Chen Xiao Bi Y. Wu Y.\u00a0K. Li Fuli Luo Yingfei Xiong and Wenfeng Liang. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming \u2013 The Rise of Code Intelligence. arxiv:https:\/\/arXiv.org\/abs\/2401.14196\u00a0[cs.SE] https:\/\/arxiv.org\/abs\/2401.14196"},{"key":"e_1_3_3_1_28_2","unstructured":"Connor Holmes Masahiro Tanaka Michael Wyatt Ammar\u00a0Ahmad Awan Jeff Rasley Samyam Rajbhandari Reza\u00a0Yazdani Aminabadi Heyang Qin Arash Bakhtiari Lev Kurilenko and Yuxiong He. 2024. DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference. arxiv:https:\/\/arXiv.org\/abs\/2401.08671\u00a0[cs.PF] https:\/\/arxiv.org\/abs\/2401.08671"},{"key":"e_1_3_3_1_29_2","unstructured":"Connor Holmes Minjia Zhang Yuxiong He and Bo Wu. 2021. Nxmtransformer: Semi-structured sparsification for natural language understanding via admm. Advances in neural information processing systems 34 (2021) 1818\u20131830."},{"key":"e_1_3_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3208040.3208062"},{"key":"e_1_3_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3293883.3295712"},{"key":"e_1_3_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00076"},{"key":"e_1_3_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i23.34592"},{"key":"e_1_3_3_1_34_2","unstructured":"Binyuan Hui Jian Yang Zeyu Cui Jiaxi Yang Dayiheng Liu et\u00a0al. 2024. Qwen2.5-Coder Technical Report. arxiv:https:\/\/arXiv.org\/abs\/2409.12186\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2409.12186"},{"key":"e_1_3_3_1_35_2","unstructured":"jcaip. 2024. [Feature Request] Add 2:4 sparsity support #6891. GitHub Issue. https:\/\/github.com\/openai\/triton\/issues\/6891 Accessed: [August 10 2025]."},{"key":"e_1_3_3_1_36_2","doi-asserted-by":"crossref","unstructured":"Huiqiang Jiang Yucheng Li Chengruidong Zhang Qianhui Wu Xufang Luo Surin Ahn Zhenhua Han Amir Abdi Dongsheng Li and Chin-Yew Lin. 1. others. 2024. Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention. Advances in Neural Information Processing Systems 37 (1) 52481\u201352515.","DOI":"10.52202\/079017-1663"},{"key":"e_1_3_3_1_37_2","doi-asserted-by":"crossref","unstructured":"Huiqiang Jiang Yucheng Li Chengruidong Zhang Qianhui Wu Xufang Luo Surin Ahn Zhenhua Han Amir\u00a0H. Abdi Dongsheng Li Chin-Yew Lin Yuqing Yang and Lili Qiu. 2024. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention. arxiv:https:\/\/arXiv.org\/abs\/2407.02490\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2407.02490","DOI":"10.52202\/079017-1663"},{"key":"e_1_3_3_1_38_2","volume-title":"Speech and Language Processing (3rd ed. draft)","author":"Jurafsky Dan","year":"2023","unstructured":"Dan Jurafsky and James\u00a0H. Martin. 2023. Speech and Language Processing (3rd ed. draft). Prentice Hall."},{"key":"e_1_3_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_3_3_1_40_2","unstructured":"Xunhao Lai Jianqiao Lu Yao Luo Yiyuan Ma and Xun Zhou. 2025. Flexprefill: A context-aware sparse attention mechanism for efficient long-sequence inference. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2502.20766 (2025)."},{"key":"e_1_3_3_1_41_2","unstructured":"Zeyu Li Chuanfu Xiao Yang Wang Xiang Liu Zhenheng Tang Baotong Lu Mao Yang Xinyu Chen and Xiaowen Chu. 2025. AnTKV: Anchor Token-Aware Sub-Bit Vector Quantization for KV Cache in Large Language Models. arxiv:https:\/\/arXiv.org\/abs\/2506.19505\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2506.19505"},{"key":"e_1_3_3_1_42_2","first-page":"513","volume-title":"Proceedings of Machine Learning and Systems","volume":"5","author":"Lin Bin","year":"2023","unstructured":"Bin Lin, Ningxin Zheng, Lei Wang, Shijie Cao, Lingxiao Ma, Quanlu Zhang, Yi Zhu, Ting Cao, Jilong Xue, Yuqing Yang, and Fan Yang. 2023. Efficient GPU Kernels for N:M-Sparse Weights in Deep Learning. In Proceedings of Machine Learning and Systems , D.\u00a0Song, M.\u00a0Carbin, and T.\u00a0Chen (Eds.), Vol.\u00a05. Curan, 513\u2013525. https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2023\/file\/a10deb4d5227a8ea307ea8ff3cb712f4-Paper-mlsys2023.pdf"},{"key":"e_1_3_3_1_43_2","unstructured":"Yujun Lin Haotian Tang Shang Yang Zhekai Zhang Guangxuan Xiao Chuang Gan and Song Han. 2025. QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving. arxiv:https:\/\/arXiv.org\/abs\/2405.04532\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2405.04532"},{"key":"e_1_3_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA61900.2025.00075"},{"key":"e_1_3_3_1_45_2","unstructured":"Xiang Liu Zhenheng Tang Hong Chen Peijie Dong Zeyu Li Xiuze Zhou Bo Li Xuming Hu and Xiaowen Chu. 2025. Can LLMs Maintain Fundamental Abilities under KV Cache Compression? arxiv:https:\/\/arXiv.org\/abs\/2502.01941\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2502.01941"},{"key":"e_1_3_3_1_46_2","unstructured":"Xiang Liu Zhenheng Tang Peijie Dong Zeyu Li Yue Liu Bo Li Xuming Hu and Xiaowen Chu. 2025. ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference. arxiv:https:\/\/arXiv.org\/abs\/2502.00299\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2502.00299"},{"key":"e_1_3_3_1_47_2","first-page":"32332","volume-title":"International Conference on Machine Learning, ICML 2024","author":"Liu Zirui","year":"2024","unstructured":"Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu. 2024. KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache. In International Conference on Machine Learning, ICML 2024. PMLR, 32332\u201332344."},{"key":"e_1_3_3_1_48_2","unstructured":"Enzhe Lu Zhejun Jiang Jingyuan Liu Yulun Du Tao Jiang Chao Hong Shaowei Liu Weiran He Enming Yuan Yuzhi Wang Zhiqi Huang Huan Yuan Suting Xu Xinran Xu Guokun Lai Yanru Chen Huabin Zheng Junjie Yan Jianlin Su Yuxin Wu Neo\u00a0Y. Zhang Zhilin Yang Xinyu Zhou Mingxing Zhang and Jiezhong Qiu. 2025. MoBA: Mixture of Block Attention for Long-Context LLMs. arxiv:https:\/\/arXiv.org\/abs\/2502.13189\u00a0[cs.LG] https:\/\/arxiv.org\/abs\/2502.13189"},{"key":"e_1_3_3_1_49_2","unstructured":"Weile Luo Ruibo Fan Zeyu Li Dayou Du Qiang Wang and Xiaowen Chu. 2024. Benchmarking and Dissecting the Nvidia Hopper GPU Architecture. arxiv:https:\/\/arXiv.org\/abs\/2402.13499\u00a0[cs.AR] https:\/\/arxiv.org\/abs\/2402.13499"},{"key":"e_1_3_3_1_50_2","unstructured":"Cong Ma Du Wu Zhelang Deng Jiang Chen Xiaowen Huang Jintao Meng Wenxi Zhu Bingqiang Wang Amelie\u00a0Chi Zhou Peng Chen et\u00a0al. 2025. Nm-spmm: Accelerating matrix multiplication using n: M sparsity with gpgpu. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2503.01253 (2025)."},{"key":"e_1_3_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2017.241"},{"key":"e_1_3_3_1_52_2","unstructured":"Asit Mishra Jorge\u00a0Albericio Latorre Jeff Pool Darko Stosic Dusan Stosic Ganesh Venkatesh Chong Yu and Paulius Micikevicius. 2021. Accelerating sparse deep neural networks. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2104.08378 (2021)."},{"key":"e_1_3_3_1_53_2","volume-title":"GPU Technology Conference","author":"Naumov Maxim","year":"2010","unstructured":"Maxim Naumov, L Chien, Philippe Vandermersch, and Ujval Kapasi. 2010. Cusparse library. In GPU Technology Conference."},{"key":"e_1_3_3_1_54_2","unstructured":"Alexander Novikov Ng\u00e2n V\u0169 Marvin Eisenberger Emilien Dupont Po-Sen Huang Adam\u00a0Zsolt Wagner Sergey Shirobokov Borislav Kozlovskii Francisco J.\u00a0R. Ruiz Abbas Mehrabian M.\u00a0Pawan Kumar Abigail See Swarat Chaudhuri George Holland Alex Davies Sebastian Nowozin Pushmeet Kohli and Matej Balog. 2025. AlphaEvolve: A coding agent for scientific and algorithmic discovery. arxiv:https:\/\/arXiv.org\/abs\/2506.13131\u00a0[cs.AI] https:\/\/arxiv.org\/abs\/2506.13131"},{"key":"e_1_3_3_1_55_2","unstructured":"NVIDIA. 2017. NVIDIA Volta GPU Architecture Whitepaper. https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf."},{"key":"e_1_3_3_1_56_2","unstructured":"NVIDIA. 2020. NVIDIA Ampere GA102 GPU Architecture Whitepaper. https:\/\/www.nvidia.com\/content\/PDF\/nvidia-ampere-ga-102-gpu-architecture-whitepaper-v2.pdf."},{"key":"e_1_3_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3627535.3638470"},{"key":"e_1_3_3_1_58_2","first-page":"90","volume-title":"Proceedings of the AAAI Spring Symposium on Logical Formalizations of Commonsense Reasoning","author":"Roemmele Melissa","year":"2011","unstructured":"Melissa Roemmele, Cosmin\u00a0Adrian Bejan, and Andrew\u00a0S. Gordon. 2011. Choice of Plausible Alternatives: An Evaluation of Commonsense Causal Reasoning. In Proceedings of the AAAI Spring Symposium on Logical Formalizations of Commonsense Reasoning. 90\u201395."},{"key":"e_1_3_3_1_59_2","doi-asserted-by":"crossref","unstructured":"Jay Shah Ganesh Bikshandi Ying Zhang Vijay Thakkar Pradeep Ramani and Tri Dao. 2024. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision. arxiv:https:\/\/arXiv.org\/abs\/2407.08608\u00a0[cs.LG] https:\/\/arxiv.org\/abs\/2407.08608","DOI":"10.52202\/079017-2193"},{"key":"e_1_3_3_1_60_2","doi-asserted-by":"crossref","unstructured":"Wei Sun Ang Li Tong Geng Sander Stuijk and Henk Corporaal. 2022. Dissecting Tensor Cores via Microbenchmarks: Latency Throughput and Numeric Behaviors. IEEE Transactions on Parallel and Distributed Systems 34 1 (2022) 246\u2013261.","DOI":"10.1109\/TPDS.2022.3217824"},{"key":"e_1_3_3_1_61_2","unstructured":"Rohan Taori Ishaan Gulrajani Tianyi Zhang Yann Dubois Xuechen Li Carlos Guestrin Percy Liang and Tatsunori\u00a0B. Hashimoto. 2023. Stanford Alpaca: An Instruction-following LLaMA Model. https:\/\/github.com\/tatsu-lab\/stanford_alpaca."},{"key":"e_1_3_3_1_62_2","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan\u00a0N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_3_3_1_63_2","first-page":"149","volume-title":"2023 USENIX Annual Technical Conference (USENIX ATC 23)","author":"Wang Yuke","year":"2023","unstructured":"Yuke Wang, Boyuan Feng, Zheng Wang, Guyue Huang, and Yufei Ding. 2023. TC-GNN: Bridging Sparse GNN Computation and Dense Tensor Cores on GPUs. In 2023 USENIX Annual Technical Conference (USENIX ATC 23). 149\u2013164."},{"key":"e_1_3_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3689031.3717455"},{"key":"e_1_3_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3689031.3717455"},{"key":"e_1_3_3_1_66_2","doi-asserted-by":"publisher","unstructured":"Haojun Xia Zhen Zheng Yuchao Li Donglin Zhuang Zhongzhu Zhou Xiafei Qiu Yong Li Wei Lin and Shuaiwen\u00a0Leon Song. 2023. Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity. Proc. VLDB Endow. 17 2 (Oct. 2023) 211\u2013224. 10.14778\/3626292.3626303","DOI":"10.14778\/3626292.3626303"},{"key":"e_1_3_3_1_67_2","unstructured":"Guangxuan Xiao Jiaming Tang Jingwei Zuo Junxian Guo Shang Yang Haotian Tang Yao Fu and Song Han. 2024. DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads. arXiv (2024)."},{"key":"e_1_3_3_1_68_2","unstructured":"Guangxuan Xiao Yuandong Tian Beidi Chen Song Han and Mike Lewis. 2023. Efficient streaming language models with attention sinks. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2309.17453 (2023)."},{"key":"e_1_3_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-96983-1_48"},{"key":"e_1_3_3_1_70_2","unstructured":"Jingyang Yuan Huazuo Gao Damai Dai Junyu Luo Liang Zhao Zhengyan Zhang Zhenda Xie Y.\u00a0X. Wei Lean Wang Zhiping Xiao Yuqing Wang Chong Ruan Ming Zhang Wenfeng Liang and Wangding Zeng. 2025. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention. arxiv:https:\/\/arXiv.org\/abs\/2502.11089\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2502.11089"},{"key":"e_1_3_3_1_71_2","unstructured":"Jingyang Yuan Huazuo Gao Damai Dai Junyu Luo Liang Zhao Zhengyan Zhang Zhenda Xie Y.\u00a0X. Wei Lean Wang Zhiping Xiao Yuqing Wang Chong Ruan Ming Zhang Wenfeng Liang and Wangding Zeng. 2025. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention. arxiv:https:\/\/arXiv.org\/abs\/2502.11089\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2502.11089"},{"key":"e_1_3_3_1_72_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1472"},{"key":"e_1_3_3_1_73_2","unstructured":"Jintao Zhang Haofeng Huang Pengle Zhang Jia Wei Jun Zhu and Jianfei Chen. 2025. SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization. arxiv:https:\/\/arXiv.org\/abs\/2411.10958\u00a0[cs.LG] https:\/\/arxiv.org\/abs\/2411.10958"},{"key":"e_1_3_3_1_74_2","unstructured":"Jintao Zhang Jia Wei Pengle Zhang Xiaoming Xu Haofeng Huang Haoxu Wang Kai Jiang Jun Zhu and Jianfei Chen. 2025. SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training. arxiv:https:\/\/arXiv.org\/abs\/2505.11594\u00a0[cs.LG] https:\/\/arxiv.org\/abs\/2505.11594"},{"key":"e_1_3_3_1_75_2","volume-title":"International Conference on Learning Representations (ICLR)","author":"Zhang Jintao","year":"2025","unstructured":"Jintao Zhang, Jia Wei, Pengle Zhang, Jun Zhu, and Jianfei Chen. 2025. SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_3_1_76_2","unstructured":"Yilong Zhao Chien-Yu Lin Kan Zhu Zihao Ye Lequn Chen Size Zheng Luis Ceze Arvind Krishnamurthy Tianqi Chen and Baris Kasikci. 2024. Atom: Low-bit Quantization for Efficient and Accurate LLM Serving. arxiv:https:\/\/arXiv.org\/abs\/2310.19102\u00a0[cs.LG] https:\/\/arxiv.org\/abs\/2310.19102"},{"key":"e_1_3_3_1_77_2","series-title":"(NIPS \u201924)","volume-title":"Proceedings of the 38th International Conference on Neural Information Processing Systems","author":"Zheng Lianmin","year":"2025","unstructured":"Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody\u00a0Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph\u00a0E. Gonzalez, Clark Barrett, and Ying Sheng. 2025. SGLang: efficient execution of structured language model programs. In Proceedings of the 38th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS \u201924). Curran Associates Inc., Red Hook, NY, USA, Article 2000, 27\u00a0pages."},{"key":"e_1_3_3_1_78_2","first-page":"213","volume-title":"16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Zheng Ningxin","year":"2022","unstructured":"Ningxin Zheng, Bin Lin, Quanlu Zhang, Lingxiao Ma, Yuqing Yang, Fan Yang, Yang Wang, Mao Yang, and Lidong Zhou. 2022. SparTA: Deep-Learning Model Sparsity via Tensor-with-Sparsity-Attribute. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). 213\u2013232."},{"key":"e_1_3_3_1_79_2","volume-title":"International Conference on Learning Representations","author":"Zhou Aojun","unstructured":"Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li. [n. d.]. Learning N: M Fine-grained Structured Sparse Neural Networks From Scratch. In International Conference on Learning Representations."},{"key":"e_1_3_3_1_80_2","volume-title":"Eighth Conference on Machine Learning and Systems","author":"Zhu Qianchao","unstructured":"Qianchao Zhu, Jiangfei Duan, Chang Chen, Siran Liu, Xiuhong Li, Guanyu Feng, Xin Lv, Xiao Chuanfu, Dahua Lin, and Chao Yang. [n. d.]. SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention. In Eighth Conference on Machine Learning and Systems."}],"event":{"name":"ICS '26: 2026 International Conference on Supercomputing","location":"Belfast United Kingdom","acronym":"ICS '26","sponsor":["SIGHPC ACM Special Interest Group on High Performance Computing, Special Interest Group on High Performance Computing","SIGARCH ACM Special Interest Group on Computer Architecture"]},"container-title":["Proceedings of the 40th ACM International Conference on Supercomputing"],"original-title":[],"deposited":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T12:56:06Z","timestamp":1782996966000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3797905.3800548"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,5]]},"references-count":79,"alternative-id":["10.1145\/3797905.3800548","10.1145\/3797905"],"URL":"https:\/\/doi.org\/10.1145\/3797905.3800548","relation":{},"subject":[],"published":{"date-parts":[[2026,7,5]]},"assertion":[{"value":"2026-07-05","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}