{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T14:47:47Z","timestamp":1782571667872,"version":"3.54.5"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T00:00:00Z","timestamp":1782518400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2023YFB3001503"],"award-info":[{"award-number":["2023YFB3001503"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62421002 and 62302505"],"award-info":[{"award-number":["62421002 and 62302505"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Transformer-based large language models (LLM) are increasingly deployed in high-performance computing environments, where the attention mechanism often becomes a key bottleneck during inference. Although state-of-the-art attention algorithms (e.g., FlashAttention) achieve high efficiency on GPUs, they are ill-suited to emerging heterogeneous many-core processors. In this work, we focus on MT-3000, a representative architecture deployed in the new-generation Tianhe supercomputer, and identify three principal challenges in realizing high-performance attention: complex multi-tier memory requiring manual data movement, excessive reduction overhead caused by sub-tile softmax operations, and static execution pipelines that fail to adapt to inference phases and sequence lengths. To overcome these challenges, we propose\n                    <jats:sc>DeferAttention<\/jats:sc>\n                    , a high-performance attention implementation designed for the MT-3000 many-core processor.\n                    <jats:sc>DeferAttention<\/jats:sc>\n                    introduces a novel deferred-reduction attention strategy to decouple reduction from the fused compute pipeline, enabling more efficient aggregation over large tiles. Moreover,\n                    <jats:sc>DeferAttention<\/jats:sc>\n                    adopts a memory-centric operator design, including data tiling, multi-level software pipelining, and modular micro-kernels, to maximize data reuse and execution throughput. Finally, to support runtime-adaptive execution,\n                    <jats:sc>DeferAttention<\/jats:sc>\n                    integrates a lightweight kernel selection strategy guided by an analytical cost model. Experimental results show that\n                    <jats:sc>DeferAttention<\/jats:sc>\n                    achieves up to 98% of the theoretical peak at the micro-kernel level and 85% at the operator level, outperforming baseline implementations and significantly accelerating end-to-end inference.\n                  <\/jats:p>","DOI":"10.1145\/3807449","type":"journal-article","created":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T11:14:23Z","timestamp":1776078863000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Optimizing Attention for Large Language Model Inference on the MT-3000 Many-Core Processor"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8316-2934","authenticated-orcid":false,"given":"Xinxin","family":"Qi","sequence":"first","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3542-4869","authenticated-orcid":false,"given":"Jianbin","family":"Fang","sequence":"additional","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8364-9793","authenticated-orcid":false,"given":"Peng","family":"Zhang","sequence":"additional","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6906-4940","authenticated-orcid":false,"given":"Yonggang","family":"Che","sequence":"additional","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,27]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"[n. d.]. TensorRT-LLM: High-performance Large Language Model Inference. Retrieved from https:\/\/docs.nvidia.com\/tensorrt-llm\/. Accessed 2025."},{"key":"e_1_3_2_3_2","unstructured":"Shantanu Acharya Fei Jia and Boris Ginsburg. 2024. Star Attention: Efficient LLM Inference over Long Sequences. arxiv:2411.17116 [cs.CL] https:\/\/arxiv.org\/abs\/2411.17116"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41404.2022.00051"},{"key":"e_1_3_2_5_2","first-page":"5831","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Beltagy Iz","year":"2020","unstructured":"Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 5831\u20135841."},{"key":"e_1_3_2_6_2","article-title":"On the opportunities and risks of foundation models","author":"Bommasani Rishi","year":"2021","unstructured":"Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et\u00a0al. 2021. On the opportunities and risks of foundation models. arXiv:2108.07258 (2021).","journal-title":"58"},{"key":"e_1_3_2_7_2","unstructured":"Tom B. Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et\u00a0al. 2020. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS\u201920). Curran Associates Inc. Red Hook NY USA Article 159 (2020) 1877\u20131901."},{"key":"e_1_3_2_8_2","volume-title":"International Conference on Learning Representations","author":"Choromanski Krzysztof","year":"2021","unstructured":"Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Ankit Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et\u00a0al. 2021. Rethinking attention with performers. In International Conference on Learning Representations."},{"key":"e_1_3_2_9_2","volume-title":"International Conference on Learning Representations (ICLR)","author":"Dao Tri","year":"2024","unstructured":"Tri Dao. 2024. FlashAttention-2: Faster attention with better parallelism and work partitioning. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_10_2","series-title":"NIPS\u201922","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems","author":"Dao Tri","year":"2024","unstructured":"Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher R\u00e9. 2024. FLASHATTENTION: Fast and memory-efficient exact attention with IO-awareness. In Proceedings of the 36th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS\u201922). Curran Associates Inc., Red Hook, NY, USA, Article 1189, 16 pages."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n19-1423"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3524059.3532372"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1631\/FITEE.2200359"},{"key":"e_1_3_2_14_2","unstructured":"Aaron Grattafiori et\u00a0al. 2024. The Llama 3 Herd of Models. arxiv:2407.21783 [cs.AI] https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","unstructured":"Desta Haileselassie Hagos Rick Battle and Danda B. Rawat. 2024. Recent advances in generative AI and large language models: Current status challenges and perspectives. In IEEE Transactions on Artificial Intelligence 5 12 (Dec. 2024) 5873\u20135893. DOI:10.1109\/TAI.2024.3444742","DOI":"10.1109\/TAI.2024.3444742"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3620666.3651380"},{"key":"e_1_3_2_17_2","first-page":"148","volume-title":"Proceedings of Machine Learning and Systems","volume":"6","author":"Hong Ke","year":"2024","unstructured":"Ke Hong, Guohao Dai, Jiaming Xu, Qiuli Mao, Xiuhong Li, Jun Liu, kangdi chen, Yuhan Dong, and Yu Wang. 2024. FlashDecoding++: Faster large language model inference with asynchronization, flat GEMM optimization, and heuristics. In Proceedings of Machine Learning and Systems, P. Gibbons, G. Pekhimenko, and C. De Sa (Eds.). Vol. 6. 148\u2013161. Retrieved from https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2024\/file\/5321b1dabcd2be188d796c21b733e8c7-Paper-Conference.pdf"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2023.3280805"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3545008.3545022"},{"key":"e_1_3_2_20_2","unstructured":"Alexander Kolesnikov Alexey Dosovitskiy Dirk Weissenborn Georg Heigold Jakob Uszkoreit Lucas Beyer Matthias Minderer Mostafa Dehghani Neil Houlsby Sylvain Gelly et\u00a0al. 2021. An image is worth \\(16\\times 16\\) words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_3_2_22_2","series-title":"NIPS\u201923","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Liu Hao","year":"2024","unstructured":"Hao Liu and Pieter Abbeel. 2024. Blockwise parallel transformers for large context models. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS\u201923). Curran Associates Inc., Red Hook, NY, USA, Article 386, 17 pages."},{"key":"e_1_3_2_23_2","article-title":"World model on million-length video and language with blockwise ringattention","volume":"2402","author":"Liu Hao","year":"2024","unstructured":"Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel. 2024. World model on million-length video and language with blockwise ringattention. ArXiv abs\/2402.08268 (2024). https:\/\/api.semanticscholar.org\/CorpusID:267637090","journal-title":"ArXiv"},{"key":"e_1_3_2_24_2","unstructured":"Hao Liu Matei Zaharia and Pieter Abbeel. 2023. Ring Attention with Blockwise Transformers for Near-Infinite Context. arxiv:2310.01889 [cs.CL] https:\/\/arxiv.org\/abs\/2310.01889"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/s42514-022-00095-y"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2906891"},{"key":"e_1_3_2_27_2","unstructured":"Maxim Milakov and Natalia Gimelshein. 2018. Online normalizer calculation for softmax. arxiv:1805.02867 [cs.PF] https:\/\/arxiv.org\/abs\/1805.02867"},{"key":"e_1_3_2_28_2","unstructured":"Nvidia. 2024. cuBLAS. Retrieved from https:\/\/developer.nvidia.com\/cublas"},{"key":"e_1_3_2_29_2","unstructured":"Nvidia. 2024. cuDNN. Retrieved from https:\/\/developer.nvidia.com\/cublas"},{"key":"e_1_3_2_30_2","unstructured":"Nvidia. 2024. CUTLASS. Retrieved from https:\/\/github.com\/NVIDIA\/cutlass"},{"key":"e_1_3_2_31_2","unstructured":"OpenAI OpenAI. 2023. GPT-4 Technical Report. (Mar2023)."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-024-05927-y"},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","unstructured":"Jay Shah Ganesh Bikshandi Ying Zhang Vijay Thakkar Pradeep Ramani and Tri Dao. 2024. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision. arxiv:2407.08608 [cs.LG] https:\/\/arxiv.org\/abs\/2407.08608","DOI":"10.52202\/079017-2193"},{"key":"e_1_3_2_34_2","series-title":"ICML\u201923","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Sheng Ying","year":"2023","unstructured":"Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher R\u00e9, Ion Stoica, and Ce Zhang. 2023. FlexGen: High-throughput generative inference of large language models with a single GPU. In Proceedings of the 40th International Conference on Machine Learning (Honolulu, Hawaii, USA) (ICML\u201923). JMLR.org, Article 1288, 23 pages."},{"key":"e_1_3_2_35_2","unstructured":"Gemini Team Rohan Anil Sebastian Borgeaud Jean-Baptiste Alayrac Jiahui Yu Radu Soricut Johan Schalkwyk Andrew M. Dai Anja Hauth Katie Millican et\u00a0al. 2023. Gemini: A family of highly capable multimodal models. arXiv:2312.11805 (2023)."},{"key":"e_1_3_2_36_2","unstructured":"Hugo Touvron et\u00a0al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. arxiv:2307.09288 [cs.CL] https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_2_37_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et\u00a0al. 2023. LLaMA: Open and Efficient Foundation Language Models. arxiv:2302.13971 [cs.CL] https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_2_39_2","first-page":"5512","volume-title":"Advances in Neural Information Processing Systems","author":"Wang Sinong","year":"2020","unstructured":"Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. 2020. Linformer: Self-attention with linear complexity. In Advances in Neural Information Processing Systems 33 (2020), 5512\u20135523."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","unstructured":"Teven Le Scao Angela Fan Christopher Akiki et\u00a0al. 2022. BLOOM: A 176B-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100. DOI:10.48550\/arXiv.2211.05100","DOI":"10.48550\/arXiv.2211.05100"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS57955.2024.00090"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS57955.2024.00090"},{"key":"e_1_3_2_43_2","unstructured":"Susan Zhang Stephen Roller Naman Goyal Mikel Artetxe Moya Chen Shuohui Chen Christopher Dewan Mona Diab Xian Li Xi Victoria Lin et\u00a0al. 2022. OPT: Open Pre-trained Transformer Language Models. arxiv:2205.01068 [cs.CL] https:\/\/arxiv.org\/abs\/2205.01068"},{"key":"e_1_3_2_44_2","unstructured":"Susan Zhang Stephen Roller Naman Goyal Mikel Artetxe Moya Chen Shuohui Chen Christopher Dewan Mona Diab Xian Li Xi Victoria Lin et\u00a0al. 2022. Opt: Open pre-trained transformer language models. arXiv:2205.01068 (2022)."},{"key":"e_1_3_2_45_2","first-page":"162","volume-title":"Proceedings of Machine Learning and Systems","volume":"6","author":"Zhao Xuanlei","year":"2024","unstructured":"Xuanlei Zhao, Bin Jia, Haotian Zhou, Ziming Liu, Shenggan Cheng, and Yang You. 2024. HeteGen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices. In Proceedings of Machine Learning and Systems, P. Gibbons, G. Pekhimenko, and C. De Sa (Eds.). Vol. 6. 162\u2013172. Retrieved from https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2024\/file\/5431dca75a8d2abc1fb51e89e8324f10-Paper-Conference.pdf"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS60453.2023.00252"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3807449","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T14:16:14Z","timestamp":1782569774000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3807449"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,27]]},"references-count":45,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3807449"],"URL":"https:\/\/doi.org\/10.1145\/3807449","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,27]]},"assertion":[{"value":"2025-07-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}