{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T12:52:11Z","timestamp":1782996731325,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":51,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,7,5]],"date-time":"2026-07-05T00:00:00Z","timestamp":1783209600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,7,6]]},"DOI":"10.1145\/3797905.3800537","type":"proceedings-article","created":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T11:50:37Z","timestamp":1782993037000},"page":"342-352","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["SPPO: Making Million-Token LLM Training Practical on Modest GPU Clusters"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-0779-3305","authenticated-orcid":false,"given":"qiaoling","family":"chen","sequence":"first","affiliation":[{"name":"Nanyang Technology University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2037-2496","authenticated-orcid":false,"given":"Shenggui","family":"Li","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7048-1722","authenticated-orcid":false,"given":"Wei","family":"Gao","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hongkong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8456-0491","authenticated-orcid":false,"given":"Peng","family":"Sun","sequence":"additional","affiliation":[{"name":"Shanghai AI Laboratory, SHANG HAI, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2751-5114","authenticated-orcid":false,"given":"Yonggang","family":"Wen","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6595-6650","authenticated-orcid":false,"given":"Tianwei","family":"Zhang","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,5]]},"reference":[{"key":"e_1_3_3_2_2_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia\u00a0Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et\u00a0al. 2023. Gpt-4 technical report. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2303.08774 (2023)."},{"key":"e_1_3_3_2_3_2","unstructured":"Microsoft\u00a0Research AI4Science and Microsoft\u00a0Azure Quantum. 2023. The impact of large language models on scientific discovery: a preliminary study using gpt-4. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2311.07361 (2023)."},{"key":"e_1_3_3_2_4_2","doi-asserted-by":"crossref","unstructured":"Kaifeng Bi Lingxi Xie Hengheng Zhang Xin Chen Xiaotao Gu and Qi Tian. 2023. Accurate medium-range global weather forecasting with 3D neural networks. Nature 619 7970 (2023) 533\u2013538.","DOI":"10.1038\/s41586-023-06185-3"},{"key":"e_1_3_3_2_5_2","unstructured":"Qiaoling Chen Diandian Gu Guoteng Wang Xun Chen YingTong Xiong Ting Huang Qinghao Hu Xin Jin Yonggang Wen Tianwei Zhang et\u00a0al. 2024. Internevo: Efficient long-sequence large language model training via hybrid parallelism and redundant sharding. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2401.09149 (2024)."},{"key":"e_1_3_3_2_6_2","unstructured":"Qiaoling Chen Qinghao Hu Zhisheng Ye Guoteng Wang Peng Sun Yonggang Wen and Tianwei Zhang. 2023. AMSP: Super-Scaling LLM Training via Advanced Model States Partitioning. CoRR abs\/2311.00257 (2023)."},{"key":"e_1_3_3_2_7_2","unstructured":"Tianqi Chen Bing Xu Chiyuan Zhang and Carlos Guestrin. 2016. Training deep nets with sublinear memory cost. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/1604.06174 (2016)."},{"key":"e_1_3_3_2_8_2","unstructured":"Tianqi Chen Bing Xu Chiyuan Zhang and Carlos Guestrin. 2016. Training deep nets with sublinear memory cost. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/1604.06174 (2016)."},{"key":"e_1_3_3_2_9_2","unstructured":"Tri Dao and Albert Gu. 2024. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2405.21060 (2024)."},{"key":"e_1_3_3_2_10_2","unstructured":"Kimi Developers. 2024. kimi. https:\/\/kimi.moonshot.cn\/"},{"key":"e_1_3_3_2_11_2","unstructured":"NVIDIA Developers. 2023. NVIDIA Transformer Engine Offloading strategy. https:\/\/docs.nvidia.com\/deeplearning\/transformer-engine\/user-guide\/api\/pytorch.html#transformer_engine.pytorch.get_cpu_offload_context."},{"key":"e_1_3_3_2_12_2","unstructured":"NVIDIA Developers. 2023. TransformerEngine. https:\/\/github.com\/NVIDIA\/ TransformerEngine"},{"key":"e_1_3_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3437801.3441593"},{"key":"e_1_3_3_2_14_2","unstructured":"Yao Fu Rameswar Panda Xinyao Niu Xiang Yue Hannaneh Hajishirzi Yoon Kim and Hao Peng. 2024. Data engineering for scaling language models to 128k context. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2402.10171 (2024)."},{"key":"e_1_3_3_2_15_2","unstructured":"Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2312.00752 (2023)."},{"key":"e_1_3_3_2_16_2","unstructured":"Diandian Gu Peng Sun Qinghao Hu Ting Huang Xun Chen Yingtong Xiong Guoteng Wang Qiaoling Chen Shangchun Zhao Jiarui Fang et\u00a0al. 2024. Loongtrain: Efficient training of long-sequence llms with head-context parallelism. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2406.18485 (2024)."},{"key":"e_1_3_3_2_17_2","unstructured":"Daya Guo Qihao Zhu Dejian Yang Zhenda Xie Kai Dong Wentao Zhang Guanting Chen Xiao Bi Yu Wu YK Li et\u00a0al. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming\u2013The Rise of Code Intelligence. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2401.14196 (2024)."},{"key":"e_1_3_3_2_18_2","series-title":"(NIPS\u201919)","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"Huang Yanping","year":"2019","unstructured":"Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia\u00a0Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc\u00a0V. Le, Yonghui Wu, and Zhifeng Chen. 2019. GPipe: efficient training of giant neural networks using pipeline parallelism. In Proceedings of the 33rd International Conference on Neural Information Processing Systems(NIPS\u201919). Curran Associates Inc."},{"key":"e_1_3_3_2_19_2","unstructured":"Sam\u00a0Ade Jacobs Masahiro Tanaka Chengming Zhang Minjia Zhang Shuaiwen\u00a0Leon Song Samyam Rajbhandari and Yuxiong He. 2023. DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models. CoRR abs\/2309.14509 (2023)."},{"key":"e_1_3_3_2_20_2","unstructured":"Albert\u00a0Q Jiang Alexandre Sablayrolles Arthur Mensch Chris Bamford Devendra\u00a0Singh Chaplot Diego de\u00a0las Casas Florian Bressand Gianna Lengyel Guillaume Lample Lucile Saulnier et\u00a0al. 2023. Mistral 7B. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2310.06825 (2023)."},{"key":"e_1_3_3_2_21_2","unstructured":"Marisa Kirisame Steven Lyubomirsky Altan Haan Jennifer Brennan Mike He Jared Roesch Tianqi Chen and Zachary Tatlock. 2020. Dynamic tensor rematerialization. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2006.09616 (2020)."},{"key":"e_1_3_3_2_22_2","unstructured":"Vijay\u00a0Anand Korthikanti Jared Casper Sangkug Lym Lawrence McAfee Michael Andersch Mohammad Shoeybi and Bryan Catanzaro. 2023. Reducing activation recomputation in large transformer models. Proceedings of Machine Learning and Systems 5 (2023) 341\u2013353."},{"key":"e_1_3_3_2_23_2","unstructured":"Dacheng Li Rulin Shao Anze Xie Eric\u00a0P Xing Xuezhe Ma Ion Stoica Joseph\u00a0E Gonzalez and Hao Zhang. 2023. Distflashattn: Distributed memory-efficient attention for long-context llms training. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2310.03294 (2023)."},{"key":"e_1_3_3_2_24_2","unstructured":"Shenggui Li Fuzhao Xue Chaitanya Baranwal Yongbin Li and Yang You. 2021. Sequence parallelism: Long sequence training from system perspective. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2105.13120 (2021)."},{"key":"e_1_3_3_2_25_2","unstructured":"Shen Li Yanli Zhao Rohan Varma Omkar Salpekar Pieter Noordhuis Teng Li Adam Paszke Jeff Smith Brian Vaughan Pritam Damania and Soumith Chintala. 2020. PyTorch Distributed: Experiences on Accelerating Data Parallel Training. CoRR abs\/2006.15704 (2020)."},{"key":"e_1_3_3_2_26_2","unstructured":"Zhuohan Li Siyuan Zhuang Shiyuan Guo Danyang Zhuo Hao Zhang Dawn Song and Ion Stoica. 2021. TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models. CoRR abs\/2102.07988 (2021)."},{"key":"e_1_3_3_2_27_2","unstructured":"Hao Liu Matei Zaharia and Pieter Abbeel. 2023. Ring Attention with Blockwise Transformers for Near-Infinite Context. CoRR abs\/2310.01889 (2023)."},{"key":"e_1_3_3_2_28_2","unstructured":"LLaMA2-7B-32K. 2022. LLaMA2-7B-32K. https:\/\/huggingface.co\/togethercomputer\/ LLaMA- 2- 7B- 32K"},{"key":"e_1_3_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359646"},{"key":"e_1_3_3_2_30_2","doi-asserted-by":"crossref","unstructured":"Deepak Narayanan Mohammad Shoeybi Jared Casper Patrick LeGresley Mostofa Patwary Vijay\u00a0Anand Korthikanti Dmitri Vainbrand Prethvi Kashinkunti Julie Bernauer Bryan Catanzaro Amar Phanishayee and Matei Zaharia. 2021. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM. CoRR abs\/2104.04473 (2021).","DOI":"10.1145\/3458817.3476209"},{"key":"e_1_3_3_2_31_2","unstructured":"Houlsby Neil and Weissenborn Dirk. 2020. Transformers for Image Recognition at Scale. Online: https:\/\/ai.googleblog.com\/2020\/12\/transformers-for-image-recognitionat.html (2020)."},{"key":"e_1_3_3_2_32_2","unstructured":"PCIeGen5. 2024. PCIeGen5. https:\/\/www.intel.com\/content\/www\/us\/en\/docs\/programmable\/683501\/23-1-9-0-0\/features.html"},{"key":"e_1_3_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"e_1_3_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378505"},{"key":"e_1_3_3_2_35_2","unstructured":"PyTorch. 2023. Accelerating Large Language Models with Accelerated Transformers. https:\/\/pytorch.org\/blog\/accelerating-large-language-models\/."},{"key":"e_1_3_3_2_36_2","unstructured":"Penghui Qi Xinyi Wan Guangxing Huang and Min Lin. 2023. Zero bubble pipeline parallelism. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2401.10241 (2023)."},{"key":"e_1_3_3_2_37_2","unstructured":"qwen Developers. 2024. qwen. https:\/\/tongyi.aliyun.com\/qianwen\/"},{"key":"e_1_3_3_2_38_2","series-title":"(SC \u201920)","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Rajbhandari Samyam","year":"2020","unstructured":"Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. ZeRO: memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis(SC \u201920). IEEE Press."},{"key":"e_1_3_3_2_39_2","first-page":"551","volume-title":"2021 USENIX Annual Technical Conference (USENIX ATC 21)","author":"Ren Jie","year":"2021","unstructured":"Jie Ren, Samyam Rajbhandari, Reza\u00a0Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021. { Zero-offload} : Democratizing { billion-scale} model training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21). 551\u2013564."},{"key":"e_1_3_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.5555\/3195638.3195660"},{"key":"e_1_3_3_2_41_2","unstructured":"Baptiste Roziere Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing\u00a0Ellen Tan Yossi Adi Jingyu Liu Romain Sauvestre Tal Remez et\u00a0al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2308.12950 (2023)."},{"key":"e_1_3_3_2_42_2","doi-asserted-by":"crossref","unstructured":"Ludan Ruan and Qin Jin. 2022. Survey: Transformer based video-language pre-training. AI Open 3 (2022) 1\u201313.","DOI":"10.1016\/j.aiopen.2022.01.001"},{"key":"e_1_3_3_2_43_2","doi-asserted-by":"crossref","unstructured":"Ao Sun Weilin Zhao Xu Han Cheng Yang Xinrong Zhang Zhiyuan Liu Chuan Shi and Maosong Sun. 2024. Seq1f1b: Efficient sequence-level pipeline parallelism for large language model training. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2406.03488 (2024).","DOI":"10.18653\/v1\/2025.naacl-long.454"},{"key":"e_1_3_3_2_44_2","unstructured":"Rohan Taori Ishaan Gulrajani Tianyi Zhang Yann Dubois Xuechen Li Carlos Guestrin Percy Liang and Tatsunori\u00a0B Hashimoto. 2023. Alpaca: A strong replicable instruction-following model. Stanford Center for Research on Foundation Models. https:\/\/crfm.stanford.edu\/2023\/03\/13\/alpaca.html 3 6 (2023) 7."},{"key":"e_1_3_3_2_45_2","series-title":"(NeurIPS \u201921)","volume-title":"Advances in Neural Information Processing Systems","author":"Tarnawski Jakub\u00a0M.","year":"2021","unstructured":"Jakub\u00a0M. Tarnawski, Deepak Narayanan, and Amar Phanishayee. 2021-12-06. Piper: Multidimensional Planner for DNN Parallelization. In Advances in Neural Information Processing Systems(NeurIPS \u201921)."},{"key":"e_1_3_3_2_46_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et\u00a0al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2307.09288 (2023)."},{"key":"e_1_3_3_2_47_2","unstructured":"Patrick von Platen. 2023. Optimizing your LLM in production. https:\/\/huggingface.co\/blog\/optimize-llm."},{"key":"e_1_3_3_2_48_2","unstructured":"Fuzhao Xue Yukang Chen Dacheng Li Qinghao Hu Ligeng Zhu Xiuyu Li Yunhao Fang Haotian Tang Shang Yang Zhijian Liu et\u00a0al. 2024. Longvila: Scaling long-context visual language models for long videos. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2408.10188 (2024)."},{"key":"e_1_3_3_2_49_2","unstructured":"Jinghan Yao Sam\u00a0Ade Jacobs Masahiro Tanaka Olatunji Ruwase Aamir Shafi Hari Subramoni and Dhabaleswar\u00a0K Panda. 2024. Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2408.16978 (2024)."},{"key":"e_1_3_3_2_50_2","first-page":"545","volume-title":"2024 USENIX Annual Technical Conference (USENIX ATC 24)","author":"Yuan Tailing","year":"2024","unstructured":"Tailing Yuan, Yuliang Liu, Xucheng Ye, Shenglong Zhang, Jianchao Tan, Bin Chen, Chengru Song, and Di Zhang. 2024. Accelerating the training of large language models using efficient activation rematerialization and optimal hybrid parallelism. In 2024 USENIX Annual Technical Conference (USENIX ATC 24). 545\u2013561."},{"key":"e_1_3_3_2_51_2","unstructured":"Pinxue Zhao Hailin Zhang Fangcheng Fu Xiaonan Nie Qibin Liu Fang Yang Yuanbo Peng Dian Jiao Shuaipeng Li Jinbao Xue et\u00a0al. 2024. Efficiently Training 7B LLM with 1 Million Sequence Length on 8 GPUs. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2407.12117 (2024)."},{"key":"e_1_3_3_2_52_2","unstructured":"Zangwei Zheng Xiangyu Peng Tianji Yang Chenhui Shen Shenggui Li Hongxin Liu Yukun Zhou Tianyi Li and Yang You. 2024. Open-sora: Democratizing efficient video production for all. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2412.20404 (2024)."}],"event":{"name":"ICS '26: 2026 International Conference on Supercomputing","location":"Belfast United Kingdom","acronym":"ICS '26","sponsor":["SIGHPC ACM Special Interest Group on High Performance Computing, Special Interest Group on High Performance Computing","SIGARCH ACM Special Interest Group on Computer Architecture"]},"container-title":["Proceedings of the 40th ACM International Conference on Supercomputing"],"original-title":[],"deposited":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T12:41:38Z","timestamp":1782996098000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3797905.3800537"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,5]]},"references-count":51,"alternative-id":["10.1145\/3797905.3800537","10.1145\/3797905"],"URL":"https:\/\/doi.org\/10.1145\/3797905.3800537","relation":{},"subject":[],"published":{"date-parts":[[2026,7,5]]},"assertion":[{"value":"2026-07-05","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}