{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,15]],"date-time":"2026-03-15T15:31:09Z","timestamp":1773588669671,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":52,"publisher":"ACM","funder":[{"name":"National Science Foundation","award":["2145471"],"award-info":[{"award-number":["2145471"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,3,22]]},"DOI":"10.1145\/3779212.3790178","type":"proceedings-article","created":{"date-parts":[[2026,3,10]],"date-time":"2026-03-10T13:55:26Z","timestamp":1773150926000},"page":"1022-1039","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["It Takes Two to Entangle"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-3234-6155","authenticated-orcid":false,"given":"Zhanghan","family":"Wang","sequence":"first","affiliation":[{"name":"New York University, New York, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-3785-5058","authenticated-orcid":false,"given":"Ding","family":"Ding","sequence":"additional","affiliation":[{"name":"New York University, New York, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2974-9113","authenticated-orcid":false,"given":"Hang","family":"Zhu","sequence":"additional","affiliation":[{"name":"ByteDance Seed, Bellevue, WA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4879-5335","authenticated-orcid":false,"given":"Haibin","family":"Lin","sequence":"additional","affiliation":[{"name":"ByteDance Seed, Bellevue, WA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9664-4377","authenticated-orcid":false,"given":"Aurojit","family":"Panda","sequence":"additional","affiliation":[{"name":"New York University, New York, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,3,22]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, et al., 2016. TensorFlow: a system for large-scale machine learning. In Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation (Savannah, GA, USA) (OSDI'16). USENIX Association, USA, 265\u2013283."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3620665.3640366"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3704865"},{"key":"e_1_3_2_1_4_1","unstructured":"Clark Barrett Pascal Fontaine and Cesare Tinelli. 2016. The Satisfiability Modulo Theories Library (SMT-LIB). www.SMT-LIB.org."},{"key":"e_1_3_2_1_5_1","volume-title":"Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang.","author":"Bradbury James","year":"2018","unstructured":"James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. 2018. JAX: composable transformations of PythonNumPy programs. http:\/\/github.com\/jax-ml\/jax"},{"key":"e_1_3_2_1_6_1","first-page":"16344","article-title":"FlashAttention: Fast and memory-efficient exact attention with io-awareness","volume":"35","author":"Dao Tri","year":"2022","unstructured":"Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher R\u00e9. 2022. FlashAttention: Fast and memory-efficient exact attention with io-awareness. NeurIPS, Vol. 35 (2022), 16344-16359.","journal-title":"NeurIPS"},{"key":"e_1_3_2_1_7_1","unstructured":"DeepSeek-AI Aixin Liu Bei Feng Bing Xue Bingxuan Wang et al. 2025. DeepSeek-V3 Technical Report. arXiv:2412.19437 [cs.CL] https:\/\/arxiv.org\/abs\/2412.19437"},{"key":"e_1_3_2_1_8_1","unstructured":"Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Amy Yang Angela Fan et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)."},{"key":"e_1_3_2_1_9_1","unstructured":"Dmitry Duplyakin Robert Ricci Aleksander Maricq Gary Wong Jonathon Duerig Eric Eide Leigh Stoller Mike Hibler David Johnson Kirk Webb et al. 2019. The design and operation of CloudLab. In USENIX ATC."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3437801.3441593"},{"key":"e_1_3_2_1_11_1","article-title":"Switch transformers: scaling to trillion parameter models with simple and efficient sparsity","volume":"23","author":"Fedus William","year":"2022","unstructured":"William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch transformers: scaling to trillion parameter models with simple and efficient sparsity. J. Mach. Learn. Res., Vol. 23, 1, Article 120 (Jan. 2022), 39 pages.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_2_1_12_1","unstructured":"Jiaao He Jiezhong Qiu Aohan Zeng Zhilin Yang Jidong Zhai and Jie Tang. 2021. FastMoE: A Fast Mixture-of-Expert Training System. arXiv:2103.13262 [cs.LG] https:\/\/arxiv.org\/abs\/2103.13262"},{"key":"e_1_3_2_1_13_1","volume-title":"Gpipe: Efficient training of giant neural networks using pipeline parallelism. Advances in neural information processing systems","author":"Huang Yanping","year":"2019","unstructured":"Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al., 2019. Gpipe: Efficient training of giant neural networks using pipeline parallelism. Advances in neural information processing systems, Vol. 32 (2019)."},{"key":"e_1_3_2_1_14_1","unstructured":"Bytedance Inc. 2024. https:\/\/volcengine.github.io\/veScaleWeb\/blog\/mlsys2024.html"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359630"},{"key":"e_1_3_2_1_16_1","unstructured":"Haitian Jiang Shaowei Zhu Zhen Zhang Zhenyu Song Xinwei Fu Zhen Jia Yida Wang and Jinyang Li. 2025. TTrace: Lightweight Error Checking and Diagnosis for Distributed Training. arXiv:2506.09280 [cs.DC] https:\/\/arxiv.org\/abs\/2506.09280"},{"key":"e_1_3_2_1_17_1","unstructured":"Vijay Korthikanti Jared Casper Sangkug Lym Lawrence McAfee Michael Andersch Mohammad Shoeybi and Bryan Catanzaro. 2022. Reducing Activation Recomputation in Large Transformer Models. arXiv:2205.05198 [cs.LG] https:\/\/arxiv.org\/abs\/2205.05198"},{"key":"e_1_3_2_1_18_1","volume-title":"Proceedings of Machine Learning and Systems","volume":"5","author":"Korthikanti Vijay Anand","year":"2023","unstructured":"Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. 2023. Reducing activation recomputation in large transformer models. Proceedings of Machine Learning and Systems, Vol. 5 (2023)."},{"key":"e_1_3_2_1_19_1","unstructured":"Dmitry Lepikhin HyoukJoong Lee Yuanzhong Xu Dehao Chen Orhan Firat Yanping Huang Maxim Krikun Noam Shazeer and Zhifeng Chen. 2020. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. arXiv:2006.16668 [cs.CL] https:\/\/arxiv.org\/abs\/2006.16668"},{"key":"e_1_3_2_1_20_1","first-page":"347","volume-title":"18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Lin Zhiqi","year":"2024","unstructured":"Zhiqi Lin, Youshan Miao, Quanlu Zhang, Fan Yang, Yi Zhu, Cheng Li, Saeed Maleki, Xu Cao, Ning Shang, Yilei Yang, Weijiang Xu, Mao Yang, Lintao Zhang, and Lidong Zhou. 2024. nnScaler: Constraint-Guided Parallelization Plan Generation for Deep Learning Training. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 347-363."},{"key":"e_1_3_2_1_21_1","unstructured":"Hao Liu Matei Zaharia and Pieter Abbeel. 2023b. Ring Attention with Blockwise Transformers for Near-Infinite Context. arXiv:2310.01889 [cs.CL] https:\/\/arxiv.org\/abs\/2310.01889"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3575693.3575707"},{"key":"e_1_3_2_1_23_1","unstructured":"Yunchi Lu Youshan Miao Cheng Tan Peng Huang Yi Zhu Xian Zhang and Fan Yang. 2025. TrainVerify: Equivalence-Based Verification for Distributed LLM Training. arXiv:2506.15961 [cs.DC] https:\/\/arxiv.org\/abs\/2506.15961"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00037"},{"key":"e_1_3_2_1_25_1","unstructured":"Benjamin Marie. 2024. https:\/\/github.com\/huggingface\/trl\/issues\/2175"},{"key":"e_1_3_2_1_26_1","unstructured":"Tim Moon and Megatron-LM Team. 2025. https:\/\/github.com\/NVIDIA\/TransformerEngine\/pull\/1528"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476209"},{"key":"e_1_3_2_1_28_1","volume-title":"Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia.","author":"Narayanan Deepak","year":"2021","unstructured":"Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Anand Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. 2021b. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM. arXiv:2104.04473 [cs.CL] https:\/\/arxiv.org\/abs\/2104.04473"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Samyam Rajbhandari Jeff Rasley Olatunji Ruwase and Yuxiong He. 2020. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. arXiv:1910.02054 [cs.LG] https:\/\/arxiv.org\/abs\/1910.02054","DOI":"10.1109\/SC41405.2020.00024"},{"key":"e_1_3_2_1_30_1","unstructured":"RookieHong and Megatron-LM Team. 2023. https:\/\/github.com\/NVIDIA\/Megatron-LM\/issues\/599"},{"key":"e_1_3_2_1_31_1","unstructured":"Noam Shazeer Azalia Mirhoseini Krzysztof Maziarz Andy Davis Quoc Le Geoffrey Hinton and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. arXiv:1701.06538 [cs.LG] https:\/\/arxiv.org\/abs\/1701.06538"},{"key":"e_1_3_2_1_32_1","unstructured":"Mohammad Shoeybi Mostofa Patwary Raul Puri Patrick LeGresley Jared Casper and Bryan Catanzaro. 2020. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv:1909.08053 [cs.CL] https:\/\/arxiv.org\/abs\/1909.08053"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1168857.1168907"},{"key":"e_1_3_2_1_34_1","unstructured":"Jianlin Su Yu Lu Shengfeng Pan Ahmed Murtadha Bo Wen and Yunfeng Liu. 2023. RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv:2104.09864 [cs.CL] https:\/\/arxiv.org\/abs\/2104.09864"},{"key":"e_1_3_2_1_35_1","unstructured":"The Huggingface Team. 2025a. https:\/\/github.com\/huggingface\/transformers\/blob\/main\/tests\/trainer\/test_trainer.py"},{"key":"e_1_3_2_1_36_1","unstructured":"The Megatron-LM Team. 2025b. https:\/\/github.com\/NVIDIA\/Megatron-LM\/blob\/main\/examples\/run_simple_mcore_train_loop.py"},{"key":"e_1_3_2_1_37_1","unstructured":"The NVIDIA Team. 2016. https:\/\/developer.nvidia.com\/tensorrt"},{"key":"e_1_3_2_1_38_1","unstructured":"The NVIDIA Team. 2024. https:\/\/docs.nvidia.com\/megatron-core\/developer-guide\/latest\/api-guide\/context_parallel.html"},{"key":"e_1_3_2_1_39_1","unstructured":"The PyTorch Team. 2023. https:\/\/pytorch.org\/docs\/stable\/torch.compiler_ir.html"},{"key":"e_1_3_2_1_40_1","unstructured":"The PyTorch Team. 2023a. https:\/\/github.com\/pytorch\/pytorch\/issues\/109505"},{"key":"e_1_3_2_1_41_1","unstructured":"The PyTorch Team. 2023b. https:\/\/github.com\/pytorch\/pytorch\/issues\/107861#issuecomment-1696058500"},{"key":"e_1_3_2_1_42_1","unstructured":"trintamaki and Megatron-LM Team. 2024. https:\/\/github.com\/NVIDIA\/Megatron-LM\/commit\/5fffdfc737f14297bc3781dfc9e273199d1df52e#diff-855adbcea94c997a151e12312a282117853f541a11989febe40db2ad12fa38c6"},{"key":"e_1_3_2_1_43_1","first-page":"37","volume-title":"PET: Optimizing Tensor Programs with Partially Equivalent Transformations and Automated Corrections. In 15th USENIX Symposium on Operating Systems Design and Implementation (OSDI 21)","author":"Wang Haojie","year":"2021","unstructured":"Haojie Wang, Jidong Zhai, Mingyu Gao, Zixuan Ma, Shizhi Tang, Liyan Zheng, Yuanzhi Li, Kaiyuan Rong, Yuanyong Chen, and Zhihao Jia. 2021. PET: Optimizing Tensor Programs with Partially Equivalent Transformations and Automated Corrections. In 15th USENIX Symposium on Operating Systems Design and Implementation (OSDI 21). USENIX Association, 37-54. https:\/\/www.usenix.org\/conference\/osdi21\/presentation\/wang"},{"key":"e_1_3_2_1_44_1","first-page":"788","article-title":"Deep learning library testing via effective model generation","author":"Wang Zan","year":"2020","unstructured":"Zan Wang, Ming Yan, Junjie Chen, Shuang Liu, and Dongdi Zhang. 2020a. Deep learning library testing via effective model generation. In ESEC\/SIGSOFT FSE. ACM, 788-799.","journal-title":"ESEC\/SIGSOFT FSE. ACM"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409761"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3434304"},{"key":"e_1_3_2_1_47_1","volume-title":"Mirage: A Multi-Level Superoptimizer for Tensor Programs. arXiv:2405.05751 [cs.LG] https:\/\/arxiv.org\/abs\/2405.05751","author":"Wu Mengdi","year":"2024","unstructured":"Mengdi Wu, Xinhao Cheng, Shengyu Liu, Chunan Shi, Jianan Ji, Kit Ao, Praveen Velliengiri, Xupeng Miao, Oded Padon, and Zhihao Jia. 2024. Mirage: A Multi-Level Superoptimizer for Tensor Programs. arXiv:2405.05751 [cs.LG] https:\/\/arxiv.org\/abs\/2405.05751"},{"key":"e_1_3_2_1_48_1","unstructured":"Zhaofeng Wu. 2021. https:\/\/github.com\/huggingface\/transformers\/issues\/14638"},{"key":"e_1_3_2_1_49_1","volume-title":"Yisu Remy Wang, Max Willsey, Sudip Roy, and Jacques Pienaar.","author":"Yang Yichen","year":"2021","unstructured":"Yichen Yang, Phitchaya Mangpo Phothilimtha, Yisu Remy Wang, Max Willsey, Sudip Roy, and Jacques Pienaar. 2021. Equality Saturation for Tensor Graph Superoptimization. arXiv:2101.01332 [cs.AI] https:\/\/arxiv.org\/abs\/2101.01332"},{"key":"e_1_3_2_1_50_1","volume-title":"Jacky Wai Keung, and Xin Xia","author":"Yu Xiao","year":"2025","unstructured":"Xiao Yu, Haoxuan Chen, Feifei Niu, Xing Hu, Jacky Wai Keung, and Xin Xia. 2025. Towards Understanding Bugs in Distributed Training and Inference Frameworks for Large Language Models. arXiv:2506.10426 [cs.SE] https:\/\/arxiv.org\/abs\/2506.10426"},{"key":"e_1_3_2_1_51_1","volume-title":"Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning. arXiv:2201.12023 [cs.LG] https:\/\/arxiv.org\/abs\/2201.12023","author":"Zheng Lianmin","year":"2022","unstructured":"Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P. Xing, Joseph E. Gonzalez, and Ion Stoica. 2022. Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning. arXiv:2201.12023 [cs.LG] https:\/\/arxiv.org\/abs\/2201.12023"},{"key":"e_1_3_2_1_52_1","volume-title":"Verifying Semantic Equivalence of Large Models with Equality Saturation. In The 5th Workshop on Machine Learning and Systems (EuroMLSys'25)","author":"Zulkifli Kahfi Soobhan","year":"2025","unstructured":"Kahfi Soobhan Zulkifli, Wenbo Qian, Shaowei Zhu, Yuan Zhou, Zhen Zhang, and Chang Lou. 2025. Verifying Semantic Equivalence of Large Models with Equality Saturation. In The 5th Workshop on Machine Learning and Systems (EuroMLSys'25) (Rotterdam, Netherlands). New York, NY, USA."}],"event":{"name":"ASPLOS '26: 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems","location":"Pittsburgh PA USA","sponsor":["SIGOPS ACM Special Interest Group on Operating Systems","SIGPLAN ACM Special Interest Group on Programming Languages","SIGARCH ACM Special Interest Group on Computer Architecture","SIGBED ACM Special Interest Group on Embedded Systems"]},"container-title":["Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2"],"original-title":[],"deposited":{"date-parts":[[2026,3,15]],"date-time":"2026-03-15T14:03:57Z","timestamp":1773583437000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3779212.3790178"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,22]]},"references-count":52,"alternative-id":["10.1145\/3779212.3790178","10.1145\/3779212"],"URL":"https:\/\/doi.org\/10.1145\/3779212.3790178","relation":{},"subject":[],"published":{"date-parts":[[2026,3,22]]},"assertion":[{"value":"2026-03-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}