{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,21]],"date-time":"2026-03-21T19:14:14Z","timestamp":1774120454435,"version":"3.50.1"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"2","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62025208 and 62421002"],"award-info":[{"award-number":["62025208 and 62421002"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2025,6,30]]},"abstract":"<jats:p>\n            Pipeline parallelism is a crucial technique for large-scale model training, enabling parameter splitting and performance enhancement. However, creating effective pipeline schedules often requires significant manual effort and coding skills, leading to practical inconveniences and complex debugging. Major frameworks such as DeepSpeed and ColossalAI simplify the process by adopting predefined pipeline schedule strategies, such as GPipe and 1F1B. The use of predefined schedules offers limited flexibility and suboptimal training efficiency, as the limited number of manually set candidates cannot provide the optimal strategy for arbitrary model training. To deal with the issue, this article aims to automatically search for the optimal strategy with high efficiency. Since current frameworks only support a limited set of fixed strategies, lacking the technical capability to create a comprehensive strategy search space, we first design a novel domain-specific language (DSL) for pipeline schedule development. The DSL exhibits great understandability, agility, and reusability, supporting the development of all known pipeline schedule strategies and their variants. Second, we are the first to model the complete pipeline schedule strategy space via the DSL, enabling an automated end-to-end globally optimal pipeline schedule searching, while past work may get stuck in a local optimum. Finally, we propose to optimize pipeline performance by modeling and solving the pipeline schedule as a Binary-Tree-Traversing (BTT) optimization problem. Based on the formalization, we further adopt a Dynamic Try-Test Genetic Algorithm\u00a0to search for the best pipeline schedule strategy, which overwhelms a variety of pre-defined ones. Experimental results show that Koala achieves an enhanced performance by up to\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(1.53\\times\\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            over state-of-the-art approaches. Besides, the pipeline schedule strategy searched by Koala outperforms pre-defined pipeline schedule strategies by\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(1.10\\times \\sim 1.55\\times\\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            . Moreover, Koala has superior scalability and effectiveness in combining with data parallelism and tensor parallelism.\n          <\/jats:p>","DOI":"10.1145\/3722113","type":"journal-article","created":{"date-parts":[[2025,3,7]],"date-time":"2025-03-07T11:14:16Z","timestamp":1741346056000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Koala: Efficient Pipeline Training through Automated Schedule Searching on Domain-Specific Language"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8595-1547","authenticated-orcid":false,"given":"Yu","family":"Tang","sequence":"first","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2916-0756","authenticated-orcid":false,"given":"Lujia","family":"Yin","sequence":"additional","affiliation":[{"name":"College of Systems Engineering, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4579-4268","authenticated-orcid":false,"given":"Qiao","family":"Li","sequence":"additional","affiliation":[{"name":"Xiamen University","place":["Xiamen, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-4274-6659","authenticated-orcid":false,"given":"Hongyu","family":"Zhu","sequence":"additional","affiliation":[{"name":"Tsinghua University","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1986-1799","authenticated-orcid":false,"given":"Hengjie","family":"Li","sequence":"additional","affiliation":[{"name":"Shanghai Artificial Intelligence Laboratory","place":["Shanghai, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8525-0608","authenticated-orcid":false,"given":"Xingcheng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Shanghai Artificial Intelligence Laboratory","place":["Shanghai, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8285-2738","authenticated-orcid":false,"given":"Linbo","family":"Qiao","sequence":"additional","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9743-2034","authenticated-orcid":false,"given":"Dongsheng","family":"Li","sequence":"additional","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8963-2535","authenticated-orcid":false,"given":"Jiaxin","family":"Li","sequence":"additional","affiliation":[{"name":"National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,7]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"265","volume-title":"Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201916)","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et\u00a0al. 2016. TensorFlow: A system for large-scale machine learning. In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201916). 265\u2013283."},{"key":"e_1_3_1_3_2","first-page":"1118","volume-title":"Proceedings of the IEEE 38th International Conference on Distributed Computing Systems (ICDCS\u201918)","author":"Ahn Shinyoung","year":"2018","unstructured":"Shinyoung Ahn, Joongheon Kim, Eunji Lim, Wan Choi, Aziz Mohaisen, and Sungwon Kang. 2018. ShmCaffe: A distributed deep learning platform with shared memory buffer for HPC architecture. In Proceedings of the IEEE 38th International Conference on Distributed Computing Systems (ICDCS\u201918). IEEE, 1118\u20131128."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1214\/ss\/1177011077"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/2656877.2656890"},{"key":"e_1_3_1_6_2","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et\u00a0al. 2020. Language models are few-shot learners. Advan. Neural Inf. Process. Syst. 33 (2020), 1877\u20131901.","journal-title":"Advan. Neural Inf. Process. Syst."},{"key":"e_1_3_1_7_2","article-title":"MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems","author":"Chen Tianqi","year":"2015","unstructured":"Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. 2015. MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems. arXiv preprint arXiv:1512.01274 (2015).","journal-title":"arXiv preprint arXiv:1512.01274"},{"key":"e_1_3_1_8_2","article-title":"Training deep nets with sublinear memory cost","author":"Chen Tianqi","year":"2016","unstructured":"Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174 (2016).","journal-title":"arXiv preprint arXiv:1604.06174"},{"key":"e_1_3_1_9_2","first-page":"571","volume-title":"Proceedings of the 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201914)","author":"Chilimbi Trishul","year":"2014","unstructured":"Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman. 2014. Project Adam: Building an efficient and scalable deep learning training system. In Proceedings of the 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201914). 571\u2013582."},{"key":"e_1_3_1_10_2","first-page":"1","volume-title":"Proceedings of the IEEE Hot Chips 32 Symposium (HCS\u201920)","author":"Choquette Jack","year":"2020","unstructured":"Jack Choquette and Wish Gandhi. 2020. Nvidia a100 GPU: Performance & innovation for GPU computing. In Proceedings of the IEEE Hot Chips 32 Symposium (HCS\u201920). IEEE Computer Society, 1\u201343."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3629523"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/2967938.2967969"},{"key":"e_1_3_1_13_2","first-page":"1","volume-title":"Recent Advances in Parallel Virtual Machine and Message Passing Interface: 8th European PVM\/MPI Users\u2019 Group Meeting Santorini\/Thera, Greece, September 23\u201326, 2001 Proceedings 8","author":"Darema Frederica","year":"2001","unstructured":"Frederica Darema. 2001. The SPMD model: Past, present and future. In Recent Advances in Parallel Virtual Machine and Message Passing Interface: 8th European PVM\/MPI Users\u2019 Group Meeting Santorini\/Thera, Greece, September 23\u201326, 2001 Proceedings 8. Springer, 1\u20131."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.5555\/2886366"},{"key":"e_1_3_1_15_2","volume-title":"Composition and Interoperability for External Domain-specific Language Engineering","author":"Degueule Thomas","year":"2016","unstructured":"Thomas Degueule. 2016. Composition and Interoperability for External Domain-specific Language Engineering. Ph.D. Dissertation. Universit\u00e9 de Rennes 1."},{"key":"e_1_3_1_16_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201920)","author":"Dosovitskiy Alexey","year":"2020","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et\u00a0al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR\u201920)."},{"key":"e_1_3_1_17_2","volume-title":"Domain-driven Design: Tackling Complexity in the Heart of Software","author":"Evans Eric","year":"2004","unstructured":"Eric Evans. 2004. Domain-driven Design: Tackling Complexity in the Heart of Software. Addison-Wesley Professional."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3437801.3441593"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.5555\/3586589.3586709"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.5555\/1809745"},{"key":"e_1_3_1_21_2","first-page":"1","volume-title":"Proceedings of the 18th Conference on Pattern Languages of Programs","author":"G\u00fcnther Sebastian","year":"2011","unstructured":"Sebastian G\u00fcnther. 2011. Development of internal domain-specific languages: Design principles and design patterns. In Proceedings of the 18th Conference on Pattern Languages of Programs. 1\u201325."},{"key":"e_1_3_1_22_2","article-title":"GPipe: Efficient training of giant neural networks using pipeline parallelism","volume":"32","author":"Huang Yanping","year":"2019","unstructured":"Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, et\u00a0al. 2019. GPipe: Efficient training of giant neural networks using pipeline parallelism. Advan. Neural Inf. Process. Syst. 32 (2019), 103\u2013112.","journal-title":"Advan. Neural Inf. Process. Syst."},{"key":"e_1_3_1_23_2","volume-title":"GPU Technology Conference (GTC\u201917)","volume":"2","author":"Jeaugey Sylvain","year":"2017","unstructured":"Sylvain Jeaugey. 2017. NCCL 2.0. In GPU Technology Conference (GTC\u201917), Vol. 2."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654889"},{"key":"e_1_3_1_25_2","first-page":"1","article-title":"Beyond data and model parallelism for deep neural networks.","volume":"1","author":"Jia Zhihao","year":"2019","unstructured":"Zhihao Jia, Matei Zaharia, and Alex Aiken. 2019. Beyond data and model parallelism for deep neural networks. Proc. Mach. Learn. Syst. 1 (2019), 1\u201313.","journal-title":"Proc. Mach. Learn. Syst."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2023.3247001"},{"key":"e_1_3_1_27_2","article-title":"Breadth-first pipeline parallelism","volume":"5","author":"Lamy-Poirier Joel","year":"2023","unstructured":"Joel Lamy-Poirier. 2023. Breadth-first pipeline parallelism. Proc. Mach. Learn. Syst. 5 (2023), 48\u201367.","journal-title":"Proc. Mach. Learn. Syst."},{"key":"e_1_3_1_28_2","article-title":"On model parallelization and scheduling strategies for distributed machine learning","volume":"27","author":"Lee Seunghak","year":"2014","unstructured":"Seunghak Lee, Jin Kyu Kim, Xun Zheng, Qirong Ho, Garth A. Gibson, and Eric P. Xing. 2014. On model parallelization and scheduling strategies for distributed machine learning. Advan. Neural Inf. Process. Syst. 27 (2014), 2834\u20132842.","journal-title":"Advan. Neural Inf. Process. Syst."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2928289"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/2640087.2644155"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476145"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3605573.3605613"},{"issue":"12","key":"e_1_3_1_33_2","article-title":"PyTorch distributed: Experiences on accelerating data parallel training","volume":"13","author":"Li Shen","unstructured":"Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, et\u00a0al. 2020. PyTorch distributed: Experiences on accelerating data parallel training. Proce. VLDB Endow. 13, 12 (2020), 3005\u20133018.","journal-title":"Proce. VLDB Endow."},{"key":"e_1_3_1_34_2","article-title":"A survey on auto-parallelism of large-scale deep learning training","author":"Liang Peng","year":"2023","unstructured":"Peng Liang, Yu Tang, Xiaoda Zhang, Youhui Bai, Teng Su, Zhiquan Lai, Linbo Qiao, and Dongsheng Li. 2023. A survey on auto-parallelism of large-scale deep learning training. IEEE Trans. Parallel Distrib. Syst. 34, 8 (2023), 3005\u20133018.","journal-title":"IEEE Trans. Parallel Distrib. Syst."},{"key":"e_1_3_1_35_2","article-title":"Tessel: Boosting distributed execution of large DNN models via flexible schedule search","author":"Lin Zhiqi","year":"2023","unstructured":"Zhiqi Lin, Youshan Miao, Guanbin Xu, Cheng Li, Olli Saarikivi, Saeed Maleki, and Fan Yang. 2023. Tessel: Boosting distributed execution of large DNN models via flexible schedule search. arXiv preprint arXiv:2311.15269 (2023).","journal-title":"arXiv preprint arXiv:2311.15269"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581784.3607073"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/1118890.1118892"},{"issue":"3","key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"470","DOI":"10.14778\/3570690.3570697","article-title":"Galvatron: Efficient transformer training over multiple GPUs using automatic parallelism","volume":"16","author":"Miao Xupeng","year":"2022","unstructured":"Xupeng Miao, Yujie Wang, Youhe Jiang, Chunan Shi, Xiaonan Nie, Hailin Zhang, and Bin Cui. 2022. Galvatron: Efficient transformer training over multiple GPUs using automatic parallelism. Proc. VLDB Endow. 16, 3 (2022), 470\u2013479.","journal-title":"Proc. VLDB Endow."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/2786763.2694364"},{"key":"e_1_3_1_40_2","first-page":"1","volume-title":"Proceedings of the 27th ACM Symposium on Operating Systems Principles (SOSP\u201919)","author":"Narayanan Deepak","year":"2019","unstructured":"Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R. Devanur, Gregory R. Ganger, Phillip B. Gibbons, and Matei Zaharia. 2019. PipeDream: Generalized pipeline parallelism for DNN training. In Proceedings of the 27th ACM Symposium on Operating Systems Principles (SOSP\u201919). 1\u201315."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476209"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-018-09679-z"},{"key":"e_1_3_1_43_2","article-title":"PyTorch: An imperative style, high-performance deep learning library","volume":"32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et\u00a0al. 2019. PyTorch: An imperative style, high-performance deep learning library. Advan. Neural Inf. Process. Syst. 32 (2019), 8024\u20138035.","journal-title":"Advan. Neural Inf. Process. Syst."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3107953"},{"key":"e_1_3_1_45_2","article-title":"Zero bubble pipeline parallelism","author":"Qi Penghui","year":"2023","unstructured":"Penghui Qi, Xinyi Wan, Guangxing Huang, and Min Lin. 2023. Zero bubble pipeline parallelism. arXiv preprint arXiv:2401.10241 (2023).","journal-title":"arXiv preprint arXiv:2401.10241"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11431-020-1647-3"},{"issue":"8","key":"e_1_3_1_47_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"issue":"140","key":"e_1_3_1_48_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 140 (2020), 1\u201367.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/2499370.2462176"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00024"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476205"},{"key":"e_1_3_1_52_2","doi-asserted-by":"crossref","first-page":"3505","DOI":"10.1145\/3394486.3406703","volume-title":"Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD\u201920)","author":"Rasley Jeff","year":"2020","unstructured":"Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD\u201920). 3505\u20133506."},{"key":"e_1_3_1_53_2","article-title":"ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope","author":"Ray Partha Pratim","year":"2023","unstructured":"Partha Pratim Ray. 2023. ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet Things Cyber-phys. Syst. 3 (2023), 1021\u2013154.","journal-title":"Internet Things Cyber-phys. Syst."},{"key":"e_1_3_1_54_2","first-page":"551","volume-title":"Proceedings of the USENIX Annual Technical Conference (USENIX ATC\u201921)","author":"Ren Jie","year":"2021","unstructured":"Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021. ZeRO-Offload: Democratizing billion-scale model training. In Proceedings of the USENIX Annual Technical Conference (USENIX ATC\u201921). 551\u2013564."},{"key":"e_1_3_1_55_2","article-title":"Mesh-TensorFlow: Deep learning for supercomputers","volume":"31","author":"Shazeer Noam","year":"2018","unstructured":"Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, et\u00a0al. 2018. Mesh-TensorFlow: Deep learning for supercomputers. Advan. Neural Inf. Process. Syst. 31 (2018), 10435\u201310444.","journal-title":"Advan. Neural Inf. Process. Syst."},{"key":"e_1_3_1_56_2","article-title":"Megatron-LM: Training multi-billion parameter language models using model parallelism","author":"Shoeybi Mohammad","year":"2019","unstructured":"Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-LM: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053 (2019).","journal-title":"arXiv preprint arXiv:1909.08053"},{"key":"e_1_3_1_57_2","first-page":"1","article-title":"Reducing bias through directed acyclic graphs","volume":"8","author":"Shrier Ian","year":"2008","unstructured":"Ian Shrier and Robert W. Platt. 2008. Reducing bias through directed acyclic graphs. BMC Med. Res. Methodol. 8 (2008), 1\u201315.","journal-title":"BMC Med. Res. Methodol."},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3406117"},{"key":"e_1_3_1_59_2","volume-title":"Programming DSLs in Kotlin","author":"Subramaniam Venkat","year":"2021","unstructured":"Venkat Subramaniam. 2021. Programming DSLs in Kotlin. Pragmatic Bookshelf."},{"key":"e_1_3_1_60_2","article-title":"LLaMA: Open and efficient foundation language models","author":"Touvron Hugo","year":"2023","unstructured":"Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth\u00e9e Lacroix, Baptiste Rozi\u00e8re, Naman Goyal, Eric Hambro, Faisal Azhar, et\u00a0al. 2023. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023).","journal-title":"arXiv preprint arXiv:2302.13971"},{"issue":"6","key":"e_1_3_1_61_2","doi-asserted-by":"crossref","first-page":"26","DOI":"10.1145\/352029.352035","article-title":"Domain-specific languages: An annotated bibliography","volume":"35","author":"Deursen Arie Van","year":"2000","unstructured":"Arie Van Deursen, Paul Klint, and Joost Visser. 2000. Domain-specific languages: An annotated bibliography. ACM SIGPLAN Not. 35, 6 (2000), 26\u201336.","journal-title":"ACM SIGPLAN Not."},{"key":"e_1_3_1_62_2","article-title":"GSPMD: General and scalable parallelization for ML computation graphs","author":"Xu Yuanzhong","year":"2021","unstructured":"Yuanzhong Xu, HyoukJoong Lee, Dehao Chen, Blake Hechtman, Yanping Huang, Rahul Joshi, Maxim Krikun, Dmitry Lepikhin, Andy Ly, Marcello Maggioni, et\u00a0al. 2021. GSPMD: General and scalable parallelization for ML computation graphs. arXiv preprint arXiv:2105.04663 (2021).","journal-title":"arXiv preprint arXiv:2105.04663"},{"key":"e_1_3_1_63_2","first-page":"269","article-title":"PipeMare: Asynchronous pipeline parallel DNN training","volume":"3","author":"Yang Bowen","year":"2021","unstructured":"Bowen Yang, Jian Zhang, Jonathan Li, Christopher R\u00e9, Christopher Aberger, and Christopher De Sa. 2021. PipeMare: Asynchronous pipeline parallel DNN training. Proc. Mach. Learn. Syst. 3 (2021), 269\u2013296.","journal-title":"Proc. Mach. Learn. Syst."},{"key":"e_1_3_1_64_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201921)","author":"Yang PengCheng","year":"2021","unstructured":"PengCheng Yang, Xiaoming Zhang, Wenpeng Zhang, Ming Yang, and Hong Wei. 2021. Group-based interleaved pipeline parallelism for large-scale DNN training. In Proceedings of the International Conference on Learning Representations (ICLR\u201921)."},{"key":"e_1_3_1_65_2","first-page":"559","volume-title":"Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201922)","author":"Zheng Lianmin","year":"2022","unstructured":"Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P. Xing, et\u00a0al. 2022. Alpa: Automating inter-and intra-operator parallelism for distributed deep learning. In Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201922). 559\u2013578."},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123939.3123978"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3722113","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,1]],"date-time":"2025-07-01T12:32:53Z","timestamp":1751373173000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3722113"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,30]]},"references-count":65,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,6,30]]}},"alternative-id":["10.1145\/3722113"],"URL":"https:\/\/doi.org\/10.1145\/3722113","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,30]]},"assertion":[{"value":"2024-09-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-26","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-01","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}