{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,24]],"date-time":"2026-01-24T22:22:17Z","timestamp":1769293337360,"version":"3.49.0"},"reference-count":37,"publisher":"Wiley","issue":"2","license":[{"start":{"date-parts":[[2026,1,12]],"date-time":"2026-01-12T00:00:00Z","timestamp":1768176000000},"content-version":"vor","delay-in-days":11,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T00:00:00Z","timestamp":1767225600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Concurrency and Computation"],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>As model parameters increase exponentially, distributed training has become essential for advancing modern deep neural networks. Megatron\u2010LM, an efficient distributed training framework developed by NVIDIA, enables the training of trillion\u2010parameter models on thousands of GPUs by integrating tensor, pipeline, and data parallelism. Its computational efficiency has established it as a foundational tool for training large\u2010scale models. Rapid identification of optimal parallel configurations for specific GPU clusters is critical for maximizing computational resource utilization, with    training time prediction serving as a key evaluation metric. The high cost and limited availability of high\u2010performance GPUs, particularly those based on NVIDIA architectures, have made the construction of large\u2010scale heterogeneous clusters a practical solution to resource and cost constraints. However, existing prediction methods do not reliably or efficiently account for the computational and communication complexities inherent in heterogeneous GPU clusters. To address this gap, HATP (Heterogeneous\u2010Aware Time Predictor) is introduced as a novel performance prediction method specifically designed for heterogeneous GPU clusters. For any given parallel configuration, HATP rapidly and accurately simulates execution times to inform the optimization of parallel strategies. To address communication differences among heterogeneous GPUs, comprehensive experimental analyses are conducted and analytical expressions are derived to characterize the communication frequency patterns in Megatron\u2010LM's parallel strategies. This work presents the first systematic quantification of communication operations within Megatron\u2010LM framework, ensuring that performance predictions remain highly accurate even in complex, heterogeneous environments. Furthermore, to account for computational differences among heterogeneous GPUs, a layer\u2010level computational performance acquisition scheme is proposed to reduce the impact of fine\u2010grained operator overlap and additional memory operations. Experimental results demonstrate that HATP achieves an average prediction accuracy of 97.41% in isomorphic environments, surpassing the current state\u2010of\u2010the\u2010art method, ACEso. HATP also attains an average accuracy of 96.04% in heterogeneous data parallel and pipeline parallel configurations, representing the first extension of training time prediction capabilities to heterogeneous environments.<\/jats:p>","DOI":"10.1002\/cpe.70500","type":"journal-article","created":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T06:13:11Z","timestamp":1768284791000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Communication Frequency in Megatron\u2010LM: Experimental Insights Applied to Heterogeneous Distributed Training Time Prediction"],"prefix":"10.1002","volume":"38","author":[{"given":"HaoRan","family":"Zhang","sequence":"first","affiliation":[{"name":"China Mobile Qilu Innovation Research Institute  Shandong China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-5970-0638","authenticated-orcid":false,"given":"Yanzhao","family":"Feng","sequence":"additional","affiliation":[{"name":"China Mobile Qilu Innovation Research Institute  Shandong China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhengwei","family":"Chen","sequence":"additional","affiliation":[{"name":"China Mobile Qilu Innovation Research Institute  Shandong China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yutong","family":"Tian","sequence":"additional","affiliation":[{"name":"China Mobile Qilu Innovation Research Institute  Shandong China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoli","family":"Zheng","sequence":"additional","affiliation":[{"name":"China Mobile Qilu Innovation Research Institute  Shandong China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cong","family":"Liu","sequence":"additional","affiliation":[{"name":"China Mobile Qilu Innovation Research Institute  Shandong China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sheng","family":"Wang","sequence":"additional","affiliation":[{"name":"China Mobile Research Institute  Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jie","family":"Ren","sequence":"additional","affiliation":[{"name":"China Mobile Qilu Innovation Research Institute  Shandong China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yucong","family":"Li","sequence":"additional","affiliation":[{"name":"China Mobile Qilu Innovation Research Institute  Shandong China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rui","family":"Zhu","sequence":"additional","affiliation":[{"name":"China Mobile Qilu Innovation Research Institute  Shandong China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,1,12]]},"reference":[{"key":"e_1_2_10_2_1","unstructured":"J.Achiam S.Adler S.Agarwal et al. \u201cGpt\u20104 Technical Report \u201d arXiv preprint arXiv:2303.08774 (2023)."},{"key":"e_1_2_10_3_1","unstructured":"A.Yang B.Yang B.Zhang et al. \u201cQwen2. 5 Technical Report \u201d arXiv preprint arXiv:2412.15115 (2024)."},{"key":"e_1_2_10_4_1","unstructured":"Meta AI \u201cLLaMA 4 Behemoth: Multimodal Intelligence \u201d2025 https:\/\/ai.meta.com\/blog\/llama\u20104\u2010multimodal\u2010intelligence."},{"key":"e_1_2_10_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3406703"},{"key":"e_1_2_10_6_1","unstructured":"M.Shoeybi M.Patwary R.Puri P.LeGresley J.Casper andB.Catanzaro \u201cMegatron\u2010Lm: Training Multi\u2010Billion Parameter Language Models Using Model Parallelism \u201d arXiv preprint arXiv:1909.08053 (2019)."},{"key":"e_1_2_10_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476209"},{"key":"e_1_2_10_8_1","first-page":"341","article-title":"Reducing Activation Recomputation in Large Transformer Models","volume":"5","author":"Korthikanti V. A.","year":"2023","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_2_10_9_1","doi-asserted-by":"crossref","unstructured":"Z.Mo J.Liao H.Xu Z.Zhou andC.Xu \u201cHetis: Serving LLMs in Heterogeneous GPU Clusters With Fine\u2010Grained and Dynamic Parallelism \u201d arXiv preprint arXiv:2509.08309 (2025).","DOI":"10.1145\/3712285.3759784"},{"key":"e_1_2_10_10_1","unstructured":"K.Zhang H.Liao andG.Tang \u201cLiteGD: Lightweight and Dynamic GPU Dispatching for Large\u2010Scale Heterogeneous Clusters \u201d arXiv preprint arXiv:2506.15595 (2025)."},{"key":"e_1_2_10_11_1","first-page":"163","volume-title":"Aceso: Efficient Parallel Dnn Training Through Iterative Bottleneck Alleviation","author":"Liu G.","year":"2024"},{"key":"e_1_2_10_12_1","unstructured":"C.Zhou \u201cFlagScale \u201d2024."},{"key":"e_1_2_10_13_1","unstructured":"S.Xu Z.Huang Y.Zeng et al. \u201cHETHUB: A Distributed Training System With Heterogeneous Cluster for Large\u2010Scale Models \u201d arXiv preprint arXiv:2405.16256 (2024)."},{"key":"e_1_2_10_14_1","unstructured":"A.SergeevandM.Del Balso \u201cHorovod: Fast and Easy Distributed Deep Learning in TensorFlow \u201d arXiv preprint arXiv:1802.05799 (2018)."},{"key":"e_1_2_10_15_1","first-page":"1","volume-title":"Zero: Memory Optimizations Toward Training Trillion Parameter Models","author":"Rajbhandari S.","year":"2020"},{"key":"e_1_2_10_16_1","first-page":"2142","volume-title":"OSDP: Optimal Sharded Data Parallel for Distributed Deep Learning","author":"Jiang Y.","year":"2023"},{"key":"e_1_2_10_17_1","first-page":"103","article-title":"Gpipe: Efficient Training of Giant Neural Networks Using Pipeline Parallelism","volume":"32","author":"Huang Y.","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_10_18_1","doi-asserted-by":"crossref","unstructured":"A.Harlap D.Narayanan A.Phanishayee et al. \u201cPipedream: Fast and Efficient Pipeline Parallel Dnn Training \u201d arXiv preprint arXiv:1806.03377 (2018).","DOI":"10.1145\/3341301.3359646"},{"key":"e_1_2_10_19_1","first-page":"7937","volume-title":"Memory\u2010Efficient Pipeline\u2010Parallel Dnn Training","author":"Narayanan D.","year":"2021"},{"key":"e_1_2_10_20_1","first-page":"342","volume-title":"Accpar: Tensor Partitioning for Heterogeneous Deep Learning Accelerators","author":"Song L.","year":"2020"},{"key":"e_1_2_10_21_1","unstructured":"X.Jia L.Jiang A.Wang et al. \u201cWhale: Efficient Giant Model Training Over Heterogeneous GPUs \u201d arXiv preprint arXiv:2011.09208 (2020)."},{"key":"e_1_2_10_22_1","unstructured":"Y.Liu Q.Xu andY. C.Hu \u201cCronus: Efficient LLM Inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill \u201d arXiv preprint arXiv:2509.17357 (2025)."},{"key":"e_1_2_10_23_1","unstructured":"H.Touvron L.Martin K.Stone et al. \u201cLlama 2: Open Foundation and Fine\u2010Tuned Chat Models \u201d arXiv preprint arXiv:2307.09288 (2023)."},{"key":"e_1_2_10_24_1","first-page":"1737","volume-title":"Deep Learning With Limited Numerical Precision","author":"Gupta S.","year":"2015"},{"key":"e_1_2_10_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2008.09.002"},{"issue":"8","key":"e_1_2_10_26_1","first-page":"9","article-title":"Language Models Are Unsupervised Multitask Learners","volume":"1","author":"Radford A.","year":"2019","journal-title":"Open AI Blog"},{"key":"e_1_2_10_27_1","unstructured":"J.Huang S.Di X.Yu et al. \u201cZCCL: Significantly Improving Collective Communication With Error\u2010Bounded Lossy Compression \u201d arXiv preprint arXiv:2502.18554 (2025)."},{"key":"e_1_2_10_28_1","unstructured":"China Mobile Qilu Innovation Institute \u201cHeterogeneous Distributed Training Framework (Core_0.8.0) \u201d2025."},{"key":"e_1_2_10_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.2017.2756959"},{"key":"e_1_2_10_30_1","unstructured":"R.AbbasiandS.Lim \u201cSuperpipeline: A Universal Approach for Reducing GPU Memory Usage in Large Models \u201d arXiv preprint arXiv:2410.08791 (2024)."},{"key":"e_1_2_10_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2942129"},{"key":"e_1_2_10_32_1","unstructured":"J.Lin W.Wang L.Yin andY.Han \u201cKAITIAN: A Unified Communication Framework for Enabling Efficient Collaboration Across Heterogeneous Accelerators in Embodied AI Systems \u201d arXiv preprint arXiv:2505.10183 (2025)."},{"key":"e_1_2_10_33_1","article-title":"Attention Is All You Need","volume":"30","author":"Vaswani A.","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"140","key":"e_1_2_10_34_1","first-page":"1","article-title":"Exploring the Limits of Transfer Learning With a Unified Text\u2010To\u2010Text Transformer","volume":"21","author":"Raffel C.","year":"2020","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_10_35_1","first-page":"6223","volume-title":"AdaMoE: Token\u2010Adaptive Routing With Null Experts for Mixture\u2010Of\u2010Experts Language Models","author":"Zeng Z.","year":"2024"},{"key":"e_1_2_10_36_1","unstructured":"Z.Doucet R.Sharma d M.Vos R.Pires A. M.Kermarrec andO.Balmau \u201cHarMoEny: Efficient Multi\u2010GPU Inference of MoE Models \u201d arXiv preprint arXiv:2506.12417 (2025)."},{"key":"e_1_2_10_37_1","unstructured":"A. Q.Jiang A.Sablayrolles A.Roux et al. \u201cMixtral of Experts \u201d arXiv preprint arXiv:2401.04088 (2024)."},{"key":"e_1_2_10_38_1","first-page":"27730","article-title":"Training Language Models to Follow Instructions With Human Feedback","volume":"35","author":"Ouyang L.","year":"2022","journal-title":"Advances in Neural Information Processing Systems"}],"container-title":["Concurrency and Computation: Practice and Experience"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.70500","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1002\/cpe.70500","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.70500","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,23]],"date-time":"2026-01-23T12:29:59Z","timestamp":1769171399000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/cpe.70500"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1]]},"references-count":37,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["10.1002\/cpe.70500"],"URL":"https:\/\/doi.org\/10.1002\/cpe.70500","archive":["Portico"],"relation":{},"ISSN":["1532-0626","1532-0634"],"issn-type":[{"value":"1532-0626","type":"print"},{"value":"1532-0634","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1]]},"assertion":[{"value":"2025-08-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70500"}}