{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,12]],"date-time":"2026-01-12T07:48:06Z","timestamp":1768204086304,"version":"3.49.0"},"reference-count":40,"publisher":"Wiley","issue":"1","license":[{"start":{"date-parts":[[2025,12,2]],"date-time":"2025-12-02T00:00:00Z","timestamp":1764633600000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2022YFC3005702"],"award-info":[{"award-number":["2022YFC3005702"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Concurrency and Computation"],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>Large\u2010scale model training in distributed data centers plays a crucial role in deep learning. Still, it faces significant challenges, including resource fragmentation, low bandwidth utilization, and complex task flow management. The problem is exacerbated by high\u2010speed, high\u2010capacity parameter synchronization, often exceeding several hundred Gbps, which leads to reduced throughput and computational inefficiencies. To address these challenges, this paper proposes an innovative approach that combines data parallelism, model parallelism, and dynamic load balancing. By integrating a Gray Markov Chain (GMC) and Markov Decision Process (MDP) model, the approach dynamically schedules resources and balances computational loads. The GMC model is used to predict future node loads, facilitating optimal weight matrix decomposition, while the MDP model adjusts data transmission paths to optimize network traffic management. The combination of these two models enhances both resource allocation and data flow optimization. Experimental results demonstrate that this integrated approach significantly improves throughput, resource utilization, and computational efficiency compared to traditional methods. The findings suggest that this hybrid approach performs exceptionally well in optimizing large\u2010scale distributed training tasks in multidata\u2010center environments, significantly improving the scalability and performance of deep learning workloads. This research shows promising implications for enhancing the efficiency and effectiveness of distributed training systems in high\u2010demand applications.<\/jats:p>","DOI":"10.1002\/cpe.70456","type":"journal-article","created":{"date-parts":[[2025,12,3]],"date-time":"2025-12-03T04:05:06Z","timestamp":1764734706000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Dynamic Load Balancing for Distributed Large Model Training: A Hybrid Framework of Gray Markov Chain and\n                    <scp>MDP<\/scp>"],"prefix":"10.1002","volume":"38","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-1836-2450","authenticated-orcid":false,"given":"Yonggang","family":"Li","sequence":"first","affiliation":[{"name":"School of Communications and Information Engineering Chongqing University of Posts and Telecommunications  Chongqing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rui","family":"Ji","sequence":"additional","affiliation":[{"name":"School of Communications and Information Engineering Chongqing University of Posts and Telecommunications  Chongqing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yaotong","family":"Su","sequence":"additional","affiliation":[{"name":"School of Communications and Information Engineering Chongqing University of Posts and Telecommunications  Chongqing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuanjin","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Communications and Information Engineering Chongqing University of Posts and Telecommunications  Chongqing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andong","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Communications and Information Engineering Chongqing University of Posts and Telecommunications  Chongqing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Longjiang","family":"Li","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering University of Electronic Science and Technology  Chengdu China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2025,12,2]]},"reference":[{"key":"e_1_2_11_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11280-024-01291-2"},{"key":"e_1_2_11_3_1","doi-asserted-by":"crossref","unstructured":"J.Devlin M.\u2010W.Chang K.Lee andK.Toutanova \u201cBERT: Pre\u2010Training of Deep Bidirectional Transformers for Language Understanding \u201d2019In Proceedings of NAACL\u2010HLT 2019 4171\u20134186 ACL.","DOI":"10.18653\/v1\/N19-1423"},{"issue":"8","key":"e_1_2_11_4_1","first-page":"9","article-title":"Language Models Are Unsupervised Multitask Learners","volume":"1","author":"Radford A.","year":"2019","journal-title":"QRRQA Blog"},{"key":"e_1_2_11_5_1","doi-asserted-by":"crossref","unstructured":"J. H.LuoandJ.Wu \u201cNeural Network Pruning With Residual\u2010Connections and Limited\u2010Data \u201d2020In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Seattle WA USA: IEEE 1458\u20131467.","DOI":"10.1109\/CVPR42600.2020.00153"},{"key":"e_1_2_11_6_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pdig.0000417"},{"key":"e_1_2_11_7_1","unstructured":"D. U. M. I. T. R. U.Erhan \u201cWhy Does Unsupervised Pre\u2010Training Help Deep Learning? \u201d2010Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics PMLR 9 201\u2013208."},{"key":"e_1_2_11_8_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6428"},{"key":"e_1_2_11_9_1","first-page":"1877","article-title":"Language Models Are Few\u2010Shot Learners","volume":"33","author":"Brown T.","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_11_10_1","unstructured":"O.Shahid S.Pouriyeh R. M.Parizi et al. \u201cCommunication Efficiency in Federated Learning: Achievements and Challenges \u201darXiv 2021 2107.10996."},{"key":"e_1_2_11_11_1","unstructured":"D.Driess F.Xia M. S. M.Sajjadi et al. \u201cPalm\u2010e: An Embodied Multimodal Language Model \u201d2023."},{"key":"e_1_2_11_12_1","doi-asserted-by":"publisher","DOI":"10.1364\/JOCN.440845"},{"key":"e_1_2_11_13_1","unstructured":"M.Shoeybi M.Patwary R.Puri et al. \u201cMegatron\u2010lm:Training Multi\u2010Billion Parameter Language Models Using Model Parallelism \u201darXiv 2019 1909.08053."},{"issue":"10","key":"e_1_2_11_14_1","first-page":"103","article-title":"Gpipe: Efficienttraining of Giant Neural Networks Using Pipeline Parallelism","volume":"32","author":"Huang Y.","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_11_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3234761"},{"key":"e_1_2_11_16_1","doi-asserted-by":"crossref","unstructured":"T. S.KhanandH.Nazariopuya \u201cAnalyzing the Implementation of the Newton\u2010Raphson Based Power Flow Formulation in CPU+GPU Computing Environment \u201d2023In 2023 North American Power Symposium (NAPS) IEEE 1\u20135.","DOI":"10.1109\/NAPS58826.2023.10318707"},{"key":"e_1_2_11_17_1","doi-asserted-by":"crossref","unstructured":"J.\u2010F.Dollinger K.Bouhouch andI.Bouzarkouna \u201cBenchmarking OpenStack for Edge Computing Applications \u201d2023In: 2023 IEEE\/ACIS 8th International Conference on Big Data Cloud Computing and Data Science (BCD) Ho Chi Minh City Vietnam IEEE 295\u2013302. 10.1109\/BCD57833.2023.10466323.","DOI":"10.1109\/BCD57833.2023.10466323"},{"key":"e_1_2_11_18_1","doi-asserted-by":"crossref","unstructured":"Y.Shi I. J.Lingareddy K.Suo andT. N.Nguyen \u201cOptimizing Resource Allocation in Cloud \u201d2022IEEE 13th Annual Information Technology Electronics and Mobile Communication Conference (IEMCON) Vancouver BC Canada 2022 0465\u20130470 doi: 10.1109\/IEMCON56893.2022.9946472.","DOI":"10.1109\/IEMCON56893.2022.9946472"},{"key":"e_1_2_11_19_1","doi-asserted-by":"crossref","unstructured":"D.Direm Z.Maamar andA.Benna \u201cDEMO of: A Policy\u2010Based System for Service Provisioning in Edge Computing \u201d20247th Conference on Cloud and Internet of Things (CIoT) Montreal QC Canada 2024 1\u20132 Doi: 10.1109\/CIoT63799.2024.10756976.","DOI":"10.1109\/CIoT63799.2024.10756976"},{"key":"e_1_2_11_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM52122.2024.10621164"},{"key":"e_1_2_11_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCC.2022.3211439"},{"key":"e_1_2_11_22_1","doi-asserted-by":"crossref","unstructured":"N.Su C.Hu andB.Li \u201cTITANIC: Towards Production Federated Learning With Large Language Models[C] \u201d2024\/\/IEEEINFOCOM.","DOI":"10.1109\/INFOCOM52122.2024.10621164"},{"key":"e_1_2_11_23_1","unstructured":"J.Zhu S.Li andY.You \u201cSky Computing: Accelerating Geo\u2010Distributed Computing in Federated Learning \u201darXiv 2022 2202.11836."},{"key":"e_1_2_11_24_1","unstructured":"C.He S.Li J.So et al. \u201cFedml: A Research Library and Benchmark for Federated Machine Learning \u201darXiv 2020 2007.13518."},{"key":"e_1_2_11_25_1","doi-asserted-by":"crossref","unstructured":"D.Narayanan M.Shoeybi J.Casper et al. \u201cEfficient Large\u2010Scale Language Model Training on GPU Clusters Using Megatron\u2010lm \u201d2021Proceedings of the International Conference for High Performance Computing Networking Storage and Analysis 1\u201315.","DOI":"10.1145\/3458817.3476209"},{"key":"e_1_2_11_26_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415530"},{"key":"e_1_2_11_27_1","unstructured":"Z.Jia M.Zaharia andA.Aiken \u201cBeyond Data and Model Parallelism for Deep Neural Networks \u201d2019Proceedings of Machine Learning and Systems 1 1\u201313."},{"key":"e_1_2_11_28_1","unstructured":"A.Harlap D.Narayanan A.Phanishayee et al. \u201cPipedream: Fast and Efficient Pipeline Parallel DNN Training \u201darXiv 2018 1806.03377."},{"key":"e_1_2_11_29_1","unstructured":"K.Nagrecha \u201cSystems for Parallel and Distributed Large\u2010Model Deep Learning Training \u201darXiv 2023 2301.02691."},{"key":"e_1_2_11_30_1","unstructured":"B.Liu \u201cTLPipe: An Efficient Two\u2010Level Optimization Pipeline Model Parallelism for Large Model Training \u201d20252025 6th International Conference on Computer Engineering and Application (ICCEA) IEEE."},{"key":"e_1_2_11_31_1","doi-asserted-by":"publisher","DOI":"10.1088\/1361-6560\/ac97d9"},{"key":"e_1_2_11_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2024.3515804"},{"key":"e_1_2_11_33_1","doi-asserted-by":"crossref","unstructured":"E.LiuandH.Zhu \u201cA Quick Employment of Markov Decision Process (MDP) in Partially Unknown Three\u2010Dimensional Discrete Space \u201d20232023 International Conference on Intelligent Computing and Control (IC&C) 48\u201358 doi: 10.1109\/IC\u2010C57619.2023.00016.","DOI":"10.1109\/IC-C57619.2023.00016"},{"key":"e_1_2_11_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSII.2021.3061777"},{"key":"e_1_2_11_35_1","doi-asserted-by":"crossref","unstructured":"Y.Huang L.Li andH.Xu \u201cA Load\u2010Balancing Strategy Based on Multi\u2010Task Learning in a Distributed Training Environment \u201d2023International Conference on Advances in Electrical Engineering and Computer Applications (AEECA) pp. 862\u2013868 2023.","DOI":"10.1109\/AEECA59734.2023.00158"},{"key":"e_1_2_11_36_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12532-012-0044-1"},{"key":"e_1_2_11_37_1","doi-asserted-by":"crossref","unstructured":"M.Zhang H.Chen C.Shen et al. \u201cLoraprune: Pruning Meets Low\u2010Rank Parameter\u2010Efficient Fine\u2010Tuning \u201d2023.","DOI":"10.18653\/v1\/2024.findings-acl.178"},{"key":"e_1_2_11_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/LCOMM.2020.3041453"},{"key":"e_1_2_11_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3335139"},{"key":"e_1_2_11_40_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-024-10915-y"},{"key":"e_1_2_11_41_1","unstructured":"Y.Huang Y.Cheng D.Chen et al. GPipe: Efficient Training of Giant Neural Networks Using Pipeline Parallelism. CoRR 2018 *abs\/1811.06965*https:\/\/arxiv.org\/abs\/1811.06965."}],"container-title":["Concurrency and Computation: Practice and Experience"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.70456","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,12]],"date-time":"2026-01-12T04:48:26Z","timestamp":1768193306000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/cpe.70456"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,2]]},"references-count":40,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["10.1002\/cpe.70456"],"URL":"https:\/\/doi.org\/10.1002\/cpe.70456","archive":["Portico"],"relation":{},"ISSN":["1532-0626","1532-0634"],"issn-type":[{"value":"1532-0626","type":"print"},{"value":"1532-0634","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,2]]},"assertion":[{"value":"2025-07-11","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-10","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-02","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70456"}}