{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T09:50:24Z","timestamp":1773481824143,"version":"3.50.1"},"reference-count":19,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,1,2]],"date-time":"2023-01-02T00:00:00Z","timestamp":1672617600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,1,2]],"date-time":"2023-01-02T00:00:00Z","timestamp":1672617600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Data Sci. Eng."],"published-print":{"date-parts":[[2023,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Large-scale distributed training mainly consists of sub-model parallel training and parameter synchronization. With the expansion of training workers, the efficiency of parameter synchronization will be affected. To tackle this problem, we first propose 2D-TGA, a <jats:bold>g<\/jats:bold>rouping <jats:bold>A<\/jats:bold>llReduce method based on the two-dimensional <jats:bold>t<\/jats:bold>orus topology. This method synchronizes the model parameters by grouping and makes full use of bandwidth. Secondly, we propose a distributed algorithm, 2D-TGA-ADMM, which combines the 2D-TGA with the alternating direction method of multipliers (ADMM). It focuses on sub-model training and reduces the wait time among workers in the synchronization process. Finally, experimental results on the Tianhe-2 supercomputing platform show that compared with the <jats:inline-formula><jats:alternatives><jats:tex-math>$${\\mathtt {MPI\\_Allreduce}}$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mrow>\n                    <mml:mi>MPI<\/mml:mi>\n                    <mml:mi>_<\/mml:mi>\n                    <mml:mi>Allreduce<\/mml:mi>\n                  <\/mml:mrow>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>, the 2D-TGA could shorten the synchronization wait time by <jats:inline-formula><jats:alternatives><jats:tex-math>$$33\\%$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mrow>\n                    <mml:mn>33<\/mml:mn>\n                    <mml:mo>%<\/mml:mo>\n                  <\/mml:mrow>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>.<\/jats:p>","DOI":"10.1007\/s41019-022-00202-7","type":"journal-article","created":{"date-parts":[[2023,1,2]],"date-time":"2023-01-02T11:08:13Z","timestamp":1672657693000},"page":"61-72","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["A Communication Efficient ADMM-based Distributed Algorithm Using Two-Dimensional Torus Grouping AllReduce"],"prefix":"10.1007","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5260-3458","authenticated-orcid":false,"given":"Guozheng","family":"Wang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongmei","family":"Lei","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zeyu","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cunlu","family":"Peng","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,1,2]]},"reference":[{"key":"202_CR1","doi-asserted-by":"crossref","unstructured":"Chen Y, Blum RS, Sadler BM (2022) Communication efficient federated learning via ordered admm in a fully decentralized setting. arXiv preprint arXiv:2202.02580","DOI":"10.1109\/CISS53076.2022.9751166"},{"key":"202_CR2","doi-asserted-by":"publisher","first-page":"4226","DOI":"10.1109\/TSP.2020.3009007","volume":"68","author":"X Wang","year":"2020","unstructured":"Wang X, Ishii H, Du L, Cheng P, Chen J (2020) Privacy-preserving distributed machine learning via local randomization and admm perturbation. IEEE Trans Signal Proc 68:4226\u20134241","journal-title":"IEEE Trans Signal Proc"},{"issue":"7","key":"202_CR3","doi-asserted-by":"publisher","first-page":"4385","DOI":"10.1109\/TITS.2020.3036071","volume":"22","author":"G Raja","year":"2020","unstructured":"Raja G, Anbalagan S, Vijayaraghavan G, Theerthagiri S, Suryanarayan SV, Wu X-W (2020) Sp-cids: secure and private collaborative ids for vanets. IEEE Trans Int Trans Syst 22(7):4385\u20134393","journal-title":"IEEE Trans Int Trans Syst"},{"key":"202_CR4","doi-asserted-by":"crossref","unstructured":"Steck H, Dimakopoulou M, Riabov N, Jebara T (2020) Admm slim: sparse recommendations for many users. In: Proceedings of the 13th international conference on web search and data mining, pp 555\u2013563","DOI":"10.1145\/3336191.3371774"},{"issue":"2","key":"202_CR5","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3377454","volume":"53","author":"J Verbraeken","year":"2020","unstructured":"Verbraeken J, Wolting M, Katzy J, Kloppenburg J, Verbelen T, Rellermeyer JS (2020) A survey on distributed machine learning. ACM Comput Surv (CSUR) 53(2):1\u201333","journal-title":"ACM Comput Surv (CSUR)"},{"issue":"2","key":"202_CR6","doi-asserted-by":"publisher","first-page":"713","DOI":"10.1007\/s11227-016-1779-7","volume":"73","author":"K Hasanov","year":"2017","unstructured":"Hasanov K, Lastovetsky A (2017) Hierarchical redesign of classic mpi reduction algorithms. J Supercomput 73(2):713\u2013725","journal-title":"J Supercomput"},{"issue":"8","key":"202_CR7","doi-asserted-by":"publisher","first-page":"8111","DOI":"10.1007\/s11227-020-03590-7","volume":"77","author":"D Wang","year":"2021","unstructured":"Wang D, Lei Y, Xie J, Wang G (2021) Hsac-aladmm: an asynchronous lazy admm algorithm based on hierarchical sparse allreduce communication. J Supercomput 77(8):8111\u20138134","journal-title":"J Supercomput"},{"key":"202_CR8","doi-asserted-by":"crossref","unstructured":"Xie J, Lei Y (2019) Admmlib: a library of communication-efficient ad-admm for distributed machine learning. In: IFIP international conference on network and parallel computing. Springer, pp 322\u2013326","DOI":"10.1007\/978-3-030-30709-7_27"},{"issue":"12","key":"202_CR9","doi-asserted-by":"publisher","first-page":"581","DOI":"10.1016\/j.parco.2009.09.001","volume":"35","author":"P Sanders","year":"2009","unstructured":"Sanders P, Speck J, Tr\u00e4ff JL (2009) Two-tree algorithms for full bandwidth broadcast, reduction and scan. Parallel Comput 35(12):581\u2013594","journal-title":"Parallel Comput"},{"issue":"01","key":"202_CR10","doi-asserted-by":"publisher","first-page":"79","DOI":"10.1142\/S0129626407002880","volume":"17","author":"RL Graham","year":"2007","unstructured":"Graham RL, Barrett BW, Shipman GM, Woodall TS, Bosilca G (2007) Open mpi: A high performance, flexible implementation of mpi point-to-point communications. Parallel Process Lett 17(01):79\u201388","journal-title":"Parallel Process Lett"},{"issue":"2","key":"202_CR11","doi-asserted-by":"publisher","first-page":"117","DOI":"10.1016\/j.jpdc.2008.09.002","volume":"69","author":"P Patarasuk","year":"2009","unstructured":"Patarasuk P, Yuan X (2009) Bandwidth optimal all-reduce algorithms for clusters of workstations. J Parallel Distrib Comput 69(2):117\u2013124","journal-title":"J Parallel Distrib Comput"},{"key":"202_CR12","unstructured":"Research B (2017) baidu-allreduce. [Online]. https:\/\/github.com\/baidu-research\/baidu-allreduce"},{"key":"202_CR13","unstructured":"Mikami H, Suganuma H, Tanaka Y, Kageyama Y, et al (2018) Massively distributed sgd: imagenet\/resnet-50 training in a flash. arXiv preprint arXiv:1811.05233"},{"key":"202_CR14","unstructured":"Ying C, Kumar S, Chen D, Wang T, Cheng Y (2018) Image classification at supercomputer scale.  arXiv preprint arXiv:1811.06992"},{"key":"202_CR15","unstructured":"Jia X, Song S, He W, Wang Y, Rong H, Zhou F, Xie L, Guo Z, Yang Y, Yu L, et al Highly scalable deep learning training system with mixed-precision: training imagenet in four minutes. arXiv preprint arXiv:1807.11205"},{"key":"202_CR16","unstructured":"Goyal P, Doll\u00e1r P, Girshick R, Noordhuis P, Wesolowski L, Kyrola A, Tulloch A, Jia Y, He K (2017) Accurate, large minibatch sgd: training imagenet in 1 hour. arXiv preprint arXiv:1706.02677"},{"key":"202_CR17","doi-asserted-by":"crossref","unstructured":"Ueno Y, Yokota R (2019) Exhaustive study of hierarchical allreduce patterns for large messages between gpus. In: 2019 19th IEEE\/ACM international symposium on cluster, cloud and grid computing (CCGRID). IEEE, pp 430\u2013439","DOI":"10.1109\/CCGRID.2019.00057"},{"key":"202_CR18","doi-asserted-by":"crossref","unstructured":"Sun DL, Fevotte C (2014) Alternating direction method of multipliers for non-negative matrix factorization with the beta-divergence. In: 2014 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, pp 6201\u20136205","DOI":"10.1109\/ICASSP.2014.6854796"},{"key":"202_CR19","unstructured":"Lin C-J, Weng RC, Keerthi SS (2008) Trust region newton method for large-scale logistic regression. J Mach Learn Res, 9(4):627-650"}],"container-title":["Data Science and Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s41019-022-00202-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s41019-022-00202-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s41019-022-00202-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,28]],"date-time":"2023-02-28T15:04:56Z","timestamp":1677596696000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s41019-022-00202-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,2]]},"references-count":19,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2023,3]]}},"alternative-id":["202"],"URL":"https:\/\/doi.org\/10.1007\/s41019-022-00202-7","relation":{},"ISSN":["2364-1185","2364-1541"],"issn-type":[{"value":"2364-1185","type":"print"},{"value":"2364-1541","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,2]]},"assertion":[{"value":"8 April 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 October 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 December 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 January 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}