{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,22]],"date-time":"2026-01-22T01:25:31Z","timestamp":1769045131654,"version":"3.49.0"},"reference-count":29,"publisher":"Wiley","issue":"1","license":[{"start":{"date-parts":[[2023,2,21]],"date-time":"2023-02-21T00:00:00Z","timestamp":1676937600000},"content-version":"vor","delay-in-days":51,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2020YFB2104000"],"award-info":[{"award-number":["2020YFB2104000"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61872132"],"award-info":[{"award-number":["61872132"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100014472","name":"Scientific Research Foundation of Hunan Provincial Education Department","doi-asserted-by":"publisher","award":["21A0535"],"award-info":[{"award-number":["21A0535"]}],"id":[{"id":"10.13039\/100014472","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004735","name":"Natural Science Foundation of Hunan Province","doi-asserted-by":"publisher","award":["2021JJ40635"],"award-info":[{"award-number":["2021JJ40635"]}],"id":[{"id":"10.13039\/501100004735","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100018579","name":"Training Program for Excellent Young Innovators of Changsha","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100018579","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["International Journal of Intelligent Systems"],"published-print":{"date-parts":[[2023,1]]},"abstract":"<jats:p>Distributed deep learning systems effectively respond to the increasing demand for large\u2010scale data processing in recent years. However, the significant investment in building distributed learning systems with powerful computing nodes places a huge financial burden on developers and researchers. It will be good to predict the precise benefit, i.e., how many times of speedup it can get compared with training on single machine (or a few), before actually building such big learning systems. To address this problem, this paper presents a novel performance model on training iteration time for heterogeneous distributed deep learning systems based on the characteristics of the parameter server (PS) system with bulk synchronous parallel (BSP) synchronization style. The accuracy of our performance model is demonstrated by comparing real measurement results on TensorFlow when training different neural networks with various kinds of hardware testbeds: the prediction accuracy is higher than 90% in most cases.<\/jats:p>","DOI":"10.1155\/2023\/2663115","type":"journal-article","created":{"date-parts":[[2023,2,21]],"date-time":"2023-02-21T23:35:06Z","timestamp":1677022506000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Modeling the Training Iteration Time for Heterogeneous Distributed Deep Learning Systems"],"prefix":"10.1155","volume":"2023","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2966-2521","authenticated-orcid":false,"given":"Yifu","family":"Zeng","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bowei","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pulin","family":"Pan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kenli","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6069-6869","authenticated-orcid":false,"given":"Guo","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2023,2,21]]},"reference":[{"key":"e_1_2_11_1_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2020.101898"},{"key":"e_1_2_11_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10044-020-00921-5"},{"key":"e_1_2_11_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/tnnls.2020.2979670"},{"key":"e_1_2_11_4_2","doi-asserted-by":"crossref","unstructured":"JooH. T.andKimK. J. Visualization of deep reinforcement learning using grad-cam: how ai plays atari games? Proceedings of the 2019 IEEE Conference on Games (CoG) August 2019 London UK 1\u20132.","DOI":"10.1109\/CIG.2019.8847950"},{"key":"e_1_2_11_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/tii.2021.3071405"},{"key":"e_1_2_11_6_2","doi-asserted-by":"publisher","DOI":"10.3389\/fgene.2021.607471"},{"key":"e_1_2_11_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/tnnls.2021.3054778"},{"key":"e_1_2_11_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/tmc.2019.2941492"},{"key":"e_1_2_11_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/tpds.2021.3134247"},{"key":"e_1_2_11_10_2","doi-asserted-by":"publisher","DOI":"10.14778\/3467861.3467867"},{"key":"e_1_2_11_11_2","unstructured":"AlqahtaniS.andDemirbasM. Performance analysis and comparison of distributed machine learning systems 2019 http:\/\/arxiv.org\/abs\/1909.02061."},{"key":"e_1_2_11_12_2","doi-asserted-by":"crossref","unstructured":"Castell\u00f3A. DolzM. F. Quintana-Ort\u00edE. S. andDuatoJ. Analysis of model parallelism for distributed neural networks Proceedings of the 26th European MPI Users\u2019 Group Meeting September 2019 Zurich Switzerland 1\u201310.","DOI":"10.1145\/3343211.3343218"},{"key":"e_1_2_11_13_2","doi-asserted-by":"crossref","unstructured":"OyamaY. NomuraA. SatoI. NishimuraH. TamatsuY. andMatsuokaS. Predicting statistics of asynchronous sgd parameters for a large-scale distributed deep learning system on gpu supercomputers Proceedings of the 2016 IEEE International Conference on Big Data (Big Data) December 2016 Washington DC USA 66\u201375.","DOI":"10.1109\/BigData.2016.7840590"},{"key":"e_1_2_11_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/access.2019.2916550"},{"key":"e_1_2_11_15_2","unstructured":"QiH. SparksE. R. andTalwalkarA. Paleo: A performance model for deep neural networks Proceedings of the 5th International Conference on Learning Representations April 2017 Toulon France."},{"key":"e_1_2_11_16_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2021.107914"},{"key":"e_1_2_11_17_2","doi-asserted-by":"crossref","unstructured":"YanF. RuwaseO. HeY. andChilimbiT. Performance modeling and scalability optimization of distributed deep learning systems Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining April 2015 Sydney Australia 1355\u20131364.","DOI":"10.1145\/2783258.2783270"},{"key":"e_1_2_11_18_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2019.06.034"},{"key":"e_1_2_11_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/tpds.2015.2414943"},{"key":"e_1_2_11_20_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6207"},{"key":"e_1_2_11_21_2","unstructured":"AbadiM. AgarwalA. BarhamP. BrevdoE. ChenZ. CitroC. CorradoG. S. DavisA. DeanJ. andDevinM. Tensorflow: large-scale machine learning on heterogeneous distributed systems 2016 http:\/\/arxiv.org\/abs\/1603.04467."},{"key":"e_1_2_11_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/tnet.2021.3087221"},{"key":"e_1_2_11_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/79173.79181"},{"key":"e_1_2_11_24_2","unstructured":"JiangY. ZhuY. LanC. YiB. CuiY. andGuoC. A unified architecture for accelerating distributed {DNN} training in heterogeneous GPU\/CPU clusters Proceedings of the 14th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 20) November 2020 Carlsbad CA USA 463\u2013479."},{"key":"e_1_2_11_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eng.2016.02.008"},{"key":"e_1_2_11_26_2","doi-asserted-by":"publisher","DOI":"10.1147\/jrd.2019.2947013"},{"key":"e_1_2_11_27_2","doi-asserted-by":"crossref","unstructured":"ShiH. ZhaoY. ZhangB. YoshigoeK. andVasilakosA. V. A free stale synchronous parallel strategy for distributed machine learning Proceedings of the 2019 International Conference on Big Data Engineering June 2019 Hong Kong China 23\u201329.","DOI":"10.1145\/3341620.3341625"},{"key":"e_1_2_11_28_2","doi-asserted-by":"crossref","unstructured":"ZhaoX. AnA. LiuJ. andChenB. X. Dynamic stale synchronous parallel distributed training for deep learning Proceedings of the 2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS) July 2019 Dallas Texas USA 1507\u20131517.","DOI":"10.1109\/ICDCS.2019.00150"},{"key":"e_1_2_11_29_2","doi-asserted-by":"crossref","unstructured":"DeanJ. Large-scale deep learning for building intelligent computer systems Proceedings of the Ninth ACM International Conference on Web Search and Data Mining February 2016 San Francisco CA USA.","DOI":"10.1145\/2835776.2835844"}],"container-title":["International Journal of Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/downloads.hindawi.com\/journals\/ijis\/2023\/2663115.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/ijis\/2023\/2663115.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1155\/2023\/2663115","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,31]],"date-time":"2024-12-31T05:20:12Z","timestamp":1735622412000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1155\/2023\/2663115"}},"subtitle":[],"editor":[{"given":"Tao","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2023,1]]},"references-count":29,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2023,1]]}},"alternative-id":["10.1155\/2023\/2663115"],"URL":"https:\/\/doi.org\/10.1155\/2023\/2663115","archive":["Portico"],"relation":{},"ISSN":["0884-8173","1098-111X"],"issn-type":[{"value":"0884-8173","type":"print"},{"value":"1098-111X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1]]},"assertion":[{"value":"2022-08-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-10-14","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"2663115"}}