{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T18:16:11Z","timestamp":1785521771803,"version":"3.56.0"},"reference-count":25,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2021,8,30]],"date-time":"2021-08-30T00:00:00Z","timestamp":1630281600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,8,30]],"date-time":"2021-08-30T00:00:00Z","timestamp":1630281600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Ministerio de Ciencia, Innovaci\u00f3n y Universidades","award":["TIN2017-82972-R"],"award-info":[{"award-number":["TIN2017-82972-R"]}]},{"DOI":"10.13039\/501100011596","name":"Conselleria d\u2019Educaci\u00f3, Investigaci\u00f3, Cultura i Esport","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100011596","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100011596","name":"Conselleria d\u2019Educaci\u00f3, Investigaci\u00f3, Cultura i Esport","doi-asserted-by":"publisher","award":["PROMETEO\/2019\/109"],"award-info":[{"award-number":["PROMETEO\/2019\/109"]}],"id":[{"id":"10.13039\/501100011596","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Computing"],"published-print":{"date-parts":[[2023,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In this work, we build a general piece-wise model to analyze data-parallel (DP) training costs of convolutional neural networks (CNNs) on clusters of GPUs. This general model is based on <jats:italic>i)<\/jats:italic> multi-layer perceptrons (MLPs) in charge of modeling the NVIDIA cuDNN\/cuBLAS library kernels involved in the training of some of the state-of-the-art CNNs; and <jats:italic>ii)<\/jats:italic> an analytical model in charge of modeling the NVIDIA NCCL Allreduce collective primitive using the Ring algorithm. The CNN training scalability study performed using this model in combination with the Roofline technique on varying batch sizes, node (floating-point) arithmetic performance, node memory bandwidth, network link bandwidth, and cluster dimension unveil some crucial bottlenecks at both GPU and cluster level. To provide evidence of this analysis, we validate the accuracy of the proposed model against a Python library for distributed deep learning training.\n<\/jats:p>","DOI":"10.1007\/s00607-021-00997-9","type":"journal-article","created":{"date-parts":[[2021,8,30]],"date-time":"2021-08-30T11:02:52Z","timestamp":1630321372000},"page":"915-934","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Using machine learning to model the training scalability of convolutional neural networks on clusters of GPUs"],"prefix":"10.1007","volume":"105","author":[{"given":"Sergio","family":"Barrachina","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Adri\u00e1n","family":"Castell\u00f3","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mar","family":"Catal\u00e1n","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9466-3398","authenticated-orcid":false,"given":"Manuel F.","family":"Dolz","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jose I.","family":"Mestre","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,8,30]]},"reference":[{"key":"997_CR1","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1016\/j.parco.2019.03.005","volume":"85","author":"AA Awan","year":"2019","unstructured":"Awan AA, Manian KV, Chu CH, Subramoni H, Panda DK (2019) Optimized large-message broadcast for deep learning workloads: MPI, MPI+ NCCL, or NCCL2? Parallel Comput 85:141\u2013152","journal-title":"Parallel Comput"},{"key":"997_CR2","doi-asserted-by":"crossref","unstructured":"Barrachina S, Castell\u00f3 A, Catal\u00e1n M, Dolz MF, Mestre J (2021) A flexible research-oriented framework for distributed training of deep neural networks. In: 2021 IEEE international symposium on parallel distributed processing, workshops and Phd Forum, pp 730\u2013739","DOI":"10.1109\/IPDPSW52791.2021.00110"},{"key":"997_CR3","doi-asserted-by":"crossref","unstructured":"Barrachina S, Castell\u00f3 A, Catal\u00e1n M, Dolz MF, Mestre J (2021) Pydtnn: a user-friendly and extensible framework for distributed deep learning. J Supercomput 77(9):9971\u20139987","DOI":"10.1007\/s11227-021-03673-z"},{"issue":"4","key":"997_CR4","first-page":"65:1","volume":"52","author":"T Ben-Nun","year":"2019","unstructured":"Ben-Nun T, Hoefler T (2019) Demystifying parallel and distributed deep learning: an in-depth concurrency analysis. ACM Comput Surv 52(4):65:1-65:43","journal-title":"ACM Comput Surv"},{"key":"997_CR5","doi-asserted-by":"crossref","unstructured":"Castell\u00f3 A, Dolz MF, Quintana-Ort\u00ed ES, Duato J (2019) Analysis of model parallelism for distributed neural networks. In: EuroMPI \u201919. Association for Computing Machinery, New York, NY, Article 7, pp 1\u201310","DOI":"10.1145\/3343211.3343218"},{"key":"997_CR6","doi-asserted-by":"crossref","unstructured":"Castell\u00f3 A, Catal\u00e1n M, Dolz MF, Mestre JI, Quintana-Ort\u00ed ES, Duato J (2021) Performance modeling for distributed training of convolutional neural networks. In: 2021 29th Euromicro international conference on parallel, distributed and network-based processing (PDP), pp 99\u2013108","DOI":"10.1109\/PDP52278.2021.00024"},{"key":"997_CR7","doi-asserted-by":"crossref","unstructured":"Castell\u00f3 A, Dolz MF, Quintana-Ort\u00ed ES, Duato J (2019) Theoretical scalability analysis of distributed deep convolutional neural networks. In: 2019 19th IEEE\/ACM international symposium on cluster, cloud and grid computing (CCGRID), pp 534\u2013541","DOI":"10.1109\/CCGRID.2019.00068"},{"issue":"13","key":"997_CR8","doi-asserted-by":"publisher","first-page":"1749","DOI":"10.1002\/cpe.1206","volume":"19","author":"E Chan","year":"2007","unstructured":"Chan E, Heimlich M, Purkayastha A, van de Geijn R (2007) Collective communication: theory, practice, and experience: research articles. Concurr Comput Pract Exper 19(13):1749\u20131783","journal-title":"Concurr Comput Pract Exper"},{"key":"997_CR9","unstructured":"Chellapilla K, Puri S, Simard P (2006) High performance convolutional neural networks for document processing. In: International workshop on frontiers in handwriting recognition"},{"key":"997_CR10","doi-asserted-by":"crossref","unstructured":"Gholami A, Azad A, Jin P, Keutzer K, Buluc A (2018) Integrated model, batch, and domain parallelism in training neural networks, pp 77\u201386","DOI":"10.1145\/3210377.3210394"},{"key":"997_CR11","unstructured":"Google Inc. Tensorflow benchmarks"},{"issue":"2","key":"997_CR12","doi-asserted-by":"publisher","first-page":"713","DOI":"10.1007\/s11227-016-1779-7","volume":"73","author":"K Hasanov","year":"2017","unstructured":"Hasanov K, Lastovetsky A (2017) Hierarchical redesign of classic MPI reduction algorithms. J Supercomput 73(2):713\u2013725","journal-title":"J Supercomput"},{"key":"997_CR13","doi-asserted-by":"crossref","unstructured":"Higham CF, Higham DJ (2018) Deep learning: an introduction for applied mathematicians. arXiv:1801.05894","DOI":"10.1137\/18M1165748"},{"key":"997_CR14","unstructured":"Jia Z, Zaharia M, Aiken A (2018) Beyond data and model parallelism for deep neural networks. CoRR, arXiv:1807.05358"},{"key":"997_CR15","doi-asserted-by":"crossref","unstructured":"Justus D, Brennan J, Bonner S, McGough AS (2018) Predicting the computational cost of deep learning models. In: IEEE international conference on big data, Big Data 2018, Seattle, WA, USA, December 10\u201313. pp 3873\u20133882. IEEE","DOI":"10.1109\/BigData.2018.8622396"},{"key":"997_CR16","unstructured":"NVIDIA (2021) NCCL Tests. https:\/\/github.com\/NVIDIA\/nccl-tests"},{"key":"997_CR17","unstructured":"NVIDIA (2021) The NVIDIA Collective Communication Library (NCCL). https:\/\/developer.nvidia.com\/nccl"},{"issue":"5","key":"997_CR18","first-page":"92:1","volume":"51","author":"S Pouyanfar","year":"2018","unstructured":"Pouyanfar S, Sadiq S, Yan Y, Tian H, Tao Y, Reyes MP, Shyu M-L, Chen S-C, Iyengar SS (2018) A survey on deep learning: algorithms, techniques, and applications. ACM Comput Surv 51(5):92:1-92:36","journal-title":"ACM Comput Surv"},{"key":"997_CR19","unstructured":"Qi H, Sparks ER, Talwalkar A (2017) Paleo: a performance model for deep neural networks. In: Proceedings of the international conference on learning representations"},{"key":"997_CR20","unstructured":"Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition"},{"issue":"12","key":"997_CR21","doi-asserted-by":"publisher","first-page":"2295","DOI":"10.1109\/JPROC.2017.2761740","volume":"105","author":"V Sze","year":"2017","unstructured":"Sze V, Chen Y-H, Yang T-J, Emer JS (2017) Efficient processing of deep neural networks: a tutorial and survey. Proc IEEE 105(12):2295\u20132329","journal-title":"Proc IEEE"},{"issue":"1","key":"997_CR22","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1177\/1094342005051521","volume":"19","author":"R Thakur","year":"2005","unstructured":"Thakur R, Rabenseifner R, Gropp W (2005) Optimization of collective communication operations in MPICH. Int J High Perform Comput Appl 19(1):49\u201366","journal-title":"Int J High Perform Comput Appl"},{"key":"997_CR23","doi-asserted-by":"crossref","unstructured":"Williams S, Patterson D, Oliker L, Shalf J, Yelick K (2008) The roofline model: a pedagogical tool for program analysis and optimization. In: 2008 IEEE hot chips 20 symposium (HCS), pp 1\u201371","DOI":"10.1109\/HOTCHIPS.2008.7476531"},{"key":"997_CR24","doi-asserted-by":"crossref","unstructured":"You Y, Demmel J, Keutzer K, Hsieh C-J, Ying C, Hseu J (2018) Large-batch training for LSTM and beyond. Technical Report UCB\/EECS-2018-138, Electrical Engineering and Computer Sciences, University of California at Berkeley","DOI":"10.1145\/3295500.3356137"},{"key":"997_CR25","unstructured":"You Y, Gitman I, Ginsburg B (2017) Scaling SGD batch size to 32k for ImageNet training. arXiv:1708.03888"}],"container-title":["Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00607-021-00997-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00607-021-00997-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00607-021-00997-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,4,24]],"date-time":"2023-04-24T15:06:47Z","timestamp":1682348807000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00607-021-00997-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,30]]},"references-count":25,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2023,5]]}},"alternative-id":["997"],"URL":"https:\/\/doi.org\/10.1007\/s00607-021-00997-9","relation":{},"ISSN":["0010-485X","1436-5057"],"issn-type":[{"value":"0010-485X","type":"print"},{"value":"1436-5057","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,8,30]]},"assertion":[{"value":"30 April 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 August 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 August 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Open Access funding provided thanks to the CRUE-CSIC agreement with Springer Nature. This research was partially sponsored by projects TIN2017-82972-R of <i>Ministerio de Ciencia, Innovaci\u00f3n y Universidades<\/i> and Prometeo\/2019\/109 of the <i>Generalitat Valenciana<\/i>. Manuel F. Dolz was also supported by the Plan GenT project CDEIGENT\/2018\/014 of the <i>Generalitat Valenciana<\/i>.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Funding"}},{"value":"Not applicable","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of interest"}},{"value":"Not applicable","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Availability of data and material"}},{"value":"Not applicable","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Code availability"}}]}}