{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T01:17:01Z","timestamp":1782782221072,"version":"3.54.5"},"reference-count":29,"publisher":"Springer Science and Business Media LLC","issue":"2-3","license":[{"start":{"date-parts":[[2022,12,12]],"date-time":"2022-12-12T00:00:00Z","timestamp":1670803200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,12,12]],"date-time":"2022-12-12T00:00:00Z","timestamp":1670803200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Parallel Prog"],"published-print":{"date-parts":[[2023,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The increasing size of deep neural networks (DNNs) raises a high demand for distributed training. An expert could find good hybrid parallelism strategies, but designing suitable strategies is time and labor-consuming. Therefore, automating parallelism strategy generation is crucial and desirable for DNN designers. Some automatic searching approaches have recently been studied to free the experts from the heavy parallel strategy conception. However, these approaches all rely on a numerical cost model, which requires heavy profiling results that lack portability. These profiling-based approaches cannot lighten the strategy generation work due to the non-reusable profiling value. Our intuition is that there is no need to estimate the actual execution time of the distributed training but to compare the relative cost of different strategies. We propose SMSG (Symbolic Modeling for Strategy Generation), which analyses the cost based on the communication and computation semantics. With SMSG, the parallel cost analyses are decoupled from hardware characteristics. SMSG defines cost functions for each kind of operator to quantitatively evaluate the amount of data for computation and communication, which eliminates the heavy profiling tasks. Besides, SMSG introduces how to apply functional transformation by using the Third Homomorphism theorem to control the high searching complexity. Our\u00a0experiments show that SMSG can find good hybrid parallelism strategies to generate an efficient training performance similar to\u00a0the state\u00a0of\u00a0the\u00a0art. Moreover, SMSG covers a wide variety of DNN models with good scalability. SMSG provides good portability when changing training configurations that a profiling-based approach cannot.<\/jats:p>","DOI":"10.1007\/s10766-022-00741-6","type":"journal-article","created":{"date-parts":[[2022,12,12]],"date-time":"2022-12-12T16:02:56Z","timestamp":1670860976000},"page":"109-127","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["SMSG: Profiling-Free Parallelism Modeling for Distributed Training of DNN"],"prefix":"10.1007","volume":"51","author":[{"given":"Haoran","family":"Wang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thibaut","family":"Tachon","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chong","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sophie","family":"Robert","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"S\u00e9bastien","family":"Limet","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,12,12]]},"reference":[{"key":"741_CR1","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown, T., Mann, B., Ryder, N.: Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 33, 1877\u20131901 (2020)","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"741_CR2","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770\u2013778 (2016)","DOI":"10.1109\/CVPR.2016.90"},{"key":"741_CR3","doi-asserted-by":"crossref","unstructured":"Cheng, H.-T., Koc, L., Harmsen, J., Shaked, T., Chandra, T., Aradhye, H., Anderson, G., Corrado, G., Chai, W., Ispir, M., Anil, R., Haque, Z., Hong, L., Jain, V., Liu, X., Shah, H.: Wide & deep learning for recommender systems. In: Proceedings of the 1st Workshop on Deep Learning for Recommender Systems. DLRS 2016, pp. 7\u201310. Association for Computing Machinery, New York, (2016)","DOI":"10.1145\/2988450.2988454"},{"issue":"6","key":"741_CR4","first-page":"1097","volume":"60","author":"A Krizhevsky","year":"2012","unstructured":"Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. Adv. Neural Inf. Process. Syst. 60(6), 1097\u20131105 (2012)","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"741_CR5","unstructured":"Dean, J., Corrado, G.S., Monga, R.: Large scale distributed deep networks. In: Proceedings of the 25th International Conference on Neural Information Processing Systems - Vol. 1. NIPS\u201912, pp. 1223\u20131231. Curran Associates Inc., Red Hook, (2012)"},{"key":"741_CR6","volume-title":"GPipe: Efficient training of giant neural networks using pipeline parallelism","author":"Y Huang","year":"2019","unstructured":"Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, M.X., Chen, D., Lee, H., Ngiam, J., Le, Q.V., Wu, Y., Chen, Z.: GPipe: Efficient training of giant neural networks using pipeline parallelism. Curran Associates Inc., Red Hook (2019)"},{"key":"741_CR7","unstructured":"Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., Catanzaro, B.: Megatron-LM: training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053 (2019)"},{"issue":"8","key":"741_CR8","doi-asserted-by":"publisher","first-page":"1967","DOI":"10.1109\/TPDS.2021.3132413","volume":"33","author":"Z Cai","year":"2022","unstructured":"Cai, Z., Yan, X., Ma, K., Wu, Y., Huang, Y., Cheng, J., Su, T., Yu, F.: Tensoropt: exploring the tradeoffs in distributed DNN training with auto-parallelism. IEEE Trans. Parallel. Distrib. Syst. 33(8), 1967\u20131981 (2022)","journal-title":"IEEE Trans. Parallel. Distrib. Syst."},{"key":"741_CR9","doi-asserted-by":"crossref","unstructured":"Fan, S., Rong, Y., Meng, C., Cao, Z., Wang, S., Zheng, Z., Wu, C., Long, G., Yang, J., Xia, L., Diao, L., Liu, X., Lin, W.: DAPPLE: a pipelined data parallel approach for training large models. In: Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. PPoPP \u201921, pp. 431\u2013445. Association for Computing Machinery, New York, (2021)","DOI":"10.1145\/3437801.3441593"},{"key":"741_CR10","unstructured":"Tarnawski, J.M., Narayanan, D., Phanishayee, A.: Piper: Multidimensional planner for DNN parallelization. In: Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P.S., Vaughan, J.W. (eds.) Advances in Neural Information Processing Systems, vol. 34, pp. 24829\u201324840 (2021)"},{"key":"741_CR11","unstructured":"Jia, Z., Zaharia, M., Aiken, A.: Beyond data and model parallelism for deep neural networks. In: Talwalkar, A., Smith, V., Zaharia, M. (eds.) Proceedings of Machine Learning and Systems, vol. 1, pp. 1\u201313 (2019)"},{"key":"741_CR12","unstructured":"Jia, Z., Lin, S., Qi, C.R., Aiken, A.: Exploring hidden dimensions in accelerating convolutional neural networks. In: Dy, J., Krause, A. (eds.) International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 80, pp. 2274\u20132283 (2018)"},{"key":"741_CR13","volume-title":"Supporting Very Large Models using Automatic Dataflow Graph Partitioning. EuroSys \u201919","author":"M Wang","year":"2019","unstructured":"Wang, M., Huang, C.-C., Li, J.: Supporting Very Large Models using Automatic Dataflow Graph Partitioning. EuroSys \u201919. Association for Computing Machinery, New York (2019)"},{"key":"741_CR14","unstructured":"Devlin, J., Chang, M.-W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol. 1 (Long and Short Papers), pp. 4171\u20134186. Association for Computational Linguistics, Minneapolis, Minnesota (2019)"},{"issue":"140","key":"741_CR15","first-page":"1","volume":"21","author":"C Raffel","year":"2020","unstructured":"Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21(140), 1\u201367 (2020)","journal-title":"J. Mach. Learn. Res."},{"key":"741_CR16","unstructured":"Abadi, M., Barham, P., Chen, J., : Tensorflow: a system for large-scale machine learning. In: 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16), pp. 265\u2013283. USENIX Association, Savannah, GA (2016)"},{"key":"741_CR17","first-page":"8026","volume":"32","author":"A Paszke","year":"2019","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L.: Pytorch: an imperative style, high-performance deep learning library. Adv. Neural Inf. Process. Syst. 32, 8026\u20138037 (2019)","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"741_CR18","unstructured":"Huawei: MindSpore. Huawei. https:\/\/www.mindspore.cn\/"},{"key":"741_CR19","doi-asserted-by":"crossref","unstructured":"Wang, H., Li, C., Tachon, T., Wang, H., Yang, S., Limet, S., Robert, S.: Efficient and systematic partitioning of large and deep neural networks for parallelization. In: European Conference on Parallel Processing, pp. 201\u2013216, Springer, (2021).","DOI":"10.1007\/978-3-030-85665-6_13"},{"issue":"1","key":"741_CR20","doi-asserted-by":"publisher","first-page":"154","DOI":"10.1016\/j.jcss.2010.06.012","volume":"77","author":"LG Valiant","year":"2011","unstructured":"Valiant, L.G.: A bridging model for multi-core computing. J. Comput. Syst. Sci. 77(1), 154\u2013166 (2011)","journal-title":"J. Comput. Syst. Sci."},{"issue":"1","key":"741_CR21","doi-asserted-by":"publisher","first-page":"177","DOI":"10.1145\/1594834.1480905","volume":"44","author":"A Morihata","year":"2009","unstructured":"Morihata, A., Matsuzaki, K., Hu, Z., Takeichi, M.: The third homomorphism theorem on trees: downward & upward lead to divide-and-conquer. SIGPLAN Not. 44(1), 177\u2013185 (2009)","journal-title":"SIGPLAN Not."},{"issue":"4","key":"741_CR22","doi-asserted-by":"publisher","first-page":"657","DOI":"10.1017\/S0956796800001908","volume":"6","author":"J Gibbons","year":"1996","unstructured":"Gibbons, J.: Functional pearls: the third homomorphism theorem. J. Funct. Program. 6(4), 657\u2013665 (1996)","journal-title":"J. Funct. Program."},{"issue":"5","key":"741_CR23","doi-asserted-by":"publisher","first-page":"549","DOI":"10.1017\/S0956796897002864","volume":"7","author":"G Huet","year":"1997","unstructured":"Huet, G.: The zipper. J. Funct. Program. 7(5), 549\u2013554 (1997)","journal-title":"J. Funct. Program."},{"key":"741_CR24","unstructured":"Schwing, A., Urtasun, R.: Fully connected deep structured networks (2015)"},{"key":"741_CR25","unstructured":"Zeng, W., Ren, X., Su, T., Wang, H., Liao, Y., Wang, Z., Jiang, X., Yang, Z., Wang, K., Zhang, X., Li, C., Gong, Z., Yao, Y., Huang, X., Wang, J., Yu, J., Guo, Q., Yu, Y., Zhang, Y., Wang, J., Tao, H., Yan, D., Yi, Z., Peng, F., Jiang, F., Zhang, H., Deng, L., Zhang, Y., Lin, Z., Zhang, C., Zhang, S., Guo, M., Gu, S., Fan, G., Wang, Y., Jin, X., Liu, Q., Tian, Y.: PanGu-$$\\alpha$$: Large-scale autoregressive pretrained chinese language models with auto-parallel computation. CoRR (2021) arXiv:2104.12369"},{"key":"741_CR26","unstructured":"Huawei: Atlas900. https:\/\/e.huawei.com\/en\/products\/cloud-computing-dc\/atlas\/atlas-900-ai"},{"key":"741_CR27","unstructured":"Krizhevsky, A.: One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv:1404.5997 (2014)"},{"key":"741_CR28","unstructured":"Oldridge, E., Perez, J., Frederickson, B., Koumchatzky, N., Lee, M., Wang, Z., Wu, L., Yu, F., Zamora, R., Yilmaz, O., et al.: Merlin: a gpu accelerated recommendation framework. Proceedings of IRS (2020)"},{"key":"741_CR29","unstructured":"Zeng, W., Ren, X., Su, T., Wang, H., Liao, Y., Wang, Z., Jiang, X., Yang, Z., Wang, K., Zhang, X., et al.: Pangu-$$\\alpha$$: Large-scale autoregressive pretrained chinese language models with auto-parallel computation. arXiv preprint arXiv:2104.12369 (2021)"}],"container-title":["International Journal of Parallel Programming"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-022-00741-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10766-022-00741-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-022-00741-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,3,30]],"date-time":"2023-03-30T11:13:18Z","timestamp":1680174798000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10766-022-00741-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,12]]},"references-count":29,"journal-issue":{"issue":"2-3","published-print":{"date-parts":[[2023,6]]}},"alternative-id":["741"],"URL":"https:\/\/doi.org\/10.1007\/s10766-022-00741-6","relation":{},"ISSN":["0885-7458","1573-7640"],"issn-type":[{"value":"0885-7458","type":"print"},{"value":"1573-7640","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,12]]},"assertion":[{"value":"10 September 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 November 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 December 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}]}}