{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,14]],"date-time":"2026-05-14T23:45:41Z","timestamp":1778802341311,"version":"3.51.4"},"reference-count":37,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2022,3,24]],"date-time":"2022-03-24T00:00:00Z","timestamp":1648080000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62072146"],"award-info":[{"award-number":["62072146"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61972376"],"award-info":[{"award-number":["61972376"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62072431"],"award-info":[{"award-number":["62072431"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"The Key Research and Development Program of Zhejiang Province under Grant","award":["2019C01059"],"award-info":[{"award-number":["2019C01059"]}]},{"name":"The Key Research and Development Program of Zhejiang Province under Grant","award":["2019C03135"],"award-info":[{"award-number":["2019C03135"]}]},{"name":"The Key Research and Development Program of Zhejiang Province under Grant","award":["2019C03134"],"award-info":[{"award-number":["2019C03134"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Deep learning, with increasingly large datasets and complex neural networks, is widely used in computer vision and natural language processing. A resulting trend is to split and train large-scale neural network models across multiple devices in parallel, known as parallel model training. Existing parallel methods are mainly based on expert design, which is inefficient and requires specialized knowledge. Although automatically implemented parallel methods have been proposed to solve these problems, these methods only consider a single optimization aspect of run time. In this paper, we present Trinity, an adaptive distributed parallel training method based on reinforcement learning, to automate the search and tuning of parallel strategies. We build a multidimensional performance evaluation model and use proximal policy optimization to co-optimize multiple optimization aspects. Our experiment used the CIFAR10 and PTB datasets based on InceptionV3, NMT, NASNet and PNASNet models. Compared with Google\u2019s Hierarchical method, Trinity achieves up to 5% reductions in runtime, communication, and memory overhead, and up to a 40% increase in parallel strategy search speeds.<\/jats:p>","DOI":"10.3390\/a15040108","type":"journal-article","created":{"date-parts":[[2022,3,25]],"date-time":"2022-03-25T00:05:18Z","timestamp":1648166718000},"page":"108","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Trinity: Neural Network Adaptive Distributed Parallel Training Method Based on Reinforcement Learning"],"prefix":"10.3390","volume":"15","author":[{"given":"Yan","family":"Zeng","sequence":"first","affiliation":[{"name":"School of Computing Science, Hangzhou Danzi University, Hangzhou 310018, China"},{"name":"Key Laboratory for Modeling and Simulation of Complex Systems, Ministry of Education, Hangzhou 310018, China"},{"name":"Data Security Governance Zhejiang Engineering Research Center, Hangzhou 310018, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiyang","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Computing Science, Hangzhou Danzi University, Hangzhou 310018, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jilin","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computing Science, Hangzhou Danzi University, Hangzhou 310018, China"},{"name":"Key Laboratory for Modeling and Simulation of Complex Systems, Ministry of Education, Hangzhou 310018, China"},{"name":"Data Security Governance Zhejiang Engineering Research Center, Hangzhou 310018, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongjian","family":"Ren","sequence":"additional","affiliation":[{"name":"School of Computing Science, Hangzhou Danzi University, Hangzhou 310018, China"},{"name":"Key Laboratory for Modeling and Simulation of Complex Systems, Ministry of Education, Hangzhou 310018, China"},{"name":"Data Security Governance Zhejiang Engineering Research Center, Hangzhou 310018, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yunquan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology Chinese Academy of Sciences, Beijing 100086, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,3,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Zhong, Z., Jin, L., and Huang, S. (2017, January 19). Deeptext: A new approach for text proposal generation and text detection in natural images. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7952348"},{"key":"ref_2","unstructured":"Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., and Le, Q.V. (2019). Xlnet: Generalized autoregressive pretraining for language understanding. arXiv."},{"key":"ref_3","unstructured":"Yuan, Y., Chen, X., and Wang, J. (2019). Object-contextual representations for semantic segmentation. arXiv."},{"key":"ref_4","unstructured":"Zhou, X., Wang, D., and Kr\u00e4henb\u00fchl, P. (2019). Objects as points. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Liu, B., Zhu, C., Li, G., Zhang, W., Lai, J., Tang, R., He, X., Li, Z., and Yu, Y. (2020, January 6\u201310). Autofis: Automatic feature interaction selection in factorization models for click-through rate prediction. Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Virtual Event.","DOI":"10.1145\/3394486.3403314"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Song, Q., Cheng, D., Zhou, H., Yang, J., Tian, Y., and Hu, X. (2020, January 6\u201310). Towards automated neural interaction discovery for click-through rate prediction. Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Virtual Event.","DOI":"10.1145\/3394486.3403137"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Peters, M.E., Ammar, W., Bhagavatula, C., and Power, R. (2017). Semi-supervised sequence tagging with bidirectional language models. arXiv.","DOI":"10.18653\/v1\/P17-1161"},{"key":"ref_8","unstructured":"Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., and Askell, A. (2020). Language models are few-shot learners. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"514","DOI":"10.1109\/ACCESS.2014.2325029","article-title":"Big data deep learning: Challenges and perspectives","volume":"2","author":"Chen","year":"2014","journal-title":"IEEE Access"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Lee, H., Hsieh, C.J., and Lee, J.S. (2021). Local critic training for model-parallel learning of deep neural networks. IEEE Trans. Neural Netw. Learn. Syst.","DOI":"10.1109\/TNNLS.2021.3057380"},{"key":"ref_11","unstructured":"Yu, H., Yang, S., and Zhu, S. (February, January 27). Parallel restarted SGD with faster convergence and less communication: Demystifying why model averaging works for deep learning. Proceedings of the AAAI Conference on Artificial Intelligence, Hilton Hawaiian Village, Honolulu, HI, USA."},{"key":"ref_12","unstructured":"Wu, Y., Schuster, M., Chen, Z., Le, Q.V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., and Macherey, K. (2016). Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. arXiv."},{"key":"ref_13","unstructured":"Sutskever, I., Vinyals, O., and Le, Q.V. (2014). Sequence to sequence learning with neural networks. arXiv."},{"key":"ref_14","unstructured":"Sun, S., Chen, W., Bian, J., Liu, X., and Liu, T.Y. (2018, January 10\u201315). Slim-DP: A multi-agent system for communication-efficient distributed deep learning. Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, Stockholm, Sweden."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Ballard, G., Buluc, A., Demmel, J., Grigori, L., Lipshitz, B., Schwartz, O., and Toledo, S. (2013, January 23\u201325). Communication optimal parallel multiplication of sparse random matrices. Proceedings of the Twenty-Fifth Annual ACM Symposium on Parallelism in Algorithms and Architectures, Montreal, QC, Canada.","DOI":"10.1145\/2486159.2486196"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Demmel, J., Eliahu, D., Fox, A., Kamil, S., Lipshitz, B., Schwartz, O., and Spillinger, O. (2013, January 20\u201324). Communication-optimal parallel recursive rectangular matrix multiplication. Proceedings of the 2013 IEEE 27th International Symposium on Parallel and Distributed Processing, Cambridge, MA, USA.","DOI":"10.1109\/IPDPS.2013.80"},{"key":"ref_17","unstructured":"Mirhoseini, A., Pham, H., Le, Q.V., Steiner, B., Larsen, R., Zhou, Y., Kumar, N., Norouzi, M., Bengio, S., and Dean, J. (2017, January 6\u201311). Device placement optimization with reinforcement learning. Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_18","unstructured":"Pellegrini, F., and Roman, J. (1996). Experimental analysis of the dual recursive bipartitioning algorithm for static mapping. TR 1038-96, LaBRI, URA CNRS 1304, Univ. Bordeaux I, Citeseer."},{"key":"ref_19","unstructured":"Schloss Dagstuhl-Leibniz-Zentrum, D.S.P. (2009). Distillating knowledge about Scotch. Dagstuhl Seminar Proceedings, F\u00fcr Informatik."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"101","DOI":"10.1002\/cpe.4330060203","article-title":"Fast multilevel implementation of recursive spectral bisection for partitioning unstructured problems","volume":"6","author":"Barnard","year":"1994","journal-title":"Concurr. Pract. Exp."},{"key":"ref_21","unstructured":"Jia, Z., Zaharia, M., and Aiken, A. (2018). Beyond data and model parallelism for deep neural networks. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Peng, Y., Bao, Y., Chen, Y., Wu, C., and Guo, C. (2018, January 23\u201326). Optimus: An efficient dynamic resource scheduler for deep learning clusters. Proceedings of the Thirteenth EuroSys Conference, Porto, Portugal.","DOI":"10.1145\/3190508.3190517"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wang, M., Huang, C.c., and Li, J. (2019, January 25\u201328). Supporting very large models using automatic dataflow graph partitioning. Proceedings of the Fourteenth EuroSys Conference 2019, Dresden, Germany.","DOI":"10.1145\/3302424.3303953"},{"key":"ref_24","unstructured":"Cai, Z., Ma, K., Yan, X., Wu, Y., Huang, Y., Cheng, J., Su, T., and Yu, F. (2020). TensorOpt: Exploring the Tradeoffs in Distributed DNN Training with Auto-Parallelism. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Yi, X., Luo, Z., Meng, C., Wang, M., Long, G., Wu, C., Yang, J., and Lin, W. (2020, January 7\u201311). Fast Training of Deep Learning Models over Multiple GPUs. Proceedings of the 21st International Middleware Conference, Delft, The Netherlands.","DOI":"10.1145\/3423211.3425675"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1095","DOI":"10.1214\/12-AOAS549","article-title":"Tree-guided group lasso for multi-response regression with structured sparsity, with an application to eQTL mapping","volume":"6","author":"Kim","year":"2012","journal-title":"Ann. Appl. Stat."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Frazier, P.I. (2018). A tutorial on Bayesian optimization. arXiv.","DOI":"10.1287\/educ.2018.0188"},{"key":"ref_28","unstructured":"Mirhoseini, A., Goldie, A., Pham, H., Steiner, B., Le, Q.V., and Dean, J. (May, January 30). A hierarchical model for device placement. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canda."},{"key":"ref_29","unstructured":"Gao, Y., Chen, L., and Li, B. (2018, January 3\u20135). Post: Device placement with cross-entropy minimization and proximal policy optimization. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_30","unstructured":"Addanki, R., Venkatakrishnan, S.B., Gupta, S., Mao, H., and Alizadeh, M. (2019). Placeto: Learning generalizable device placement algorithms for distributed machine learning. arXiv."},{"key":"ref_31","unstructured":"Paliwal, A., Gimeno, F., Nair, V., Li, Y., Lubin, M., Kohli, P., and Vinyals, O. (2019). Reinforced genetic algorithm learning for optimizing computation graphs. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Fiduccia, C.M., and Mattheyses, R.M. (1982, January 14\u201316). A linear-time heuristic for improving network partitions. Proceedings of the 19th Design Automation Conference, Las Vegas, NV, USA.","DOI":"10.1109\/DAC.1982.1585498"},{"key":"ref_33","unstructured":"Li, M., Andersen, D.G., Park, J.W., Smola, A.J., Ahmed, A., Josifovski, V., Long, J., Shekita, E.J., and Su, B.Y. (2014, January 6\u20138). Scaling distributed machine learning with the parameter server. Proceedings of the 11th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 14), Broomfield, CO, USA."},{"key":"ref_34","unstructured":"Li, M., Andersen, D.G., Smola, A.J., and Yu, K. (2014, January 8\u201313). Communication Efficient Distributed Machine Learning with the Parameter Server. Proceedings of the Advances in Neural Information Processing Systems (NIPS), New York, NY, USA."},{"key":"ref_35","unstructured":"Li, M., Zhou, L., Yang, Z., Li, A., Xia, F., Andersen, D.G., and Smola, A. (2013, January 9\u201310). Parameter server for distributed machine learning. Proceedings of the Big Learning NIPS Workshop, Lake Tahoe, SN, USA."},{"key":"ref_36","unstructured":"Bello, I., Pham, H., Le, Q.V., Norouzi, M., and Bengio, S. (2016). Neural combinatorial optimization with reinforcement learning. arXiv."},{"key":"ref_37","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/15\/4\/108\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:42:20Z","timestamp":1760136140000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/15\/4\/108"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,24]]},"references-count":37,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2022,4]]}},"alternative-id":["a15040108"],"URL":"https:\/\/doi.org\/10.3390\/a15040108","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,3,24]]}}}