{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T01:43:49Z","timestamp":1782870229739,"version":"3.54.5"},"reference-count":77,"publisher":"Springer Science and Business Media LLC","issue":"11","license":[{"start":{"date-parts":[[2022,10,7]],"date-time":"2022-10-07T00:00:00Z","timestamp":1665100800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,10,7]],"date-time":"2022-10-07T00:00:00Z","timestamp":1665100800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100010661","name":"Horizon 2020 Framework Programme","doi-asserted-by":"publisher","award":["687691"],"award-info":[{"award-number":["687691"]}],"id":[{"id":"10.13039\/100010661","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2022,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Stochastic gradient descent (SGD) is a widely adopted iterative method for optimizing differentiable objective functions. In this paper, we propose and discuss a novel approach to scale up SGD in applications involving non-convex functions and large datasets. We address the bottleneck problem arising when using both shared and distributed memory. Typically, the former is bounded by limited computation resources and bandwidth whereas the latter suffers from communication overheads. We propose a unified distributed and parallel implementation of SGD (named DPSGD) that relies on both asynchronous distribution and lock-free parallelism. By combining two strategies into a unified framework, DPSGD is able to strike a better trade-off between local computation and communication. The convergence properties of DPSGD are studied for non-convex problems such as those arising in statistical modelling and machine learning. Our theoretical analysis shows that DPSGD leads to speed-up with respect to the number of cores and number of workers while guaranteeing an asymptotic convergence rate of<jats:inline-formula><jats:alternatives><jats:tex-math>$$O(1\/\\sqrt{T})$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mrow><mml:mi>O<\/mml:mi><mml:mo>(<\/mml:mo><mml:mn>1<\/mml:mn><mml:mo>\/<\/mml:mo><mml:msqrt><mml:mi>T<\/mml:mi><\/mml:msqrt><mml:mo>)<\/mml:mo><\/mml:mrow><\/mml:math><\/jats:alternatives><\/jats:inline-formula>given that the number of cores is bounded by<jats:inline-formula><jats:alternatives><jats:tex-math>$$T^{1\/4}$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:msup><mml:mi>T<\/mml:mi><mml:mrow><mml:mn>1<\/mml:mn><mml:mo>\/<\/mml:mo><mml:mn>4<\/mml:mn><\/mml:mrow><\/mml:msup><\/mml:math><\/jats:alternatives><\/jats:inline-formula>and the number of workers is bounded by<jats:inline-formula><jats:alternatives><jats:tex-math>$$T^{1\/2}$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:msup><mml:mi>T<\/mml:mi><mml:mrow><mml:mn>1<\/mml:mn><mml:mo>\/<\/mml:mo><mml:mn>2<\/mml:mn><\/mml:mrow><\/mml:msup><\/mml:math><\/jats:alternatives><\/jats:inline-formula>where<jats:italic>T<\/jats:italic>is the number of iterations. The potential gains that can be achieved by DPSGD are demonstrated empirically on a stochastic variational inference problem (Latent Dirichlet Allocation) and on a deep reinforcement learning (DRL) problem (advantage actor critic - A2C) resulting in two algorithms: DPSVI and HSA2C. Empirical results validate our theoretical findings. Comparative studies are conducted to show the performance of the proposed DPSGD against the state-of-the-art DRL algorithms.<\/jats:p>","DOI":"10.1007\/s10994-022-06243-3","type":"journal-article","created":{"date-parts":[[2022,10,7]],"date-time":"2022-10-07T16:04:12Z","timestamp":1665158652000},"page":"4039-4079","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Scaling up stochastic gradient descent for non-convex optimisation"],"prefix":"10.1007","volume":"111","author":[{"given":"Saad","family":"Mohamad","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hamad","family":"Alamri","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1980-5517","authenticated-orcid":false,"given":"Abdelhamid","family":"Bouchachia","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,10,7]]},"reference":[{"key":"6243_CR1","first-page":"265","volume":"16","author":"M Abadi","year":"2016","unstructured":"Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., et al. (2016). Tensorflow: A system for large-scale machine learning. OSDI, 16, 265\u2013283.","journal-title":"OSDI"},{"key":"6243_CR2","doi-asserted-by":"crossref","unstructured":"Abbeel, P., Coates, A., Quigley, M., & Ng, A.\u00a0Y. (2007). An application of reinforcement learning to aerobatic helicopter flight. In Advances in neural information processing systems (pp. 1\u20138).","DOI":"10.7551\/mitpress\/7503.003.0006"},{"key":"6243_CR3","doi-asserted-by":"crossref","unstructured":"Adamski, I., Adamski, R., Grel, T., J\u0119drych, A., Kaczmarek, K., & Michalewski, H. (2018). Distributed deep reinforcement learning: Learn how to play atari games in 21 minutes. arXiv preprint arXiv:1801.02852.","DOI":"10.1007\/978-3-319-92040-5_19"},{"key":"6243_CR4","doi-asserted-by":"crossref","unstructured":"Adamski, R., Grel, T., Klimek, M., & Michalewski, H. (2017). Atari games and intel processors. Workshop on Computer Games (pp. 1\u201318). Springer.","DOI":"10.1007\/978-3-319-75931-9_1"},{"key":"6243_CR5","doi-asserted-by":"crossref","unstructured":"Agarwal, A., & Duchi, J.C. (2011). Distributed delayed stochastic optimization. In Neural Information Processing Systems.","DOI":"10.1109\/CDC.2012.6426626"},{"key":"6243_CR6","unstructured":"Ba, J., Grosse, R., & Martens, J. (2016). Distributed second-order optimization using kronecker-factored approximations."},{"key":"6243_CR7","unstructured":"Babaeizadeh, M., Frosio, I., Tyree, S., Clemons, J., & Kautz, J. (2016). Reinforcement learning through asynchronous advantage actor-critic on a gpu. arXiv preprint arXiv:1611.06256."},{"key":"6243_CR8","doi-asserted-by":"publisher","first-page":"253","DOI":"10.1613\/jair.3912","volume":"47","author":"MG Bellemare","year":"2013","unstructured":"Bellemare, M. G., Naddaf, Y., Veness, J., & Bowling, M. (2013). The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 47, 253\u2013279.","journal-title":"Journal of Artificial Intelligence Research"},{"key":"6243_CR9","doi-asserted-by":"publisher","first-page":"859","DOI":"10.1080\/01621459.2017.1285773","volume":"112","author":"DM Blei","year":"2017","unstructured":"Blei, D. M., Kucukelbir, A., & McAuliffe, J. D. (2017). Variational inference: A review for statisticians. Journal of the American Statistical Association, 112, 859\u2013877.","journal-title":"Journal of the American Statistical Association"},{"key":"6243_CR10","first-page":"993","volume":"3","author":"DM Blei","year":"2003","unstructured":"Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent dirichlet allocation. Journal of Machine Learning research, 3, 993\u20131022.","journal-title":"Journal of Machine Learning research"},{"key":"6243_CR11","doi-asserted-by":"crossref","unstructured":"Bottou, L. (2010). Large-scale machine learning with stochastic gradient descent. In Proceedings of COMPSTAT\u20192010 (pp. 177\u2013186). Springer.","DOI":"10.1007\/978-3-7908-2604-3_16"},{"issue":"2","key":"6243_CR12","doi-asserted-by":"publisher","first-page":"223","DOI":"10.1137\/16M1080173","volume":"60","author":"L Bottou","year":"2018","unstructured":"Bottou, L., Curtis, F. E., & Nocedal, J. (2018). Optimization methods for large-scale machine learning. SIAM Review, 60(2), 223\u2013311.","journal-title":"SIAM Review"},{"key":"6243_CR13","unstructured":"Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). Openai gym. arXiv preprint arXiv:1606.01540."},{"key":"6243_CR14","unstructured":"Chilimbi, T.\u00a0M., Suzue, Y., Apacible, J., & Kalyanaraman, K. (2014). Project adam: Building an efficient and scalable deep learning training system. In OSDI."},{"key":"6243_CR15","unstructured":"Clemente, A.V., Castej\u00f3n, H.N., & Chandra, A. (2017). Efficient parallel methods for deep reinforcement learning. arXiv preprint arXiv:1705.04862."},{"key":"6243_CR16","unstructured":"Crane, R., & Roosta, F. (2019). Dingo: Distributed newton-type method for gradient-norm optimization. arXiv preprint arXiv:1901.05134."},{"key":"6243_CR17","first-page":"2656","volume":"28","author":"C De Sa","year":"2015","unstructured":"De Sa, C., Zhang, C., Olukotun, K., & R\u00e9, C. (2015). Taming the wild: A unified analysis of hogwild!-style algorithms. Advances in Neural Information Processing Systems, 28, 2656.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"6243_CR18","unstructured":"Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Senior, A., Tucker, P., Yang, K., & Le, Q.V., et\u00a0al. (2012). Large scale distributed deep networks. In Advances in neural information processing systems (pp. 1223\u20131231)."},{"key":"6243_CR19","first-page":"165","volume":"13","author":"O Dekel","year":"2012","unstructured":"Dekel, O., Gilad-Bachrach, R., Shamir, O., & Xiao, L. (2012). Optimal distributed online prediction using mini-batches. Journal of Machine Learning Research, 13, 165\u2013202.","journal-title":"Journal of Machine Learning Research"},{"key":"6243_CR20","unstructured":"Duchi, J.C., Chaturapruek, S., & R\u00e9, C. (2015). Asynchronous stochastic convex optimization. arXiv preprint arXiv:1508.00882."},{"key":"6243_CR21","doi-asserted-by":"publisher","first-page":"164","DOI":"10.1109\/TCOMM.2020.3026398","volume":"69","author":"A Elgabli","year":"2020","unstructured":"Elgabli, A., Park, J., Bedi, A. S., Issaid, C. B., Bennis, M., & Aggarwal, V. (2020). Q-GADMM: Quantized group ADMM for communication efficient decentralized machine learning. IEEE Transactions on Communications, 69, 164\u2013181.","journal-title":"IEEE Transactions on Communications"},{"key":"6243_CR22","unstructured":"Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., & Dunning, I., et\u00a0al. (2018). Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures. arXiv preprint arXiv:1802.01561."},{"key":"6243_CR23","doi-asserted-by":"crossref","unstructured":"Fang, C., & Lin, Z. (2017). Parallel asynchronous stochastic variance reduction for nonconvex optimization. In AAAI.","DOI":"10.1609\/aaai.v31i1.10651"},{"key":"6243_CR24","doi-asserted-by":"publisher","first-page":"2341","DOI":"10.1137\/120880811","volume":"23","author":"S Ghadimi","year":"2013","unstructured":"Ghadimi, S., & Lan, G. (2013). Stochastic first-and zeroth-order methods for nonconvex stochastic programming. SIAM Journal on Optimization, 23, 2341\u20132368.","journal-title":"SIAM Journal on Optimization"},{"key":"6243_CR25","first-page":"1303","volume":"14","author":"MD Hoffman","year":"2013","unstructured":"Hoffman, M. D., Blei, D. M., Wang, C., & Paisley, J. (2013). Stochastic variational inference. Journal of Machine Learning Research, 14, 1303\u20131347.","journal-title":"Journal of Machine Learning Research"},{"key":"6243_CR26","unstructured":"Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., van Hasselt, H., & Silver, D. (2018). Distributed prioritized experience replay. arXiv preprint arXiv:1803.00933."},{"key":"6243_CR27","unstructured":"Hsieh, C.-J., Yu, H.-F., & Dhillon, I. (2015). Passcode: Parallel asynchronous stochastic dual co-ordinate descent. In International Conference on Machine Learning (pp. 2370\u20132379). PMLR."},{"key":"6243_CR28","doi-asserted-by":"crossref","unstructured":"Huo, Z., & Huang, H. (2016). Asynchronous stochastic gradient descent with variance reduction for non-convex optimization. arXiv preprint arXiv:1604.03584.","DOI":"10.1609\/aaai.v31i1.10940"},{"key":"6243_CR29","doi-asserted-by":"crossref","unstructured":"Huo, Z., & Huang, H. (2017). Asynchronous mini-batch gradient descent with variance reduction for non-convex optimization. In AAAI.","DOI":"10.1609\/aaai.v31i1.10940"},{"key":"6243_CR30","doi-asserted-by":"publisher","first-page":"39","DOI":"10.1145\/1394608.1382172","volume":"36","author":"E Ipek","year":"2008","unstructured":"Ipek, E., Mutlu, O., Mart\u00ednez, J. F., & Caruana, R. (2008). Self-optimizing memory controllers: A reinforcement learning approach. ACM SIGARCH Computer Architecture News, 36, 39\u201350.","journal-title":"ACM SIGARCH Computer Architecture News"},{"key":"6243_CR31","unstructured":"Jahani, M., He, X., Ma, C., Mokhtari, A., Mudigere, D., Ribeiro, A., & Tak\u00e1c, M. (2020a). Efficient distributed hessian free algorithm for large-scale empirical risk minimization via accumulating sample strategy. In International Conference on Artificial Intelligence and Statistics (pp. 2634\u20132644). PMLR."},{"key":"6243_CR32","doi-asserted-by":"crossref","unstructured":"Jahani, M., Nazari, M., Rusakov, S., Berahas, A.\u00a0S., & Tak\u00e1\u010d, M. (2020b). Scaling up quasi-newton algorithms: Communication efficient distributed sr1. In International Conference on Machine Learning, Optimization, and Data Science (pp. 41\u201354). Springer.","DOI":"10.1007\/978-3-030-64583-0_5"},{"key":"6243_CR33","doi-asserted-by":"publisher","first-page":"183","DOI":"10.1023\/A:1007665907178","volume":"37","author":"MI Jordan","year":"1999","unstructured":"Jordan, M. I., Ghahramani, Z., Jaakkola, T. S., & Saul, L. K. (1999). An introduction to variational methods for graphical models. Machine Learning, 37, 183\u2013233.","journal-title":"Machine Learning"},{"key":"6243_CR34","unstructured":"Krizhevsky, A., Sutskever, I., & Hinton, G.\u00a0E. (2012). Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems (pp. 1097\u20131105)."},{"key":"6243_CR35","unstructured":"Langford, J., Smola, A.J., & Zinkevich, M. (2009). Slow learners are fast. Neural Information Processing Systems."},{"key":"6243_CR36","unstructured":"Leblond, R., Pedregosa, F., & Lacoste-Julien, S. (2017). Asaga: asynchronous parallel saga. In Artificial Intelligence and Statistics (pp. 46\u201354). PMLR."},{"key":"6243_CR37","doi-asserted-by":"crossref","unstructured":"Li, M., Andersen, D.\u00a0G., Park, J.\u00a0W., Smola, A.\u00a0J., Ahmed, A., Josifovski, V., Long, J., Shekita, E.\u00a0J., & Su, B.-Y. (2014a). Scaling distributed machine learning with the parameter server. In OSDI.","DOI":"10.1145\/2640087.2644155"},{"key":"6243_CR38","doi-asserted-by":"crossref","unstructured":"Li, M., Zhang, T., Chen, Y., & Smola, A.\u00a0J. (2014b). Efficient mini-batch training for stochastic optimization. In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 661\u2013670). ACM.","DOI":"10.1145\/2623330.2623612"},{"key":"6243_CR39","unstructured":"Lian, X., Huang, Y., Li, Y., & Liu, J. (2015). Asynchronous parallel stochastic gradient for nonconvex optimization. In Neural Information Processing Systems."},{"key":"6243_CR40","unstructured":"Lian, X., Zhang, W., Zhang, C., & Liu, J. (2018). Asynchronous decentralized parallel stochastic gradient descent. In International Conference on Machine Learning (pp. 3043\u20133052). PMLR."},{"key":"6243_CR41","unstructured":"Lichman, M. (2013). UCI machine learning repository."},{"key":"6243_CR42","unstructured":"Lin, T., Stich, S.\u00a0U., Patel, K.\u00a0K., & Jaggi, M. (2018). Don\u2019t use large mini-batches, use local sgd. arXiv preprint arXiv:1808.07217."},{"key":"6243_CR43","unstructured":"Mnih, V., Badia, A.\u00a0P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., & Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning (pp. 1928\u20131937)."},{"key":"6243_CR44","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., & Riedmiller, M. (2013). Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602."},{"issue":"7540","key":"6243_CR45","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529.","journal-title":"Nature"},{"key":"6243_CR46","doi-asserted-by":"crossref","unstructured":"Mohamad, S., Bouchachia, A., & Sayed-Mouchaweh, M. (2018). Asynchronous stochastic variational inference. arXiv preprint arXiv:1801.04289.","DOI":"10.1007\/978-3-030-16841-4_31"},{"key":"6243_CR47","unstructured":"Nair, A., Srinivasan, P., Blackwell, S., Alcicek, C., Fearon, R., De\u00a0Maria, A., Panneershelvam, V., Suleyman, M., Beattie, C., & Petersen, S., et\u00a0al. (2015). Massively parallel methods for deep reinforcement learning. arXiv preprint arXiv:1507.04296."},{"key":"6243_CR48","unstructured":"Neiswanger, W., Wang, C., & Xing, E. (2015). Embarrassingly parallel variational inference in nonconjugate models. arXiv preprint arXiv:1510.04163."},{"issue":"4","key":"6243_CR49","doi-asserted-by":"publisher","first-page":"1574","DOI":"10.1137\/070704277","volume":"19","author":"A Nemirovski","year":"2009","unstructured":"Nemirovski, A., Juditsky, A., Lan, G., & Shapiro, A. (2009). Robust stochastic approximation approach to stochastic programming. SIAM Journal on Optimization, 19(4), 1574\u20131609.","journal-title":"SIAM Journal on Optimization"},{"key":"6243_CR50","unstructured":"Niu, F., Recht, B., R\u00e9, C., & Wright, S.\u00a0J. (2011). Hogwild!: A lock-free approach to parallelizing stochastic gradient descent. arXiv preprint arXiv:1106.5730."},{"key":"6243_CR51","unstructured":"Ong, H.\u00a0Y., Chavez, K., & Hong, A. (2015). Distributed deep q-learning. arXiv preprint arXiv:1508.04186."},{"key":"6243_CR52","unstructured":"Paine, T., Jin, H., Yang, J., Lin, Z., & Huang, T. (2013). Gpu asynchronous stochastic gradient descent to speed up neural network training. arXiv preprint arXiv:1312.6186."},{"key":"6243_CR53","unstructured":"Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., & Lerer, A. (2017). Automatic differentiation in pytorch."},{"key":"6243_CR54","unstructured":"Recht, B., Re, C., Wright, S., & Niu, F. (2011). Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Neural Information Processing Systems."},{"key":"6243_CR55","unstructured":"Reddi, S.\u00a0J., Hefny, A., Sra, S., Poczos, B., & Smola, A. (2015). On variance reduction in stochastic gradient descent and its asynchronous variants. arXiv preprint arXiv:1506.06840."},{"key":"6243_CR56","doi-asserted-by":"publisher","first-page":"400","DOI":"10.1214\/aoms\/1177729586","volume":"22","author":"H Robbins","year":"1951","unstructured":"Robbins, H., & Monro, S. (1951). A stochastic approximation method. The Annals of Mathematical Statistics, 22, 400\u2013407.","journal-title":"The Annals of Mathematical Statistics"},{"key":"6243_CR57","unstructured":"Ruder, S. (2016). An overview of gradient descent optimization algorithms. arXiv preprint arXiv:1609.04747."},{"key":"6243_CR58","unstructured":"Schaul, T., Quan, J., Antonoglou, I., & Silver, D. (2015). Prioritized experience replay. arXiv preprint arXiv:1511.05952."},{"key":"6243_CR59","unstructured":"Shamir, O., Srebro, N., & Zhang, T. (2014). Communication-efficient distributed optimization using an approximate newton-type method. In International Conference on Machine Learning (pp. 1000\u20131008). PMLR."},{"issue":"7587","key":"6243_CR60","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., et al. (2016). Mastering the game of go with deep neural networks and tree search. Nature, 529(7587), 484\u2013489.","journal-title":"Nature"},{"key":"6243_CR61","unstructured":"Stich, S.\u00a0U. (2018). Local sgd converges fast and communicates little. arXiv preprint arXiv:1805.09767."},{"key":"6243_CR62","unstructured":"Stooke, A., & Abbeel, P. (2018). Accelerated methods for deep reinforcement learning. arXiv preprint arXiv:1803.02811."},{"key":"6243_CR63","doi-asserted-by":"crossref","unstructured":"Sutton, R. S., & Barto, A. G. (1998). Reinforcement learning: An introduction (Vol. 1). MIT Press.","DOI":"10.1109\/TNN.1998.712192"},{"key":"6243_CR64","unstructured":"Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction. MIT press."},{"key":"6243_CR65","first-page":"1057","volume":"12","author":"RS Sutton","year":"2000","unstructured":"Sutton, R. S., McAllester, D. A., Singh, S. P., & Mansour, Y. (2000). Policy gradient methods for reinforcement learning with function approximation. Advances in Neural Information Processing Systems, 12, 1057\u20131063.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"6243_CR66","doi-asserted-by":"crossref","unstructured":"Theocharous, G., Thomas, P.S., & Ghavamzadeh, M. (2015). Personalized ad recommendation systems for life-time value optimization with guarantees. In IJCAI (pp. 1806\u20131812).","DOI":"10.1145\/2740908.2741998"},{"issue":"9","key":"6243_CR67","doi-asserted-by":"publisher","first-page":"803","DOI":"10.1109\/TAC.1986.1104412","volume":"31","author":"J Tsitsiklis","year":"1986","unstructured":"Tsitsiklis, J., Bertsekas, D., & Athans, M. (1986). Distributed asynchronous deterministic and stochastic gradient optimization algorithms. IEEE Transactions on Automatic Control, 31(9), 803\u2013812.","journal-title":"IEEE Transactions on Automatic Control"},{"key":"6243_CR68","doi-asserted-by":"crossref","unstructured":"Wainwright, M.J., & Jordan, M.I., et\u00a0al. (2008). Graphical models, exponential families, and variational inference. Foundations and Trends\u00ae in Machine Learning.","DOI":"10.1561\/9781601981851"},{"key":"6243_CR69","doi-asserted-by":"crossref","unstructured":"Wang, J., Sahu, A.\u00a0K., Yang, Z., Joshi, G., & Kar, S. (2019). Matcha: Speeding up decentralized sgd via matching decomposition sampling. In 2019 Sixth Indian Control Conference (ICC) (pp. 299\u2013300). IEEE.","DOI":"10.1109\/ICC47138.2019.9123209"},{"key":"6243_CR70","doi-asserted-by":"crossref","unstructured":"Williams, R.J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. In Reinforcement Learning (pp. 5\u201332). Springer.","DOI":"10.1007\/978-1-4615-3618-5_2"},{"key":"6243_CR71","doi-asserted-by":"crossref","unstructured":"Xing, E.\u00a0P., Ho, Q., Dai, W., Kim, J.\u00a0K., Wei, J., Lee, S., Zheng, X., Xie, P., Kumar, A., & Yu, Y. (2015). Petuum: A new platform for distributed machine learning on big data. IEEE Transactions on Big Data.","DOI":"10.1145\/2783258.2783323"},{"key":"6243_CR72","doi-asserted-by":"publisher","first-page":"5693","DOI":"10.1609\/aaai.v33i01.33015693","volume":"33","author":"H Yu","year":"2019","unstructured":"Yu, H., Yang, S., & Zhu, S. (2019). Parallel restarted SGD with faster convergence and less communication: Demystifying why model averaging works for deep learning. Proceedings of the AAAI Conference on Artificial Intelligence, 33, 5693\u20135700.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"6243_CR73","unstructured":"Zhang, S., Choromanska, A.\u00a0E., & LeCun, Y. (2015). Deep learning with elastic averaging sgd. In Neural Information Processing Systems."},{"key":"6243_CR74","doi-asserted-by":"crossref","unstructured":"Zhao, S.-Y., & Li, W.-J. (2016). Fast asynchronous parallel stochastic gradient descent: A lock-free approach with convergence guarantee. In AAAI.","DOI":"10.1609\/aaai.v30i1.10305"},{"key":"6243_CR75","doi-asserted-by":"crossref","unstructured":"Zhao, S.-Y., Zhang, G.-D., & Li, W.-J. (2017). Lock-free optimization for non-convex problems. In AAAI.","DOI":"10.1609\/aaai.v31i1.10921"},{"key":"6243_CR76","doi-asserted-by":"crossref","unstructured":"Zhou, F., & Cong, G. (2017). On the convergence properties of a $$k$$-step averaging stochastic gradient descent algorithm for nonconvex optimization. arXiv preprint arXiv:1708.01012.","DOI":"10.24963\/ijcai.2018\/447"},{"key":"6243_CR77","unstructured":"Zinkevich, M., Weimer, M., Li, L., & Smola, A.\u00a0J. (2010). Parallelized stochastic gradient descent. In Neural Information Processing Systems."}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06243-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-022-06243-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06243-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,28]],"date-time":"2023-11-28T16:10:26Z","timestamp":1701187826000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-022-06243-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,7]]},"references-count":77,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2022,11]]}},"alternative-id":["6243"],"URL":"https:\/\/doi.org\/10.1007\/s10994-022-06243-3","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,7]]},"assertion":[{"value":"5 July 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 April 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 July 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 October 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"None.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"Not applicable.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"Code could be make available at a later stage.","order":6,"name":"Ethics","group":{"name":"EthicsHeading","label":"Code availability"}}]}}