{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T03:41:01Z","timestamp":1782790861467,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":39,"publisher":"ACM","license":[{"start":{"date-parts":[[2017,8,13]],"date-time":"2017-08-13T00:00:00Z","timestamp":1502582400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2017,8,13]]},"DOI":"10.1145\/3097983.3098147","type":"proceedings-article","created":{"date-parts":[[2017,8,4]],"date-time":"2017-08-04T18:35:54Z","timestamp":1501871754000},"page":"1275-1284","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Small Batch or Large Batch?"],"prefix":"10.1145","author":[{"given":"Peifeng","family":"Yin","sequence":"first","affiliation":[{"name":"IBM Almaden Research Center, San Jose, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ping","family":"Luo","sequence":"additional","affiliation":[{"name":"Chinese Academy of Sciences &amp; University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Taiga","family":"Nakamura","sequence":"additional","affiliation":[{"name":"IBM Almaden Research Center, San Jose, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2017,8,13]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"et al. Tensorflow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi , Ashish Agarwal , Paul Barham , Eugene Brevdo , Zhifeng Chen , Craig Citro , Greg S Corrado , Andy Davis , Jeffrey Dean , Matthieu Devin , et al. Tensorflow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467 , 2016 . Mart\u00edn Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al. Tensorflow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467, 2016."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2124295.2124312"},{"key":"e_1_3_2_1_3_1","unstructured":"Rajen Bhatt and Abhinav Dhall. Skin segmentation dataset. UCI Machine Learning Repository.  Rajen Bhatt and Abhinav Dhall. Skin segmentation dataset. UCI Machine Learning Repository."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-7908-2604-3_16"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1561\/2200000016"},{"key":"e_1_3_2_1_6_1","volume-title":"The direct extension of admm for multi-block convex minimization problems is not necessarily convergent. Mathematical Programming, 155(1--2):57--79","author":"Chen Caihua","year":"2016","unstructured":"Caihua Chen , Bingsheng He , Yinyu Ye , and Xiaoming Yuan . The direct extension of admm for multi-block convex minimization problems is not necessarily convergent. Mathematical Programming, 155(1--2):57--79 , 2016 . Caihua Chen, Bingsheng He, Yinyu Ye, and Xiaoming Yuan. The direct extension of admm for multi-block convex minimization problems is not necessarily convergent. Mathematical Programming, 155(1--2):57--79, 2016."},{"key":"e_1_3_2_1_7_1","first-page":"1223","volume-title":"Advances in Neural Information Processing Systems","author":"Dean Jeffrey","year":"2012","unstructured":"Jeffrey Dean , Greg Corrado , Rajat Monga , Kai Chen , Matthieu Devin , Mark Mao , Andrew Senior , Paul Tucker , Ke Yang , Quoc V Le , Large scale distributed deep networks . In Advances in Neural Information Processing Systems , pages 1223 -- 1231 , 2012 . Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al. Large scale distributed deep networks. In Advances in Neural Information Processing Systems, pages 1223--1231, 2012."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1327452.1327492"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611971200"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2021068"},{"key":"e_1_3_2_1_11_1","unstructured":"The Apache Software Foundation. Apache hadoop. http:\/\/hadoop.apache.org\/core\/ 2009.  The Apache Software Foundation. Apache hadoop. http:\/\/hadoop.apache.org\/core\/ 2009."},{"key":"e_1_3_2_1_12_1","unstructured":"The Apache Software Foundation. Mahout project. http:\/\/mahout.apache.org 2012.  The Apache Software Foundation. Mahout project. http:\/\/mahout.apache.org 2012."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/0898-1221(76)90003-1"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/1870568.1870593"},{"issue":"2","key":"e_1_3_2_1_15_1","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1051\/m2an\/197509R200411","article-title":"par \u00e9l\u00e9ments finis d'ordre un, et la r\u00e9solution, par p\u00e9nalisation-dualit\u00e9 d'une classe de probl\u00e8mes de dirichlet non lin\u00e9aires. Revue franccaise d'automatique, informatique, recherche op\u00e9rationnelle","volume":"9","author":"Glowinski Roland","year":"1975","unstructured":"Roland Glowinski and A Marroco . Sur l'approximation , par \u00e9l\u00e9ments finis d'ordre un, et la r\u00e9solution, par p\u00e9nalisation-dualit\u00e9 d'une classe de probl\u00e8mes de dirichlet non lin\u00e9aires. Revue franccaise d'automatique, informatique, recherche op\u00e9rationnelle . Analyse num\u00e9rique , 9 ( 2 ): 41 -- 76 , 1975 . Roland Glowinski and A Marroco. Sur l'approximation, par \u00e9l\u00e9ments finis d'ordre un, et la r\u00e9solution, par p\u00e9nalisation-dualit\u00e9 d'une classe de probl\u00e8mes de dirichlet non lin\u00e9aires. Revue franccaise d'automatique, informatique, recherche op\u00e9rationnelle. Analyse num\u00e9rique, 9(2):41--76, 1975.","journal-title":"Analyse num\u00e9rique"},{"key":"e_1_3_2_1_16_1","first-page":"1223","volume-title":"Phillip B Gibbons, Garth A Gibson, Greg Ganger, and Eric P Xing. More effective distributed ml via a stale synchronous parallel parameter server. In Advances in neural information processing systems","author":"Ho Qirong","year":"2013","unstructured":"Qirong Ho , James Cipar , Henggang Cui , Seunghak Lee , Jin Kyu Kim , Phillip B Gibbons, Garth A Gibson, Greg Ganger, and Eric P Xing. More effective distributed ml via a stale synchronous parallel parameter server. In Advances in neural information processing systems , pages 1223 -- 1231 , 2013 . Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B Gibbons, Garth A Gibson, Greg Ganger, and Eric P Xing. More effective distributed ml via a stale synchronous parallel parameter server. In Advances in neural information processing systems, pages 1223--1231, 2013."},{"key":"e_1_3_2_1_17_1","first-page":"315","volume-title":"Advances in Neural Information Processing Systems","author":"Johnson Rie","year":"2013","unstructured":"Rie Johnson and Tong Zhang . Accelerating stochastic gradient descent using predictive variance reduction . In Advances in Neural Information Processing Systems , pages 315 -- 323 , 2013 . Rie Johnson and Tong Zhang. Accelerating stochastic gradient descent using predictive variance reduction. In Advances in Neural Information Processing Systems, pages 315--323, 2013."},{"key":"e_1_3_2_1_18_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik","year":"2014","unstructured":"Diederik Kingma and Jimmy Ba . Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 , 2014 . Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/2685048.2685095"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2623330.2623612"},{"key":"e_1_3_2_1_21_1","unstructured":"M. Lichman. UCI machine learning repository. http:\/\/archive.ics.uci.edu\/ml 2013.  M. Lichman. UCI machine learning repository. http:\/\/archive.ics.uci.edu\/ml 2013."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.14778\/2212351.2212354"},{"key":"e_1_3_2_1_23_1","first-page":"2915","volume-title":"Advances in Neural Information Processing Systems","author":"McMahan Brendan","year":"2014","unstructured":"Brendan McMahan and Matthew Streeter . Delay-tolerant algorithms for asynchronous distributed online learning . In Advances in Neural Information Processing Systems , pages 2915 -- 2923 , 2014 . Brendan McMahan and Matthew Streeter. Delay-tolerant algorithms for asynchronous distributed online learning. In Advances in Neural Information Processing Systems, pages 2915--2923, 2014."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.dss.2014.03.001"},{"key":"e_1_3_2_1_25_1","first-page":"63","volume-title":"A statistical study of on-line learning. Online Learning and Neural Networks","author":"Murata Noboru","year":"1998","unstructured":"Noboru Murata . A statistical study of on-line learning. Online Learning and Neural Networks . Cambridge University Press , Cambridge, UK , pages 63 -- 92 , 1998 . Noboru Murata. A statistical study of on-line learning. Online Learning and Neural Networks. Cambridge University Press, Cambridge, UK, pages 63--92, 1998."},{"key":"e_1_3_2_1_26_1","first-page":"543","volume-title":"Doklady an SSSR","author":"Nesterov Yurii","year":"1983","unstructured":"Yurii Nesterov . A method for unconstrained convex minimization problem with the rate of convergence o (1\/k2) . In Doklady an SSSR , volume 269 , pages 543 -- 547 , 1983 . Yurii Nesterov. A method for unconstrained convex minimization problem with the rate of convergence o (1\/k2). In Doklady an SSSR, volume 269, pages 543--547, 1983."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1137\/0330046"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0893-6080(98)00116-6"},{"key":"e_1_3_2_1_29_1","first-page":"693","volume-title":"Advances in Neural Information Processing Systems","author":"Recht Benjamin","year":"2011","unstructured":"Benjamin Recht , Christopher Re , Stephen Wright , and Feng Niu . Hogwild : A lock-free approach to parallelizing stochastic gradient descent . In Advances in Neural Information Processing Systems , pages 693 -- 701 , 2011 . Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu. Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Advances in Neural Information Processing Systems, pages 693--701, 2011."},{"key":"e_1_3_2_1_30_1","first-page":"2663","volume-title":"Advances in Neural Information Processing Systems","author":"Roux Nicolas L","year":"2012","unstructured":"Nicolas L Roux , Mark Schmidt , and Francis R Bach . A stochastic gradient method with an exponential convergence rate for finite training sets . In Advances in Neural Information Processing Systems , pages 2663 -- 2671 , 2012 . Nicolas L Roux, Mark Schmidt, and Francis R Bach. A stochastic gradient method with an exponential convergence rate for finite training sets. In Advances in Neural Information Processing Systems, pages 2663--2671, 2012."},{"key":"e_1_3_2_1_31_1","first-page":"378","volume-title":"Advances in Neural Information Processing Systems","author":"Shalev-Shwartz Shai","year":"2013","unstructured":"Shai Shalev-Shwartz and Tong Zhang . Accelerated mini-batch stochastic dual coordinate ascent . In Advances in Neural Information Processing Systems , pages 378 -- 385 , 2013 . Shai Shalev-Shwartz and Tong Zhang. Accelerated mini-batch stochastic dual coordinate ascent. In Advances in Neural Information Processing Systems, pages 378--385, 2013."},{"key":"e_1_3_2_1_32_1","volume-title":"Stochastic dual coordinate ascent methods for regularized loss minimization. Journal of Machine Learning Research, 14(Feb):567--599","author":"Shalev-Shwartz Shai","year":"2013","unstructured":"Shai Shalev-Shwartz and Tong Zhang . Stochastic dual coordinate ascent methods for regularized loss minimization. Journal of Machine Learning Research, 14(Feb):567--599 , 2013 . Shai Shalev-Shwartz and Tong Zhang. Stochastic dual coordinate ascent methods for regularized loss minimization. Journal of Machine Learning Research, 14(Feb):567--599, 2013."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2013.158"},{"key":"e_1_3_2_1_34_1","first-page":"2179","volume-title":"Advances in Neural Information Processing Systems","author":"Vainsencher Daniel","year":"2015","unstructured":"Daniel Vainsencher , Han Liu , and Tong Zhang . Local smoothness in variance reduced optimization . In Advances in Neural Information Processing Systems , pages 2179 -- 2187 , 2015 . Daniel Vainsencher, Han Liu, and Tong Zhang. Local smoothness in variance reduced optimization. In Advances in Neural Information Processing Systems, pages 2179--2187, 2015."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1137\/140961791"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2007.12.020"},{"key":"e_1_3_2_1_37_1","first-page":"10","article-title":"Cluster computing with working sets","volume":"10","author":"Zaharia Matei","year":"2010","unstructured":"Matei Zaharia , Mosharaf Chowdhury , Michael J Franklin , Scott Shenker , and Ion Stoica . Spark : Cluster computing with working sets . HotCloud , 10 : 10 -- 10 , 2010 . Matei Zaharia, Mosharaf Chowdhury, Michael J Franklin, Scott Shenker, and Ion Stoica. Spark: Cluster computing with working sets. HotCloud, 10:10--10, 2010.","journal-title":"HotCloud"},{"key":"e_1_3_2_1_38_1","volume-title":"Adadelta: an adaptive learning rate method. arXiv preprint arXiv:1212.5701","author":"Zeiler Matthew D","year":"2012","unstructured":"Matthew D Zeiler . Adadelta: an adaptive learning rate method. arXiv preprint arXiv:1212.5701 , 2012 . Matthew D Zeiler. Adadelta: an adaptive learning rate method. arXiv preprint arXiv:1212.5701, 2012."},{"key":"e_1_3_2_1_39_1","first-page":"2595","volume-title":"Advances in neural information processing systems","author":"Zinkevich Martin","year":"2010","unstructured":"Martin Zinkevich , Markus Weimer , Lihong Li , and Alex J Smola . Parallelized stochastic gradient descent . In Advances in neural information processing systems , pages 2595 -- 2603 , 2010 . Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola. Parallelized stochastic gradient descent. In Advances in neural information processing systems, pages 2595--2603, 2010."}],"event":{"name":"KDD '17: The 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","location":"Halifax NS Canada","acronym":"KDD '17","sponsor":["SIGMOD ACM Special Interest Group on Management of Data","SIGKDD ACM Special Interest Group on Knowledge Discovery in Data"]},"container-title":["Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3097983.3098147","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3097983.3098147","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:30:02Z","timestamp":1750217402000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3097983.3098147"}},"subtitle":["Gaussian Walk with Rebound Can Teach"],"short-title":[],"issued":{"date-parts":[[2017,8,13]]},"references-count":39,"alternative-id":["10.1145\/3097983.3098147","10.1145\/3097983"],"URL":"https:\/\/doi.org\/10.1145\/3097983.3098147","relation":{},"subject":[],"published":{"date-parts":[[2017,8,13]]},"assertion":[{"value":"2017-08-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}