{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T09:05:12Z","timestamp":1775639112683,"version":"3.50.1"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"7","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2019,3]]},"abstract":"<jats:p>A number of popular systems, most notably Google's TensorFlow, have been implemented from the ground up to support machine learning tasks. We consider how to make a very small set of changes to a modern relational database management system (RDBMS) to make it suitable for distributed learning computations. Changes include adding better support for recursion, and optimization and execution of very large compute plans. We also show that there are key advantages to using an RDBMS as a machine learning platform. In particular, learning based on a database management system allows for trivial scaling to large data sets and especially large models, where different computational units operate on different parts of a model that may be too large to fit into RAM.<\/jats:p>","DOI":"10.14778\/3317315.3317323","type":"journal-article","created":{"date-parts":[[2019,5,1]],"date-time":"2019-05-01T13:27:58Z","timestamp":1556717278000},"page":"822-835","source":"Crossref","is-referenced-by-count":32,"title":["Declarative recursive computation on an RDBMS"],"prefix":"10.14778","volume":"12","author":[{"given":"Dimitrije","family":"Jankov","sequence":"first","affiliation":[{"name":"Rice University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shangyu","family":"Luo","sequence":"additional","affiliation":[{"name":"Rice University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Binhang","family":"Yuan","sequence":"additional","affiliation":[{"name":"Rice University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhuhua","family":"Cai","sequence":"additional","affiliation":[{"name":"Rice University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jia","family":"Zou","sequence":"additional","affiliation":[{"name":"Rice University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chris","family":"Jermaine","sequence":"additional","affiliation":[{"name":"Rice University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zekai J.","family":"Gao","sequence":"additional","affiliation":[{"name":"Rice University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,3]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Bigdl. https:\/\/bigdl-project.github.io\/master\/ 2017. Accessed Sep 1 2018.  Bigdl. https:\/\/bigdl-project.github.io\/master\/ 2017. Accessed Sep 1 2018."},{"key":"e_1_2_1_2_1","unstructured":"Caffe2. https:\/\/caffe2.ai 2017. Accessed Sep 1 2018.  Caffe2. https:\/\/caffe2.ai 2017. Accessed Sep 1 2018."},{"key":"e_1_2_1_3_1","unstructured":"Chainerj. https:\/\/chainer.org\/ 2017. Accessed Sep 1 2018.  Chainerj. https:\/\/chainer.org\/ 2017. Accessed Sep 1 2018."},{"key":"e_1_2_1_4_1","unstructured":"Gluon. https:\/\/github.com\/gluon-api\/gluon-api 2017. Accessed Sep 1 2018.  Gluon. https:\/\/github.com\/gluon-api\/gluon-api 2017. Accessed Sep 1 2018."},{"key":"e_1_2_1_5_1","unstructured":"Introducing apache spark datasets. https:\/\/databricks.com\/blog\/2016\/01\/04\/introducing-apache-spark-datasets.html 2017. Accessed Sep 1 2018.  Introducing apache spark datasets. https:\/\/databricks.com\/blog\/2016\/01\/04\/introducing-apache-spark-datasets.html 2017. Accessed Sep 1 2018."},{"key":"e_1_2_1_6_1","unstructured":"Keras. https:\/\/keras.io\/ 2017. Accessed Sep 1 2018.  Keras. https:\/\/keras.io\/ 2017. Accessed Sep 1 2018."},{"key":"e_1_2_1_7_1","unstructured":"Pytorch. http:\/\/pytorch.org 2017. Accessed Sep 1 2018.  Pytorch. http:\/\/pytorch.org 2017. Accessed Sep 1 2018."},{"key":"e_1_2_1_8_1","unstructured":"Deeplearning4j. https:\/\/deeplearning4j.org\/ 2018. Accessed Sep 1 2018.  Deeplearning4j. https:\/\/deeplearning4j.org\/ 2018. Accessed Sep 1 2018."},{"key":"e_1_2_1_9_1","unstructured":"M. Abadi A. Agarwal P. Barham E. Brevdo Z. Chen C. Citro G. S. Corrado A. Davis J. Dean M. Devin S. Ghemawat I. Goodfellow A. Harp G. Irving M. Isard Y. Jia R. Jozefowicz L. Kaiser M. Kudlur J. Levenberg D. Mane R. Monga S. Moore D. Murray C. Olah M. Schuster J. Shlens B. Steiner I. Sutskever K. Talwar P. Tucker V. Vanhoucke V. Vasudevan F. Viegas O. Vinyals P. Warden M. Wattenberg M. Wicke Y. Yu and X. Zheng. TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. arXivpreprint arXiv:1603.04467 2016.  M. Abadi A. Agarwal P. Barham E. Brevdo Z. Chen C. Citro G. S. Corrado A. Davis J. Dean M. Devin S. Ghemawat I. Goodfellow A. Harp G. Irving M. Isard Y. Jia R. Jozefowicz L. Kaiser M. Kudlur J. Levenberg D. Mane R. Monga S. Moore D. Murray C. Olah M. Schuster J. Shlens B. Steiner I. Sutskever K. Talwar P. Tucker V. Vanhoucke V. Vasudevan F. Viegas O. Vinyals P. Warden M. Wattenberg M. Wicke Y. Yu and X. Zheng. TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. arXivpreprint arXiv:1603.04467 2016."},{"key":"e_1_2_1_10_1","first-page":"265","volume-title":"OSDI","author":"Abadi M.","year":"2016"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/567752.567763"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/237814.237823"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2742797"},{"key":"e_1_2_1_14_1","volume-title":"NIPS","author":"Bergstra J.","year":"2011"},{"key":"e_1_2_1_15_1","doi-asserted-by":"crossref","unstructured":"L. S. Blackford J. Choi A. Cleary E. D'Azevedo J. Demmel I. Dhillon J. Dongarra S. Hammarling G. Henry A. Petitet etal ScaLAPACK users' guide volume 4. 1997.  L. S. Blackford J. Choi A. Cleary E. D'Azevedo J. Demmel I. Dhillon J. Dongarra S. Hammarling G. Henry A. Petitet et al. ScaLAPACK users' guide volume 4. 1997.","DOI":"10.1137\/1.9780898719642"},{"key":"e_1_2_1_16_1","volume-title":"NIPS","author":"Blei D. M.","year":"2003"},{"key":"e_1_2_1_17_1","unstructured":"R. Burkard T. B\u00f6nniger G. Katzakidis and U. Derigs. Assignment and Matching Problems: Solution Methods with FORTRAN-Programs. Lecture Notes in Economics and Mathematical Systems. 2013.  R. Burkard T. B\u00f6nniger G. Katzakidis and U. Derigs. Assignment and Matching Problems: Solution Methods with FORTRAN-Programs. Lecture Notes in Economics and Mathematical Systems. 2013."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463676.2465283"},{"key":"e_1_2_1_19_1","first-page":"28","article-title":"Apache flink\u2122: Stream and batch processing in a single engine","volume":"38","author":"Carbone P.","year":"2015","journal-title":"IEEE Data Eng. Bull."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/275487.275492"},{"key":"e_1_2_1_21_1","unstructured":"J. Chen X. Pan R. Monga S. Bengio and R. Jozefowicz. Revisiting distributed synchronous sgd. arXiv preprint arXiv:1604.00981 2016.  J. Chen X. Pan R. Monga S. Bengio and R. Jozefowicz. Revisiting distributed synchronous sgd. arXiv preprint arXiv:1604.00981 2016."},{"key":"e_1_2_1_22_1","unstructured":"T. Chen M. Li Y. Li M. Lin N. Wang M. Wang T. Xiao B. Xu C. Zhang and Z. Zhang. MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. arXiv preprint arXiv:1512.01274 2015.  T. Chen M. Li Y. Li M. Lin N. Wang M. Wang T. Xiao B. Xu C. Zhang and Z. Zhang. MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. arXiv preprint arXiv:1512.01274 2015."},{"key":"e_1_2_1_23_1","first-page":"571","volume-title":"OSDI","author":"Chilimbi T.","year":"2014"},{"key":"e_1_2_1_24_1","volume-title":"ICML","author":"Coates A.","year":"2013"},{"key":"e_1_2_1_25_1","volume-title":"NIPS","author":"Collobert R.","year":"2011"},{"key":"e_1_2_1_26_1","first-page":"1223","volume-title":"NIPS","author":"Dean J.","year":"2012"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2094114.2094126"},{"key":"e_1_2_1_28_1","unstructured":"A. L. Gaunt M. A. Johnson M. Riechert D. Tarlow R. Tomioka D. Vytiniotis and S. Webster. AMPNet: Asynchronous Model-Parallel Training for Dynamic Neural Networks. arXiv preprint arXiv:1705.09786 2017.  A. L. Gaunt M. A. Johnson M. Riechert D. Tarlow R. Tomioka D. Vytiniotis and S. Webster. AMPNet: Asynchronous Model-Parallel Training for Dynamic Neural Networks. arXiv preprint arXiv:1705.09786 2017."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2011.5767930"},{"key":"e_1_2_1_30_1","unstructured":"P. Goyal P. Doll\u00e1r R. Girshick P. Noordhuis L. Wesolowski A. Kyrola A. Tulloch Y. Jia and K. He. Accurate large minibatch sgd: training imagenet in 1 hour. arXiv preprint arXiv:1706.02677 2017.  P. Goyal P. Doll\u00e1r R. Girshick P. Noordhuis L. Wesolowski A. Kyrola A. Tulloch Y. Jia and K. He. Accurate large minibatch sgd: training imagenet in 1 hour. arXiv preprint arXiv:1706.02677 2017."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/7902.7903"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/0893-6080(89)90020-8"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654889"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/276305.276315"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2011.5767867"},{"key":"e_1_2_1_36_1","unstructured":"A. Krizhevsky. One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv:1404.5997 2014.  A. Krizhevsky. One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv:1404.5997 2014."},{"key":"e_1_2_1_37_1","unstructured":"C.-G. Lee and Z. Ma. The generalized quadratic assignment problem. 2004.  C.-G. Lee and Z. Ma. The generalized quadratic assignment problem. 2004."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.5555\/2685048.2685095"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2017.108"},{"key":"e_1_2_1_40_1","first-page":"581","volume-title":"EDBT","author":"May N.","year":"2015"},{"key":"e_1_2_1_41_1","unstructured":"T. Mikolov K. Chen G. S. Corrado and J. Dean. Efficient estimation of word representations in vector space. CoRR abs\/1301.3781 2013.  T. Mikolov K. Chen G. S. Corrado and J. Dean. Efficient estimation of word representations in vector space. CoRR abs\/1301.3781 2013."},{"key":"e_1_2_1_42_1","first-page":"3111","volume-title":"NIPS","author":"Mikolov T.","year":"2013"},{"key":"e_1_2_1_43_1","unstructured":"G. Neubig C. Dyer Y. Goldberg A. Matthews W. Ammar A. Anastasopoulos M. Ballesteros D. Chiang D. Clothiaux T. Cohn K. Duh M. Faruqui C. Gan D. Garrette Y. Ji L. Kong A. Kuncoro G. Kumar C. Malaviya P. Michel Y. Oda M. Richardson N. Saphra S. Swayamdipta and P. Yin. DyNet: The Dynamic Neural Network Toolkit. arXiv preprint arXiv:1701.03980 2017.  G. Neubig C. Dyer Y. Goldberg A. Matthews W. Ammar A. Anastasopoulos M. Ballesteros D. Chiang D. Clothiaux T. Cohn K. Duh M. Faruqui C. Gan D. Garrette Y. Ji L. Kong A. Kuncoro G. Kumar C. Malaviya P. Michel Y. Oda M. Richardson N. Saphra S. Swayamdipta and P. Yin. DyNet: The Dynamic Neural Network Toolkit. arXiv preprint arXiv:1701.03980 2017."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2749436"},{"key":"e_1_2_1_45_1","first-page":"84","volume-title":"EDBT","author":"Passing L.","year":"2017"},{"key":"e_1_2_1_46_1","first-page":"693","volume-title":"NIPS","author":"Recht B.","year":"2011"},{"key":"e_1_2_1_47_1","unstructured":"S. Ruder. An overview of gradient descent optimization algorithms. arXiv preprint arXiv:1609.04747 2016.  S. Ruder. An overview of gradient descent optimization algorithms. arXiv preprint arXiv:1609.04747 2016."},{"key":"e_1_2_1_48_1","unstructured":"N. Shazeer A. Mirhoseini K. Maziarz A. Davis Q. V. Le G. E. Hinton and J. Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. CoRR abs\/1701.06538 2017.  N. Shazeer A. Mirhoseini K. Maziarz A. Davis Q. V. Le G. E. Hinton and J. Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. CoRR abs\/1701.06538 2017."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920931"},{"issue":"2","key":"e_1_2_1_50_1","first-page":"49","article-title":"Petuum: A new platform for distributed machine learning on big data","volume":"1","author":"Xing E. P.","year":"2015","journal-title":"KDD"},{"key":"e_1_2_1_51_1","unstructured":"D. Yu A. Eversole M. Seltzer K. Yao O. Kuchaiev Y. Zhang F. Seide Z. Huang B. Guenter H. Wang J. Droppo G. Zweig C. Rossbach J. Gao A. Stolcke J. Currey M. Slaney G. Chen A. Agarwal C. Basoglu M. Padmilac A. Kamenev V. Ivanov S. Cypher H. Parthasarathi B. Mitra B. Peng and X. Huang. An introduction to computational networks and the computational network toolkit. Technical report 2014.  D. Yu A. Eversole M. Seltzer K. Yao O. Kuchaiev Y. Zhang F. Seide Z. Huang B. Guenter H. Wang J. Droppo G. Zweig C. Rossbach J. Gao A. Stolcke J. Currey M. Slaney G. Chen A. Agarwal C. Basoglu M. Padmilac A. Kamenev V. Ivanov S. Cypher H. Parthasarathi B. Mitra B. Peng and X. Huang. An introduction to computational networks and the computational network toolkit. Technical report 2014."},{"key":"e_1_2_1_52_1","first-page":"1","volume-title":"HotCloud","author":"Zaharia M.","year":"2010"},{"key":"e_1_2_1_53_1","unstructured":"H. Zhang Z. Hu J. Wei P. Xie G. Kim Q. Ho and E. Xing. Poseidon: A system architecture for efficient gpu-based deep learning on multiple machines. arXiv preprint arXiv:1512.06216 2015.  H. Zhang Z. Hu J. Wei P. Xie G. Kim Q. Ho and E. Xing. Poseidon: A system architecture for efficient gpu-based deep learning on multiple machines. arXiv preprint arXiv:1512.06216 2015."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3317315.3317323","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:08:16Z","timestamp":1672225696000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3317315.3317323"}},"subtitle":["or, why you should use a database for distributed machine learning"],"short-title":[],"issued":{"date-parts":[[2019,3]]},"references-count":53,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2019,3]]}},"alternative-id":["10.14778\/3317315.3317323"],"URL":"https:\/\/doi.org\/10.14778\/3317315.3317323","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2019,3]]}}}