{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,5]],"date-time":"2026-05-05T02:21:21Z","timestamp":1777947681985,"version":"3.51.4"},"reference-count":117,"publisher":"Association for Computing Machinery (ACM)","issue":"10","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,6]]},"abstract":"<jats:p>Deep learning (DL) is growing in popularity for many data analytics applications, including among enterprises. Large business-critical datasets in such settings typically reside in RDBMSs or other data systems. The DB community has long aimed to bring machine learning (ML) to DBMS-resident data. Given past lessons from in-DBMS ML and recent advances in scalable DL systems, DBMS and cloud vendors are increasingly interested in adding more DL support for DB-resident data. Recently, a new parallel DL model selection execution approach called Model Hopper Parallelism (MOP) was proposed. In this paper, we characterize the particular suitability of MOP for DL on data systems, but to bring MOP-based DL to DB-resident data, we show that there is no single \"best\" approach, and an interesting tradeoff space of approaches exists. We explain four canonical approaches and build prototypes upon Greenplum Database, compare them analytically on multiple criteria (e.g., runtime efficiency and ease of governance) and compare them empirically with large-scale DL workloads. Our experiments and analyses show that it is non-trivial to meet all practical desiderata well and there is a Pareto frontier; for instance, some approaches are 3x-6x faster but fare worse on governance and portability. Our results and insights can help DBMS and cloud vendors design better DL support for DB users. All of our source code, data, and other artifacts are available at https:\/\/github.com\/makemebitter\/cerebro-ds.<\/jats:p>","DOI":"10.14778\/3467861.3467867","type":"journal-article","created":{"date-parts":[[2021,10,26]],"date-time":"2021-10-26T16:17:12Z","timestamp":1635265032000},"page":"1769-1782","source":"Crossref","is-referenced-by-count":29,"title":["Distributed deep learning on data systems"],"prefix":"10.14778","volume":"14","author":[{"given":"Yuhao","family":"Zhang","sequence":"first","affiliation":[{"name":"University of California"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Frank","family":"McQuillan","sequence":"additional","affiliation":[{"name":"VMware, Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nandish","family":"Jayaram","sequence":"additional","affiliation":[{"name":"Intuit, Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nikhil","family":"Kak","sequence":"additional","affiliation":[{"name":"VMware, Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ekta","family":"Khanna","sequence":"additional","affiliation":[{"name":"VMware, Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Orhan","family":"Kislal","sequence":"additional","affiliation":[{"name":"VMware, Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Domino","family":"Valdano","sequence":"additional","affiliation":[{"name":"VMware, Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arun","family":"Kumar","sequence":"additional","affiliation":[{"name":"University of California"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,26]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Cerebro Documentation. https:\/\/adalabucsd.github.io\/cerebro-system\/.  Cerebro Documentation. https:\/\/adalabucsd.github.io\/cerebro-system\/."},{"key":"e_1_2_1_2_1","unstructured":"First hand knowledge from the authors.  First hand knowledge from the authors."},{"key":"e_1_2_1_3_1","volume-title":"Accessed","author":"Learning Deploy Machine","year":"2020"},{"key":"e_1_2_1_4_1","volume-title":"Accessed","author":"Deep Neural The CREATE MODEL","year":"2020"},{"key":"e_1_2_1_5_1","unstructured":"Script for Tensorflow Model Averaging Accessed January 31 2020. https:\/\/github.com\/tensorflow\/tensor2tensor\/blob\/master\/tensor2tensor\/utils\/avg_checkpoints.py.  Script for Tensorflow Model Averaging Accessed January 31 2020. https:\/\/github.com\/tensorflow\/tensor2tensor\/blob\/master\/tensor2tensor\/utils\/avg_checkpoints.py."},{"key":"e_1_2_1_6_1","unstructured":"Code Release of This Work Accessed November 19 2020. https:\/\/github.com\/makemebitter\/cerebro-ds.  Code Release of This Work Accessed November 19 2020. https:\/\/github.com\/makemebitter\/cerebro-ds."},{"key":"e_1_2_1_7_1","unstructured":"About Greenplum Query Processing Accessed October 31 2020. https:\/\/gpdb.docs.pivotal.io\/560\/admin_guide\/query\/topics\/parallel-proc.html.  About Greenplum Query Processing Accessed October 31 2020. https:\/\/gpdb.docs.pivotal.io\/560\/admin_guide\/query\/topics\/parallel-proc.html."},{"key":"e_1_2_1_8_1","unstructured":"Google BigQuery ML Accessed October 31 2020. https:\/\/cloud.google.com\/bigquery-ml\/docs.  Google BigQuery ML Accessed October 31 2020. https:\/\/cloud.google.com\/bigquery-ml\/docs."},{"key":"e_1_2_1_9_1","unstructured":"Google BigQuery ML TensorFlow integration Accessed October 31 2020. https:\/\/cloud.google.com\/bigquery-ml\/docs\/making-predictions-with-imported-tensorflow-models.  Google BigQuery ML TensorFlow integration Accessed October 31 2020. https:\/\/cloud.google.com\/bigquery-ml\/docs\/making-predictions-with-imported-tensorflow-models."},{"key":"e_1_2_1_10_1","unstructured":"Horovod on Spark Accessed October 31 2020. https:\/\/github.com\/horovod\/horovod\/blob\/master\/docs\/spark.rst.  Horovod on Spark Accessed October 31 2020. https:\/\/github.com\/horovod\/horovod\/blob\/master\/docs\/spark.rst."},{"key":"e_1_2_1_11_1","unstructured":"MADlib Deep Learning Accessed October 31 2020. https:\/\/madlib.apache.org\/docs\/latest\/group__grp__dl.html.  MADlib Deep Learning Accessed October 31 2020. https:\/\/madlib.apache.org\/docs\/latest\/group__grp__dl.html."},{"key":"e_1_2_1_12_1","unstructured":"MADlib Model Selection Accessed October 31 2020. https:\/\/madlib.apache.org\/docs\/latest\/group__grp__keras__run__model__selection.html.  MADlib Model Selection Accessed October 31 2020. https:\/\/madlib.apache.org\/docs\/latest\/group__grp__keras__run__model__selection.html."},{"key":"e_1_2_1_13_1","unstructured":"Microsoft SQL Server Machine Learning Services Accessed October 31 2020. https:\/\/docs.microsoft.com\/en-us\/sql\/machine-learning\/sql-server-machine-learning-services?view=sql-server-2017.  Microsoft SQL Server Machine Learning Services Accessed October 31 2020. https:\/\/docs.microsoft.com\/en-us\/sql\/machine-learning\/sql-server-machine-learning-services?view=sql-server-2017."},{"key":"e_1_2_1_14_1","unstructured":"Oracle Data Mining Accessed October 31 2020. https:\/\/www.oracle.com\/database\/technologies\/advanced-analytics\/odm.html.  Oracle Data Mining Accessed October 31 2020. https:\/\/www.oracle.com\/database\/technologies\/advanced-analytics\/odm.html."},{"key":"e_1_2_1_15_1","unstructured":"Oracle Machine Learning Accessed October 31 2020. https:\/\/www.oracle.com\/data-science\/machine-learning.html.  Oracle Machine Learning Accessed October 31 2020. https:\/\/www.oracle.com\/data-science\/machine-learning.html."},{"key":"e_1_2_1_16_1","unstructured":"TensorFrames Accessed October 31 2020. https:\/\/github.com\/databricks\/tensorframes.  TensorFrames Accessed October 31 2020. https:\/\/github.com\/databricks\/tensorframes."},{"key":"e_1_2_1_17_1","unstructured":"TOAST Tables in Postgres Accessed October 31 2020. https:\/\/wiki.postgresql.org\/wiki\/TOAST.  TOAST Tables in Postgres Accessed October 31 2020. https:\/\/wiki.postgresql.org\/wiki\/TOAST."},{"key":"e_1_2_1_18_1","volume-title":"CIDR. www.cidrdb.org","author":"Agrawal A.","year":"2020"},{"key":"e_1_2_1_19_1","volume-title":"Accessed","author":"D.","year":"2020"},{"key":"e_1_2_1_20_1","first-page":"1","volume-title":"ICIS","author":"Akita R.","year":"2016"},{"key":"e_1_2_1_21_1","volume-title":"Accessed","year":"2020"},{"key":"e_1_2_1_22_1","first-page":"127","volume":"21","author":"Anil R.","year":"2020","journal-title":"Apache Mahout: Machine Learning on Distributed Dataflow Systems. J. Mach. Learn. Res."},{"key":"e_1_2_1_23_1","first-page":"223","volume-title":"DOOD","author":"Atkinson M. P.","year":"1989"},{"key":"e_1_2_1_24_1","volume-title":"Rmsprop and equilibrated adaptive learning rates for nonconvex optimization. corr abs\/1502.04390","author":"Bengio Y.","year":"2015"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.5555\/3042817.3042832"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.14778\/3007263.3007279"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.14778\/3229863.3229865"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732286.2732292"},{"key":"e_1_2_1_29_1","volume-title":"Inria Saclay Ile de France","author":"Bouthillier X.","year":"2020"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/UIC-ATC.2017.8397411"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/248603.248616"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213936"},{"key":"e_1_2_1_33_1","unstructured":"E. Commission. GDPR Accessed October 31 2020. https:\/\/ec.europa.eu\/info\/law\/law-topic\/data-protection\/eu-data-protection-rules_en.  E. Commission. GDPR Accessed October 31 2020. https:\/\/ec.europa.eu\/info\/law\/law-topic\/data-protection\/eu-data-protection-rules_en."},{"key":"e_1_2_1_34_1","volume-title":"Accessed","year":"2020"},{"key":"e_1_2_1_35_1","first-page":"31","volume":"4","author":"Databricks","year":"2020","journal-title":"Introducing Apache Spark 2"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389715"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.14778\/3236187.3236194"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994515"},{"key":"e_1_2_1_40_1","volume-title":"Accessed","year":"2020"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3386137"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213874"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2018.2873325"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098043"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.5555\/3086952"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.5555\/3172077.3172127"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.14778\/2367502.2367510"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3187009.3177734"},{"key":"e_1_2_1_49_1","volume-title":"Accessed","year":"2020"},{"key":"e_1_2_1_50_1","volume-title":"Population Based Training of Neural Networks. arXiv preprint arXiv:1711.09846","author":"Jaderberg M.","year":"2017"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3422648.3422659"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3380575"},{"key":"e_1_2_1_53_1","first-page":"13","volume":"2020","author":"Kaggle","year":"2021","journal-title":"Kaggle Survey"},{"key":"e_1_2_1_54_1","first-page":"31","volume":"2019","author":"Kaggle","year":"2020","journal-title":"State of Data Science and Machine Learning"},{"key":"e_1_2_1_55_1","volume-title":"CIDR. www.cidrdb.org","author":"Karanasos K.","year":"2020"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3196959.3196960"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661829.2661864"},{"key":"e_1_2_1_58_1","volume-title":"Adam: A Method for Stochastic Optimization. arXiv preprint arXiv:1412.6980","author":"Kingma D. P.","year":"2014"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342276"},{"key":"e_1_2_1_60_1","unstructured":"Kubeflow. Kubeflow Accessed November 26 2020. https:\/\/www.kubeflow.org\/.  Kubeflow. Kubeflow Accessed November 26 2020. https:\/\/www.kubeflow.org\/."},{"key":"e_1_2_1_61_1","volume-title":"Accessed","author":"Kumar A.","year":"2020"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/2935694.2935698"},{"key":"e_1_2_1_63_1","volume-title":"CIDR. www.cidrdb.org","author":"Kumar A.","year":"2021"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342633"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3300070"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735479.2735488"},{"key":"e_1_2_1_67_1","volume-title":"MLSys. mlsys.org","author":"Li L.","year":"2020"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.5555\/2685048.2685095"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3220023"},{"key":"e_1_2_1_70_1","volume-title":"Tune: A research platform for distributed model selection and training. arXiv preprint arXiv:1807.05118","author":"Liaw R.","year":"2018"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.124"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3360319"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2017.108"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.5555\/2946645.2946679"},{"key":"e_1_2_1_75_1","volume-title":"Accessed","year":"2020"},{"key":"e_1_2_1_76_1","unstructured":"MLflow. MLflow Accessed November 26 2020. https:\/\/mlflow.org\/.  MLflow. MLflow Accessed November 26 2020. https:\/\/mlflow.org\/."},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.5555\/3291168.3291210"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389709"},{"key":"e_1_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/3329486.3329496"},{"key":"e_1_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407816"},{"key":"e_1_2_1_81_1","volume-title":"Cerebro: A Data System for Optimized Deep Learning Model Selection. https:\/\/adalabucsd.github.io\/papers\/TR_2020_Cerebro.pdf","author":"Nakandala S.","year":"2020"},{"key":"e_1_2_1_82_1","volume-title":"Accessed","author":"S.","year":"2020"},{"key":"e_1_2_1_83_1","doi-asserted-by":"publisher","DOI":"10.1145\/2733373.2807410"},{"key":"e_1_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2006.31"},{"key":"e_1_2_1_85_1","doi-asserted-by":"publisher","DOI":"10.1145\/342009.336589"},{"key":"e_1_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.5555\/3277355.3277416"},{"key":"e_1_2_1_87_1","first-page":"473","volume-title":"EDBT","author":"Raasveldt M.","year":"2018"},{"key":"e_1_2_1_88_1","doi-asserted-by":"publisher","DOI":"10.14778\/3352063.3352110"},{"key":"e_1_2_1_89_1","volume-title":"MLSys. mlsys.org","author":"Renggli C.","year":"2019"},{"key":"e_1_2_1_90_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407796"},{"key":"e_1_2_1_91_1","volume-title":"Introducing Cloudlab: Scientific Infrastructure for Advancing Cloud Architectures and Applications","author":"Ricci R.","year":"2014"},{"key":"e_1_2_1_92_1","doi-asserted-by":"publisher","DOI":"10.1145\/3328519.3329134"},{"key":"e_1_2_1_93_1","doi-asserted-by":"publisher","DOI":"10.1145\/3329486.3329494"},{"key":"e_1_2_1_94_1","volume-title":"Horovod: Fast and Easy Distributed Deep Learning in TF. arXiv preprint arXiv:1802.05799","author":"Sergeev A.","year":"2018"},{"key":"e_1_2_1_95_1","doi-asserted-by":"publisher","DOI":"10.5555\/2621980"},{"key":"e_1_2_1_96_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3319863"},{"key":"e_1_2_1_97_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2017.109"},{"key":"e_1_2_1_98_1","volume-title":"Experiments on Parallel Training of Deep Neural Network using Model Averaging. CoRR, abs\/1507.01239","author":"Su H.","year":"2015"},{"key":"e_1_2_1_99_1","doi-asserted-by":"publisher","DOI":"10.1109\/69.109106"},{"key":"e_1_2_1_100_1","volume-title":"Accessed","author":"Pivotal Mware","year":"2020"},{"key":"e_1_2_1_101_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939753"},{"key":"e_1_2_1_102_1","doi-asserted-by":"publisher","DOI":"10.1145\/2783258.2783273"},{"key":"e_1_2_1_103_1","volume-title":"ACM","author":"Wang R.","year":"2017"},{"key":"e_1_2_1_104_1","doi-asserted-by":"publisher","DOI":"10.1145\/2733373.2806232"},{"key":"e_1_2_1_105_1","doi-asserted-by":"publisher","DOI":"10.14778\/3282495.3282499"},{"key":"e_1_2_1_106_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-015-0391-4"},{"key":"e_1_2_1_107_1","doi-asserted-by":"publisher","DOI":"10.1145\/2987550.2987586"},{"key":"e_1_2_1_108_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390303"},{"key":"e_1_2_1_109_1","doi-asserted-by":"publisher","DOI":"10.14778\/3297753.3297763"},{"key":"e_1_2_1_110_1","doi-asserted-by":"publisher","DOI":"10.1109\/69.755617"},{"key":"e_1_2_1_111_1","volume-title":"Tensor Relational Algebra for Machine Learning System Design. CoRR, abs\/2009.00524","author":"Yuan B.","year":"2020"},{"key":"e_1_2_1_112_1","volume-title":"CIDR. www.cidrdb.org","author":"Zaharia M.","year":"2021"},{"key":"e_1_2_1_113_1","volume-title":"Parallel SGD: When does averaging help? CoRR, abs\/1606.07365","author":"Zhang J.","year":"2016"},{"key":"e_1_2_1_114_1","doi-asserted-by":"publisher","DOI":"10.1631\/FITEE.1700808"},{"key":"e_1_2_1_115_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3314038"},{"key":"e_1_2_1_116_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2019.00194"},{"key":"e_1_2_1_117_1","doi-asserted-by":"publisher","DOI":"10.5555\/2997046.2997185"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3467861.3467867","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:37:07Z","timestamp":1672223827000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3467861.3467867"}},"subtitle":["a comparative analysis of approaches"],"short-title":[],"issued":{"date-parts":[[2021,6]]},"references-count":117,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2021,6]]}},"alternative-id":["10.14778\/3467861.3467867"],"URL":"https:\/\/doi.org\/10.14778\/3467861.3467867","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,6]]}}}