{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,1,10]],"date-time":"2023-01-10T10:53:15Z","timestamp":1673347995124},"reference-count":16,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2019,8]]},"abstract":"<jats:p>Running machine learning (ML) workloads at scale is as much a data management problem as a model engineering problem. Big performance challenges exist when data management systems invoke ML classifiers as user-defined functions (UDFs) or when stand-alone ML frameworks interact with data stores for data loading and pre-processing (ETL). In particular, UDFs can be precompiled or simply a black box for the data management system and the data layout may be completely different from the native layout, thus adding overheads at the boundaries. In this demo, we will show how bottlenecks between existing systems can be eliminated when their engines are designed around runtime compilation and native code generation, which is the case for many state-of-the-art relational engines as well as ML frameworks. We demonstrate an integration of Flare (an accelerator for Spark SQL), and Lantern (an accelerator for TensorFlow and PyTorch) that results in a highly optimized end-to-end compiled data path, switching between SQL and ML processing with negligible overhead.<\/jats:p>","DOI":"10.14778\/3352063.3352097","type":"journal-article","created":{"date-parts":[[2019,9,18]],"date-time":"2019-09-18T18:36:11Z","timestamp":1568831771000},"page":"1910-1913","source":"Crossref","is-referenced-by-count":3,"title":["Flare &amp; lantern"],"prefix":"10.14778","volume":"12","author":[{"given":"Gr\u00e9gory","family":"Essertel","sequence":"first","affiliation":[{"name":"Purdue University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ruby Y.","family":"Tahboub","sequence":"additional","affiliation":[{"name":"Purdue University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fei","family":"Wang","sequence":"additional","affiliation":[{"name":"Purdue University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"James","family":"Decker","sequence":"additional","affiliation":[{"name":"Purdue University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tiark","family":"Rompf","sequence":"additional","affiliation":[{"name":"Purdue University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,8]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"M. Abadi A. Agarwal P. Barham E. Brevdo Z. Chen C. Citro G. Corrado A. Davis J. Dean M. Devin S. Ghemawat I. Goodfellow A. Harp G. Irving M. Isard Y. Jia R. Jozefowicz L. Kaiser M. Kudlur J. Levenberg D. Man\u00e9 R. Monga S. Moore D. Murray C. Olah M. Schuster J. Shlens B. Steiner I. Sutskever K. Talwar P. Tucker V. Vanhoucke V. Vasudevan F. Vi\u00e9gas O. Vinyals P. Warden M. Wattenberg M. Wicke Y. Yu and X. Zheng. TensorFlow: Large-scale machine learning on heterogeneous distributed systems 2015.  M. Abadi A. Agarwal P. Barham E. Brevdo Z. Chen C. Citro G. Corrado A. Davis J. Dean M. Devin S. Ghemawat I. Goodfellow A. Harp G. Irving M. Isard Y. Jia R. Jozefowicz L. Kaiser M. Kudlur J. Levenberg D. Man\u00e9 R. Monga S. Moore D. Murray C. Olah M. Schuster J. Shlens B. Steiner I. Sutskever K. Talwar P. Tucker V. Vanhoucke V. Vasudevan F. Vi\u00e9gas O. Vinyals P. Warden M. Wattenberg M. Wicke Y. Yu and X. Zheng. TensorFlow: Large-scale machine learning on heterogeneous distributed systems 2015."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2742797"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2854038.2854042"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_5_1","first-page":"799","volume-title":"OSDI","author":"Essertel G. M.","year":"2018"},{"key":"e_1_2_1_6_1","article-title":"Partial evaluation of computation process --- an approach to a compiler-compiler","author":"Futamura Y.","year":"1971","journal-title":"Transactions of the Institute of Electronics and Communication Engineers of Japan, 54-C(8):721--728"},{"key":"e_1_2_1_7_1","unstructured":"J. Leskovec and A. Krevl. SNAP Datasets: Stanford large network dataset collection. http:\/\/snap.stanford.edu\/data June 2014.  J. Leskovec and A. Krevl. SNAP Datasets: Stanford large network dataset collection. http:\/\/snap.stanford.edu\/data June 2014."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.14778\/2002938.2002940"},{"key":"e_1_2_1_9_1","volume-title":"CIDR","author":"Palkar S.","year":"2017"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2184319.2184345"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2584665"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196893"},{"key":"e_1_2_1_13_1","unstructured":"The Transaction Processing Council. TPC-H Version 2.15.0.  The Transaction Processing Council. TPC-H Version 2.15.0."},{"key":"e_1_2_1_14_1","first-page":"10201","volume-title":"NeurIPS","author":"Wang F.","year":"2018"},{"key":"e_1_2_1_15_1","volume-title":"ICLR Workshop Track","author":"Wang F.","year":"2018"},{"key":"e_1_2_1_16_1","unstructured":"F. Wang X. Wu G. M. Essertel J. M. Decker and T. Rompf. Demystifying differentiable programming: Shift\/reset the penultimate backpropagator. CoRR abs\/1803.10228 2018.  F. Wang X. Wu G. M. Essertel J. M. Decker and T. Rompf. Demystifying differentiable programming: Shift\/reset the penultimate backpropagator. CoRR abs\/1803.10228 2018."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3352063.3352097","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:33:04Z","timestamp":1672223584000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3352063.3352097"}},"subtitle":["efficiently swapping horses midstream"],"short-title":[],"issued":{"date-parts":[[2019,8]]},"references-count":16,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2019,8]]}},"alternative-id":["10.14778\/3352063.3352097"],"URL":"https:\/\/doi.org\/10.14778\/3352063.3352097","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2019,8]]}}}