{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,5,7]],"date-time":"2024-05-07T22:59:21Z","timestamp":1715122761377},"reference-count":40,"publisher":"Association for Computing Machinery (ACM)","issue":"13","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2017,9]]},"abstract":"<jats:p>\n            Linear algebra operations are at the core of many Machine Learning (ML) programs. At the same time, a considerable amount of the effort for solving data analytics problems is spent in data preparation. As a result, end-to-end ML pipelines often consist of (\n            <jats:italic>i<\/jats:italic>\n            ) relational operators used for joining the input data, (\n            <jats:italic>ii<\/jats:italic>\n            ) user defined functions used for feature extraction and vectorization, and (\n            <jats:italic>iii<\/jats:italic>\n            ) linear algebra operators used for model training and cross-validation. Often, these pipelines need to scale out to large datasets. In this case, these pipelines are usually implemented on top of dataflow engines like Hadoop, Spark, or Flink. These dataflow engines implement relational operators on\n            <jats:italic>row-partitioned<\/jats:italic>\n            datasets. However, efficient linear algebra operators use\n            <jats:italic>block-partitioned<\/jats:italic>\n            matrices. As a result, pipelines combining both kinds of operators require rather expensive changes to the physical representation, in particular re-partitioning steps. In this paper, we investigate the potential of reducing shuffling costs by fusing relational and linear algebra operations into specialized physical operators. We present\n            <jats:italic>BlockJoin<\/jats:italic>\n            , a distributed join algorithm which directly produces block-partitioned results. To minimize shuffling costs, BlockJoin applies database techniques known from columnar processing, such as index-joins and late materialization, in the context of parallel dataflow engines. Our experimental evaluation shows speedups up to 6\u00d7 and the skew resistance of BlockJoin compared to state-of-the-art pipelines implemented in Spark.\n          <\/jats:p>","DOI":"10.14778\/3151106.3151110","type":"journal-article","created":{"date-parts":[[2017,10,19]],"date-time":"2017-10-19T12:30:08Z","timestamp":1508416208000},"page":"2061-2072","source":"Crossref","is-referenced-by-count":10,"title":["Blockjoin"],"prefix":"10.14778","volume":"10","author":[{"given":"Andreas","family":"Kunft","sequence":"first","affiliation":[{"name":"Technische Universit\u00e4t Berlin"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Asterios","family":"Katsifodimos","sequence":"additional","affiliation":[{"name":"Delft University of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sebastian","family":"Schelter","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Berlin"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tilmann","family":"Rabl","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Berlin and German Research Center for Artificial Intelligence (DFKI)"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Volker","family":"Markl","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Berlin and German Research Center for Artificial Intelligence (DFKI)"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,9]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376712"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1739041.1739056"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.14778\/2336664.2336678"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2750543"},{"key":"e_1_2_1_5_1","unstructured":"Apache Hadoop http:\/\/hadoop.apache.org.  Apache Hadoop http:\/\/hadoop.apache.org."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/276305.276386"},{"issue":"3","key":"e_1_2_1_7_1","first-page":"52","article-title":"SystemML's optimizer: Plan generation for large-scale machine learning programs","volume":"37","author":"Boehm M.","year":"2014","unstructured":"M. Boehm SystemML's optimizer: Plan generation for large-scale machine learning programs . IEEE Data Eng. Bull. , 37 ( 3 ): 52 -- 62 , 2014 . M. Boehm et al. SystemML's optimizer: Plan generation for large-scale machine learning programs. IEEE Data Eng. Bull., 37(3):52--62, 2014.","journal-title":"IEEE Data Eng. Bull."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.14778\/3007263.3007279"},{"key":"e_1_2_1_9_1","volume-title":"VLDB","author":"Boncz P. A.","year":"1999","unstructured":"P. A. Boncz , S. Manegold , M. L. Kersten , Database architecture optimized for the new bottleneck: Memory access . In VLDB , 1999 . P. A. Boncz, S. Manegold, M. L. Kersten, et al. Database architecture optimized for the new bottleneck: Memory access. In VLDB, 1999."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939675"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807167.1807271"},{"key":"e_1_2_1_12_1","volume-title":"VLDB","author":"Chaudhuri S.","year":"1994","unstructured":"S. Chaudhuri and K. Shim . Including group-by in query optimization . In VLDB , 1994 . S. Chaudhuri and K. Shim. Including group-by in query optimization. In VLDB, 1994."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/FMPC.1992.234898"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687553.1687576"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/971697.602261"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994515"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2011.5767930"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463676.2465273"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2009.14"},{"key":"e_1_2_1_20_1","first-page":"2","volume-title":"CIDR","volume":"1","author":"Kraska T.","year":"2013","unstructured":"T. Kraska : A distributed machine-learning system . In CIDR , volume 1 , pages 2 -- 1 , 2013 . T. Kraska et al. Mlbase: A distributed machine-learning system. In CIDR, volume 1, pages 2--1, 2013."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2723713"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2926534.2926540"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1081870.1081893"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/s007780050071"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/191246.191258"},{"issue":"34","key":"e_1_2_1_26_1","first-page":"1","article-title":"Mllib: Machine learning in apache spark","volume":"17","author":"Meng X.","year":"2016","unstructured":"X. Meng Mllib: Machine learning in apache spark . JMLR , 17 ( 34 ): 1 -- 7 , 2016 . X. Meng et al. Mllib: Machine learning in apache spark. JMLR, 17(34):1--7, 2016.","journal-title":"JMLR"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/32.52778"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1989323.1989423"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2588555.2610521"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2016.7498324"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/69.109109"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2015.7113367"},{"key":"e_1_2_1_33_1","volume-title":"NIPS Workshop MLSystems","author":"Schelter S.","year":"2016","unstructured":"S. Schelter : Declarative machine learning on distributed dataflow systems . In NIPS Workshop MLSystems , 2016 . S. Schelter et al. Samsara: Declarative machine learning on distributed dataflow systems. In NIPS Workshop MLSystems, 2016."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2013.158"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/71.250116"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559845.1559854"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2038916.2038928"},{"key":"e_1_2_1_38_1","volume-title":"Spark: Cluster computing with working sets. HotCloud, 10(10--10):95","author":"Zaharia M.","year":"2010","unstructured":"M. Zaharia Spark: Cluster computing with working sets. HotCloud, 10(10--10):95 , 2010 . M. Zaharia et al. Spark: Cluster computing with working sets. HotCloud, 10(10--10):95, 2010."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2877204"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2010.5447802"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3151106.3151110","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:09:03Z","timestamp":1672225743000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3151106.3151110"}},"subtitle":["efficient matrix partitioning through joins"],"short-title":[],"issued":{"date-parts":[[2017,9]]},"references-count":40,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2017,9]]}},"alternative-id":["10.14778\/3151106.3151110"],"URL":"https:\/\/doi.org\/10.14778\/3151106.3151110","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2017,9]]}}}