{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T07:11:09Z","timestamp":1779174669663,"version":"3.51.4"},"reference-count":71,"publisher":"Association for Computing Machinery (ACM)","issue":"11","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2019,7]]},"abstract":"<jats:p>Machine learning (ML) pipelines for model training and validation typically include preprocessing, such as data cleaning and feature engineering, prior to training an ML model. Preprocessing combines relational algebra and user-defined functions (UDFs), while model training uses iterations and linear algebra. Current systems are tailored to either of the two. As a consequence, preprocessing and ML steps are optimized in isolation. To enable holistic optimization of ML training pipelines, we present Lara, a declarative domain-specific language for collections and matrices. Lara's inter-mediate representation (IR) reflects on the complete program, i.e., UDFs, control flow, and both data types. Two views on the IR enable diverse optimizations. Monads enable operator pushdown and fusion across type and loop boundaries. Combinators provide the semantics of domain-specific operators and optimize data access and cross-validation of ML algorithms. Our experiments on preprocessing pipelines and selected ML algorithms show the effects of our proposed optimizations on dense and sparse data, which achieve speedups of up to an order of magnitude.<\/jats:p>","DOI":"10.14778\/3342263.3342633","type":"journal-article","created":{"date-parts":[[2019,9,18]],"date-time":"2019-09-18T18:36:11Z","timestamp":1568831771000},"page":"1553-1567","source":"Crossref","is-referenced-by-count":36,"title":["An intermediate representation for optimizing machine learning pipelines"],"prefix":"10.14778","volume":"12","author":[{"given":"Andreas","family":"Kunft","sequence":"first","affiliation":[{"name":"TU Berlin"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Asterios","family":"Katsifodimos","sequence":"additional","affiliation":[{"name":"Delft University of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sebastian","family":"Schelter","sequence":"additional","affiliation":[{"name":"New York University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sebastian","family":"Bre\u00df","sequence":"additional","affiliation":[{"name":"DFKI and TU Berlin"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tilmann","family":"Rabl","sequence":"additional","affiliation":[{"name":"Universit\u00e4t Potsdam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Volker","family":"Markl","sequence":"additional","affiliation":[{"name":"DFKI"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,7]]},"reference":[{"key":"e_1_2_1_1_1","first-page":"265","volume-title":"12th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2016","author":"Abadi M.","year":"2016","unstructured":"M. Abadi , P. Barham , J. Chen , Z. Chen , A. Davis , J. Dean , M. Devin , S. Ghemawat , G. Irving , M. Isard , M. Kudlur , J. Levenberg , R. Monga , S. Moore , D. G. Murray , B. Steiner , P. A. Tucker , V. Vasudevan , P. Warden , M. Wicke , Y. Yu , and X. Zheng . Tensorflow: A system for large-scale machine learning . In 12th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2016 , Savannah, GA, USA , November 2-4, 2016 ., pages 265 -- 283 , 2016. M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. A. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng. Tensorflow: A system for large-scale machine learning. In 12th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2016, Savannah, GA, USA, November 2-4, 2016., pages 265--283, 2016."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-014-0357-y"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3281629"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2750543"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2899396"},{"key":"e_1_2_1_6_1","volume-title":"Morgan Kaufmann","author":"Allen R.","year":"2001","unstructured":"R. Allen and K. Kennedy . Optimizing Compilers for Modern Architectures: A Dependence-based Approach . Morgan Kaufmann , 2001 . R. Allen and K. Kennedy. Optimizing Compilers for Modern Architectures: A Dependence-based Approach. Morgan Kaufmann, 2001."},{"key":"e_1_2_1_7_1","unstructured":"Apache Hadoop http:\/\/hadoop.apache.org.  Apache Hadoop http:\/\/hadoop.apache.org."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/278283.278285"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098021"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/256095.256116"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.14778\/3007263.3007279"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/3229863.3229865"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732286.2732292"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137765.3137775"},{"issue":"3","key":"e_1_2_1_15_1","first-page":"52","article-title":"Systemml's optimizer: Plan generation for large-scale machine learning programs","volume":"37","author":"B\u00f6hm M.","year":"2014","unstructured":"M. B\u00f6hm , D. R. Burdick , A. V. Evfimievski , B. Reinwald , F. R. Reiss , P. Sen , S. Tatikonda , and Y. Tian . Systemml's optimizer: Plan generation for large-scale machine learning programs . IEEE Data Eng. Bull. , 37 ( 3 ): 52 -- 62 , 2014 . M. B\u00f6hm, D. R. Burdick, A. V. Evfimievski, B. Reinwald, F. R. Reiss, P. Sen, S. Tatikonda, and Y. Tian. Systemml's optimizer: Plan generation for large-scale machine learning programs. IEEE Data Eng. Bull., 37(3):52--62, 2014.","journal-title":"IEEE Data Eng. Bull."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-018-0512-y"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2489837.2489840"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1941553.1941561"},{"key":"e_1_2_1_19_1","volume-title":"VLDB'94","author":"Chaudhuri S.","year":"1994","unstructured":"S. Chaudhuri and K. Shim . Including group-by in query optimization . In VLDB'94 , Proceedings of 20th International Conference on Very Large Data Bases , September 12-15, 1994 , Santiago de Chile, Chile, pages 354--366, 1994. S. Chaudhuri and K. Shim. Including group-by in query optimization. In VLDB'94, Proceedings of 20th International Conference on Very Large Data Bases, September 12-15, 1994, Santiago de Chile, Chile, pages 354--366, 1994."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137628.3137633"},{"key":"e_1_2_1_21_1","volume-title":"Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems. CoRR, abs\/1512.01274","author":"Chen T.","year":"2015","unstructured":"T. Chen , M. Li , Y. Li , M. Lin , N. Wang , M. Wang , T. Xiao , B. Xu , C. Zhang , and Z. Zhang . Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems. CoRR, abs\/1512.01274 , 2015 . T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang. Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems. CoRR, abs\/1512.01274, 2015."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1291151.1291199"},{"key":"e_1_2_1_23_1","volume-title":"CIDR 2017, 8th Biennial Conference on Innovative Data Systems Research, Chaminade, CA, USA, January 8-11, 2017, Online Proceedings","author":"Elgamal T.","year":"2017","unstructured":"T. Elgamal , S. Luo , M. Boehm , A. V. Evfimievski , S. Tatikonda , B. Reinwald , and P. Sen . SPOOF: sum-product optimization and operator fusion for large-scale machine learning . In CIDR 2017, 8th Biennial Conference on Innovative Data Systems Research, Chaminade, CA, USA, January 8-11, 2017, Online Proceedings , 2017 . T. Elgamal, S. Luo, M. Boehm, A. V. Evfimievski, S. Tatikonda, B. Reinwald, and P. Sen. SPOOF: sum-product optimization and operator fusion for large-scale machine learning. In CIDR 2017, 8th Biennial Conference on Innovative Data Systems Research, Chaminade, CA, USA, January 8-11, 2017, Online Proceedings, 2017."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994515"},{"key":"e_1_2_1_25_1","volume-title":"Courier Corporation","author":"Eves H. W.","year":"1980","unstructured":"H. W. Eves . Elementary matrix theory . Courier Corporation , 1980 . H. W. Eves. Elementary matrix theory. Courier Corporation, 1980."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.14778\/2350229.2350245"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/377674.377676"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183760"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2011.5767930"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628136.2628138"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.5555\/645478.757691"},{"key":"e_1_2_1_32_1","volume-title":"Data science from scratch: first principles with python. \"O'Reilly Media","author":"Grus J.","year":"2015","unstructured":"J. Grus . Data science from scratch: first principles with python. \"O'Reilly Media , Inc .\", 2015 . J. Grus. Data science from scratch: first principles with python. \"O'Reilly Media, Inc.\", 2015."},{"key":"e_1_2_1_33_1","volume-title":"University of Konstanz","author":"Grust T.","year":"1999","unstructured":"T. Grust . Comprehending queries. PhD thesis , University of Konstanz , Germany , 1999 . T. Grust. Comprehending queries. PhD thesis, University of Konstanz, Germany, 1999."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1008705026446"},{"key":"e_1_2_1_35_1","volume-title":"Morgan Kaufmann","author":"Harris D.","year":"2010","unstructured":"D. Harris and S. Harris . Digital design and computer architecture . Morgan Kaufmann , 2010 . D. Harris and S. Harris. Digital design and computer architecture. Morgan Kaufmann, 2010."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/2350229.2350244"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133901"},{"key":"e_1_2_1_38_1","first-page":"1137","volume-title":"Proceedings of the Fourteenth International Joint Conference on Artificial Intelligence, IJCAI 95","author":"Kohavi R.","year":"1995","unstructured":"R. Kohavi . A study of cross-validation and bootstrap for accuracy estimation and model selection . In Proceedings of the Fourteenth International Joint Conference on Artificial Intelligence, IJCAI 95 , Montr\u00e9al Qu\u00e9bec, Canada , August 20-25 1995 , 2 Volumes, pages 1137 -- 1145 , 1995. R. Kohavi. A study of cross-validation and bootstrap for accuracy estimation and model selection. In Proceedings of the Fourteenth International Joint Conference on Artificial Intelligence, IJCAI 95, Montr\u00e9al Qu\u00e9bec, Canada, August 20-25 1995, 2 Volumes, pages 1137--1145, 1995."},{"key":"e_1_2_1_39_1","volume-title":"CIDR 2013, Sixth Biennial Conference on Innovative Data Systems Research, Asilomar, CA, USA, January 6-9, 2013, Online Proceedings","author":"Kraska T.","year":"2013","unstructured":"T. Kraska , A. Talwalkar , J. C. Duchi , R. Griffith , M. J. Franklin , and M. I. Jordan . Mlbase: A distributed machine-learning system . In CIDR 2013, Sixth Biennial Conference on Innovative Data Systems Research, Asilomar, CA, USA, January 6-9, 2013, Online Proceedings , 2013 . T. Kraska, A. Talwalkar, J. C. Duchi, R. Griffith, M. J. Franklin, and M. I. Jordan. Mlbase: A distributed machine-learning system. In CIDR 2013, Sixth Biennial Conference on Innovative Data Systems Research, Asilomar, CA, USA, January 6-9, 2013, Online Proceedings, 2013."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2935694.2935698"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2723713"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2926534.2926540"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.14778\/3151106.3151110"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/355841.355847"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.5555\/2787930"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.14778\/2350229.2350239"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3277006.3277013"},{"key":"e_1_2_1_48_1","volume-title":"Springer Science & Business Media","author":"Lane S. Mac","year":"2013","unstructured":"S. Mac Lane . Categories for the working mathematician, volume 5 . Springer Science & Business Media , 2013 . S. Mac Lane. Categories for the working mathematician, volume 5. Springer Science & Business Media, 2013."},{"key":"e_1_2_1_49_1","volume-title":"NumPy, and IPython. \"O'Reilly Media","author":"McKinney W.","year":"2012","unstructured":"W. McKinney . Python for data analysis : Data wrangling with Pandas , NumPy, and IPython. \"O'Reilly Media , Inc .\", 2012 . W. McKinney. Python for data analysis: Data wrangling with Pandas, NumPy, and IPython. \"O'Reilly Media, Inc.\", 2012."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2001269.2001285"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2847538.2847541"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.14778\/2002938.2002940"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.14778\/3213880.3213890"},{"key":"e_1_2_1_54_1","volume-title":"CIDR 2017, 8th Biennial Conference on Innovative Data Systems Research, Chaminade, CA, USA, January 8-11, 2017, Online Proceedings","author":"Palkar S.","year":"2017","unstructured":"S. Palkar , J. J. Thomas , A. Shanbhag , M. Schwarzkopf , S. P. Amarasinghe , and M. Zaharia . A common runtime for high performance data analysis . In CIDR 2017, 8th Biennial Conference on Innovative Data Systems Research, Chaminade, CA, USA, January 8-11, 2017, Online Proceedings , 2017 . S. Palkar, J. J. Thomas, A. Shanbhag, M. Schwarzkopf, S. P. Amarasinghe, and M. Zaharia. A common runtime for high performance data analysis. In CIDR 2017, 8th Biennial Conference on Innovative Data Systems Research, Chaminade, CA, USA, January 8-11, 2017, Online Proceedings, 2017."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3158101"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559845.1559865"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10990-013-9096-9"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/1868294.1868314"},{"issue":"4","key":"e_1_2_1_60_1","first-page":"5","article-title":"On challenges in machine learning model management","volume":"41","author":"Schelter S.","year":"2018","unstructured":"S. Schelter , F. Bie\u00dfmann , T. Januschowski , D. Salinas , S. Seufert , and G. Szarvas . On challenges in machine learning model management . IEEE Data Eng. Bull. , 41 ( 4 ): 5 -- 15 , 2018 . S. Schelter, F. Bie\u00dfmann, T. Januschowski, D. Salinas, S. Seufert, and G. Szarvas. On challenges in machine learning model management. IEEE Data Eng. Bull., 41(4):5--15, 2018.","journal-title":"IEEE Data Eng. Bull."},{"key":"e_1_2_1_61_1","volume-title":"Machine Learning Systems workshop at NeurIPS","author":"Schelter S.","year":"2016","unstructured":"S. Schelter , A. Palumbo , S. Quinn , S. Marthi , and A. Musselman . Samsara: Declarative machine learning on distributed dataflow systems . In Machine Learning Systems workshop at NeurIPS , 2016 . S. Schelter, A. Palumbo, S. Quinn, S. Marthi, and A. Musselman. Samsara: Declarative machine learning on distributed dataflow systems. In Machine Learning Systems workshop at NeurIPS, 2016."},{"key":"e_1_2_1_62_1","volume-title":"Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015","author":"Sculley D.","year":"2015","unstructured":"D. Sculley , G. Holt , D. Golovin , E. Davydov , T. Phillips , D. Ebner , V. Chaudhary , M. Young , J. Crespo , and D. Dennison . Hidden technical debt in machine learning systems . In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015 , December 7-12, 2015 , Montreal, Quebec, Canada, pages 2503--2511 , 2015. D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V. Chaudhary, M. Young, J. Crespo, and D. Dennison. Hidden technical debt in machine learning systems. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, pages 2503--2511, 2015."},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/2806777.2806945"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2017.109"},{"key":"e_1_2_1_65_1","first-page":"609","volume-title":"Proceedings of the 28th International Conference on Machine Learning, ICML 2011","author":"Sujeeth A. K.","year":"2011","unstructured":"A. K. Sujeeth , H. Lee , K. J. Brown , T. Rompf , H. Chafi , M. Wu , A. R. Atreya , M. Odersky , and K. Olukotun . Optiml: An implicitly parallel domain-specific language for machine learning . In Proceedings of the 28th International Conference on Machine Learning, ICML 2011 , Bellevue, Washington, USA, June 28 - July 2, 2011 , pages 609 -- 616 , 2011. A. K. Sujeeth, H. Lee, K. J. Brown, T. Rompf, H. Chafi, M. Wu, A. R. Atreya, M. Odersky, and K. Olukotun. Optiml: An implicitly parallel domain-specific language for machine learning. In Proceedings of the 28th International Conference on Machine Learning, ICML 2011, Bellevue, Washington, USA, June 28 - July 2, 2011, pages 609--616, 2011."},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196893"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.14778\/3275366.3284963"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.5555\/645387.651542"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/1553374.1553516"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/3190508.3190551"},{"key":"e_1_2_1_71_1","volume-title":"2nd USENIX Workshop on Hot Topics in Cloud Computing, HotCloud'10","author":"Zaharia M.","year":"2010","unstructured":"M. Zaharia , M. Chowdhury , M. J. Franklin , S. Shenker , and I. Stoica . Spark: Cluster computing with working sets . In 2nd USENIX Workshop on Hot Topics in Cloud Computing, HotCloud'10 , Boston, MA, USA , June 22, 2010 , 2010. M. Zaharia, M. Chowdhury, M. J. Franklin, S. Shenker, and I. Stoica. Spark: Cluster computing with working sets. In 2nd USENIX Workshop on Hot Topics in Cloud Computing, HotCloud'10, Boston, MA, USA, June 22, 2010, 2010."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3342263.3342633","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T09:55:48Z","timestamp":1672221348000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3342263.3342633"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,7]]},"references-count":71,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2019,7]]}},"alternative-id":["10.14778\/3342263.3342633"],"URL":"https:\/\/doi.org\/10.14778\/3342263.3342633","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2019,7]]}}}