{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T18:44:44Z","timestamp":1774982684848,"version":"3.50.1"},"reference-count":31,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2024,12,18]],"date-time":"2024-12-18T00:00:00Z","timestamp":1734480000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"NSF","award":["IIS-2107150"],"award-info":[{"award-number":["IIS-2107150"]}]},{"name":"NIH","award":["2U24DK097771-11"],"award-info":[{"award-number":["2U24DK097771-11"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2024,12,18]]},"abstract":"<jats:p>Data analytics tasks are often formulated as data workflows represented as directed acyclic graphs (DAGs) of operators. The recent trend of adopting machine learning (ML) techniques in workflows results in increasingly complicated DAGs with many operators and edges. Compared to the operator-at-a-time execution paradigm, pipelined execution has benefits of reducing the materialization cost of intermediate results and allowing operators to produce results early, which are critical in iterative analysis on large data volumes. Correctly scheduling a workflow DAG for pipelined execution is non-trivial due to the richer semantics of operators and the increasing complexity of DAGs. Several existing data systems adopt simple heuristics to solve the problem without considering costs such as materialization sizes. In this paper, we systematically study the problem of scheduling a workflow DAG for pipelined execution, and develop a novel cost-based optimizer called Pasta for generating a high-quality schedule. The Pasta optimizer is not only general and applicable to a wide variety of cost functions, but also capable of utilizing properties inherent in a broad class of cost functions to improve its performance significantly. We conducted a thorough evaluation of developed techniques on real-world workflows and show the efficiency and efficacy of these solutions.<\/jats:p>","DOI":"10.1145\/3698832","type":"journal-article","created":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T16:40:35Z","timestamp":1734712835000},"page":"1-26","source":"Crossref","is-referenced-by-count":1,"title":["Pasta: A Cost-Based Optimizer for Generating Pipelining Schedules for Dataflow DAGs"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-5346-7028","authenticated-orcid":false,"given":"Xiaozhen","family":"Liu","sequence":"first","affiliation":[{"name":"University of California, Irvine, Irvine, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1186-4803","authenticated-orcid":false,"given":"Yicong","family":"Huang","sequence":"additional","affiliation":[{"name":"University of California, Irvine, Irvine, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7935-0035","authenticated-orcid":false,"given":"Xinyuan","family":"Lin","sequence":"additional","affiliation":[{"name":"University of California, Irvine, Irvine, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-9327-3906","authenticated-orcid":false,"given":"Avinash","family":"Kumar","sequence":"additional","affiliation":[{"name":"University of California, Irvine, Irvine, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3928-690X","authenticated-orcid":false,"given":"Sadeem","family":"Alsudais","sequence":"additional","affiliation":[{"name":"King Saud University, Riyadh, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8015-6870","authenticated-orcid":false,"given":"Chen","family":"Li","sequence":"additional","affiliation":[{"name":"University of California, Irvine, Irvine, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,12,20]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"Alteryx 2024. AI Analytics Platform - Alteryx https:\/\/www.alteryx.com\/."},{"key":"e_1_2_2_2_1","unstructured":"Apache Flink 2024. Apache Flink\u00ae - Stateful Computations over Data Streams | Apache Flink https:\/\/flink.apache.org."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2011.5767921"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/645922.673330"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/375551.375561"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3526186"},{"key":"e_1_2_2_7_1","unstructured":"Docker Swarm 2024. Swarm mode | Docker Docs https:\/\/docs.docker.com\/engine\/swarm\/."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE55515.2023.00276"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE51399.2021.00127"},{"key":"e_1_2_2_10_1","volume-title":"GRAPHENE: Packing and Dependency-Aware Scheduling for Data-Parallel Clusters. In 12th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2016","author":"Grandl Robert","year":"2016","unstructured":"Robert Grandl, Srikanth Kandula, Sriram Rao, Aditya Akella, and Janardhan Kulkarni. 2016. GRAPHENE: Packing and Dependency-Aware Scheduling for Data-Parallel Clusters. In 12th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2016, Savannah, GA, USA, November 2--4, 2016, Kimberly Keeton and Timothy Roscoe (Eds.). USENIX Association, 81--97. https:\/\/www.usenix.org\/conference\/osdi16\/technical-sessions\/presentation\/grandl_graphene"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/645920.672835"},{"key":"e_1_2_2_12_1","volume-title":"Proceedings of the 8th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2011","author":"Hindman Benjamin","year":"2011","unstructured":"Benjamin Hindman, Andy Konwinski, Matei Zaharia, Ali Ghodsi, Anthony D. Joseph, Randy H. Katz, Scott Shenker, and Ion Stoica. 2011. Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center. In Proceedings of the 8th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2011, Boston, MA, USA, March 30 - April 1, 2011, David G. Andersen and Sylvia Ratnasamy (Eds.). USENIX Association. https:\/\/www.usenix.org\/conference\/nsdi11\/mesos-platform-fine-grained-resource-sharing-data-center"},{"key":"e_1_2_2_13_1","unstructured":"KNIME 2024. Open for Innovation | KNIME https:\/\/www.knime.com\/."},{"key":"e_1_2_2_14_1","unstructured":"KNIME Community Workflows 2024. Workflows | KNIME Community Hub https:\/\/hub.knime.com\/search?type=Workflow."},{"key":"e_1_2_2_15_1","unstructured":"Kubernetes 2024. Kubernetes https:\/\/kubernetes.io\/."},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/3377369.3377381"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/S00453-007--9004-Y"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.14778\/3554821.3554888"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/170035.170053"},{"key":"e_1_2_2_20_1","unstructured":"Pipelined Regions in Apache Flink 2020. Improvements in task scheduling for batch workloads in Apache Flink 1.12 https:\/\/flink.apache.org\/2020\/12\/02\/improvements-in-task-scheduling-for-batch-workloads-in-apache-flink-1.12\/."},{"key":"e_1_2_2_21_1","volume-title":"Database Management Systems","author":"Ramakrishnan Raghu","unstructured":"Raghu Ramakrishnan and Johannes Gehrke. 2002. Database Management Systems. WCB\/McGraw-Hill."},{"key":"e_1_2_2_22_1","unstructured":"RapidMiner 2024. Data Analytics and AI Platform | Altair RapidMiner https:\/\/rapidminer.com\/."},{"key":"e_1_2_2_23_1","unstructured":"Vladislav Shkapenyuk Ryan Williams Stavros Harizopoulos and Anastassia Ailamaki. 2005. Deadlock resolution in pipelined query graphs. Carnegie Mellon University Technical Report."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2017.109"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2523616.2523633"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.14778\/3681954.3682022"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.14778\/3611540.3611580"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS47924.2020.00047"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.14778\/3547305.3547310"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2854006.2854012"},{"key":"e_1_2_2_31_1","volume-title":"Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2012","author":"Zaharia Matei","year":"2012","unstructured":"Matei Zaharia, Mosharaf Chowdhury, Tathagata Das, Ankur Dave, Justin Ma, Murphy McCauley, Michael J. Franklin, Scott Shenker, and Ion Stoica. 2012. Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing. In Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2012, San Jose, CA, USA, April 25--27, 2012. 15--28."}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3698832","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3698832","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T17:45:13Z","timestamp":1774979113000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3698832"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,18]]},"references-count":31,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2024,12,18]]}},"alternative-id":["10.1145\/3698832"],"URL":"https:\/\/doi.org\/10.1145\/3698832","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,18]]}}}