{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T07:12:42Z","timestamp":1779174762756,"version":"3.51.4"},"reference-count":14,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2023,8]]},"abstract":"<jats:p>\n            Software systems that learn from data with machine learning (ML) are used in critical decision-making processes. Unfortunately, real-world experience shows that the pipelines for data preparation, feature encoding and model training in ML systems are often brittle with respect to their input data. As a consequence, data scientists have to run different kinds of\n            <jats:italic toggle=\"yes\">data centric what-if analyses<\/jats:italic>\n            to evaluate the robustness and reliability of such pipelines, e.g., with respect to data errors or preprocessing techniques. These what-if analyses follow a common pattern: they take an existing ML pipeline, create a pipeline variant by introducing a small change, and execute this variant to see how the change impacts the pipeline's output score.\n          <\/jats:p>\n          <jats:p>We recently proposed mlwhatif, a library that enables data scientists to declaratively specify what-if analyses for an ML pipeline, and to automatically generate, optimize and execute the required pipeline variants. We demonstrate how data scientists can leverage mlwhatif for a variety of pipelines and three different what-if analyses focusing on the robustness of a pipeline against data errors, the impact of data cleaning operations, and the impact of data preprocessing operations on fairness. In particular, we demonstrate step-by-step how mlwhatif generates and optimizes the required execution plans for the pipeline analyses. Our library is publicly available at https:\/\/github.com\/stefan-grafberger\/mlwhatif.<\/jats:p>","DOI":"10.14778\/3611540.3611606","type":"journal-article","created":{"date-parts":[[2023,9,15]],"date-time":"2023-09-15T11:32:37Z","timestamp":1694777557000},"page":"4002-4005","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["mlwhatif: What If You Could Stop Re-Implementing Your Machine Learning Pipeline Analyses over and over?"],"prefix":"10.14778","volume":"16","author":[{"given":"Stefan","family":"Grafberger","sequence":"first","affiliation":[{"name":"AIRLab, University of Amsterdam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shubha","family":"Guha","sequence":"additional","affiliation":[{"name":"University of Amsterdam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paul","family":"Groth","sequence":"additional","affiliation":[{"name":"University of Amsterdam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sebastian","family":"Schelter","sequence":"additional","affiliation":[{"name":"University of Amsterdam"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,8]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipeline. ESEC\/FSE","author":"Biswas Sumon","year":"2021","unstructured":"Sumon Biswas and Hridesh Rajan. 2021. Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipeline. ESEC\/FSE (2021)."},{"key":"e_1_2_1_2_1","volume-title":"Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines. SIGMOD","author":"Grafberger Stefan","year":"2023","unstructured":"Stefan Grafberger, Paul Groth, and Sebastian Schelter. 2023. Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines. SIGMOD (2023)."},{"key":"e_1_2_1_3_1","volume-title":"Data distribution debugging in machine learning pipelines. VLDBJ","author":"Grafberger Stefan","year":"2022","unstructured":"Stefan Grafberger, Paul Groth, Julia Stoyanovich, and Sebastian Schelter. 2022. Data distribution debugging in machine learning pipelines. VLDBJ (2022)."},{"key":"e_1_2_1_4_1","volume-title":"MLINSPECT: A Data Distribution Debugger for Machine Learning Pipelines. SIGMOD","author":"Grafberger Stefan","year":"2021","unstructured":"Stefan Grafberger, Shubha Guha, Julia Stoyanovich, and Sebastian Schelter. 2021. MLINSPECT: A Data Distribution Debugger for Machine Learning Pipelines. SIGMOD (2021)."},{"key":"e_1_2_1_5_1","volume-title":"Julia Stoyanovich, and Sebastian Schelter.","author":"Guha Shubha","year":"2023","unstructured":"Shubha Guha, Falaah Arif Khan, Julia Stoyanovich, and Sebastian Schelter. 2023. Automated Data Cleaning Can Hurt Fairness in Machine Learning-based Decision Making. ICDE (2023)."},{"key":"e_1_2_1_6_1","volume-title":"Nezihe Merve Gurel, Bo Li, Ce Zhang, Costas J Spanos, and Dawn Song.","author":"Jia Ruoxi","year":"2019","unstructured":"Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nezihe Merve Gurel, Bo Li, Ce Zhang, Costas J Spanos, and Dawn Song. 2019. Efficient task-specific data valuation for nearest neighbor algorithms. VLDB (2019)."},{"key":"e_1_2_1_7_1","volume-title":"Cleanml: A benchmark for joint data cleaning and machine learning. ICDE","author":"Li Peng","year":"2019","unstructured":"Peng Li, Xi Rao, Jennifer Blase, Yue Zhang, Xu Chu, and Ce Zhang. 2019. Cleanml: A benchmark for joint data cleaning and machine learning. ICDE (2019)."},{"key":"e_1_2_1_8_1","volume-title":"Confident learning: Estimating uncertainty in dataset labels. JAIR 70","author":"Northcutt Curtis","year":"2021","unstructured":"Curtis Northcutt, Lu Jiang, and Isaac Chuang. 2021. Confident learning: Estimating uncertainty in dataset labels. JAIR 70 (2021)."},{"key":"e_1_2_1_9_1","volume-title":"Steven Euijong Whang, and Martin Zinkevich","author":"Polyzotis Neoklis","year":"2018","unstructured":"Neoklis Polyzotis, Sudip Roy, Steven Euijong Whang, and Martin Zinkevich. 2018. Data lifecycle challenges in production machine learning: a survey. SIGMOD Record 47, 2 (2018)."},{"key":"e_1_2_1_10_1","volume-title":"On challenges in machine learning model management","author":"Schelter Sebastian","year":"2018","unstructured":"Sebastian Schelter, Felix Biessmann, Tim Januschowski, David Salinas, Stephan Seufert, and Gyuri Szarvas. 2018. On challenges in machine learning model management. IEEE Data Engineering Bulletin (2018)."},{"key":"e_1_2_1_11_1","volume-title":"Proactively Screening Machine Learning Pipelines with ArgusEyes. SIGMOD","author":"Schelter Sebastian","year":"2023","unstructured":"Sebastian Schelter, Stefan Grafberger, Shubha Guha, Bojan Karla\u0161, and Ce Zhang. 2023. Proactively Screening Machine Learning Pipelines with ArgusEyes. SIGMOD (2023)."},{"key":"e_1_2_1_12_1","volume-title":"Screening Native ML Pipelines with \"ArgusEyes\". CIDR","author":"Schelter Sebastian","year":"2022","unstructured":"Sebastian Schelter, Stefan Grafberger, Shubha Guha, Olivier Sprangers, Bojan Karla\u0161, and Ce Zhang. 2022. Screening Native ML Pipelines with \"ArgusEyes\". CIDR (2022)."},{"key":"e_1_2_1_13_1","volume-title":"JENGA - A Framework to Study the Impact of Data Errors on the Predictions of Machine Learning Models. EDBT","author":"Schelter Sebastian","year":"2021","unstructured":"Sebastian Schelter, Tammo Rukat, and Felix Biessmann. 2021. JENGA - A Framework to Study the Impact of Data Errors on the Predictions of Machine Learning Models. EDBT (2021)."},{"key":"e_1_2_1_14_1","volume-title":"Responsible Data Management. Commun. ACM","author":"Stoyanovich Julia","year":"2022","unstructured":"Julia Stoyanovich, Bill Howe, Serge Abiteboul, H.V. Jagadish, and Sebastian Schelter. 2022. Responsible Data Management. Commun. ACM (2022)."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3611540.3611606","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,10]],"date-time":"2025-09-10T22:33:08Z","timestamp":1757543588000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3611540.3611606"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8]]},"references-count":14,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2023,8]]}},"alternative-id":["10.14778\/3611540.3611606"],"URL":"https:\/\/doi.org\/10.14778\/3611540.3611606","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2023,8]]},"assertion":[{"value":"2023-08-01","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}