{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,17]],"date-time":"2025-09-17T03:16:30Z","timestamp":1758078990595,"version":"3.44.0"},"reference-count":22,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2025,8]]},"abstract":"<jats:p>\n            Data scientists develop ML pipelines in an iterative manner: they repeatedly screen a pipeline for potential issues, debug it, and then revise and improve its code according to their findings. However, this manual process is tedious and error-prone. To address this challenge, we propose to assist data scientists with automatically derived\n            <jats:italic toggle=\"yes\">interactive suggestions for pipeline improvements<\/jats:italic>\n            during this development cycle. We demonstrate mlidea, a library to generate interactive suggestions with so-called\n            <jats:italic toggle=\"yes\">shadow pipelines<\/jats:italic>\n            , hidden variants of the original pipeline that modify it to auto-detect potential issues, try out modifications for improvements, and suggest and explain these modifications to the user. Our system uses incremental view maintenance to enable data scientists to quickly iterate on their code and to ensure low-latency maintenance of the shadow pipelines. We demonstrate how our system improves code for various domains with three interactive shadow pipelines: fixing mislabeled rows, enhancing robustness against data quality problems, and improving pipeline performance on data slices with subpar predictions.\n          <\/jats:p>","DOI":"10.14778\/3750601.3750671","type":"journal-article","created":{"date-parts":[[2025,9,16]],"date-time":"2025-09-16T13:38:05Z","timestamp":1758029885000},"page":"5359-5362","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["mlidea: Interactively Improving ML Data Preparation Code via \"Shadow Pipelines\""],"prefix":"10.14778","volume":"18","author":[{"given":"Stefan","family":"Grafberger","sequence":"first","affiliation":[{"name":"BIFOLD &amp; TU Berlin"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paul","family":"Groth","sequence":"additional","affiliation":[{"name":"University of Amsterdam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sebastian","family":"Schelter","sequence":"additional","affiliation":[{"name":"BIFOLD &amp; TU Berlin"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,16]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Data Validation for Machine Learning. MLSys","author":"Breck Eric","year":"2019","unstructured":"Eric Breck, Neoklis Polyzotis, Sudip Roy, Steven Whang, and Martin Zinkevich. 2019. Data Validation for Machine Learning. MLSys (2019)."},{"key":"e_1_2_1_2_1","doi-asserted-by":"crossref","unstructured":"Zaheer Chothia John Liagouris Frank McSherry and Timothy Roscoe. 2016. Explaining outputs in modern data analytics. Technical Report. ETH Zurich.","DOI":"10.14778\/2994509.2994530"},{"key":"e_1_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Yeounoh Chung et al. 2019. Slice finder: Automated data slicing for model validation. ICDE (2019).","DOI":"10.1109\/ICDE.2019.00139"},{"key":"e_1_2_1_4_1","unstructured":"GitHub. 2021. GitHub Copilot \u00b7 Your AI pair programmer. https:\/\/copilot.github.com\/."},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Stefan Grafberger et al. 2023. Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines. SIGMOD (2023).","DOI":"10.1145\/3589273"},{"key":"e_1_2_1_6_1","volume-title":"DEEM workshop @ SIGMOD","author":"Stefan","year":"2024","unstructured":"Stefan Grafberger et al. 2024. Towards Interactively Improving ML Data Preparation Code via\" Shadow Pipelines\". DEEM workshop @ SIGMOD (2024)."},{"key":"e_1_2_1_7_1","volume-title":"Data distribution debugging in machine learning pipelines. VLDBJ","author":"Grafberger Stefan","year":"2022","unstructured":"Stefan Grafberger, Paul Groth, Julia Stoyanovich, and Sebastian Schelter. 2022. Data distribution debugging in machine learning pipelines. VLDBJ (2022)."},{"key":"e_1_2_1_8_1","volume-title":"Towards Declarative Systems for Data-Centric ML. DMLR workshop @ ICML","author":"Grafberger Stefan","year":"2023","unstructured":"Stefan Grafberger, Bojan Karla\u0161, Paul Groth, and Sebastian Schelter. 2023. Towards Declarative Systems for Data-Centric ML. DMLR workshop @ ICML (2023)."},{"key":"e_1_2_1_9_1","unstructured":"Grammarly. [n.d.]. Demo. https:\/\/demo.grammarly.com\/."},{"key":"e_1_2_1_10_1","doi-asserted-by":"crossref","unstructured":"Kenneth Holstein et al. 2019. Improving fairness in machine learning systems: What do industry practitioners need? CHI (2019).","DOI":"10.1145\/3290605.3300830"},{"key":"e_1_2_1_11_1","unstructured":"Jetbrains. [n.d.]. Code inspections. https:\/\/www.jetbrains.com\/help\/idea\/code-inspection.html#access-inspections-and-settings."},{"key":"e_1_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Ruoxi Jia et al. 2019. Efficient task-specific data valuation for nearest neighbor algorithms. VLDB (2019).","DOI":"10.14778\/3342263.3342637"},{"key":"e_1_2_1_13_1","unstructured":"Bojan Karla\u0161 et al. 2023. Data Debugging with Shapley Importance over Machine Learning Pipelines. ICLR (2023)."},{"key":"e_1_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Sanjay Krishnan et al. 2016. ActiveClean: interactive data cleaning for statistical modeling. VLDB (2016).","DOI":"10.1145\/2882903.2899409"},{"key":"e_1_2_1_15_1","volume-title":"Rebecca Isaacs, and Michael Isard.","author":"McSherry Frank","year":"2013","unstructured":"Frank McSherry, Derek Gordon Murray, Rebecca Isaacs, and Michael Isard. 2013. Differential Dataflow. CIDR (2013)."},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Neoklis Polyzotis et al. 2018. Data lifecycle challenges in production machine learning: a survey. SIGMOD Record (2018).","DOI":"10.1145\/3299887.3299891"},{"key":"e_1_2_1_17_1","volume-title":"Interpretable data-based explanations for fairness debugging. SIGMOD","author":"Pradhan Romila","year":"2022","unstructured":"Romila Pradhan, Jiongli Zhu, Boris Glavic, and Babak Salimi. 2022. Interpretable data-based explanations for fairness debugging. SIGMOD (2022)."},{"key":"e_1_2_1_18_1","volume-title":"Sliceline: Fast, linear-algebra-based slice finding for ml model debugging. SIGMOD","author":"Sagadeeva Svetlana","year":"2021","unstructured":"Svetlana Sagadeeva and Matthias Boehm. 2021. Sliceline: Fast, linear-algebra-based slice finding for ml model debugging. SIGMOD (2021)."},{"key":"e_1_2_1_19_1","unstructured":"Sebastian Schelter et al. 2018. On challenges in machine learning model management. IEEE Data Engineering Bulletin (2018)."},{"key":"e_1_2_1_20_1","unstructured":"Sebastian Schelter et al. 2021. JENGA - A Framework to Study the Impact of Data Errors on the Predictions of Machine Learning Models. EDBT (2021)."},{"key":"e_1_2_1_21_1","volume-title":"Proactively Screening ML Pipelines with ArgusEyes. SIGMOD","author":"Schelter Sebastian","year":"2023","unstructured":"Sebastian Schelter, Stefan Grafberger, Shubha Guha, Bojan Karla\u0161, and Ce Zhang. 2023. Proactively Screening ML Pipelines with ArgusEyes. SIGMOD (2023)."},{"key":"e_1_2_1_22_1","volume-title":"Responsible Data Management. Commun. ACM","author":"Stoyanovich Julia","year":"2022","unstructured":"Julia Stoyanovich, Bill Howe, Serge Abiteboul, H.V. Jagadish, and Sebastian Schelter. 2022. Responsible Data Management. Commun. ACM (2022)."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3750601.3750671","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,16]],"date-time":"2025-09-16T13:42:39Z","timestamp":1758030159000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3750601.3750671"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8]]},"references-count":22,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2025,8]]}},"alternative-id":["10.14778\/3750601.3750671"],"URL":"https:\/\/doi.org\/10.14778\/3750601.3750671","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2025,8]]},"assertion":[{"value":"2025-09-16","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}