{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T18:36:31Z","timestamp":1779129391888,"version":"3.51.4"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2024,9,30]],"date-time":"2024-09-30T00:00:00Z","timestamp":1727654400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,9,30]],"date-time":"2024-09-30T00:00:00Z","timestamp":1727654400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100006764","name":"Technische Universit\u00e4t Berlin","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100006764","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Datenbank Spektrum"],"published-print":{"date-parts":[[2024,11]]},"abstract":"<jats:title>Summary<\/jats:title><jats:p>Software systems that learn from data via machine learning (ML) are being deployed in increasing numbers in real world application scenarios. These ML applications contain complex data preparation pipelines, which take several raw inputs, integrate, filter and encode them to produce the input data for model training. This is in stark contrast to academic studies and benchmarks, which typically work with static, already prepared datasets. It is a difficult and tedious task to ensure at development time that the data preparation pipelines for such ML applications adhere to sound experimentation practices and compliance requirements. Identifying potential correctness issues currently requires a high degree of discipline, knowledge, and time from data scientists, and they often only implement one-off solutions, based on specialised frameworks that are incompatible with the rest of the data science ecosystem.<\/jats:p><jats:p>We discuss how to model data preparation pipelines as dataflow computations from relational inputs to matrix outputs, and propose techniques that use record-level provenance to automatically screen these pipelines for many common correctness issues (e.g., data leakage between train and test data). We design a prototypical system to screen such data preparation pipelines and furthermore enable the automatic computation of important metadata such as group fairness metrics. We discuss how to extract the semantics and the data provenance of common artifacts in supervised learning tasks and evaluate our system on several example pipelines with real-world data.<\/jats:p>","DOI":"10.1007\/s13222-024-00483-4","type":"journal-article","created":{"date-parts":[[2024,9,30]],"date-time":"2024-09-30T10:01:51Z","timestamp":1727690511000},"page":"187-196","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Automated Provenance-Based Screening of ML Data Preparation Pipelines"],"prefix":"10.1007","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4722-5840","authenticated-orcid":false,"given":"Sebastian","family":"Schelter","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shubha","family":"Guha","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stefan","family":"Grafberger","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,9,30]]},"reference":[{"key":"483_CR1","volume-title":"Hidden technical debt in machine learning systems","author":"D Sculley","year":"2015","unstructured":"Sculley\u00a0D, Holt\u00a0G, Golovin\u00a0D et\u00a0al (2015) Hidden technical debt in machine learning systems. NeurIPS"},{"key":"483_CR2","doi-asserted-by":"crossref","unstructured":"Polyzotis N, Roy S, Whang SE, Zinkevich M (2018) Data lifecycle challenges in production machine learning: a survey. SIGMOD Rec 47:","DOI":"10.1145\/3299887.3299891"},{"key":"483_CR3","volume-title":"On challenges in machine learning model management","author":"S Schelter","year":"2018","unstructured":"Schelter\u00a0S et\u00a0al (2018) On challenges in machine learning model management. IEEE Data Engineering Bulletin"},{"key":"483_CR4","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415570","volume-title":"Responsible data management","author":"J Stoyanovich","year":"2020","unstructured":"Stoyanovich\u00a0J, Howe\u00a0B, Jagadish\u00a0H (2020) Responsible data management. PVLDB"},{"key":"483_CR5","first-page":"1","volume":"14","author":"W McKinney","year":"2011","unstructured":"McKinney\u00a0W et\u00a0al (2011) Pandas: a foundational python library for data analysis and statistics. Pyth Hi Perf Sci Comp 14:1\u20139","journal-title":"Pyth Hi Perf Sci Comp"},{"key":"483_CR6","first-page":"15","volume-title":"Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing","author":"M Zaharia","year":"2012","unstructured":"Zaharia\u00a0M et\u00a0al (2012) Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. 9th {USENIX} Symposium on Networked Systems Design and Implementation ({NSDI} 12), S\u00a015\u201328"},{"key":"483_CR7","volume-title":"Scikit-learn: Machine learning in python","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa\u00a0F, Varoquaux\u00a0G et\u00a0al (2011) Scikit-learn: Machine learning in python. JMLR"},{"key":"483_CR8","first-page":"1235","volume":"17","author":"X Meng","year":"2016","unstructured":"Meng\u00a0X et\u00a0al (2016) Mllib: Machine learning in apache spark. J\u00a0Mach Learn Res 17:1235\u20131241","journal-title":"J Mach Learn Res"},{"key":"483_CR9","doi-asserted-by":"crossref","unstructured":"Baylor D et\u00a0al (2017) Tfx: A tensorflow-based production-scale machine learning platform. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp\u00a01387\u20131395","DOI":"10.1145\/3097983.3098021"},{"key":"483_CR10","volume-title":"Why is my classifier discriminatory?","author":"I Chen","year":"2018","unstructured":"Chen\u00a0I, Johansson\u00a0FD, Sontag\u00a0D (2018) Why is my classifier discriminatory? NeurIPS"},{"key":"483_CR11","volume-title":"\u201camnesia\u201d-a selection of machine learning models that can forget user data very fast","author":"S Schelter","year":"2020","unstructured":"Schelter\u00a0S (2020) \u201camnesia\u201d-a selection of machine learning models that can forget user data very fast. CIDR"},{"key":"483_CR12","unstructured":"Bellamy RK, Dey K et\u00a0al (2018) Ai fairness 360: An extensible toolkit for detecting, understanding, and mitigating unwanted algorithmic bias. arXiv:1810.01943"},{"key":"483_CR13","doi-asserted-by":"crossref","unstructured":"Herschel M, Diestelk\u00e4mper R, Lahmar HB (2017) A survey on provenance: What for? what form? what from? VLDB\u00a0J","DOI":"10.1007\/s00778-017-0486-1"},{"key":"483_CR14","doi-asserted-by":"publisher","DOI":"10.1145\/1265530.1265535","volume-title":"Provenance semirings","author":"TJ Green","year":"2007","unstructured":"Green\u00a0TJ, Karvounarakis\u00a0G, Tannen\u00a0V (2007) Provenance semirings. PODS"},{"key":"483_CR15","volume-title":"Lightweight inspection of data preprocessing in native machine learning pipelines","author":"S Grafberger","year":"2021","unstructured":"Grafberger\u00a0S, Stoyanovich\u00a0J, Schelter\u00a0S (2021) Lightweight inspection of data preprocessing in native machine learning pipelines. CIDR"},{"key":"483_CR16","doi-asserted-by":"crossref","unstructured":"Grafberger S, Groth P, Stoyanovich J, Schelter S (2022) Data distribution debugging for machine learning pipelines. Int J Very Large Datab (vldbj)","DOI":"10.1007\/s00778-021-00726-w"},{"key":"483_CR17","volume-title":"Screening native machine learning pipelines with arguseyes","author":"S Schelter","year":"2022","unstructured":"Schelter\u00a0S et\u00a0al (2022) Screening native machine learning pipelines with arguseyes. Conference on Innovative Data Systems Research (CIDR)"},{"key":"483_CR18","doi-asserted-by":"crossref","unstructured":"Schelter S, Grafberger S, Guha S, Karlas B Zhang C Proactively screening machine learning pipelines with arguseyes. Companion of the 2023 International Conference on Management of Data, S\u00a091\u201394","DOI":"10.1145\/3555041.3589682"},{"key":"483_CR19","volume-title":"Accelerating the machine learning lifecycle with mlflow","author":"M Zaharia","year":"2018","unstructured":"Zaharia\u00a0M, Chen\u00a0A, Davidson\u00a0A et\u00a0al (2018) Accelerating the machine learning lifecycle with mlflow. IEEE, Data Engineering Bulletin"},{"key":"483_CR20","doi-asserted-by":"publisher","first-page":"59","DOI":"10.1007\/s13222-021-00399-3","volume":"22","author":"M Klettke","year":"2022","unstructured":"Klettke\u00a0M, St\u00f6rl\u00a0U (2022) Four generations in data engineering for data science. Datenb-Spek 22:59\u201366. https:\/\/doi.org\/10.1007\/s13222-021-00399-3","journal-title":"Datenb-Spek"},{"key":"483_CR21","doi-asserted-by":"publisher","first-page":"40","DOI":"10.48786\/edbt.2023.04","volume-title":"Proceedings 26th International Conference on Extending Database Technology, EDBT 2023, Ioannina, Greece, March 28\u201331, 2023","author":"ME Sch\u00fcle","year":"2023","unstructured":"Sch\u00fcle\u00a0ME, Scalerandi\u00a0L, Kemper\u00a0A, Neumann\u00a0T (2023) Blue elephants inspecting pandas: Inspection and execution of machine learning pipelines in SQL. In: Stoyanovich\u00a0J et\u00a0al (Hrsg) Proceedings 26th International Conference on Extending Database Technology, EDBT 2023, Ioannina, Greece, March 28\u201331, 2023. OpenProceedings.org, S\u00a040\u201352 https:\/\/doi.org\/10.48786\/edbt.2023.04"},{"key":"483_CR22","doi-asserted-by":"publisher","first-page":"121","DOI":"10.1007\/s13222-022-00413-2","volume":"22","author":"F Neutatz","year":"2022","unstructured":"Neutatz\u00a0F, Chen\u00a0B, Alkhatib\u00a0Y, Ye\u00a0J, Abedjan\u00a0Z (2022) Data cleaning and automl: Would an optimizer choose to clean? Datenb-Spek 22:121\u2013130. https:\/\/doi.org\/10.1007\/s13222-022-00413-2","journal-title":"Datenb-Spek"},{"key":"483_CR23","doi-asserted-by":"publisher","unstructured":"Kumar A, Boehm M, Yang J (2017) Data management in machine learning: Challenges, techniques, and systems. Proceedings of the 2017 ACM International Conference on Management of Data, pp\u00a01717\u20131722. https:\/\/doi.org\/10.1145\/3035918.3054775","DOI":"10.1145\/3035918.3054775"},{"key":"483_CR24","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-024-00835-2","author":"S Redyuk","year":"2024","unstructured":"Redyuk\u00a0S, Kaoudi\u00a0Z, Schelter\u00a0S, Markl\u00a0V (2024) Assisted design of data science pipelines. VLDB\u00a0J. https:\/\/doi.org\/10.1007\/s00778-024-00835-2","journal-title":"VLDB J"},{"key":"483_CR25","volume-title":"Red onions, soft cheese and data: From food safety to data traceability for responsible ai","author":"S Grafberger","year":"2024","unstructured":"Grafberger\u00a0S, Zhang\u00a0Z, Schelter\u00a0S, Zhang\u00a0C (2024) Red onions, soft cheese and data: From food safety to data traceability for responsible ai. IEEE Data Engineering Bulletin"},{"key":"483_CR26","volume-title":"Towards interactively improving ml data preparation code via \u201cshadow pipelines\u201d","author":"S Grafberger","year":"2022","unstructured":"Grafberger\u00a0S, Groth\u00a0P, Schelter\u00a0S (2022) Towards interactively improving ml data preparation code via \u201cshadow pipelines\u201d. DEEM workshop @ SIGMOD"},{"key":"483_CR27","doi-asserted-by":"publisher","DOI":"10.1145\/3644385","author":"A Chapman","year":"2024","unstructured":"Chapman\u00a0A, Lauro\u00a0L, Missier\u00a0P, Torlone\u00a0R (2024) Supporting better insights of data science pipelines with fine-grained provenance. ACM Trans Database Syst. https:\/\/doi.org\/10.1145\/3644385","journal-title":"ACM Trans Database Syst"},{"key":"483_CR28","doi-asserted-by":"publisher","DOI":"10.1145\/3589273","volume-title":"Automating and optimizing data-centric what-if analyses on\u00a0native\u00a0 machine\u00a0learning\u00a0pipelines","author":"S Grafberger","year":"2023","unstructured":"Grafberger\u00a0S, Groth\u00a0P, Schelter\u00a0S (2023) Automating and optimizing data-centric what-if analyses on\u00a0native\u00a0 machine\u00a0learning\u00a0pipelines. SIGMOD"},{"key":"483_CR29","doi-asserted-by":"publisher","first-page":"4002","DOI":"10.14778\/3611540.3611606","volume":"16","author":"S Grafberger","year":"2023","unstructured":"Grafberger\u00a0S, Guha\u00a0S, Groth\u00a0P, Schelter\u00a0S (2023) mlwhatif: What if you could stop re-implementing your machine learning pipeline analyses over and over? Proc Vldb Endow 16:4002\u20134005. https:\/\/doi.org\/10.14778\/3611540.3611606","journal-title":"Proc Vldb Endow"},{"key":"483_CR30","volume-title":"Reconstructing and querying ml pipeline intermediates","author":"S Schelter","year":"2023","unstructured":"Schelter\u00a0S (2023) Reconstructing and querying ml pipeline intermediates. CIDR"},{"key":"483_CR31","volume-title":"Failing loudly: An empirical study of methods for detecting dataset shift","author":"S Rabanser","year":"2019","unstructured":"Rabanser\u00a0S, G\u00fcnnemann\u00a0S, Lipton\u00a0ZC (2019) Failing loudly: An empirical study of methods for detecting dataset shift. NeurIPS"},{"key":"483_CR32","doi-asserted-by":"crossref","unstructured":"Jia R, Dao D, Wang B et\u00a0al (2019) Efficient task-specific data valuation for nearest neighbor algorithms. PVLDB 12:","DOI":"10.14778\/3342263.3342637"},{"key":"483_CR33","volume-title":"Vamsa: Automated provenance tracking in data science scripts","author":"MH Namaki","year":"2020","unstructured":"Namaki\u00a0MH, Floratou\u00a0A, Psallidas\u00a0F et\u00a0al (2020) Vamsa: Automated provenance tracking in data science scripts. KDD"},{"key":"483_CR34","volume-title":"Extending relational query processing with ML inference","author":"K Karanasos","year":"2021","unstructured":"Karanasos\u00a0K et\u00a0al (2021) Extending relational query processing with ML inference. CIDR"},{"key":"483_CR35","unstructured":"Onnx (2019) http:\/\/onnx.ai\/"},{"key":"483_CR36","doi-asserted-by":"crossref","unstructured":"Mehrabi N, Morstatter F, Saxena N, Lerman K, Galstyan A (2021) A survey on bias and fairness in machine learning. ACM Comput Surv 54:","DOI":"10.1145\/3457607"},{"key":"483_CR37","doi-asserted-by":"crossref","unstructured":"Grafberger S, Guha S, Stoyanovich J, Schelter S (2021) Mlinspect: A data distribution debugger for machine learning pipelines. Proceedings of the 2021 International Conference on Management of Data, pp\u00a02736\u20132739","DOI":"10.1145\/3448016.3452759"},{"key":"483_CR38","unstructured":"Bellamy RKE et\u00a0al (2018) AI Fairness 360: An extensible toolkit for detecting, understanding, and mitigating unwanted algorithmic bias. https:\/\/arxiv.org\/abs\/1810.01943"}],"container-title":["Datenbank-Spektrum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13222-024-00483-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s13222-024-00483-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13222-024-00483-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,19]],"date-time":"2024-11-19T16:43:15Z","timestamp":1732034595000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s13222-024-00483-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,30]]},"references-count":38,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,11]]}},"alternative-id":["483"],"URL":"https:\/\/doi.org\/10.1007\/s13222-024-00483-4","relation":{},"ISSN":["1618-2162","1610-1995"],"issn-type":[{"value":"1618-2162","type":"print"},{"value":"1610-1995","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,9,30]]},"assertion":[{"value":"13 May 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 August 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 September 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}