{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T05:08:46Z","timestamp":1755839326138,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":25,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,6,12]],"date-time":"2022-06-12T00:00:00Z","timestamp":1654992000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100008536","name":"Amazon Web Services","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100008536","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"NSF (National Science Foundation)","doi-asserted-by":"publisher","award":["1845638,1740305,2008295,2106197,2103794"],"award-info":[{"award-number":["1845638,1740305,2008295,2106197,2103794"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100006785","name":"Google","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100006785","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,6,12]]},"DOI":"10.1145\/3533028.3533305","type":"proceedings-article","created":{"date-parts":[[2022,5,23]],"date-time":"2022-05-23T22:19:46Z","timestamp":1653344386000},"page":"1-5","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["How I stopped worrying about training data bugs and started complaining"],"prefix":"10.1145","author":[{"given":"Lampros","family":"Flokas","sequence":"first","affiliation":[{"name":"Columbia University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weiyuan","family":"Wu","sequence":"additional","affiliation":[{"name":"Simon Fraser University, Burnaby, BC, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiannan","family":"Wang","sequence":"additional","affiliation":[{"name":"Simon Fraser University, Burnaby, BC, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nakul","family":"Verma","sequence":"additional","affiliation":[{"name":"Columbia University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eugene","family":"Wu","sequence":"additional","affiliation":[{"name":"Columbia University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,6,12]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Xu Chu Ihab F Ilyas Sanjay Krishnan and Jiannan Wang. 2016. Data cleaning: Overview and emerging challenges. In SIGMOD.  Xu Chu Ihab F Ilyas Sanjay Krishnan and Jiannan Wang. 2016. Data cleaning: Overview and emerging challenges. In SIGMOD.","DOI":"10.1145\/2882903.2912574"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"crossref","unstructured":"Anna Fariha Suman Nath and Alexandra Meliou. 2020. Causality-Guided Adaptive Interventional Debugging. In SIGMOD.  Anna Fariha Suman Nath and Alexandra Meliou. 2020. Causality-Guided Adaptive Interventional Debugging. In SIGMOD.","DOI":"10.1145\/3318464.3389694"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Lampros Flokas Weiyuan Wu Yejia Liu Jiannan Wang Nakul Verma and Eugene Wu. 2022. Complaint-Driven Training Data Debugging at Interactive Speeds. In SIGMOD.  Lampros Flokas Weiyuan Wu Yejia Liu Jiannan Wang Nakul Verma and Eugene Wu. 2022. Complaint-Driven Training Data Debugging at Interactive Speeds. In SIGMOD.","DOI":"10.1145\/3514221.3517849"},{"key":"e_1_3_2_1_4_1","volume-title":"DataExposer: Exposing Disconnect between Data and Systems. arxiv","author":"Galhotra Sainyam","year":"2021","unstructured":"Sainyam Galhotra , Anna Fariha , Raoni Louren\u00e7o , Juliana Freire , Alexandra Meliou , and Divesh Srivastava . 2021. DataExposer: Exposing Disconnect between Data and Systems. arxiv ( 2021 ). arXiv:2105.06058 Sainyam Galhotra, Anna Fariha, Raoni Louren\u00e7o, Juliana Freire, Alexandra Meliou, and Divesh Srivastava. 2021. DataExposer: Exposing Disconnect between Data and Systems. arxiv (2021). arXiv:2105.06058"},{"volume-title":"Meet Michelangelo: Uber's machine learning platform. https:\/\/eng.uber.com\/michelangelo.","author":"Hermann J.","key":"e_1_3_2_1_5_1","unstructured":"J. Hermann and M. D. Balso . [n.d.] . Meet Michelangelo: Uber's machine learning platform. https:\/\/eng.uber.com\/michelangelo. J. Hermann and M. D. Balso. [n.d.]. Meet Michelangelo: Uber's machine learning platform. https:\/\/eng.uber.com\/michelangelo."},{"key":"e_1_3_2_1_6_1","volume-title":"Muhammad Ali Gulzar, Seunghyun Yoo, Miryung Kim, Todd Millstein, and Tyson Condie.","author":"Interlandi Matteo","year":"2015","unstructured":"Matteo Interlandi , Kshitij Shah , Sai Deep Tetali , Muhammad Ali Gulzar, Seunghyun Yoo, Miryung Kim, Todd Millstein, and Tyson Condie. 2015 . Titian : Data Provenance Support in Spark. VLDB ( 2015). Matteo Interlandi, Kshitij Shah, Sai Deep Tetali, Muhammad Ali Gulzar, Seunghyun Yoo, Miryung Kim, Todd Millstein, and Tyson Condie. 2015. Titian: Data Provenance Support in Spark. VLDB (2015)."},{"key":"e_1_3_2_1_7_1","volume-title":"Xu Chu, Wentao Wu, and Ce Zhang.","author":"Karlas Bojan","year":"2020","unstructured":"Bojan Karlas , Peng Li , Renzhi Wu , Nezihe Merve G\u00fcrel , Xu Chu, Wentao Wu, and Ce Zhang. 2020 . Nearest Neighbor Classifiers over Incomplete Information : From Certain Answers to Certain Predictions. VLDB ( 2020). Bojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve G\u00fcrel, Xu Chu, Wentao Wu, and Ce Zhang. 2020. Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions. VLDB (2020)."},{"key":"e_1_3_2_1_8_1","unstructured":"Pang Wei Koh and Percy Liang. 2017. Understanding Black-box Predictions via Influence Functions. In ICML.  Pang Wei Koh and Percy Liang. 2017. Understanding Black-box Predictions via Influence Functions. In ICML."},{"key":"e_1_3_2_1_9_1","volume-title":"BoostClean: Automated Error Detection and Repair for Machine Learning. arixv","author":"Krishnan Sanjay","year":"2017","unstructured":"Sanjay Krishnan , Michael J. Franklin , Ken Goldberg , and Eugene Wu. 2017. BoostClean: Automated Error Detection and Repair for Machine Learning. arixv ( 2017 ). arXiv:1711.01299 Sanjay Krishnan, Michael J. Franklin, Ken Goldberg, and Eugene Wu. 2017. BoostClean: Automated Error Detection and Repair for Machine Learning. arixv (2017). arXiv:1711.01299"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2899409"},{"key":"e_1_3_2_1_11_1","volume-title":"AlphaClean: Automatic Generation of Data Cleaning Pipelines. arxiv","author":"Krishnan Sanjay","year":"2019","unstructured":"Sanjay Krishnan and Eugene Wu. 2019. AlphaClean: Automatic Generation of Data Cleaning Pipelines. arxiv ( 2019 ). arXiv:1904.11827 Sanjay Krishnan and Eugene Wu. 2019. AlphaClean: Automatic Generation of Data Cleaning Pipelines. arxiv (2019). arXiv:1904.11827"},{"key":"e_1_3_2_1_12_1","unstructured":"Seokki Lee Bertram Lud\u00e4scher and Boris Glavic. 2019. PUG: a framework and practical implementation for why and why-not provenance. VLDB J. (2019).  Seokki Lee Bertram Lud\u00e4scher and Boris Glavic. 2019. PUG: a framework and practical implementation for why and why-not provenance. VLDB J. (2019)."},{"key":"e_1_3_2_1_13_1","volume-title":"Picket: Self-supervised Data Diagnostics for ML Pipelines. arxiv","author":"Liu Zifan","year":"2020","unstructured":"Zifan Liu , Zhechun Zhou , and Theodoros Rekatsinas . 2020 . Picket: Self-supervised Data Diagnostics for ML Pipelines. arxiv (2020). arXiv:2006.04730 Zifan Liu, Zhechun Zhou, and Theodoros Rekatsinas. 2020. Picket: Self-supervised Data Diagnostics for ML Pipelines. arxiv (2020). arXiv:2006.04730"},{"key":"e_1_3_2_1_14_1","volume-title":"Explaining Inference Queries with Bayesian Optimization. arXiv","author":"Lockhart Brandon","year":"2021","unstructured":"Brandon Lockhart , Jinglin Peng , Weiyuan Wu , Jiannan Wang , and Eugene Wu. 2021. Explaining Inference Queries with Bayesian Optimization. arXiv ( 2021 ). Brandon Lockhart, Jinglin Peng, Weiyuan Wu, Jiannan Wang, and Eugene Wu. 2021. Explaining Inference Queries with Bayesian Optimization. arXiv (2021)."},{"key":"e_1_3_2_1_15_1","volume-title":"Shasha","author":"Louren\u00e7o Raoni","year":"2020","unstructured":"Raoni Louren\u00e7o , Juliana Freire , and Dennis E . Shasha . 2020 . BugDoc: Algorithms to Debug Computational Processes. In SIGMOD. Raoni Louren\u00e7o, Juliana Freire, and Dennis E. Shasha. 2020. BugDoc: Algorithms to Debug Computational Processes. In SIGMOD."},{"key":"e_1_3_2_1_16_1","unstructured":"MLTrace [n.d.]. Home - mltrace 0.16 documentation. https:\/\/mltrace.readthedocs.io\/. Accessed: 2021-08-31.  MLTrace [n.d.]. Home - mltrace 0.16 documentation. https:\/\/mltrace.readthedocs.io\/. Accessed: 2021-08-31."},{"key":"e_1_3_2_1_17_1","volume-title":"From Cleaning before ML to Cleaning for ML. Data Engineering","author":"Neutatz Felix","year":"2021","unstructured":"Felix Neutatz , Binger Chen , Ziawasch Abedjan , and Eugene Wu. 2021. From Cleaning before ML to Cleaning for ML. Data Engineering ( 2021 ). Felix Neutatz, Binger Chen, Ziawasch Abedjan, and Eugene Wu. 2021. From Cleaning before ML to Cleaning for ML. Data Engineering (2021)."},{"key":"e_1_3_2_1_18_1","volume-title":"Smoke: Fine-grained Lineage at Interactive Speed. VLDB","author":"Psallidas Fotis","year":"2018","unstructured":"Fotis Psallidas and Eugene Wu . 2018 . Smoke: Fine-grained Lineage at Interactive Speed. VLDB (2018). Fotis Psallidas and Eugene Wu. 2018. Smoke: Fine-grained Lineage at Interactive Speed. VLDB (2018)."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357223.3362702"},{"key":"e_1_3_2_1_20_1","volume-title":"Overton: A Data System for Monitoring and Improving Machine-Learned Products. In CIDR.","author":"R\u00e9 Christopher","year":"2020","unstructured":"Christopher R\u00e9 . 2020 . Overton: A Data System for Monitoring and Improving Machine-Learned Products. In CIDR. Christopher R\u00e9. 2020. Overton: A Data System for Monitoring and Improving Machine-Learned Products. In CIDR."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Nithya Sambasivan Shivani Kapania Hannah Highfill Diana Akrong Praveen Paritosh and Lora M Aroyo. 2021. \"Everyone wants to do the model work not the data work\": Data Cascades in High-Stakes AI. In CHI.  Nithya Sambasivan Shivani Kapania Hannah Highfill Diana Akrong Praveen Paritosh and Lora M Aroyo. 2021. \"Everyone wants to do the model work not the data work\": Data Cascades in High-Stakes AI. In CHI.","DOI":"10.1145\/3411764.3445518"},{"key":"e_1_3_2_1_22_1","volume-title":"MODELDB: A System for Machine Learning Model Management. In CIDR.","author":"Vartak Manasi","year":"2017","unstructured":"Manasi Vartak . 2017 . MODELDB: A System for Machine Learning Model Management. In CIDR. Manasi Vartak. 2017. MODELDB: A System for Machine Learning Model Management. In CIDR."},{"key":"e_1_3_2_1_23_1","unstructured":"Wentao Wu and Ce Zhang. 2021. Towards understanding end-to-end learning in the context of data: machine learning dancing over semirings & Codd's table. In DEEM@SIGMOD.  Wentao Wu and Ce Zhang. 2021. Towards understanding end-to-end learning in the context of data: machine learning dancing over semirings & Codd's table. In DEEM@SIGMOD."},{"key":"e_1_3_2_1_24_1","volume-title":"Davidson","author":"Wu Yinjun","year":"2021","unstructured":"Yinjun Wu , James Weimer , and Susan B . Davidson . 2021 . CHEF : A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties. VLDB ( 2021). Yinjun Wu, James Weimer, and Susan B. Davidson. 2021. CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties. VLDB (2021)."},{"key":"e_1_3_2_1_25_1","volume-title":"Andy Konwinski, Siddharth Murching, Tomas Nykodym, Paul Ogilvie, Mani Parkhe, et al.","author":"Zaharia Matei","year":"2018","unstructured":"Matei Zaharia , Andrew Chen , Aaron Davidson , Ali Ghodsi , Sue Ann Hong , Andy Konwinski, Siddharth Murching, Tomas Nykodym, Paul Ogilvie, Mani Parkhe, et al. 2018 . Accelerating the machine learning lifecycle with MLflow. IEEE Data Eng. Bull . (2018). Matei Zaharia, Andrew Chen, Aaron Davidson, Ali Ghodsi, Sue Ann Hong, Andy Konwinski, Siddharth Murching, Tomas Nykodym, Paul Ogilvie, Mani Parkhe, et al. 2018. Accelerating the machine learning lifecycle with MLflow. IEEE Data Eng. Bull. (2018)."}],"event":{"name":"SIGMOD\/PODS '22: International Conference on Management of Data","sponsor":["SIGMOD ACM Special Interest Group on Management of Data"],"location":"Philadelphia Pennsylvania","acronym":"SIGMOD\/PODS '22"},"container-title":["Proceedings of the Sixth Workshop on Data Management for End-To-End Machine Learning"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3533028.3533305","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/abs\/10.1145\/3533028.3533305","content-type":"text\/html","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3533028.3533305","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3533028.3533305","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:38Z","timestamp":1750186838000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3533028.3533305"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,12]]},"references-count":25,"alternative-id":["10.1145\/3533028.3533305","10.1145\/3533028"],"URL":"https:\/\/doi.org\/10.1145\/3533028.3533305","relation":{},"subject":[],"published":{"date-parts":[[2022,6,12]]},"assertion":[{"value":"2022-06-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}