{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,18]],"date-time":"2025-11-18T11:48:38Z","timestamp":1763466518811},"reference-count":16,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2015,8]]},"abstract":"<jats:p>Analysts report spending upwards of 80% of their time on problems in data cleaning. The data cleaning process is inherently iterative, with evolving cleaning workflows that start with basic exploratory data analysis on small samples of dirty data, then refine analysis with more sophisticated\/expensive cleaning operators (e.g., crowdsourcing), and finally apply the insights to a full dataset. While an analyst often knows at a logical level what operations need to be done, they often have to manage a large search space of physical operators and parameters. We present Wisteria, a system designed to support the iterative development and optimization of data cleaning workflows, especially ones that utilize the crowd. Wisteria separates logical operations from physical implementations, and driven by analyst feedback, suggests optimizations and\/or replacements to the analyst's choice of physical implementation. We highlight research challenges in sampling, in-flight operator replacement, and crowdsourcing. We overview the system architecture and these techniques, then provide a demonstration designed to showcase how Wisteria can improve iterative data analysis and cleaning. The code is available at: http:\/\/www.sampleclean.org.<\/jats:p>","DOI":"10.14778\/2824032.2824122","type":"journal-article","created":{"date-parts":[[2015,9,16]],"date-time":"2015-09-16T12:18:17Z","timestamp":1442405897000},"page":"2004-2007","source":"Crossref","is-referenced-by-count":27,"title":["Wisteria"],"prefix":"10.14778","volume":"8","author":[{"given":"Daniel","family":"Haas","sequence":"first","affiliation":[{"name":"UC Berkeley"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sanjay","family":"Krishnan","sequence":"additional","affiliation":[{"name":"UC Berkeley"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiannan","family":"Wang","sequence":"additional","affiliation":[{"name":"UC Berkeley"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael J.","family":"Franklin","sequence":"additional","affiliation":[{"name":"UC Berkeley"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eugene","family":"Wu","sequence":"additional","affiliation":[{"name":"Columbia University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,8]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Apache falcon. http:\/\/falcon.apache.org.  Apache falcon. http:\/\/falcon.apache.org."},{"key":"e_1_2_1_2_1","unstructured":"Informatica. https:\/\/www.informatica.com.  Informatica. https:\/\/www.informatica.com."},{"key":"e_1_2_1_3_1","unstructured":"Talend. https:\/\/www.talend.com\/solutions\/etl-analytics.  Talend. https:\/\/www.talend.com\/solutions\/etl-analytics."},{"key":"e_1_2_1_4_1","unstructured":"Trifacta. http:\/\/www.trifacta.com.  Trifacta. http:\/\/www.trifacta.com."},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","first-page":"1126","DOI":"10.1145\/2623330.2623617","volume-title":"Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining","author":"Chen Z.","year":"2014","unstructured":"Z. Chen and M. Cafarella . Integrating spreadsheet data via accurate and low-effort extraction . In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining , pages 1126 -- 1135 . ACM, 2014 . 10.1145\/2623330.2623617 Z. Chen and M. Cafarella. Integrating spreadsheet data via accurate and low-effort extraction. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 1126--1135. ACM, 2014. 10.1145\/2623330.2623617"},{"key":"e_1_2_1_6_1","first-page":"541","volume-title":"SIGMOD Conference","author":"Dallachiesa M.","year":"2013","unstructured":"M. Dallachiesa , A. Ebaid , A. Eldawy , A. K. Elmagarmid , I. F. Ilyas , M. Ouzzani , and N. Tang . Nadeef: a commodity data cleaning system . In SIGMOD Conference , pages 541 -- 552 , 2013 . 10.1145\/2463676.2465327 M. Dallachiesa, A. Ebaid, A. Eldawy, A. K. Elmagarmid, I. F. Ilyas, M. Ouzzani, and N. Tang. Nadeef: a commodity data cleaning system. In SIGMOD Conference, pages 541--552, 2013. 10.1145\/2463676.2465327"},{"key":"e_1_2_1_7_1","volume-title":"SIGMOD","author":"Gokhale C.","year":"2014","unstructured":"C. Gokhale , S. Das , A. Doan , J. F. Naughton , N. Rampalli , J. Shavlik , and X. Zhu . Corleone: Hands-off crowdsourcing for entity matching . In SIGMOD , 2014 . 10.1145\/2588555.2588576 C. Gokhale, S. Das, A. Doan, J. F. Naughton, N. Rampalli, J. Shavlik, and X. Zhu. Corleone: Hands-off crowdsourcing for entity matching. In SIGMOD, 2014. 10.1145\/2588555.2588576"},{"key":"e_1_2_1_8_1","first-page":"3363","volume-title":"CHI","author":"Kandel S.","year":"2011","unstructured":"S. Kandel , A. Paepcke , J. Hellerstein , and J. Heer . Wrangler: interactive visual specification of data transformation scripts . In CHI , pages 3363 -- 3372 , 2011 . 10.1145\/1978942.1979444 S. Kandel, A. Paepcke, J. Hellerstein, and J. Heer. Wrangler: interactive visual specification of data transformation scripts. In CHI, pages 3363--3372, 2011. 10.1145\/1978942.1979444"},{"key":"e_1_2_1_9_1","volume-title":"VAST","author":"Kandel S.","year":"2012","unstructured":"S. Kandel , A. Paepcke , J. Hellerstein , and H. Jeffrey . Enterprise data analysis and visualization: An interview study . VAST , 2012 . S. Kandel, A. Paepcke, J. Hellerstein, and H. Jeffrey. Enterprise data analysis and visualization: An interview study. VAST, 2012."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824037"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807167.1807178"},{"key":"e_1_2_1_12_1","volume-title":"SIGMOD","author":"Park H.","year":"2014","unstructured":"H. Park and J. Widom . Crowdfill: Collecting structured data from the crowd . In SIGMOD , 2014 . 10.1145\/2588555.2610503 H. Park and J. Widom. Crowdfill: Collecting structured data from the crowd. In SIGMOD, 2014. 10.1145\/2588555.2610503"},{"key":"e_1_2_1_13_1","volume-title":"CIDR","author":"Stonebraker M.","year":"2013","unstructured":"M. Stonebraker , D. Bruckner , I. F. Ilyas , G. Beskales , M. Cherniack , S. B. Zdonik , A. Pagan , and S. Xu . Data curation at scale: The data tamer system . In CIDR , 2013 . M. Stonebraker, D. Bruckner, I. F. Ilyas, G. Beskales, M. Cherniack, S. B. Zdonik, A. Pagan, and S. Xu. Data curation at scale: The data tamer system. In CIDR, 2013."},{"key":"e_1_2_1_14_1","first-page":"301","volume-title":"Proceedings of the 11th USENIX conference on Operating Systems Design and Implementation","author":"Venkataraman S.","year":"2014","unstructured":"S. Venkataraman , A. Panda , G. Ananthanarayanan , M. J. Franklin , and I. Stoica . The power of choice in data-aware cluster scheduling . In Proceedings of the 11th USENIX conference on Operating Systems Design and Implementation , pages 301 -- 316 . USENIX Association , 2014 . S. Venkataraman, A. Panda, G. Ananthanarayanan, M. J. Franklin, and I. Stoica. The power of choice in data-aware cluster scheduling. In Proceedings of the 11th USENIX conference on Operating Systems Design and Implementation, pages 301--316. USENIX Association, 2014."},{"key":"e_1_2_1_15_1","volume-title":"Using OpenRefine","author":"Verborgh R.","year":"2013","unstructured":"R. Verborgh and M. De Wilde . Using OpenRefine . Packt Publishing Ltd , 2013 . R. Verborgh and M. De Wilde. Using OpenRefine. Packt Publishing Ltd, 2013."},{"key":"e_1_2_1_16_1","first-page":"469","volume-title":"SIGMOD Conference","author":"Wang J.","year":"2014","unstructured":"J. Wang , S. Krishnan , M. J. Franklin , K. Goldberg , T. Kraska , and T. Milo . A sample-and-clean framework for fast and accurate query processing on dirty data . In SIGMOD Conference , pages 469 -- 480 , 2014 . 10.1145\/2588555.2610505 J. Wang, S. Krishnan, M. J. Franklin, K. Goldberg, T. Kraska, and T. Milo. A sample-and-clean framework for fast and accurate query processing on dirty data. In SIGMOD Conference, pages 469--480, 2014. 10.1145\/2588555.2610505"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/2824032.2824122","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:04:42Z","timestamp":1672221882000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/2824032.2824122"}},"subtitle":["nurturing scalable data cleaning infrastructure"],"short-title":[],"issued":{"date-parts":[[2015,8]]},"references-count":16,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2015,8]]}},"alternative-id":["10.14778\/2824032.2824122"],"URL":"https:\/\/doi.org\/10.14778\/2824032.2824122","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2015,8]]}}}