{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T03:06:07Z","timestamp":1775012767983,"version":"3.50.1"},"reference-count":7,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2012,8]]},"abstract":"<jats:p>\n            As data analytics becomes mainstream, and the complexity of the underlying data and computation grows, it will be increasingly important to provide tools that help analysts understand the underlying reasons when they encounter errors in the result. While data provenance has been a large step in providing tools to help debug complex workflows, its current form has limited utility when debugging aggregation operators that compute a single output from a large collection of inputs. Traditional provenance will return the entire input collection, which has very low precision. In contrast, users are seeking\n            <jats:italic>precise descriptions<\/jats:italic>\n            of the inputs that caused the errors. We propose a\n            <jats:italic>Ranked Provenance System<\/jats:italic>\n            , which identifies subsets of inputs that influenced the output error, describes each subset with human readable predicates and orders them by contribution to the error. In this demonstration, we will present DBWipes, a novel data cleaning system that allows users to execute aggregate queries, and interactively detect, understand, and clean errors in the query results. Conference attendees will explore anomalies in campaign donations from the current US presidential election and in readings from a 54-node sensor deployment.\n          <\/jats:p>","DOI":"10.14778\/2367502.2367531","type":"journal-article","created":{"date-parts":[[2014,6,24]],"date-time":"2014-06-24T12:17:57Z","timestamp":1403612277000},"page":"1894-1897","source":"Crossref","is-referenced-by-count":9,"title":["A demonstration of DBWipes"],"prefix":"10.14778","volume":"5","author":[{"given":"Eugene","family":"Wu","sequence":"first","affiliation":[{"name":"MIT CSAIL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Samuel","family":"Madden","sequence":"additional","affiliation":[{"name":"MIT CSAIL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael","family":"Stonebraker","sequence":"additional","affiliation":[{"name":"MIT CSAIL"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,8]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"D. Huynh and S. Mazzocchi. Google refine. http:\/\/code.google.com\/p\/google-refine\/.  D. Huynh and S. Mazzocchi. Google refine. http:\/\/code.google.com\/p\/google-refine\/."},{"key":"e_1_2_1_2_1","first-page":"841","volume-title":"SIGMOD","author":"Kanagal B.","year":"2011","unstructured":"B. Kanagal , J. Li , and A. Deshpande . Sensitivity analysis and explanations for robust query evaluation in probabilistic databases . In SIGMOD , pages 841 -- 852 , 2011 . 10.1145\/1989323.1989411 B. Kanagal, J. Li, and A. Deshpande. Sensitivity analysis and explanations for robust query evaluation in probabilistic databases. In SIGMOD, pages 841--852, 2011. 10.1145\/1989323.1989411"},{"key":"e_1_2_1_3_1","first-page":"3363","volume-title":"CHI","author":"Kandel S.","year":"2011","unstructured":"S. Kandel , A. Paepcke , J. Hellerstein , and J. Heer . Wrangler: interactive visual specification of data transformation scripts . In CHI , pages 3363 -- 3372 , 2011 . 10.1145\/1978942.1979444 S. Kandel, A. Paepcke, J. Hellerstein, and J. Heer. Wrangler: interactive visual specification of data transformation scripts. In CHI, pages 3363--3372, 2011. 10.1145\/1978942.1979444"},{"key":"e_1_2_1_4_1","first-page":"153","volume":"5","author":"Lavra N.","year":"2004","unstructured":"N. Lavra , B. Kavek , P. Flach , and L. Todorovski . Subgroup discovery with cn2-sd. In JMLR , volume 5 , pages 153 -- 188 , Feb. 2004 . N. Lavra, B. Kavek, P. Flach, and L. Todorovski. Subgroup discovery with cn2-sd. In JMLR, volume 5, pages 153--188, Feb. 2004.","journal-title":"Subgroup discovery with cn2-sd. In JMLR"},{"key":"e_1_2_1_5_1","first-page":"59","volume":"33","author":"Meliou A.","year":"2010","unstructured":"A. Meliou , W. Gatterbauer , J. Y. Halpern , C. Koch , K. F. Moore , and D. Suciu . Causality in databases. In IEEE Data Eng. Bull. , volume 33 , pages 59 -- 67 , 2010 . A. Meliou, W. Gatterbauer, J. Y. Halpern, C. Koch, K. F. Moore, and D. Suciu. Causality in databases. In IEEE Data Eng. Bull., volume 33, pages 59--67, 2010.","journal-title":"Causality in databases. In IEEE Data Eng. Bull."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.v18:10"},{"key":"e_1_2_1_8_1","first-page":"381","volume-title":"VLDB","author":"Raman V.","year":"2001","unstructured":"V. Raman and J. M. Hellerstein . Potter's wheel: An interactive data cleaning system . In VLDB , pages 381 -- 390 , 2001 . V. Raman and J. M. Hellerstein. Potter's wheel: An interactive data cleaning system. In VLDB, pages 381--390, 2001."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/2367502.2367531","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:49:13Z","timestamp":1672224553000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/2367502.2367531"}},"subtitle":["clean as you query"],"short-title":[],"issued":{"date-parts":[[2012,8]]},"references-count":7,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2012,8]]}},"alternative-id":["10.14778\/2367502.2367531"],"URL":"https:\/\/doi.org\/10.14778\/2367502.2367531","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2012,8]]}}}