{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T17:10:44Z","timestamp":1784913044928,"version":"3.55.0"},"reference-count":37,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2016,8]]},"abstract":"<jats:p>Analysts often clean dirty data iteratively--cleaning some data, executing the analysis, and then cleaning more data based on the results. We explore the iterative cleaning process in the context of statistical model training, which is an increasingly popular form of data analytics. We propose ActiveClean, which allows for progressive and iterative cleaning in statistical modeling problems while preserving convergence guarantees. ActiveClean supports an important class of models called convex loss models (e.g., linear regression and SVMs), and prioritizes cleaning those records likely to affect the results. We evaluate ActiveClean on five real-world datasets UCI Adult, UCI EEG, MNIST, IMDB, and Dollars For Docs with both real and synthetic errors. The results show that our proposed optimizations can improve model accuracy by up-to 2.5x for the same amount of data cleaned. Furthermore for a fixed cleaning budget and on all real dirty datasets, ActiveClean returns more accurate models than uniform sampling and Active Learning.<\/jats:p>","DOI":"10.14778\/2994509.2994514","type":"journal-article","created":{"date-parts":[[2016,9,6]],"date-time":"2016-09-06T15:27:03Z","timestamp":1473175623000},"page":"948-959","source":"Crossref","is-referenced-by-count":193,"title":["ActiveClean"],"prefix":"10.14778","volume":"9","author":[{"given":"Sanjay","family":"Krishnan","sequence":"first","affiliation":[{"name":"UC Berkeley"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiannan","family":"Wang","sequence":"additional","affiliation":[{"name":"Simon Fraser University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eugene","family":"Wu","sequence":"additional","affiliation":[{"name":"Columbia University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Michael J.","family":"Franklin","sequence":"additional","affiliation":[{"name":"UC Berkeley"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ken","family":"Goldberg","sequence":"additional","affiliation":[{"name":"UC Berkeley"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2016,8]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Apache spark survey. https:\/\/databricks.com\/blog\/2015\/09\/24\/spark-survey-results-2015-are-now-available.html.  Apache spark survey. https:\/\/databricks.com\/blog\/2015\/09\/24\/spark-survey-results-2015-are-now-available.html."},{"key":"e_1_2_1_2_1","unstructured":"Dollars for docs. https:\/\/projects.propublica.org\/docdollars\/.  Dollars for docs. https:\/\/projects.propublica.org\/docdollars\/."},{"key":"e_1_2_1_3_1","unstructured":"Keystone ml. http:\/\/keystone-ml.org\/.  Keystone ml. http:\/\/keystone-ml.org\/."},{"key":"e_1_2_1_4_1","unstructured":"Tensor flow. https:\/\/www.tensorflow.org\/.  Tensor flow. https:\/\/www.tensorflow.org\/."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732967.2732975"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2737786"},{"key":"e_1_2_1_7_1","volume-title":"CoRR","author":"Bertsekas D. P.","year":"2015","unstructured":"D. P. Bertsekas . Incremental gradient, subgradient, and proximal methods for convex optimization: A survey . In CoRR , 2015 . D. P. Bertsekas. Incremental gradient, subgradient, and proximal methods for convex optimization: A survey. In CoRR, 2015."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-35289-8_25"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/861869"},{"key":"e_1_2_1_10_1","volume-title":"JMLR","author":"Dekel O.","year":"2012","unstructured":"O. Dekel , R. Gilad-Bachrach , O. Shamir , and L. Xiao . Optimal distributed online prediction using mini-batches . In JMLR , 2012 . O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao. Optimal distributed online prediction using mini-batches. In JMLR, 2012."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/1316689.1316741"},{"key":"e_1_2_1_12_1","volume-title":"JMLR","author":"Drineas P.","year":"2012","unstructured":"P. Drineas , M. Magdon-Ismail , M. W. Mahoney , and D. P. Woodruff . Fast approximation of matrix coherence and statistical leverage . In JMLR , 2012 . P. Drineas, M. Magdon-Ismail, M. W. Mahoney, and D. P. Woodruff. Fast approximation of matrix coherence and statistical leverage. In JMLR, 2012."},{"key":"e_1_2_1_13_1","volume-title":"Foundations of Data Quality Management. Synthesis Lectures on Data Management","author":"Fan W.","year":"2012","unstructured":"W. Fan and F. Geerts . Foundations of Data Quality Management. Synthesis Lectures on Data Management . 2012 . W. Fan and F. Geerts. Foundations of Data Quality Management. Synthesis Lectures on Data Management. 2012."},{"issue":"3","key":"e_1_2_1_14_1","first-page":"37","article-title":"From data mining to knowledge discovery in databases","volume":"17","author":"Fayyad U.","year":"1996","unstructured":"U. Fayyad , G. Piatetsky-Shapiro , and P. Smyth . From data mining to knowledge discovery in databases . AI magazine , 17 ( 3 ): 37 , 1996 . U. Fayyad, G. Piatetsky-Shapiro, and P. Smyth. From data mining to knowledge discovery in databases. AI magazine, 17(3):37, 1996.","journal-title":"AI magazine"},{"key":"e_1_2_1_15_1","volume-title":"NIPS","author":"Feng J.","year":"2014","unstructured":"J. Feng , H. Xu , S. Mannor , and S. Yan . Robust logistic regression and classification . In NIPS , 2014 . J. Feng, H. Xu, S. Mannor, and S. Yan. Robust logistic regression and classification. In NIPS, 2014."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2588555.2588576"},{"key":"e_1_2_1_17_1","volume-title":"AISTATS","author":"Guillory A.","year":"2009","unstructured":"A. Guillory , E. Chastain , and J. Bilmes . Active learning as non-convex optimization . In AISTATS , 2009 . A. Guillory, E. Chastain, and J. Bilmes. Active learning as non-convex optimization. In AISTATS, 2009."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661829.2661885"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/11748625_6"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2012.219"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2899409"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939502.2939511"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2645710.2645740"},{"key":"e_1_2_1_24_1","volume-title":"Arxiv","author":"Krishnan S.","year":"2015","unstructured":"S. Krishnan , J. Wang , E. Wu , M. J. Franklin , and K. Goldberg . Activeclean: Interactive data cleaning while learning convex loss models . In Arxiv , 2015 . S. Krishnan, J. Wang, E. Wu, M. J. Franklin, and K. Goldberg. Activeclean: Interactive data cleaning while learning convex loss models. In Arxiv, 2015."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735471.2735474"},{"key":"e_1_2_1_26_1","volume-title":"JMLR","author":"Nelson B.","year":"2012","unstructured":"B. Nelson , B. I. P. Rubinstein , L. Huang , A. D. Joseph , S. J. Lee , S. Rao , and J. D. Tygar . Query strategies for evading convex-inducing classifiers . In JMLR , 2012 . B. Nelson, B. I. P. Rubinstein, L. Huang, A. D. Joseph, S. J. Lee, S. Rao, and J. D. Tygar. Query strategies for evading convex-inducing classifiers. In JMLR, 2012."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2009.191"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2014.2359666"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1017\/S0266466603004110"},{"key":"e_1_2_1_30_1","volume-title":"Data cleaning: Problems and current approaches","author":"Rahm E.","year":"2000","unstructured":"E. Rahm and H. H. Do . Data cleaning: Problems and current approaches . In IEEE Data Eng. Bull ., 2000 . E. Rahm and H. H. Do. Data cleaning: Problems and current approaches. In IEEE Data Eng. Bull., 2000."},{"key":"e_1_2_1_31_1","volume-title":"Active learning literature survey","author":"Settles B.","year":"2010","unstructured":"B. Settles . Active learning literature survey . In University of Wisconsin , Madison , 2010 . B. Settles. Active learning literature survey. In University of Wisconsin, Madison, 2010."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.2517-6161.1951.tb00088.x"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2588555.2610505"},{"key":"e_1_2_1_34_1","volume-title":"ICML","author":"Xiao H.","year":"2015","unstructured":"H. Xiao , B. Biggio , G. Brown , G. Fumera , C. Eckert , and F. Roli . Is feature selection secure against training data poisoning ? In ICML , 2015 . H. Xiao, B. Biggio, G. Brown, G. Fumera, C. Eckert, and F. Roli. Is feature selection secure against training data poisoning? In ICML, 2015."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463676.2463706"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/1952376.1952378"},{"key":"e_1_2_1_37_1","volume-title":"ICML","author":"Zhao P.","year":"2015","unstructured":"P. Zhao and T. Zhang . Stochastic optimization with importance sampling for regularized loss minimization . In ICML , 2015 . P. Zhao and T. Zhang. Stochastic optimization with importance sampling for regularized loss minimization. In ICML, 2015."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/2994509.2994514","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:51:58Z","timestamp":1672224718000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/2994509.2994514"}},"subtitle":["interactive data cleaning for statistical modeling"],"short-title":[],"issued":{"date-parts":[[2016,8]]},"references-count":37,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2016,8]]}},"alternative-id":["10.14778\/2994509.2994514"],"URL":"https:\/\/doi.org\/10.14778\/2994509.2994514","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2016,8]]}}}