{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T18:20:40Z","timestamp":1775067640009,"version":"3.50.1"},"reference-count":76,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,11,1]],"date-time":"2023-11-01T00:00:00Z","timestamp":1698796800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["J. Data and Information Quality"],"published-print":{"date-parts":[[2023,12,31]]},"abstract":"<jats:p>Democratisation of machine learning (ML) has been an important theme in the research community for the last several years with notable progress made by the model-building community with automated machine learning models. However, data play a central role in building ML models and there is a need to focus on data-centric AI innovations. In this article, we first map the steps taken by data scientists for the data preparation phase and identify open areas and pain points via user interviews. We then propose a framework and four novel algorithms for exploratory data analysis and data quality for AI steps addressing the pain points from user interviews. We also validate our algorithms with open source datasets and show the effectiveness of our proposed methods. Next, we build a tool that automatically generates python code encompassing the above algorithms and study the usefulness of these algorithms via two user studies with data scientists. We observe from the first study results that the participants who used the tool were able to gain 2\u00d7 productivity and \u00a06% model improvement over the control group. The second study is performed in a more realistic environment to understand how the tool would be used in real-world scenarios. The results from this study are coherent with the first study and show an average of 30\u201350% of time savings that can be attributed to the tool.<\/jats:p>","DOI":"10.1145\/3603709","type":"journal-article","created":{"date-parts":[[2023,6,26]],"date-time":"2023-06-26T12:04:31Z","timestamp":1687781071000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":25,"title":["A Data-centric AI Framework for Automating Exploratory Data Analysis and Data Quality Tasks"],"prefix":"10.1145","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0031-8428","authenticated-orcid":false,"given":"Hima","family":"Patel","sequence":"first","affiliation":[{"name":"IBM Research, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-1982-024X","authenticated-orcid":false,"given":"Shanmukha","family":"Guttula","sequence":"additional","affiliation":[{"name":"IBM Research, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0177-6292","authenticated-orcid":false,"given":"Nitin","family":"Gupta","sequence":"additional","affiliation":[{"name":"IBM Research, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4986-0688","authenticated-orcid":false,"given":"Sandeep","family":"Hans","sequence":"additional","affiliation":[{"name":"IBM Research, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-2649-6844","authenticated-orcid":false,"given":"Ruhi Sharma","family":"Mittal","sequence":"additional","affiliation":[{"name":"IBM Research, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1322-7218","authenticated-orcid":false,"given":"Lokesh","family":"N","sequence":"additional","affiliation":[{"name":"IIT Bombay, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,11]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Kaggle Repository. Retrieved from https:\/\/www.kaggle.com."},{"key":"e_1_3_1_3_2","unstructured":"2020. Great Expectations. Retrieved from https:\/\/github.com\/great-expectations\/great_expectations."},{"key":"e_1_3_1_4_2","unstructured":"2020. Test Driven Data Analysis. Retrieved from https:\/\/github.com\/tdda."},{"key":"e_1_3_1_5_2","unstructured":"2022. Retrieved from https:\/\/pypi.org\/project\/nbformat\/."},{"key":"e_1_3_1_6_2","unstructured":"2022. Automated Machine Learning. Retrieved from https:\/\/www.datarobot.com\/platform\/automated-machine-learning\/."},{"key":"e_1_3_1_7_2","unstructured":"2022. Automated Machine Learning with Scikit-learn. Retrieved from https:\/\/github.com\/automl\/auto-sklearn."},{"key":"e_1_3_1_8_2","unstructured":"2022. AutoML Vision\u2014Google Cloud AutoML. Retrieved from https:\/\/cloud.google.com\/\/automl."},{"key":"e_1_3_1_9_2","unstructured":"2022. Cleaning Big Data: Most Time-Consuming Least Enjoyable Data Science Task Survey Says. Retrieved from https:\/\/www.forbes.com\/sites\/gilpress\/2016\/03\/23\/data-preparation-most-time-consuming-least-enjoyable-data-science-task-survey-says\/?sh=1304712a6f63."},{"key":"e_1_3_1_10_2","unstructured":"2022. Data Preparation: Time Consuming and Tedious? Retrieved from https:\/\/rapidminer.com\/blog\/data-prep-time-consuming-tedious\/."},{"key":"e_1_3_1_11_2","unstructured":"2022. IBM AutoAI. Retrieved from https:\/\/www.ibm.com\/in-en\/cloud\/watson-studio\/autoai."},{"key":"e_1_3_1_12_2","unstructured":"2022. A Python Automated Machine Learning Tool that Optimizes Machine Learning Pipelines Using Genetic Programming. Retrieved from https:\/\/github.com\/EpistasisLab\/tpot."},{"key":"e_1_3_1_13_2","unstructured":"2022. Scalable AutoML in H2O-3 Open Source. Retrieved from https:\/\/h2o.ai\/platform\/h2o-automl\/."},{"key":"e_1_3_1_14_2","volume-title":"SMDS\u201921","author":"Afzal Shazia","year":"2021","unstructured":"Shazia Afzal, C. Rajmohan, Manish Kesarwani, Sameep Mehta, and Hima Patel. 2021. Data readiness report. In SMDS\u201921. IEEE."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1080\/713827181"},{"issue":"175","key":"e_1_3_1_16_2","first-page":"1","article-title":"DataWig: Missing value imputation for tables","volume":"20","author":"Biessmann Felix","year":"2019","unstructured":"Felix Biessmann, Tammo Rukat, Phillipp Schmidt, Prathik Naidu, Sebastian Schelter, Andrey Taptunov, Dustin Lange, and David Salinas. 2019. DataWig: Missing value imputation for tables. J. Mach. Learn. Res. 20, 175 (2019), 1\u20136. http:\/\/jmlr.org\/papers\/v20\/18-753.html","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_1_17_2","volume-title":"CIKM\u201918","author":"Bie\u00dfmann Felix","year":"2018","unstructured":"Felix Bie\u00dfmann, David Salinas, Sebastian Schelter, Philipp Schmidt, and Dustin Lange. 2018. \u201cDeep\u201d learning for missing value imputationin tables with non-numerical data. In CIKM\u201918."},{"key":"e_1_3_1_18_2","volume-title":"CIDR\u201917","author":"Binnig Carsten","year":"2017","unstructured":"Carsten Binnig, Lorenzo De Stefani, Tim Kraska, Eli Upfal, Emanuel Zgraggen, and Zheguang Zhao. 2017. Toward sustainable insights, or why polygamy is bad for you. In CIDR\u201917."},{"key":"e_1_3_1_19_2","volume-title":"MLSys\u201919","author":"Breck Eric","year":"2019","unstructured":"Eric Breck, Neoklis Polyzotis, Sudip Roy, Steven Whang, and Martin Zinkevich. 2019. Data validation for machine learning. In MLSys\u201919."},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","DOI":"10.1613\/jair.606","article-title":"Identifying mislabeled training data","author":"Brodley Carla E.","year":"1999","unstructured":"Carla E. Brodley and Mark A. Friedl. 1999. Identifying mislabeled training data. J. Artif. Intell. Res. 11 (1999), 131\u2013167.","journal-title":"J. Artif. Intell. Res."},{"key":"e_1_3_1_21_2","article-title":"pandas-profiling: Exploratory Data Analysis for Python","author":"Brugman Simon","year":"2019","unstructured":"Simon Brugman. 2019. pandas-profiling: Exploratory Data Analysis for Python. Retrieved from https:\/\/github.com\/pandas-profiling\/pandas-profiling.","journal-title":"https:\/\/github.com\/pandas-profiling\/pandas-profiling"},{"key":"e_1_3_1_22_2","unstructured":"Dheeru Dua and Casey Graff. 2017. UCI Machine Learning Repository. Retrieved from http:\/\/archive.ics.uci.edu\/ml."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/2854006.2854008"},{"key":"e_1_3_1_24_2","unstructured":"Chen Feng Georgios Tzimiropoulos and Ioannis Patras. 2022. SSR: An Efficient and Robust Framework for Learning with Unknown Label Noise. arXiv preprint arXiv:2111.11288 (2021)."},{"key":"e_1_3_1_25_2","volume-title":"IDEAL\u201906","author":"Garc\u00eda Vicente","year":"2006","unstructured":"Vicente Garc\u00eda, Roberto Alejo, Jos\u00e9 Salvador S\u00e1nchez, Jos\u00e9 Mart\u00ednez Sotoca, and Ram\u00f3n Alberto Mollineda. 2006. Combined effects of class imbalance and class overlap on instance-based classification. In IDEAL\u201906."},{"key":"e_1_3_1_26_2","volume-title":"AAAI\u201917","author":"Ghosh Aritra","year":"2017","unstructured":"Aritra Ghosh, Himanshu Kumar, and P. Shanti Sastry. 2017. Robust loss functions under label noise for deep neural networks. In AAAI\u201917."},{"key":"e_1_3_1_27_2","volume-title":"PAKDD\u201918","author":"Gondara Lovedeep","year":"2018","unstructured":"Lovedeep Gondara and Ke Wang. 2018. MIDA: Multiple imputation using denoising autoencoders. In PAKDD\u201918."},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","first-page":"148","DOI":"10.1214\/aoms\/1177705148","article-title":"Snowball sampling","author":"Goodman Leo A.","year":"1961","unstructured":"Leo A. Goodman. 1961. Snowball sampling. Ann. Math. Stat. (1961), 148\u2013170.","journal-title":"Ann. Math. Stat."},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","unstructured":"Sandeep Hans Diptikalyan Saha and Aniya Aggarwal. 2022. Explainable data imputation using constraints.","DOI":"10.1145\/3570991.3571009"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1561\/1900000045"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2017.2685998"},{"key":"e_1_3_1_32_2","unstructured":"Ronny Kohavi and Barry Becker. 1996. Census Income Data. Retrieved from https:\/\/archive.ics.uci.edu\/ml\/datasets\/adult."},{"key":"e_1_3_1_33_2","volume-title":"Magellan: Toward Building Entity Matching Management Systems","author":"Konda Pradap Venkatramanan","year":"2018","unstructured":"Pradap Venkatramanan Konda. 2018. Magellan: Toward Building Entity Matching Management Systems. The University of Wisconsin\u2014Madison."},{"issue":"8","key":"e_1_3_1_34_2","article-title":"Matrix factorization techniques for recommender systems","volume":"42","author":"Koren Yehuda","year":"2009","unstructured":"Yehuda Koren, Robert M. Bell, and Chris Volinsky. 2009. Matrix factorization techniques for recommender systems. IEEE Comput. 42, 8 (2009).","journal-title":"IEEE Comput."},{"key":"e_1_3_1_35_2","article-title":"Boostclean: Automated error detection and repair for machine learning","author":"Krishnan Sanjay","year":"2017","unstructured":"Sanjay Krishnan, Michael J. Franklin, Ken Goldberg, and Eugene Wu. 2017. Boostclean: Automated error detection and repair for machine learning. arXiv:1711.01299. Retrieved from https:\/\/arxiv.org\/abs\/1711.01299.","journal-title":"arXiv:1711.01299"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2018.01.008"},{"key":"e_1_3_1_37_2","first-page":"316","volume-title":"CVPR\u201922","author":"Li Shikun","year":"2022","unstructured":"Shikun Li, Xiaobo Xia, Shiming Ge, and Tongliang Liu. 2022. Selective-supervised contrastive learning with noisy labels. In CVPR\u201922. 316\u2013325."},{"key":"e_1_3_1_38_2","first-page":"9089","volume-title":"CVPR\u201922","author":"Liang Kevin J.","year":"2022","unstructured":"Kevin J. Liang, Samrudhdhi B. Rangrej, Vladan Petrovic, and Tal Hassner. 2022. Few-shot learning with noisy labels. In CVPR\u201922. 9089\u20139098."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1002\/9781119013563"},{"key":"e_1_3_1_40_2","volume-title":"AAAI\u201920","author":"Liu Sijia","year":"2020","unstructured":"Sijia Liu, Parikshit Ram, Deepak Vijaykeerthy, Djallel Bouneffouf, Gregory Bramble, Horst Samulowitz, Dakuo Wang, Andrew Conn, and Alexander Gray. 2020. An ADMM based framework for automl pipeline configuration. In AAAI\u201920."},{"key":"e_1_3_1_41_2","volume-title":"ICML\u201919","author":"Mattei Pierre-Alexandre","year":"2019","unstructured":"Pierre-Alexandre Mattei and Jes Frellsen. 2019. MIWAE: Deep generative modelling and imputation of incomplete data sets. In ICML\u201919."},{"key":"e_1_3_1_42_2","article-title":"Spectral regularization algorithms for learning large incomplete matrices","author":"Mazumder Rahul","year":"2010","unstructured":"Rahul Mazumder, Trevor Hastie, and Robert Tibshirani. 2010. Spectral regularization algorithms for learning large incomplete matrices. J. Mach. Learn. Res. (2010).","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_1_43_2","volume-title":"ICLR\u201919","author":"Menon Aditya Krishna","year":"2019","unstructured":"Aditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, and Sanjiv Kumar. 2019. Can gradient clipping mitigate label noise? In ICLR\u201919."},{"key":"e_1_3_1_44_2","unstructured":"Tomas Mikolov Kai Chen Greg Corrado and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space."},{"issue":"3","key":"e_1_3_1_45_2","doi-asserted-by":"crossref","first-page":"411","DOI":"10.2307\/2980740","article-title":"Quota sampling","volume":"115","author":"Moser Claus Adolf","year":"1952","unstructured":"Claus Adolf Moser. 1952. Quota sampling. J. Roy. Stat. Soc. Ser. A (Gen.) 115, 3 (1952), 411\u2013423.","journal-title":"J. Roy. Stat. Soc. Ser. A (Gen.)"},{"key":"e_1_3_1_46_2","volume-title":"NIPS\u201913","author":"Natarajan Nagarajan","year":"2013","unstructured":"Nagarajan Natarajan, Inderjit S. Dhillon, Pradeep K. Ravikumar, and Ambuj Tewari. 2013. Learning with noisy labels. In NIPS\u201913."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/375360.375365"},{"key":"e_1_3_1_48_2","doi-asserted-by":"crossref","unstructured":"Alfredo Naz\u00e1bal Pablo M. Olmos Zoubin Ghahramani and Isabel Valera. 2018. Handling incomplete heterogeneous data using VAEs. Pattern Recognition 107 (2020) 107501.","DOI":"10.1016\/j.patcog.2020.107501"},{"key":"e_1_3_1_49_2","doi-asserted-by":"crossref","unstructured":"Curtis G. Northcutt Lu Jiang and Isaac L. Chuang. 2021. Confident learning: Estimating uncertainty in dataset labels. Journal of Artificial Intelligence Research 70 (2021) 1373\u20131411.","DOI":"10.1613\/jair.1.12125"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3056285"},{"key":"e_1_3_1_51_2","first-page":"1","article-title":"Stratified sampling","author":"Parsons Van L.","year":"2014","unstructured":"Van L. Parsons. 2014. Stratified sampling. In Wiley StatsRef: Statistics Reference Online, 1\u201311.","journal-title":"Wiley StatsRef: Statistics Reference Online"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1080\/14786440009463897"},{"key":"e_1_3_1_53_2","doi-asserted-by":"crossref","unstructured":"Kathy Razmadze Yael Amsterdamer Amit Somech Susan B. Davidson and Tova Milo. 2022. Selecting Sub-tables for Data Exploration. arXiv preprint (2022).","DOI":"10.1109\/ICDE55515.2023.00192"},{"key":"e_1_3_1_54_2","article-title":"Holoclean: Holistic data repairs with probabilistic inference","author":"Rekatsinas Theodoros","year":"2017","unstructured":"Theodoros Rekatsinas, Xu Chu, Ihab F. Ilyas, and Christopher R\u00e9. 2017. Holoclean: Holistic data repairs with probabilistic inference. arXiv:1702.00820. Retrieved from https:\/\/arxiv.org\/abs\/1702.00820.","journal-title":"arXiv:1702.00820"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2925300"},{"key":"e_1_3_1_56_2","volume-title":"CODS COMAD","author":"Saha Diptikalyan","year":"2022","unstructured":"Diptikalyan Saha, Aniya Aggarwal, and Sandeep Hans. 2022. Data synthesis for testing black-box machine learning models. In CODS COMAD."},{"key":"e_1_3_1_57_2","unstructured":"J. Sayyad Shirabad and T. J. Menzies. 2005. The PROMISE Repository of Software Engineering Databases. School of Information Technology and Engineering University of Ottawa Canada. Retrieved fromhttps:\/\/www.openml.org\/search?type=data&sort=runs&id=1063&status=active."},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1201\/9781439821862"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.14778\/3229863.3229867"},{"key":"e_1_3_1_60_2","article-title":"Convenience sampling","volume":"347","author":"Sedgwick Philip","year":"2013","unstructured":"Philip Sedgwick. 2013. Convenience sampling. Br. Med. J. 347 (2013).","journal-title":"Br. Med. J."},{"key":"e_1_3_1_61_2","first-page":"1644","volume-title":"Big Data\u201920","author":"Shrivastava Shrey","year":"2020","unstructured":"Shrey Shrivastava, Dhaval Patel, Nianjun Zhou, Arun Iyengar, and Anuradha Bhamidipaty. 2020. DQLearn: A toolkit for structured data quality learning. In Big Data\u201920. IEEE, 1644\u20131653."},{"key":"e_1_3_1_62_2","article-title":"Effortless data exploration with zenvisage: An expressive and interactive visual analytics system","author":"Siddiqui Tarique","year":"2016","unstructured":"Tarique Siddiqui, Albert Kim, John Lee, Karrie Karahalios, and Aditya Parameswaran. 2016. Effortless data exploration with zenvisage: An expressive and interactive visual analytics system. arXiv:1604.03583 (2016). Retrieved from https:\/\/arxiv.org\/abs\/1604.03583.","journal-title":"arXiv:1604.03583"},{"issue":"2","key":"e_1_3_1_63_2","doi-asserted-by":"crossref","first-page":"163","DOI":"10.1016\/0378-3758(77)90021-0","article-title":"New systematic sampling","volume":"1","author":"Singh D.","year":"1977","unstructured":"D. Singh and Padam Singh. 1977. New systematic sampling. J. Stat. Plan. Infer. 1, 2 (1977), 163\u2013177.","journal-title":"J. Stat. Plan. Infer."},{"key":"e_1_3_1_64_2","article-title":"Enriching data imputation under similarity rule constraints","author":"Song Shaoxu","year":"2018","unstructured":"Shaoxu Song, Yu Sun, Aoqian Zhang, Lei Chen, and Jianmin Wang. 2018. Enriching data imputation under similarity rule constraints. In TKDE\u201918.","journal-title":"TKDE\u201918"},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","DOI":"10.1093\/bioinformatics\/btr597","article-title":"MissForest\u2014Non-parametric missing value imputation for mixed-type data","author":"Stekhoven Daniel J.","year":"2012","unstructured":"Daniel J. Stekhoven and Peter B\u00fchlmann. 2012. MissForest\u2014Non-parametric missing value imputation for mixed-type data. Bioinformatics (2012).","journal-title":"Bioinformatics"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1086\/224909"},{"key":"e_1_3_1_67_2","article-title":"Missing value estimation methods for DNA microarrays","volume":"17","author":"Troyanskaya Olga G.","year":"2001","unstructured":"Olga G. Troyanskaya, Michael N. Cantor, Gavin Sherlock, Patrick O. Brown, Trevor Hastie, Robert Tibshirani, David Botstein, and Russ B. Altman. 2001. Missing value estimation methods for DNA microarrays. Bioinformatics 17 (2001).","journal-title":"Bioinformatics"},{"key":"e_1_3_1_68_2","volume-title":"VLDB\u201915","author":"Vartak Manasi","year":"2015","unstructured":"Manasi Vartak, Sajjadur Rahman, Samuel Madden, Aditya Parameswaran, and Neoklis Polyzotis. 2015. Seedb: Efficient data-driven visualization recommendations to support visual analytics. In VLDB\u201915."},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/358105.893"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2020.106631"},{"key":"e_1_3_1_71_2","volume-title":"ICEBI\u201910","author":"Xiong Haitao","year":"2010","unstructured":"Haitao Xiong, Junjie Wu, and Lu Liu. 2010. Classification with class overlapping: A systematic study. In ICEBI\u201910."},{"key":"e_1_3_1_72_2","first-page":"553","volume-title":"SIGMOD\u201913","author":"Yakout Mohamed","year":"2013","unstructured":"Mohamed Yakout, Laure Berti-\u00c9quille, and Ahmed K. Elmagarmid. 2013. Don\u2019t be scared: Use scalable automatic repairing with maximal likelihood and bounded changes. In SIGMOD\u201913. 553\u2013564."},{"key":"e_1_3_1_73_2","first-page":"7017","volume-title":"CVPR\u201919","author":"Yi Kun","year":"2019","unstructured":"Kun Yi and Jianxin Wu. 2019. Probabilistic end-to-end noise correction for learning with noisy labels. In CVPR\u201919. 7017\u20137025."},{"key":"e_1_3_1_74_2","volume-title":"ICML\u201918","author":"Yoon Jinsung","year":"2018","unstructured":"Jinsung Yoon, James Jordon, and Mihaela van der Schaar. 2018. GAIN: Missing data imputation using generative adversarial nets. In ICML\u201918."},{"key":"e_1_3_1_75_2","volume-title":"ICBCB\u201917","author":"Yu HaiYue","year":"2017","unstructured":"HaiYue Yu and KunHong Liu. 2017. Classification of multi-class microarray datasets using a minimizing class-overlapping based ECOC algorithm. In ICBCB\u201917."},{"key":"e_1_3_1_76_2","unstructured":"Hongbao Zhang Pengtao Xie and Eric P. Xing. 2018. Missing value imputation based on deep generative models."},{"key":"e_1_3_1_77_2","first-page":"527","volume-title":"SIGMOD\u201917","author":"Zhao Zheguang","year":"2017","unstructured":"Zheguang Zhao, Lorenzo De Stefani, Emanuel Zgraggen, Carsten Binnig, Eli Upfal, and Tim Kraska. 2017. Controlling false discoveries during interactive data exploration. In SIGMOD\u201917. 527\u2013540."}],"container-title":["Journal of Data and Information Quality"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3603709","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3603709","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:21Z","timestamp":1750178241000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3603709"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11]]},"references-count":76,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,12,31]]}},"alternative-id":["10.1145\/3603709"],"URL":"https:\/\/doi.org\/10.1145\/3603709","relation":{},"ISSN":["1936-1955","1936-1963"],"issn-type":[{"value":"1936-1955","type":"print"},{"value":"1936-1963","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11]]},"assertion":[{"value":"2022-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-04-11","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}