{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,30]],"date-time":"2026-01-30T01:49:18Z","timestamp":1769737758011,"version":"3.49.0"},"reference-count":28,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:p>Data quality assessment is an essential process of any data analysis process including machine learning. The process is time-consuming as it involves multiple independent data quality checks that are performed iteratively at scale on evolving data resulting from exploratory data analysis (EDA). Existing solutions that provide computational optimizations for data quality assessment often separate the data structure from its data quality which then requires efforts from users to explicitly maintain state-like information. They demand a certain level of distributed system knowledge to ensure high-level pipeline optimizations from data analysts who should instead be focusing on analyzing the data. We, therefore, propose data-quality-aware dataframes, a data quality management system embedded as part of a data analyst's familiar data structure, such as a Python dataframe. The framework automatically detects changes in datasets' metadata and exploits the context of each of the quality checks to provide efficient data quality assessment on ever-changing data. We demonstrate in our experiment that our approach can reduce the overall data quality evaluation runtime by 40-80% in both local and distributed setups with less than 10% increase in memory usage.<\/jats:p>","DOI":"10.14778\/3503585.3503602","type":"journal-article","created":{"date-parts":[[2022,4,14]],"date-time":"2022-04-14T22:18:07Z","timestamp":1649974687000},"page":"949-957","source":"Crossref","is-referenced-by-count":3,"title":["DQDF"],"prefix":"10.14778","volume":"15","author":[{"given":"Phanwadee","family":"Sinthong","sequence":"first","affiliation":[{"name":"University of California"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dhaval","family":"Patel","sequence":"additional","affiliation":[{"name":"IBM TJ Watson Research Center"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nianjun","family":"Zhou","sequence":"additional","affiliation":[{"name":"IBM TJ Watson Research Center"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shrey","family":"Shrivastava","sequence":"additional","affiliation":[{"name":"IBM TJ Watson Research Center"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arun","family":"Iyengar","sequence":"additional","affiliation":[{"name":"IBM TJ Watson Research Center"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anuradha","family":"Bhamidipaty","sequence":"additional","affiliation":[{"name":"IBM TJ Watson Research Center"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,4,14]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2021. Apache Spark. http:\/\/spark.apache.org\/.  2021. Apache Spark. http:\/\/spark.apache.org\/."},{"key":"e_1_2_1_2_1","unstructured":"2021. Dask. http:\/\/dask.org\/.  2021. Dask. http:\/\/dask.org\/."},{"key":"e_1_2_1_3_1","unstructured":"2021. Dask Common Uses. https:\/\/docs.dask.org\/en\/latest\/dataframe.html#common-uses-and-anti-uses\/.  2021. Dask Common Uses. https:\/\/docs.dask.org\/en\/latest\/dataframe.html#common-uses-and-anti-uses\/."},{"key":"e_1_2_1_4_1","unstructured":"2021. Data Cleaning in Python. https:\/\/towardsdatascience.com\/data-cleaning-in-python-the-ultimate-guide-2020-c63b88bf0a0d.  2021. Data Cleaning in Python. https:\/\/towardsdatascience.com\/data-cleaning-in-python-the-ultimate-guide-2020-c63b88bf0a0d."},{"key":"e_1_2_1_5_1","unstructured":"2021. DQDF Case Study. https:\/\/github.com\/psinthong\/DQDF_Case_Study.  2021. DQDF Case Study. https:\/\/github.com\/psinthong\/DQDF_Case_Study."},{"key":"e_1_2_1_6_1","unstructured":"2021. The incredible growth of Python. https:\/\/stackoverflow.blog\/2017\/09\/06\/incredible-growth-python\/.  2021. The incredible growth of Python. https:\/\/stackoverflow.blog\/2017\/09\/06\/incredible-growth-python\/."},{"key":"e_1_2_1_7_1","unstructured":"2021. Koalas. http:\/\/koalas.readthedocs.io.  2021. Koalas. http:\/\/koalas.readthedocs.io."},{"key":"e_1_2_1_8_1","unstructured":"2021. Modin. https:\/\/modin.readthedocs.io\/en\/latest\/.  2021. Modin. https:\/\/modin.readthedocs.io\/en\/latest\/."},{"key":"e_1_2_1_9_1","unstructured":"2021. Numpy. https:\/\/numpy.org\/.  2021. Numpy. https:\/\/numpy.org\/."},{"key":"e_1_2_1_10_1","unstructured":"2021. Pandas. http:\/\/pandas.pydata.org\/.  2021. Pandas. http:\/\/pandas.pydata.org\/."},{"key":"e_1_2_1_11_1","unstructured":"2021. Pandas Makes Python Better. https:\/\/towardsdatascience.com\/pandas-makes-python-better-ec6cc1e30233\/.  2021. Pandas Makes Python Better. https:\/\/towardsdatascience.com\/pandas-makes-python-better-ec6cc1e30233\/."},{"key":"e_1_2_1_12_1","unstructured":"2021. PandasSchema. https:\/\/tmiguelt.github.io\/PandasSchema\/.  2021. PandasSchema. https:\/\/tmiguelt.github.io\/PandasSchema\/."},{"key":"e_1_2_1_13_1","unstructured":"2021. Sberbank Russian Housing Market. https:\/\/www.kaggle.com\/c\/sberbank-russian-housing-market\/overview\/description.  2021. Sberbank Russian Housing Market. https:\/\/www.kaggle.com\/c\/sberbank-russian-housing-market\/overview\/description."},{"key":"e_1_2_1_14_1","unstructured":"2021. Scikit-learn. https:\/\/scikit-learn.org\/stable\/\/.  2021. Scikit-learn. https:\/\/scikit-learn.org\/stable\/\/."},{"key":"e_1_2_1_15_1","unstructured":"2021. Vaex. https:\/\/github.com\/vaexio\/vaex.  2021. Vaex. https:\/\/github.com\/vaexio\/vaex."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824080"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.25080\/Majora-342d178e-010"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098021"},{"key":"e_1_2_1_19_1","volume-title":"Data Infrastructure for Machine Learning. In SysML Conference.","author":"Breck Eric","year":"2018","unstructured":"Eric Breck , Neoklis Polyzotis , Sudip Roy , Steven Euijong Whang , and Martin Zinkevich . 2018 . Data Infrastructure for Machine Learning. In SysML Conference. Eric Breck, Neoklis Polyzotis, Sudip Roy, Steven Euijong Whang, and Martin Zinkevich. 2018. Data Infrastructure for Machine Learning. In SysML Conference."},{"key":"e_1_2_1_20_1","unstructured":"Jason Brownlee. 2020. Data preparation for machine learning: data cleaning feature selection and data transforms in Python. Machine Learning Mastery.  Jason Brownlee. 2020. Data preparation for machine learning: data cleaning feature selection and data transforms in Python. Machine Learning Mastery."},{"key":"e_1_2_1_21_1","volume-title":"The challenges of data quality and data quality assessment in the big data era. Data science journal 14","author":"Cai Li","year":"2015","unstructured":"Li Cai and Yangyong Zhu . 2015. The challenges of data quality and data quality assessment in the big data era. Data science journal 14 ( 2015 ). Li Cai and Yangyong Zhu. 2015. The challenges of data quality and data quality assessment in the big data era. Data science journal 14 (2015)."},{"key":"e_1_2_1_22_1","volume-title":"TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines. In ACM International Conference on Management of Data (SIGMOD). 2793--2796","author":"Caveness Emily","year":"2020","unstructured":"Emily Caveness , Paul Suganthan GC , Zhuo Peng , Neoklis Polyzotis , Sudip Roy , and Martin Zinkevich . 2020 . TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines. In ACM International Conference on Management of Data (SIGMOD). 2793--2796 . Emily Caveness, Paul Suganthan GC, Zhuo Peng, Neoklis Polyzotis, Sudip Roy, and Martin Zinkevich. 2020. TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines. In ACM International Conference on Management of Data (SIGMOD). 2793--2796."},{"key":"e_1_2_1_23_1","unstructured":"David J DeWitt. 1993. The Wisconsin benchmark: Past present and future. In The Benchmark Handbook J. Gray (Ed.). Morgan Kaufmann.  David J DeWitt. 1993. The Wisconsin benchmark: Past present and future. In The Benchmark Handbook J. Gray (Ed.). Morgan Kaufmann."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407807"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/505248.506010"},{"key":"e_1_2_1_26_1","volume-title":"Differential Data Quality Verification on Partitioned Data. In IEEE International Conference on Data Engineering (ICDE). 1940--1945","author":"Schelter Sebastian","year":"2019","unstructured":"Sebastian Schelter , Stefan Grafberger , Philipp Schmidt , Tammo Rukat , Mario Kiessling , Andrey Taptunov , Felix Biessmann , and Dustin Lange . 2019 . Differential Data Quality Verification on Partitioned Data. In IEEE International Conference on Data Engineering (ICDE). 1940--1945 . Sebastian Schelter, Stefan Grafberger, Philipp Schmidt, Tammo Rukat, Mario Kiessling, Andrey Taptunov, Felix Biessmann, and Dustin Lange. 2019. Differential Data Quality Verification on Partitioned Data. In IEEE International Conference on Data Engineering (ICDE). 1940--1945."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.14778\/3229863.3229867"},{"key":"e_1_2_1_28_1","volume-title":"Automated and Interactive Data Quality Advisor. In IEEE International Conference on Big Data (Big Data). 2913--2922","author":"Shrivastava Shrey","year":"2019","unstructured":"Shrey Shrivastava , Dhaval Patel , Anuradha Bhamidipaty , Wesley M Gifford , Stuart A Siegel , Venkata Sitaramagiridharganesh Ganapavarapu , and Jayant R Kalagnanam . 2019 . DQA: Scalable , Automated and Interactive Data Quality Advisor. In IEEE International Conference on Big Data (Big Data). 2913--2922 . Shrey Shrivastava, Dhaval Patel, Anuradha Bhamidipaty, Wesley M Gifford, Stuart A Siegel, Venkata Sitaramagiridharganesh Ganapavarapu, and Jayant R Kalagnanam. 2019. DQA: Scalable, Automated and Interactive Data Quality Advisor. In IEEE International Conference on Big Data (Big Data). 2913--2922."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3503585.3503602","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:31:38Z","timestamp":1672223498000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3503585.3503602"}},"subtitle":["data-quality-aware dataframes"],"short-title":[],"issued":{"date-parts":[[2021,12]]},"references-count":28,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["10.14778\/3503585.3503602"],"URL":"https:\/\/doi.org\/10.14778\/3503585.3503602","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,12]]}}}