{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T04:27:55Z","timestamp":1777350475815,"version":"3.51.4"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"7","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2022,3]]},"abstract":"<jats:p>The lack of sufficient labeled data is a key bottleneck for practitioners in many real-world supervised machine learning (ML) tasks. In this paper, we study a new problem, namely<jats:italic>selective data acquisition in the wild for model charging<\/jats:italic>: given a supervised ML task and data in the wild (e.g., enterprise data warehouses, online data repositories, data markets, and so on), the problem is to select labeled data points from the data in the wild as additional train data that can help the ML task. It consists of two steps (Fig. 1). The first step is to discover relevant datasets (<jats:italic>e.g.<\/jats:italic>, tables with similar relational schema), which will result in a set of candidate datasets. Because these candidate datasets come from different sources and may follow different distributions, not all data points they contain can help. The second step is to select which data points from these candidate datasets should be used. We build an end-to-end solution. For step 1, we piggyback off-the-shelf data discovery tools. Technically, our focus is on step 2, for which we propose a solution framework called<jats:bold>AutoData.<\/jats:bold>It first clusters all data points from candidate datasets such that each cluster contains similar data points from different sources. It then iteratively picks which cluster to use, samples data points (<jats:italic>i.e.<\/jats:italic>, a mini-batch) from the picked cluster, evaluates the mini-batch, and then revises the search criteria by learning from the feedback (<jats:italic>i.e.<\/jats:italic>, reward) based on the evaluation. We propose a multi-armed bandit based solution and a Deep Q Networks-based reinforcement learning solution. Experiments using both relational and image datasets show the effectiveness of our solutions.<\/jats:p>","DOI":"10.14778\/3523210.3523223","type":"journal-article","created":{"date-parts":[[2022,6,22]],"date-time":"2022-06-22T22:23:21Z","timestamp":1655936601000},"page":"1466-1478","source":"Crossref","is-referenced-by-count":47,"title":["Selective data acquisition in the wild for model charging"],"prefix":"10.14778","volume":"15","author":[{"given":"Chengliang","family":"Chai","sequence":"first","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiabin","family":"Liu","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nan","family":"Tang","sequence":"additional","affiliation":[{"name":"QCRI, Doha, Qatar"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guoliang","family":"Li","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuyu","family":"Luo","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,6,22]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1890\/13-1452.1"},{"key":"e_1_2_1_2_1","doi-asserted-by":"crossref","unstructured":"Peter Auer. 2000. Using Upper Confidence Bounds for Online Learning. In FOCS. 270--279. Peter Auer. 2000. Using Upper Confidence Bounds for Online Learning. In FOCS . 270--279.","DOI":"10.1109\/SFCS.2000.892116"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2500490"},{"key":"e_1_2_1_4_1","unstructured":"Bing API. 2022. https:\/\/docs.microsoft.com\/en-us\/. Accessed: 2022-03-14. Bing API. 2022. https:\/\/docs.microsoft.com\/en-us\/. Accessed: 2022-03-14."},{"key":"e_1_2_1_5_1","volume-title":"Human-in-the-loop Outlier Detection. In SIGMOD Conference","author":"Chai Chengliang","year":"2020","unstructured":"Chengliang Chai , Lei Cao , Guoliang Li , Jian Li , Yuyu Luo , and Samuel Madden . 2020 . Human-in-the-loop Outlier Detection. In SIGMOD Conference 2020. 19--33. Chengliang Chai, Lei Cao, Guoliang Li, Jian Li, Yuyu Luo, and Samuel Madden. 2020. Human-in-the-loop Outlier Detection. In SIGMOD Conference 2020. 19--33."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2020.2972543"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2915252"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-018-0509-6"},{"key":"e_1_2_1_9_1","volume-title":"Data Management for Machine Learning: A Survey. TKDE","author":"Chai Chengliang","year":"2022","unstructured":"Chengliang Chai , Jiayi Wang , Yuyu Luo , Zeping Niu , and Guoliang Li. 2022. Data Management for Machine Learning: A Survey. TKDE ( 2022 ), 1--1. Chengliang Chai, Jiayi Wang, Yuyu Luo, Zeping Niu, and Guoliang Li. 2022. Data Management for Machine Learning: A Survey. TKDE (2022), 1--1."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.400568"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.14778\/3397230.3397235"},{"key":"e_1_2_1_12_1","first-page":"125","article-title":"Informativeness-Based Active Learning for Entity Resolution","volume":"1168","author":"Christen Victor","year":"2019","unstructured":"Victor Christen , Peter Christen , and Erhard Rahm . 2019 . Informativeness-Based Active Learning for Entity Resolution . In ECML PKDD , Vol. 1168. 125 -- 141 . Victor Christen, Peter Christen, and Erhard Rahm. 2019. Informativeness-Based Active Learning for Entity Resolution. In ECML PKDD, Vol. 1168. 125--141.","journal-title":"ECML PKDD"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s41019-021-00164-2"},{"key":"e_1_2_1_14_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1111\/j.2517-6161.1977.tb01600.x","article-title":"Maximum likelihood from incomplete data via the EM algorithm","volume":"39","author":"Dempster Arthur P","year":"1977","unstructured":"Arthur P et al. Dempster . 1977 . Maximum likelihood from incomplete data via the EM algorithm . Journal of the Royal Statistical Society 39 , 1 (1977), 1 -- 22 . Arthur P et al. Dempster. 1977. Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society 39, 1 (1977), 1--22.","journal-title":"Journal of the Royal Statistical Society"},{"key":"e_1_2_1_15_1","unstructured":"Martin Ester and Hans-Peter Kriegel etal 1996. A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise. In KDD-96. 226--231. Martin Ester and Hans-Peter Kriegel et al. 1996. A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise. In KDD-96 . 226--231."},{"key":"e_1_2_1_16_1","volume-title":"Aurum: A Data Discovery System. In ICDE. 1001--1012.","author":"Fernandez Raul Castro","year":"2018","unstructured":"Raul Castro Fernandez and Ziawasch Abedjan et al. 2018 . Aurum: A Data Discovery System. In ICDE. 1001--1012. Raul Castro Fernandez and Ziawasch Abedjan et al. 2018. Aurum: A Data Discovery System. In ICDE. 1001--1012."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.990138"},{"key":"e_1_2_1_18_1","unstructured":"Google Dataset Search API. 2022. https:\/\/developers.google.com\/search. Google Dataset Search API. 2022. https:\/\/developers.google.com\/search."},{"key":"e_1_2_1_19_1","volume-title":"Outdated Fact Detection in Knowledge Bases. In ICDE 2020","author":"Hao Shuang","year":"2020","unstructured":"Shuang Hao , Chengliang Chai , Guoliang Li , Nan Tang , Ning Wang , and Xiang Yu . 2020 . Outdated Fact Detection in Knowledge Bases. In ICDE 2020 . IEEE, 1890--1893. Shuang Hao, Chengliang Chai, Guoliang Li, Nan Tang, Ning Wang, and Xiang Yu. 2020. Outdated Fact Detection in Knowledge Bases. In ICDE 2020. IEEE, 1890--1893."},{"key":"e_1_2_1_20_1","first-page":"1599","article-title":"Q-learning algorithm using an adaptive-sized Q-table","volume":"2","author":"Hirashima Y.","year":"1999","unstructured":"Y. Hirashima , Y. Iiguni , A. Inoue , and S. Masuda . 1999 . Q-learning algorithm using an adaptive-sized Q-table . In IEEE CDC , Vol. 2. 1599 -- 1604 vol.2. Y. Hirashima, Y. Iiguni, A. Inoue, and S. Masuda. 1999. Q-learning algorithm using an adaptive-sized Q-table. In IEEE CDC, Vol. 2. 1599--1604 vol.2.","journal-title":"IEEE CDC"},{"key":"e_1_2_1_21_1","unstructured":"ImageNet. 2022. https:\/\/image-net.org\/. Accessed: 2022-03-14. ImageNet. 2022. https:\/\/image-net.org\/. Accessed: 2022-03-14."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.57"},{"key":"e_1_2_1_23_1","first-page":"1926","article-title":"CDB","volume":"11","author":"Li Guoliang","year":"2018","unstructured":"Guoliang Li and Chengliang Chai et al. 2018 . CDB : A Crowd-Powered Database System. PVLDB 11 , 12 (2018), 1926 -- 1929 . Guoliang Li and Chengliang Chai et al. 2018. CDB: A Crowd-Powered Database System. PVLDB 11, 12 (2018), 1926--1929.","journal-title":"A Crowd-Powered Database System. PVLDB"},{"key":"e_1_2_1_24_1","volume-title":"Adaptive Active Learning for Image Classification","author":"Li Xin","unstructured":"Xin Li and Yuhong Guo . 2013. Adaptive Active Learning for Image Classification . In IEEE CVPR. 859--866. Xin Li and Yuhong Guo. 2013. Adaptive Active Learning for Image Classification. In IEEE CVPR. 859--866."},{"key":"e_1_2_1_25_1","volume-title":"Data Acquisition for Improving Machine Learning Models. CoRR abs\/2105.14107","author":"Li Yifan","year":"2021","unstructured":"Yifan Li , Xiaohui Yu , and Nick Koudas . 2021. Data Acquisition for Improving Machine Learning Models. CoRR abs\/2105.14107 ( 2021 ). Yifan Li, Xiaohui Yu, and Nick Koudas. 2021. Data Acquisition for Improving Machine Learning Models. CoRR abs\/2105.14107 (2021)."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476311.3476333"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE48307.2020.00069"},{"key":"e_1_2_1_28_1","doi-asserted-by":"crossref","unstructured":"Yuyu Luo Nan Tang Guoliang Li Chengliang Chai Wenbo Li and Xuedi Qin. 2021. Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL Benchmarks. In SIGMOD' 21. 1235--1247. Yuyu Luo Nan Tang Guoliang Li Chengliang Chai Wenbo Li and Xuedi Qin. 2021. Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL Benchmarks. In SIGMOD' 21. 1235--1247.","DOI":"10.1145\/3448016.3457261"},{"key":"e_1_2_1_29_1","first-page":"6950","article-title":"Coresets for Data-efficient Training of Machine Learning Models","volume":"119","author":"Mirzasoleiman Baharan","year":"2020","unstructured":"Baharan Mirzasoleiman and Jeff A. Bilmes 2020 . Coresets for Data-efficient Training of Machine Learning Models . In ICML , Vol. 119. 6950 -- 6960 . Baharan Mirzasoleiman and Jeff A. Bilmes et al. 2020. Coresets for Data-efficient Training of Machine Learning Models. In ICML, Vol. 119. 6950--6960.","journal-title":"ICML"},{"key":"e_1_2_1_30_1","doi-asserted-by":"crossref","unstructured":"Volodymyr Mnih and Kavukcuoglu et al. 2015. Human-level control through deep reinforcement learning. nature 518 7540 (2015) 529--533. Volodymyr Mnih and Kavukcuoglu et al. 2015. Human-level control through deep reinforcement learning. nature 518 7540 (2015) 529--533.","DOI":"10.1038\/nature14236"},{"key":"e_1_2_1_31_1","unstructured":"Volodymyr Mnih and Koray Kavukcuoglusilvesilver et al. 2013. Playing Atari with Deep Reinforcement Learning. CoRR abs\/1312.5602 (2013). Volodymyr Mnih and Koray Kavukcuoglusilvesilver et al. 2013. Playing Atari with Deep Reinforcement Learning. CoRR abs\/1312.5602 (2013)."},{"key":"e_1_2_1_32_1","volume-title":"Stepleton","author":"Munos R\u00e9mi","year":"2016","unstructured":"R\u00e9mi Munos and Tom et al. Stepleton . 2016 . Safe and efficient off-policy reinforcement learning. arXiv preprint arXiv:1606.02647 (2016). R\u00e9mi Munos and Tom et al. Stepleton. 2016. Safe and efficient off-policy reinforcement learning. arXiv preprint arXiv:1606.02647 (2016)."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476299"},{"key":"e_1_2_1_34_1","volume-title":"Pu et al","author":"Nargesian Fatemeh","year":"2020","unstructured":"Fatemeh Nargesian and Ken Q . Pu et al . 2020 . Organizing Data Lakes for Navigation. In SIGMOD. 1939--1950. Fatemeh Nargesian and Ken Q. Pu et al. 2020. Organizing Data Lakes for Navigation. In SIGMOD. 1939--1950."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.14778\/3192965.3192973"},{"key":"e_1_2_1_36_1","volume-title":"From Cleaning before ML to Cleaning for ML","author":"Neutatz Felix","year":"2021","unstructured":"Felix Neutatz , Binger Chen , Ziawasch Abedjan , and Eugene Wu. 2021. From Cleaning before ML to Cleaning for ML . IEEE Data Eng. Bull . ( 2021 ). Felix Neutatz, Binger Chen, Ziawasch Abedjan, and Eugene Wu. 2021. From Cleaning before ML to Cleaning for ML. IEEE Data Eng. Bull. (2021)."},{"key":"e_1_2_1_37_1","unstructured":"Andrew Ng. 2021. MLOPs: From Model-centric to Data-centric AI. Andrew Ng. 2021. MLOPs: From Model-centric to Data-centric AI."},{"key":"e_1_2_1_38_1","unstructured":"NYU Auctus. 2022. https:\/\/auctus.vida-nyu.org\/. Accessed: 2022-01-14. NYU Auctus. 2022. https:\/\/auctus.vida-nyu.org\/. Accessed: 2022-01-14."},{"key":"e_1_2_1_39_1","volume-title":"Sample surveys: design, methods and applications","author":"Pfeffermann Danny","unstructured":"Danny Pfeffermann and Calyampudi Radhakrishna Rao . 2009. Sample surveys: design, methods and applications . Elsevier . Danny Pfeffermann and Calyampudi Radhakrishna Rao. 2009. Sample surveys: design, methods and applications. Elsevier."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3384695"},{"key":"e_1_2_1_41_1","volume-title":"ICDE 2021","author":"Qin Xuedi","year":"1973","unstructured":"Xuedi Qin and Chengliang Chai et al. 2021. Ranking Desired Tuples by Database Exploration . In ICDE 2021 . IEEE, 1973 --1978. Xuedi Qin and Chengliang Chai et al. 2021. Ranking Desired Tuples by Database Exploration. In ICDE 2021. IEEE, 1973--1978."},{"key":"e_1_2_1_42_1","volume-title":"Data Programming: Creating Large Training Sets, Quickly. arXiv:1605.07723","author":"Ratner Alexander","year":"2017","unstructured":"Alexander Ratner and Christopher De Sa 2017 . Data Programming: Creating Large Training Sets, Quickly. arXiv:1605.07723 Alexander Ratner and Christopher De Sa et al. 2017. Data Programming: Creating Large Training Sets, Quickly. arXiv:1605.07723"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-019-00552-1"},{"key":"e_1_2_1_44_1","unstructured":"Nicholas Roy and Andrew McCallum. 2001. Toward Optimal Active Learning through Sampling Estimation of Error Reduction. In (ICML. 441--448. Nicholas Roy and Andrew McCallum. 2001. Toward Optimal Active Learning through Sampling Estimation of Error Reduction. In (ICML. 441--448."},{"key":"e_1_2_1_45_1","doi-asserted-by":"crossref","unstructured":"Sunita Sarawagi and Anuradha Bhamidipaty. 2002. Interactive deduplication using active learning. In SIGKDD. ACM 269--278. Sunita Sarawagi and Anuradha Bhamidipaty. 2002. Interactive deduplication using active learning. In SIGKDD. ACM 269--278.","DOI":"10.1145\/775047.775087"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3068335"},{"key":"e_1_2_1_47_1","unstructured":"Ozan Sener and Silvio Savarese. 2018. Active Learning for Convolutional Neural Networks: A Core-Set Approach. In ICLR. Ozan Sener and Silvio Savarese. 2018. Active Learning for Convolutional Neural Networks: A Core-Set Approach. In ICLR."},{"key":"e_1_2_1_48_1","doi-asserted-by":"crossref","unstructured":"David Silver and Aja Huang et al. 2016. Mastering the game of Go with deep neural networks and tree search. Nat. 529 7587 (2016) 484--489. David Silver and Aja Huang et al. 2016. Mastering the game of Go with deep neural networks and tree search. Nat. 529 7587 (2016) 484--489.","DOI":"10.1038\/nature16961"},{"key":"e_1_2_1_49_1","unstructured":"Sklearn. 2022. https:\/\/scikit-learn.org\/stable\/modules\/generated\/sklearn.cluster.estimate_bandwidth.html. Accessed: 2022-03-14. Sklearn. 2022. https:\/\/scikit-learn.org\/stable\/modules\/generated\/sklearn.cluster.estimate_bandwidth.html. Accessed: 2022-03-14."},{"key":"e_1_2_1_50_1","volume-title":"Barto","author":"Sutton Richard S.","year":"1998","unstructured":"Richard S. Sutton and Andrew G . Barto . 1998 . Reinforcement learning - an introduction. MIT Press . Richard S. Sutton and Andrew G. Barto. 1998. Reinforcement learning - an introduction. MIT Press."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3452792"},{"key":"e_1_2_1_52_1","first-page":"66","article-title":"Dimensionality reduction: a comparative","volume":"10","author":"Der Maaten Laurens Van","year":"2009","unstructured":"Laurens Van Der Maaten , Eric Postma , Jaap Van den Herik , 2009 . Dimensionality reduction: a comparative . J Mach Learn Res 10 , 66 -- 71 (2009), 13. Laurens Van Der Maaten, Eric Postma, Jaap Van den Herik, et al. 2009. Dimensionality reduction: a comparative. J Mach Learn Res 10, 66--71 (2009), 13.","journal-title":"J Mach Learn Res"},{"key":"e_1_2_1_53_1","first-page":"437","article-title":"Multi-armed Bandit Algorithms and Empirical Evaluation","volume":"3720","author":"Vermorel Joann\u00e8s","year":"2005","unstructured":"Joann\u00e8s Vermorel and Mehryar Mohri . 2005 . Multi-armed Bandit Algorithms and Empirical Evaluation . In ECML , Vol. 3720. 437 -- 448 . Joann\u00e8s Vermorel and Mehryar Mohri. 2005. Multi-armed Bandit Algorithms and Empirical Evaluation. In ECML, Vol. 3720. 437--448.","journal-title":"ECML"},{"key":"e_1_2_1_54_1","first-page":"10842","article-title":"Data Valuation using Reinforcement Learning","volume":"119","author":"Yoon Jinsung","year":"2020","unstructured":"Jinsung Yoon , Sercan \u00d6mer Arik , and Tomas Pfister . 2020 . Data Valuation using Reinforcement Learning . In ICML , Vol. 119. 10842 -- 10851 . Jinsung Yoon, Sercan \u00d6mer Arik, and Tomas Pfister. 2020. Data Valuation using Reinforcement Learning. In ICML, Vol. 119. 10842--10851.","journal-title":"ICML"},{"key":"e_1_2_1_55_1","volume-title":"Agarwal et al","author":"Yu Hai","year":"2004","unstructured":"Hai Yu and Pankaj K . Agarwal et al . 2004 . Practical methods for shape fitting and kinetic data structures using core sets. In SCG. ACM , 263--272. Hai Yu and Pankaj K. Agarwal et al. 2004. Practical methods for shape fitting and kinetic data structures using core sets. In SCG. ACM, 263--272."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3523210.3523223","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,27]],"date-time":"2024-09-27T15:03:55Z","timestamp":1727449435000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3523210.3523223"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3]]},"references-count":55,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2022,3]]}},"alternative-id":["10.14778\/3523210.3523223"],"URL":"https:\/\/doi.org\/10.14778\/3523210.3523223","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2022,3]]}}}