{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T01:29:04Z","timestamp":1779326944262,"version":"3.51.4"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2022,9]]},"abstract":"<jats:p>Successful machine learning (ML) needs to learn from good data. However, one common issue about train data for ML practitioners is the lack of good features. To mitigate this problem, feature augmentation is often employed by joining with (or enriching features from) multiple tables, so as to become feature-rich ML. A consequent problem is that the enriched train data may contain too many tuples, especially if the feature augmentation is obtained through 1 (or many)-to-many or fuzzy joins. Training an ML model with a very large train dataset is data-inefficient. Coreset is often used to achieve data-efficient ML training, which selects a small subset of train data that can theoretically and practically perform similarly as using the full dataset. However, coreset selection over a large train dataset is also known to be time-consuming.<\/jats:p>\n          <jats:p>In this paper, we aim at achieving both feature-rich ML through feature augmentation and data-efficient ML through coreset selection. In order to avoid time-consuming coreset selection over a feature augmented (or fully materialized) table, we propose to efficiently select the coreset without materializing the augmented table. Note that coreset selection typically uses weighted gradients of the subset to approximate the full gradient of the entire train dataset. Our key idea is that the gradient computation for coreset selection of the augmented table can be pushed down to partial feature similarity of tuples within each individual table, without join materialization. These partial feature similarity values can be aggregated to estimate the gradient of the augmented table, which is upper bounded with provable theoretical guarantees. Extensive experiments show that our method can improve the efficiency by nearly 2 orders of magnitudes, while keeping almost the same accuracy as training with the fully augmented train data.<\/jats:p>","DOI":"10.14778\/3561261.3561267","type":"journal-article","created":{"date-parts":[[2022,11,16]],"date-time":"2022-11-16T15:32:50Z","timestamp":1668612770000},"page":"64-76","source":"Crossref","is-referenced-by-count":22,"title":["Coresets over multiple tables for feature-rich and data-efficient machine learning"],"prefix":"10.14778","volume":"16","author":[{"given":"Jiayi","family":"Wang","sequence":"first","affiliation":[{"name":"Tsinghua University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengliang","family":"Chai","sequence":"additional","affiliation":[{"name":"Tsinghua University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nan","family":"Tang","sequence":"additional","affiliation":[{"name":"QCRI"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiabin","family":"Liu","sequence":"additional","affiliation":[{"name":"Tsinghua University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guoliang","family":"Li","sequence":"additional","affiliation":[{"name":"Tsinghua University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,11,16]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2022. https:\/\/www.kaggle.com\/datasets\/olistbr\/brazilian-ecommerce\/. Accessed: 2022-04-28.  2022. https:\/\/www.kaggle.com\/datasets\/olistbr\/brazilian-ecommerce\/. Accessed: 2022-04-28."},{"key":"e_1_2_1_2_1","unstructured":"2022. Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning [Technical Report]. https:\/\/github.com\/for0nething\/RECON-TR\/blob\/main\/main.pdf. Last accessed: 2022-09-15.  2022. Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning [Technical Report]. https:\/\/github.com\/for0nething\/RECON-TR\/blob\/main\/main.pdf. Last accessed: 2022-09-15."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-021-00419-9"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.bdr.2015.04.001"},{"key":"e_1_2_1_5_1","volume-title":"Exploiting the structure: Stochastic gradient methods using raw clusters. Advances in Neural Information Processing Systems 29","author":"Allen-Zhu Zeyuan","year":"2016","unstructured":"Zeyuan Allen-Zhu , Yang Yuan , and Karthik Sridharan . 2016. Exploiting the structure: Stochastic gradient methods using raw clusters. Advances in Neural Information Processing Systems 29 ( 2016 ). Zeyuan Allen-Zhu, Yang Yuan, and Karthik Sridharan. 2016. Exploiting the structure: Stochastic gradient methods using raw clusters. Advances in Neural Information Processing Systems 29 (2016)."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.14778\/3007263.3007279"},{"key":"e_1_2_1_7_1","volume-title":"Convex optimization","author":"Boyd Stephen","unstructured":"Stephen Boyd , Stephen P Boyd , and Lieven Vandenberghe . 2004. Convex optimization . Cambridge university press . Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe. 2004. Convex optimization. Cambridge university press."},{"key":"e_1_2_1_8_1","volume-title":"New Frameworks for Offline and Streaming Coreset Constructions. CoRR abs\/1612.00889","author":"Braverman Vladimir","year":"2016","unstructured":"Vladimir Braverman , Dan Feldman , and Harry Lang . 2016. New Frameworks for Offline and Streaming Coreset Constructions. CoRR abs\/1612.00889 ( 2016 ). Vladimir Braverman, Dan Feldman, and Harry Lang. 2016. New Frameworks for Offline and Streaming Coreset Constructions. CoRR abs\/1612.00889 (2016)."},{"key":"e_1_2_1_9_1","volume-title":"ICML","volume":"80","author":"Campbell Trevor","year":"2018","unstructured":"Trevor Campbell and Tamara Broderick . 2018 . Bayesian Coreset Construction via Greedy Iterative Geodesic Ascent . In ICML 2018, Vol. 80 . PMLR, 697--705. Trevor Campbell and Tamara Broderick. 2018. Bayesian Coreset Construction via Greedy Iterative Geodesic Ascent. In ICML 2018, Vol. 80. PMLR, 697--705."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389772"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2915252"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/3523210.3523223"},{"key":"e_1_2_1_13_1","volume-title":"Data management for machine learning: A survey","author":"Chai Chengliang","year":"2022","unstructured":"Chengliang Chai , Jiayi Wang , Yuyu Luo , Zeping Niu , and Guoliang Li. 2022. Data management for machine learning: A survey . IEEE Transactions on Knowledge and Data Engineering ( 2022 ). Chengliang Chai, Jiayi Wang, Yuyu Luo, Zeping Niu, and Guoliang Li. 2022. Data management for machine learning: A survey. IEEE Transactions on Knowledge and Data Engineering (2022)."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137628.3137633"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.14778\/3397230.3397235"},{"key":"e_1_2_1_16_1","volume-title":"On a stochastic approximation method. The Annals of Mathematical Statistics","author":"Chung Kai Lai","year":"1954","unstructured":"Kai Lai Chung . 1954. On a stochastic approximation method. The Annals of Mathematical Statistics ( 1954 ), 463--483. Kai Lai Chung. 1954. On a stochastic approximation method. The Annals of Mathematical Statistics (1954), 463--483."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687553.1687576"},{"key":"e_1_2_1_18_1","volume-title":"Introduction to Core-sets: an Updated Survey. CoRR abs\/2011.09384","author":"Feldman Dan","year":"2020","unstructured":"Dan Feldman . 2020. Introduction to Core-sets: an Updated Survey. CoRR abs\/2011.09384 ( 2020 ). Dan Feldman. 2020. Introduction to Core-sets: an Updated Survey. CoRR abs\/2011.09384 (2020)."},{"key":"e_1_2_1_19_1","volume-title":"Computers and intractability","author":"Garey Michael R","unstructured":"Michael R Garey and David S Johnson . 1979. Computers and intractability . Vol. 174 . freeman San Francisco . Michael R Garey and David S Johnson. 1979. Computers and intractability. Vol. 174. freeman San Francisco."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944968"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.14778\/2367502.2367510"},{"key":"e_1_2_1_22_1","volume-title":"Variance reduced stochastic gradient descent with neighbors. Advances in Neural Information Processing Systems 28","author":"Hofmann Thomas","year":"2015","unstructured":"Thomas Hofmann , Aurelien Lucchi , Simon Lacoste-Julien , and Brian McWilliams . 2015. Variance reduced stochastic gradient descent with neighbors. Advances in Neural Information Processing Systems 28 ( 2015 ). Thomas Hofmann, Aurelien Lucchi, Simon Lacoste-Julien, and Brian McWilliams. 2015. Variance reduced stochastic gradient descent with neighbors. Advances in Neural Information Processing Systems 28 (2015)."},{"key":"e_1_2_1_23_1","volume-title":"ICML","volume":"139","author":"Huang Jiawei","year":"2021","unstructured":"Jiawei Huang , Ruomin Huang , Wenjie Liu , Nikolaos M. Freris , and Hu Ding . 2021 . A Novel Sequential Coreset Method for Gradient Descent Algorithms . In ICML 2021, Vol. 139 . PMLR, 4412--4422. Jiawei Huang, Ruomin Huang, Wenjie Liu, Nikolaos M. Freris, and Hu Ding. 2021. A Novel Sequential Coreset Method for Gradient Descent Algorithms. In ICML 2021, Vol. 139. PMLR, 4412--4422."},{"key":"e_1_2_1_24_1","volume-title":"Bilmes","author":"Iyer Rishabh K.","year":"2013","unstructured":"Rishabh K. Iyer and Jeff A . Bilmes . 2013 . Submodular Optimization with Submodular Cover and Submodular Knapsack Constraints. In NeurIPS 2013. 2436--2444. Rishabh K. Iyer and Jeff A. Bilmes. 2013. Submodular Optimization with Submodular Cover and Submodular Knapsack Constraints. In NeurIPS 2013. 2436--2444."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476311.3476372"},{"key":"e_1_2_1_26_1","volume-title":"Iyer","author":"Killamsetty KrishnaTeja","year":"2021","unstructured":"KrishnaTeja Killamsetty , Durga Sivasubramanian , Ganesh Ramakrishnan , and Rishabh K . Iyer . 2021 . GLISTER : Generalization based Data Subset Selection for Efficient and Robust Learning. In AAAI 2021,. AAAI Press , 8110--8118. KrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, and Rishabh K. Iyer. 2021. GLISTER: Generalization based Data Subset Selection for Efficient and Robust Learning. In AAAI 2021,. AAAI Press, 8110--8118."},{"key":"e_1_2_1_27_1","volume-title":"Submodularity for Data Selection in Machine Translation. In EMNLP","author":"Kirchhoff Katrin","year":"2014","unstructured":"Katrin Kirchhoff and Jeff A. Bilmes . 2014 . Submodularity for Data Selection in Machine Translation. In EMNLP 2014 . ACL, 131--141. Katrin Kirchhoff and Jeff A. Bilmes. 2014. Submodularity for Data Selection in Machine Translation. In EMNLP 2014. ACL, 131--141."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824087"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2723713"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2882952"},{"key":"e_1_2_1_31_1","first-page":"86","article-title":"A Survey on Advancing the DBMS Query Optimizer: Cardinality Estimation, Cost Model, and Plan Enumeration. Data Sci","volume":"6","author":"Lan Hai","year":"2021","unstructured":"Hai Lan , Zhifeng Bao , and Yuwei Peng . 2021 . A Survey on Advancing the DBMS Query Optimizer: Cardinality Estimation, Cost Model, and Plan Enumeration. Data Sci . Eng. 6 , 1 (2021), 86 -- 101 . Hai Lan, Zhifeng Bao, and Yuwei Peng. 2021. A Survey on Advancing the DBMS Query Optimizer: Cardinality Estimation, Cost Model, and Plan Enumeration. Data Sci. Eng. 6, 1 (2021), 86--101.","journal-title":"Eng."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.14778\/2850583.2850594"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3064036"},{"key":"e_1_2_1_34_1","volume-title":"Enabling and Optimizing NonLinear Feature Interactions in Factorized Linear Algebra. In SIGMOD","author":"Li Side","year":"2019","unstructured":"Side Li , Lingjiao Chen , and Arun Kumar . 2019 . Enabling and Optimizing NonLinear Feature Interactions in Factorized Linear Algebra. In SIGMOD 2019. 1571--1588. Side Li, Lingjiao Chen, and Arun Kumar. 2019. Enabling and Optimizing NonLinear Feature Interactions in Factorized Linear Algebra. In SIGMOD 2019. 1571--1588."},{"key":"e_1_2_1_35_1","volume-title":"Feature Augmentation with Reinforcement Learning. In ICDE 2022","author":"Liu Jiabin","year":"2022","unstructured":"Jiabin Liu , Chengliang Chai , Yuyu Luo , Yin Lou , Jianhua Feng , and Nan Tang . 2022 . Feature Augmentation with Reinforcement Learning. In ICDE 2022 , Kuala Lumpur, Malaysia, May 9--12 , 2022. IEEE, 3360--3372. Jiabin Liu, Chengliang Chai, Yuyu Luo, Yin Lou, Jianhua Feng, and Nan Tang. 2022. Feature Augmentation with Reinforcement Learning. In ICDE 2022, Kuala Lumpur, Malaysia, May 9--12, 2022. IEEE, 3360--3372."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476311.3476333"},{"key":"e_1_2_1_37_1","first-page":"1275","article-title":"Bao","volume":"2021","author":"Marcus Ryan","year":"2021","unstructured":"Ryan Marcus , Parimarjan Negi , Hongzi Mao , Nesime Tatbul , Mohammad Alizadeh , and Tim Kraska . 2021 . Bao : Making Learned Query Optimization Practical. In SIGMOD 2021. 1275 -- 1288 . Ryan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul, Mohammad Alizadeh, and Tim Kraska. 2021. Bao: Making Learned Query Optimization Practical. In SIGMOD 2021. 1275--1288.","journal-title":"Making Learned Query Optimization Practical. In SIGMOD"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v29i1.9486"},{"key":"e_1_2_1_39_1","volume-title":"Coresets for Data-efficient Training of Machine Learning Models. In ICML","volume":"119","author":"Mirzasoleiman Baharan","year":"2020","unstructured":"Baharan Mirzasoleiman , Jeff A. Bilmes , and Jure Leskovec . 2020 . Coresets for Data-efficient Training of Machine Learning Models. In ICML 2020, Vol. 119 . 6950--6960. Baharan Mirzasoleiman, Jeff A. Bilmes, and Jure Leskovec. 2020. Coresets for Data-efficient Training of Machine Learning Models. In ICML 2020, Vol. 119. 6950--6960."},{"key":"e_1_2_1_40_1","volume-title":"NeurIPS","author":"Mirzasoleiman Baharan","year":"2020","unstructured":"Baharan Mirzasoleiman , Kaidi Cao , and Jure Leskovec . 2020. Coresets for Robust Training of Deep Neural Networks against Noisy Labels . In NeurIPS 2020 . Baharan Mirzasoleiman, Kaidi Cao, and Jure Leskovec. 2020. Coresets for Robust Training of Deep Neural Networks against Noisy Labels. In NeurIPS 2020."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/s13218-017-0519-3"},{"key":"e_1_2_1_42_1","volume-title":"Stochastic optimization: algorithms and applications","author":"Nedi\u0107 Angelia","unstructured":"Angelia Nedi\u0107 and Dimitri Bertsekas . 2001. Convergence rate of incremental subgradient algorithms . In Stochastic optimization: algorithms and applications . Springer , 223--264. Angelia Nedi\u0107 and Dimitri Bertsekas. 2001. Convergence rate of incremental subgradient algorithms. In Stochastic optimization: algorithms and applications. Springer, 223--264."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.14778\/3007263.3007312"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3384695"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE51399.2021.00186"},{"key":"e_1_2_1_46_1","volume-title":"Proceedings of the statistical data analysis based on the L1 norm conference, neuchatel, switzerland","volume":"31","author":"Rdusseeun LKPJ","year":"1987","unstructured":"LKPJ Rdusseeun and P Kaufman . 1987 . Clustering by means of medoids . In Proceedings of the statistical data analysis based on the L1 norm conference, neuchatel, switzerland , Vol. 31 . LKPJ Rdusseeun and P Kaufman. 1987. Clustering by means of medoids. In Proceedings of the statistical data analysis based on the L1 norm conference, neuchatel, switzerland, Vol. 31."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.14778\/2535573.2488340"},{"key":"e_1_2_1_48_1","volume-title":"Fair-Batch: Batch Selection for Model Fairness. In ICLR","author":"Roh Yuji","year":"2021","unstructured":"Yuji Roh , Kangwook Lee , Steven Euijong Whang , and Changho Suh . 2021 . Fair-Batch: Batch Selection for Model Fairness. In ICLR 2021. OpenReview.net. Yuji Roh, Kangwook Lee, Steven Euijong Whang, and Changho Suh. 2021. Fair-Batch: Batch Selection for Model Fairness. In ICLR 2021. OpenReview.net."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2882939"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.14778\/3157794.3157804"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.14778\/3485450.3485458"},{"key":"e_1_2_1_52_1","volume-title":"DREW: Efficient Winograd CNN Inference with Deep Reuse. In WWW (WWW '22)","author":"Wu Ruofan","year":"2022","unstructured":"Ruofan Wu , Feng Zhang , Jiawei Guan , Zhen Zheng , Xiaoyong Du , and Xipeng Shen . 2022 . DREW: Efficient Winograd CNN Inference with Deep Reuse. In WWW (WWW '22) . Association for Computing Machinery , 1807--1816. Ruofan Wu, Feng Zhang, Jiawei Guan, Zhen Zheng, Xiaoyong Du, and Xipeng Shen. 2022. DREW: Efficient Winograd CNN Inference with Deep Reuse. In WWW (WWW '22). Association for Computing Machinery, 1807--1816."},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.14778\/3421424.3421432"},{"key":"e_1_2_1_54_1","first-page":"63","article-title":"A Survey of Traffic Prediction: from Spatio-Temporal Data to Intelligent Transportation. Data Sci","volume":"6","author":"Yuan Haitao","year":"2021","unstructured":"Haitao Yuan and Guoliang Li . 2021 . A Survey of Traffic Prediction: from Spatio-Temporal Data to Intelligent Transportation. Data Sci . Eng. 6 , 1 (2021), 63 -- 85 . Haitao Yuan and Guoliang Li. 2021. A Survey of Traffic Prediction: from Spatio-Temporal Data to Intelligent Transportation. Data Sci. Eng. 6, 1 (2021), 63--85.","journal-title":"Eng."},{"key":"e_1_2_1_55_1","first-page":"459","article-title":"POCLib:A high-performance framework for enabling near orthogonal processing on compression","volume":"33","author":"Zhang Feng","year":"2022","unstructured":"Feng Zhang , Jidong Zhai , Xipeng Shen , Onur Mutlu , and Xiaoyong Du . 2022 . POCLib:A high-performance framework for enabling near orthogonal processing on compression . TPDS 33 , 2 (2022), 459 -- 475 . Feng Zhang, Jidong Zhai, Xipeng Shen, Onur Mutlu, and Xiaoyong Du. 2022. POCLib:A high-performance framework for enabling near orthogonal processing on compression. TPDS 33, 2 (2022), 459--475.","journal-title":"TPDS"},{"key":"e_1_2_1_56_1","unstructured":"Bo Zhao and Hakan Bilen. 2021. Dataset condensation with differentiable siamese augmentation. In ICML. PMLR 12674--12685.  Bo Zhao and Hakan Bilen. 2021. Dataset condensation with differentiable siamese augmentation. In ICML. PMLR 12674--12685."},{"key":"e_1_2_1_57_1","first-page":"3","article-title":"Dataset Condensation with Gradient Matching","volume":"1","author":"Zhao Bo","year":"2021","unstructured":"Bo Zhao , Konda Reddy Mopuri , and Hakan Bilen . 2021 . Dataset Condensation with Gradient Matching . ICLR 1 , 2 (2021), 3 . Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen. 2021. Dataset Condensation with Gradient Matching. ICLR 1, 2 (2021), 3.","journal-title":"ICLR"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183739"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3561261.3561267","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T09:20:50Z","timestamp":1672219250000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3561261.3561267"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9]]},"references-count":58,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,9]]}},"alternative-id":["10.14778\/3561261.3561267"],"URL":"https:\/\/doi.org\/10.14778\/3561261.3561267","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2022,9]]}}}