{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,23]],"date-time":"2026-04-23T22:57:26Z","timestamp":1776985046640,"version":"3.51.4"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,7]]},"abstract":"<jats:p>In recent years, we have witnessed the development of novel data augmentation (DA) techniques for creating additional training data needed by machine learning based solutions. In this tutorial, we will provide a comprehensive overview of techniques developed by the data management community for data preparation and data integration. In addition to surveying task-specific DA operators that leverage rules, transformations, and external knowledge for creating additional training data, we also explore the advanced DA techniques such as interpolation, conditional generation, and DA policy learning. Finally, we describe the connection between DA and other machine learning paradigms such as active learning, pre-training, and weakly-supervised learning. We hope that this discussion can shed light on future research directions for a holistic data augmentation framework for high-quality dataset creation.<\/jats:p>","DOI":"10.14778\/3476311.3476403","type":"journal-article","created":{"date-parts":[[2021,10,28]],"date-time":"2021-10-28T22:48:56Z","timestamp":1635461336000},"page":"3182-3185","source":"Crossref","is-referenced-by-count":13,"title":["Data augmentation for ML-driven data preparation and integration"],"prefix":"10.14778","volume":"14","author":[{"given":"Yuliang","family":"Li","sequence":"first","affiliation":[{"name":"Megagon Labs"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaolan","family":"Wang","sequence":"additional","affiliation":[{"name":"Megagon Labs"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhengjie","family":"Miao","sequence":"additional","affiliation":[{"name":"Duke University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wang-Chiew","family":"Tan","sequence":"additional","affiliation":[{"name":"Facebook AI"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,28]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Ateret Anaby-Tavor Boaz Carmeli Esther Goldbraich Amir Kantor George Kour Segev Shlomov Naama Tepper and Naama Zwerdling. 2020. Do Not Have Enough Data? Deep Learning to the Rescue!. In AAAI. 7383--7390.  Ateret Anaby-Tavor Boaz Carmeli Esther Goldbraich Amir Kantor George Kour Segev Shlomov Naama Tepper and Naama Zwerdling. 2020. Do Not Have Enough Data? Deep Learning to the Rescue!. In AAAI . 7383--7390.","DOI":"10.1609\/aaai.v34i05.6233"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454741"},{"key":"e_1_2_1_3_1","unstructured":"Ursin Brunner and Kurt Stockinger. 2020. Entity matching with transformer architectures-a step forward in data integration. In EDBT.  Ursin Brunner and Kurt Stockinger. 2020. Entity matching with transformer architectures-a step forward in data integration. In EDBT ."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389742"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/1325851.1325891"},{"key":"e_1_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Jiaao Chen Zichao Yang and Diyi Yang. 2020. MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text Classification. In ACL. 2147--2157.  Jiaao Chen Zichao Yang and Diyi Yang. 2020. MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text Classification. In ACL. 2147--2157.","DOI":"10.18653\/v1\/2020.acl-main.194"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2013.6544847"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2749431"},{"key":"e_1_2_1_9_1","volume-title":"Autoaugment: Learning augmentation strategies from data. In CVPR. 113--123.","author":"Cubuk Ekin D","year":"2019","unstructured":"Ekin D Cubuk , Barret Zoph , Dandelion Mane , Vijay Vasudevan , and Quoc V Le . 2019 . Autoaugment: Learning augmentation strategies from data. In CVPR. 113--123. Ekin D Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V Le. 2019. Autoaugment: Learning augmentation strategies from data. In CVPR. 113--123."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW50498.2020.00359"},{"key":"e_1_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Xiang Dai and Heike Adel. 2020. An Analysis of Simple Data Augmentation for Named Entity Recognition. In COLING. 3861--3867.  Xiang Dai and Heike Adel. 2020. An Analysis of Simple Data Augmentation for Named Entity Recognition. In COLING . 3861--3867.","DOI":"10.18653\/v1\/2020.coling-main.343"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/3430915.3442430"},{"key":"e_1_2_1_13_1","volume-title":"Shafiq R. Joty, Luo Si, and Chunyan Miao.","author":"Ding Bosheng","year":"2020","unstructured":"Bosheng Ding , Linlin Liu , Lidong Bing , Canasai Kruengkrai , Thien Hai Nguyen , Shafiq R. Joty, Luo Si, and Chunyan Miao. 2020 . DAGA : Data Augmentation with a Generation Approach for Low-resource Tagging Tasks. In EMNLP. 6045--6057. Bosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai, Thien Hai Nguyen, Shafiq R. Joty, Luo Si, and Chunyan Miao. 2020. DAGA: Data Augmentation with a Generation Approach for Low-resource Tagging Tasks. In EMNLP. 6045--6057."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/376284.375731"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/2401764"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1046456.1046463"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3197387"},{"key":"e_1_2_1_18_1","volume-title":"Data augmentation for low-resource neural machine translation. arXiv preprint arXiv:1705.00440","author":"Fadaee Marzieh","year":"2017","unstructured":"Marzieh Fadaee , Arianna Bisazza , and Christof Monz . 2017. Data augmentation for low-resource neural machine translation. arXiv preprint arXiv:1705.00440 ( 2017 ). Marzieh Fadaee, Arianna Bisazza, and Christof Monz. 2017. Data augmentation for low-resource neural machine translation. arXiv preprint arXiv:1705.00440 (2017)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407802"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58607-2_30"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969033.2969125"},{"key":"e_1_2_1_22_1","volume-title":"Augmenting data with mixup for sentence classification: An empirical study. arXiv preprint arXiv:1905.08941","author":"Guo Hongyu","year":"2019","unstructured":"Hongyu Guo , Yongyi Mao , and Richong Zhang . 2019. Augmenting data with mixup for sentence classification: An empirical study. arXiv preprint arXiv:1905.08941 ( 2019 ). Hongyu Guo, Yongyi Mao, and Richong Zhang. 2019. Augmenting data with mixup for sentence classification: An empirical study. arXiv preprint arXiv:1905.08941 (2019)."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58595-2_1"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3193539"},{"key":"e_1_2_1_25_1","unstructured":"Jeffrey Heer Joseph M. Hellerstein and Sean Kandel. 2015. Predictive Interaction for Data Transformation. In CIDR.  Jeffrey Heer Joseph M. Hellerstein and Sean Kandel. 2015. Predictive Interaction for Data Transformation. In CIDR ."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3319888"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3455699"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3064034"},{"key":"e_1_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Jungo Kasai Kun Qian Sairam Gurajada Yunyao Li and Lucian Popa. 2019. Low-resource Deep Entity Resolution with Transfer and Active Learning. In ACL. 5851--5861.  Jungo Kasai Kun Qian Sairam Gurajada Yunyao Li and Lucian Popa. 2019. Low-resource Deep Entity Resolution with Transfer and Active Learning. In ACL . 5851--5861.","DOI":"10.18653\/v1\/P19-1586"},{"key":"e_1_2_1_30_1","volume-title":"Contextual Augmentation: Data Augmentation by Words with Paradigmatic Relations. In NAACL-HLT. 452--457.","author":"Kobayashi Sosuke","year":"2018","unstructured":"Sosuke Kobayashi . 2018 . Contextual Augmentation: Data Augmentation by Words with Paradigmatic Relations. In NAACL-HLT. 452--457. Sosuke Kobayashi. 2018. Contextual Augmentation: Data Augmentation by Words with Paradigmatic Relations. In NAACL-HLT. 452--457."},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the 2nd Workshop on Life-long Learning for Spoken Language Systems. 18--26","author":"Kumar Varun","year":"2020","unstructured":"Varun Kumar , Ashutosh Choudhary , and Eunah Cho . 2020 . Data Augmentation using Pre-trained Transformer Models . In Proceedings of the 2nd Workshop on Life-long Learning for Spoken Language Systems. 18--26 . Varun Kumar, Ashutosh Choudhary, and Eunah Cho. 2020. Data Augmentation using Pre-trained Transformer Models. In Proceedings of the 2nd Workshop on Life-long Learning for Spoken Language Systems. 18--26."},{"key":"e_1_2_1_32_1","volume-title":"Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461","author":"Lewis Mike","year":"2019","unstructured":"Mike Lewis , Yinhan Liu , Naman Goyal , Marjan Ghazvininejad , Abdelrahman Mohamed , Omer Levy , Ves Stoyanov , and Luke Zettlemoyer . 2019 . Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461 (2019). Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461 (2019)."},{"key":"e_1_2_1_33_1","volume-title":"DADA: Differentiable Automatic Data Augmentation. arXiv preprint arXiv:2003.03780","author":"Li Yonggang","year":"2020","unstructured":"Yonggang Li , Guosheng Hu , Yongtao Wang , Timothy Hospedales , Neil M Robertson , and Yongxing Yang . 2020 . DADA: Differentiable Automatic Data Augmentation. arXiv preprint arXiv:2003.03780 (2020). Yonggang Li, Guosheng Hu, Yongtao Wang, Timothy Hospedales, Neil M Robertson, and Yongxing Yang. 2020. DADA: Differentiable Automatic Data Augmentation. arXiv preprint arXiv:2003.03780 (2020)."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.14778\/3421424.3421431"},{"key":"e_1_2_1_35_1","volume-title":"Improved differentiable architecture search with early stopping. arXiv preprint arXiv:1909.06035","author":"Liang Hanwen","year":"2019","unstructured":"Hanwen Liang , Shifeng Zhang , Jiacheng Sun , Xingqiu He , Weiran Huang , Kechen Zhuang , and Zhenguo Li. 2019. Darts+ : Improved differentiable architecture search with early stopping. arXiv preprint arXiv:1909.06035 ( 2019 ). Hanwen Liang, Shifeng Zhang, Jiacheng Sun, Xingqiu He, Weiran Huang, Kechen Zhuang, and Zhenguo Li. 2019. Darts+: Improved differentiable architecture search with early stopping. arXiv preprint arXiv:1909.06035 (2019)."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454885"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W16-2914"},{"key":"e_1_2_1_38_1","volume-title":"DARTS: Differentiable Architecture Search. In ICLR.","author":"Liu Hanxiao","year":"2018","unstructured":"Hanxiao Liu , Karen Simonyan , and Yiming Yang . 2018 . DARTS: Differentiable Architecture Search. In ICLR. Hanxiao Liu, Karen Simonyan, and Yiming Yang. 2018. DARTS: Differentiable Architecture Search. In ICLR."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2005.39"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407801"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3324956"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3380597"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457258"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380144"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196926"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3358018"},{"key":"e_1_2_1_47_1","doi-asserted-by":"crossref","unstructured":"Tong Niu and Mohit Bansal. 2019. Automatically Learning Data Augmentation Policies for Dialogue Tasks. In EMNLP-IJCNLP. 1317--1323.  Tong Niu and Mohit Bansal. 2019. Automatically Learning Data Augmentation Policies for Dialogue Tasks. In EMNLP-IJCNLP . 1317--1323.","DOI":"10.18653\/v1\/D19-1132"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.14778\/3231751.3231757"},{"key":"e_1_2_1_49_1","volume-title":"The effectiveness of data augmentation in image classification using deep learning. arXiv preprint arXiv:1712.04621","author":"Perez Luis","year":"2017","unstructured":"Luis Perez and Jason Wang . 2017. The effectiveness of data augmentation in image classification using deep learning. arXiv preprint arXiv:1712.04621 ( 2017 ). Luis Perez and Jason Wang. 2017. The effectiveness of data augmentation in image classification using deep learning. arXiv preprint arXiv:1712.04621 (2017)."},{"key":"e_1_2_1_50_1","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford , Jeffrey Wu , Rewon Child , David Luan , Dario Amodei , and Ilya Sutskever . 2019 . Language models are unsupervised multitask learners . OpenAI Blog 1 , 8 (2019), 9 . Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.14778\/3157794.3157797"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.5555\/3294996.3295083"},{"key":"e_1_2_1_53_1","volume-title":"Proceedings of the NIPS workshop on cost-sensitive learning","author":"Settles Burr","year":"2008","unstructured":"Burr Settles , Mark Craven , and Lewis Friedland . 2008 . Active learning with real annotation costs . In Proceedings of the NIPS workshop on cost-sensitive learning . Vancouver, CA, 1--10. Burr Settles, Mark Craven, and Lewis Friedland. 2008. Active learning with real annotation costs. In Proceedings of the NIPS workshop on cost-sensitive learning. Vancouver, CA, 1--10."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.14778\/3397230.3397237"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.14778\/3457390.3457391"},{"key":"e_1_2_1_56_1","volume-title":"Mourad Ouzzani, Nan Tang, and Shafiq Joty.","author":"Thirumuruganathan Saravanan","year":"2018","unstructured":"Saravanan Thirumuruganathan , Shameem A Puthiya Parambath , Mourad Ouzzani, Nan Tang, and Shafiq Joty. 2018 . Reuse and adaptation for entity resolution through transfer learning. arXiv preprint arXiv:1809.11084 (2018). Saravanan Thirumuruganathan, Shameem A Puthiya Parambath, Mourad Ouzzani, Nan Tang, and Shafiq Joty. 2018. Reuse and adaptation for entity resolution through transfer learning. arXiv preprint arXiv:1809.11084 (2018)."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.14778\/3291264.3291268"},{"key":"e_1_2_1_58_1","volume-title":"Wei and Kai Zou","author":"Jason","year":"2019","unstructured":"Jason W. Wei and Kai Zou . 2019 . EDA : Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks. In EMNLP-IJCNLP. Association for Computational Linguistics , 6381--6387. Jason W. Wei and Kai Zou. 2019. EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks. In EMNLP-IJCNLP. Association for Computational Linguistics, 6381--6387."},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415562"},{"key":"e_1_2_1_60_1","volume-title":"Statistical Research Division","author":"Winkler William E","unstructured":"William E Winkler . 1999. The state of record linkage and current research problems . In Statistical Research Division , US Census Bureau . Citeseer. William E Winkler. 1999. The state of record linkage and current research problems. In Statistical Research Division, US Census Bureau. Citeseer."},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-22747-0_7"},{"key":"e_1_2_1_62_1","unstructured":"Qizhe Xie Zihang Dai Eduard H. Hovy Thang Luong and Quoc Le. 2020. Unsupervised Data Augmentation for Consistency Training. In NeurIPS.  Qizhe Xie Zihang Dai Eduard H. Hovy Thang Luong and Quoc Le. 2020. Unsupervised Data Augmentation for Consistency Training. In NeurIPS ."},{"key":"e_1_2_1_63_1","unstructured":"Yan Xu Ran Jia Lili Mou Ge Li Yunchuan Chen Yangyang Lu and Zhi Jin. 2016. Improved relation classification by deep recurrent neural networks with data augmentation. In COLING. 1461--1470.  Yan Xu Ran Jia Lili Mou Ge Li Yunchuan Chen Yangyang Lu and Zhi Jin. 2016. Improved relation classification by deep recurrent neural networks with data augmentation. In COLING . 1461--1470."},{"key":"e_1_2_1_64_1","unstructured":"Hongyi Zhang Moustapha Ciss\u00e9 Yann N. Dauphin and David Lopez-Paz. 2018. mixup: Beyond Empirical Risk Minimization. In ICLR.  Hongyi Zhang Moustapha Ciss\u00e9 Yann N. Dauphin and David Lopez-Paz. 2018. mixup: Beyond Empirical Risk Minimization. In ICLR ."},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313578"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3476311.3476403","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:42:37Z","timestamp":1672227757000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3476311.3476403"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7]]},"references-count":65,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2021,7]]}},"alternative-id":["10.14778\/3476311.3476403"],"URL":"https:\/\/doi.org\/10.14778\/3476311.3476403","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,7]]}}}