{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T21:45:02Z","timestamp":1778535902354,"version":"3.51.4"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"11","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,7]]},"abstract":"<jats:p>Cutting-edge machine learning techniques often require millions of labeled data objects to train a robust model. Because relying on humans to supply such a huge number of labels is rarely practical, automated methods for label generation are needed. Unfortunately, critical challenges in auto-labeling remain unsolved, including the following research questions: (1) which objects to ask humans to label, (2) how to automatically propagate labels to other objects, and (3) when to stop labeling. These three questions are not only each challenging in their own right, but they also correspond to tightly interdependent problems. Yet existing techniques provide at best isolated solutions to a subset of these challenges. In this work, we propose the first approach, called LANCET, that successfully addresses all three challenges in an integrated framework. LANCET is based on a theoretical foundation characterizing the properties that the labeled dataset must satisfy to train an effective prediction model, namely the Covariate-shift and the Continuity conditions. First, guided by the Covariate-shift condition, LANCET maps raw input data into a semantic feature space, where an unlabeled object is expected to share the same label with its near-by labeled neighbor. Next, guided by the Continuity condition, LANCET selects objects for labeling, aiming to ensure that unlabeled objects always have some sufficiently close labeled neighbors. These two strategies jointly maximize the accuracy of the automatically produced labels and the prediction accuracy of the machine learning models trained on these labels. Lastly, LANCET uses a distribution matching network to verify whether both the Covariate-shift and Continuity conditions hold, in which case it would be safe to terminate the labeling process. Our experiments on diverse public data sets demonstrate that LANCET consistently outperforms the state-of-the-art methods from Snuba to GOGGLES and other baselines by a large margin - up to 30 percentage points increase in accuracy.<\/jats:p>","DOI":"10.14778\/3476249.3476269","type":"journal-article","created":{"date-parts":[[2021,10,27]],"date-time":"2021-10-27T16:46:23Z","timestamp":1635353183000},"page":"2154-2166","source":"Crossref","is-referenced-by-count":9,"title":["LANCET"],"prefix":"10.14778","volume":"14","author":[{"given":"Huayi","family":"Zhang","sequence":"first","affiliation":[{"name":"Worcester Polytechnic Institute"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lei","family":"Cao","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Samuel","family":"Madden","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elke","family":"Rundensteiner","sequence":"additional","affiliation":[{"name":"Worcester Polytechnic Institute"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,27]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2020. LANCET: Labeling Complex Data at Scale (Extended Version). https:\/\/drive.google.com\/file\/d\/1cX8l-CyPyoqW4DGrIAgvfaTvRfV4Twx5\/view?usp=sharing.  2020. LANCET: Labeling Complex Data at Scale (Extended Version). https:\/\/drive.google.com\/file\/d\/1cX8l-CyPyoqW4DGrIAgvfaTvRfV4Twx5\/view?usp=sharing."},{"key":"e_1_2_1_2_1","first-page":"3","article-title":"A public domain dataset for human activity recognition using smartphones","volume":"3","author":"Anguita Davide","year":"2013","unstructured":"Davide Anguita , Alessandro Ghio , Luca Oneto , Xavier Parra , and Jorge Luis Reyes-Ortiz . 2013 . A public domain dataset for human activity recognition using smartphones .. In Esann , Vol. 3. 3 . Davide Anguita, Alessandro Ghio, Luca Oneto, Xavier Parra, and Jorge Luis Reyes-Ortiz. 2013. A public domain dataset for human activity recognition using smartphones.. In Esann, Vol. 3. 3.","journal-title":"Esann"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-009-5152-4"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.14778\/3352063.3352138"},{"key":"e_1_2_1_5_1","volume-title":"A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709","author":"Chen Ting","year":"2020","unstructured":"Ting Chen , Simon Kornblith , Mohammad Norouzi , and Geoffrey Hinton . 2020. A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709 ( 2020 ). Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709 (2020)."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295397"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3380592"},{"key":"e_1_2_1_8_1","doi-asserted-by":"crossref","unstructured":"J. Deng W. Dong R. Socher L.-J. Li K. Li and L. Fei-Fei. 2009. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09.  J. Deng W. Dong R. Socher L.-J. Li K. Li and L. Fei-Fei. 2009. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_9_1","volume-title":"Density estimation using real nvp. arXiv preprint arXiv:1605.08803","author":"Dinh Laurent","year":"2016","unstructured":"Laurent Dinh , Jascha Sohl-Dickstein , and Samy Bengio . 2016. Density estimation using real nvp. arXiv preprint arXiv:1605.08803 ( 2016 ). Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. 2016. Density estimation using real nvp. arXiv preprint arXiv:1605.08803 (2016)."},{"key":"e_1_2_1_10_1","volume-title":"Adversarial active learning for deep networks: a margin based approach. arXiv preprint arXiv:1802.09841","author":"Ducoffe Melanie","year":"2018","unstructured":"Melanie Ducoffe and Frederic Precioso . 2018. Adversarial active learning for deep networks: a margin based approach. arXiv preprint arXiv:1802.09841 ( 2018 ). Melanie Ducoffe and Frederic Precioso. 2018. Adversarial active learning for deep networks: a margin based approach. arXiv preprint arXiv:1802.09841 (2018)."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/3045390.3045502"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305381.3305504"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/3045118.3045244"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/2946645.2946704"},{"key":"e_1_2_1_15_1","volume-title":"Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728","author":"Gidaris Spyros","year":"2018","unstructured":"Spyros Gidaris , Praveer Singh , and Nikos Komodakis . 2018. Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728 ( 2018 ). Spyros Gidaris, Praveer Singh, and Nikos Komodakis. 2018. Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728 (2018)."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/3086952"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"e_1_2_1_18_1","unstructured":"Irina Higgins Loic Matthey Arka Pal Christopher Burgess Xavier Glorot Matthew Botvinick Shakir Mohamed and Alexander Lerchner. 2016. beta-vae: Learning basic visual concepts with a constrained variational framework. (2016).  Irina Higgins Loic Matthey Arka Pal Christopher Burgess Xavier Glorot Matthew Botvinick Shakir Mohamed and Alexander Lerchner. 2016. beta-vae: Learning basic visual concepts with a constrained variational framework. (2016)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969033.2969226"},{"key":"e_1_2_1_20_1","unstructured":"Alex Krizhevsky Geoffrey Hinton etal 2009. Learning multiple layers of features from tiny images. (2009).  Alex Krizhevsky Geoffrey Hinton et al. 2009. Learning multiple layers of features from tiny images. (2009)."},{"key":"e_1_2_1_21_1","volume-title":"Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242","author":"Laine Samuli","year":"2016","unstructured":"Samuli Laine and Timo Aila . 2016. Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242 ( 2016 ). Samuli Laine and Timo Aila. 2016. Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242 (2016)."},{"key":"e_1_2_1_22_1","volume-title":"Deep learning. nature 521, 7553","author":"LeCun Yann","year":"2015","unstructured":"Yann LeCun , Yoshua Bengio , and Geoffrey Hinton . 2015. Deep learning. nature 521, 7553 ( 2015 ), 436--444. Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015. Deep learning. nature 521, 7553 (2015), 436--444."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/2886521.2886706"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.304"},{"key":"e_1_2_1_25_1","article-title":"Nearest Neighbour Searches and the Curse of Dimensionality","volume":"24","author":"R.","year":"1979","unstructured":"R. B. MARIMONT and M. B. SHAPIRO. 1979 . Nearest Neighbour Searches and the Curse of Dimensionality . IMA Journal of Applied Mathematics 24 , 1 (08 1979), 59--70. R. B. MARIMONT and M. B. SHAPIRO. 1979. Nearest Neighbour Searches and the Curse of Dimensionality. IMA Journal of Applied Mathematics 24, 1 (08 1979), 59--70.","journal-title":"IMA Journal of Applied Mathematics"},{"key":"e_1_2_1_26_1","volume-title":"Virtual adversarial training: a regularization method for supervised and semi-supervised learning","author":"Miyato Takeru","year":"2018","unstructured":"Takeru Miyato , Shin-ichi Maeda, Masanori Koyama , and Shin Ishii . 2018. Virtual adversarial training: a regularization method for supervised and semi-supervised learning . IEEE transactions on pattern analysis and machine intelligence 41, 8 ( 2018 ), 1979--1993. Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. 2018. Virtual adversarial training: a regularization method for supervised and semi-supervised learning. IEEE transactions on pattern analysis and machine intelligence 41, 8 (2018), 1979--1993."},{"key":"e_1_2_1_27_1","unstructured":"Yuval Netzer Tao Wang Adam Coates Alessandro Bissacco Bo Wu and Andrew Y Ng. 2011. Reading digits in natural images with unsupervised feature learning. (2011).  Yuval Netzer Tao Wang Adam Coates Alessandro Bissacco Bo Wu and Andrew Y Ng. 2011. Reading digits in natural images with unsupervised feature learning. (2011)."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46466-4_5"},{"key":"e_1_2_1_29_1","volume-title":"Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434","author":"Radford Alec","year":"2015","unstructured":"Alec Radford , Luke Metz , and Soumith Chintala . 2015. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434 ( 2015 ). Alec Radford, Luke Metz, and Soumith Chintala. 2015. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434 (2015)."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.14778\/3157794.3157797"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.5555\/3157096.3157346"},{"key":"e_1_2_1_32_1","volume-title":"Active learning for convolutional neural networks: a core-set approach. arXiv preprint arXiv:1708.00489","author":"Sener Ozan","year":"2017","unstructured":"Ozan Sener and Silvio Savarese . 2017. Active learning for convolutional neural networks: a core-set approach. arXiv preprint arXiv:1708.00489 ( 2017 ). Ozan Sener and Silvio Savarese. 2017. Active learning for convolutional neural networks: a core-set approach. arXiv preprint arXiv:1708.00489 (2017)."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0378-3758(00)00115-4"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00607"},{"key":"e_1_2_1_35_1","volume-title":"Fixmatch: Simplifying semi-supervised learning with consistency and confidence. arXiv preprint arXiv:2001.07685","author":"Sohn Kihyuk","year":"2020","unstructured":"Kihyuk Sohn , David Berthelot , Chun-Liang Li , Zizhao Zhang , Nicholas Carlini , Ekin D Cubuk , Alex Kurakin , Han Zhang , and Colin Raffel . 2020 . Fixmatch: Simplifying semi-supervised learning with consistency and confidence. arXiv preprint arXiv:2001.07685 (2020). Kihyuk Sohn, David Berthelot, Chun-Liang Li, Zizhao Zhang, Nicholas Carlini, Ekin D Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel. 2020. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. arXiv preprint arXiv:2001.07685 (2020)."},{"key":"e_1_2_1_36_1","unstructured":"Casper Kaae S\u00f8nderby Tapani Raiko Lars Maal\u00f8e S\u00f8ren Kaae S\u00f8nderby and Ole Winther.2016. Ladder variational autoencoders. In Advances in neural information processing systems. 3738--3746.  Casper Kaae S\u00f8nderby Tapani Raiko Lars Maal\u00f8e S\u00f8ren Kaae S\u00f8nderby and Ole Winther.2016. Ladder variational autoencoders. In Advances in neural information processing systems. 3738--3746."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1162\/153244302760185243"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.14778\/3291264.3291268"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390294"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2016.2589879"},{"key":"e_1_2_1_41_1","unstructured":"Pete Warden. 2017. Speech commands: A public dataset for single-word speech recognition. (2017).  Pete Warden. 2017. Speech commands: A public dataset for single-word speech recognition. (2017)."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00018"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.5555\/3042573.3042721"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3446776"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46487-9_40"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3476249.3476269","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T09:59:40Z","timestamp":1672221580000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3476249.3476269"}},"subtitle":["labeling complex data at scale"],"short-title":[],"issued":{"date-parts":[[2021,7]]},"references-count":45,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2021,7]]}},"alternative-id":["10.14778\/3476249.3476269"],"URL":"https:\/\/doi.org\/10.14778\/3476249.3476269","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,7]]}}}