{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,23]],"date-time":"2026-01-23T10:57:49Z","timestamp":1769165869130,"version":"3.49.0"},"reference-count":35,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2022,5,23]],"date-time":"2022-05-23T00:00:00Z","timestamp":1653264000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["J. Data and Information Quality"],"published-print":{"date-parts":[[2022,9,30]]},"abstract":"<jats:p>Few-shot learning (FSL) aims at learning to generalize from only a small number of labeled examples for a given target task. Most current state-of-the-art FSL methods typically have two limitations. First, they usually require access to a source dataset (in a similar domain) with abundant labeled examples, which may not always be possible due to privacy concerns and copyright issues. Second, they typically do not offer any estimation of the generalization error on the target FSL task, because the handful of labeled examples must be used for training and cannot spare a validation subset. In this article, we propose a cluster-then-label approach to perform few-shot learning. Our approach does not require access to the labeled source dataset and provides an estimation of generalization error. We show empirically, on four benchmark datasets, that our approach provides competitive predictive performance to state-of-the-art FSL approaches and our generalization error estimation is accurate. Finally, we explore the application of our proposed method to automatic image data labeling. We compare our method with existing automatic data labeling systems. The end-to-end performance of our method outperforms the state-of-the-art automatic data labeling system Snuba by 26% and is only 7% away from the fully supervised upper bound.<\/jats:p>","DOI":"10.1145\/3491232","type":"journal-article","created":{"date-parts":[[2022,3,4]],"date-time":"2022-03-04T22:24:08Z","timestamp":1646432648000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["A Cluster-then-label Approach for Few-shot Learning with Application to Automatic Image Data Labeling"],"prefix":"10.1145","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9144-8999","authenticated-orcid":false,"given":"Renzhi","family":"Wu","sequence":"first","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nilaksh","family":"Das","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sanya","family":"Chaba","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sakshi","family":"Gandhi","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Duen Horng","family":"Chau","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xu","family":"Chu","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, GA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,5,23]]},"reference":[{"key":"e_1_3_2_2_2","article-title":"Data augmentation generative adversarial networks","author":"Antoniou Antreas","year":"2017","unstructured":"Antreas Antoniou, Amos Storkey, and Harrison Edwards. 2017. Data augmentation generative adversarial networks. arXiv:1711.04340. Retrieved from https:\/\/arxiv.org\/abs\/1711.04340.","journal-title":"arXiv:1711.04340"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.5555\/1841234"},{"key":"e_1_3_2_4_2","article-title":"A closer look at few-shot classification","author":"Chen Wei-Yu","year":"2019","unstructured":"Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, and Jia-Bin Huang. 2019. A closer look at few-shot classification. arXiv:1904.04232. Retrieved from https:\/\/arxiv.org\/abs\/1904.04232.","journal-title":"arXiv:1904.04232"},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","unstructured":"Navneet Dalal and Bill Triggs. 2005. Histograms of oriented gradients for human detection. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905) vol. 1. 886\u2013893.","DOI":"10.1109\/CVPR.2005.177"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3380592"},{"key":"e_1_3_2_7_2","first-page":"4171","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 4171\u20134186."},{"key":"e_1_3_2_8_2","first-page":"647","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Donahue Jeff","year":"2014","unstructured":"Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell. 2014. Decaf: A deep convolutional activation feature for generic visual recognition. In Proceedings of the International Conference on Machine Learning. 647\u2013655."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.5555\/3305381.3305498"},{"key":"e_1_3_2_10_2","first-page":"169","volume-title":"Artificial Intelligence and Statistics","author":"Goldberg Andrew","year":"2009","unstructured":"Andrew Goldberg, Xiaojin Zhu, Aarti Singh, Zhiting Xu, and Robert Nowak. 2009. Multi-manifold semi-supervised learning. In Artificial Intelligence and Statistics. 169\u2013176."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMI.2013.2284099"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF02278710"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cell.2018.02.010"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/AINA.2017.166"},{"key":"e_1_3_2_15_2","first-page":"4372","volume-title":"Advances in Neural Information Processing Systems","author":"Lee Joshua","year":"2019","unstructured":"Joshua Lee, Prasanna Sattigeri, and Gregory Wornell. 2019. Learning new tricks from old dogs: Multi-source transfer learning from pre-trained networks. In Advances in Neural Information Processing Systems. 4372\u20134382."},{"key":"e_1_3_2_16_2","article-title":"Lgm-net: Learning to generate matching networks for few-shot learning","author":"Li Huaiyu","year":"2019","unstructured":"Huaiyu Li, Weiming Dong, Xing Mei, Chongyang Ma, Feiyue Huang, and Bao-Gang Hu. 2019. Lgm-net: Learning to generate matching networks for few-shot learning. arXiv:1905.06331. Retrieved from https:\/\/arxiv.org\/abs\/1905.06331.","journal-title":"arXiv:1905.06331"},{"key":"e_1_3_2_17_2","volume-title":"Automated Surface Finish Inspection Using Convolutional Neural Networks","author":"Louhichi Wafa","year":"2019","unstructured":"Wafa Louhichi. 2019. Automated Surface Finish Inspection Using Convolutional Neural Networks. Ph.D. Dissertation. Georgia Institute of Technology."},{"key":"e_1_3_2_18_2","article-title":"On first-order meta-learning algorithms","author":"Nichol Alex","year":"2018","unstructured":"Alex Nichol, Joshua Achiam, and John Schulman. 2018. On first-order meta-learning algorithms. arXiv:1803.02999. Retrieved from https:\/\/arxiv.org\/abs\/1803.02999.","journal-title":"arXiv:1803.02999"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.222"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-018-24876-0"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.14778\/3157794.3157797"},{"key":"e_1_3_2_22_2","first-page":"3567","volume-title":"Advances in Neural Information Processing Systems","author":"Ratner Alexander J.","year":"2016","unstructured":"Alexander J. Ratner, Christopher M. De Sa, Sen Wu, Daniel Selsam, and Christopher R\u00e9. 2016. Data programming: Creating large training sets, quickly. In Advances in Neural Information Processing Systems. 3567\u20133575."},{"key":"e_1_3_2_23_2","article-title":"Meta-learning for semi-supervised few-shot classification","author":"Ren Mengye","year":"2018","unstructured":"Mengye Ren, Eleni Triantafillou, Sachin Ravi, Jake Snell, Kevin Swersky, Joshua B. Tenenbaum, Hugo Larochelle, and Richard S. Zemel. 2018. Meta-learning for semi-supervised few-shot classification. arXiv:1803.00676. Retrieved from https:\/\/arxiv.org\/abs\/1803.00676.","journal-title":"arXiv:1803.00676"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_25_2","first-page":"1842","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Santoro Adam","year":"2016","unstructured":"Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap. 2016. Meta-learning with memory-augmented neural networks. In Proceedings of the International Conference on Machine Learning. 1842\u20131850."},{"key":"e_1_3_2_26_2","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556. Retrieved from https:\/\/arxiv.org\/abs\/1409.1556.","journal-title":"arXiv:1409.1556"},{"key":"e_1_3_2_27_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_28_2","first-page":"4077","volume-title":"Advances in Neural Information Processing Systems","author":"Snell Jake","year":"2017","unstructured":"Jake Snell, Kevin Swersky, and Richard Zemel. 2017. Prototypical networks for few-shot learning. In Advances in Neural Information Processing Systems. 4077\u20134087."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.14778\/3291264.3291268"},{"key":"e_1_3_2_30_2","first-page":"3630","volume-title":"Advances in Neural Information Processing Systems","author":"Vinyals Oriol","year":"2016","unstructured":"Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et\u00a0al. 2016. Matching networks for one shot learning. In Advances in Neural Information Processing Systems. 3630\u20133638."},{"key":"e_1_3_2_31_2","unstructured":"Catherine Wah Steve Branson Peter Welinder Pietro Perona and Serge Belongie. 2011. The caltech-ucsd birds-200-2011 dataset."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00760"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1016\/0169-7439(87)80084-9"},{"key":"e_1_3_2_34_2","article-title":"Tapnet: Neural network augmented with task-adaptive projection for few-shot learning","author":"Yoon Sung Whan","year":"2019","unstructured":"Sung Whan Yoon, Jun Seo, and Jaekyun Moon. 2019. Tapnet: Neural network augmented with task-adaptive projection for few-shot learning. arXiv:1905.06549. Retrieved from https:\/\/arxiv.org\/abs\/1905.06549.","journal-title":"arXiv:1905.06549"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10590-1_53"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.2200\/S00196ED1V01Y200906AIM006"}],"container-title":["Journal of Data and Information Quality"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3491232","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3491232","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:09:19Z","timestamp":1750183759000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3491232"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,5,23]]},"references-count":35,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,9,30]]}},"alternative-id":["10.1145\/3491232"],"URL":"https:\/\/doi.org\/10.1145\/3491232","relation":{},"ISSN":["1936-1955","1936-1963"],"issn-type":[{"value":"1936-1955","type":"print"},{"value":"1936-1963","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,5,23]]},"assertion":[{"value":"2021-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-05-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}