{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T09:22:09Z","timestamp":1781774529756,"version":"3.54.5"},"reference-count":78,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2022,8,11]],"date-time":"2022-08-11T00:00:00Z","timestamp":1660176000000},"content-version":"vor","delay-in-days":222,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,8,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>This paper studies the use of language models as a source of synthetic unlabeled text for NLP. We formulate a general framework called \u201cgenerate, annotate, and learn (GAL)\u201d to take advantage of synthetic text within knowledge distillation, self-training, and few-shot learning applications. To generate high-quality task-specific text, we either fine-tune LMs on inputs from the task of interest, or prompt large LMs with few examples. We use the best available classifier to annotate synthetic text with soft pseudo labels for knowledge distillation and self-training, and use LMs to obtain hard labels for few-shot learning. We train new supervised models on the combination of labeled and pseudo-labeled data, which results in significant gains across several applications. We investigate key components of GAL and present theoretical and empirical arguments against the use of class-conditional LMs to generate synthetic labeled text instead of unlabeled text. GAL achieves new state-of-the-art knowledge distillation results for 6-layer transformers on the GLUE leaderboard.<\/jats:p>","DOI":"10.1162\/tacl_a_00492","type":"journal-article","created":{"date-parts":[[2022,8,11]],"date-time":"2022-08-11T19:43:54Z","timestamp":1660247034000},"page":"826-842","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":30,"title":["Generate, Annotate, and Learn: NLP with Synthetic Text"],"prefix":"10.1162","volume":"10","author":[{"given":"Xuanli","family":"He","sequence":"first","affiliation":[{"name":"Monash University, Australia. xuanli.he1@monash.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Islam","family":"Nassar","sequence":"additional","affiliation":[{"name":"Monash University, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jamie","family":"Kiros","sequence":"additional","affiliation":[{"name":"Google Research, Brain Team, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gholamreza","family":"Haffari","sequence":"additional","affiliation":[{"name":"Monash University, Australia. gholamreza.haffari@monash.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohammad","family":"Norouzi","sequence":"additional","affiliation":[{"name":"Google Research, Brain Team, Canada. mnorouzi@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2022,8,12]]},"reference":[{"issue":"3","key":"2022081119434531100_bib1","doi-asserted-by":"publisher","first-page":"365","DOI":"10.1162\/0891201041850876","article-title":"Understanding the Yarowsky algorithm","volume":"30","author":"Abney","year":"2004","journal-title":"Computational Linguistics"},{"issue":"4","key":"2022081119434531100_bib2","doi-asserted-by":"publisher","first-page":"373","DOI":"10.1109\/TIT.1970.1054472","article-title":"Learning with a probabilistic teacher","volume":"16","author":"Agrawala","year":"1970","journal-title":"IEEE Transactions on Information Theory"},{"key":"2022081119434531100_bib3","doi-asserted-by":"publisher","first-page":"6168","DOI":"10.18653\/v1\/P19-1620","article-title":"Synthetic QA corpora generation with roundtrip consistency","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Alberti","year":"2019"},{"key":"2022081119434531100_bib4","doi-asserted-by":"publisher","first-page":"7432","DOI":"10.1609\/aaai.v34i05.6239","article-title":"PIQA: Reasoning about physical commonsense in natural language","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Bisk","year":"2020"},{"key":"2022081119434531100_bib5","article-title":"Language models are few-shot learners","author":"Brown","year":"2020","journal-title":"arXiv:2005.14165"},{"key":"2022081119434531100_bib6","doi-asserted-by":"publisher","first-page":"535","DOI":"10.1145\/1150402.1150464","article-title":"Model compression","author":"Bucilu\u0103","year":"2006","journal-title":"Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining"},{"key":"2022081119434531100_bib7","article-title":"Unlabeled data improves adversarial robustness","volume":"32","author":"Carmon","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022081119434531100_bib8","article-title":"Vicinal risk minimization","author":"Chapelle","year":"2001","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022081119434531100_bib9","article-title":"Big self-supervised models are strong semi-supervised learners","author":"Chen","year":"2020","journal-title":"NeurIPS"},{"key":"2022081119434531100_bib10","article-title":"Self-training avoids using spurious features under domain shift","volume-title":"Advances in Neural Information Processing Systems 33: Annual Conferenc\u201312, 2020, virtual","author":"Chen","year":"2020"},{"key":"2022081119434531100_bib11","article-title":"Electra: Pre-training text encoders as discriminators rather than generators","author":"Clark","year":"2020","journal-title":"International Conference on Learning Representations"},{"key":"2022081119434531100_bib12","article-title":"Self- training improves pre-training for natural language understanding","author":"Jingfei","year":"2020","journal-title":"arXiv:2010.02194"},{"key":"2022081119434531100_bib13","doi-asserted-by":"publisher","first-page":"395","DOI":"10.3115\/1220575.1220625","article-title":"Bootstrapping without the boot","volume-title":"Proceedings of Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing","author":"Eisner","year":"2005"},{"key":"2022081119434531100_bib14","doi-asserted-by":"crossref","first-page":"968","DOI":"10.18653\/v1\/2021.findings-acl.84","article-title":"A survey of data augmentation approaches for NLP","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Feng","year":"2021"},{"key":"2022081119434531100_bib15","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.1967.1053952","article-title":"Learning to recognize patterns without a teacher","author":"Fralick","year":"1967","journal-title":"IEEE Transactions on Information Theory"},{"key":"2022081119434531100_bib16","first-page":"1607","article-title":"Born again neural networks","author":"Furlanello","year":"2018","journal-title":"International Conference on Machine Learning"},{"key":"2022081119434531100_bib17","article-title":"A framework for few-shot language model evaluation","author":"Gao","year":"2021"},{"key":"2022081119434531100_bib18","article-title":"Improving robustness using generated data","volume":"34","author":"Gowal","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022081119434531100_bib19","doi-asserted-by":"publisher","first-page":"93","DOI":"10.1162\/tacl_a_00302","article-title":"A knowledge- enhanced pretraining model for commonsense story generation","volume":"8","author":"Guan","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2022081119434531100_bib20","first-page":"159","article-title":"Analysis of semi-supervised learning with the yarowsky algorithm","volume-title":"UAI 2007, Proceedings of the Twenty-Third Conference on Uncertainty in Artificial Intelligence, Vancouver, BC, Canada, July 19-22, 2007","author":"Haffari","year":"2007"},{"key":"2022081119434531100_bib21","article-title":"Deberta: Decoding- enhanced BERT with disentangled attention","author":"He","year":"2020","journal-title":"arXiv:2006.03654"},{"key":"2022081119434531100_bib22","article-title":"Scaling laws for transfer","author":"Hernandez","year":"2021"},{"key":"2022081119434531100_bib23","article-title":"Distilling the knowledge in a neural network","author":"Hinton","year":"2015","journal-title":"arXiv:1503.02531"},{"key":"2022081119434531100_bib24","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.372","article-title":"TinyBERT: Distilling BERT for natural language understanding","author":"Jiao","year":"2019","journal-title":"arXiv:1909 .10351"},{"key":"2022081119434531100_bib25","doi-asserted-by":"publisher","first-page":"1317","DOI":"10.18653\/v1\/D16-1139","article-title":"Sequence-level knowledge distillation","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing","author":"Kim","year":"2016"},{"key":"2022081119434531100_bib26","doi-asserted-by":"publisher","first-page":"452","DOI":"10.18653\/v1\/N18-2072","article-title":"Contextual augmentation: Data augmentation by words with paradigmatic relations","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers)","author":"Kobayashi","year":"2018"},{"key":"2022081119434531100_bib27","first-page":"5468","article-title":"Understanding self-training for gradual domain adaptation","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Kumar","year":"2020"},{"key":"2022081119434531100_bib28","article-title":"Data augmentation using pre- trained transformer models","author":"Kumar","year":"2020","journal-title":"arXiv:2003.02245"},{"key":"2022081119434531100_bib29","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1333","article-title":"Corpora generation for grammatical error correction","author":"Lichtarge","year":"2019","journal-title":"arXiv:1904.05780"},{"key":"2022081119434531100_bib30","article-title":"Roberta: A robustly optimized bert pretraining approach","author":"Liu","year":"2019","journal-title":"arXiv:1907.11692"},{"key":"2022081119434531100_bib31","doi-asserted-by":"publisher","first-page":"1275","DOI":"10.18653\/v1\/P18-1118","article-title":"Document context neural machine translation with memory networks","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Maruf","year":"2018"},{"key":"2022081119434531100_bib32","article-title":"Unnatural language processing: Bridging the gap between synthetic and natural language data","volume":"abs\/2004.13645","author":"Marzoev","year":"2020","journal-title":"ArXiv"},{"key":"2022081119434531100_bib33","doi-asserted-by":"publisher","first-page":"2947","DOI":"10.18653\/v1\/D18-1325","article-title":"Document- level neural machine translation with hierarchical attention networks","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Miculicich","year":"2018"},{"key":"2022081119434531100_bib34","article-title":"Self-distillation amplifies regularization in hilbert space","volume-title":"Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual","author":"Mobahi","year":"2020"},{"key":"2022081119434531100_bib35","doi-asserted-by":"publisher","first-page":"280","DOI":"10.18653\/v1\/K16-1028","article-title":"Abstractive text summarization using sequence-to-sequence RNNs and beyond","volume-title":"Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning","author":"Nallapati","year":"2016"},{"key":"2022081119434531100_bib36","doi-asserted-by":"crossref","first-page":"314","DOI":"10.18653\/v1\/W19-5333","article-title":"Facebook fair\u2019s wmt19 news translation task submission","volume-title":"Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1)","author":"Ng","year":"2019"},{"key":"2022081119434531100_bib37","article-title":"Exemplar vaes for exemplar based generation and data augmentation","author":"Norouzi","year":"2020","journal-title":"arXiv:2004.04795"},{"key":"2022081119434531100_bib38","doi-asserted-by":"publisher","first-page":"2329","DOI":"10.18653\/v1\/2020.coling-main.211","article-title":"Facts2Story: Controlling text generation by key facts","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Orbach","year":"2020"},{"key":"2022081119434531100_bib39","doi-asserted-by":"publisher","first-page":"48","DOI":"10.18653\/v1\/N19-4009","article-title":"fairseq: A fast, extensible toolkit for sequence modeling","author":"Ott","year":"2019","journal-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations)"},{"key":"2022081119434531100_bib40","article-title":"Statistical and algorithmic insights for semi- supervised learning with self-training","author":"Oymak","year":"2020","journal-title":"CoRR"},{"key":"2022081119434531100_bib41","article-title":"Language models are unsupervised multitask learners","author":"Radford","year":"2019"},{"key":"2022081119434531100_bib42","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"Journal of Machine Learning Research"},{"key":"2022081119434531100_bib43","first-page":"7909","article-title":"Understanding and mitigating the tradeoff between robustness and accuracy","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Raghunathan","year":"2020"},{"key":"2022081119434531100_bib44","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.86","article-title":"Mate-kd: Masked adversarial text, a companion to knowledge distillation","author":"Rashid","year":"2021","journal-title":"arXiv preprint arXiv:2105.05912"},{"key":"2022081119434531100_bib45","first-page":"12268","article-title":"Classification accuracy score for conditional generative models","author":"Ravuri","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022081119434531100_bib46","doi-asserted-by":"crossref","first-page":"379","DOI":"10.18653\/v1\/D15-1044","article-title":"A neural attention model for abstractive sentence summarization","volume-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing","author":"Rush","year":"2015"},{"key":"2022081119434531100_bib47","article-title":"DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter","author":"Sanh","year":"2019","journal-title":"ArXiv"},{"key":"2022081119434531100_bib48","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.1965.1053799","article-title":"Probability of error of some adaptive pattern-recognition machines","author":"Scudder","year":"1965","journal-title":"IEEE Transactions on Information Theory"},{"key":"2022081119434531100_bib49","doi-asserted-by":"publisher","first-page":"86","DOI":"10.18653\/v1\/P16-1009","article-title":"Improving neural machine translation models with monolingual data","volume-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Sennrich","year":"2016"},{"key":"2022081119434531100_bib50","article-title":"A simple but tough-to-beat data augmentation approach for natural language understanding and generation","author":"Shen","year":"2020","journal-title":"arXiv preprint arXiv:2009.13818"},{"key":"2022081119434531100_bib51","article-title":"Low resource text classification with ulmfit and backtranslation","author":"Shleifer","year":"2019","journal-title":"arXiv preprint arXiv:1903.09244"},{"key":"2022081119434531100_bib52","article-title":"Fixmatch: Simplifying semi-supervised learning with consistency and confidence","author":"Sohn","year":"2020","journal-title":"arXiv:2001.07685"},{"key":"2022081119434531100_bib53","doi-asserted-by":"publisher","first-page":"4314","DOI":"10.18653\/v1\/D19-1441","article-title":"Patient knowledge distillation for BERT model compression","author":"Sun","year":"2019","journal-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)"},{"key":"2022081119434531100_bib54","article-title":"Ernie: Enhanced representation through knowledge integration","author":"Sun","year":"2019","journal-title":"arXiv preprint arXiv:1904.09223"},{"key":"2022081119434531100_bib55","first-page":"25","article-title":"Transductive learning for statistical machine translation","volume-title":"Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics","author":"Ueffing","year":"2007"},{"key":"2022081119434531100_bib56","article-title":"Principles of risk minimization for learning theory","author":"Vapnik","year":"1992","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022081119434531100_bib57","doi-asserted-by":"publisher","first-page":"1264","DOI":"10.18653\/v1\/P18-1117","article-title":"Context-aware neural machine translation learns anaphora resolution","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Voita","year":"2018"},{"key":"2022081119434531100_bib58","first-page":"3335","article-title":"Generalised unsupervised domain adaptation of neural machine translation with cross-lingual data selection","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Thuy","year":"2021"},{"key":"2022081119434531100_bib59","doi-asserted-by":"publisher","first-page":"5715","DOI":"10.18653\/v1\/2021.emnlp-main.462","article-title":"Strata: Self- training with task augmentation for better few- shot learning","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Tu","year":"2021"},{"key":"2022081119434531100_bib60","doi-asserted-by":"publisher","first-page":"4465","DOI":"10.18653\/v1\/P19-1439","article-title":"Can you tell me how to get past sesame street? Sentence-level pretraining beyond language modeling","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Wang","year":"2019"},{"key":"2022081119434531100_bib61","article-title":"SuperGLUE: A stickier benchmark for general- purpose language understanding systems","author":"Wang","year":"2019","journal-title":"arXiv: 1905.00537"},{"key":"2022081119434531100_bib62","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-5446","article-title":"GLUE: A multi-task benchmark and analysis platform for natural language understanding","author":"Wang","year":"2019","journal-title":"International Conference on Learning Representations"},{"key":"2022081119434531100_bib63","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.220","article-title":"Learning to synthesize data for semantic parsing","volume-title":"Proceedings of the Meeting of the North-American Chapter of Association for Computational Linguistics (NAACL)","author":"Wang","year":"2021"},{"key":"2022081119434531100_bib64","article-title":"GPT-J-6B: A 6 billion parameter autoregressive language model","author":"Wang","year":"2021"},{"key":"2022081119434531100_bib65","doi-asserted-by":"publisher","first-page":"2557","DOI":"10.18653\/v1\/D15-1306","article-title":"That\u2019s so annoying!!!: A lexical and frame- semantic embedding based data augmentation approach to automatic categorization of annoying behaviors using# petpeeve tweets","author":"Wang","year":"2015","journal-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing"},{"key":"2022081119434531100_bib66","article-title":"Kdgan: Knowledge distillation with generative adversarial networks.","author":"Wang","year":"2018","journal-title":"NeurIPS"},{"key":"2022081119434531100_bib67","doi-asserted-by":"publisher","first-page":"1332","DOI":"10.3115\/v1\/P15-1129","article-title":"Building a semantic parser overnight","volume-title":"Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Wang","year":"2015"},{"key":"2022081119434531100_bib68","article-title":"Theoretical analysis of self-training with deep networks on unlabeled data","volume-title":"International Conference on Learning Representations","author":"Wei","year":"2021"},{"key":"2022081119434531100_bib69","doi-asserted-by":"publisher","first-page":"6382","DOI":"10.18653\/v1\/D19-1670","article-title":"Eda: Easy data augmentation techniques for boosting performance on text classification tasks","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Wei","year":"2019"},{"key":"2022081119434531100_bib70","doi-asserted-by":"publisher","first-page":"38","DOI":"10.18653\/v1\/2020.emnlp-demos.6","article-title":"Transformers: State-of-the-art natural language processing","author":"Wolf","year":"2020","journal-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations"},{"key":"2022081119434531100_bib71","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1007\/978-3-030-22747-0_7","article-title":"Conditional BERT contextual augmentation","volume-title":"International Conference on Computational Science","author":"Xing","year":"2019"},{"key":"2022081119434531100_bib72","doi-asserted-by":"publisher","first-page":"10684","DOI":"10.1109\/CVPR42600.2020.01070","article-title":"Self-training with noisy student improves imagenet classification","author":"Xie","year":"2020","journal-title":"2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"2022081119434531100_bib73","first-page":"7859","article-title":"Bert-of-theseus: Compressing BERT by progressive module replacing","author":"Canwen","year":"2020","journal-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)"},{"key":"2022081119434531100_bib74","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.90","article-title":"G-daug: Generative data augmentation for commonsense reasoning","author":"Yang","year":"2020","journal-title":"arXiv:2004.11546"},{"key":"2022081119434531100_bib75","doi-asserted-by":"publisher","first-page":"189","DOI":"10.3115\/981658.981684","article-title":"Unsupervised word sense disambiguation rivaling supervised methods","author":"Yarowsky","year":"1995","journal-title":"33rd Annual Meeting of the Association for Computational Linguistics"},{"key":"2022081119434531100_bib76","article-title":"QANet: Combining local convolution with global self- attention for reading comprehension","author":"Adams Wei","year":"2018","journal-title":"ICLR"},{"key":"2022081119434531100_bib77","article-title":"mixup: Beyond empirical risk minimization","author":"Zhang","year":"2018","journal-title":"ICLR"},{"key":"2022081119434531100_bib78","doi-asserted-by":"publisher","first-page":"3713","DOI":"10.1109\/ICCV.2019.00381","article-title":"Be your own teacher: Improve the performance of convolutional neural networks via self distillation","author":"Zhang","year":"2019","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00492\/2038511\/tacl_a_00492.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00492\/2038511\/tacl_a_00492.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,8,11]],"date-time":"2022-08-11T19:44:36Z","timestamp":1660247076000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00492\/112605\/Generate-Annotate-and-Learn-NLP-with-Synthetic"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022]]},"references-count":78,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00492","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022]]},"published":{"date-parts":[[2022]]}}}