{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T19:13:47Z","timestamp":1784142827660,"version":"3.55.0"},"reference-count":83,"publisher":"Association for Computing Machinery (ACM)","issue":"13s","license":[{"start":{"date-parts":[[2023,7,13]],"date-time":"2023-07-13T00:00:00Z","timestamp":1689206400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100003593","name":"CNPq","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100003593","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100002322","name":"CAPES","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100002322","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100004901","name":"FAPEMIG","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004901","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100008536","name":"Amazon Web Services","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100008536","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100007065","name":"NVIDIA","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100007065","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Google Research Awards"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2023,12,31]]},"abstract":"<jats:p>\n            Progress in natural language processing has been dictated by the\n            <jats:italic>rule of more<\/jats:italic>\n            : more data, more computing power, more complexity, best exemplified by deep learning Transformers. However, training (or fine-tuning) large dense models for specific applications usually requires significant amounts of computing resources. One way to ameliorate this problem is through data engineering instead of the algorithmic or hardware perspectives. Our focus here is an under-investigated data engineering technique, with enormous potential in the current scenario \u2013\n            <jats:italic>Instance Selection<\/jats:italic>\n            (IS) (a.k.a. Selective Sampling, Prototype Selection). The IS goal is to reduce the training set size by removing noisy or redundant instances while maintaining or improving the effectiveness (accuracy) of the trained models and reducing the training process cost. We survey classical and recent state-of-the-art IS techniques and provide a scientifically sound comparison of IS methods applied to an essential natural language processing task\u2014Automatic Text Classification (ATC). IS methods have been normally applied to small tabular datasets and have not been systematically compared in ATC. We consider several neural and non-neural state-of-the-art ATC solutions and many datasets. We answer several research questions based on tradeoffs induced by a tripod (training set reduction, effectiveness, and efficiency). Our answers reveal an enormous unfulfilled potential for IS solutions. Specially, we show that in 12 out of 19 datasets, specific IS methods\u2014namely, Condensed Nearest Neighbor, Local Set-based Smoother, and Local Set Border Selector\u2014can reduce the size of the training set without effectiveness losses. Furthermore, in the case of fine-tuning the Transformer methods, the IS methods reduce the amount of data needed, without losing effectiveness and with considerable training-time gains.\n          <\/jats:p>","DOI":"10.1145\/3582000","type":"journal-article","created":{"date-parts":[[2023,1,24]],"date-time":"2023-01-24T12:04:08Z","timestamp":1674561848000},"page":"1-52","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":55,"title":["A Comparative Survey of Instance Selection Methods applied to Non-Neural and Transformer-Based Text Classification"],"prefix":"10.1145","volume":"55","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1988-8412","authenticated-orcid":false,"given":"Washington","family":"Cunha","sequence":"first","affiliation":[{"name":"Federal University of Minas Gerais"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8121-8607","authenticated-orcid":false,"given":"Felipe","family":"Viegas","sequence":"additional","affiliation":[{"name":"Federal University of Minas Gerais"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0251-7172","authenticated-orcid":false,"given":"Celso","family":"Fran\u00e7a","sequence":"additional","affiliation":[{"name":"Federal University of Minas Gerais"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7117-3994","authenticated-orcid":false,"given":"Thierson","family":"Rosa","sequence":"additional","affiliation":[{"name":"Federal University of Goi\u00e1s"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4913-4902","authenticated-orcid":false,"given":"Leonardo","family":"Rocha","sequence":"additional","affiliation":[{"name":"Federal University of S\u00e3o Jo\u00e3o Del-Rei"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2075-3363","authenticated-orcid":false,"given":"Marcos Andr\u00e9","family":"Gon\u00e7alves","sequence":"additional","affiliation":[{"name":"Federal University of Minas Gerais"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,7,13]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3548772"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF00153759"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.5555\/553876"},{"issue":"2","key":"e_1_3_3_5_2","article-title":"Impact of instance selection on kNN-based text categorization.","volume":"14","author":"Barigou Fatiha","year":"2018","unstructured":"Fatiha Barigou. 2018. Impact of instance selection on kNN-based text categorization. Journal of Information Processing Systems 14, 2 (2018).","journal-title":"Journal of Information Processing Systems"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/RCIS.2009.5089308"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1014043630878"},{"key":"e_1_3_3_8_2","first-page":"1877","volume-title":"Advances in Neural Information Processing Systems","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 1877\u20131901. https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf."},{"key":"e_1_3_3_9_2","article-title":"On the role of text preprocessing in neural network architectures: An evaluation study on text categorization and sentiment analysis","author":"Camacho-Collados Jose","year":"2017","unstructured":"Jose Camacho-Collados and Mohammad Taher Pilehvar. 2017. On the role of text preprocessing in neural network architectures: An evaluation study on text categorization and sentiment analysis. arXiv preprint arXiv:1707.01780 (2017).","journal-title":"arXiv preprint arXiv:1707.01780"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331239"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2018.2820051"},{"key":"e_1_3_3_12_2","doi-asserted-by":"crossref","first-page":"228","DOI":"10.1007\/978-3-319-64283-3_17","volume-title":"Big Data Analytics and Knowledge Discovery","author":"Carbonera Joel Lu\u00eds","year":"2017","unstructured":"Joel Lu\u00eds Carbonera. 2017. An efficient approach for instance selection. In Big Data Analytics and Knowledge Discovery, Ladjel Bellatreche and Sharma Chakravarthy (Eds.). Springer International Publishing, Cham, 228\u2013243."},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICTAI.2015.114"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICTAI.2016.0090"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICTAI.2018.00053"},{"issue":"3","key":"e_1_3_3_16_2","first-page":"27:1\u201327:27","article-title":"LIBSVM: A library for support vector machines","volume":"2","author":"Chang Chih-Chung","year":"2011","unstructured":"Chih-Chung Chang and Chih-Jen Lin. 2011. LIBSVM: A library for support vector machines. ACM Transactions on Intelligent Systems and Technology (TIST) 2, 3 (2011), 27:1\u201327:27.","journal-title":"ACM Transactions on Intelligent Systems and Technology (TIST)"},{"key":"e_1_3_3_17_2","first-page":"22243","article-title":"Big self-supervised models are strong semi-supervised learners","volume":"33","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E. Hinton. 2020. Big self-supervised models are strong semi-supervised learners. Advances in Neural Information Processing Systems 33 (2020), 22243\u201322255.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_18_2","article-title":"Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems","author":"Chen Tianqi","year":"2015","unstructured":"Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. 2015. Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems. arXiv preprint arXiv:1512.01274 (2015).","journal-title":"arXiv preprint arXiv:1512.01274"},{"key":"e_1_3_3_19_2","volume-title":"International Conference on Learning Representations","author":"Chou Yu-Ying","year":"2021","unstructured":"Yu-Ying Chou, Hsuan-Tien Lin, and Tyng-Luh Liu. 2021. Adaptive and generative zero-shot learning. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=ahAUv8TI2Mz."},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2020.102263"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2020.102481"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-37629-1_2"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/N15-1146"},{"key":"e_1_3_3_24_2","first-page":"4171","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (HLT-NAACL)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (HLT-NAACL). 4171\u20134186."},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.3389\/frai.2020.00004"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.5555\/1390681.1442794"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2011.142"},{"key":"e_1_3_3_28_2","article-title":"Towards robustness to label noise in text classification via noise modeling","volume":"2101","author":"Garg Siddhant","year":"2021","unstructured":"Siddhant Garg, Goutham Ramakrishnan, and Varun Thumbe. 2021. Towards robustness to label noise in text classification via noise modeling. CoRR abs\/2101.11214 (2021). arXiv:2101.11214https:\/\/arxiv.org\/abs\/2101.11214.","journal-title":"CoRR"},{"key":"e_1_3_3_29_2","first-page":"155","volume-title":"Proceedings of the Eleventh International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research)","volume":"2","author":"Goldberg Andrew B.","year":"2007","unstructured":"Andrew B. Goldberg, Xiaojin Zhu, and Stephen Wright. 2007. Dissimilarity in graph-based semi-supervised classification. In Proceedings of the Eleventh International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research), Marina Meila and Xiaotong Shen (Eds.), Vol. 2. PMLR, San Juan, Puerto Rico, 155\u2013162. https:\/\/proceedings.mlr.press\/v2\/goldberg07a.html."},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-acl.350"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.1968.1054155"},{"key":"e_1_3_3_32_2","first-page":"192","volume-title":"SIGIR\u201994","author":"Hersh William","year":"1994","unstructured":"William Hersh, Chris Buckley, T. J. Leone, and David Hickam. 1994. OHSUMED: An interactive retrieval evaluation and new large test collection for research. In SIGIR\u201994, Bruce W. Croft and C. J. van Rijsbergen (Eds.). Springer London, London, 192\u2013201."},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/75.4.800"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/160688.160758"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-57077-4_10"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/E17-2068"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICMLA.2017.0-134"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICMLA.2017.0-134"},{"key":"e_1_3_3_39_2","article-title":"Unsupervised machine translation using monolingual corpora only","author":"Lample Guillaume","year":"2017","unstructured":"Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc\u2019Aurelio Ranzato. 2017. Unsupervised machine translation using monolingual corpora only. arXiv preprint arXiv:1711.00043 (2017).","journal-title":"arXiv preprint arXiv:1711.00043"},{"key":"e_1_3_3_40_2","article-title":"Albert: A lite bert for self-supervised learning of language representations","author":"Lan Zhenzhong","year":"2019","unstructured":"Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942 (2019).","journal-title":"arXiv preprint arXiv:1909.11942"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_3_3_42_2","volume-title":"NLDB\u201915","author":"Lev Guy","year":"2015","unstructured":"Guy Lev, Benjamin Klein, and Lior Wolf. 2015. NLDB\u201915. Chapter In Defense of Word Embedding for Generic Text Representation."},{"key":"e_1_3_3_43_2","article-title":"Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension","author":"Lewis Mike","year":"2019","unstructured":"Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461 (2019).","journal-title":"arXiv preprint arXiv:1910.13461"},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2014.10.001"},{"issue":"2","key":"e_1_3_3_45_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3495162","article-title":"A survey on text classification: From traditional to deep learning","volume":"13","author":"Li Qian","year":"2022","unstructured":"Qian Li, Hao Peng, Jianxin Li, Congying Xia, Renyu Yang, Lichao Sun, Philip S. Yu, and Lifang He. 2022. A survey on text classification: From traditional to deep learning. ACM Transactions on Intelligent Systems and Technology (TIST) 13, 2 (2022), 1\u201341.","journal-title":"ACM Transactions on Intelligent Systems and Technology (TIST)"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.3115\/1072228.1072378"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1017\/9781108639286.005"},{"key":"e_1_3_3_48_2","article-title":"Roberta: A robustly optimized bert pretraining approach","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019).","journal-title":"arXiv preprint arXiv:1907.11692"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3178876.3186168"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0221152"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2020.113297"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.143"},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.670"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3412180"},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3439726"},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2022.07.025"},{"key":"e_1_3_3_57_2","article-title":"Nuts and bolts of building AI applications using deep learning","author":"Ng Andrew","year":"2016","unstructured":"Andrew Ng. 2016. Nuts and bolts of building AI applications using deep learning. NIPS Keynote Talk (2016).","journal-title":"NIPS Keynote Talk"},{"key":"e_1_3_3_58_2","doi-asserted-by":"publisher","DOI":"10.3115\/1118693.1118704"},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"issue":"8","key":"e_1_3_3_60_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et\u00a0al. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_3_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292522.3326027"},{"key":"e_1_3_3_63_2","doi-asserted-by":"publisher","DOI":"10.1140\/epjds\/s13688-016-0085-1"},{"key":"e_1_3_3_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/2366316.2366320"},{"key":"e_1_3_3_65_2","doi-asserted-by":"publisher","DOI":"10.3390\/fi11090202"},{"key":"e_1_3_3_66_2","article-title":"DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter","author":"Sanh Victor","year":"2019","unstructured":"Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108 (2019).","journal-title":"arXiv preprint arXiv:1910.01108"},{"key":"e_1_3_3_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/505282.505283"},{"key":"e_1_3_3_68_2","first-page":"1631","volume-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing","author":"Socher Richard","year":"2013","unstructured":"Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Seattle, Washington, USA, 1631\u20131642. https:\/\/aclanthology.org\/D13-1170."},{"key":"e_1_3_3_69_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2009.03.002"},{"key":"e_1_3_3_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/2783258.2783307"},{"key":"e_1_3_3_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/1401890.1402008"},{"key":"e_1_3_3_72_2","doi-asserted-by":"publisher","DOI":"10.5555\/2747013.2747139"},{"key":"e_1_3_3_73_2","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331259"},{"key":"e_1_3_3_74_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289600.3291032"},{"key":"e_1_3_3_75_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1248"},{"key":"e_1_3_3_76_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-42911-3_49"},{"key":"e_1_3_3_77_2","doi-asserted-by":"publisher","DOI":"10.1145\/3293318"},{"key":"e_1_3_3_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.1972.4309137"},{"key":"e_1_3_3_79_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007626913721"},{"key":"e_1_3_3_80_2","first-page":"5754","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems (NIPS)","volume":"32","author":"Yang Zhilin","year":"2019","unstructured":"Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R. Salakhutdinov, and Quoc V. Le. 2019. XLNet: Generalized autoregressive pretraining for language understanding. In Proceedings of the 33rd International Conference on Neural Information Processing Systems (NIPS), Vol. 32. 5754\u20135764."},{"key":"e_1_3_3_81_2","article-title":"Dive into deep learning","author":"Zhang Aston","year":"2021","unstructured":"Aston Zhang, Zachary C. Lipton, Mu Li, and Alexander J. Smola. 2021. Dive into deep learning. arXiv preprint arXiv:2106.11342 (2021).","journal-title":"arXiv preprint arXiv:2106.11342"},{"key":"e_1_3_3_82_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2021.06.008"},{"key":"e_1_3_3_83_2","first-page":"649","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Zhang Xiang","year":"2016","unstructured":"Xiang Zhang, Junbo Zhao, and Yann LeCun. 2016. Character-level convolutional networks for text classification. In Advances in Neural Information Processing Systems (NIPS). Vol. 28. Curran Associates, Inc., 649\u2013657. http:\/\/papers.nips.cc\/paper\/5782-character-level-convolutional-networks-for-text-classification.pdf."},{"key":"e_1_3_3_84_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.75"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3582000","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3582000","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:44Z","timestamp":1750178804000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3582000"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,13]]},"references-count":83,"journal-issue":{"issue":"13s","published-print":{"date-parts":[[2023,12,31]]}},"alternative-id":["10.1145\/3582000"],"URL":"https:\/\/doi.org\/10.1145\/3582000","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,13]]},"assertion":[{"value":"2022-02-10","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-01-18","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}