{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T21:34:54Z","timestamp":1784237694337,"version":"3.55.0"},"reference-count":154,"publisher":"Association for Computing Machinery (ACM)","issue":"14s","license":[{"start":{"date-parts":[[2023,7,17]],"date-time":"2023-07-17T00:00:00Z","timestamp":1689552000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2023,12,31]]},"abstract":"<jats:p>In the past few years, it has become increasingly evident that deep neural networks are not resilient enough to withstand adversarial perturbations in input data, leaving them vulnerable to attack. Various authors have proposed strong adversarial attacks for computer vision and Natural Language Processing (NLP) tasks. As a response, many defense mechanisms have also been proposed to prevent these networks from failing. The significance of defending neural networks against adversarial attacks lies in ensuring that the model\u2019s predictions remain unchanged even if the input data is perturbed. Several methods for adversarial defense in NLP have been proposed, catering to different NLP tasks such as text classification, named entity recognition, and natural language inference. Some of these methods not only defend neural networks against adversarial attacks but also act as a regularization mechanism during training, saving the model from overfitting. This survey aims to review the various methods proposed for adversarial defenses in NLP over the past few years by introducing a novel taxonomy. The survey also highlights the fragility of advanced deep neural networks in NLP and the challenges involved in defending them.<\/jats:p>","DOI":"10.1145\/3593042","type":"journal-article","created":{"date-parts":[[2023,4,20]],"date-time":"2023-04-20T12:15:21Z","timestamp":1681992921000},"page":"1-39","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":138,"title":["A Survey of Adversarial Defenses and Robustness in NLP"],"prefix":"10.1145","volume":"55","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3629-3701","authenticated-orcid":false,"given":"Shreya","family":"Goyal","sequence":"first","affiliation":[{"name":"Robert Bosch Centre for Data Science and AI, Indian Institute of Technology Madras, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4248-646X","authenticated-orcid":false,"given":"Sumanth","family":"Doddapaneni","sequence":"additional","affiliation":[{"name":"Robert Bosch Centre for Data Science and AI, Indian Institute of Technology Madras, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3687-9922","authenticated-orcid":false,"given":"Mitesh M.","family":"Khapra","sequence":"additional","affiliation":[{"name":"Robert Bosch Centre for Data Science and AI, Indian Institute of Technology Madras, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5364-7639","authenticated-orcid":false,"given":"Balaraman","family":"Ravindran","sequence":"additional","affiliation":[{"name":"Robert Bosch Centre for Data Science and AI, Indian Institute of Technology Madras, India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,7,17]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2807385"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/d18-1316"},{"key":"e_1_3_1_4_2","first-page":"274","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Athalye Anish","year":"2018","unstructured":"Anish Athalye, Nicholas Carlini, and David Wagner. 2018. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In Proceedings of the International Conference on Machine Learning. PMLR, 274\u2013283."},{"key":"e_1_3_1_5_2","article-title":"Imperceptible adversarial attacks on tabular data","author":"Ballet Vincent","year":"2019","unstructured":"Vincent Ballet, Xavier Renard, Jonathan Aigrain, Thibault Laugel, Pascal Frossard, and Marcin Detyniecki. 2019. Imperceptible adversarial attacks on tabular data. arXiv preprint arXiv:1911.03274 (2019).","journal-title":"arXiv preprint arXiv:1911.03274"},{"key":"e_1_3_1_6_2","article-title":"Defending pre-trained language models from adversarial word substitutions without performance sacrifice","author":"Bao Rongzhou","year":"2021","unstructured":"Rongzhou Bao, Jiayi Wang, and Hai Zhao. 2021. Defending pre-trained language models from adversarial word substitutions without performance sacrifice. arXiv preprint arXiv:2105.14553 (2021).","journal-title":"arXiv preprint arXiv:2105.14553"},{"key":"e_1_3_1_7_2","unstructured":"Yonatan Belinkov and Yonatan Bisk. 2018. Synthetic and Natural Noise Both Break Neural Machine Translation. arxiv:1711.02173 [cs.CL]."},{"key":"e_1_3_1_8_2","unstructured":"Petr B\u011blohl\u00e1vek. 2017. Using adversarial examples in natural language processing. In Proceedings of the Association for Computational Linguistics Univerzita Karlova Matematicko-fyzik\u00e1ln\u00ed fakulta. https:\/\/aclanthology.org\/L18-1584.pdf."},{"key":"e_1_3_1_9_2","article-title":"Data-driven mitigation of adversarial text perturbation","author":"Bhalerao Rasika","year":"2022","unstructured":"Rasika Bhalerao, Mohammad Al-Rubaie, Anand Bhaskar, and Igor Markov. 2022. Data-driven mitigation of adversarial text perturbation. arXiv preprint arXiv:2202.09483 (2022).","journal-title":"arXiv preprint arXiv:2202.09483"},{"key":"e_1_3_1_10_2","article-title":"A survey of black-box adversarial attacks on computer vision models","author":"Bhambri Siddhant","year":"2019","unstructured":"Siddhant Bhambri, Sumanyu Muku, Avinash Tulasi, and Arun Balaji Buduru. 2019. A survey of black-box adversarial attacks on computer vision models. arXiv preprint arXiv:1912.01667 (2019).","journal-title":"arXiv preprint arXiv:1912.01667"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3453483.3454056"},{"key":"e_1_3_1_12_2","first-page":"1987","volume-title":"Proceedings of the IEEE Symposium on Security and Privacy (SP\u201922)","author":"Boucher Nicholas","year":"2022","unstructured":"Nicholas Boucher, Ilia Shumailov, Ross Anderson, and Nicolas Papernot. 2022. Bad characters: Imperceptible NLP attacks. In Proceedings of the IEEE Symposium on Security and Privacy (SP\u201922). IEEE, 1987\u20132004."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.3390\/make3040048"},{"key":"e_1_3_1_14_2","article-title":"Adversarial attacks and defences: A survey","author":"Chakraborty Anirban","year":"2018","unstructured":"Anirban Chakraborty, Manaar Alam, Vishal Dey, Anupam Chattopadhyay, and Debdeep Mukhopadhyay. 2018. Adversarial attacks and defences: A survey. arXiv preprint arXiv:1810.00069 (2018).","journal-title":"arXiv preprint arXiv:1810.00069"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.777"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1425"},{"key":"e_1_3_1_17_2","article-title":"AdvAug: Robust adversarial augmentation for neural machine translation","author":"Cheng Yong","year":"2020","unstructured":"Yong Cheng, Lu Jiang, Wolfgang Macherey, and Jacob Eisenstein. 2020. AdvAug: Robust adversarial augmentation for neural machine translation. arXiv preprint arXiv:2006.11834 (2020).","journal-title":"arXiv preprint arXiv:2006.11834"},{"issue":"4","key":"e_1_3_1_18_2","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1080\/00031305.1995.10476177","article-title":"Understanding the metropolis-hastings algorithm","volume":"49","author":"Chib Siddhartha","year":"1995","unstructured":"Siddhartha Chib and Edward Greenberg. 1995. Understanding the metropolis-hastings algorithm. Amer. Statist. 49, 4 (1995), 327\u2013335.","journal-title":"Amer. Statist."},{"key":"e_1_3_1_19_2","article-title":"Privacy-preserving neural representations of text","author":"Coavoux Maximin","year":"2018","unstructured":"Maximin Coavoux, Shashi Narayan, and Shay B. Cohen. 2018. Privacy-preserving neural representations of text. arXiv preprint arXiv:1808.09408 (2018).","journal-title":"arXiv preprint arXiv:1808.09408"},{"key":"e_1_3_1_20_2","article-title":"Build it break it fix it for dialogue safety: Robustness from adversarial human attack","author":"Dinan Emily","year":"2019","unstructured":"Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019. Build it break it fix it for dialogue safety: Robustness from adversarial human attack. arXiv preprint arXiv:1908.06083 (2019).","journal-title":"arXiv preprint arXiv:1908.06083"},{"key":"e_1_3_1_21_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dong Xinshuai","year":"2020","unstructured":"Xinshuai Dong, Anh Tuan Luu, Rongrong Ji, and Hong Liu. 2020. Towards robustness against natural language word substitutions. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401209"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460120.3484538"},{"key":"e_1_3_1_24_2","article-title":"On adversarial examples for character-level neural machine translation","author":"Ebrahimi Javid","year":"2018","unstructured":"Javid Ebrahimi, Daniel Lowd, and Dejing Dou. 2018. On adversarial examples for character-level neural machine translation. arXiv preprint arXiv:1806.09030 (2018).","journal-title":"arXiv preprint arXiv:1806.09030"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-2006"},{"key":"e_1_3_1_26_2","unstructured":"David Eppstein. 1995. Zonohedra and zonotopes. (1995). https:\/\/www.ics.uci.edu\/eppstein\/pubs\/Epp-TR-95-53.pdf."},{"key":"e_1_3_1_27_2","article-title":"Defending against backdoor attacks in natural language generation","author":"Fan Chun","year":"2021","unstructured":"Chun Fan, Xiaoya Li, Yuxian Meng, Xiaofei Sun, Xiang Ao, Fei Wu, Jiwei Li, and Tianwei Zhang. 2021. Defending against backdoor attacks in natural language generation. arXiv preprint arXiv:2106.01810 (2021).","journal-title":"arXiv preprint arXiv:2106.01810"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1905334117"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/SPW.2018.00016"},{"key":"e_1_3_1_30_2","article-title":"BAE: BERT-based adversarial examples for text classification","author":"Garg Siddhant","year":"2020","unstructured":"Siddhant Garg and Goutham Ramakrishnan. 2020. BAE: BERT-based adversarial examples for text classification. arXiv preprint arXiv:2004.01970 (2020).","journal-title":"arXiv preprint arXiv:2004.01970"},{"key":"e_1_3_1_31_2","article-title":"Explaining and harnessing adversarial examples","author":"Goodfellow Ian J.","year":"2014","unstructured":"Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572 (2014).","journal-title":"arXiv preprint arXiv:1412.6572"},{"key":"e_1_3_1_32_2","unstructured":"Ian J. Goodfellow Jonathon Shlens and Christian Szegedy. 2015. Explaining and Harnessing Adversarial Examples. arxiv:1412.6572 [stat.ML]."},{"key":"e_1_3_1_33_2","article-title":"On the effectiveness of interval bound propagation for training verifiably robust models","author":"Gowal Sven","year":"2018","unstructured":"Sven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel, Chongli Qin, Jonathan Uesato, Relja Arandjelovic, Timothy Mann, and Pushmeet Kohli. 2018. On the effectiveness of interval bound propagation for training verifiably robust models. arXiv preprint arXiv:1810.12715 (2018).","journal-title":"arXiv preprint arXiv:1810.12715"},{"key":"e_1_3_1_34_2","article-title":"Towards variable-length textual adversarial attacks","author":"Guo Junliang","year":"2021","unstructured":"Junliang Guo, Zhirui Zhang, Linlin Zhang, Linli Xu, Boxing Chen, Enhong Chen, and Weihua Luo. 2021. Towards variable-length textual adversarial attacks. arXiv preprint arXiv:2104.08139 (2021).","journal-title":"arXiv preprint arXiv:2104.08139"},{"key":"e_1_3_1_35_2","article-title":"Adversarial attack and defense of structured prediction models","author":"Han Wenjuan","year":"2020","unstructured":"Wenjuan Han, Liwen Zhang, Yong Jiang, and Kewei Tu. 2020. Adversarial attack and defense of structured prediction models. arXiv preprint arXiv:2010.01610 (2020).","journal-title":"arXiv preprint arXiv:2010.01610"},{"key":"e_1_3_1_36_2","article-title":"Model extraction and adversarial transferability, your BERT is vulnerable!","author":"He Xuanli","year":"2021","unstructured":"Xuanli He, Lingjuan Lyu, Qiongkai Xu, and Lichao Sun. 2021. Model extraction and adversarial transferability, your BERT is vulnerable! arXiv preprint arXiv:2103.10013 (2021).","journal-title":"arXiv preprint arXiv:2103.10013"},{"key":"e_1_3_1_37_2","unstructured":"Hossein Hosseini Sreeram Kannan Baosen Zhang and Radha Poovendran. 2017. Deceiving Google\u2019s Perspective API Built for Detecting Toxic Comments. arxiv:1702.08138 [cs.LG]."},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"1520","DOI":"10.18653\/v1\/P19-1147","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Hsieh Yu-Lun","year":"2019","unstructured":"Yu-Lun Hsieh, Minhao Cheng, Da-Cheng Juan, Wei Wei, Wen-Lian Hsu, and Cho-Jui Hsieh. 2019. On the robustness of self-attentive models. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 1520\u20131529."},{"key":"e_1_3_1_39_2","article-title":"Achieving verified robustness to symbol substitutions via interval bound propagation","author":"Huang Po-Sen","year":"2019","unstructured":"Po-Sen Huang, Robert Stanforth, Johannes Welbl, Chris Dyer, Dani Yogatama, Sven Gowal, Krishnamurthy Dvijotham, and Pushmeet Kohli. 2019. Achieving verified robustness to symbol substitutions via interval bound propagation. arXiv preprint arXiv:1909.01492 (2019).","journal-title":"arXiv preprint arXiv:1909.01492"},{"key":"e_1_3_1_40_2","article-title":"Adversarial attacks and defense on texts: A survey","author":"Huq Aminul","year":"2020","unstructured":"Aminul Huq, Mst Pervin, et\u00a0al. 2020. Adversarial attacks and defense on texts: A survey. arXiv preprint arXiv:2005.14108 (2020).","journal-title":"arXiv preprint arXiv:2005.14108"},{"key":"e_1_3_1_41_2","article-title":"Fooling explanations in text classifiers","author":"Ivankay Adam","year":"2022","unstructured":"Adam Ivankay, Ivan Girardi, Chiara Marchiori, and Pascal Frossard. 2022. Fooling explanations in text classifiers. arXiv preprint arXiv:2206.03178 (2022).","journal-title":"arXiv preprint arXiv:2206.03178"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n18-1170"},{"key":"e_1_3_1_43_2","article-title":"Adversarial example generation with syntactically controlled paraphrase networks","author":"Iyyer Mohit","year":"2018","unstructured":"Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018. Adversarial example generation with syntactically controlled paraphrase networks. arXiv preprint arXiv:1804.06059 (2018).","journal-title":"arXiv preprint arXiv:1804.06059"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1215"},{"key":"e_1_3_1_45_2","article-title":"Adversarial examples for evaluating reading comprehension systems","author":"Jia Robin","year":"2017","unstructured":"Robin Jia and Percy Liang. 2017. Adversarial examples for evaluating reading comprehension systems. arXiv preprint arXiv:1707.07328 (2017).","journal-title":"arXiv preprint arXiv:1707.07328"},{"key":"e_1_3_1_46_2","article-title":"Certified robustness to adversarial word substitutions","author":"Jia Robin","year":"2019","unstructured":"Robin Jia, Aditi Raghunathan, Kerem G\u00f6ksel, and Percy Liang. 2019. Certified robustness to adversarial word substitutions. arXiv preprint arXiv:1909.00986 (2019).","journal-title":"arXiv preprint arXiv:1909.00986"},{"key":"e_1_3_1_47_2","article-title":"Is BERT really robust? Natural language attack on text classification and entailment","author":"Jin Di","year":"2019","unstructured":"Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2019. Is BERT really robust? Natural language attack on text classification and entailment. arXiv preprint arXiv:1907.11932 (2019).","journal-title":"arXiv preprint arXiv:1907.11932"},{"key":"e_1_3_1_48_2","article-title":"Robust encodings: A framework for combating adversarial typos","author":"Jones Erik","year":"2020","unstructured":"Erik Jones, Robin Jia, Aditi Raghunathan, and Percy Liang. 2020. Robust encodings: A framework for combating adversarial typos. arXiv preprint arXiv:2005.01229 (2020).","journal-title":"arXiv preprint arXiv:2005.01229"},{"key":"e_1_3_1_49_2","article-title":"Adventure: Adversarial training for textual entailment with knowledge-guided examples","author":"Kang Dongyeop","year":"2018","unstructured":"Dongyeop Kang, Tushar Khot, Ashish Sabharwal, and Eduard Hovy. 2018. Adventure: Adversarial training for textual entailment with knowledge-guided examples. arXiv preprint arXiv:1805.04680 (2018).","journal-title":"arXiv preprint arXiv:1805.04680"},{"key":"e_1_3_1_50_2","article-title":"Improving adversarial robustness of ensembles with diversity training","author":"Kariyappa Sanjay","year":"2019","unstructured":"Sanjay Kariyappa and Moinuddin K. Qureshi. 2019. Improving adversarial robustness of ensembles with diversity training. arXiv preprint arXiv:1901.09981 (2019).","journal-title":"arXiv preprint arXiv:1901.09981"},{"key":"e_1_3_1_51_2","article-title":"BERT-defense: A probabilistic model based on BERT to combat cognitively inspired orthographic adversarial attacks","author":"Keller Yannik","year":"2021","unstructured":"Yannik Keller, Jan Mackensen, and Steffen Eger. 2021. BERT-defense: A probabilistic model based on BERT to combat cognitively inspired orthographic adversarial attacks. arXiv preprint arXiv:2106.01452 (2021).","journal-title":"arXiv preprint arXiv:2106.01452"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICNN.1995.488968"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-022-13428-4#citeas"},{"key":"e_1_3_1_54_2","first-page":"3468","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Ko Ching-Yun","year":"2019","unstructured":"Ching-Yun Ko, Zhaoyang Lyu, Lily Weng, Luca Daniel, Ngai Wong, and Dahua Lin. 2019. POPQORN: Quantifying robustness of recurrent neural networks. In Proceedings of the International Conference on Machine Learning. PMLR, 3468\u20133477."},{"key":"e_1_3_1_55_2","article-title":"A survey on adversarial attack in the age of artificial intelligence","volume":"2021","author":"Kong Zixiao","year":"2021","unstructured":"Zixiao Kong, Jingfeng Xue, Yong Wang, Lu Huang, Zequn Niu, and Feng Li. 2021. A survey on adversarial attack in the age of artificial intelligence. Wirel. Commun. Mob. Comput. 2021 (2021), 1\u201322.","journal-title":"Wirel. Commun. Mob. Comput."},{"key":"e_1_3_1_56_2","unstructured":"Volodymyr Kuleshov Shantanu Thakoor Tingfung Lau and S. Ermon. 2018. Adversarial examples for natural language classification problems. openreview.net. https:\/\/openreview.net\/forum?id=r1QZ3zbAZ."},{"key":"e_1_3_1_57_2","first-page":"11047","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"36","author":"Malfa Emanuele La","year":"2022","unstructured":"Emanuele La Malfa and Marta Kwiatkowska. 2022. The king is naked: On the notion of robustness for natural language processing. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 11047\u201311057."},{"key":"e_1_3_1_58_2","article-title":"Assessing robustness of text classification through maximal safe radius computation","author":"Malfa Emanuele La","year":"2020","unstructured":"Emanuele La Malfa, Min Wu, Luca Laurenti, Benjie Wang, Anthony Hartshorn, and Marta Kwiatkowska. 2020. Assessing robustness of text classification through maximal safe radius computation. arXiv preprint arXiv:2010.02004 (2020).","journal-title":"arXiv preprint arXiv:2010.02004"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2019.00044"},{"key":"e_1_3_1_60_2","article-title":"Meta learning for natural language processing: A survey","author":"Lee Hung-yi","year":"2022","unstructured":"Hung-yi Lee, Shang-Wen Li, and Ngoc Thang Vu. 2022. Meta learning for natural language processing: A survey. arXiv preprint arXiv:2205.01500 (2022).","journal-title":"arXiv preprint arXiv:2205.01500"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00440"},{"key":"e_1_3_1_62_2","first-page":"4585","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Levine Alexander","year":"2020","unstructured":"Alexander Levine and Soheil Feizi. 2020. Robustness certificates for sparse adversarial attacks by randomized ablation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 4585\u20134593."},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2022.119170"},{"key":"e_1_3_1_64_2","first-page":"1381","volume-title":"Proceedings of the 29th USENIX Security Symposium (USENIX Security\u201920)","author":"Li Jinfeng","year":"2020","unstructured":"Jinfeng Li, Tianyu Du, Shouling Ji, Rong Zhang, Quan Lu, Min Yang, and Ting Wang. 2020. TextShield: Robust text classification based on multimodal embedding and neural machine translation. In Proceedings of the 29th USENIX Security Symposium (USENIX Security\u201920). 1381\u20131398."},{"key":"e_1_3_1_65_2","first-page":"7708","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201921)","author":"Li Jinfeng","year":"2021","unstructured":"Jinfeng Li, Tianyu Du, Xiangyu Liu, Rong Zhang, Hui Xue, and Shouling Ji. 2021. Enhancing model robustness by incorporating adversarial knowledge into semantic representation. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201921). IEEE, 7708\u20137712."},{"key":"e_1_3_1_66_2","article-title":"TextBugger: Generating adversarial text against real-world applications","author":"Li Jinfeng","year":"2018","unstructured":"Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. 2018. TextBugger: Generating adversarial text against real-world applications. arXiv preprint arXiv:1812.05271 (2018).","journal-title":"arXiv preprint arXiv:1812.05271"},{"key":"e_1_3_1_67_2","article-title":"Adversarial VQA: A new benchmark for evaluating the robustness of VQA models","author":"Li Linjie","year":"2021","unstructured":"Linjie Li, Jie Lei, Zhe Gan, and Jingjing Liu. 2021. Adversarial VQA: A new benchmark for evaluating the robustness of VQA models. arXiv preprint arXiv:2106.00245 (2021).","journal-title":"arXiv preprint arXiv:2106.00245"},{"key":"e_1_3_1_68_2","article-title":"TAVAT: Token-aware virtual adversarial training for language understanding","author":"Li Linyang","year":"2020","unstructured":"Linyang Li and Xipeng Qiu. 2020. TAVAT: Token-aware virtual adversarial training for language understanding. arXiv preprint arXiv:2004.14543 (2020).","journal-title":"arXiv preprint arXiv:2004.14543"},{"key":"e_1_3_1_69_2","first-page":"692","volume-title":"Proceedings of the 4th International Conference on Electronic Information Technology and Computer Engineering","author":"Li Lianjie","year":"2020","unstructured":"Lianjie Li, Zi Zhu, Dongyu Du, Shuxia Ren, Yao Zheng, and Guangsheng Chang. 2020. Adversarial convolutional neural network for text classification. In Proceedings of the 4th International Conference on Electronic Information Technology and Computer Engineering. 692\u2013696."},{"key":"e_1_3_1_70_2","article-title":"DailyDialog: A manually labelled multi-turn dialogue dataset","author":"Li Yanran","year":"2017","unstructured":"Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. DailyDialog: A manually labelled multi-turn dialogue dataset. arXiv preprint arXiv:1710.03957 (2017).","journal-title":"arXiv preprint arXiv:1710.03957"},{"key":"e_1_3_1_71_2","article-title":"Searching for an effective defender: Benchmarking defense against adversarial word substitution","author":"Li Zongyi","year":"2021","unstructured":"Zongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li, Xiaoqing Zheng, Qi Zhang, Kai-Wei Chang, and Cho-Jui Hsieh. 2021. Searching for an effective defender: Benchmarking defense against adversarial word substitution. arXiv preprint arXiv:2108.12777 (2021).","journal-title":"arXiv preprint arXiv:2108.12777"},{"key":"e_1_3_1_72_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2018\/585"},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.5555\/3304222.3304355"},{"key":"e_1_3_1_74_2","first-page":"8384","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Liu Hui","year":"2020","unstructured":"Hui Liu, Yongzheng Zhang, Yipeng Wang, Zheng Lin, and Yige Chen. 2020. Joint character-level word embedding and adversarial stability training to defend adversarial text. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 8384\u20138391."},{"key":"e_1_3_1_75_2","article-title":"Adversarial multi-task learning for text classification","author":"Liu Pengfei","year":"2017","unstructured":"Pengfei Liu, Xipeng Qiu, and Xuanjing Huang. 2017. Adversarial multi-task learning for text classification. arXiv preprint arXiv:1704.05742 (2017).","journal-title":"arXiv preprint arXiv:1704.05742"},{"key":"e_1_3_1_76_2","article-title":"Adversarial training for large neural language models","author":"Liu Xiaodong","year":"2020","unstructured":"Xiaodong Liu, Hao Cheng, Pengcheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao. 2020. Adversarial training for large neural language models. arXiv preprint arXiv:2004.08994 (2020).","journal-title":"arXiv preprint arXiv:2004.08994"},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2020.107332"},{"key":"e_1_3_1_78_2","article-title":"Adversarially regularising neural NLI models to integrate logical background knowledge","author":"Minervini Pasquale","year":"2018","unstructured":"Pasquale Minervini and Sebastian Riedel. 2018. Adversarially regularising neural NLI models to integrate logical background knowledge. arXiv preprint arXiv:1808.08609 (2018).","journal-title":"arXiv preprint arXiv:1808.08609"},{"key":"e_1_3_1_79_2","article-title":"Adversarial training methods for semi-supervised text classification","author":"Miyato Takeru","year":"2016","unstructured":"Takeru Miyato, Andrew M. Dai, and Ian Goodfellow. 2016. Adversarial training methods for semi-supervised text classification. arXiv preprint arXiv:1605.07725 (2016).","journal-title":"arXiv preprint arXiv:1605.07725"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2858821"},{"key":"e_1_3_1_81_2","article-title":"TextAttack: A framework for adversarial attacks, data augmentation, and adversarial training in NLP","author":"Morris John X.","year":"2020","unstructured":"John X. Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020. TextAttack: A framework for adversarial attacks, data augmentation, and adversarial training in NLP. arXiv preprint arXiv:2005.05909 (2020).","journal-title":"arXiv preprint arXiv:2005.05909"},{"key":"e_1_3_1_82_2","article-title":"Counter-fitting word vectors to linguistic constraints","author":"Mrk\u0161i\u0107 Nikola","year":"2016","unstructured":"Nikola Mrk\u0161i\u0107, Diarmuid O. S\u00e9aghdha, Blaise Thomson, Milica Ga\u0161i\u0107, Lina Rojas-Barahona, Pei-Hao Su, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016. Counter-fitting word vectors to linguistic constraints. arXiv preprint arXiv:1603.00892 (2016).","journal-title":"arXiv preprint arXiv:1603.00892"},{"key":"e_1_3_1_83_2","article-title":"Adversarial NLI: A new benchmark for natural language understanding","author":"Nie Yixin","year":"2019","unstructured":"Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2019. Adversarial NLI: A new benchmark for natural language understanding. arXiv preprint arXiv:1910.14599 (2019).","journal-title":"arXiv preprint arXiv:1910.14599"},{"key":"e_1_3_1_84_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.2979670"},{"key":"e_1_3_1_85_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2018.10.315"},{"key":"e_1_3_1_86_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_3_1_87_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/d14-1162"},{"key":"e_1_3_1_88_2","article-title":"Targeted adversarial training for natural language understanding","author":"Pereira Lis","year":"2021","unstructured":"Lis Pereira, Xiaodong Liu, Hao Cheng, Hoifung Poon, Jianfeng Gao, and Ichiro Kobayashi. 2021. Targeted adversarial training for natural language understanding. arXiv preprint arXiv:2104.05847 (2021).","journal-title":"arXiv preprint arXiv:2104.05847"},{"key":"e_1_3_1_89_2","article-title":"Does robustness improve fairness? Approaching fairness with word substitution robustness methods for text classification","author":"Pruksachatkun Yada","year":"2021","unstructured":"Yada Pruksachatkun, Satyapriya Krishna, Jwala Dhamala, Rahul Gupta, and Kai-Wei Chang. 2021. Does robustness improve fairness? Approaching fairness with word substitution robustness methods for text classification. arXiv preprint arXiv:2106.10826 (2021).","journal-title":"arXiv preprint arXiv:2106.10826"},{"key":"e_1_3_1_90_2","article-title":"Combating adversarial misspellings with robust word recognition","author":"Pruthi Danish","year":"2019","unstructured":"Danish Pruthi, Bhuwan Dhingra, and Zachary C. Lipton. 2019. Combating adversarial misspellings with robust word recognition. arXiv preprint arXiv:1905.11268 (2019).","journal-title":"arXiv preprint arXiv:1905.11268"},{"key":"e_1_3_1_91_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.04.020"},{"key":"e_1_3_1_92_2","article-title":"Certified defenses against adversarial examples","author":"Raghunathan Aditi","year":"2018","unstructured":"Aditi Raghunathan, Jacob Steinhardt, and Percy Liang. 2018. Certified defenses against adversarial examples. arXiv preprint arXiv:1801.09344 (2018).","journal-title":"arXiv preprint arXiv:1801.09344"},{"key":"e_1_3_1_93_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eng.2019.12.012"},{"key":"e_1_3_1_94_2","first-page":"1085","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Ren Shuhuai","year":"2019","unstructured":"Shuhuai Ren, Yihe Deng, Kun He, and Wanxiang Che. 2019. Generating natural language adversarial examples through probability weighted word saliency. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 1085\u20131097."},{"key":"e_1_3_1_95_2","article-title":"Generating natural language adversarial examples on a large scale with generative models","author":"Ren Yankun","year":"2020","unstructured":"Yankun Ren, Jianbin Lin, Siliang Tang, Jun Zhou, Shuang Yang, Yuan Qi, and Xiang Ren. 2020. Generating natural language adversarial examples on a large scale with generative models. arXiv preprint arXiv:2003.10388 (2020).","journal-title":"arXiv preprint arXiv:2003.10388"},{"key":"e_1_3_1_96_2","first-page":"1","volume-title":"Proceedings of the International Joint Conference on Neural Networks (IJCNN\u201921)","author":"Rosenberg Ishai","year":"2021","unstructured":"Ishai Rosenberg, Asaf Shabtai, Yuval Elovici, and Lior Rokach. 2021. Sequence squeezing: A defense method against adversarial examples for API call-based RNN variants. In Proceedings of the International Joint Conference on Neural Networks (IJCNN\u201921). IEEE, 1\u201310."},{"key":"e_1_3_1_97_2","doi-asserted-by":"crossref","unstructured":"Cynthia Rudin and Joanna Radin. 2019. Why are we using black box models in AI when we don\u2019t need to? A lesson from an explainable AI competition. Harvard Data Science Review 1 2 (2019) 10\u20131162. https:\/\/hdsr.mitpress.mit.edu\/pub\/f9kuryi8\/release\/1.","DOI":"10.1162\/99608f92.5a8a3a3d"},{"key":"e_1_3_1_98_2","doi-asserted-by":"crossref","first-page":"225","DOI":"10.1007\/978-3-030-81685-8_10","volume-title":"Proceedings of the International Conference on Computer-Aided Verification","author":"Ryou Wonryong","year":"2021","unstructured":"Wonryong Ryou, Jiayu Chen, Mislav Balunovic, Gagandeep Singh, Andrei Dan, and Martin Vechev. 2021. Scalable polyhedral verification of recurrent neural networks. In Proceedings of the International Conference on Computer-Aided Verification. Springer, 225\u2013248."},{"key":"e_1_3_1_99_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00347"},{"key":"e_1_3_1_100_2","volume-title":"Proceedings of the 31st AAAI Conference on Artificial Intelligence","author":"Sakaguchi Keisuke","year":"2017","unstructured":"Keisuke Sakaguchi, Kevin Duh, Matt Post, and Benjamin Van Durme. 2017. Robsut wrod reocginiton via semi-character recurrent neural network. In Proceedings of the 31st AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_1_101_2","unstructured":"Suranjana Samanta and Sameep Mehta. 2017. Towards Crafting Text Adversarial Samples. arxiv:1707.02812 [cs.LG]."},{"key":"e_1_3_1_102_2","article-title":"Interpretable adversarial perturbation in input embedding space for text","author":"Sato Motoki","year":"2018","unstructured":"Motoki Sato, Jun Suzuki, Hiroyuki Shindo, and Yuji Matsumoto. 2018. Interpretable adversarial perturbation in input embedding space for text. arXiv preprint arXiv:1805.02917 (2018).","journal-title":"arXiv preprint arXiv:1805.02917"},{"key":"e_1_3_1_103_2","article-title":"Adversarial training for free!","volume":"32","author":"Shafahi Ali","year":"2019","unstructured":"Ali Shafahi, Mahyar Najibi, Mohammad Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S. Davis, Gavin Taylor, and Tom Goldstein. 2019. Adversarial training for free! Adv. Neural Inf. Process. Syst. 32 (2019).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_104_2","article-title":"Robustness verification for transformers","author":"Shi Zhouxing","year":"2020","unstructured":"Zhouxing Shi, Huan Zhang, Kai-Wei Chang, Minlie Huang, and Cho-Jui Hsieh. 2020. Robustness verification for transformers. arXiv preprint arXiv:2002.06622 (2020).","journal-title":"arXiv preprint arXiv:2002.06622"},{"key":"e_1_3_1_105_2","article-title":"Avoiding the hypothesis-only bias in natural language inference via ensemble adversarial training","author":"Stacey Joe","year":"2020","unstructured":"Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Sebastian Riedel, and Tim Rockt\u00e4schel. 2020. Avoiding the hypothesis-only bias in natural language inference via ensemble adversarial training. arXiv preprint arXiv:2004.07790 (2020).","journal-title":"arXiv preprint arXiv:2004.07790"},{"key":"e_1_3_1_106_2","article-title":"Improving commonsense causal reasoning by adversarial training and data augmentation","author":"Stali\u016bnait\u0117 Ieva","year":"2021","unstructured":"Ieva Stali\u016bnait\u0117, Philip John Gorinski, and Ignacio Iacobacci. 2021. Improving commonsense causal reasoning by adversarial training and data augmentation. arXiv preprint arXiv:2101.04966 (2021).","journal-title":"arXiv preprint arXiv:2101.04966"},{"key":"e_1_3_1_107_2","doi-asserted-by":"publisher","DOI":"10.5555\/3294996.3295110"},{"key":"e_1_3_1_108_2","article-title":"Using random perturbations to mitigate adversarial attacks on sentiment analysis models","author":"Swenor Abigail","year":"2022","unstructured":"Abigail Swenor and Jugal Kalita. 2022. Using random perturbations to mitigate adversarial attacks on sentiment analysis models. arXiv preprint arXiv:2202.05758 (2022).","journal-title":"arXiv preprint arXiv:2202.05758"},{"key":"e_1_3_1_109_2","article-title":"Natural language processing advancements by deep learning: A survey","author":"Torfi Amirsina","year":"2020","unstructured":"Amirsina Torfi, Rouzbeh A. Shirvani, Yaser Keneshloo, Nader Tavaf, and Edward A. Fox. 2020. Natural language processing advancements by deep learning: A survey. arXiv preprint arXiv:2003.01200 (2020).","journal-title":"arXiv preprint arXiv:2003.01200"},{"key":"e_1_3_1_110_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00279"},{"key":"e_1_3_1_111_2","article-title":"Imitation attacks and defenses for black-box machine translation systems","author":"Wallace Eric","year":"2020","unstructured":"Eric Wallace, Mitchell Stern, and Dawn Song. 2020. Imitation attacks and defenses for black-box machine translation systems. arXiv preprint arXiv:2004.15015 (2020).","journal-title":"arXiv preprint arXiv:2004.15015"},{"key":"e_1_3_1_112_2","unstructured":"Matthew Wallace Rishabh Khandelwal and Brian Tang. 2022. Does IBP scale? (2022). https:\/\/www.bjaytang.com\/pdfs\/Does_IBP_Scale_.pdf."},{"key":"e_1_3_1_113_2","article-title":"InfoBERT: Improving robustness of language models from an information theoretic perspective","author":"Wang Boxin","year":"2020","unstructured":"Boxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan, Ruoxi Jia, Bo Li, and Jingjing Liu. 2020. InfoBERT: Improving robustness of language models from an information theoretic perspective. arXiv preprint arXiv:2010.02329 (2020).","journal-title":"arXiv preprint arXiv:2010.02329"},{"key":"e_1_3_1_114_2","first-page":"6555","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wang Dilin","year":"2019","unstructured":"Dilin Wang, Chengyue Gong, and Qiang Liu. 2019. Improving neural language modeling via adversarial training. In Proceedings of the International Conference on Machine Learning. PMLR, 6555\u20136565."},{"key":"e_1_3_1_115_2","article-title":"CAT-Gen: Improving robustness in NLP models via controlled adversarial text generation","author":"Wang Tianlu","year":"2020","unstructured":"Tianlu Wang, Xuezhi Wang, Yao Qin, Ben Packer, Kang Li, Jilin Chen, Alex Beutel, and Ed Chi. 2020. CAT-Gen: Improving robustness in NLP models via controlled adversarial text generation. arXiv preprint arXiv:2010.02338 (2020).","journal-title":"arXiv preprint arXiv:2010.02338"},{"key":"e_1_3_1_116_2","first-page":"1102","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Wang Wenjie","year":"2021","unstructured":"Wenjie Wang, Pengfei Tang, Jian Lou, and Li Xiong. 2021. Certified robustness to word substitution attack with differential privacy. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1102\u20131112."},{"key":"e_1_3_1_117_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3058278"},{"key":"e_1_3_1_118_2","article-title":"Towards a robust deep neural network in texts: A survey","author":"Wang Wenqi","year":"2019","unstructured":"Wenqi Wang, Run Wang, Lina Wang, Zhibo Wang, and Aoshuang Ye. 2019. Towards a robust deep neural network in texts: A survey. arXiv preprint arXiv:1902.07285 (2019).","journal-title":"arXiv preprint arXiv:1902.07285"},{"key":"e_1_3_1_119_2","article-title":"Towards a robust deep neural network against adversarial texts: A survey","author":"Wang Wenqi","year":"2023","unstructured":"Wenqi Wang, Run Wang, Lina Wang, Zhibo Wang, and Aoshuang Ye. 2023. Towards a robust deep neural network against adversarial texts: A survey. IEEE Trans. Knowl. Data Eng. 35, 3 (2023), 3159\u20133179.","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"e_1_3_1_120_2","unstructured":"Xiaosen Wang Hao Jin Yichen Yang and Kun He. 2021. Natural language adversarial defense through synonym encoding. Uncertainty in Artificial Intelligence PMLR 823\u2013833. https:\/\/www.auai.org\/uai2021\/pdf\/uai2021.315.pdf."},{"key":"e_1_3_1_121_2","article-title":"Randomized substitution and vote for textual adversarial example detection","author":"Wang Xiaosen","year":"2021","unstructured":"Xiaosen Wang, Yifeng Xiong, and Kun He. 2021. Randomized substitution and vote for textual adversarial example detection. arXiv preprint arXiv:2109.05698 (2021).","journal-title":"arXiv preprint arXiv:2109.05698"},{"key":"e_1_3_1_122_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-2091"},{"key":"e_1_3_1_123_2","article-title":"Robust machine comprehension models via adversarial training","author":"Wang Yicheng","year":"2018","unstructured":"Yicheng Wang and Mohit Bansal. 2018. Robust machine comprehension models via adversarial training. arXiv preprint arXiv:1804.06473 (2018).","journal-title":"arXiv preprint arXiv:1804.06473"},{"key":"e_1_3_1_124_2","article-title":"Evaluating the robustness of neural networks: An extreme value theory approach","author":"Weng Tsui-Wei","year":"2018","unstructured":"Tsui-Wei Weng, Huan Zhang, Pin-Yu Chen, Jinfeng Yi, Dong Su, Yupeng Gao, Cho-Jui Hsieh, and Luca Daniel. 2018. Evaluating the robustness of neural networks: An extreme value theory approach. arXiv preprint arXiv:1801.10578 (2018).","journal-title":"arXiv preprint arXiv:1801.10578"},{"key":"e_1_3_1_125_2","article-title":"ANLizing the adversarial natural language inference dataset","author":"Williams Adina","year":"2020","unstructured":"Adina Williams, Tristan Thrush, and Douwe Kiela. 2020. ANLizing the adversarial natural language inference dataset. arXiv preprint arXiv:2010.12729 (2020).","journal-title":"arXiv preprint arXiv:2010.12729"},{"key":"e_1_3_1_126_2","article-title":"Performance evaluation of adversarial attacks: Discrepancies and solutions","author":"Wu Jing","year":"2021","unstructured":"Jing Wu, Mingyi Zhou, Ce Zhu, Yipeng Liu, Mehrtash Harandi, and Li Li. 2021. Performance evaluation of adversarial attacks: Discrepancies and solutions. arXiv preprint arXiv:2104.11103 (2021).","journal-title":"arXiv preprint arXiv:2104.11103"},{"key":"e_1_3_1_127_2","first-page":"1778","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Wu Yi","year":"2017","unstructured":"Yi Wu, David Bamman, and Stuart Russell. 2017. Adversarial training for relation extraction. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. 1778\u20131783."},{"key":"e_1_3_1_128_2","article-title":"Identifying adversarial attacks on text classifiers","author":"Xie Zhouhang","year":"2022","unstructured":"Zhouhang Xie, Jonathan Brophy, Adam Noack, Wencong You, Kalyani Asthana, Carter Perkins, Sabrina Reis, Sameer Singh, and Daniel Lowd. 2022. Identifying adversarial attacks on text classifiers. arXiv preprint arXiv:2201.08555 (2022).","journal-title":"arXiv preprint arXiv:2201.08555"},{"key":"e_1_3_1_129_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11633-019-1211-x"},{"key":"e_1_3_1_130_2","first-page":"5518","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP\u201919)","author":"Xu Jingjing","year":"2019","unstructured":"Jingjing Xu, Liang Zhao, Hanqi Yan, Qi Zeng, Yun Liang, and Xu Sun. 2019. LexicalAT: Lexical-based adversarial reinforcement training for robust sentiment classification. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP\u201919). 5518\u20135527."},{"key":"e_1_3_1_131_2","article-title":"Grey-box adversarial attack and defence for sentiment classification","author":"Xu Ying","year":"2021","unstructured":"Ying Xu, Xu Zhong, Antonio Jimeno Yepes, and Jey Han Lau. 2021. Grey-box adversarial attack and defence for sentiment classification. arXiv preprint arXiv:2103.11576 (2021).","journal-title":"arXiv preprint arXiv:2103.11576"},{"key":"e_1_3_1_132_2","article-title":"Elephant in the room: An evaluation framework for assessing adversarial examples in NLP","author":"Xu Ying","year":"2020","unstructured":"Ying Xu, Xu Zhong, Antonio Jose Jimeno Yepes, and Jey Han Lau. 2020. Elephant in the room: An evaluation framework for assessing adversarial examples in NLP. arXiv preprint arXiv:2001.07820 (2020).","journal-title":"arXiv preprint arXiv:2001.07820"},{"key":"e_1_3_1_133_2","article-title":"Robust textual embedding against word-level adversarial attacks","author":"Yang Yichen","year":"2022","unstructured":"Yichen Yang, Xiaosen Wang, and Kun He. 2022. Robust textual embedding against word-level adversarial attacks. arXiv preprint arXiv:2202.13817 (2022).","journal-title":"arXiv preprint arXiv:2202.13817"},{"key":"e_1_3_1_134_2","article-title":"Adversarial training for machine reading comprehension with virtual embeddings","author":"Yang Ziqing","year":"2021","unstructured":"Ziqing Yang, Yiming Cui, Chenglei Si, Wanxiang Che, Ting Liu, Shijin Wang, and Guoping Hu. 2021. Adversarial training for machine reading comprehension with virtual embeddings. arXiv preprint arXiv:2106.04437 (2021).","journal-title":"arXiv preprint arXiv:2106.04437"},{"key":"e_1_3_1_135_2","article-title":"Robust multilingual part-of-speech tagging via adversarial training","author":"Yasunaga Michihiro","year":"2017","unstructured":"Michihiro Yasunaga, Jungo Kasai, and Dragomir Radev. 2017. Robust multilingual part-of-speech tagging via adversarial training. arXiv preprint arXiv:1711.04903 (2017).","journal-title":"arXiv preprint arXiv:1711.04903"},{"key":"e_1_3_1_136_2","article-title":"SAFER: A structure-free approach for certified robustness to adversarial word substitutions","author":"Ye Mao","year":"2020","unstructured":"Mao Ye, Chengyue Gong, and Qiang Liu. 2020. SAFER: A structure-free approach for certified robustness to adversarial word substitutions. arXiv preprint arXiv:2005.14424 (2020).","journal-title":"arXiv preprint arXiv:2005.14424"},{"key":"e_1_3_1_137_2","article-title":"Detection of word adversarial examples in text classification: Benchmark and baseline via robust density estimation","author":"Yoo KiYoon","year":"2022","unstructured":"KiYoon Yoo, Jangho Kim, Jiho Jang, and Nojun Kwak. 2022. Detection of word adversarial examples in text classification: Benchmark and baseline via robust density estimation. arXiv preprint arXiv:2203.01677 (2022).","journal-title":"arXiv preprint arXiv:2203.01677"},{"key":"e_1_3_1_138_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2018.2886017"},{"key":"e_1_3_1_139_2","article-title":"Word-level textual adversarial attacking as combinatorial optimization","author":"Zang Yuan","year":"2019","unstructured":"Yuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu, Meng Zhang, Qun Liu, and Maosong Sun. 2019. Word-level textual adversarial attacking as combinatorial optimization. arXiv preprint arXiv:1910.12196 (2019).","journal-title":"arXiv preprint arXiv:1910.12196"},{"key":"e_1_3_1_140_2","article-title":"OpenAttack: An open-source textual adversarial attack toolkit","author":"Zeng Guoyang","year":"2020","unstructured":"Guoyang Zeng, Fanchao Qi, Qianrui Zhou, Tingji Zhang, Zixian Ma, Bairu Hou, Yuan Zang, Zhiyuan Liu, and Maosong Sun. 2020. OpenAttack: An open-source textual adversarial attack toolkit. arXiv preprint arXiv:2009.09191 (2020).","journal-title":"arXiv preprint arXiv:2009.09191"},{"key":"e_1_3_1_141_2","article-title":"Certified robustness to text adversarial attacks by randomized [MASK]","author":"Zeng Jiehang","year":"2021","unstructured":"Jiehang Zeng, Xiaoqing Zheng, Jianhan Xu, Linyang Li, Liping Yuan, and Xuanjing Huang. 2021. Certified robustness to text adversarial attacks by randomized [MASK]. arXiv preprint arXiv:2105.03743 (2021).","journal-title":"arXiv preprint arXiv:2105.03743"},{"key":"e_1_3_1_142_2","article-title":"MACER: Attack-free and scalable robust training via maximizing certified radius","author":"Zhai Runtian","year":"2020","unstructured":"Runtian Zhai, Chen Dan, Di He, Huan Zhang, Boqing Gong, Pradeep Ravikumar, Cho-Jui Hsieh, and Liwei Wang. 2020. MACER: Attack-free and scalable robust training via maximizing certified radius. arXiv preprint arXiv:2001.02378 (2020).","journal-title":"arXiv preprint arXiv:2001.02378"},{"key":"e_1_3_1_143_2","article-title":"A survey on universal adversarial attack","author":"Zhang Chaoning","year":"2021","unstructured":"Chaoning Zhang, Philipp Benz, Chenguo Lin, Adil Karjauv, Jing Wu, and In So Kweon. 2021. A survey on universal adversarial attack. arXiv preprint arXiv:2103.01498 (2021).","journal-title":"arXiv preprint arXiv:2103.01498"},{"key":"e_1_3_1_144_2","article-title":"Generating fluent adversarial examples for natural languages","author":"Zhang Huangzhao","year":"2020","unstructured":"Huangzhao Zhang, Hao Zhou, Ning Miao, and Lei Li. 2020. Generating fluent adversarial examples for natural languages. arXiv preprint arXiv:2007.06174 (2020).","journal-title":"arXiv preprint arXiv:2007.06174"},{"key":"e_1_3_1_145_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2981616"},{"issue":"3","key":"e_1_3_1_146_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3374217","article-title":"Adversarial attacks on deep-learning models in natural language processing: A survey","volume":"11","author":"Zhang Wei Emma","year":"2020","unstructured":"Wei Emma Zhang, Quan Z. Sheng, Ahoud Alhazmi, and Chenliang Li. 2020. Adversarial attacks on deep-learning models in natural language processing: A survey. ACM Trans. Intell. Syst. Technol. 11, 3 (2020), 1\u201341.","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"e_1_3_1_147_2","article-title":"Certified robustness to programmable transformations in LSTMs","author":"Zhang Yuhao","year":"2021","unstructured":"Yuhao Zhang, Aws Albarghouthi, and Loris D\u2019Antoni. 2021. Certified robustness to programmable transformations in LSTMs. arXiv preprint arXiv:2102.07818 (2021).","journal-title":"arXiv preprint arXiv:2102.07818"},{"key":"e_1_3_1_148_2","article-title":"PAWS: Paraphrase adversaries from word scrambling","author":"Zhang Yuan","year":"2019","unstructured":"Yuan Zhang, Jason Baldridge, and Luheng He. 2019. PAWS: Paraphrase adversaries from word scrambling. arXiv preprint arXiv:1904.01130 (2019).","journal-title":"arXiv preprint arXiv:1904.01130"},{"key":"e_1_3_1_149_2","unstructured":"Zhengli Zhao Dheeru Dua and Sameer Singh. 2018. Generating Natural Adversarial Examples. arXiv:1710.11342 [cs.LG]."},{"key":"e_1_3_1_150_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.485"},{"key":"e_1_3_1_151_2","article-title":"Learning to discriminate perturbations for blocking adversarial attacks in text classification","author":"Zhou Yichao","year":"2019","unstructured":"Yichao Zhou, Jyun-Yu Jiang, Kai-Wei Chang, and Wei Wang. 2019. Learning to discriminate perturbations for blocking adversarial attacks in text classification. arXiv preprint arXiv:1909.03084 (2019).","journal-title":"arXiv preprint arXiv:1909.03084"},{"key":"e_1_3_1_152_2","article-title":"Defense against adversarial attacks in NLP via Dirichlet neighborhood ensemble","author":"Zhou Yi","year":"2020","unstructured":"Yi Zhou, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-wei Chang, and Xuanjing Huang. 2020. Defense against adversarial attacks in NLP via Dirichlet neighborhood ensemble. arXiv preprint arXiv:2006.11627 (2020).","journal-title":"arXiv preprint arXiv:2006.11627"},{"key":"e_1_3_1_153_2","article-title":"TREATED: Towards universal defense against textual adversarial attacks","author":"Zhu Bin","year":"2021","unstructured":"Bin Zhu, Zhaoquan Gu, Le Wang, and Zhihong Tian. 2021. TREATED: Towards universal defense against textual adversarial attacks. arXiv preprint arXiv:2109.06176 (2021).","journal-title":"arXiv preprint arXiv:2109.06176"},{"key":"e_1_3_1_154_2","article-title":"FreeLB: Enhanced adversarial training for natural language understanding","author":"Zhu Chen","year":"2019","unstructured":"Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2019. FreeLB: Enhanced adversarial training for natural language understanding. arXiv preprint arXiv:1909.11764 (2019).","journal-title":"arXiv preprint arXiv:1909.11764"},{"key":"e_1_3_1_155_2","article-title":"Adversarial training as Stackelberg game: An unrolled optimization approach","author":"Zuo Simiao","year":"2021","unstructured":"Simiao Zuo, Chen Liang, Haoming Jiang, Xiaodong Liu, Pengcheng He, Jianfeng Gao, Weizhu Chen, and Tuo Zhao. 2021. Adversarial training as Stackelberg game: An unrolled optimization approach. arXiv preprint arXiv:2104.04886 (2021).","journal-title":"arXiv preprint arXiv:2104.04886"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3593042","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3593042","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:19Z","timestamp":1750178239000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3593042"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,17]]},"references-count":154,"journal-issue":{"issue":"14s","published-print":{"date-parts":[[2023,12,31]]}},"alternative-id":["10.1145\/3593042"],"URL":"https:\/\/doi.org\/10.1145\/3593042","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,17]]},"assertion":[{"value":"2022-04-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-04-02","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}