{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,10]],"date-time":"2025-06-10T18:10:04Z","timestamp":1749579004299,"version":"3.41.0"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/100000006","name":"Office of Naval Research","doi-asserted-by":"crossref","award":["N00014-18-1-2670 and N00014-23-1-2132"],"award-info":[{"award-number":["N00014-18-1-2670 and N00014-23-1-2132"]}],"id":[{"id":"10.13039\/100000006","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000001","name":"U.S. National Science Foundation","doi-asserted-by":"crossref","award":["CNS-1822094"],"award-info":[{"award-number":["CNS-1822094"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2025,8,31]]},"abstract":"<jats:p>\n            Adversarial attacks pose significant challenges to deep neural networks (DNNs) such as Transformer models in natural language processing (NLP). This article introduces a novel defense strategy, called\n            <jats:italic>GenFighter<\/jats:italic>\n            , which enhances adversarial robustness by learning and reasoning on the training classification distribution. GenFighter identifies potentially malicious instances deviating from the distribution, transforms them into semantically equivalent instances aligned with the training data, and employs ensemble techniques for a unified and robust response. By conducting extensive experiments, we show that GenFighter outperforms state-of-the-art defenses in\n            <jats:italic>accuracy under attack<\/jats:italic>\n            and\n            <jats:italic>attack success rate<\/jats:italic>\n            metrics while maintaining the same or superior generalization capabilities. Additionally, it requires a high number of queries per attack, making the attack more challenging in real scenarios. Finally, The ablation study shows that our approach proficiently integrates transfer learning, a generative\/evolutive procedure, and an ensemble method, providing an effective defense against NLP adversarial attacks.\n          <\/jats:p>","DOI":"10.1145\/3729240","type":"journal-article","created":{"date-parts":[[2025,4,10]],"date-time":"2025-04-10T14:43:06Z","timestamp":1744296186000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["GenFighter: A Generative and Evolutive Textual Attack Removal"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-9223-6852","authenticated-orcid":false,"given":"Md Athikul","family":"Islam","sequence":"first","affiliation":[{"name":"Department of Computer Science, Boise State University, Boise, Idaho, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0689-5063","authenticated-orcid":false,"given":"Edoardo","family":"Serra","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Boise State University, Boise, Idaho, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3210-558X","authenticated-orcid":false,"given":"Sushil","family":"Jajodia","sequence":"additional","affiliation":[{"name":"Center for Secure Information Systems, George Mason University, Fairfax, Virginia, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,6,10]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1316"},{"key":"e_1_3_2_3_2","volume":"4","author":"Bishop Christopher M.","year":"2006","unstructured":"Christopher M. Bishop and Nasser M. Nasrabadi. 2006. Pattern Recognition and Machine Learning, Vol. 4, Springer.","journal-title":"Pattern Recognition and Machine Learning"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3641289"},{"key":"e_1_3_2_5_2","unstructured":"Patrick Chao Alexander Robey Edgar Dobriban Hamed Hassani George J. Pappas and Eric Wong. 2023. Jailbreaking black box large language models in twenty queries. arXiv:2310.08419. Retrieved from https:\/\/arxiv.org\/abs\/2310.08419"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_2_7_2","unstructured":"Xinshuai Dong Anh Tuan Luu Rongrong Ji and Hong Liu. 2021. Towards robustness against natural language word substitutions. arXiv:2107.13541. Retrieved from https:\/\/arxiv.org\/abs\/2107.13541"},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","unstructured":"Javid Ebrahimi Anyi Rao Daniel Lowd and Dejing Dou. 2017. HotFlip: White-box adversarial examples for text classification. arXiv:1712.06751. Retrieved from https:\/\/arxiv.org\/abs\/1712.06751","DOI":"10.18653\/v1\/P18-2006"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-2006"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/SPW.2018.00016"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.498"},{"key":"e_1_3_2_12_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations (ICLR \u201915)","author":"Goodfellow Ian J.","year":"2015","unstructured":"Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and harnessing adversarial examples. In Proceedings of the 3rd International Conference on Learning Representations (ICLR \u201915). Yoshua Bengio and Yann LeCun (Eds.). Retrieved from http:\/\/arxiv.org\/abs\/1412.6572"},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","first-page":"377","DOI":"10.1145\/1143844.1143892","volume-title":"Proceedings of the 23rd International Conference on Machine Learning (ICML \u201906)","author":"Greene Derek","year":"2006","unstructured":"Derek Greene and P\u00e1draig Cunningham. 2006. Practical solutions to the problem of diagonal dominance in kernel document clustering. In Proceedings of the 23rd International Conference on Machine Learning (ICML \u201906). ACM Press, 377\u2013384."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1419"},{"key":"e_1_3_2_15_2","first-page":"11","volume-title":"Proceedings of the 37th International Conference on Machine Learning (ICML \u201920)","volume":"428","author":"Ishida Takashi","year":"2020","unstructured":"Takashi Ishida, Ikko Yamane, Tomoya Sakai, Gang Niu, and Masashi Sugiyama. 2020. Do we need zero training loss after achieving zero training error? In Proceedings of the 37th International Conference on Machine Learning (ICML \u201920). JMLR.org, Article 428, 11 pages."},{"key":"e_1_3_2_16_2","unstructured":"Neel Jain Avi Schwarzschild Yuxin Wen Gowthami Somepalli John Kirchenbauer Ping Yeh Chiang Micah Goldblum Aniruddha Saha Jonas Geiping and Tom Goldstein. 2023. Baseline defenses for adversarial attacks against aligned language models. arXiv:2309.00614. Retrieved from https:\/\/arxiv.org\/abs\/2309.00614"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.123"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1423"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6311"},{"key":"e_1_3_2_20_2","unstructured":"Jinfeng Li Shouling Ji Tianyu Du Bo Li and Ting Wang. 2018. Textbugger: Generating adversarial text against real-world applications. arXiv:1812.05271. Retrieved from https:\/\/arxiv.org\/abs\/1812.05271"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.500"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.251"},{"key":"e_1_3_2_23_2","unstructured":"Bin Liang Hongcheng Li Miaoqiang Su Pan Bian Xirong Li and Wenchang Shi. 2017. Deep text classification can be fooled. arXiv:1704.08006. Retrieved from https:\/\/arxiv.org\/abs\/1704.08006"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.386"},{"key":"e_1_3_2_25_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2020. Ro{BERT}a: A Robustly Optimized {BERT} Pretraining Approach. Retrieved from https:\/\/openreview.net\/forum?id=SyxS0T4tvS"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.5555\/2002472.2002491"},{"key":"e_1_3_2_27_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Madry Aleksander","year":"2018","unstructured":"Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. Towards deep learning models resistant to adversarial attacks. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=rJzIBfZAb"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.661"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/IBDAP62940.2024.10689688"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.282"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.23919\/SOFTCOM.2019.8903799"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.3115\/1219840.1219855"},{"key":"e_1_3_2_33_2","unstructured":"Colin Raffel Noam Shazeer Adam Roberts Katherine Lee Sharan Narang Michael Matena Yanqi Zhou Wei Li and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv:1910.10683. Retrieved from https:\/\/arxiv.org\/abs\/1910.10683"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1103"},{"key":"e_1_3_2_35_2","unstructured":"Matthew Renze and Erhan Guven. 2024. Self-reflection in LLM agents: Effects on problem-solving performance. arXiv:2405.06682. Retrieved from https:\/\/arxiv.org\/abs\/2405.06682"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-73003-5_196"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1079"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1214\/aos\/1176344136"},{"key":"e_1_3_2_39_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Shi Zhouxing","year":"2020","unstructured":"Zhouxing Shi, Huan Zhang, Kai-Wei Chang, Minlie Huang, and Cho-Jui Hsieh. 2020. Robustness verification for transformers. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=BJxwPJHFwS"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-acl.137"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D13-1170"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.217"},{"key":"e_1_3_2_43_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et al. 2023. LLaMA: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_2_44_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Wang Boxin","year":"2021","unstructured":"Boxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan, Ruoxi Jia, Bo Li, and Jingjing Liu. 2021. Info{BERT}: Improving robustness of language models from an information theoretic perspective. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=hpH98mK5Puk"},{"key":"e_1_3_2_45_2","first-page":"823","volume-title":"Proceedings of the 37th Conference on Uncertainty in Artificial Intelligence (Proceedings of Machine Learning Research, Vol. 161)","author":"Wang Xiaosen","year":"2021","unstructured":"Xiaosen Wang, Jin Hao, Yichen Yang, and Kun He. 2021. Natural language adversarial defense through synonym encoding. In Proceedings of the 37th Conference on Uncertainty in Artificial Intelligence (Proceedings of Machine Learning Research, Vol. 161). Cassio de Campos and Marloes H. Maathuis (Eds.), PMLR, 823\u2013833. Retrieved from https:\/\/proceedings.mlr.press\/v161\/wang21a.html"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.948"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.155"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3116612"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.317"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-demo.43"},{"key":"e_1_3_2_51_2","unstructured":"Jiehang Zeng Xiaoqing Zheng Jianhan Xu Linyang Li Liping Yuan and Xuanjing Huang. 2021. Certified robustness to text adversarial attacks by randomized [MASK]. arXiv:2105.03743. Retrieved from https:\/\/arxiv.org\/abs\/2105.03743"},{"key":"e_1_3_2_52_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"28","author":"Zhang Xiang","year":"2015","unstructured":"Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Proceedings of the Advances in Neural Information Processing Systems. C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett (Eds.), Vol. 28, Curran Associates, Inc. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2015\/file\/250cf8b51c773f3f8dc8b4be867a9a02-Paper.pdf"},{"key":"e_1_3_2_53_2","volume-title":"Proceedings of the NAACL","author":"Zhang Yuan","year":"2019","unstructured":"Yuan Zhang, Jason Baldridge, and Luheng He. 2019. PAWS: Paraphrase adversaries from word scrambling. In Proceedings of the NAACL."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.426"},{"key":"e_1_3_2_55_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Zhu Chen","year":"2020","unstructured":"Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2020. FreeLB: Enhanced adversarial training for natural language understanding. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=BygzbyHFvB"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729240","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,10]],"date-time":"2025-06-10T17:54:55Z","timestamp":1749578095000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3729240"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,10]]},"references-count":54,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,8,31]]}},"alternative-id":["10.1145\/3729240"],"URL":"https:\/\/doi.org\/10.1145\/3729240","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"type":"print","value":"2157-6904"},{"type":"electronic","value":"2157-6912"}],"subject":[],"published":{"date-parts":[[2025,6,10]]},"assertion":[{"value":"2024-06-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-30","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}