{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T02:41:20Z","timestamp":1774924880866,"version":"3.50.1"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,12,28]],"date-time":"2024-12-28T00:00:00Z","timestamp":1735344000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Postdoctoral Fellowship Program of CPSF","award":["GZC20232811"],"award-info":[{"award-number":["GZC20232811"]}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"crossref","award":["2024M753357 and 2024M750154"],"award-info":[{"award-number":["2024M753357 and 2024M750154"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key Research Program of Frontier Sciences, CAS","award":["ZDBSLYJSC038"],"award-info":[{"award-number":["ZDBSLYJSC038"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2025,1,31]]},"abstract":"<jats:p>Code based adversarial attacks play a crucial role in revealing vulnerabilities of software system. Recently, pre-trained programming language models (PLMs) have demonstrated remarkable success in various significant software engineering tasks, progressively transforming the paradigm of software development. Despite their impressive capabilities, these powerful models are vulnerable to adversarial attacks. Therefore, it is necessary to carefully investigate the robustness and vulnerabilities of the PLMs by means of adversarial attacks. Adversarial attacks entail imperceptible input modifications that cause target models to make incorrect predictions. Existing approaches for attacking PLMs often employ either identifier renaming or the greedy algorithm, which may yield sub-optimal performance or lead to high inference times. In response to these limitations, we propose CARL, an unsupervised black-box attack model that leverages reinforcement learning to generate imperceptible adversarial examples. Specifically, CARL comprises a programming language encoder and a perturbation prediction layer. In order to achieve more effective and efficient attack, we cast the task as a sequence decision-making process, optimizing through policy gradient with a suite of reward functions. We conduct extensive experiments to validate the effectiveness of CARL on code summarization, code translation, and code refinement tasks, covering various programming languages and PLMs. The experimental results demonstrate that CARL surpasses state-of-the-art code attack models, achieving the highest attack success rate across multiple tasks and PLMs while maintaining high attack efficiency, imperceptibility, consistency, and fluency.<\/jats:p>","DOI":"10.1145\/3688839","type":"journal-article","created":{"date-parts":[[2024,8,14]],"date-time":"2024-08-14T15:16:35Z","timestamp":1723648595000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["CARL: Unsupervised Code-Based Adversarial Attacks for Programming Language Models via Reinforcement Learning"],"prefix":"10.1145","volume":"34","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2093-1473","authenticated-orcid":false,"given":"Kaichun","family":"Yao","sequence":"first","affiliation":[{"name":"Institute of Software, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9274-2698","authenticated-orcid":false,"given":"Hao","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5354-8630","authenticated-orcid":false,"given":"Chuan","family":"Qin","sequence":"additional","affiliation":[{"name":"PBC School of Finance, Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4570-643X","authenticated-orcid":false,"given":"Hengshu","family":"Zhu","sequence":"additional","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1823-0459","authenticated-orcid":false,"given":"Yanjun","family":"Wu","sequence":"additional","affiliation":[{"name":"Institute of Software, Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7153-6465","authenticated-orcid":false,"given":"Libo","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Software, Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,12,28]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"2655","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT \u201921).","author":"Ahmad Wasi Uddin","year":"2021","unstructured":"Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Unified pre-training for program understanding and generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT \u201921). Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-T\u00fcr, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou (Eds.), Association for Computational Linguistics, 2655\u20132668."},{"key":"e_1_3_2_3_2","unstructured":"Loubna Ben Allal Raymond Li Denis Kocetkov Chenghao Mou Christopher Akiki Carlos Munoz Ferrandis Niklas Muennighoff Mayank Mishra Alex Gu Manan Dey Logesh Kumar Umapathi Carolyn Jane Anderson Yangtian Zi Joel Lamy Poirier Hailey Schoelkopf Sergey Troshin Dmitry Abulkhanov Manuel Romero Michael Lappert Francesco De Toni Bernardo Garc\u00eda del R\u00edo Qian Liu Shamik Bose Urvashi Bhattacharyya Terry Yue Zhuo Ian Yu Paulo Villegas Marco Zocca Sourab Mangrulkar David Lansky Huu Nguyen Danish Contractor Luis Villa Jia Li Dzmitry Bahdanau Yacine Jernite Sean Hughes Daniel Fried Arjun Guha Harm de Vries and Leandro von Werra. 2023. SantaCoder: Don\u2019t reach for the stars! arXiv:2301.03988. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2301.03988"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1316"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE51524.2021.9678706"},{"key":"e_1_3_2_6_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mohammad Bavarian Clemens Winter Philippe Tillet Felipe Petroski Such Dave Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William Hebgen Guss Alex Nichol Alex Paino Nikolas Tezak Jie Tang Igor Babuschkin Suchir Balaji Shantanu Jain William Saunders Christopher Hesse Andrew N. Carr Jan Leike Josh Achiam Vedant Misra Evan Morikawa Alec Radford Matthew Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating large language models trained on code. arXiv:2107.03374. Retrieved from https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_2_7_2","series-title":"Long and Short Papers","first-page":"4171","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT \u201919)","volume":"1","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT \u201919), Vol. 1 (Long and Short Papers). Jill Burstein, Christy Doran, and Thamar Solorio (Eds.), Association for Computational Linguistics, 4171\u20134186."},{"key":"e_1_3_2_8_2","series-title":"Advances in Neural Information Processing Systems","first-page":"13042","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems (NeurIPS \u201919).","author":"Dong Li","year":"2019","unstructured":"Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. In Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems (NeurIPS \u201919). Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d\u2019Alch\u00e9-Buc, Emily B. Fox, and Roman Garnett (Eds.), Advances in Neural Information Processing Systems, 13042\u201313054."},{"key":"e_1_3_2_9_2","unstructured":"Zhengxiao Du Yujie Qian Xiao Liu Ming Ding Jiezhong Qiu Zhilin Yang and Jie Tang. 2021. Glm: General language model pretraining with autoregressive blank infilling. arXiv:2103.10360. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2103.10360"},{"key":"e_1_3_2_10_2","series-title":"Short Papers","first-page":"31","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL \u201918)","volume":"2","author":"Ebrahimi Javid","year":"2018","unstructured":"Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018. HotFlip: White-box adversarial examples for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL \u201918), Vol. 2 Short Papers. Iryna Gurevych and Yusuke Miyao (Eds.), Association for Computational Linguistics, 31\u201336."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3580305.3599894"},{"key":"e_1_3_2_12_2","doi-asserted-by":"crossref","first-page":"1536","DOI":"10.18653\/v1\/2020.findings-emnlp.139","volume-title":"Findings of the Association for Computational Linguistics (EMNLP \u201920).","author":"Feng Zhangyin","year":"2020","unstructured":"Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. CodeBERT: A pre-trained model for programming and natural languages. In Findings of the Association for Computational Linguistics (EMNLP \u201920). Trevor Cohn, Yulan He, and Yang Liu (Eds.), Association for Computational Linguistics, 1536\u20131547."},{"key":"e_1_3_2_13_2","unstructured":"Daniel Fried Armen Aghajanyan Jessy Lin Sida Wang Eric Wallace Freda Shi Ruiqi Zhong Wen-tau Yih Luke Zettlemoyer and Mike Lewis. 2022. Incoder: A generative model for code infilling and synthesis. Retrieved from https:\/\/openreview.net\/forum?id=hQwb-lbM6EL"},{"key":"e_1_3_2_14_2","first-page":"50","volume-title":"Proceedings of the IEEE Security and Privacy Workshops (SP Workshops \u201918)","author":"Gao Ji","year":"2018","unstructured":"Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi. 2018. Black-box generation of adversarial text sequences to evade deep learning classifiers. In Proceedings of the IEEE Security and Privacy Workshops (SP Workshops \u201918). IEEE Computer Society, 50\u201356."},{"key":"e_1_3_2_15_2","doi-asserted-by":"crossref","first-page":"6174","DOI":"10.18653\/v1\/2020.emnlp-main.498","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201920).","author":"Garg Siddhant","year":"2020","unstructured":"Siddhant Garg and Goutham Ramakrishnan. 2020. BAE: BERT-based adversarial examples for text classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201920). Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.), Association for Computational Linguistics, 6174\u20136181."},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.499"},{"key":"e_1_3_2_17_2","volume-title":"Proceedings of the 9th International Conference on Learning Representations (ICLR \u201921)","author":"Guo Daya","year":"2021","unstructured":"Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou. 2021. GraphCodeBERT: Pre-training code representations with data flow. In Proceedings of the 9th International Conference on Learning Representations (ICLR \u201921). OpenReview.net."},{"key":"e_1_3_2_18_2","first-page":"526","volume-title":"Proceedings of the IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER \u201922)","author":"Henke Jordan","year":"2022","unstructured":"Jordan Henke, Goutham Ramakrishnan, Zi Wang, Aws Albarghouth, Somesh Jha, and Thomas Reps. 2022. Semantic robustness of models of source code. In Proceedings of the IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER \u201922). IEEE, 526\u2013537."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1147"},{"key":"e_1_3_2_20_2","unstructured":"Hamel Husain Ho-Hsiang Wu Tiferet Gazit Miltiadis Allamanis and Marc Brockschmidt. 2019. CodeSearchNet challenge: Evaluating the state of semantic code search. arXiv:1909.09436. Retrieved from http:\/\/arxiv.org\/abs\/1909.09436"},{"key":"e_1_3_2_21_2","first-page":"14892","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"37","author":"Jha Akshita","year":"2023","unstructured":"Akshita Jha and Chandan K. Reddy. 2023. Codeattack: Code-based adversarial attacks for pre-trained programming language models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, 14892\u201314900."},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"Feihu Jiang Chuan Qin Kaichun Yao Chuyu Fang Fuzhen Zhuang Hengshu Zhu and Hui Xiong. 2024. Enhancing question answering for enterprise knowledge bases using large language models. arXiv:2404.08695. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2404.08695","DOI":"10.1007\/978-981-97-5562-2_18"},{"key":"e_1_3_2_23_2","first-page":"54","volume-title":"Proceedings of the 37th Conference on Uncertainty in Artificial Intelligence (UAI \u201921)","volume":"161","author":"Jiang Xue","year":"2021","unstructured":"Xue Jiang, Zhuoran Zheng, Chen Lyu, Liang Li, and Lei Lyu. 2021. TreeBERT: A tree-based pre-trained model for programming language. In Proceedings of the 37th Conference on Uncertainty in Artificial Intelligence (UAI \u201921), Vol. 161. Cassio P. de Campos, Marloes H. Maathuis, and Erik Quaeghebeur (Eds.), AUAI Press, 54\u201363."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6311"},{"key":"e_1_3_2_25_2","first-page":"8018","volume-title":"Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI \u201920), the 32nd Innovative Applications of Artificial Intelligence Conference (IAAI \u201920), the 10th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI \u201920)","author":"Jin Di","year":"2020","unstructured":"Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020. Is BERT really robust? A strong baseline for natural language attack on text classification and entailment. In Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI \u201920), the 32nd Innovative Applications of Artificial Intelligence Conference (IAAI \u201920), the 10th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI \u201920). AAAI Press, 8018\u20138025."},{"key":"e_1_3_2_26_2","unstructured":"Aditya Kanade Petros Maniatis Gogul Balakrishnan and Kensen Shi. 2020. Pre-trained contextual embedding of source code. arXiv:2001.00059. Retrieved from http:\/\/arxiv.org\/abs\/2001.00059"},{"key":"e_1_3_2_27_2","doi-asserted-by":"crossref","first-page":"7871","DOI":"10.18653\/v1\/2020.acl-main.703","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL \u201920)","author":"Lewis Mike","year":"2020","unstructured":"Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL \u201920). Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.), Association for Computational Linguistics, 7871\u20137880."},{"key":"e_1_3_2_28_2","volume-title":"Proceedings of the 26th Annual Network and Distributed System Security Symposium (NDSS \u201919)","author":"Li Jinfeng","year":"2019","unstructured":"Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. 2019. TextBugger: Generating adversarial text against real-world applications. In Proceedings of the 26th Annual Network and Distributed System Security Symposium (NDSS \u201919). The Internet Society."},{"key":"e_1_3_2_29_2","first-page":"6193","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP \u201920)","author":"Li Linyang","year":"2020","unstructured":"Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020. BERT-ATTACK: Adversarial attack against BERT using BERT. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP \u201920). Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.), Association for Computational Linguistics, 6193\u20136202."},{"key":"e_1_3_2_30_2","unstructured":"Raymond Li Loubna Ben Allal Yangtian Zi Niklas Muennighoff Denis Kocetkov Chenghao Mou Marc Marone Christopher Akiki Jia Li Jenny Chim Qian Liu Evgenii Zheltonozhskii Terry Yue Zhuo Thomas Wang Olivier Dehaene Mishig Davaadorj Joel Lamy-Poirier Jo\u00e3o Monteiro Oleh Shliazhko Nicolas Gontier Nicholas Meade Armel Zebaze Ming-Ho Yee Logesh Kumar Umapathi Jian Zhu Benjamin Lipkin Muhtasham Oblokulov Zhiruo Wang Rudra Murthy Jason Stillerman Siva Sankalp Patel Dmitry Abulkhanov Marco Zocca Manan Dey Zhihan Zhang Nour Fahmy Urvashi Bhattacharyya Wenhao Yu Swayam Singh Sasha Luccioni Paulo Villegas Maxim Kunakov Fedor Zhdanov Manuel Romero Tony Lee Nadav Timor Jennifer Ding Claire Schlesinger Hailey Schoelkopf Jan Ebert Tri Dao Mayank Mishra Alex Gu Jennifer Robinson Carolyn Jane Anderson Brendan Dolan-Gavitt Danish Contractor Siva Reddy Daniel Fried Dzmitry Bahdanau Yacine Jernite Carlos Mu\u00f1oz Ferrandis Sean Hughes Thomas Wolf Arjun Guha Leandro von Werra and Harm de Vries. 2023. Starcoder: May the source be with you! Retrieved from https:\/\/openreview.net\/forum?id=KoFOg41haE"},{"key":"e_1_3_2_31_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv:1907.11692. Retrieved from http:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_2_32_2","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021","author":"Lu Shuai","year":"2021","unstructured":"Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021. CodeXGLUE: A machine learning benchmark dataset for code understanding and generation. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021. Joaquin Vanschoren and Sai-Kit Yeung (Eds.), MIT Press."},{"key":"e_1_3_2_33_2","unstructured":"Ziyang Luo Can Xu Pu Zhao Qingfeng Sun Xiubo Geng Wenxiang Hu Chongyang Tao Jing Ma Qingwei Lin and Daxin Jiang. 2023. Wizardcoder: Empowering code large language models with evol-instruct. Retrieved from https:\/\/openreview.net\/forum?id=UnUwSIgK5W"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1063\/1.1699114"},{"key":"e_1_3_2_35_2","unstructured":"Erik Nijkamp Bo Pang Hiroaki Hayashi Lifu Tu Huan Wang Yingbo Zhou Silvio Savarese and Caiming Xiong. 2022. Codegen: An open large language model for code with multi-turn program synthesis. Retrieved from https:\/\/openreview.net\/forum?id=iaYcJKpY2B_"},{"key":"e_1_3_2_36_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311\u2013318."},{"key":"e_1_3_2_37_2","series-title":"Long Papers","first-page":"5582","volume-title":"Proceedings of the 57th Conference of the Association for Computational Linguistics (ACL \u201919)","volume":"1","author":"Pruthi Danish","year":"2019","unstructured":"Danish Pruthi, Bhuwan Dhingra, and Zachary C. Lipton. 2019. Combating adversarial misspellings with robust word recognition. In Proceedings of the 57th Conference of the Association for Computational Linguistics (ACL \u201919), Vol. 1 Long Papers. Anna Korhonen, David R. Traum, and Llu\u00eds M\u00e0rquez (Eds.), Association for Computational Linguistics, 5582\u20135591."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2022.3145396"},{"issue":"8","key":"e_1_3_2_39_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_2_40_2","first-page":"140:1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21 (2020), 140:1\u2013140:67.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_41_2","series-title":"Long Papers","first-page":"1085","volume-title":"Proceedings of the 57th Conference of the Association for Computational Linguistics (ACL \u201919)","volume":"1","author":"Ren Shuhuai","year":"2019","unstructured":"Shuhuai Ren, Yihe Deng, Kun He, and Wanxiang Che. 2019. Generating natural language adversarial examples through probability weighted word saliency. In Proceedings of the 57th Conference of the Association for Computational Linguistics (ACL \u201919), Vol. 1 Long Papers. Anna Korhonen, David R. Traum, and Llu\u00eds M\u00e0rquez (Eds.), Association for Computational Linguistics, 1085\u20131097."},{"key":"e_1_3_2_42_2","unstructured":"Shuo Ren Daya Guo Shuai Lu Long Zhou Shujie Liu Duyu Tang Neel Sundaresan Ming Zhou Ambrosio Blanco and Shuai Ma. 2020. CodeBLEU: A method for automatic evaluation of code synthesis. arXiv:2009.10297. Retrieved from https:\/\/arxiv.org\/abs\/2009.10297"},{"key":"e_1_3_2_43_2","unstructured":"Baptiste Roziere Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing Ellen Tan Yossi Adi Jingyu Liu Tal Remez J\u00e9r\u00e9my Rapin Artyom Kozhevnikov Ivan Evtimov Joanna Bitton Manish Bhatt Cristian Canton Ferrer Aaron Grattafiori Wenhan Xiong Alexandre D\u00e9fossez Jade Copet Faisal Azhar Hugo Touvron Louis Martin Nicolas Usunier Thomas Scialom and Gabriel Synnaeve. 2023. Code llama: Open foundation models for code. arXiv:2308.12950. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2308.12950"},{"key":"e_1_3_2_44_2","first-page":"1057","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"12","author":"Sutton Richard S.","year":"1999","unstructured":"Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour. 1999. Policy gradient methods for reinforcement learning with function approximation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 12, 1057\u20131063."},{"key":"e_1_3_2_45_2","first-page":"1433","volume-title":"Proceedings of the 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC\/FSE \u201920)","author":"Svyatkovskiy Alexey","year":"2020","unstructured":"Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020. IntelliCode compose: code generation using transformer. In Proceedings of the 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC\/FSE \u201920). Prem Devanbu, Myra B. Cohen, and Thomas Zimmermann (Eds.), ACM, 1433\u20131443."},{"key":"e_1_3_2_46_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale Dan Bikel Lukas Blecher Cristian Canton Ferrer Moya Chen Guillem Cucurull David Esiobu Jude Fernandes Jeremy Fu Wenyin Fu Brian Fuller Cynthia Gao Vedanuj Goswami Naman Goyal Anthony Hartshorn Saghar Hosseini Rui Hou Hakan Inan Marcin Kardas Viktor Kerkez Madian Khabsa Isabel Kloumann Artem Korenev Punit Singh Koura Marie-Anne Lachaux Thibaut Lavril Jenya Lee Diana Liskovich Yinghai Lu Yuning Mao Xavier Martinet Todor Mihaylov Pushkar Mishra Igor Molybog Yixin Nie Andrew Poulton Jeremy Reizenstein Rashi Rungta Kalyan Saladi Alan Schelten Ruan Silva Eric Michael Smith Ranjan Subramanian Xiaoqing Ellen Tan Binh Tang Ross Taylor Adina Williams Jian Xiang Kuan Puxin Xu Zheng Yan Iliyan Zarov Yuchen Zhang Angela Fan Melanie Kambadur Sharan Narang Aurelien Rodriguez Robert Stojnic Sergey Edunov and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2307.09288"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340544"},{"key":"e_1_3_2_48_2","unstructured":"Xin Wang Yasheng Wang Fei Mi Pingyi Zhou Yao Wan Xiao Liu Li Li Hao Wu Jin Liu and Xin Jiang. 2021. Syncobert: Syntax-guided multi-modal contrastive pre-training for code representation. arXiv:2108.04556. Retrieved from https:\/\/arxiv.org\/abs\/2108.04556"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.685"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992696"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.5555\/3455716.3455759"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510146"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482319"},{"key":"e_1_3_2_54_2","first-page":"3154","volume-title":"Proceedings of the 39th IEEE International Conference on Data Engineering (ICDE \u201923)","author":"Yao Kaichun","year":"2023","unstructured":"Kaichun Yao, Jingshuai Zhang, Chuan Qin, Xin Song, Peng Wang, Hengshu Zhu, and Hui Xiong. 2023. ResuFormer: Semantic structure understanding for resumes via multi-modal pre-training. In Proceedings of the 39th IEEE International Conference on Data Engineering (ICDE \u201923). IEEE, 3154\u20133167."},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2018.01.020"},{"key":"e_1_3_2_56_2","first-page":"1","volume-title":"Proceedings of the ACM on Programming Languages","volume":"4","author":"Yefet Noam","year":"2020","unstructured":"Noam Yefet, Uri Alon, and Eran Yahav. 2020. Adversarial examples for models of code. Proceedings of the ACM on Programming Languages 4, OOPSLA (2020), 1\u201330."},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","unstructured":"Daoguang Zan Bei Chen Dejian Yang Zeqi Lin Minsu Kim Bei Guan Yongji Wang Weizhu Chen and Jian-Guang Lou. 2022. CERT: continual pre-training on sketches for library-oriented code generation. arXiv:2206.06888. Retrieved from https:\/\/doi.org\/10.24963\/ijcai.2022\/329","DOI":"10.24963\/ijcai.2022\/329"},{"key":"e_1_3_2_58_2","first-page":"1","volume-title":"ACM Transactions on Software Engineering and Methodology (TOSEM)","volume":"31","author":"Zhang Huangzhao","year":"2022","unstructured":"Huangzhao Zhang, Zhiyi Fu, Ge Li, Lei Ma, Zhehao Zhao, Hua\u2019an Yang, Yizhe Sun, Yang Liu, and Zhi Jin. 2022. Towards robustness of deep program processing models\u2014detection, estimation, and enhancement. ACM Transactions on Software Engineering and Methodology (TOSEM) 31, 3 (2022), 1\u201340."},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i01.5469"},{"key":"e_1_3_2_60_2","doi-asserted-by":"crossref","unstructured":"Qinkai Zheng Xiao Xia Xu Zou Yuxiao Dong Shan Wang Yufei Xue Zihan Wang Lei Shen Andi Wang Yang Li Teng Su Zhilin Yang and Jie Tang. 2023. Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x. arXiv:2303.17568. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2303.17568","DOI":"10.1145\/3580305.3599790"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3501256"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3688839","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3688839","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:04:10Z","timestamp":1750291450000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3688839"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,28]]},"references-count":60,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1,31]]}},"alternative-id":["10.1145\/3688839"],"URL":"https:\/\/doi.org\/10.1145\/3688839","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,28]]},"assertion":[{"value":"2023-12-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-07-29","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}