{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,28]],"date-time":"2026-08-28T16:48:28Z","timestamp":1787935708366,"version":"build-2784847793"},"reference-count":151,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"name":"Science and Technology Development Fund of Macau SAR, China","award":["fdct0080\/2024\/RIA2"],"award-info":[{"award-number":["fdct0080\/2024\/RIA2"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,3,31]]},"abstract":"<jats:p>With the rapid development of artificial intelligence, large language models (LLMs) have made remarkable advancements in natural language processing. These models are trained on vast datasets to exhibit powerful language understanding and generation capabilities across various applications, including chatbots and agents. However, LLMs have revealed a variety of privacy and security issues throughout their life cycle, drawing significant academic and industrial attention. Moreover, the risks faced by LLMs differ significantly from those encountered by traditional language models. Given that current surveys lack a clear taxonomy of unique threat models across diverse scenarios, we emphasize the unique privacy and security threats associated with four specific scenarios: pre-training, fine-tuning, deployment, and LLM-based agents. Addressing the characteristics of each risk, this survey outlines and analyzes potential countermeasures. Research on attack and defense situations can offer feasible research directions, enabling more areas to benefit from LLMs.<\/jats:p>","DOI":"10.1145\/3764113","type":"journal-article","created":{"date-parts":[[2025,9,11]],"date-time":"2025-09-11T10:37:42Z","timestamp":1757587062000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":38,"title":["Unique Security and Privacy Threats of Large Language Models: A Comprehensive Survey"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5114-4659","authenticated-orcid":false,"given":"Shang","family":"Wang","sequence":"first","affiliation":[{"name":"School of Computer Science, University of Technology Sydney","place":["Ultimo, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0702-7102","authenticated-orcid":false,"given":"Tianqing","family":"Zhu","sequence":"additional","affiliation":[{"name":"Faculty of Data Science, City University of Macau","place":["Taipa, Macao"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3603-6617","authenticated-orcid":false,"given":"Bo","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer Science, University of Technology Sydney","place":["Ultimo, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3690-0321","authenticated-orcid":false,"given":"Ming","family":"Ding","sequence":"additional","affiliation":[{"name":"Data61","place":["Eveleigh, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7561-0992","authenticated-orcid":false,"given":"Dayong","family":"Ye","sequence":"additional","affiliation":[{"name":"School of Computer Science, University of Technology Sydney","place":["Ultimo, Australia"]},{"name":"Faculty of Data Science, City University of Macau","place":["Ultimo, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1680-2521","authenticated-orcid":false,"given":"Wanlei","family":"Zhou","sequence":"additional","affiliation":[{"name":"City University of Macau","place":["Taipa, Macao"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3491-5968","authenticated-orcid":false,"given":"Philip","family":"Yu","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Illinois at Chicago","place":["Chicago, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,10,6]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/2976749.2978318"},{"key":"e_1_3_1_3_2","doi-asserted-by":"crossref","unstructured":"John Abascal Stanley Wu Alina Oprea and Jonathan Ullman. 2024. Tmi! finetuned models leak private information from their pretraining data. Proceedings on Privacy Enhancing Technologies 2024 3 (2024) 202\u2013223.","DOI":"10.56553\/popets-2024-0075"},{"key":"e_1_3_1_4_2","unstructured":"Accountability Act. 1996. Health insurance portability and accountability act of 1996. Public Law 104 (1996) 191."},{"key":"e_1_3_1_5_2","first-page":"2255","volume-title":"Proceedings of the 30th USENIX Security Symposium (USENIX Security 21)","author":"Azizi Ahmadreza","year":"2021","unstructured":"Ahmadreza Azizi, Ibrahim Asadullah Tahmid, Asim Waheed, Neal Mangaokar, Jiameng Pu, Mobin Javed, Chandan K Reddy, and Bimal Viswanath. 2021. {T-Miner}: A generative approach to defend against trojan attacks on {DNN-based} text classification. In Proceedings of the 30th USENIX Security Symposium (USENIX Security 21). 2255\u20132272."},{"key":"e_1_3_1_6_2","unstructured":"Yuntao Bai Andy Jones Kamal Ndousse Amanda Askell Anna Chen Nova DasSarma Dawn Drain Stanislav Fort Deep Ganguli Tom Henighan et\u00a0al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv:1701.00133. Retrieved from https:\/\/arxiv.org\/abs\/1701.00133"},{"key":"e_1_3_1_7_2","volume-title":"Proceedings of the 1st Conference on Language Modeling","author":"Baumg\u00e4rtner Tim","year":"2024","unstructured":"Tim Baumg\u00e4rtner, Yang Gao, Dana Alon, and Donald Metzler. 2024. Best-of-venom: Attacking RLHF by injecting poisoned preference data. In Proceedings of the 1st Conference on Language Modeling."},{"key":"e_1_3_1_8_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Bianchi Federico","year":"2023","unstructured":"Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Rottger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou. 2023. Safety-tuned LLaMAs: Lessons from improving the safety of large language models that follow instructions. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_1_9_2","unstructured":"Xiangrui Cai Haidong Xu Sihan Xu Ying Zhang and Xiaojie Yuan. 2022. Badprompt: Backdoor attacks on continuous prompts. Advances in Neural Information Processing Systems 35 (2022) 37068\u201337080."},{"key":"e_1_3_1_10_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Carlini Nicholas","year":"2022","unstructured":"Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022. Quantifying memorization across neural language models. In Proceedings of the 11th International Conference on Learning Representations."},{"key":"e_1_3_1_11_2","volume-title":"Proceedings of the 37th Annual Conference on Neural Information Processing Systems (NeurIPS 2023)","author":"Carlini Nicholas","year":"2023","unstructured":"Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tram\u00e8r, and Ludwig Schmidt. 2023. Are aligned neural networks adversarially aligned?. In Proceedings of the 37th Annual Conference on Neural Information Processing Systems (NeurIPS 2023)."},{"key":"e_1_3_1_12_2","first-page":"2633","volume-title":"Proceedings of the 30th USENIX Security Symposium (USENIX Security 21)","author":"Carlini Nicholas","year":"2021","unstructured":"Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et\u00a0al. 2021. Extracting training data from large language models. In Proceedings of the 30th USENIX Security Symposium (USENIX Security 21). 2633\u20132650."},{"key":"e_1_3_1_13_2","unstructured":"Chi-Min Chan Jianxuan Yu Weize Chen Chunyang Jiang Xinyu Liu Weijie Shi Zhiyuan Liu Wei Xue and Yike Guo. 2024. Agentmonitor: A plug-and-play framework for predictive and secure multi-agent systems. arXiv:2408.14972 (2024)."},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","first-page":"3510","DOI":"10.18653\/v1\/2022.findings-acl.277","volume-title":"Findings of the Association for Computational Linguistics: ACL 2022","author":"Chen Tianyu","year":"2022","unstructured":"Tianyu Chen, Hangbo Bao, Shaohan Huang, Li Dong, Binxing Jiao, Daxin Jiang, Haoyi Zhou, Jianxin Li, and Furu Wei. 2022. THE-X: Privacy-preserving transformer inference with homomorphic encryption. In Findings of the Association for Computational Linguistics: ACL 2022. 3510\u20133520."},{"key":"e_1_3_1_15_2","first-page":"1","volume-title":"Proceedings of the 2023 IEEE International Conference on Electro Information Technology (eIT)","author":"Chowdhury M. D. Minhaz","year":"2023","unstructured":"M. D. Minhaz Chowdhury, Nafiz Rifat, Mostofa Ahsan, Shadman Latif, Rahul Gomes, and Md Saifur Rahman. 2023. ChatGPT: A threat against the CIA triad of cyber security. In Proceedings of the 2023 IEEE International Conference on Electro Information Technology (eIT). IEEE, 1\u20136."},{"key":"e_1_3_1_16_2","first-page":"5009","article-title":"A unified evaluation of textual backdoor learning: Frameworks and benchmarks","volume":"35","author":"Cui Ganqu","year":"2022","unstructured":"Ganqu Cui, Lifan Yuan, Bingxiang He, Yangyi Chen, Zhiyuan Liu, and Maosong Sun. 2022. A unified evaluation of textual backdoor learning: Frameworks and benchmarks. Advances in Neural Information Processing Systems 35 (2022), 5009\u20135023.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_17_2","unstructured":"Tianyu Cui Yanling Wang Chuanpu Fu Yong Xiao Sijia Li Xinhao Deng Yunpeng Liu Qinglin Zhang Ziyi Qiu Peiyang Li et\u00a0al. 2024. Risk taxonomy mitigation and assessment benchmarks of large language model systems. arXiv:2401.05778 (2024)."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3712001"},{"key":"e_1_3_1_19_2","volume-title":"Proceedings of the Network and Distributed System Security Symposium, NDSS 2024","author":"Deng Gelei","year":"2024","unstructured":"Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2024. Masterkey: Automated jailbreak across multiple large language model chatbots. In Proceedings of the Network and Distributed System Security Symposium, NDSS 2024. The Internet Society."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.88"},{"key":"e_1_3_1_21_2","volume-title":"Proceedings of the Network and Distributed System Security Symposium, NDSS 2025","author":"Dong Tian","year":"2025","unstructured":"Tian Dong, Minhui Xue, Guoxing Chen, Rayne Holland, Yan Meng, Shaofeng Li, Zhen Liu, and Haojin Zhu. 2025. The philosopher\u2019s stone: Trojaning plugins of large language models. In Proceedings of the Network and Distributed System Security Symposium, NDSS 2025. The Internet Society."},{"key":"e_1_3_1_22_2","unstructured":"Ye Dong Wen-jie Lu Yancheng Zheng Haoqi Wu Derun Zhao Jin Tan Zhicong Huang Cheng Hong Tao Wei and Wenguang Cheng. 2023. Puma: Secure inference of llama-7b in five minutes. arXiv:2307.12533 (2023)."},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"Zhichen Dong Zhanhui Zhou Chao Yang Jing Shao and Yu Qiao. 2024. Attacks defenses and evaluations for llm conversation safety: A survey. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 6734\u20136747. arXiv:2402.09283.","DOI":"10.18653\/v1\/2024.naacl-long.375"},{"key":"e_1_3_1_24_2","unstructured":"Haonan Duan Adam Dziedzic Nicolas Papernot and Franziska Boenisch. 2023. Flocks of stochastic parrots: Differentially private prompt learning for large language models. Advances in Neural Information Processing Systems 36 (2023) 76852\u201376871."},{"key":"e_1_3_1_25_2","unstructured":"Ronen Eldan and Mark Russinovich. 2023. Who\u2019s harry potter? Approximate unlearning in LLMs. arXiv:2310.02238 (2023)."},{"issue":"4","key":"e_1_3_1_26_2","first-page":"2349","article-title":"Design and evaluation of a multi-domain trojan detection method on deep neural networks","volume":"19","author":"Gao Yansong","year":"2021","unstructured":"Yansong Gao, Yeonjae Kim, Bao Gia Doan, Zhi Zhang, Gongxuan Zhang, Surya Nepal, Damith C. Ranasinghe, and Hyoungshick Kim. 2021. Design and evaluation of a multi-domain trojan detection method on deep neural networks. IEEE Transactions on Dependable and Secure Computing 19, 4 (2021), 2349\u20132364.","journal-title":"IEEE Transactions on Dependable and Secure Computing"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3605764.3623985"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.464"},{"key":"e_1_3_1_29_2","unstructured":"Daya Guo Dejian Yang Haowei Zhang Junxiao Song Ruoyu Zhang Runxin Xu Qihao Zhu Shirong Ma Peiyi Wang Xiao Bi et\u00a0al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv:2501.12948 (2025)."},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","unstructured":"Maanak Gupta CharanKumar Akiri Kshitiz Aryal Eli Parker and Lopamudra Praharaj. 2023. From chatgpt to threatgpt: Impact of generative AI in cybersecurity and privacy. IEEE Access 11 (2023) 80218\u201380245.","DOI":"10.1109\/ACCESS.2023.3300381"},{"key":"e_1_3_1_31_2","unstructured":"Feng He Tianqing Zhu Dayong Ye Bo Liu Wanlei Zhou and Philip S. Yu. 2024. The emerged security and privacy of llm agent: A survey with case studies. arXiv:2407.19354 (2024)."},{"key":"e_1_3_1_32_2","volume-title":"Proceedings of the 34th USENIX Security Symposium (USENIX Security 25)","author":"He Yu","year":"2025","unstructured":"Yu He, Boheng Li, Liu Liu, Zhongjie Ba, Wei Dong, Yiming Li, Zhan Qin, Kui Ren, and Chun Chen. 2025. Towards label-only membership inference attack against pre-trained large language models. In Proceedings of the 34th USENIX Security Symposium (USENIX Security 25)."},{"key":"e_1_3_1_33_2","unstructured":"Jen-tse Huang Jiaxu Zhou Tailin Jin Xuhui Zhou Zixi Chen Wenxuan Wang Youliang Yuan Michael R. Lyu and Maarten Sap. 2025. On the resilience of LLM-based multi-agent collaboration with faulty agents. In Proceedings of the International Conference on Machine Learning. https:\/\/openreview.net\/forum?id=bkiM54QftZ"},{"key":"e_1_3_1_34_2","unstructured":"Ken Huang Vineeth Sai Narajala John Yeoh Ramesh Raskar Youssef Harkati Jerry Huang Idan Habler and Chris Hughes. 2025. A novel zero-trust identity framework for agentic AI: Decentralized authentication and fine-grained access control. arXiv:2505.19301 (2025)."},{"key":"e_1_3_1_35_2","unstructured":"Tiansheng Huang Sihao Hu Fatih Ilhan Selim Furkan Tekin and Ling Liu. 2024. Harmful fine-tuning attacks and defenses for large language models: A survey. arXiv:2409.18169 (2024)."},{"key":"e_1_3_1_36_2","unstructured":"Yue Huang Qihui Zhang Philip S. Yu and Lichao Sun. 2023. Trustgpt: A benchmark for trustworthy and responsible large language models. arXiv:2306.11507 (2023)."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3543507.3583348"},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"3600","DOI":"10.1145\/3658644.3670370","volume-title":"Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security","author":"Hui Bo","year":"2024","unstructured":"Bo Hui, Haolin Yuan, Neil Gong, Philippe Burlina, and Yinzhi Cao. 2024. Pleak: Prompt leaking attacks against large language model applications. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 3600\u20133614."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.inlg-main.28"},{"key":"e_1_3_1_40_2","unstructured":"Abhyuday Jagannatha Bhanu Pratap Singh Rawat and Hong Yu. 2021. Membership inference attack susceptibility of clinical language models. arXiv:2104.08305 (2021)."},{"key":"e_1_3_1_41_2","unstructured":"Changyue Jiang Xudong Pan Geng Hong Chenfu Bao and Min Yang. 2025. Feedback-guided extraction of knowledge base from retrieval-augmented LLM applications. arXiv:2411.14110 (2025)."},{"key":"e_1_3_1_42_2","unstructured":"Shuyu Jiang Xingshu Chen and Rui Tang. 2023. Prompt packer: Deceiving llms through compositional instruction with hidden attacks. arXiv:2310.10077 (2023)."},{"key":"e_1_3_1_43_2","first-page":"10697","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Kandpal Nikhil","year":"2022","unstructured":"Nikhil Kandpal, Eric Wallace, and Colin Raffel. 2022. Deduplicating training data mitigates privacy risks in language models. In Proceedings of the International Conference on Machine Learning. PMLR, 10697\u201310707."},{"key":"e_1_3_1_44_2","unstructured":"Siwon Kim Sangdoo Yun Hwaran Lee Martin Gubri Sungroh Yoon and Seong Joon Oh. 2023. Propile: Probing privacy leakage in large language models. Advances in Neural Information Processing Systems 36 (2023) 20750\u201320762."},{"key":"e_1_3_1_45_2","first-page":"17061","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Kirchenbauer John","year":"2023","unstructured":"John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023. A watermark for large language models. In Proceedings of the International Conference on Machine Learning. PMLR, 17061\u201317084."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.243"},{"key":"e_1_3_1_47_2","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Li Haoran","year":"2023","unstructured":"Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song. 2023. Multi-step jailbreaking privacy attacks on ChatGPT. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing."},{"key":"e_1_3_1_48_2","first-page":"5858","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Li Haoran","year":"2022","unstructured":"Haoran Li, Yangqiu Song, and Lixin Fan. 2022. You Don\u2019t know my favorite color: Preventing dialogue representations from revealing speakers\u2019 private personas. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 5858\u20135870."},{"key":"e_1_3_1_49_2","unstructured":"Jiazhao Li Yijin Yang Zhuofeng Wu V. G. Vydiswaran and Chaowei Xiao. 2024. ChatGPT as an attack tool: Stealthy textual backdoor attack via blackbox generative model trigger. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2985\u20133004."},{"key":"e_1_3_1_50_2","first-page":"338","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Li Linyang","year":"2023","unstructured":"Linyang Li, Demin Song, and Xipeng Qiu. 2023. Text adversarial purification as defense against adversarial attacks. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 338\u2013350."},{"key":"e_1_3_1_51_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Li Xuechen","year":"2021","unstructured":"Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto. 2021. Large language models can be strong differentially private learners. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_52_2","unstructured":"Yanzhou Li Tianlin Li Kangjie Chen Jian Zhang Shangqing Liu Wenhan Wang Tianwei Zhang and Yang Liu. 2024. Badedit: Backdooring large language models by model editing. In Proceedings of the International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=duZANm2ABX"},{"key":"e_1_3_1_53_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Li Yige","year":"2020","unstructured":"Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. 2020. Neural attention distillation: Erasing backdoor triggers from deep neural networks. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_54_2","unstructured":"Yuanchun Li Hao Wen Weijun Wang Xiangyu Li Yizhen Yuan Guohong Liu Jiacheng Liu Wenxing Xu Xiang Wang Yi Sun et\u00a0al. 2024. Personal llm agents: Insights and survey about the capability efficiency and security. arXiv:2401.05459 (2024)."},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639091"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-00470-5_13"},{"key":"e_1_3_1_57_2","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Liu Xiao","year":"2022","unstructured":"Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2022. P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics."},{"key":"e_1_3_1_58_2","unstructured":"Xiaogeng Liu Nan Xu Muhao Chen and Chaowei Xiao. 2024. Autodan: Generating stealthy jailbreak prompts on aligned large language models. In Proceedings of the International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=7Jwpw4qKkb"},{"key":"e_1_3_1_59_2","doi-asserted-by":"crossref","unstructured":"Xiao Liu Yanan Zheng Zhengxiao Du Ming Ding Yujie Qian Zhilin Yang and Jie Tang. 2024. GPT understands too. AI Open 5 (2024) 208\u2013215.","DOI":"10.1016\/j.aiopen.2023.08.012"},{"key":"e_1_3_1_60_2","unstructured":"Yi Liu Gelei Deng Yuekang Li Kailong Wang Tianwei Zhang Yepang Liu Haoyu Wang Yan Zheng and Yang Liu. 2023. Prompt injection attack against LLM-integrated applications. arXiv:2306.05499 (2023)."},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3663530.3665021"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.469"},{"key":"e_1_3_1_63_2","doi-asserted-by":"crossref","first-page":"4465","DOI":"10.1145\/3658644.3670361","volume-title":"Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security","author":"Ma Hua","year":"2024","unstructured":"Hua Ma, Shang Wang, Yansong Gao, Zhi Zhang, Huming Qiu, Minhui Xue, Alsharif Abuadbba, Anmin Fu, Surya Nepal, and Derek Abbott. 2024. Watch out! simple horizontal class backdoor can trivially evade defense. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 4465\u20134479."},{"key":"e_1_3_1_64_2","unstructured":"Jimit Majmudar Christophe Dupuy Charith Peris Sami Smaili Rahul Gupta and Richard Zemel. 2022. Differentially private decoding in large language models. arXiv:2205.13621 (2022)."},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","first-page":"4860","DOI":"10.18653\/v1\/2022.emnlp-main.323","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Mattern Justus","year":"2022","unstructured":"Justus Mattern, Zhijing Jin, Benjamin Weggenmann, Bernhard Schoelkopf, and Mrinmaya Sachan. 2022. Differentially private language models for secure data sharing. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 4860\u20134873."},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.719"},{"key":"e_1_3_1_67_2","volume-title":"Proceedings of the Second Workshop on New Frontiers in Adversarial Machine Learning","author":"Maus Natalie","year":"2023","unstructured":"Natalie Maus, Patrick Chao, Eric Wong, and Jacob R. Gardner. 2023. Black box adversarial prompting for foundation models. In Proceedings of the Second Workshop on New Frontiers in Adversarial Machine Learning."},{"key":"e_1_3_1_68_2","doi-asserted-by":"crossref","first-page":"12448","DOI":"10.18653\/v1\/2023.emnlp-main.765","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Morris John Xavier","year":"2023","unstructured":"John Xavier Morris, Volodymyr Kuleshov, Vitaly Shmatikov, and Alexander M. Rush. 2023. Text embeddings reveal (almost) as much as text. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 12448\u201312460."},{"key":"e_1_3_1_69_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Morris John Xavier","year":"2024","unstructured":"John Xavier Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov, and Alexander M. Rush. 2024. Language model inversion. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3576915.3616652"},{"key":"e_1_3_1_71_2","volume-title":"Proceedings of the 13th International Conference on Learning Representations","author":"Nasr Milad","year":"2025","unstructured":"Milad Nasr, Javier Rando, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Florian Tram\u00e8r, and Katherine Lee. 2025. Scalable extraction of training data from aligned, production language models. In Proceedings of the 13th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=vjel3nWP2a"},{"key":"e_1_3_1_72_2","volume-title":"Proceedings of the Network and Distributed System Security Symposium, NDSS 2024","author":"Pei Hengzhi","year":"2024","unstructured":"Hengzhi Pei, Jinyuan Jia, Wenbo Guo, Bo Li, and Dawn Song. 2024. TextGuard: Provable defense against backdoor attacks on text classification. In Proceedings of the Network and Distributed System Security Symposium, NDSS 2024. The Internet Society."},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.423"},{"key":"e_1_3_1_74_2","volume-title":"Proceedings of the NeurIPS ML Safety Workshop","author":"Perez F\u00e1bio","year":"2022","unstructured":"F\u00e1bio Perez and Ian Ribeiro. 2022. Ignore previous prompt: Attack techniques for language models. In Proceedings of the NeurIPS ML Safety Workshop."},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.752"},{"key":"e_1_3_1_76_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Rando Javier","year":"2023","unstructured":"Javier Rando and Florian Tram\u00e8r. 2023. Universal jailbreak backdoors from poisoned human feedback. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_1_77_2","volume-title":"Proceedings of the R0-FoMo: Robustness of Few-shot and Zero-shot Learning in Large Foundation Models","author":"Robey Alexander","year":"2023","unstructured":"Alexander Robey, Eric Wong, Hamed Hassani, and George Pappas. 2023. SmoothLLM: Defending large language models against jailbreaking attacks. In Proceedings of the R0-FoMo: Robustness of Few-shot and Zero-shot Learning in Large Foundation Models."},{"key":"e_1_3_1_78_2","doi-asserted-by":"crossref","unstructured":"Sahar Sadrizadeh Ljiljana Dolamic and Pascal Frossard. 2023. TransFool: An adversarial attack against neural machine translation models. Transactions on Machine Learning Research (2023). https:\/\/openreview.net\/forum?id=sFk3aBNb81","DOI":"10.23919\/EUSIPCO58844.2023.10289979"},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.244"},{"key":"e_1_3_1_80_2","doi-asserted-by":"crossref","unstructured":"Shawn Shan Wenxin Ding Josephine Passananti Stanley Wu Haitao Zheng and Ben Y. Zhao. 2024. Nightshade: Prompt-specific poisoning attacks on text-to-image generative models. In Proceedings of the 2024 IEEE Symposium on Security and Privacy (SP). IEEE Computer Society 807\u2013825.","DOI":"10.1109\/SP54263.2024.00207"},{"key":"e_1_3_1_81_2","doi-asserted-by":"crossref","unstructured":"M. Shanahan K. McDonell and L. Reynolds. 2023. Role play with large language models. Nature 623 7987 (2023) 493\u2013498.","DOI":"10.1038\/s41586-023-06647-8"},{"key":"e_1_3_1_82_2","doi-asserted-by":"crossref","first-page":"102433","DOI":"10.1016\/j.cose.2021.102433","article-title":"Bddr: An effective defense against textual backdoor attacks","volume":"110","author":"Shao Kun","year":"2021","unstructured":"Kun Shao, Junan Yang, Yang Ai, Hui Liu, and Yu Zhang. 2021. Bddr: An effective defense against textual backdoor attacks. Computers and Security 110 (2021), 102433.","journal-title":"Computers and Security"},{"key":"e_1_3_1_83_2","first-page":"103","volume-title":"Proceedings of the 2025 IEEE Symposium on Security and Privacy (SP)","author":"Shen Guangyu","year":"2024","unstructured":"Guangyu Shen, Siyuan Cheng, Zhuo Zhang, Guanhong Tao, Kaiyuan Zhang, Hanxi Guo, Lu Yan, Xiaolong Jin, Shengwei An, Shiqing Ma, et\u00a0al. 2024. BAIT: Large language model backdoor scanning by inverting attack target. In Proceedings of the 2025 IEEE Symposium on Security and Privacy (SP). IEEE Computer Society, 103\u2013103."},{"key":"e_1_3_1_84_2","unstructured":"Lingfeng Shen Haiyun Jiang Lemao Liu and Shuming Shi. 2022. Rethink the Evaluation for Attack Strength of Backdoor Attacks in Natural Language Processing. arXiv:2201.02993 (2022)."},{"key":"e_1_3_1_85_2","doi-asserted-by":"publisher","DOI":"10.1145\/3658644.3670388"},{"key":"e_1_3_1_86_2","doi-asserted-by":"publisher","DOI":"10.1145\/3658644.3690291"},{"key":"e_1_3_1_87_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.205"},{"key":"e_1_3_1_88_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.346"},{"key":"e_1_3_1_89_2","first-page":"61836","article-title":"On the exploitability of instruction tuning","volume":"36","author":"Shu Manli","year":"2023","unstructured":"Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein. 2023. On the exploitability of instruction tuning. Advances in Neural Information Processing Systems 36 (2023), 61836\u201361856.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_90_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Staab Robin","year":"2023","unstructured":"Robin Staab, Mark Vero, Mislav Balunovic, and Martin Vechev. 2023. Beyond memorization: Violating privacy via inference with large language models. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_1_91_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.trustnlp-1.18"},{"key":"e_1_3_1_92_2","first-page":"1620","volume-title":"Proceedings of the 2025 IEEE Symposium on Security and Privacy (SP)","author":"Sun Zhen","year":"2024","unstructured":"Zhen Sun, Tianshuo Cong, Yule Liu, Chenhao Lin, Xinlei He, Rongmao Chen, Xingshuo Han, and Xinyi Huang. 2024. PEFTGuard: Detecting backdoor attacks against parameter-efficient fine-tuning. In Proceedings of the 2025 IEEE Symposium on Security and Privacy (SP). IEEE Computer Society, 1620\u20131638."},{"key":"e_1_3_1_93_2","unstructured":"Zhiqing Sun Yikang Shen Qinhong Zhou Hongxin Zhang Zhenfang Chen David Cox Yiming Yang and Chuang Gan. 2023. Principle-driven self-alignment of language models from scratch with minimal human supervision. Advances in Neural Information Processing Systems 36 (2023) 2511\u20132565."},{"key":"e_1_3_1_94_2","unstructured":"Yu Tian Xiao Yang Jingyuan Zhang Yinpeng Dong and Hang Su. 2023. Evil geniuses: Delving into the safety of LLM-based agents. arXiv:2311.11855 (2023)."},{"key":"e_1_3_1_95_2","first-page":"11117","article-title":"Seqpate: Differentially private text generation via knowledge distillation","volume":"35","author":"Tian Zhiliang","year":"2022","unstructured":"Zhiliang Tian, Yingxiu Zhao, Ziyue Huang, Yu-Xiang Wang, Nevin L. Zhang, and He He. 2022. Seqpate: Differentially private text generation via knowledge distillation. Advances in Neural Information Processing Systems 35 (2022), 11117\u201311130.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_96_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et\u00a0al. 2023. Llama: Open and efficient foundation language models. arXiv:2302.13971 (2023)."},{"issue":"1","key":"e_1_3_1_97_2","doi-asserted-by":"crossref","first-page":"37","DOI":"10.69554\/KZRS2422","article-title":"Machine unlearning for generative AI","volume":"3","author":"Viswanath Yashaswini","year":"2024","unstructured":"Yashaswini Viswanath, Sudha Jamthe, Suresh Lokiah, and Emanuele Bianchini. 2024. Machine unlearning for generative AI. Journal of AI, Robotics and Workplace Automation 3, 1 (2024), 37\u201346.","journal-title":"Journal of AI, Robotics and Workplace Automation"},{"key":"e_1_3_1_98_2","first-page":"35413","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wan Alexander","year":"2023","unstructured":"Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023. Poisoning language models during instruction tuning. In Proceedings of the International Conference on Machine Learning. PMLR, 35413\u201335425."},{"key":"e_1_3_1_99_2","doi-asserted-by":"crossref","unstructured":"Bo Wang Weiyi He Pengfei He Shenglai Zeng Zhen Xiang Yue Xing and Jiliang Tang. 2025. Unveiling privacy risks in LLM agent memory. arXiv:2502.13172 (2025).","DOI":"10.18653\/v1\/2025.acl-long.1227"},{"key":"e_1_3_1_100_2","doi-asserted-by":"crossref","unstructured":"Rupeng Zhang Haowei Wang Junjie Wang Mingyang Li Yuekai Huang Dandan Wang and Qing Wang. 2025. From allies to adversaries: Manipulating llm tool-calling through adversarial injection. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2009\u20132028.","DOI":"10.18653\/v1\/2025.naacl-long.101"},{"key":"e_1_3_1_101_2","unstructured":"Jiongxiao Wang Zichen Liu Keun Hee Park Muhao Chen and Chaowei Xiao. 2023. Adversarial demonstration attacks on large language models. arXiv:2305.14950 (2023)."},{"key":"e_1_3_1_102_2","doi-asserted-by":"crossref","first-page":"2551","DOI":"10.18653\/v1\/2024.acl-long.140","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wang Jiongxiao","year":"2024","unstructured":"Jiongxiao Wang, Junlin Wu, Muhao Chen, Yevgeniy Vorobeychik, and Chaowei Xiao. 2024. RLHFPoison: Reward poisoning attack for reinforcement learning with human feedback in large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2551\u20132570."},{"key":"e_1_3_1_103_2","unstructured":"Kun Wang Guibin Zhang Zhenhong Zhou Jiahao Wu Miao Yu Shiqian Zhao Chenlong Yin Jinhu Fu Yibo Yan Hanjun Luo et\u00a0al. 2025. A comprehensive survey in LLM (-agent) full stack safety: Data training and deployment. arXiv:2504.15585 (2025)."},{"key":"e_1_3_1_104_2","doi-asserted-by":"publisher","DOI":"10.1145\/3579856.3582829"},{"key":"e_1_3_1_105_2","doi-asserted-by":"crossref","unstructured":"Shilong Wang Guibin Zhang Miao Yu Guancheng Wan Fanci Meng Chongye Guo Kun Wang and Yang Wang. 2025. G-safeguard: A topology-guided security lens and treatment on LLM-based multi-agent systems. arXiv:2502.11127 (2025).","DOI":"10.18653\/v1\/2025.acl-long.359"},{"key":"e_1_3_1_106_2","unstructured":"ShangWang Tianqing Zhu Dayong Ye andWanlei Zhou. 2024. When machine unlearning meets retrieval-augmented generation (RAG): Keep secret or forget knowledge? arXiv:2410.15267 (2024)."},{"key":"e_1_3_1_107_2","volume-title":"Proceedings of the Network and Distributed System Security Symposium, NDSS 2024","author":"Wei Chengkun","year":"2024","unstructured":"Chengkun Wei, Wenlong Meng, Zhikun Zhang, Min Chen, Minghu Zhao, Wenjing Fang, Lei Wang, Zihui Zhang, and Wenzhi Chen. 2024. LMSanitator: Defending prompt-tuning against task-agnostic backdoors. In Proceedings of the Network and Distributed System Security Symposium, NDSS 2024. The Internet Society."},{"key":"e_1_3_1_108_2","doi-asserted-by":"crossref","unstructured":"Jiali Wei Ming Fan Wenjing Jiao Wuxia Jin and Ting Liu. 2024. Bdmmt: Backdoor sample detection for language models through model mutation testing. IEEE Transactions on Information Forensics and Security 19 (2024) 4285\u20134300.","DOI":"10.1109\/TIFS.2024.3376968"},{"key":"e_1_3_1_109_2","unstructured":"Jason Wei Yi Tay Rishi Bommasani Colin Raffel Barret Zoph Sebastian Borgeaud Dani Yogatama Maarten Bosma Denny Zhou Donald Metzler et\u00a0al. 2022. Emergent abilities of large language models. Transactions on Machine Learning Research (2022). https:\/\/openreview.net\/forum?id=yzkSU5zdwD"},{"key":"e_1_3_1_110_2","unstructured":"Zeming Wei Yifei Wang and Yisen Wang. 2023. Jailbreak and guard aligned language models with only few in-context demonstrations. arXiv:2310.06387 (2023)."},{"key":"e_1_3_1_111_2","doi-asserted-by":"crossref","first-page":"3481","DOI":"10.1145\/3658644.3690306","volume-title":"Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security","author":"Wen Rui","year":"2024","unstructured":"Rui Wen, Zheng Li, Michael Backes, and Yang Zhang. 2024. Membership inference attacks against in-context learning. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 3481\u20133495."},{"key":"e_1_3_1_112_2","unstructured":"Yotam Wolf Noam Wies Oshri Avnery Yoav Levine and Amnon Shashua. 2024. Fundamental limitations of alignment in large language models. In Proceedings of the 41st International Conference on Machine Learning. PMLR 53079\u201353112."},{"key":"e_1_3_1_113_2","doi-asserted-by":"publisher","DOI":"10.1145\/3658644.3690322"},{"key":"e_1_3_1_114_2","first-page":"94","volume-title":"Proceedings of the 2025 IEEE Symposium on Security and Privacy (SP)","author":"Wu Junlin","year":"2024","unstructured":"Junlin Wu, Jiongxiao Wang, Chaowei Xiao, Chenguang Wang, Ning Zhang, and Yevgeniy Vorobeychik. 2024. Preference poisoning attacks on reward model learning. In Proceedings of the 2025 IEEE Symposium on Security and Privacy (SP). IEEE Computer Society, 94\u201394."},{"key":"e_1_3_1_115_2","doi-asserted-by":"crossref","unstructured":"Xiaodong Wu Ran Duan and Jianbing Ni. 2023. Unveiling security privacy and ethical concerns of chatgpt. Journal of Information and Intelligence 2 2 (2023) 102\u2013115.","DOI":"10.1016\/j.jiixd.2023.10.007"},{"key":"e_1_3_1_116_2","first-page":"7867","article-title":"A unified detection framework for inference-stage backdoor defenses","volume":"36","author":"Xian Xun","year":"2023","unstructured":"Xun Xian, Ganghua Wang, Jayanth Srinivasa, Ashish Kundu, Xuan Bi, Mingyi Hong, and Jie Ding. 2023. A unified detection framework for inference-stage backdoor defenses. Advances in Neural Information Processing Systems 36 (2023), 7867\u20137894.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_117_2","doi-asserted-by":"crossref","first-page":"14179","DOI":"10.18653\/v1\/2024.emnlp-main.785","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Xiao Yijia","year":"2024","unstructured":"Yijia Xiao, Yiqiao Jin, Yushi Bai, Yue Wu, Xianjun Yang, Xiao Luo, Wenchao Yu, Xujiang Zhao, Yanchi Liu, Quanquan Gu, et\u00a0al. 2024. Large language models can be contextual privacy protection learners. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 14179\u201314201."},{"key":"e_1_3_1_118_2","unstructured":"HanXiang Xu ShenAo Wang Ningke Li Yanjie Zhao Kai Chen Kailong Wang Yang Liu Ting Yu and HaoYu Wang. 2024. Large language models for cyber security: A systematic literature review. arXiv:2405.04760 (2024)."},{"key":"e_1_3_1_119_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3623326"},{"key":"e_1_3_1_120_2","doi-asserted-by":"crossref","unstructured":"Biwei Yan Kun Li Minghui Xu Yueyan Dong Yue Zhang Zhaochun Ren and Xiuzheng Cheng. 2024. On protecting the data privacy of large language models (LLMs): A survey. arXiv:2403.05156 (2024).","DOI":"10.1109\/ICMC60390.2024.00008"},{"key":"e_1_3_1_121_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.725"},{"key":"e_1_3_1_122_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-long.337"},{"key":"e_1_3_1_123_2","first-page":"1795","volume-title":"Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24)","author":"Yan Shenao","year":"2024","unstructured":"Shenao Yan, Shen Wang, Yue Duan, Hanbin Hong, Kiho Lee, Doowon Kim, and Yuan Hong. 2024. An {LLM-assisted}{easy-to-trigger} backdoor attack on code completion models: Injecting disguised vulnerabilities against strong detection. In Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24). 1795\u20131812."},{"key":"e_1_3_1_124_2","doi-asserted-by":"crossref","unstructured":"Haomiao Yang Kunlan Xiang Mengyu Ge Hongwei Li Rongxing Lu and Shui Yu. 2024. A comprehensive overview of backdoor attacks in large language models within communication networks. IEEE Network 38 6 (2024) 211\u2013218.","DOI":"10.1109\/MNET.2024.3367788"},{"key":"e_1_3_1_125_2","doi-asserted-by":"crossref","unstructured":"Mengmeng Yang Taolin Guo Tianqing Zhu Ivan Tjuawinata Jun Zhao and Kwok-Yan Lam. 2024. Local differential privacy and its applications: A comprehensive survey. Computer Standards and Interfaces 89 (2024) 103827.","DOI":"10.1016\/j.csi.2023.103827"},{"key":"e_1_3_1_126_2","unstructured":"Meng Yang Tianqing Zhu Chi Liu WanLei Zhou Shui Yu and Philip S. Yu. 2024. New emerged security and privacy of pre-trained model: A survey and outlook. arXiv:2411.07691 (2024)."},{"key":"e_1_3_1_127_2","first-page":"100938","article-title":"Watch out for your agents! investigating backdoor threats to llm-based agents","volume":"37","author":"Yang Wenkai","year":"2024","unstructured":"Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, and Xu Sun. 2024. Watch out for your agents! investigating backdoor threats to llm-based agents. Advances in Neural Information Processing Systems 37 (2024), 100938\u2013100964.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_128_2","first-page":"123","volume-title":"Proceedings of the 2024 IEEE Symposium on Security and Privacy (SP)","author":"Yang Yuchen","year":"2024","unstructured":"Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, and Yinzhi Cao. 2024. Sneakyprompt: Jailbreaking text-to-image generative models. In Proceedings of the 2024 IEEE Symposium on Security and Privacy (SP). IEEE Computer Society, 123\u2013123."},{"key":"e_1_3_1_129_2","first-page":"7745","volume-title":"Proceedings of the ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Yao Hongwei","year":"2024","unstructured":"Hongwei Yao, Jian Lou, and Zhan Qin. 2024. Poisonprompt: Backdoor attack on prompt-based large language models. In Proceedings of the ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 7745\u20137749."},{"key":"e_1_3_1_130_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.hcc.2024.100211"},{"key":"e_1_3_1_131_2","volume-title":"Proceedings of the Socially Responsible Language Modelling Research","author":"Yao Yuanshun","year":"2023","unstructured":"Yuanshun Yao, Xiaojun Xu, and Yang Liu. 2023. Large language model unlearning. In Proceedings of the Socially Responsible Language Modelling Research."},{"key":"e_1_3_1_132_2","unstructured":"Dayong Ye Tianqing Zhu Shang Wang Bo Liu Leo Yu Zhang Wanlei Zhou and Yang Zhang. 2025. Data-free model-related attacks: Unleashing the potential of generative AI. arXiv:2501.16671 (2025)."},{"key":"e_1_3_1_133_2","doi-asserted-by":"publisher","DOI":"10.1145\/3548606.3560675"},{"key":"e_1_3_1_134_2","first-page":"4657","volume-title":"Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24)","author":"Yu Jiahao","year":"2024","unstructured":"Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. 2024. {LLM-Fuzzer}: Scaling assessment of large language model jailbreaks. In Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24). 4657\u20134674."},{"key":"e_1_3_1_135_2","first-page":"40306","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Yu Weichen","year":"2023","unstructured":"Weichen Yu, Tianyu Pang, Qian Liu, Chao Du, Bingyi Kang, Yan Huang, Min Lin, and Shuicheng Yan. 2023. Bag of tricks for training data extraction from language models. In Proceedings of the International Conference on Machine Learning. PMLR, 40306\u201340320."},{"key":"e_1_3_1_136_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Yuan Youliang","year":"2024","unstructured":"Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2024. GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_1_137_2","volume-title":"Proceedings of the Network and Distributed System Security Symposium, NDSS 2025","author":"Zeng Rui","year":"2025","unstructured":"Rui Zeng, Xi Chen, Yuwen Pu, Xuhong Zhang, Tianyu Du, and Shouling Ji. 2025. CLIBE: Detecting dynamic backdoors in transformer-based NLP models. In Proceedings of the Network and Distributed System Security Symposium, NDSS 2025. The Internet Society."},{"key":"e_1_3_1_138_2","unstructured":"Yifan Zeng Yiran Wu Xiao Zhang Huazheng Wang and Qingyun Wu. 2024. Autodefense: Multi-agent llm defense against jailbreak attacks. arXiv:2403.04783 (2024)."},{"key":"e_1_3_1_139_2","first-page":"1813","volume-title":"Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24)","author":"Zhang Ruisi","year":"2024","unstructured":"Ruisi Zhang, Shehzeen Samarah Hussain, Paarth Neekhara, and Farinaz Koushanfar. 2024. {REMARK-LLM}: A robust and efficient watermarking framework for generative large language models. In Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24). 1813\u20131830."},{"key":"e_1_3_1_140_2","doi-asserted-by":"publisher","DOI":"10.5555\/3698900.3699004"},{"key":"e_1_3_1_141_2","doi-asserted-by":"publisher","DOI":"10.1109\/TDSC.2024.3372777"},{"key":"e_1_3_1_142_2","volume-title":"Proceedings of the 1st Conference on Language Modeling","author":"Zhang Yiming","year":"2024","unstructured":"Yiming Zhang, Nicholas Carlini, and Daphne Ippolito. 2024. Effective prompt extraction from language models. In Proceedings of the 1st Conference on Language Modeling."},{"key":"e_1_3_1_143_2","doi-asserted-by":"crossref","first-page":"12674","DOI":"10.18653\/v1\/2023.acl-long.709","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhang Zhexin","year":"2023","unstructured":"Zhexin Zhang, Jiaxin Wen, and Minlie Huang. 2023. ETHICIST: Targeted training data extraction through loss smoothed soft prompting and calibrated confidence estimation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 12674\u201312687."},{"key":"e_1_3_1_144_2","doi-asserted-by":"crossref","first-page":"12303","DOI":"10.18653\/v1\/2023.emnlp-main.757","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Zhao Shuai","year":"2023","unstructured":"Shuai Zhao, Jinming Wen, Anh Luu, Junbo Zhao, and Jie Fu. 2023. Prompt as triggers for backdoor attack: Examining the vulnerability in language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 12303\u201312317."},{"key":"e_1_3_1_145_2","unstructured":"Wayne Xin Zhao Kun Zhou Junyi Li Tianyi Tang Xiaolei Wang Yupeng Hou Yingqian Min Beichen Zhang Junjie Zhang Zican Dong et\u00a0al. 2023. A survey of large language models. arXiv:2303.18223 (2023)."},{"key":"e_1_3_1_146_2","doi-asserted-by":"crossref","unstructured":"Shuai Zhou Chi Liu Dayong Ye Tianqing Zhu Wanlei Zhou and Philip S. Yu. 2022. Adversarial attacks and defenses in deep learning: From a perspective of cybersecurity. ACM Computing Surveys 55 8 (2022) 1\u201339.","DOI":"10.1145\/3547330"},{"key":"e_1_3_1_147_2","unstructured":"Zhenhong Zhou Zherui Li Jie Zhang Yuanhe Zhang Kun Wang Yang Liu and Qing Guo. 2025. CORBA: Contagious recursive blocking attacks on multi-agent systems based on large language models. arXiv:2502.14529 (2025)."},{"key":"e_1_3_1_148_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2022.3233190"},{"key":"e_1_3_1_149_2","first-page":"19","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Zhu Yukun","year":"2015","unstructured":"Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE International Conference on Computer Vision. 19\u201327."},{"key":"e_1_3_1_150_2","unstructured":"Daniel M. Ziegler Nisan Stiennon Jeffrey Wu Tom B. Brown Alec Radford Dario Amodei Paul Christiano and Geoffrey Irving. 2019. Fine-tuning language models from human preferences. arXiv:1909.08593 (2019)."},{"key":"e_1_3_1_151_2","unstructured":"Andy Zou Zifan Wang J. Zico Kolter and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv:2307.15043 (2023)."},{"key":"e_1_3_1_152_2","unstructured":"Wei Zou Runpeng Geng Binghui Wang and Jinyuan Jia. 2024. PoisonedRAG: Knowledge poisoning attacks to retrieval-augmented generation of large language models. In Proceedings of the 34th USENIX Security Symposium (USENIX Security 25). 3827\u20133844."}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3764113","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,6]],"date-time":"2025-10-06T13:58:58Z","timestamp":1759759138000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3764113"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,6]]},"references-count":151,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,3,31]]}},"alternative-id":["10.1145\/3764113"],"URL":"https:\/\/doi.org\/10.1145\/3764113","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,6]]},"assertion":[{"value":"2024-06-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-12","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-06","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}