{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T20:16:29Z","timestamp":1785442589110,"version":"3.56.0"},"reference-count":145,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"name":"National Key Research and Development Program of China","award":["2023YFB3107400"],"award-info":[{"award-number":["2023YFB3107400"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62006181, 62132011, 62161160337, 62206217, U20A20177, U21B2018"],"award-info":[{"award-number":["62006181, 62132011, 62161160337, 62206217, U20A20177, U21B2018"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Shaanxi Province Key Industry Innovation Program","award":["2021ZDLGY01-02 and 2023-ZDLGY-38"],"award-info":[{"award-number":["2021ZDLGY01-02 and 2023-ZDLGY-38"]}]},{"name":"National Research Foundation, Singapore, the Cyber Security Agency under its National Cybersecurity R&D Programme","award":["NCRP25-P04-TAICeN"],"award-info":[{"award-number":["NCRP25-P04-TAICeN"]}]},{"name":"DSO National Laboratories under the AI Singapore Programme","award":["AISG2-GC-2023-008"],"award-info":[{"award-number":["AISG2-GC-2023-008"]}]},{"name":"National Research Foundation, Prime Minister\u2019s Office, Singapore under the Campus for Research Excellence and Technological Enterprise (CREATE) programme"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2026,1,31]]},"abstract":"<jats:p>\n                    The systems and software powered by Large Language Models (LLMs) and Multi-Modal Large Language Models (MLLMs) have played a critical role in numerous scenarios. However, current LLM systems are vulnerable to prompt-based attacks, with jailbreaking attacks enabling the LLM system to generate harmful content, while hijacking attacks manipulate the LLM system to perform attacker-desired tasks, underscoring the necessity for detection tools. Unfortunately, existing detecting approaches are usually tailored to specific attacks, resulting in poor generalization in detecting various attacks across different modalities. To address it, we propose\n                    <jats:sc>JailGuard<\/jats:sc>\n                    , a universal detection framework deployed on top of LLM systems for prompt-based attacks across text and image modalities.\n                    <jats:sc>JailGuard<\/jats:sc>\n                    operates on the principle that attacks are inherently less robust than benign ones. Specifically,\n                    <jats:sc>JailGuard<\/jats:sc>\n                    mutates untrusted inputs to generate variants and leverages the discrepancy of the variants\u2019 responses on the target model to distinguish attack samples from benign samples. We implement 18 mutators for text and image inputs and design a mutator combination policy to further improve detection generalization. The evaluation on the dataset containing 15 known attack types suggests that\n                    <jats:sc>JailGuard<\/jats:sc>\n                    achieves the best detection accuracy of 86.14%\/82.90% on text and image inputs, outperforming state-of-the-art methods by 11.81\u201325.73% and 12.20\u201321.40%.\n                  <\/jats:p>","DOI":"10.1145\/3724393","type":"journal-article","created":{"date-parts":[[2025,3,19]],"date-time":"2025-03-19T13:39:13Z","timestamp":1742391553000},"page":"1-40","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["<scp>JailGuard<\/scp>\n                    : A Universal Detection Framework for Prompt-based Attacks on LLM Systems"],"prefix":"10.1145","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7010-6749","authenticated-orcid":false,"given":"Xiaoyu","family":"Zhang","sequence":"first","affiliation":[{"name":"Xi\u2019an Jiaotong University, Xi\u2019an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5603-1322","authenticated-orcid":false,"given":"Cen","family":"Zhang","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2207-1622","authenticated-orcid":false,"given":"Tianlin","family":"Li","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5784-770X","authenticated-orcid":false,"given":"Yihao","family":"Huang","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2018-9344","authenticated-orcid":false,"given":"Xiaojun","family":"Jia","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5058-4660","authenticated-orcid":false,"given":"Ming","family":"Hu","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4230-1077","authenticated-orcid":false,"given":"Jie","family":"Zhang","sequence":"additional","affiliation":[{"name":"CFAR, A*STAR, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7300-9215","authenticated-orcid":false,"given":"Yang","family":"Liu","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1551-8948","authenticated-orcid":false,"given":"Shiqing","family":"Ma","sequence":"additional","affiliation":[{"name":"University of Massachusetts Amherst, Amherst, Massachusetts, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6959-0569","authenticated-orcid":false,"given":"Chao","family":"Shen","sequence":"additional","affiliation":[{"name":"Xi\u2019an Jiaotong University, Xi\u2019an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,12,11]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"2023. AuditNLG: Auditing Generative AI Language Modeling for Trustworthiness. Retrieved from https:\/\/github.com\/salesforce\/AuditNLG"},{"key":"e_1_3_3_3_2","unstructured":"Azure AI Content Safety. 2023. Retrieved from https:\/\/azure.microsoft.com\/en-us\/products\/ai-services\/ai-content-safety"},{"key":"e_1_3_3_4_2","unstructured":"ChatGPT Plugins. 2023. Retrieved from https:\/\/openai.com\/index\/chatgpt-plugins\/"},{"key":"e_1_3_3_5_2","unstructured":"DALLE-3 Masterclass: Everything You Didn t Know (Complete DALLE 3 Tutorial). 2023. Retrieved from https:\/\/midjourney.fm\/blog-DALLE3-Masterclass-Everything-You-Didnt-Know-Complete-DALLE-3-Tutorial-38611"},{"key":"e_1_3_3_6_2","unstructured":"GPT-4 System Card. 2023. Retrieved from https:\/\/cdn.openai.com\/papers\/gpt-4-system-card.pdf"},{"key":"e_1_3_3_7_2","unstructured":"GPT-4(v) System Card. 2023. Retrieved from https:\/\/cdn.openai.com\/papers\/GPTV_System_Card.pdf"},{"key":"e_1_3_3_8_2","unstructured":"Hands-on AI Demos for Human Resources. 2024. Retrieved from https:\/\/labs.hrflow.ai\/"},{"key":"e_1_3_3_9_2","unstructured":"spaCy: Industrial-Strength Natural Language Processing in Python. 2024. Retrieved from https:\/\/spacy.io\/"},{"key":"e_1_3_3_10_2","unstructured":"The Website of JailGuard. 2024. Retrieved from https:\/\/sites.google.com\/view\/jailguard"},{"key":"e_1_3_3_11_2","unstructured":"Wizard-Vicuna-13B-Uncensored. 2024. Retrieved from https:\/\/huggingface.co\/cognitivecomputations\/Wizard-Vicuna-13B-Uncensored"},{"key":"e_1_3_3_12_2","first-page":"79","volume-title":"Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security","author":"Abdelnabi Sahar","year":"2023","unstructured":"Sahar Abdelnabi, Kai Greshake, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023. Not what you\u2019ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, 79\u201390."},{"key":"e_1_3_3_13_2","volume-title":"International Conference on Learning Representations (ICLR)","author":"Abdullah Hadi","year":"2021","unstructured":"Hadi Abdullah, Aditya Karlekar, Vincent Bindschaedler, and Patrick Traynor. 2021. Demystifying limited adversarial transferability in automatic speech recognition systems. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_3_14_2","unstructured":"Gabriel Alon and Michael Kamfonas. 2023. Detecting language model attacks with perplexity. arXiv:2308.14132. Retrieved from https:\/\/arxiv.org\/abs\/2308.14132"},{"key":"e_1_3_3_15_2","unstructured":"Maksym Andriushchenko Francesco Croce and Nicolas Flammarion. 2024. Jailbreaking leading safety-aligned LLMs with simple adaptive attacks. arXiv:2404.02151. Retrieved from https:\/\/arxiv.org\/abs\/2404.02151"},{"key":"e_1_3_3_16_2","first-page":"484","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part XXIII","author":"Andriushchenko Maksym","year":"2020","unstructured":"Maksym Andriushchenko, Francesco Croce, Nicolas Flammarion, and Matthias Hein. 2020. Square attack: A query-efficient black-box adversarial attack via random search. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part XXIII. Springer, 484\u2013501."},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3583780.3615111"},{"key":"e_1_3_3_18_2","first-page":"274","volume-title":"International Conference on Machine Learning","author":"Athalye Anish","year":"2018","unstructured":"Anish Athalye, Nicholas Carlini, and David Wagner. 2018. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In International Conference on Machine Learning. PMLR, 274\u2013283."},{"key":"e_1_3_3_19_2","first-page":"16692","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Bai Yalong","year":"2022","unstructured":"Yalong Bai, Yifan Yang, Wei Zhang, and Tao Mei. 2022. Directional self-supervised learning for heavy image augmentations. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 16692\u201316701."},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3544558"},{"key":"e_1_3_3_21_2","unstructured":"Yuzhe Cai Shaoguang Mao Wenshan Wu Zehua Wang Yaobo Liang Tao Ge Chenfei Wu Wang You Ting Song Yan Xia et al. 2023. Low-code llm: Visual programming over llms. arXiv:2304.08103. Retrieved from https:\/\/arxiv.org\/abs\/2304.08103"},{"key":"e_1_3_3_22_2","unstructured":"Patrick Chao Alexander Robey Edgar Dobriban Hamed Hassani George J. Pappas and Eric Wong. 2023. Jailbreaking black box large language models in twenty queries. arXiv:2310.08419. Retrieved from https:\/\/arxiv.org\/abs\/2310.08419"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i16.29728"},{"key":"e_1_3_3_24_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman et al. 2021. Evaluating large language models trained on code. arXiv:2107.03374. Retrieved from https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_3_25_2","unstructured":"Sizhe Chen Julien Piet Chawin Sitawarin and David Wagner. 2024. StruQ: Defending against prompt injection with structured queries. arXiv:2402.06363. Retrieved from https:\/\/arxiv.org\/abs\/2402.06363"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324916000334"},{"key":"e_1_3_3_27_2","first-page":"1310","volume-title":"International Conference on Machine Learning","author":"Cohen Jeremy","year":"2019","unstructured":"Jeremy Cohen, Elan Rosenfeld, and Zico Kolter. 2019. Certified adversarial robustness via randomized smoothing. In International Conference on Machine Learning. PMLR, 1310\u20131320."},{"key":"e_1_3_3_28_2","first-page":"2206","volume-title":"International Conference on Machine Learning","author":"Croce Francesco","year":"2020","unstructured":"Francesco Croce and Matthias Hein. 2020. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In International Conference on Machine Learning. PMLR, 2206\u20132216."},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW50498.2020.00359"},{"key":"e_1_3_3_30_2","unstructured":"Justin Cui Wei-Lin Chiang Ion Stoica and Cho-Jui Hsieh. 2024. OR-bench: An over-refusal benchmark for large language models. arXiv:2405.20947. Retrieved from https:\/\/arxiv.org\/abs\/2405.20947"},{"key":"e_1_3_3_31_2","first-page":"28772","article-title":"Achieving rotational invariance with bessel-convolutional neural networks","volume":"34","author":"Delchevalerie Valentin","year":"2021","unstructured":"Valentin Delchevalerie, Adrien Bibal, Beno\u00eet Fr\u00e9nay, and Alexandre Mayer. 2021. Achieving rotational invariance with bessel-convolutional neural networks. Advances in Neural Information Processing Systems 34 (2021), 28772\u201328783.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_32_2","unstructured":"Gelei Deng Yi Liu Yuekang Li Kailong Wang Ying Zhang Zefeng Li Haoyu Wang Tianwei Zhang and Yang Liu. 2023. 2023. MasterKey: Automated jailbreak across multiple large language model chatbots. arXiv:2307.08715. Retrieved from https:\/\/arxiv.org\/abs\/2307.08715"},{"key":"e_1_3_3_33_2","volume-title":"The 12th International Conference on Learning Representations","author":"Deng Yue","year":"2023","unstructured":"Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing. 2023. Multilingual jailbreak challenges in large language models. In The 12th International Conference on Learning Representations."},{"key":"e_1_3_3_34_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from http:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_3_35_2","unstructured":"Gavin Weiguang Ding Yash Sharma Kry Yik Chau Lui and Ruitong Huang. 2018. Mma training: Direct input space margin maximization through adversarial training. arXiv:1812.02637. Retrieved from https:\/\/arxiv.org\/abs\/1812.02637"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01617"},{"key":"e_1_3_3_37_2","first-page":"30039","article-title":"Alpacafarm: A simulation framework for methods that learn from human feedback","volume":"36","author":"Dubois Yann","year":"2024","unstructured":"Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S. Liang, and Tatsunori B. Hashimoto. 2024. Alpacafarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing Systems 36 (2024), 30039\u201330069.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-90-481-8847-5_10"},{"key":"e_1_3_3_39_2","first-page":"8404","article-title":"Certified defense to image transformations via randomized smoothing","volume":"33","author":"Fischer Marc","year":"2020","unstructured":"Marc Fischer, Maximilian Baader, and Martin Vechev. 2020. Certified defense to image transformations via randomized smoothing. Advances in Neural Information Processing Systems 33 (2020), 8404\u20138417.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_40_2","first-page":"2150","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Gao Ruijun","year":"2022","unstructured":"Ruijun Gao, Qing Guo, Felix Juefei-Xu, Hongkai Yu, Huazhu Fu, Wei Feng, Yang Liu, and Song Wang. 2022. Can you spot the chameleon? Adversarially camouflaging images from co-salient object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2150\u20132159."},{"key":"e_1_3_3_41_2","unstructured":"Jonas Geiping Alex Stein Manli Shu Khalid Saifullah Yuxin Wen and Tom Goldstein. 2024. Coercing LLMs to do and reveal (almost) anything. arXiv:2402.14020. Retrieved from https:\/\/arxiv.org\/abs\/2402.14020"},{"key":"e_1_3_3_42_2","unstructured":"Spyros Gidaris Praveer Singh and Nikos Komodakis. 2018. Unsupervised representation learning by predicting image rotations. arXiv:1803.07728. Retrieved from https:\/\/arxiv.org\/abs\/1803.07728"},{"key":"e_1_3_3_43_2","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1007\/978-3-031-70239-6_26","article-title":"MaskPure: Improving the defense of text adversaries with stochastic purification","volume":"19","author":"Gietz Harrison","year":"2024","unstructured":"Harrison Gietz and Jugal Kalita. 2024. MaskPure: Improving the defense of text adversaries with stochastic purification. Natural Language Processing and Information Systems 19 (2024), 379\u2013393.","journal-title":"Natural Language Processing and Information Systems"},{"key":"e_1_3_3_44_2","unstructured":"Yunpeng Gong Liqing Huang and Lifei Chen. 2021. Eliminate deviation with deviation for data augmentation and a general multi-modal data learning method. arXiv:2101.08533. Retrieved from https:\/\/arxiv.org\/abs\/2101.08533"},{"key":"e_1_3_3_45_2","unstructured":"Ian J. Goodfellow Jonathon Shlens and Christian Szegedy. 2014. Explaining and harnessing adversarial examples. arXiv:1412.6572. Retrieved from https:\/\/arxiv.org\/abs\/1412.6572"},{"key":"e_1_3_3_46_2","unstructured":"Riley Goodside. 2022. Prompt Injection Attacks against GPT-3. Retrieved from https:\/\/simonwillison.net\/2022\/Sep\/12\/prompt-injection\/"},{"key":"e_1_3_3_47_2","unstructured":"Yunhao Gou Kai Chen Zhili Liu Lanqing Hong Hang Xu Zhenguo Li Dit-Yan Yeung James T. Kwok and Yu Zhang. 2024. Eyes closed safety on: Protecting multimodal llms via image-to-text transformation. arXiv:2403.09572. Retrieved from https:\/\/arxiv.org\/abs\/2403.09572"},{"key":"e_1_3_3_48_2","unstructured":"Chuan Guo Mayank Rana Moustapha Cisse and Laurens Van Der Maaten. 2017. Countering adversarial images using input transformations. arXiv:1711.00117. Retrieved from https:\/\/arxiv.org\/abs\/1711.00117"},{"key":"e_1_3_3_49_2","first-page":"975","article-title":"Watch out! Motion is blurring the vision of your deep neural networks","volume":"33","author":"Guo Qing","year":"2020","unstructured":"Qing Guo, Felix Juefei-Xu, Xiaofei Xie, Lei Ma, Jian Wang, Bing Yu, Wei Feng, and Yang Liu. 2020. Watch out! Motion is blurring the vision of your deep neural networks. Advances in Neural Information Processing Systems 33 (2020), 975\u2013985.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_50_2","first-page":"8465","volume-title":"International Conference on Machine Learning","author":"Hao Zhongkai","year":"2022","unstructured":"Zhongkai Hao, Chengyang Ying, Yinpeng Dong, Hang Su, Jian Song, and Jun Zhu. 2022. Gsmooth: Certified robustness against semantic transformations via generalized randomized smoothing. In International Conference on Machine Learning. PMLR, 8465\u20138483."},{"key":"e_1_3_3_51_2","unstructured":"Adam Hare Yu Chen Yinan Liu Zhenming Liu and Christopher G. Brinton. 2020. On extending NLP techniques from the categorical to the latent space: KL divergence Zipf\u2019s law and similarity search. arXiv:2012.01941. Retrieved from https:\/\/arxiv.org\/abs\/2012.01941"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3663529.3663849"},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_3_54_2","unstructured":"Dan Hendrycks Norman Mu Ekin D. Cubuk Barret Zoph Justin Gilmer and Balaji Lakshminarayanan. 2019. Augmix: A simple data processing method to improve robustness and uncertainty. arXiv:1912.02781. Retrieved from https:\/\/arxiv.org\/abs\/1912.02781"},{"key":"e_1_3_3_55_2","unstructured":"Chih-Hui Ho and Nuno Vasconcelos. 2022. DISCO: Adversarial defense with local implicit functions. arXiv:2212.05630. Retrieved from https:\/\/arxiv.org\/abs\/2212.05630"},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01181"},{"key":"e_1_3_3_57_2","first-page":"9","volume-title":"Proceedings of the Sixth New Zealand Computer Science Research Student Conference (NZCSRSC2008)","author":"Huang Anna","year":"2008","unstructured":"Anna Huang. 2008. Similarity measures for text document clustering. In Proceedings of the Sixth New Zealand Computer Science Research Student Conference (NZCSRSC2008), 9\u201356."},{"key":"e_1_3_3_58_2","unstructured":"Hai Huang Zhengyu Zhao Michael Backes Yun Shen and Yang Zhang. 2023. Composite backdoor attacks against large language models. arXiv:2310.07676. Retrieved from https:\/\/arxiv.org\/abs\/2310.07676"},{"key":"e_1_3_3_59_2","volume-title":"The 12th International Conference on Learning Representations","author":"Huang Yangsibo","year":"2023","unstructured":"Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen. 2023. Catastrophic jailbreak of open-source LLMs via exploiting generation. In The 12th International Conference on Learning Representations."},{"key":"e_1_3_3_60_2","unstructured":"Neel Jain Avi Schwarzschild Yuxin Wen Gowthami Somepalli John Kirchenbauer Ping-Yeh Chiang Micah Goldblum Aniruddha Saha Jonas Geiping and Tom Goldstein. 2023. Baseline defenses for adversarial attacks against aligned language models. arXiv:2309.00614. Retrieved from https:\/\/arxiv.org\/abs\/2309.00614"},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413976"},{"key":"e_1_3_3_62_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-emnlp.234"},{"key":"e_1_3_3_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3571884.3604313"},{"key":"e_1_3_3_64_2","doi-asserted-by":"crossref","first-page":"32","DOI":"10.25080\/Majora-14bd3278-006","volume-title":"Scipy","author":"Komer Brent","year":"2014","unstructured":"Brent Komer, James Bergstra, and Chris Eliasmith. 2014. Hyperopt-Sklearn: Automatic hyperparameter configuration for Scikit-Learn. In Scipy, 32\u201337."},{"key":"e_1_3_3_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3660853.3660878"},{"key":"e_1_3_3_66_2","unstructured":"Aounon Kumar Chirag Agarwal Suraj Srinivas Soheil Feizi and Hima Lakkaraju.2023. 2023. Certifying llm safety against adversarial prompting. arXiv:2309.02705. Retrieved from https:\/\/arxiv.org\/abs\/2309.02705"},{"issue":"6","key":"e_1_3_3_67_2","doi-asserted-by":"crossref","first-page":"2638","DOI":"10.1109\/TAI.2023.3323918","article-title":"Kullback-Leibler divergence based regularized normalization for low resource tasks","volume":"5","author":"Kumar Neeraj","year":"2023","unstructured":"Neeraj Kumar, Ankur Narang, and Brejesh Lall. 2023. Kullback-Leibler divergence based regularized normalization for low resource tasks. IEEE Transactions on Artificial Intelligence 5, 6 (2023), 2638\u20132650.","journal-title":"IEEE Transactions on Artificial Intelligence"},{"key":"e_1_3_3_68_2","doi-asserted-by":"publisher","DOI":"10.1615\/JMachLearnModelComput.2023049518"},{"key":"e_1_3_3_69_2","doi-asserted-by":"publisher","DOI":"10.1201\/9781351251389-8"},{"key":"e_1_3_3_70_2","unstructured":"Xuechen Li Tianyi Zhang Yann Dubois Rohan Taori Ishaan Gulrajani Carlos Guestrin Percy Liang and Tatsunori B. Hashimoto. 2023. AlpacaEval: An Automatic Evaluator of Instruction-following Models. Retrieved from https:\/\/github.com\/tatsu-lab\/alpaca_eval"},{"key":"e_1_3_3_71_2","unstructured":"Xuan Li Zhanke Zhou Jianing Zhu Jiangchao Yao Tongliang Liu and Bo Han. 2023. Deepinception: Hypnotize large language model to be jailbreaker. arXiv:2311.03191. Retrieved from https:\/\/arxiv.org\/abs\/2311.03191"},{"key":"e_1_3_3_72_2","first-page":"34892","article-title":"Visual instruction tuning","volume":"36","author":"Liu Haotian","year":"2024","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. Advances in Neural Information Processing Systems 36 (2024), 34892\u201334916.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_73_2","unstructured":"Xiaogeng Liu Zhiyuan Yu Yizhe Zhang Ning Zhang and Chaowei Xiao. 2024. Automatic and universal prompt injection attacks against large language models. arXiv:2403.04957. Retrieved from https:\/\/arxiv.org\/abs\/2403.04957"},{"key":"e_1_3_3_74_2","unstructured":"Xin Liu Yichen Zhu Jindong Gu Yunshi Lan Chao Yang and Yu Qiao.2023. MM-SafetyBench: A benchmark for safety evaluation of multimodal large language models. arXiv:2311.17600. Retrieved from https:\/\/arxiv.org\/abs\/2311.17600"},{"key":"e_1_3_3_75_2","unstructured":"Yi Liu Gelei Deng Yuekang Li Kailong Wang Tianwei Zhang Yepang Liu Haoyu Wang Yan Zheng and Yang Liu. 2023. Prompt injection attack against LLM-integrated applications. arXiv:2306.05499. Retrieved from https:\/\/arxiv.org\/abs\/2306.05499"},{"key":"e_1_3_3_76_2","unstructured":"Yi Liu Gelei Deng Zhengzi Xu Yuekang Li Yaowen Zheng Ying Zhang Lida Zhao Tianwei Zhang and Yang Liu. 2023. Jailbreaking chatgpt via prompt engineering: An empirical study. arXiv:2305.13860. Retrieved from https:\/\/arxiv.org\/abs\/2305.13860"},{"key":"e_1_3_3_77_2","doi-asserted-by":"publisher","DOI":"10.1145\/3663530.3665021"},{"key":"e_1_3_3_78_2","first-page":"1831","volume-title":"USENIX Security Symposium","author":"Liu Yupei","year":"2024","unstructured":"Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, and Neil Zhenqiang Gong. 2024. Formalizing and benchmarking prompt injection attacks and defenses. In USENIX Security Symposium, 1831\u20131847."},{"key":"e_1_3_3_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01458"},{"key":"e_1_3_3_80_2","unstructured":"Raphael Gontijo Lopes Dong Yin Ben Poole Justin Gilmer and Ekin D. Cubuk. 2019. Improving robustness without sacrificing accuracy with patch gaussian augmentation. arXiv:1906.02611. Retrieved from https:\/\/arxiv.org\/abs\/1906.02611"},{"key":"e_1_3_3_81_2","unstructured":"Renze Lou Kai Zhang and Wenpeng Yin. 2023. Is prompt all you need? No. a comprehensive and broader view of instruction learning. arXiv:2303.10475. Retrieved from https:\/\/arxiv.org\/abs\/2303.10475"},{"key":"e_1_3_3_82_2","unstructured":"Aleksander Madry Aleksandar Makelov Ludwig Schmidt Dimitris Tsipras and Adrian Vladu. 2017. Towards deep learning models resistant to adversarial attacks. arXiv:1706.06083. Retrieved from https:\/\/arxiv.org\/abs\/1706.06083"},{"key":"e_1_3_3_83_2","doi-asserted-by":"publisher","DOI":"10.1145\/3487043"},{"key":"e_1_3_3_84_2","doi-asserted-by":"publisher","DOI":"10.11613\/BM.2012.031"},{"key":"e_1_3_3_85_2","doi-asserted-by":"crossref","first-page":"11576","DOI":"10.1109\/ICRA48891.2023.10160396","volume-title":"2023 IEEE International Conference on Robotics and Automation (ICRA)","author":"Mees Oier","year":"2023","unstructured":"Oier Mees, Jessica Borja-Diaz, and Wolfram Burgard. 2023. Grounding language with visual affordances over unstructured data. In 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 11576\u201311582."},{"key":"e_1_3_3_86_2","unstructured":"Anay Mehrotra Manolis Zampetakis Paul Kassianik Blaine Nelson Hyrum Anderson Yaron Singer and Amin Karbasi. 2023. Tree of attacks: Jailbreaking black-box llms automatically. arXiv:2312.02119. Retrieved from https:\/\/arxiv.org\/abs\/2312.02119"},{"key":"e_1_3_3_87_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2022.3233726"},{"key":"e_1_3_3_88_2","unstructured":"Meta. 2024. Meet Your New Assistant: Meta AI Built With Llama 3. Retrieved from https:\/\/about.fb.com\/news\/2024\/04\/meta-ai-assistant-built-with-llama-3"},{"key":"e_1_3_3_89_2","unstructured":"Jan Hendrik Metzen Tim Genewein Volker Fischer and Bastian Bischoff. 2017. On detecting adversarial perturbations. arXiv:1702.04267. Retrieved from https:\/\/arxiv.org\/abs\/1702.04267"},{"key":"e_1_3_3_90_2","unstructured":"Microsoft. 2024. Microsoft Copilot. Retrieved from https:\/\/www.microsoft.com\/en-us\/bing"},{"key":"e_1_3_3_91_2","doi-asserted-by":"publisher","DOI":"10.1145\/219717.219748"},{"key":"e_1_3_3_92_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.array.2022.100258"},{"key":"e_1_3_3_93_2","unstructured":"Weili Nie Brandon Guo Yujia Huang Chaowei Xiao Arash Vahdat and Anima Anandkumar. 2022. Diffusion models for adversarial purification. arXiv:2205.07460. Retrieved from https:\/\/arxiv.org\/abs\/2205.07460"},{"key":"e_1_3_3_94_2","unstructured":"Zhenxing Niu Haodong Ren Xinbo Gao Gang Hua and Rong Jin. 2024. Jailbreaking attack against multimodal large language model. arXiv:2402.02309. Retrieved from https:\/\/arxiv.org\/abs\/2402.02309"},{"key":"e_1_3_3_95_2","unstructured":"David A. Noever and Samantha E. Miller Noever. 2021. Reading isn\u2019t believing: Adversarial attacks on multi-modal neurons. arXiv:2103.10480. Retrieved from https:\/\/arxiv.org\/abs\/2103.10480"},{"key":"e_1_3_3_96_2","unstructured":"OpenAI. 2024. Hello GPT-4o. Retrieved from https:\/\/openai.com\/index\/hello-gpt-4o\/"},{"key":"e_1_3_3_97_2","unstructured":"Chris Parnin Gustavo Soares Rahul Pandita Sumit Gulwani Jessica Rich and Austin Z. Henley. 2023. Building your own product copilot: Challenges opportunities and needs. arXiv:2312.14231. Retrieved from https:\/\/arxiv.org\/abs\/2312.14231"},{"key":"e_1_3_3_98_2","volume-title":"NeurIPS ML Safety Workshop","author":"Perez F\u00e1bio","year":"2022","unstructured":"F\u00e1bio Perez and Ian Ribeiro. 2022. Ignore previous prompt: Attack techniques for language models. In NeurIPS ML Safety Workshop."},{"key":"e_1_3_3_99_2","doi-asserted-by":"publisher","DOI":"10.1016\/J.JSS.2020.110657"},{"key":"e_1_3_3_100_2","unstructured":"Renjie Pi Tianyang Han Yueqi Xie Rui Pan Qing Lian Hanze Dong Jipeng Zhang and Tong Zhang. 2024. MLLM-potector: Ensuring MLLM\u2019s safety without hurting performance. arXiv:2401.02906. Retrieved from https:\/\/arxiv.org\/abs\/2401.02906"},{"key":"e_1_3_3_101_2","doi-asserted-by":"crossref","first-page":"521","DOI":"10.1007\/978-94-024-0881-2_20","volume-title":"Handbook of Linguistic Annotation","author":"Pradhan Sameer","year":"2017","unstructured":"Sameer Pradhan and Lance Ramshaw. 2017. Ontonotes: Large scale multi-layer, multi-lingual, distributed annotation. In Handbook of Linguistic Annotation. Springer, 521\u2013554."},{"key":"e_1_3_3_102_2","unstructured":"The Associated Press. 2025. Man who exploded Cybertruck in Las Vegas used ChatGPT in planning police say. Retrieved from https:\/\/www.npr.org\/2025\/01\/07\/nx-s1-5251611\/cybertruck-explosion-las-vegas-chatgpt-ai"},{"key":"e_1_3_3_103_2","first-page":"21527","volume-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence","author":"Qi Xiangyu","year":"2024","unstructured":"Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. 2024. Visual adversarial examples jailbreak aligned large language models. In Proceedings of the 38th AAAI Conference on Artificial Intelligence, 21527\u201321536."},{"key":"e_1_3_3_104_2","unstructured":"Xiangyu Qi Yi Zeng Tinghao Xie Pin-Yu Chen Ruoxi Jia Prateek Mittal and Peter Henderson. 2023. Fine-tuning aligned language models compromises safety even when users do not intend to! arXiv:2310.03693. Retrieved from https:\/\/arxiv.org\/abs\/2310.03693"},{"key":"e_1_3_3_105_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN55064.2022.9892117"},{"key":"e_1_3_3_106_2","volume-title":"ICML 2021 Workshop on Adversarial Machine Learning","author":"Rade Rahul","year":"2021","unstructured":"Rahul Rade and Seyed-Mohsen Moosavi-Dezfooli. 2021. Helper-based adversarial training: Reducing excessive margin to achieve a better accuracy vs. robustness trade-off. In ICML 2021 Workshop on Adversarial Machine Learning."},{"issue":"8","key":"e_1_3_3_107_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_3_108_2","unstructured":"Alexander Robey Eric Wong Hamed Hassani and George J. Pappas. 2023. Smoothllm: Defending large language models against jailbreaking attacks. arXiv:2310.03684. Retrieved from https:\/\/arxiv.org\/abs\/2310.03684"},{"key":"e_1_3_3_109_2","first-page":"1","volume-title":"2020 International Conference on Emerging Trends in Smart Technologies (ICETST)","author":"Saeed Sumaira","year":"2020","unstructured":"Sumaira Saeed, Sajjad Haider, and Quratulain Rajput. 2020. On finding similar verses from the Holy Quran using word embeddings. In 2020 International Conference on Emerging Trends in Smart Technologies (ICETST). IEEE, 1\u20136."},{"key":"e_1_3_3_110_2","unstructured":"Vikash Sehwag Saeed Mahloujifar Tinashe Handina Sihui Dai Chong Xiang Mung Chiang and Prateek Mittal. 2021. Robust learning meets generative models: Can proxy distributions improve adversarial robustness? arXiv:2104.09425. Retrieved from https:\/\/arxiv.org\/abs\/2104.09425"},{"key":"e_1_3_3_111_2","unstructured":"Zeyang Sha and Yang Zhang. 2024. Prompt stealing attacks against large language models. arXiv:2402.12959. Retrieved from https:\/\/arxiv.org\/abs\/2402.12959"},{"key":"e_1_3_3_112_2","unstructured":"Guangyu Shen Siyuan Cheng Kaiyuan Zhang Guanhong Tao Shengwei An Lu Yan Zhuo Zhang Shiqing Ma and Xiangyu Zhang. 2024. Rapid optimization for jailbreaking LLMs via subconscious exploitation and echopraxia. arXiv:2402.05467. Retrieved from https:\/\/arxiv.org\/abs\/2402.05467"},{"key":"e_1_3_3_113_2","first-page":"1671","volume-title":"ACM SIGSAC Conference on Computer and Communications Security (CCS)","author":"Shen Xinyue","year":"2024","unstructured":"Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2024. \u201cDo anything now\u201d: Characterizing and evaluating in-the-wild jailbreak prompts on large language models. In ACM SIGSAC Conference on Computer and Communications Security (CCS), 1671\u20131685."},{"key":"e_1_3_3_114_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01171"},{"key":"e_1_3_3_115_2","doi-asserted-by":"crossref","unstructured":"Binyu Tian Felix Juefei-Xu Qing Guo Xiaofei Xie Xiaohong Li and Yang Liu. 2021. AVA: Adversarial vignetting attack against visual recognition. arXiv:2105.05558. Retrieved from https:\/\/arxiv.org\/abs\/2105.05558","DOI":"10.24963\/ijcai.2021\/145"},{"key":"e_1_3_3_116_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_3_117_2","doi-asserted-by":"publisher","DOI":"10.1109\/FUZZY.2010.5584447"},{"key":"e_1_3_3_118_2","first-page":"15097","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Wang Xueping","year":"2021","unstructured":"Xueping Wang, Shasha Li, Min Liu, Yaonan Wang, and Amit K. Roy-Chowdhury. 2021. Multi-expert adversarial attack detection in person re-identification using context inconsistency. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 15097\u201315107."},{"key":"e_1_3_3_119_2","first-page":"2877","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics","author":"Wang Yijun","year":"2021","unstructured":"Yijun Wang, Changzhi Sun, Yuanbin Wu, Hao Zhou, Lei Li, and Junchi Yan. 2021. ENPAR: Enhancing entity and entity pair representations for joint entity relation extraction. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics, 2877\u20132887."},{"key":"e_1_3_3_120_2","unstructured":"Irene Weber. 2024. Large language models as software components: A taxonomy for LLM-integrated applications. arXiv:2406.10300. Retrieved from https:\/\/arxiv.org\/abs\/2406.10300"},{"key":"e_1_3_3_121_2","first-page":"80079","article-title":"Jailbroken: How does llm safety training fail","volume":"36","author":"Wei Alexander","year":"2024","unstructured":"Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2024. Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems 36 (2024), 80079\u201380110.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_122_2","unstructured":"Zeming Wei Yifei Wang and Yisen Wang. 2023. Jailbreak and guard aligned language models with only few in-context demonstrations. arXiv:2310.06387. Retrieved from https:\/\/arxiv.org\/abs\/2310.06387"},{"key":"e_1_3_3_123_2","unstructured":"Simon Willison. 2023. Delimiters Won\u2019t Save You from Prompt Injection. Retrieved from https:\/\/simonwillison.net\/2023\/May\/11\/delimiters-wont-save-you\/"},{"key":"e_1_3_3_124_2","unstructured":"Davey Winder. 2023. Hacker Reveals Microsoft s New AI-Powered Bing Chat Search Secrets. Retrieved from https:\/\/www.forbes.com\/sites\/daveywinder\/2023\/02\/13\/hacker-reveals-microsofts-new-ai-powered-bing-chat-search-secrets\/"},{"key":"e_1_3_3_125_2","unstructured":"Daoyuan Wu Shuai Wang Yang Liu and Ning Liu. 2024. LLMs can defend themselves against jailbreaking in a practical manner: A vision paper. arXiv:2402.15727. Retrieved from https:\/\/arxiv.org\/abs\/2402.15727"},{"key":"e_1_3_3_126_2","first-page":"247","volume-title":"32nd USENIX Security Symposium (USENIX Security \u201923)","author":"Wu Xinghui","year":"2023","unstructured":"Xinghui Wu, Shiqing Ma, Chao Shen, Chenhao Lin, Qian Wang, Qi Li, and Yuan Rao. 2023. \\(\\{\\) KENKU \\(\\}\\) Towards efficient and stealthy black-box adversarial attacks against. \\(\\{\\) ASR \\(\\}\\) Systems. In 32nd USENIX Security Symposium (USENIX Security \u201923), 247\u2013264."},{"key":"e_1_3_3_127_2","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-023-00765-8"},{"key":"e_1_3_3_128_2","unstructured":"Jiashu Xu Mingyu Derek Ma Fei Wang Chaowei Xiao and Muhao Chen. 2023. Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models. arXiv:2305.14710. Retrieved from https:\/\/arxiv.org\/abs\/2305.14710"},{"key":"e_1_3_3_129_2","unstructured":"Weilin Xu David Evans and Yanjun Qi. 2017. Feature squeezing: Detecting adversarial examples in deep neural networks. arXiv:1704.01155. Retrieved from https:\/\/arxiv.org\/abs\/1704.01155"},{"key":"e_1_3_3_130_2","unstructured":"Zihao Xu Yi Liu Gelei Deng Yuekang Li and Stjepan Picek. 2024. LLM Jailbreak attack versus defense techniques\u2014A comprehensive study. arXiv:2402.13457. Retrieved from https:\/\/arxiv.org\/abs\/2402.13457"},{"key":"e_1_3_3_131_2","unstructured":"Jun Yan Vikas Yadav Shiyang Li Lichang Chen Zheng Tang Hai Wang Vijay Srinivasan Xiang Ren and Hongxia Jin. 2023. Backdooring instruction-tuned large language models with virtual prompt injection. arXiv:2307.16888. Retrieved from https:\/\/arxiv.org\/abs\/2307.16888"},{"key":"e_1_3_3_132_2","first-page":"3465","volume-title":"Annual Meeting of the Association for Computational Linguistics (ACL)","author":"Ye Mao","year":"2020","unstructured":"Mao Ye, Chengyue Gong, and Qiang Liu. 2020. SAFER: A structure-free approach for certified robustness to adversarial word substitutions. In Annual Meeting of the Association for Computational Linguistics (ACL), 3465\u20133475."},{"key":"e_1_3_3_133_2","unstructured":"Jingwei Yi Yueqi Xie Bin Zhu Keegan Hines Emre Kiciman Guangzhong Sun Xing Xie and Fangzhao Wu. 2023. Benchmarking and defending against indirect prompt injection attacks on large language models. arXiv:2312.14197. Retrieved from https:\/\/arxiv.org\/abs\/2312.14197"},{"key":"e_1_3_3_134_2","first-page":"895","article-title":"On the dimensionality of word embedding","volume":"31","author":"Yin Zi","year":"2018","unstructured":"Zi Yin and Yuanyuan Shen. 2018. On the dimensionality of word embedding. Advances in Neural Information Processing Systems 31 (2018), 895\u2013906.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_135_2","unstructured":"Jiahao Yu Xingwei Lin and Xinyu Xing. 2023. Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts. arXiv:2309.10253. Retrieved from https:\/\/arxiv.org\/abs\/2309.10253"},{"key":"e_1_3_3_136_2","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00476"},{"key":"e_1_3_3_137_2","doi-asserted-by":"crossref","unstructured":"Chong Zhang Mingyu Jin Qinkai Yu Chengzhi Liu Haochen Xue and Xiaobo Jin. 2024. Goal-guided generative prompt injection attack on large language models. arXiv:2404.07234. Retrieved from https:\/\/arxiv.org\/abs\/2404.07234","DOI":"10.1109\/ICDM59182.2024.00119"},{"key":"e_1_3_3_138_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13042-010-0001-0"},{"key":"e_1_3_3_139_2","first-page":"26958","volume-title":"International Conference on Machine Learning","author":"Zhao Haiteng","year":"2022","unstructured":"Haiteng Zhao, Chang Ma, Xinshuai Dong, Anh Tuan Luu, Zhi-Hong Deng, and Hanwang Zhang. 2022. Certified robustness against natural language attacks by causal intervention. In International Conference on Machine Learning. PMLR, 26958\u201326970."},{"key":"e_1_3_3_140_2","unstructured":"Lianmin Zheng Wei-Lin Chiang Ying Sheng Siyuan Zhuang Zhanghao Wu Yonghao Zhuang Zi Lin Zhuohan Li Dacheng Li Eric P. Xing et al. 2023. Judging LLM-as-a-judge with MT-Bench and Chatbot arena. arXiv:2306.05685. Retrieved from https:\/\/arxiv.org\/abs\/2306.05685"},{"key":"e_1_3_3_141_2","unstructured":"Andy Zhou Bo Li and Haohan Wang. 2024. Robust prompt optimization for defending language models against jailbreaking attacks. arXiv:2401.17263. Retrieved from https:\/\/arxiv.org\/abs\/2401.17263"},{"key":"e_1_3_3_142_2","unstructured":"Jeffrey Zhou Tianjian Lu Swaroop Mishra Siddhartha Brahma Sujoy Basu Yi Luan Denny Zhou and Le Hou. 2023. Instruction-following evaluation for large language models. arXiv:2311.07911. Retrieved from https:\/\/arxiv.org\/abs\/2311.07911"},{"key":"e_1_3_3_143_2","unstructured":"Deyao Zhu Jun Chen Xiaoqian Shen Xiang Li and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv:2304.10592. Retrieved from https:\/\/arxiv.org\/abs\/2304.10592"},{"key":"e_1_3_3_144_2","unstructured":"Sicheng Zhu Ruiyi Zhang Bang An Gang Wu Joe Barrow Zichao Wang Furong Huang Ani Nenkova and Tong Sun. 2023. AutoDAN: Automatic and interpretable adversarial attacks on large language models. arXiv:2310.15140. Retrieved from https:\/\/arxiv.org\/abs\/2310.15140"},{"key":"e_1_3_3_145_2","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331232"},{"key":"e_1_3_3_146_2","unstructured":"Andy Zou Zifan Wang J. Zico Kolter and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv:2307.15043. Retrieved from https:\/\/arxiv.org\/abs\/2307.15043"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3724393","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,11]],"date-time":"2025-12-11T15:57:36Z","timestamp":1765468656000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3724393"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,11]]},"references-count":145,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1,31]]}},"alternative-id":["10.1145\/3724393"],"URL":"https:\/\/doi.org\/10.1145\/3724393","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,11]]},"assertion":[{"value":"2024-05-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}