{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T16:30:51Z","timestamp":1783009851049,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":72,"publisher":"ACM","funder":[{"name":"NSF","award":["2229876"],"award-info":[{"award-number":["2229876"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,10,13]]},"DOI":"10.1145\/3733799.3762981","type":"proceedings-article","created":{"date-parts":[[2025,12,30]],"date-time":"2025-12-30T11:38:49Z","timestamp":1767094729000},"page":"230-241","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-7447-0910","authenticated-orcid":false,"given":"Julien","family":"Piet","sequence":"first","affiliation":[{"name":"University of California, Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-5009-8424","authenticated-orcid":false,"given":"Xiao","family":"Huang","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-9086-8684","authenticated-orcid":false,"given":"Dennis","family":"Jacob","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8494-189X","authenticated-orcid":false,"given":"Annabella","family":"Chow","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-5360-1624","authenticated-orcid":false,"given":"Maha","family":"Alrashed","sequence":"additional","affiliation":[{"name":"KACST, Riyadh, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0750-356X","authenticated-orcid":false,"given":"Geng","family":"Zhao","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3746-1447","authenticated-orcid":false,"given":"Zhanhao","family":"Hu","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4949-9661","authenticated-orcid":false,"given":"Chawin","family":"Sitawarin","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0494-2586","authenticated-orcid":false,"given":"Basel","family":"Alomair","sequence":"additional","affiliation":[{"name":"KACST, Riyadh, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9944-9232","authenticated-orcid":false,"given":"David","family":"Wagner","sequence":"additional","affiliation":[{"name":"University of California, Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,12,30]]},"reference":[{"key":"e_1_3_3_2_2_2","unstructured":"Alex Albert. 2023. JailbreakChat. http:\/\/web.archive.org\/web\/20230220011306\/https:\/\/www.jailbreakchat.com\/ Accessed through Internet Archive Wayback Machine archived on February 20 2023."},{"key":"e_1_3_3_2_3_2","unstructured":"Anthropic. 2023. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2204.05862 (2023). https:\/\/arxiv.org\/abs\/2204.05862"},{"key":"e_1_3_3_2_4_2","doi-asserted-by":"publisher","unstructured":"Yoshua Bengio Geoffrey Hinton Andrew Yao Dawn Song Pieter Abbeel Trevor Darrell Yuval\u00a0Noah Harari Ya-Qin Zhang Lan Xue Shai Shalev-Shwartz Gillian Hadfield Jeff Clune Tegan Maharaj Frank Hutter At\u0131l\u0131m\u00a0G\u00fcne\u015f Baydin Sheila McIlraith Qiqi Gao Ashwin Acharya David Krueger Anca Dragan Philip Torr Stuart Russell Daniel Kahneman Jan Brauner and S\u00f6ren Mindermann. 2024. Managing Extreme AI Risks amid Rapid Progress. Science 384 6698 (2024) 842\u2013845. 10.1126\/science.adn0117 arXiv:https:\/\/www.science.org\/doi\/pdf\/10.1126\/science.adn0117","DOI":"10.1126\/science.adn0117"},{"key":"e_1_3_3_2_5_2","unstructured":"Patrick Chao Alexander Robey Edgar Dobriban Hamed Hassani George\u00a0J. Pappas and Eric Wong. 2023. Jailbreaking Black Box Large Language Models in Twenty Queries. arxiv:https:\/\/arXiv.org\/abs\/2310.08419\u00a0[cs] http:\/\/arxiv.org\/abs\/2310.08419"},{"key":"e_1_3_3_2_6_2","unstructured":"Sizhe Chen Julien Piet Chawin Sitawarin and David Wagner. 2024. Struq: Defending against prompt injection with structured queries. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2402.06363 (2024)."},{"key":"e_1_3_3_2_7_2","unstructured":"Sizhe Chen Arman Zharmagambetov Saeed Mahloujifar Kamalika Chaudhuri and Chuan Guo. 2024. Aligning llms to be robust against prompt injection. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.05451 (2024)."},{"key":"e_1_3_3_2_8_2","first-page":"1127","volume-title":"32nd USENIX Security Symposium (USENIX Security 23)","author":"Chen Yizheng","year":"2023","unstructured":"Yizheng Chen, Zhoujie Ding, and David Wagner. 2023. Continuous Learning for Android Malware Detection. In 32nd USENIX Security Symposium (USENIX Security 23). USENIX Association, Anaheim, CA, 1127\u20131144. https:\/\/www.usenix.org\/conference\/usenixsecurity23\/presentation\/chen-yizheng"},{"key":"e_1_3_3_2_9_2","unstructured":"Junjie Chu Yugeng Liu Ziqing Yang Xinyue Shen Michael Backes and Yang Zhang. 2024. Comprehensive Assessment of Jailbreak Attacks Against LLMs. arxiv:https:\/\/arXiv.org\/abs\/2402.05668\u00a0[cs.CR] https:\/\/arxiv.org\/abs\/2402.05668"},{"key":"e_1_3_3_2_10_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Deng Yue","year":"2024","unstructured":"Yue Deng, Wenxuan Zhang, Sinno\u00a0Jialin Pan, and Lidong Bing. 2024. Multilingual Jailbreak Challenges in Large Language Models. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=vESNKdEMGp"},{"key":"e_1_3_3_2_11_2","unstructured":"Entire_Comparison783. 2023. DAN Prompt. www.reddit.com\/r\/ChatGPT\/comments\/10x1nux\/dan_prompt\/"},{"key":"e_1_3_3_2_12_2","unstructured":"Google DeepMind. 2024. Gemini 2.0 Flash."},{"key":"e_1_3_3_2_13_2","unstructured":"Xingang Guo Fangxu Yu Huan Zhang Lianhui Qin and Bin Hu. 2024. COLD-attack: Jailbreaking LLMs with Stealthiness and Controllability. arxiv:https:\/\/arXiv.org\/abs\/2402.08679\u00a0[cs] http:\/\/arxiv.org\/abs\/2402.08679"},{"key":"e_1_3_3_2_14_2","unstructured":"Xiaomeng Hu Pin-Yu Chen and Tsung-Yi Ho. 2024. Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes. arxiv:https:\/\/arXiv.org\/abs\/2403.00867\u00a0[cs] http:\/\/arxiv.org\/abs\/2403.00867"},{"key":"e_1_3_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-0557"},{"key":"e_1_3_3_2_16_2","volume-title":"Neurips Safe Generative AI Workshop 2024","author":"Huang David","year":"2024","unstructured":"David Huang, Avidan Shah, Alexandre Araujo, David Wagner, and Chawin Sitawarin. 2024. Stronger Universal and Transfer Attacks by Suppressing Refusals. In Neurips Safe Generative AI Workshop 2024. https:\/\/openreview.net\/forum?id=eIBWRAbhND"},{"key":"e_1_3_3_2_17_2","unstructured":"Hakan Inan Kartikeya Upasani Jianfeng Chi Rashi Rungta Krithika Iyer Yuning Mao Michael Tontchev Qing Hu Brian Fuller Davide Testuggine et\u00a0al. 2023. Llama guard: LLM-based input-output safeguard for human-AI conversations. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2312.06674 (2023)."},{"key":"e_1_3_3_2_18_2","unstructured":"Internet Archive. 1996. Wayback Machine. https:\/\/web.archive.org\/. Accessed: 2025-04-13."},{"key":"e_1_3_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2307.04657"},{"key":"e_1_3_3_2_20_2","unstructured":"Albert\u00a0Q. Jiang Alexandre Sablayrolles Arthur Mensch Chris Bamford Devendra\u00a0Singh Chaplot Diego de\u00a0las Casas Florian Bressand Gianna Lengyel Guillaume Lample Lucile Saulnier L\u00e9lio\u00a0Renard Lavaud Marie-Anne Lachaux Pierre Stock Teven\u00a0Le Scao Thibaut Lavril Thomas Wang Timoth\u00e9e Lacroix and William\u00a0El Sayed. 2023. Mistral 7B. arxiv:https:\/\/arXiv.org\/abs\/2310.06825\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2310.06825"},{"key":"e_1_3_3_2_21_2","unstructured":"Liwei Jiang Kavel Rao Seungju Han Allyson Ettinger Faeze Brahman Sachin Kumar Niloofar Mireshghallah Ximing Lu Maarten Sap Yejin Choi and Nouha Dziri. 2024. WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models. arxiv:https:\/\/arXiv.org\/abs\/2406.18510\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2406.18510"},{"key":"e_1_3_3_2_22_2","unstructured":"Erik Jones Anca Dragan Aditi Raghunathan and Jacob Steinhardt. 2023. Automatically Auditing Large Language Models via Discrete Optimization. arxiv:https:\/\/arXiv.org\/abs\/2303.04381\u00a0[cs] http:\/\/arxiv.org\/abs\/2303.04381"},{"key":"e_1_3_3_2_23_2","series-title":"(Sec \u201924)","volume-title":"Proceedings of the 33rd USENIX Conference on Security Symposium","author":"Lin Zilong","year":"2024","unstructured":"Zilong Lin, Jian Cui, Xiaojing Liao, and XiaoFeng Wang. 2024. Malla: Demystifying Real-World Large Language Model Integrated Malicious Services. In Proceedings of the 33rd USENIX Conference on Security Symposium(Sec \u201924). USENIX Association, Philadelphia, PA, USA and USA, Article 263."},{"key":"e_1_3_3_2_24_2","volume-title":"The Twelfth International Conference on Learning Representations","author":"Liu Xiaogeng","year":"2024","unstructured":"Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. 2024. AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models. In The Twelfth International Conference on Learning Representations. arxiv:https:\/\/arXiv.org\/abs\/2310.04451\u00a0[cs] http:\/\/arxiv.org\/abs\/2310.04451"},{"key":"e_1_3_3_2_25_2","unstructured":"Weidi Luo Siyuan Ma Xiaogeng Liu Xiaoyu Guo and Chaowei Xiao. 2024. JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks. arxiv:https:\/\/arXiv.org\/abs\/2404.03027\u00a0[cs.CR]"},{"key":"e_1_3_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2404.03027"},{"key":"e_1_3_3_2_27_2","doi-asserted-by":"publisher","unstructured":"Marcello Martina and Gian Foresti. 2021. A Continuous Learning Approach for Real-Time Network Intrusion Detection. International Journal of Neural Systems 31 (11 2021). 10.1142\/S012906572150060X","DOI":"10.1142\/S012906572150060X"},{"key":"e_1_3_3_2_28_2","unstructured":"Anay Mehrotra Manolis Zampetakis Paul Kassianik Blaine Nelson Hyrum Anderson Yaron Singer and Amin Karbasi. 2023. Tree of Attacks: Jailbreaking Black-Box LLMs Automatically. arxiv:https:\/\/arXiv.org\/abs\/2312.02119\u00a0[cs stat] http:\/\/arxiv.org\/abs\/2312.02119"},{"key":"e_1_3_3_2_29_2","unstructured":"Meta AI. 2023. Prompt Guard - Llama Model Cards and Prompt Formats. https:\/\/www.llama.com\/docs\/model-cards-and-prompt-formats\/prompt-guard\/. https:\/\/www.llama.com\/docs\/model-cards-and-prompt-formats\/prompt-guard\/ Accessed on April 11 2025."},{"key":"e_1_3_3_2_30_2","doi-asserted-by":"publisher","unstructured":"Seyedreza Mohseni Seyedali Mohammadi Deepa Tilwani Yash Saxena Gerald\u00a0Ketu Ndawula Sriram Vema Edward Raff and Manas Gaur. 2025. Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation. 10.48550\/arXiv.2412.16135 arxiv:https:\/\/arXiv.org\/abs\/2412.16135\u00a0[cs]","DOI":"10.48550\/arXiv.2412.16135"},{"key":"e_1_3_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.17"},{"key":"e_1_3_3_2_32_2","doi-asserted-by":"publisher","unstructured":"Stephen Moskal Sam Laney Erik Hemberg and Una-May O\u2019Reilly. 2023. LLMs Killed the Script Kiddie: How Agents Supported by Large Language Models Change the Landscape of Network Threat Testing. 10.48550\/arXiv.2310.06936 arxiv:https:\/\/arXiv.org\/abs\/2310.06936\u00a0[cs]","DOI":"10.48550\/arXiv.2310.06936"},{"key":"e_1_3_3_2_33_2","doi-asserted-by":"publisher","unstructured":"Maximilian Mozes Xuanli He Bennett Kleinberg and Lewis\u00a0D. Griffin. 2023. Use of LLMs for Illicit Purposes: Threats Prevention Measures and Vulnerabilities. 10.48550\/arXiv.2308.12833 arxiv:https:\/\/arXiv.org\/abs\/2308.12833\u00a0[cs]","DOI":"10.48550\/arXiv.2308.12833"},{"key":"e_1_3_3_2_34_2","unstructured":"OpenAI. 2025. Moderation - OpenAI API. https:\/\/platform.openai.com\/docs\/guides\/moderation\/overview"},{"key":"e_1_3_3_2_35_2","unstructured":"Long Ouyang Jeff Wu Xu Jiang Diogo Almeida Carroll\u00a0L. Wainwright Pamela Mishkin Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray John Schulman Jacob Hilton Fraser Kelton Luke Miller Maddie Simens Amanda Askell Peter Welinder Paul Christiano Jan Leike and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. arxiv:https:\/\/arXiv.org\/abs\/2203.02155\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2203.02155"},{"key":"e_1_3_3_2_36_2","unstructured":"Nicolas Papernot Patrick McDaniel and Ian Goodfellow. 2016. Transferability in Machine Learning: From Phenomena to Black-Box Attacks Using Adversarial Samples. arXiv:https:\/\/arXiv.org\/abs\/1605.07277 [cs] (May 2016). arxiv:https:\/\/arXiv.org\/abs\/1605.07277\u00a0[cs] http:\/\/arxiv.org\/abs\/1605.07277"},{"key":"e_1_3_3_2_37_2","unstructured":"Anselm Paulus Arman Zharmagambetov Chuan Guo Brandon Amos and Yuandong Tian. 2024. AdvPrompter: Fast Adaptive Adversarial Prompting for Llms. arxiv:https:\/\/arXiv.org\/abs\/2404.16873\u00a0[cs] http:\/\/arxiv.org\/abs\/2404.16873"},{"key":"e_1_3_3_2_38_2","unstructured":"Alwin Peng Julian Michael Henry Sleight Ethan Perez and Mrinank Sharma. 2024. Rapid Response: Mitigating LLM Jailbreaks with a Few Examples. arxiv:https:\/\/arXiv.org\/abs\/2411.07494\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2411.07494"},{"key":"e_1_3_3_2_39_2","doi-asserted-by":"publisher","unstructured":"Mary Phuong et\u00a0al. 2024. Evaluating Frontier Models for Dangerous Capabilities. 10.48550\/arXiv.2403.13793 arxiv:https:\/\/arXiv.org\/abs\/2403.13793\u00a0[cs]","DOI":"10.48550\/arXiv.2403.13793"},{"key":"e_1_3_3_2_40_2","unstructured":"Mansi Phute Alec Helbling Matthew Hull ShengYun Peng Sebastian Szyller Cory Cornelius and Duen\u00a0Horng Chau. 2024. LLM Self Defense: By Self Examination LLMs Know They Are Being Tricked. arxiv:https:\/\/arXiv.org\/abs\/2308.07308\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2308.07308"},{"key":"e_1_3_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-70879-4_6"},{"key":"e_1_3_3_2_42_2","unstructured":"Pengzhen Ren Yun Xiao Xiaojun Chang Po-Yao Huang Zhihui Li Brij\u00a0B. Gupta Xiaojiang Chen and Xin Wang. 2021. A Survey of Deep Active Learning. arxiv:https:\/\/arXiv.org\/abs\/2009.00236\u00a0[cs.LG] https:\/\/arxiv.org\/abs\/2009.00236"},{"key":"e_1_3_3_2_43_2","unstructured":"Sayak\u00a0Saha Roy Poojitha Thota Krishna\u00a0Vamsi Naragam and Shirin Nilizadeh. 2023. From Chatbots to PhishBots? \u2013 Preventing Phishing Scams Created Using ChatGPT Google Bard and Claude. arxiv:https:\/\/arXiv.org\/abs\/2310.19181\u00a0[cs] http:\/\/arxiv.org\/abs\/2310.19181"},{"key":"e_1_3_3_2_44_2","doi-asserted-by":"publisher","unstructured":"Mrinank Sharma et\u00a0al. 2025. Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming. 10.48550\/arXiv.2501.18837 arxiv:https:\/\/arXiv.org\/abs\/2501.18837\u00a0[cs]","DOI":"10.48550\/arXiv.2501.18837"},{"key":"e_1_3_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2308.03825"},{"key":"e_1_3_3_2_46_2","unstructured":"Chawin Sitawarin Norman Mu David Wagner and Alexandre Araujo. 2024. PAL: Proxy-Guided Black-Box Attack on Large Language Models. arxiv:https:\/\/arXiv.org\/abs\/2402.09674\u00a0[cs] http:\/\/arxiv.org\/abs\/2402.09674"},{"key":"e_1_3_3_2_47_2","unstructured":"Gemini Team and Google. 2023. Gemini: A Family of Highly Capable Multimodal Models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2312.11805 (2023). https:\/\/arxiv.org\/abs\/2312.11805"},{"key":"e_1_3_3_2_48_2","unstructured":"Llama Team. 2024. The Llama 3 Herd of Models. arxiv:https:\/\/arXiv.org\/abs\/2407.21783\u00a0[cs.AI] https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_3_2_49_2","unstructured":"Hugo Touvron et\u00a0al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2307.09288 (2023)."},{"key":"e_1_3_3_2_50_2","unstructured":"Apurv Verma Satyapriya Krishna Sebastian Gehrmann Madhavan Seshadri Anu Pradhan Tom Ault Leslie Barrett David Rabinowitz John Doucette and NhatHai Phan. 2024. Operationalizing a threat model for red-teaming large language models (llms). arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2407.14937 (2024)."},{"key":"e_1_3_3_2_51_2","doi-asserted-by":"publisher","unstructured":"Shengye Wan Cyrus Nikolaidis Daniel Song David Molnar James Crnkovich Jayson Grace Manish Bhatt Sahana Chennabasappa Spencer Whitman Stephanie Ding Vlad Ionescu Yue Li and Joshua Saxe. 2024. CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models. 10.48550\/arXiv.2408.01605 arxiv:https:\/\/arXiv.org\/abs\/2408.01605\u00a0[cs]","DOI":"10.48550\/arXiv.2408.01605"},{"key":"e_1_3_3_2_52_2","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Wang Boxin","year":"2023","unstructured":"Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang\u00a0T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. 2023. DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models. In Proceedings of the 37th International Conference on Neural Information Processing Systems. Curran Associates, Inc., Article 1361. arxiv:https:\/\/arXiv.org\/abs\/2306.11698\u00a0[cs] http:\/\/arxiv.org\/abs\/2306.11698"},{"key":"e_1_3_3_2_53_2","doi-asserted-by":"publisher","unstructured":"Liyuan Wang Xingxing Zhang Hang Su and Jun Zhu. 2024. A Comprehensive Survey of Continual Learning: Theory Method and Application. IEEE Transactions on Pattern Analysis and Machine Intelligence 46 8 (2024) 5362\u20135383. 10.1109\/TPAMI.2024.3367329","DOI":"10.1109\/TPAMI.2024.3367329"},{"key":"e_1_3_3_2_54_2","unstructured":"Peiran Wang Xiaogeng Liu and Chaowei Xiao. 2024. RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process. arxiv:https:\/\/arXiv.org\/abs\/2410.08660\u00a0[cs.CR] https:\/\/arxiv.org\/abs\/2410.08660"},{"key":"e_1_3_3_2_55_2","doi-asserted-by":"publisher","unstructured":"Xunguang Wang Wenxuan Wang Zhenlan Ji Zongjie Li Pingchuan Ma Daoyuan Wu and Shuai Wang. 2025. STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models. 10.48550\/arXiv.2503.17932 arxiv:https:\/\/arXiv.org\/abs\/2503.17932\u00a0[cs]","DOI":"10.48550\/arXiv.2503.17932"},{"key":"e_1_3_3_2_56_2","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Wei Alexander","year":"2023","unstructured":"Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How Does LLM Safety Training Fail?. In Proceedings of the 37th International Conference on Neural Information Processing Systems. Curran Associates Inc. arxiv:https:\/\/arXiv.org\/abs\/2307.02483\u00a0[cs] http:\/\/arxiv.org\/abs\/2307.02483"},{"key":"e_1_3_3_2_57_2","unstructured":"Yueqi Xie Minghong Fang Renjie Pi and Neil Gong. 2024. GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis. arxiv:https:\/\/arXiv.org\/abs\/2402.13494\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2402.13494"},{"key":"e_1_3_3_2_58_2","doi-asserted-by":"publisher","unstructured":"Rongwu Xu Xiaojian Li Shuo Chen and Wei Xu. 2025. Nuclear Deployed: Analyzing Catastrophic Risks in Decision-Making of Autonomous LLM Agents. 10.48550\/arXiv.2502.11355 arxiv:https:\/\/arXiv.org\/abs\/2502.11355\u00a0[cs]","DOI":"10.48550\/arXiv.2502.11355"},{"key":"e_1_3_3_2_59_2","first-page":"2327","volume-title":"30th USENIX Security Symposium (USENIX Security 21)","author":"Yang Limin","year":"2021","unstructured":"Limin Yang, Wenbo Guo, Qingying Hao, Arridhana Ciptadi, Ali Ahmadzadeh, Xinyu Xing, and Gang Wang. 2021. CADE: Detecting and Explaining Concept Drift Samples for Security Applications. In 30th USENIX Security Symposium (USENIX Security 21). USENIX Association, 2327\u20132344. https:\/\/www.usenix.org\/conference\/usenixsecurity21\/presentation\/yang-limin"},{"key":"e_1_3_3_2_60_2","unstructured":"Zheng-Xin Yong Cristina Menghini and Stephen\u00a0H. Bach. 2023. Low-Resource Languages Jailbreak GPT-4. arxiv:https:\/\/arXiv.org\/abs\/2310.02446\u00a0[cs] http:\/\/arxiv.org\/abs\/2310.02446"},{"key":"e_1_3_3_2_61_2","unstructured":"Jiahao Yu Xingwei Lin Zheng Yu and Xinyu Xing. 2024. GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts. arxiv:https:\/\/arXiv.org\/abs\/2309.10253\u00a0[cs] http:\/\/arxiv.org\/abs\/2309.10253"},{"key":"e_1_3_3_2_62_2","unstructured":"Zhuowen Yuan Zidi Xiong Yi Zeng Ning Yu Ruoxi Jia Dawn Song and Bo Li. 2024. RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content. arxiv:https:\/\/arXiv.org\/abs\/2403.13031\u00a0[cs.CR] https:\/\/arxiv.org\/abs\/2403.13031"},{"key":"e_1_3_3_2_63_2","unstructured":"Wenjun Zeng Yuchi Liu Ryan Mullins Ludovic Peran Joe Fernandez Hamza Harkous Karthik Narasimhan Drew Proud Piyush Kumar Bhaktipriya Radharapu Olivia Sturman and Oscar Wahltinez. 2024. ShieldGemma: Generative AI Content Moderation Based on Gemma. arxiv:https:\/\/arXiv.org\/abs\/2407.21772\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2407.21772"},{"key":"e_1_3_3_2_64_2","unstructured":"Yi Zeng Hongpeng Lin Jingwen Zhang Diyi Yang Ruoxi Jia and Weiyan Shi. 2024. How Johnny Can Persuade Llms to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing Llms. arxiv:https:\/\/arXiv.org\/abs\/2401.06373\u00a0[cs] http:\/\/arxiv.org\/abs\/2401.06373"},{"key":"e_1_3_3_2_65_2","doi-asserted-by":"publisher","unstructured":"Peiyuan Zhang Guangtao Zeng Tianduo Wang and Wei Lu. 2024. TinyLlama: An Open-Source Small Language Model. 10.48550\/arXiv.2401.02385 arxiv:https:\/\/arXiv.org\/abs\/2401.02385\u00a0[cs]","DOI":"10.48550\/arXiv.2401.02385"},{"key":"e_1_3_3_2_66_2","doi-asserted-by":"publisher","unstructured":"Shenyi Zhang Yuchen Zhai Keyan Guo Hongxin Hu Shengnan Guo Zheng Fang Lingchen Zhao Chao Shen Cong Wang and Qian Wang. 2025. JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation. 10.48550\/arXiv.2502.07557 arxiv:https:\/\/arXiv.org\/abs\/2502.07557\u00a0[cs]","DOI":"10.48550\/arXiv.2502.07557"},{"key":"e_1_3_3_2_67_2","doi-asserted-by":"publisher","unstructured":"Xiaoyu Zhang Cen Zhang Tianlin Li Yihao Huang Xiaojun Jia Ming Hu Jie Zhang Yang Liu Shiqing Ma and Chao Shen. 2025. JailGuard: A Universal Detection Framework for Prompt-Based Attacks on LLM Systems. ACM Transactions on Software Engineering and Methodology (March 2025) 3724393. 10.1145\/3724393","DOI":"10.1145\/3724393"},{"key":"e_1_3_3_2_68_2","unstructured":"Zhexin Zhang Yida Lu Jingyuan Ma Di Zhang Rui Li Pei Ke Hao Sun Lei Sha Zhifang Sui Hongning Wang and Minlie Huang. 2024. ShieldLM: Empowering LLMs as Aligned Customizable and Explainable Safety Detectors. arxiv:https:\/\/arXiv.org\/abs\/2402.16444\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2402.16444"},{"key":"e_1_3_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2405.01470"},{"key":"e_1_3_3_2_70_2","unstructured":"Lianmin Zheng Wei-Lin Chiang Ying Sheng Tianle Li Siyuan Zhuang Zhanghao Wu Yonghao Zhuang Zhuohan Li Zi Lin Eric\u00a0P. Xing Joseph\u00a0E. Gonzalez Ion Stoica and Hao Zhang. 2024. LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset. arxiv:https:\/\/arXiv.org\/abs\/2309.11998\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2309.11998"},{"key":"e_1_3_3_2_71_2","volume-title":"First Conference on Language Modeling","author":"Zhu Sicheng","year":"2024","unstructured":"Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun. 2024. AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models. In First Conference on Language Modeling. https:\/\/openreview.net\/forum?id=INivcBeIDK"},{"key":"e_1_3_3_2_72_2","unstructured":"Andy Zou Long Phan Justin Wang Derek Duenas Maxwell Lin Maksym Andriushchenko Rowan Wang Zico Kolter Matt Fredrikson and Dan Hendrycks. 2024. Improving Alignment and Robustness with Circuit Breakers. arxiv:https:\/\/arXiv.org\/abs\/2406.04313\u00a0[cs.LG]"},{"key":"e_1_3_3_2_73_2","unstructured":"Andy Zou Zifan Wang Nicholas Carlini Milad Nasr J.\u00a0Zico Kolter and Matt Fredrikson. 2023. Universal and Transferable Adversarial Attacks on Aligned Language Models. arxiv:https:\/\/arXiv.org\/abs\/2307.15043\u00a0[cs] http:\/\/arxiv.org\/abs\/2307.15043"}],"event":{"name":"AISec '25: Proceedings of the 2025 Workshop on Artificial Intelligence and Security","location":"Taipei , Taiwan","acronym":"AISec '25","sponsor":["SIGSAC ACM Special Interest Group on Security, Audit, and Control"]},"container-title":["Proceedings of the 18th ACM Workshop on Artificial Intelligence and Security"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3733799.3762981","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,30]],"date-time":"2025-12-30T11:52:35Z","timestamp":1767095555000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3733799.3762981"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,13]]},"references-count":72,"alternative-id":["10.1145\/3733799.3762981","10.1145\/3733799"],"URL":"https:\/\/doi.org\/10.1145\/3733799.3762981","relation":{},"subject":[],"published":{"date-parts":[[2025,10,13]]},"assertion":[{"value":"2025-12-30","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}