{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T10:11:32Z","timestamp":1784542292776,"version":"3.55.0"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"ISSTA","funder":[{"name":"the State Key Laboratory of Industrial Control Technology, China","award":["ICT2024C01"],"award-info":[{"award-number":["ICT2024C01"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,22]]},"abstract":"<jats:p>Generative large language models (LLMs) have revolutionized natural language processing with their transformative and emergent capabilities. However, recent evidence indicates that LLMs can produce harmful content that violates social norms, raising significant concerns regarding the safety and ethical ramifications of deploying these advanced models. Thus, it is both critical and imperative to perform a rigorous and comprehensive safety evaluation of LLMs before deployment. Despite this need, owing to the extensiveness of LLM generation space, it still lacks a unified and standardized risk taxonomy to systematically reflect the LLM content safety, as well as automated safety assessment techniques to explore the potential risks efficiently.<\/jats:p>\n          <jats:p>\n            To bridge the striking gap, we propose S-Eval, a novel LLM-based automated Safety Evaluation framework with a newly defined comprehensive risk taxonomy. S-Eval incorporates two key components, i.e., an expert testing LLM\n            <jats:italic toggle=\"yes\">M<\/jats:italic>\n            <jats:sub>\n              <jats:italic toggle=\"yes\">t<\/jats:italic>\n            <\/jats:sub>\n            and a novel safety critique LLM\n            <jats:italic toggle=\"yes\">M<\/jats:italic>\n            <jats:sub>\n              <jats:italic toggle=\"yes\">c<\/jats:italic>\n            <\/jats:sub>\n            . The expert testing LLM\n            <jats:italic toggle=\"yes\">M<\/jats:italic>\n            <jats:sub>\n              <jats:italic toggle=\"yes\">t<\/jats:italic>\n            <\/jats:sub>\n            is responsible for automatically generating test cases in accordance with the proposed risk management (including 8 risk dimensions and a total of 102 subdivided risks). The safety critique LLM\n            <jats:italic toggle=\"yes\">M<\/jats:italic>\n            <jats:sub>\n              <jats:italic toggle=\"yes\">c<\/jats:italic>\n            <\/jats:sub>\n            can provide quantitative and explainable safety evaluations for better risk awareness of LLMs. In contrast to prior works, S-Eval differs in significant ways: (i)\n            <jats:italic toggle=\"yes\">efficient<\/jats:italic>\n            \u2013 we construct a multi-dimensional and open-ended benchmark comprising 220,000 test cases across 102 risks utilizing\n            <jats:italic toggle=\"yes\">M<\/jats:italic>\n            <jats:sub>\n              <jats:italic toggle=\"yes\">t<\/jats:italic>\n            <\/jats:sub>\n            and conduct safety evaluations for 21 influential LLMs via\n            <jats:italic toggle=\"yes\">M<\/jats:italic>\n            <jats:sub>\n              <jats:italic toggle=\"yes\">c<\/jats:italic>\n            <\/jats:sub>\n            on our benchmark. The entire process is fully automated and requires no human involvement. (ii)\n            <jats:italic toggle=\"yes\">effective<\/jats:italic>\n            \u2013 extensive validations show S-Eval facilitates a more thorough assessment and better perception of potential LLM risks, and\n            <jats:italic toggle=\"yes\">M<\/jats:italic>\n            <jats:sub>\n              <jats:italic toggle=\"yes\">c<\/jats:italic>\n            <\/jats:sub>\n            not only accurately quantifies the risks of LLMs but also provides explainable and in-depth insights into their safety, surpassing comparable models such as LLaMA-Guard-2. (iii)\n            <jats:italic toggle=\"yes\">adaptive<\/jats:italic>\n            \u2013 S-Eval can be flexibly configured and adapted to the rapid evolution of LLMs and accompanying new safety threats, test generation methods and safety critique methods thanks to the LLM-based architecture. We further study the impact of hyper-parameters and language environments on model safety, which may lead to promising directions for future research. S-Eval has been deployed in our industrial partner for the automated safety evaluation of multiple LLMs serving millions of users, demonstrating its effectiveness in real-world scenarios.\n          <\/jats:p>","DOI":"10.1145\/3728971","type":"journal-article","created":{"date-parts":[[2025,6,22]],"date-time":"2025-06-22T10:52:56Z","timestamp":1750589576000},"page":"2136-2157","source":"Crossref","is-referenced-by-count":7,"title":["S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-4866-245X","authenticated-orcid":false,"given":"Xiaohan","family":"Yuan","sequence":"first","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-8280-897X","authenticated-orcid":false,"given":"Jinfeng","family":"Li","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9812-3911","authenticated-orcid":false,"given":"Dongxia","family":"Wang","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9027-3421","authenticated-orcid":false,"given":"Yuefeng","family":"Chen","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4486-556X","authenticated-orcid":false,"given":"Xiaofeng","family":"Mao","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0517-1592","authenticated-orcid":false,"given":"Longtao","family":"Huang","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4322-4285","authenticated-orcid":false,"given":"Jialuo","family":"Chen","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2093-2839","authenticated-orcid":false,"given":"Hui","family":"Xue","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-6662-9020","authenticated-orcid":false,"given":"Xiaoxia","family":"Liu","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1936-2840","authenticated-orcid":false,"given":"Wenhai","family":"Wang","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3441-6277","authenticated-orcid":false,"given":"Kui","family":"Ren","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7113-7635","authenticated-orcid":false,"given":"Jingyi","family":"Wang","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,22]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Diogo Almeida, Janko Altenschmidt, Sam Altman, and Shyamal Anadkat.","author":"Achiam Josh","year":"2023","unstructured":"Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, and Shyamal Anadkat. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774."},{"key":"e_1_2_1_2_1","doi-asserted-by":"crossref","unstructured":"NIST AI. 2023. Artificial Intelligence Risk Management Framework (AI RMF 1.0).","DOI":"10.6028\/NIST.AI.100-1.jpn"},{"key":"e_1_2_1_3_1","unstructured":"AI@Meta. 2024. Llama 3 Model Card. https:\/\/github.com\/meta-llama\/llama3\/blob\/main\/MODEL_CARD.md"},{"key":"e_1_2_1_4_1","unstructured":"Jinze Bai Shuai Bai Yunfei Chu Zeyu Cui Kai Dang Xiaodong Deng Yang Fan Wenbin Ge Yu Han and Fei Huang. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609."},{"key":"e_1_2_1_5_1","unstructured":"Baidu. 2023. ErnieBot. https:\/\/yiyan.baidu.com\/"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics. 675\u2013718","author":"Bang Yejin","year":"2023","unstructured":"Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, and Willy Chung. 2023. A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics. 675\u2013718."},{"key":"e_1_2_1_7_1","volume-title":"Risk society: Towards a new modernity","author":"Beck Ulrich","unstructured":"Ulrich Beck. 1992. Risk society: Towards a new modernity. Sage."},{"key":"e_1_2_1_8_1","unstructured":"Rishabh Bhardwaj and Soujanya Poria. 2023. Language model unalignment: Parametric red-teaming to expose hidden harms and biases. arXiv preprint arXiv:2310.14303."},{"key":"e_1_2_1_9_1","unstructured":"Rishabh Bhardwaj and Soujanya Poria. 2023. Red-teaming large language models using chain of utterances for safety-alignment. arXiv preprint arXiv:2308.09662."},{"key":"e_1_2_1_10_1","volume-title":"Varieties of criminal behavior: Summary and policy implications","author":"Chaiken Jan M","unstructured":"Jan M Chaiken, Marcia R Chaiken, and Joyce E Peterson. 1982. Varieties of criminal behavior: Summary and policy implications. Rand Santa Monica, CA."},{"key":"e_1_2_1_11_1","unstructured":"Justin Cui Wei-Lin Chiang Ion Stoica and Cho-Jui Hsieh. 2024. OR-Bench: An Over-Refusal Benchmark for Large Language Models. arXiv preprint arXiv:2405.20947."},{"key":"e_1_2_1_12_1","first-page":"14","article-title":"Misuse of the Internet by pedophiles: Implications for law enforcement and probation practice","volume":"61","author":"Durkin Keith F","year":"1997","unstructured":"Keith F Durkin. 1997. Misuse of the Internet by pedophiles: Implications for law enforcement and probation practice. Fed. Probation, 61 (1997), 14.","journal-title":"Fed. Probation"},{"key":"e_1_2_1_13_1","unstructured":"Deep Ganguli Liane Lovitt Jackson Kernion Amanda Askell Yuntao Bai Saurav Kadavath Ben Mann Ethan Perez Nicholas Schiefer and Kamal Ndousse. 2022. Red teaming language models to reduce harms: Methods scaling behaviors and lessons learned. arXiv preprint arXiv:2209.07858."},{"key":"e_1_2_1_14_1","volume-title":"Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462.","author":"Gehman Samuel","year":"2020","unstructured":"Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462."},{"key":"e_1_2_1_15_1","volume-title":"Chatglm: A family of large language models from glm-130b to glm-4 all tools. arXiv preprint arXiv:2406.12793.","author":"Aohan Zeng Team GLM","year":"2024","unstructured":"Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, and Hanlin Zhao. 2024. Chatglm: A family of large language models from glm-130b to glm-4 all tools. arXiv preprint arXiv:2406.12793."},{"key":"e_1_2_1_16_1","unstructured":"Google. 2023. Generative AI Prohibited Use Policy. https:\/\/policies.google.com\/terms\/generative-ai\/use-policy"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Hendrycks Dan","year":"2021","unstructured":"Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021. Aligning AI With Shared Human Values. Proceedings of the International Conference on Learning Representations."},{"key":"e_1_2_1_18_1","volume-title":"Flames: Benchmarking value alignment of chinese large language models. arXiv preprint arXiv:2311.06899.","author":"Huang Kexin","year":"2023","unstructured":"Kexin Huang, Xiangyang Liu, Qianyu Guo, Tianxiang Sun, Jiawei Sun, Yaru Wang, Zeyang Zhou, Yixu Wang, Yan Teng, and Xipeng Qiu. 2023. Flames: Benchmarking value alignment of chinese large language models. arXiv preprint arXiv:2311.06899."},{"key":"e_1_2_1_19_1","volume-title":"Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, and Lucile Saulnier.","author":"Jiang Albert Q","year":"2023","unstructured":"Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, and Lucile Saulnier. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825."},{"key":"e_1_2_1_20_1","unstructured":"Shuyu Jiang Xingshu Chen and Rui Tang. 2023. Prompt packer: Deceiving llms through compositional instruction with hidden attacks. arXiv preprint arXiv:2310.10077."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/SPW63631.2024.00018"},{"key":"e_1_2_1_22_1","volume-title":"Critiquellm: Scaling llm-as-critic for effective and explainable evaluation of large language model generation. arXiv preprint arXiv:2311.18702.","author":"Ke Pei","year":"2023","unstructured":"Pei Ke, Bosi Wen, Zhuoer Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, and Hongning Wang. 2023. Critiquellm: Scaling llm-as-critic for effective and explainable evaluation of large language model generation. arXiv preprint arXiv:2311.18702."},{"key":"e_1_2_1_23_1","doi-asserted-by":"crossref","unstructured":"Lijun Li Bowen Dong Ruohui Wang Xuhao Hu Wangmeng Zuo Dahua Lin Yu Qiao and Jing Shao. 2024. SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models. arXiv preprint arXiv:2402.05044.","DOI":"10.18653\/v1\/2024.findings-acl.235"},{"key":"e_1_2_1_24_1","volume-title":"Deepinception: Hypnotize large language model to be jailbreaker. arXiv preprint arXiv:2311.03191.","author":"Li Xuan","year":"2023","unstructured":"Xuan Li, Zhanke Zhou, Jianing Zhu, Jiangchao Yao, Tongliang Liu, and Bo Han. 2023. Deepinception: Hypnotize large language model to be jailbreaker. arXiv preprint arXiv:2311.03191."},{"key":"e_1_2_1_25_1","unstructured":"Percy Liang Rishi Bommasani Tony Lee Dimitris Tsipras Dilara Soylu Michihiro Yasunaga Yian Zhang Deepak Narayanan Yuhuai Wu and Ananya Kumar. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_2_1_27_1","unstructured":"Xiaoxia Liu Jingyi Wang Jun Sun Xiaohan Yuan Guoliang Dong Peng Di Wenhai Wang and Dongxia Wang. 2023. Prompting frameworks for large language models: A survey. arXiv preprint arXiv:2311.12785."},{"key":"e_1_2_1_28_1","doi-asserted-by":"crossref","unstructured":"Yi Liu Gelei Deng Zhengzi Xu Yuekang Li Yaowen Zheng Ying Zhang Lida Zhao Tianwei Zhang and Yang Liu. 2023. Jailbreaking chatgpt via prompt engineering: An empirical study. arXiv preprint arXiv:2305.13860.","DOI":"10.1145\/3663530.3665021"},{"key":"e_1_2_1_29_1","unstructured":"OpenAI. 2024. Hello GPT-4o. https:\/\/openai.com\/index\/hello-gpt-4o"},{"key":"e_1_2_1_30_1","unstructured":"OpenAI. 2024. Moderation. https:\/\/platform.openai.com\/docs\/guides\/moderation"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","unstructured":"D Wayne Osgood. 2010. Statistical models of life events and criminal behavior. Handbook of quantitative criminology 375\u2013396. https:\/\/doi.org\/10.1007\/978-0-387-77650-7_19 10.1007\/978-0-387-77650-7_19","DOI":"10.1007\/978-0-387-77650-7_19"},{"key":"e_1_2_1_32_1","unstructured":"European Parliament. 2021. Artificial Intelligence Act. https:\/\/artificialintelligenceact.com\/"},{"key":"e_1_2_1_33_1","volume-title":"Phu Mon Htut, and Samuel R Bowman","author":"Parrish Alicia","year":"2021","unstructured":"Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman. 2021. BBQ: A hand-built bias benchmark for question answering. arXiv preprint arXiv:2110.08193."},{"key":"e_1_2_1_34_1","doi-asserted-by":"crossref","unstructured":"Emily Sheng Kai-Wei Chang Premkumar Natarajan and Nanyun Peng. 2021. Societal biases in language generation: Progress and challenges. arXiv preprint arXiv:2105.04054.","DOI":"10.18653\/v1\/2021.acl-long.330"},{"key":"e_1_2_1_35_1","doi-asserted-by":"crossref","unstructured":"Taylor Shin Yasaman Razeghi Robert L. Logan IV Eric Wallace and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Empirical Methods in Natural Language Processing.","DOI":"10.18653\/v1\/2020.emnlp-main.346"},{"key":"e_1_2_1_36_1","unstructured":"Guijin Son Hanearl Jung Moonjeong Hahm Keonju Na and Sol Jin. 2023. Beyond classification: Financial reasoning in state-of-the-art language models. arXiv preprint arXiv:2305.01505."},{"key":"e_1_2_1_37_1","unstructured":"Hao Sun Zhexin Zhang Jiawen Deng Jiale Cheng and Minlie Huang. 2023. Safety Assessment of Chinese Large Language Models. arXiv preprint arXiv:2304.10436."},{"key":"e_1_2_1_38_1","volume-title":"Trustllm: Trustworthiness in large language models. arXiv preprint arXiv:2401.05561.","author":"Sun Lichao","year":"2024","unstructured":"Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, and Xiner Li. 2024. Trustllm: Trustworthiness in large language models. arXiv preprint arXiv:2401.05561."},{"key":"e_1_2_1_39_1","unstructured":"Ruixiang Tang Xiaotian Han Xiaoqian Jiang and Xia Hu. 2023. Does synthetic data generation of llms help clinical text mining? arXiv preprint arXiv:2303.04360."},{"key":"e_1_2_1_40_1","unstructured":"Gemini Team Rohan Anil Sebastian Borgeaud Yonghui Wu Jean-Baptiste Alayrac Jiahui Yu Radu Soricut Johan Schalkwyk Andrew M Dai and Anja Hauth. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805."},{"key":"e_1_2_1_41_1","volume-title":"Mihir Sanjay Kale, and Juliette Love","author":"Team Gemma","year":"2024","unstructured":"Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi\u00e8re, Mihir Sanjay Kale, and Juliette Love. 2024. Gemma: Open Models Based on Gemini Research and Technology. arXiv preprint arXiv:2403.08295."},{"key":"e_1_2_1_42_1","unstructured":"Llama Team. 2024. Meta Llama Guard 2. https:\/\/github.com\/meta-llama\/PurpleLlama\/blob\/main\/Llama-Guard2\/MODEL_CARD.md"},{"key":"e_1_2_1_43_1","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava and Shruti Bhosale. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288."},{"key":"e_1_2_1_44_1","volume-title":"\u0141 ukasz Kaiser, and Illia Polosukhin","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141 ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30 (2017)."},{"key":"e_1_2_1_45_1","volume-title":"DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models. Advances in Neural Information Processing Systems, 36","author":"Wang Boxin","year":"2024","unstructured":"Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, and Rylan Schaeffer. 2024. DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models. Advances in Neural Information Processing Systems, 36 (2024)."},{"key":"e_1_2_1_46_1","volume-title":"Do-not-answer: A dataset for evaluating safeguards in llms. arXiv preprint arXiv:2308.13387.","author":"Wang Yuxia","year":"2023","unstructured":"Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. 2023. Do-not-answer: A dataset for evaluating safeguards in llms. arXiv preprint arXiv:2308.13387."},{"key":"e_1_2_1_47_1","volume-title":"Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36","author":"Wei Alexander","year":"2024","unstructured":"Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2024. Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36 (2024)."},{"key":"e_1_2_1_48_1","unstructured":"Zeming Wei Yifei Wang and Yisen Wang. 2023. Jailbreak and guard aligned language models with only few in-context demonstrations. arXiv preprint arXiv:2310.06387."},{"key":"e_1_2_1_49_1","unstructured":"Jules White Quchen Fu Sam Hays Michael Sandborn Carlos Olea Henry Gilbert Ashraf Elnashar Jesse Spencer-Smith and Douglas C Schmidt. 2023. A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382."},{"key":"e_1_2_1_50_1","volume-title":"Cvalues: Measuring the values of chinese large language models from safety to responsibility. arXiv preprint arXiv:2307.09705.","author":"Xu Guohai","year":"2023","unstructured":"Guohai Xu, Jiayi Liu, Ming Yan, Haotian Xu, Jinghui Si, Zhuoran Zhou, Peng Yi, Xing Gao, Jitao Sang, and Rong Zhang. 2023. Cvalues: Measuring the values of chinese large language models from safety to responsibility. arXiv preprint arXiv:2307.09705."},{"key":"e_1_2_1_51_1","unstructured":"Aiyuan Yang Bin Xiao Bingning Wang Borong Zhang Ce Bian Chao Yin Chenxu Lv Da Pan Dian Wang and Dong Yan. 2023. Baichuan 2: Open large-scale language models. arXiv preprint arXiv:2309.10305."},{"key":"e_1_2_1_52_1","volume-title":"Yi: Open Foundation Models by 01. AI. arXiv preprint arXiv:2403.04652.","author":"Young Alex","year":"2024","unstructured":"Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, and Jing Chang. 2024. Yi: Open Foundation Models by 01. AI. arXiv preprint arXiv:2403.04652."},{"key":"e_1_2_1_53_1","volume-title":"Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts. arXiv preprint arXiv:2309.10253.","author":"Yu Jiahao","year":"2023","unstructured":"Jiahao Yu, Xingwei Lin, and Xinyu Xing. 2023. Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts. arXiv preprint arXiv:2309.10253."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-emnlp.79"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/IAEAC.2017.8054419"},{"key":"e_1_2_1_56_1","volume-title":"Safetybench: Evaluating the safety of large language models with multiple choice questions. arXiv preprint arXiv:2309.07045.","author":"Zhang Zhexin","year":"2023","unstructured":"Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2023. Safetybench: Evaluating the safety of large language models with multiple choice questions. arXiv preprint arXiv:2309.07045."},{"key":"e_1_2_1_57_1","volume-title":"Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36","author":"Zheng Lianmin","year":"2024","unstructured":"Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, and Eric Xing. 2024. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36 (2024)."},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1080\/00141840902940492"},{"key":"e_1_2_1_59_1","unstructured":"Andy Zou Zifan Wang J Zico Kolter and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043."}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728971","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,16]],"date-time":"2025-07-16T16:46:10Z","timestamp":1752684370000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728971"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,22]]},"references-count":59,"journal-issue":{"issue":"ISSTA","published-print":{"date-parts":[[2025,6,22]]}},"alternative-id":["10.1145\/3728971"],"URL":"https:\/\/doi.org\/10.1145\/3728971","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,22]]}}}