{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T19:31:35Z","timestamp":1784662295241,"version":"3.55.0"},"reference-count":259,"publisher":"Springer Science and Business Media LLC","issue":"12","license":[{"start":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T00:00:00Z","timestamp":1760659200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T00:00:00Z","timestamp":1760659200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100000266","name":"Engineering and Physical Sciences Research Council","doi-asserted-by":"publisher","award":["EP\/T026995\/1"],"award-info":[{"award-number":["EP\/T026995\/1"]}],"id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>In the burgeoning field of Large Language Models (LLMs), developing a robust safety mechanism, colloquially known as \u201csafeguards\u201d or \u201cguardrails\u201d, has become imperative to ensure the ethical use of LLMs within prescribed boundaries. This article provides a systematic literature review on the current status of this critical mechanism. It discusses its major challenges and how it can be enhanced into a comprehensive mechanism dealing with ethical issues in various contexts. First, the paper elucidates the current landscape of safeguarding mechanisms that major LLM service providers and the open-source community employ. This is followed by the techniques to evaluate, analyze, and enhance some (un)desirable properties that a guardrail might want to enforce, such as hallucinations, fairness, privacy, and so on. Based on them, we review techniques to circumvent these controls (i.e., attacks), to defend the attacks, and to reinforce the guardrails. While the techniques mentioned above represent the current status and the active research trends, we also discuss several challenges that cannot be easily dealt with by the methods and present our vision on how to implement a comprehensive guardrail through the full consideration of multi-disciplinary approach, neural-symbolic method, and systems development lifecycle.<\/jats:p>","DOI":"10.1007\/s10462-025-11389-2","type":"journal-article","created":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T02:48:29Z","timestamp":1760669309000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":72,"title":["Safeguarding large language models: a survey"],"prefix":"10.1007","volume":"58","author":[{"given":"Yi","family":"Dong","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ronghui","family":"Mu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanghao","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Siqi","family":"Sun","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tianle","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Changshun","family":"Wu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gaojie","family":"Jin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi","family":"Qi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinwei","family":"Hu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jie","family":"Meng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Saddek","family":"Bensalem","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaowei","family":"Huang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,10,17]]},"reference":[{"key":"11389_CR1","doi-asserted-by":"publisher","unstructured":"Abadi M, Chu A, Goodfellow I, McMahan HB, Mironov I, Talwar K, Zhang L (2016) Deep learning with differential privacy. In: Proc. 2016 ACM SIGSAC Conf. Comput. Commun. Secur. CCS \u201916, pp. 308\u2013318. Association for Computing Machinery, New York, NY, USA . https:\/\/doi.org\/10.1145\/2976749.2978318","DOI":"10.1145\/2976749.2978318"},{"key":"11389_CR2","unstructured":"ActiveFence (2023) LLM Safety Review: Benchmarks and Analysis"},{"key":"11389_CR3","unstructured":"Alon G, Kamfonas M (2023) Detecting language model attacks with perplexity. arXiv prepr. arxiv:2308.14132"},{"issue":"5","key":"11389_CR4","doi-asserted-by":"publisher","first-page":"997","DOI":"10.1037\/0022-3514.43.5.997","volume":"43","author":"TM Amabile","year":"1982","unstructured":"Amabile TM (1982) Social psychology of creativity: a consensual assessment technique. J Pers Soc Psychol 43(5):997","journal-title":"J Pers Soc Psychol"},{"key":"11389_CR5","unstructured":"Anil R, Dai AM, Firat O, Johnson M, Lepikhin D, Passos A, Shakeri S, Taropa E, Bailey P, Chen Z, et\u00a0al (2023) Palm 2 technical report. arXiv prepr. arxiv:2305.10403"},{"key":"11389_CR6","doi-asserted-by":"crossref","unstructured":"Arora U, Huang W, He H (2021) Types of out-of-distribution texts and how to detect them. arXiv prepr. arxiv:2109.06827","DOI":"10.18653\/v1\/2021.emnlp-main.835"},{"key":"11389_CR7","first-page":"1","volume":"10","author":"G Arun","year":"2025","unstructured":"Arun G, Syam R, Nair AA, Vaidya S (2025) An integrated framework for ethical healthcare chatbots using langchain and nemo guardrails. AI Ethics 10:1\u201312","journal-title":"AI Ethics"},{"key":"11389_CR8","unstructured":"Ayyamperumal SG, Ge L (2024) Current state of llm risks and ai guardrails. arXiv preprint arXiv:2406.12934"},{"key":"11389_CR9","doi-asserted-by":"crossref","unstructured":"Badyal N, Jacoby D, Coady Y (2023) Intentional biases in LLM responses. In: 2023 IEEE 14th Annu. Ubiquitous Comput. Electron. Mob. Commun. Conf. (UEMCON), pp. 0502\u20130506. IEEE","DOI":"10.1109\/UEMCON59035.2023.10316060"},{"key":"11389_CR10","unstructured":"Bai Y, Kadavath S, Kundu S, Askell A, Kernion J, Jones A, Chen A, Goldie A, Mirhoseini A, McKinnon C, et\u00a0al (2022) Constitutional AI: Harmlessness from AI feedback. arXiv prepr. arxiv:2212.08073"},{"key":"11389_CR11","unstructured":"Barrantes M, Herudek B, Wang R (2020) Adversarial nli for factual correctness in text summarisation models. arXiv prepr. arxiv:2005.11739"},{"key":"11389_CR12","doi-asserted-by":"crossref","unstructured":"Bassani E, Sanchez I (2024) Guardbench: A large-scale benchmark for guardrail models. In: Proceedings of the 2024 conference on empirical methods in natural language processing, pp. 18393\u201318409","DOI":"10.18653\/v1\/2024.emnlp-main.1022"},{"issue":"1","key":"11389_CR13","doi-asserted-by":"publisher","first-page":"291","DOI":"10.1038\/s41467-023-44371-z","volume":"15","author":"PRAS Bassi","year":"2024","unstructured":"Bassi PRAS, Dertkigil SSJ, Cavalli A (2024) Improving deep neural network generalization and robustness to background bias via layer-wise relevance propagation optimization. Nat Commun 15(1):291. https:\/\/doi.org\/10.1038\/s41467-023-44371-z","journal-title":"Nat Commun"},{"key":"11389_CR14","doi-asserted-by":"publisher","first-page":"1946","DOI":"10.1145\/3591300","volume":"7","author":"L Beurer-Kellner","year":"2023","unstructured":"Beurer-Kellner L, Fischer M, Vechev M (2023) Prompting is programming: a query language for large language models. Proceed ACM Program Lang 7:1946\u20131969","journal-title":"Proceed ACM Program Lang"},{"key":"11389_CR15","unstructured":"Beurer-Kellner L, Fischer M, Vechev M (2023) Lmql chat: Scripted chatbot development"},{"key":"11389_CR16","unstructured":"Bianchi F, Suzgun M, Attanasio G, R\u00f6ttger P, Jurafsky D, Hashimoto T, Zou J (2023) Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions. arXiv prepr. arxiv:2309.07875"},{"issue":"5","key":"11389_CR17","doi-asserted-by":"publisher","first-page":"277","DOI":"10.1038\/s42254-023-00581-4","volume":"5","author":"A Birhane","year":"2023","unstructured":"Birhane A, Kasirzadeh A, Leslie D, Wachter S (2023) Science in the age of large language models. Nat Rev Phys 5(5):277\u2013280. https:\/\/doi.org\/10.1038\/s42254-023-00581-4","journal-title":"Nat Rev Phys"},{"key":"11389_CR18","doi-asserted-by":"publisher","unstructured":"Blodgett SL, Barocas S, III HD, Wallach HM (2020) Language (technology) is power: A critical survey of \"Bias\" in NLP. In: Jurafsky, D., Chai, J., Schluter, N., Tetreault, J.R. (eds.) Proc. 58th Annu. Meet. Assoc. Comput. Linguist., pp. 5454\u20135476. Association for Computational Linguistics. https:\/\/doi.org\/10.18653\/V1\/2020.ACL-MAIN.485","DOI":"10.18653\/V1\/2020.ACL-MAIN.485"},{"key":"11389_CR19","unstructured":"Bodhankar A (2024) Content Moderation and Safety Checks with NVIDIA NeMo Guardrails. https:\/\/developer.nvidia.com\/blog\/content-moderation-and-safety-checks-with-nvidia-nemo-guardrails\/?utm_source=chatgpt.coms"},{"key":"11389_CR20","unstructured":"Botsihhin G, Boccaccia L (2025) Enhancing LLM Capabilities with NeMo Guardrails on Amazon SageMaker JumpStart. https:\/\/aws.amazon.com\/cn\/blogs\/machine-learning\/enhancing-llm-capabilities-with-nemo-guardrails-on-amazon-sagemaker-jumpstart\/?utm_source=chatgpt.com"},{"issue":"12","key":"11389_CR21","doi-asserted-by":"publisher","first-page":"0188418","DOI":"10.1371\/journal.pone.0188418","volume":"12","author":"SL Brand","year":"2017","unstructured":"Brand SL, Thompson Coon J, Fleming LE, Carroll L, Bethel A, Wyatt K (2017) Whole-system approaches to improving the health and wellbeing of healthcare workers: a systematic review. PLoS ONE 12(12):0188418. https:\/\/doi.org\/10.1371\/journal.pone.0188418","journal-title":"PLoS ONE"},{"key":"11389_CR22","unstructured":"Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, Agarwal S, Herbert-Voss A, Krueger G, Henighan T, Child R, Ramesh A, Ziegler DM, Wu J, Winter C, Hesse C, Chen M, Sigler E, Litwin M, Gray S, Chess B, Clark J, Berner C, McCandlish S, Radford A, Sutskever I, Amodei D (2020) Language models are few-shot learners. In: Proc. 34th Int. Conf. Neural Inf. Process. Syst. NIPS\u201920. Curran Associates Inc., Red Hook, NY, USA"},{"key":"11389_CR23","first-page":"37068","volume":"35","author":"X Cai","year":"2022","unstructured":"Cai X, Xu H, Xu S, Zhang Y et al (2022) Badprompt: backdoor attacks on continuous prompts. NeurIPS 35:37068\u201337080","journal-title":"NeurIPS"},{"key":"11389_CR24","unstructured":"Cao B, Cao Y, Lin L, Chen J (2023) Defending against alignment-breaking attacks via robustly aligned llm. arXiv prepr. arxiv:2309.14348"},{"key":"11389_CR25","doi-asserted-by":"crossref","unstructured":"Chakrabarty T, Laban P, Agarwal D, Muresan S, Wu C-S (2023) Art or artifice? large language models and the false promise of creativity. arXiv prepr. arxiv:2309.14556","DOI":"10.1145\/3613904.3642731"},{"key":"11389_CR26","unstructured":"Chao P, Robey A, Dobriban E, Hassani H, Pappas GJ, Wong E (2023) Jailbreaking black box large language models in twenty queries. arXiv prepr. arxiv:2310.08419"},{"key":"11389_CR27","unstructured":"Cheng Q, Sun T, Zhang W, Wang S, Liu X, Zhang M, He J, Huang M, Yin Z, Chen K, et\u00a0al (2023) Evaluating hallucinations in chinese large language models. arXiv prepr. arxiv:2310.03368"},{"key":"11389_CR28","doi-asserted-by":"crossref","unstructured":"Chen B, Paliwal A, Yan Q (2023) Jailbreaker in jail: Moving target defense for large language models. In: Proc. 10th ACM Workshop Mov. Target Def., pp. 29\u201332","DOI":"10.1145\/3605760.3623764"},{"key":"11389_CR29","unstructured":"Chen Z, Pinto F, Pan M, Li B. Safewatch: An efficient safety-policy following video guardrail model with transparent explanations. In: The thirteenth international conference on learning representations"},{"key":"11389_CR30","doi-asserted-by":"crossref","unstructured":"Chen X, Salem A, Chen D, Backes M, Ma S, Shen Q, Wu Z, Zhang Y (2021) Badnl: Backdoor attacks against nlp models with semantic-preserving improvements. In: Proc. 37th Annu. Comput. Secur. Appl. Conf., pp. 554\u2013569","DOI":"10.1145\/3485832.3485837"},{"key":"11389_CR31","doi-asserted-by":"crossref","unstructured":"Chen X, Tang S, Zhu R, Yan S, Jin L, Wang Z, Su L, Wang X, Tang H (2023) The janus interface: How fine-tuning in large language models amplifies the privacy risks. arXiv prepr. arxiv:2310.15469","DOI":"10.1145\/3658644.3690325"},{"key":"11389_CR32","unstructured":"Chen L, Zaharia M, Zou J (2023) How is ChatGPT\u2019s behavior changing over time? arXiv prepr. arxiv:2307.09009"},{"key":"11389_CR33","unstructured":"Chern I, Chern S, Chen S, Yuan W, Feng K, Zhou C, He J, Neubig G, Liu P, et\u00a0al (2023) FacTool: Factuality detection in generative AI\u2013A tool augmented framework for multi-task and multi-domain scenarios. arXiv prepr. arxiv:2307.13528"},{"key":"11389_CR34","unstructured":"Cohen J, Rosenfeld E, Kolter Z (2019) Certified adversarial robustness via randomized smoothing. In: 36th Int. Conf. Mach. Learn. (ICML 2019), pp. 1310\u20131320. PMLR"},{"issue":"4","key":"11389_CR35","doi-asserted-by":"publisher","first-page":"1058","DOI":"10.2337\/dc10-1145","volume":"34","author":"BF Crabtree","year":"2011","unstructured":"Crabtree BF, Miller WL, Stange KC (2011) The chronic care model and diabetes management in US primary care settings: a systematic review. Diabetes Care 34(4):1058\u20131063. https:\/\/doi.org\/10.2337\/dc10-1145","journal-title":"Diabetes Care"},{"key":"11389_CR36","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li L-J, Li K, Fei-Fei L (2009) Imagenet: A large-scale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248\u2013255. IEEE","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"11389_CR37","doi-asserted-by":"crossref","unstructured":"Deng B, Wang W, Feng F, Deng Y, Wang Q, He X (2023) Attack prompt generation for red teaming and defending large language models. arXiv prepr. arxiv:2310.12505","DOI":"10.18653\/v1\/2023.findings-emnlp.143"},{"key":"11389_CR38","unstructured":"Deng Y, Zhang W, Pan SJ, Bing L (2023) Multilingual jailbreak challenges in large language models. In: 12th Int. Conf. Learn. Represent. (ICLR 2024)"},{"key":"11389_CR39","doi-asserted-by":"crossref","unstructured":"Deshpande A, Murahari V, Rajpurohit T, Kalyan A, Narasimhan K (2023) Toxicity in chatgpt: Analyzing persona-assigned language models. In: Bouamor, H., Pino, J., Bali, K. (eds.) Find. Assoc. Comput. Linguist.: EMNLP 2023, pp. 1236\u20131270. Association for Computational Linguistics, ???","DOI":"10.18653\/v1\/2023.findings-emnlp.88"},{"key":"11389_CR40","doi-asserted-by":"publisher","unstructured":"Devlin J, Chang M-W, Lee K, Toutanova K (2019) BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proc. 2019 Conf. N. Am. Chapter Assoc. Comput. Linguist.: Hum. Lang. Technol., pp. 4171\u20134186. Association for Computational Linguistics, Minneapolis, Minnesota. https:\/\/doi.org\/10.18653\/v1\/N19-1423","DOI":"10.18653\/v1\/N19-1423"},{"key":"11389_CR41","doi-asserted-by":"crossref","unstructured":"Dinan E, Humeau S, Chintagunta B, Weston J (2019) Build it break it fix it for dialogue safety: Robustness from adversarial human attack. In: Proc. 2019 Conf. Empir. Methods Nat. Lang. Process. 9th Int. Jt. Conf. Nat. Lang. Process. (EMNLP-IJCNLP), pp. 4537\u20134546","DOI":"10.18653\/v1\/D19-1461"},{"key":"11389_CR42","unstructured":"Ding P, Kuang J, Ma D, Cao X, Xian Y, Chen J, Huang S (2023) A wolf in sheep\u2019s clothing: Generalized nested jailbreak prompts can fool large language models easily. arXiv prepr. arxiv:2311.08268"},{"issue":"4","key":"11389_CR43","doi-asserted-by":"publisher","first-page":"754","DOI":"10.1111\/isj.12370","volume":"32","author":"M Dolata","year":"2022","unstructured":"Dolata M, Feuerriegel S, Schwabe G (2022) A sociotechnical view of algorithmic fairness. Inf Syst J 32(4):754\u2013818","journal-title":"Inf Syst J"},{"key":"11389_CR44","unstructured":"Dong Y, Mu R, Jin G, Qi Y, Hu J, Zhao X, Meng J, Ruan W, Huang X (2024) Building guardrails for large language models. In: 41st Int. Conf. Mach. Learn. (ICML 2024). PMLR"},{"key":"11389_CR45","unstructured":"Duan H, Dziedzic A, Papernot N, Boenisch F (2023) Flocks of stochastic parrots: Differentially private prompt learning for large language models. arXiv prepr. arxiv:2305.15594"},{"issue":"4","key":"11389_CR46","doi-asserted-by":"publisher","first-page":"25","DOI":"10.21659\/rupkatha.v15n4.10","volume":"15","author":"S Dwivedi","year":"2023","unstructured":"Dwivedi S, Ghosh S, Dwivedi S (2023) Breaking the bias: Gender fairness in LLMs using prompt engineering and in-context learning. Rupkatha J Interdiscip Stud Humanit 15(4):25","journal-title":"Rupkatha J Interdiscip Stud Humanit"},{"key":"11389_CR47","unstructured":"Ernst JS, Marton S, Brinkmann J, Vellasques E, Foucard D, Kraemer M, Lambert M (2023) Bias mitigation for large language models using adversarial learning. In: ECAI 2023 Workshop Fairness Bias AI"},{"issue":"05","key":"11389_CR48","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1109\/MIC.2023.3310110","volume":"27","author":"F Filgueiras","year":"2023","unstructured":"Filgueiras F, Mendonca R, Almeida V (2023) Governing artificial intelligence through a sociotechnical lens. IEEE Internet Comput 27(05):49\u201352. https:\/\/doi.org\/10.1109\/MIC.2023.3310110","journal-title":"IEEE Internet Comput"},{"key":"11389_CR49","unstructured":"Freitas BAT, Lotufo RdA (2024) Retail-gpt: leveraging retrieval augmented generation (rag) for building e-commerce chat assistants. arXiv preprint arXiv:2408.08925"},{"key":"11389_CR50","unstructured":"Gangavarapu A (2024) Enhancing guardrails for safe and secure healthcare ai. arXiv preprint arXiv:2409.17190"},{"key":"11389_CR51","unstructured":"Ganguli D, Lovitt L, Kernion J, Askell A, Bai Y, Kadavath S, Mann B, Perez E, Schiefer N, Ndousse K, et\u00a0al (2022) Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv prepr. arxiv:2209.07858"},{"key":"11389_CR52","unstructured":"Gao M, Ruan J, Sun R, Yin X, Yang S, Wan X (2023) Human-like summarization evaluation with chatgpt. arXiv prepr. arxiv:2304.02554"},{"issue":"7","key":"11389_CR53","doi-asserted-by":"publisher","first-page":"3184","DOI":"10.3390\/app11073184","volume":"11","author":"I Garrido-Mu\u00f1oz","year":"2021","unstructured":"Garrido-Mu\u00f1oz I, Montejo-R\u00e1ez A, Mart\u00ednez-Santiago F, Ure\u00f1a-L\u00f3pez LA (2021) A survey on bias in deep NLP. Appl Sci 11(7):3184","journal-title":"Appl Sci"},{"issue":"1","key":"11389_CR54","first-page":"139","volume":"45","author":"M Gaur","year":"2024","unstructured":"Gaur M, Sheth A (2024) Building trustworthy neurosymbolic ai systems: Consistency, reliability, explainability, and safety. AI Mag 45(1):139\u2013155","journal-title":"AI Mag"},{"key":"11389_CR55","first-page":"3356","volume":"2020","author":"S Gehman","year":"2020","unstructured":"Gehman S, Gururangan S, Sap M, Choi Y, Smith NA (2020) Realtoxicityprompts: evaluating neural toxic degeneration in language models. Find Assoc Comput Ling EMNLP 2020:3356\u20133369","journal-title":"Find Assoc Comput Ling EMNLP"},{"key":"11389_CR56","doi-asserted-by":"publisher","unstructured":"Gehman S, Gururangan S, Sap M, Choi Y, Smith NA (2020) RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In: Cohn, T., He, Y., Liu, Y. (eds.) Find. Assoc. Comput. Linguist.: EMNLP 2020, pp. 3356\u20133369. Association for Computational Linguistics, Online . https:\/\/doi.org\/10.18653\/v1\/2020.findings-emnlp.301","DOI":"10.18653\/v1\/2020.findings-emnlp.301"},{"key":"11389_CR57","unstructured":"Geisler S, Wollschl\u00e4ger T, Abdalla MHI, Gasteiger J, G\u00fcnnemann S (2024) Attacking large language models with projected gradient descent. arXiv prepr. arxiv:2402.09154"},{"key":"11389_CR58","unstructured":"Ge S, Zhou C, Hou R, Khabsa M, Wang Y-C, Wang Q, Han J, Mao Y (2023) Mart: improving llm safety with multi-round automatic red-teaming. arXiv prepr. arxiv:2311.07689"},{"key":"11389_CR59","unstructured":"Glukhov D, Shumailov I, Gal Y, Papernot N, Papyan V (2023) Llm censorship: A machine learning challenge or a computer security problem? arXiv prepr. arxiv:2307.10719"},{"key":"11389_CR60","doi-asserted-by":"crossref","unstructured":"Goodrich B, Rao V, Liu PJ, Saleh M (2019) Assessing the factual accuracy of generated text. In: Proc. 25th ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., pp. 166\u2013175","DOI":"10.1145\/3292500.3330955"},{"key":"11389_CR61","doi-asserted-by":"publisher","DOI":"10.1145\/3593042","author":"S Goyal","year":"2023","unstructured":"Goyal S, Doddapaneni S, Khapra MM, Ravindran B (2023) A survey of adversarial defenses and robustness in NLP. ACM Comput Surv. https:\/\/doi.org\/10.1145\/3593042","journal-title":"ACM Comput Surv"},{"key":"11389_CR62","doi-asserted-by":"publisher","DOI":"10.1145\/3555088","author":"N Goyal","year":"2022","unstructured":"Goyal N, Kivlichan ID, Rosen R, Vasserman L (2022) Is your toxicity my toxicity? Exploring the impact of rater identity on toxicity annotation. Proc ACM Hum Comput Interact. https:\/\/doi.org\/10.1145\/3555088","journal-title":"Proc ACM Hum Comput Interact"},{"key":"11389_CR63","doi-asserted-by":"crossref","unstructured":"Greshake K, Abdelnabi S, Mishra S, Endres C, Holz T, Fritz M (2023) Not what you\u2019ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In: Proc. 16th ACM Workshop Artif. Intell. Secur., pp. 79\u201390","DOI":"10.1145\/3605764.3623985"},{"key":"11389_CR64","unstructured":"Guo Z, Jin R, Liu C, Huang Y, Shi D, Supryadi Yu L, Liu Y, Li J, Xiong B, Xiong D (2023) Evaluating large language models: A comprehensive survey. arXiv prepr. arxiv:2310.19736v3 [cs.CL]"},{"key":"11389_CR65","doi-asserted-by":"crossref","unstructured":"Guo C, Sablayrolles A, J\u00e9gou H, Kiela D (2021) Gradient-based adversarial attacks against text transformers. arXiv prepr. arxiv:2104.13733","DOI":"10.18653\/v1\/2021.emnlp-main.464"},{"key":"11389_CR66","unstructured":"Guo X, Yu F, Zhang H, Qin L, Hu B (2024) Cold-attack: Jailbreaking llms with stealthiness and controllability. arXiv prepr. arxiv:2402.08679"},{"key":"11389_CR67","unstructured":"Helbling A, Phute M, Hull M, Chau DH (2023) Llm self defense: By self examination, llms know they are being tricked. arXiv prepr. arxiv:2308.07308"},{"key":"11389_CR68","unstructured":"Helff L, Friedrich F, Brack M, Schramowski P, Kersting K (2024) Llavaguard: Vlm-based safeguard for vision dataset curation and safety assessment. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 8322\u20138326"},{"key":"11389_CR69","unstructured":"Hendrycks D, Gimpel K (2016) A baseline for detecting misclassified and out-of-distribution examples in neural networks. In: 4th Int. Conf. Learn. Represent. (ICLR 2016)"},{"key":"11389_CR70","unstructured":"Hosseini H, Kannan S, Zhang B, Poovendran R (2017) Deceiving google\u2019s perspective api built for detecting toxic comments. arXiv prepr. arxiv:1702.08138"},{"key":"11389_CR71","unstructured":"Huang D, Bu Q, Zhang J, Xie X, Chen J, Cui H (2023) Bias assessment and mitigation in llm-based code generation. arXiv prepr. arxiv:2309.14345"},{"key":"11389_CR72","doi-asserted-by":"crossref","unstructured":"Huang X, Ruan W, Huang W, Jin G, Dong Y, Wu C, Bensalem S, Mu R, Qi Y, Zhao X, et\u00a0al (2023) A survey of safety and trustworthiness of large language models through the lens of verification and validation. arXiv prepr. arxiv:2305.11391","DOI":"10.1007\/s10462-024-10824-0"},{"key":"11389_CR73","unstructured":"Huang L, Yu W, Ma W, Zhong W, Feng Z, Wang H, Chen Q, Peng W, Feng X, Qin B, et\u00a0al (2023) A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv prepr. arxiv:2311.05232"},{"key":"11389_CR74","unstructured":"Hu J, Dong Y, Huang X (2024) Adaptive guardrails for large language models via trust modeling and in-context learning. arXiv preprint arXiv:2408.08959"},{"key":"11389_CR75","unstructured":"Hu T, Zhou X-H (2024) Unveiling llm evaluation focused on metrics: challenges and solutions. arXiv preprint arXiv:2404.09135"},{"key":"11389_CR76","doi-asserted-by":"crossref","unstructured":"Igamberdiev T, Habernal I (2023) DP-BART for privatized text rewriting under local differential privacy. arXiv prepr. arxiv:2302.07636","DOI":"10.18653\/v1\/2023.findings-acl.874"},{"key":"11389_CR77","unstructured":"Inan H, Upasani K, Chi J, Rungta R, Iyer K, Mao Y, Tontchev M, Hu Q, Fuller B, Testuggine D, et\u00a0al (2023) Llama guard: Llm-based input-output safeguard for human-ai conversations. arxiv:2312.06674"},{"key":"11389_CR78","unstructured":"Jain N, Schwarzschild A, Wen Y, Somepalli G, Kirchenbauer J, Chiang P-y, Goldblum M, Saha A, Geiping J, Goldstein T (2023) Baseline defenses for adversarial attacks against aligned language models. arXiv prepr. arxiv:2309.00614"},{"key":"11389_CR79","unstructured":"Jang E, Gu S, Poole B (2016) Categorical reparameterization with gumbel-softmax. arXiv prepr. arxiv:1611.01144"},{"key":"11389_CR80","doi-asserted-by":"crossref","unstructured":"Jha SK, Jha S, Lincoln P, Bastian ND, Velasquez A, Ewetz R, Neema S (2023) Counterexample guided inductive synthesis using large language models and satisfiability solving. In: 2023 IEEE Mil. Commun. Conf. (MILCOM 2023), pp. 944\u2013949. IEEE, ???","DOI":"10.1109\/MILCOM58377.2023.10356332"},{"key":"11389_CR81","unstructured":"Jiang S, Chen X, Tang R (2023) Prompt packer: Deceiving llms through compositional instruction with hidden attacks. arXiv prepr. arxiv:2310.10077"},{"key":"11389_CR82","unstructured":"Jiang M, Ruan Y, Huang S, Liao S, Pitis S, Grosse RB, Ba J (2023) Calibrating language models via augmented prompt ensembles. In: ICML 2023 Workshop Chall. Deployable Gener. AI"},{"key":"11389_CR83","unstructured":"Jin H, Chen R, Zhou A, Chen J, Zhang Y, Wang H (2024) GUARD: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models. arXiv prepr. arxiv:2402.03299"},{"key":"11389_CR84","unstructured":"Jr, DM, Prabhakaran V, Kuhlberg J, Smart A, Isaac WS (2020) Extending the machine learning abstraction boundary: A complex systems approach to incorporate societal context. CoRR abs\/2006.09663 arxiv:2006.09663"},{"key":"11389_CR85","unstructured":"Kadavath S, Conerly T, Askell A, Henighan T, Drain D, Perez E, Schiefer N, Hatfield-Dodds Z, DasSarma N, Tran-Johnson E, et\u00a0al (2022) Language models (mostly) know what they know. arXiv prepr. arxiv:2207.05221"},{"key":"11389_CR86","doi-asserted-by":"crossref","unstructured":"Kang D, Li X, Stoica I, Guestrin C, Zaharia M, Hashimoto T (2023) Exploiting programmatic behavior of LLMs: Dual-use through standard security attacks. arXiv prepr. arxiv:2302.05733","DOI":"10.1109\/SPW63631.2024.00018"},{"key":"11389_CR87","unstructured":"Kaushik D, Hovy E, Lipton Z (2019) Learning the difference that makes a difference with counterfactually-augmented data. In: 7th Int. Conf. Learn. Represent. (ICLR 2019)"},{"key":"11389_CR88","unstructured":"Kirchenbauer J, Geiping J, Wen Y, Katz J, Miers I, Goldstein T (2023) A watermark for large language models. In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J. (eds.) 40th Int. Conf. Mach. Learn. (ICML 2023). Proceedings of Machine Learning Research, vol. 202, pp. 17061\u201317084. PMLR, ??? (2023-07-23\/2023-07-29)"},{"key":"11389_CR89","doi-asserted-by":"crossref","unstructured":"Kirk HR, Birhane A, Vidgen B, Derczynski L (2023) Handling and presenting harmful text in NLP research. arXiv prepr. arxiv:2204.14256v3 [cs.CL]","DOI":"10.18653\/v1\/2022.findings-emnlp.35"},{"key":"11389_CR90","unstructured":"Koh NH, Plata J, Chai J (2023) BAD: BiAs Detection for Large Language Models in the context of candidate screening. arXiv prepr. arxiv:2305.10407"},{"key":"11389_CR91","unstructured":"Koh H, Kim D, Lee M, Jung K (2024) Can LLMs recognize toxicity? Structured toxicity investigation framework and semantic-based metric. arXiv prepr. arXiv:2402,06900v2arxiv:2402.06900v2 [cs.CL]"},{"key":"11389_CR92","unstructured":"Kuhn L, Gal Y, Farquhar S (2022) Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In: 10th Int. Conf. Learn. Represent. (ICLR 2022)"},{"key":"11389_CR93","doi-asserted-by":"crossref","unstructured":"Kukreja S, Kumar T, Purohit A, Dasgupta A, Guha D (2024) A literature survey on open source large language models. In: Proceedings of the 2024 7th International Conference on Computers in Management and Business, pp. 133\u2013143","DOI":"10.1145\/3647782.3647803"},{"key":"11389_CR94","doi-asserted-by":"crossref","unstructured":"Kumar S, Balachandran V, Njoo L, Anastasopoulos A, Tsvetkov Y (2023) Language generation models can cause harm: So what can we do about it? An actionable survey. In: Proc. 17th Conf. Eur. Chapter Assoc. Comput. Linguist., pp. 3299\u20133321","DOI":"10.18653\/v1\/2023.eacl-main.241"},{"key":"11389_CR95","unstructured":"Lab ES (2023) Documentation of LMQL. https:\/\/lmql.ai"},{"key":"11389_CR96","unstructured":"Lamb LC, d\u2019Avila Garcez A, Gori M, Prates MOR, Avelar PHC, Vardi MY (2021) Graph neural networks meet neural-symbolic computing: A survey and perspective. In: Proc. 29th Int. Jt. Conf. Artif. Intell. (IJCAI 2021). IJCAI\u201920, Yokohama, Yokohama, Japan"},{"key":"11389_CR97","doi-asserted-by":"crossref","unstructured":"Lapid R, Langberg R, Sipper M (2023) Open sesame! universal black box jailbreaking of large language models. arXiv prepr. arxiv:2309.01446","DOI":"10.3390\/app14167150"},{"key":"11389_CR98","unstructured":"Liang P, Bommasani R, Lee T, Tsipras D, Soylu D, Yasunaga M, Zhang Y, Narayanan D, Wu Y, Kumar A, Newman B, Yuan B, Yan B, Zhang C, Cosgrove C, Manning CD, R\u00e9 C, Acosta-Navas D, Hudson DA, Zelikman E, Durmus E, Ladhak F, Rong F, Ren H, Yao H, Wang J, Santhanam K, Orr L, Zheng L, Yuksekgonul M, Suzgun M, Kim N, Guha N, Chatterji N, Khattab O, Henderson P, Huang Q, Chi R, Xie SM, Santurkar S, Ganguli S, Hashimoto T, Icard T, Zhang T, Chaudhary V, Wang W, Li X, Mai Y, Zhang Y, Koreeda Y Holistic evaluation of language models. arXiv prepr. arxiv:2211.09110v2 [cs.CL]"},{"key":"11389_CR99","unstructured":"Li H, Chen Y, Luo J, Kang Y, Zhang X, Hu Q, Chan C, Song Y (2023) Privacy in large language models: Attacks, defenses and future directions. arXiv prepr. arxiv:2310.10383"},{"key":"11389_CR100","doi-asserted-by":"crossref","unstructured":"Li X, Liu M, Gao S, Buntine W (2023) A survey on out-of-distribution evaluation of neural NLP models. In: Proc. 32th Int. Jt. Conf. Artif. Intell. (IJCAI 2023), pp. 6683\u20136691","DOI":"10.24963\/ijcai.2023\/749"},{"key":"11389_CR101","unstructured":"Limisiewicz T, Mare\u010dek D, Musil T (2023) Debiasing algorithm through model adaptation. arXiv prepr. arxiv:2310.18913"},{"key":"11389_CR102","unstructured":"Lin S, Hilton J, Evans O (2022) Teaching models to express their uncertainty in words. arXiv prepr. arxiv:2205.14334"},{"key":"11389_CR103","unstructured":"Li X, Tramer F, Liang P, Hashimoto T (2022) Large language models can be strong differentially private learners. In: 10th Int. Conf. Learn. Represent. (ICLR 2022)"},{"key":"11389_CR104","unstructured":"Liu Y, Deng G, Li Y, Wang K, Zhang T, Liu Y, Wang H, Zheng Y, Liu Y (2023) Prompt injection attack against LLM-integrated applications. arXiv prepr. arxiv:2306.05499"},{"key":"11389_CR105","unstructured":"Liu A, Pan L, Hu X, Meng S, Wen L (2024) A semantic invariant robust watermark for large language models. In: 12th Int. Conf. Learn. Represent. (ICLR 2024)"},{"key":"11389_CR106","unstructured":"Liu X, Xu N, Chen M, Xiao C (2023) Autodan: Generating stealthy jailbreak prompts on aligned large language models. arXiv prepr. arxiv:2310.04451"},{"key":"11389_CR107","unstructured":"Liu Y, Zeng X, Meng F, Zhou J (2023) Instruction position matters in sequence generation with large language models. arXiv prepr. arxiv:2308.12097"},{"key":"11389_CR108","unstructured":"Liu T, Zhang Y, Zhao Z, Dong Y, Meng G, Chen K (2024) Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction. arXiv prepr. arxiv:2402.18104"},{"key":"11389_CR109","unstructured":"Li Y, Wei F, Zhao J, Zhang C, Zhang H (2023) Rain: Your language models can align themselves without finetuning. arXiv prepr. arxiv:2309.07124"},{"key":"11389_CR110","unstructured":"Li Z, Zhang S, Zhao H, Yang Y, Yang D (2023) Batgpt: A bidirectional autoregessive talker from generative pre-trained transformer. arXiv prepr. arxiv:2307.00360"},{"key":"11389_CR111","unstructured":"Li X, Zhou Z, Zhu J, Yao J, Liu T, Han B (2023) Deepinception: Hypnotize large language model to be jailbreaker. arXiv prepr. arxiv:2311.03191"},{"key":"11389_CR112","unstructured":"Lou R, Zhang K, Yin W Is prompt all you need? no. A comprehensive and broader view of instruction learning. arXiv prepr. arxiv:2303.10475"},{"key":"11389_CR113","unstructured":"Luo Z, Xie Q, Ananiadou S (2023) Chatgpt as a factual inconsistency evaluator for abstractive text summarization. arXiv prepr. arxiv:2303.15621"},{"key":"11389_CR114","unstructured":"Lv H, Wang X, Zhang Y, Huang C, Dou S, Ye J, Gui T, Zhang Q, Huang X (2024) CodeChameleon: Personalized encryption framework for jailbreaking large language models. arXiv prepr. arxiv:2402.16717"},{"key":"11389_CR115","unstructured":"Lyu C, Xu J, Wang L (2023) New trends in machine translation using large language models: Case examples with chatgpt. arXiv prepr. arxiv:2305.01181"},{"key":"11389_CR116","unstructured":"Malik A (2023) Evaluating large language models through gender and racial stereotypes. arXiv prepr. arxiv:2311.14788"},{"key":"11389_CR117","doi-asserted-by":"crossref","unstructured":"Manakul P, Liusie A, Gales MJ (2023) Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. arXiv prepr. arxiv:2303.08896","DOI":"10.18653\/v1\/2023.emnlp-main.557"},{"key":"11389_CR118","doi-asserted-by":"crossref","unstructured":"Mangaokar N, Hooda A, Choi J, Chandrashekaran S, Fawaz K, Jha S, Prakash A (2024) PRP: Propagating universal perturbations to attack large language model guard-rails. arXiv prepr. arxiv:2402.15911","DOI":"10.18653\/v1\/2024.acl-long.591"},{"key":"11389_CR119","doi-asserted-by":"publisher","unstructured":"Ma X, Sap M, Rashkin H, Choi Y (2020) PowerTransformer: Unsupervised controllable revision for biased language correction. In: Webber, B., Cohn, T., He, Y., Liu, Y. (eds.) Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 7426\u20137441. Association for Computational Linguistics, Online. https:\/\/doi.org\/10.18653\/v1\/2020.emnlp-main.602","DOI":"10.18653\/v1\/2020.emnlp-main.602"},{"key":"11389_CR120","unstructured":"Maus N, Chao P, Wong E, Gardner JR (2023) Black box adversarial prompting for foundation models. In: 2nd Workshop New Front. Advers. Mach. Learn"},{"key":"11389_CR121","unstructured":"Mbiazi D, Bhange M, Babaei M, Sheth I, Kenfack PJ (2023) Survey on AI ethics: A socio-technical perspective. arXiv prepr. arxiv:2311.17228 [cs.CY]"},{"issue":"1","key":"11389_CR122","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1038\/s41746-024-01083-y","volume":"7","author":"N Mehandru","year":"2024","unstructured":"Mehandru N, Miao BY, Almaraz ER, Sushil M, Butte AJ, Alaa A (2024) Evaluating large language models as agents in the clinic. NPJ Dig Med 7(1):84","journal-title":"NPJ Dig Med"},{"key":"11389_CR123","unstructured":"Mehrotra A, Zampetakis M, Kassianik P, Nelson B, Anderson H, Singer Y, Karbasi A (2023) Tree of attacks: Jailbreaking black-box llms automatically. arXiv prepr. arxiv:2312.02119"},{"key":"11389_CR124","unstructured":"Menini S, Aprosio AP, Tonelli S (2021) Abuse is contextual, what about NLP? The role of context in abusive language annotation and detection. arXiv prepr. arxiv:2103.14916 [cs.CL]"},{"issue":"4","key":"11389_CR125","doi-asserted-by":"publisher","first-page":"371","DOI":"10.1037\/h0040525","volume":"67","author":"S Milgram","year":"1963","unstructured":"Milgram S (1963) Behavioral study of obedience. J Abnorm Soc Psychol 67(4):371","journal-title":"J Abnorm Soc Psychol"},{"issue":"6","key":"11389_CR126","doi-asserted-by":"crossref","first-page":"617","DOI":"10.2307\/2064025","volume":"4","author":"S Milgram","year":"1975","unstructured":"Milgram S (1975) Obedience to authority: an experimental view. Contemp Sociol 4(6):617","journal-title":"Contemp Sociol"},{"key":"11389_CR127","doi-asserted-by":"crossref","unstructured":"Min S, Krishna K, Lyu X, Lewis M, Yih W-t, Koh PW, Iyyer M, Zettlemoyer L, Hajishirzi H (2023) Factscore: Fine-grained atomic evaluation of factual precision in long form text generation. arXiv prepr. arxiv:2305.14251","DOI":"10.18653\/v1\/2023.emnlp-main.741"},{"key":"11389_CR128","first-page":"29468","volume":"35","author":"F Mireshghallah","year":"2022","unstructured":"Mireshghallah F, Backurs A, Inan HA, Wutschitz L, Kulkarni J (2022) Differentially private model compression. NeurIPS 35:29468\u201329483","journal-title":"NeurIPS"},{"key":"11389_CR129","doi-asserted-by":"crossref","unstructured":"Mishra A, Patel D, Vijayakumar A, Li XL, Kapanipathi P, Talamadupula K (2021) Looking beyond sentence-level natural language inference for question answering and text summarization. In: Proc. 2021 Conf. N. Am. Chapter Assoc. Comput. Linguist.: Hum. Lang. Technol., pp. 1322\u20131336","DOI":"10.18653\/v1\/2021.naacl-main.104"},{"key":"11389_CR130","unstructured":"Mohapatra J, Ko C-Y, Weng L, Chen P-Y, Liu S, Daniel L. Hidden cost of randomized smoothing. In: Banerjee, A., Fukumizu, K. (eds.) Proc. 24th Int. Conf. Artif. Intell. Stat. Proceedings of Machine Learning Research, vol. 130, pp. 4033\u20134041. PMLR, (2021-04-13\/2021-04-15)"},{"key":"11389_CR131","doi-asserted-by":"crossref","unstructured":"Motoki F, Pinho Neto V, Rodrigues V (2023) More human than human: Measuring chatgpt political bias. Available SSRN 4372349","DOI":"10.1007\/s11127-023-01097-2"},{"key":"11389_CR132","unstructured":"Naihin S, Atkinson D, Green M, Hamadi M, Swift C, Schonholtz D, Kalai AT, Bau D (2023) Testing language model agents safely in the wild. arXiv prepr. arxiv:2311.10538"},{"key":"11389_CR133","doi-asserted-by":"crossref","unstructured":"Nan F, Nallapati R, Wang Z, Santos CN, Zhu H, Zhang D, McKeown K, Xiang B (2021) Entity-level factual consistency of abstractive text summarization. arXiv prepr. arxiv:2102.09130","DOI":"10.18653\/v1\/2021.eacl-main.235"},{"key":"11389_CR134","doi-asserted-by":"crossref","unstructured":"Narayanan A, Kapoor S (2023) Is GPT-4 getting worse over time? AI Snake Oil","DOI":"10.1515\/9780691249643"},{"key":"11389_CR135","unstructured":"Narayanan D, Shoeybi M, Casper J, LeGresley P, Patwary M, Korthikanti V, Vainbrand D, Catanzaro B (2021) Scaling language model training to a trillion parameters using megatron. arXiv prepr. arxiv:2104.04473v5"},{"key":"11389_CR136","doi-asserted-by":"publisher","DOI":"10.1109\/ISAP.2005.1599245","author":"P Ngatchou","year":"2005","unstructured":"Ngatchou P, Zarei A, El-Sharkawi A (2005) Pareto multi objective optimization. Proc 13th Int Conf Intell Syst Appl Power Syst. https:\/\/doi.org\/10.1109\/ISAP.2005.1599245","journal-title":"Proc 13th Int Conf Intell Syst Appl Power Syst"},{"key":"11389_CR137","unstructured":"Nguyen TT, Huynh TT, Nguyen PL, Liew AW-C, Yin H, Nguyen QVH (2022) A survey of machine unlearning. arXiv prepr. arxiv:2209.02299v5 [cs.LG]"},{"key":"11389_CR138","unstructured":"Nvidia: Colang (2023)"},{"key":"11389_CR139","unstructured":"Oba D, Kaneko M, Bollegala D (2023) In-contextual bias suppression for large language models. arXiv prepr. arxiv:2309.07251"},{"key":"11389_CR140","unstructured":"Oh S, Jin Y, Sharma M, Kim D, Ma E, Verma G, Kumar S (2024) Uniguard: Towards universal safety guardrails for jailbreak attacks on multimodal large language models. arXiv preprint arXiv:2411.01703"},{"key":"11389_CR141","doi-asserted-by":"publisher","unstructured":"Oh C, Won H, So J, Kim T, Kim Y, Choi H, Song K (2022) Learning fair representation via distributional contrastive disentanglement. In: Zhang, A., Rangwala, H. (eds.) 28th ACM SIGKDD Conf. Knowl. Discov. Data Min. (KDD 2022), pp. 1295\u20131305. ACM. https:\/\/doi.org\/10.1145\/3534678.3539232","DOI":"10.1145\/3534678.3539232"},{"key":"11389_CR142","unstructured":"OpenAI: GPT-4 technical report. arXiv e-prints (2023) arxiv:2303.08774"},{"key":"11389_CR143","unstructured":"Oppermann A (2023) What Is the V-model in Software Development?"},{"key":"11389_CR144","first-page":"27730","volume":"35","author":"L Ouyang","year":"2022","unstructured":"Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A et al (2022) Training language models to follow instructions with human feedback. NeurIPS 35:27730\u201327744","journal-title":"NeurIPS"},{"key":"11389_CR145","unstructured":"Ovalle A, Mehrabi N, Goyal P, Dhamala J, Chang K-W, Zemel RS, Galstyan A, Gupta R (2023) Are you talking to [\u2019xem\u2019] or [\u2019x\u2019, \u2019em\u2019]? On tokenization and addressing misgendering in LLMs with pronoun tokenization parity. CoRR abs\/2312.11779"},{"key":"11389_CR146","doi-asserted-by":"crossref","unstructured":"Ozdayi MS, Peris C, Fitzgerald J, Dupuy C, Majmudar J, Khan H, Parikh R, Gupta R (2023) Controlling the extraction of memorized data from large language models via prompt-tuning. arXiv prepr. arxiv:2305.11759","DOI":"10.18653\/v1\/2023.acl-short.129"},{"key":"11389_CR147","doi-asserted-by":"crossref","unstructured":"Paduraru C, Patilea C, Stefanescu A (2024) Cyberguardian: An interactive assistant for cybersecurity specialists using large language models. In: Proceedings of the 19th International Conference on Software Technologies (ICSOFT 2024), Dijon, France, pp. 8\u201310","DOI":"10.5220\/0012811700003753"},{"key":"11389_CR148","doi-asserted-by":"publisher","DOI":"10.1007\/s10439-023-03306-x","author":"S Pal","year":"2023","unstructured":"Pal S, Bhattacharya M, Lee S-S, Chakraborty C (2023) A domain-specific next-generation large language model (LLM) or ChatGPT is required for biomedical engineering and research. Ann Biomed Eng. https:\/\/doi.org\/10.1007\/s10439-023-03306-x","journal-title":"Ann Biomed Eng"},{"key":"11389_CR149","unstructured":"Pantha N, Ramasubramanian M, Gurung I, Maskey M, Ramachandran R (2024) Challenges in guardrailing large language models for science. arXiv preprint arXiv:2411.08181"},{"key":"11389_CR150","doi-asserted-by":"publisher","unstructured":"Pavlopoulos J, Sorensen J, Dixon L, Thain N, Androutsopoulos I (2020) Toxicity detection: Does context really matter? In: Jurafsky, D., Chai, J., Schluter, N., Tetreault, J. (eds.) Proc. 58th Annu. Meet. Assoc. Comput. Linguist., pp. 4296\u20134305. Association for Computational Linguistics, Online. https:\/\/doi.org\/10.18653\/v1\/2020.acl-main.396","DOI":"10.18653\/v1\/2020.acl-main.396"},{"key":"11389_CR151","unstructured":"Pelrine K, Taufeeque M, Zaj\u0105c M, McLean E, Gleave A (2023) Exploiting novel gpt-4 apis. arXiv prepr. arxiv:2312.14302"},{"key":"11389_CR152","doi-asserted-by":"crossref","unstructured":"Perez E, Huang S, Song F, Cai T, Ring R, Aslanides J, Glaese A, McAleese N, Irving G (2022) Red teaming language models with language models. arXiv prepr. arxiv:2202.03286","DOI":"10.18653\/v1\/2022.emnlp-main.225"},{"key":"11389_CR153","unstructured":"Perez F, Ribeiro I (2022) Ignore previous prompt: Attack techniques for language models. In: NeuIPS Workshop Mach. Learn. Saf"},{"key":"11389_CR154","doi-asserted-by":"crossref","unstructured":"Peri SDB, Santhanalakshmi S, Radha R (2024) Chatbot to chat with medical books using retrieval-augmented generation model. In: 2024 IEEE North Karnataka Subsection Flagship International Conference (NKCon), pp. 1\u20135. IEEE","DOI":"10.1109\/NKCon62728.2024.10774900"},{"key":"11389_CR155","first-page":"10018","volume":"2024","author":"R Peri","year":"2024","unstructured":"Peri R, Jayanthi SM, Ronanki S, Bhatia A, Mundnich K, Dingliwal S, Das N, Hou Z, Vishnubhotla HGS et al (2024) Speechguard: exploring the adversarial robustness of multi-modal large language models. Find Assoc Comput Ling ACL 2024:10018\u201310035","journal-title":"Find Assoc Comput Ling ACL"},{"key":"11389_CR156","doi-asserted-by":"crossref","unstructured":"Plant R, Giuffrida V, Gkatzia D (2022) You are what you write: Preserving privacy in the era of large language models. arXiv prepr. arxiv:2204.09391","DOI":"10.2139\/ssrn.4417900"},{"key":"11389_CR157","doi-asserted-by":"publisher","unstructured":"Qian R, Ross C, Fernandes J, Smith EM, Kiela D, Williams A (2022) Perturbation augmentation for fairer NLP. In: Goldberg, Y., Kozareva, Z., Zhang, Y. (eds.) Proc. 2022 Conf. Empir. Methods Nat. Lang. Process. (EMNLP 2022), pp. 9496\u20139521. Association for Computational Linguistics. https:\/\/doi.org\/10.18653\/V1\/2022.EMNLP-MAIN.646","DOI":"10.18653\/V1\/2022.EMNLP-MAIN.646"},{"key":"11389_CR158","unstructured":"Qi X, Wei B, Carlini N, Huang Y, Xie T, He L, Jagielski M, Nasr M, Mittal P, Henderson P (2024) On evaluating the durability of safeguards for open-weight llms. arXiv preprint arXiv:2412.07097"},{"key":"11389_CR159","unstructured":"Qi X, Zeng Y, Xie T, Chen P-Y, Jia R, Mittal P, Henderson P (2024) Fine-tuning aligned language models compromises safety, even when users do not intend to! In: 12th Int. Conf. Learn. Represent. (ICLR 2024)"},{"key":"11389_CR160","doi-asserted-by":"crossref","unstructured":"Rahman MA, Alqahtani L, Albooq A, Ainousah A (2024) A survey on security and privacy of large multimodal deep learning models: Teaching and learning perspective. In: 21st Learn. Technol. Conf. (L &T 2024), pp. 13\u201318. IEEE, ???","DOI":"10.1109\/LT60077.2024.10469434"},{"issue":"1","key":"11389_CR161","doi-asserted-by":"publisher","first-page":"43","DOI":"10.4236\/jsea.2024.171003","volume":"17","author":"P Rai","year":"2024","unstructured":"Rai P, Sood S, Madisetti VK, Bahga A (2024) Guardian: a multi-tiered defense architecture for thwarting prompt injection attacks on llms. J Softw Eng Appl 17(1):43\u201368","journal-title":"J Softw Eng Appl"},{"key":"11389_CR162","unstructured":"Rajpal S (2023) Guardrails AI"},{"key":"11389_CR163","first-page":"71095","volume":"36","author":"A Rame","year":"2023","unstructured":"Rame A, Couairon G, Dancette C, Gaya J-B, Shukor M, Soulier L, Cord M (2023) Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards. Adv Neural Inf Process Syst 36:71095\u201371134","journal-title":"Adv Neural Inf Process Syst"},{"key":"11389_CR164","doi-asserted-by":"crossref","unstructured":"Ramezani A, Xu Y (2023) Knowledge of cultural moral norms in large language models. arXiv prepr. arxiv:2306.01857","DOI":"10.18653\/v1\/2023.acl-long.26"},{"key":"11389_CR165","doi-asserted-by":"crossref","unstructured":"Ranaldi L, Ruzzetti ES, Venditti D, Onorati D, Zanzotto FM (2023) A trip towards fairness: Bias and de-biasing in large language models. arXiv prepr. arxiv:2305.13862","DOI":"10.18653\/v1\/2024.starsem-1.30"},{"key":"11389_CR166","doi-asserted-by":"crossref","unstructured":"Rebedea T, Dinu R, Sreedhar M, Parisien C, Cohen J (2023) Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails. arXiv prepr. arxiv:2310.10501","DOI":"10.18653\/v1\/2023.emnlp-demo.40"},{"key":"11389_CR167","unstructured":"Ren AZ, Dixit A, Bodrova A, Singh S, Tu S, Brown N, Xu P, Takayama L, Xia F, Varley J, et\u00a0al (2023) Robots that ask for help: Uncertainty alignment for large language model planners. In: 2023 Conf. Robot Learn., pp. 661\u2013682. PMLR, ???"},{"key":"11389_CR168","unstructured":"Ren J, Luo J, Zhao Y, Krishna K, Saleh M, Lakshminarayanan B, Liu PJ (2023) Out-of-distribution detection and selective generation for conditional language models. In: 11th Int. Conf. Learn. Represent. (ICLR 2023)"},{"key":"11389_CR169","unstructured":"Robey A, Wong E, Hassani H, Pappas GJ (2023) Smoothllm: Defending large language models against jailbreaking attacks. arXiv prepr. arxiv:2310.03684"},{"key":"11389_CR170","doi-asserted-by":"publisher","unstructured":"Rosenblatt L, Piedras L, Wilkins J (2022) Critical perspectives: A benchmark revealing pitfalls in PerspectiveAPI. In: Biester, L., Demszky, D., Jin, Z., Sachan, M., Tetreault, J., Wilson, S., Xiao, L., Zhao, J. (eds.) Proc. 2nd Workshop NLP Posit. Impact (NLP4PI), pp. 15\u201324. Association for Computational Linguistics, Abu Dhabi, United Arab Emirates (Hybrid) . https:\/\/doi.org\/10.18653\/v1\/2022.nlp4pi-1.2","DOI":"10.18653\/v1\/2022.nlp4pi-1.2"},{"key":"11389_CR171","doi-asserted-by":"crossref","unstructured":"R\u00f6ttger P, Kirk HR, Vidgen B, Attanasio G, Bianchi F, Hovy D (2023) Xstest: A test suite for identifying exaggerated safety behaviours in large language models. arXiv prepr. arxiv:2308.01263","DOI":"10.18653\/v1\/2024.naacl-long.301"},{"key":"11389_CR172","unstructured":"Ruan Y, Dong H, Wang A, Pitis S, Zhou Y, Ba J, Dubois Y, Maddison CJ, Hashimoto T (2023) Identifying the risks of lm agents with an lm-emulated sandbox. arXiv prepr. arxiv:2309.15817"},{"key":"11389_CR173","doi-asserted-by":"publisher","unstructured":"Sap M, Gabriel S, Qin L, Jurafsky D, Smith NA, Choi Y (2020) Social bias frames: Reasoning about social and power implications of language. In: Jurafsky, D., Chai, J., Schluter, N., Tetreault, J. (eds.) Proc. 58th Annu. Meet. Assoc. Comput. Linguist., pp. 5477\u20135490. Association for Computational Linguistics, Online . https:\/\/doi.org\/10.18653\/v1\/2020.acl-main.486","DOI":"10.18653\/v1\/2020.acl-main.486"},{"key":"11389_CR174","doi-asserted-by":"crossref","unstructured":"Sarker IH (2024) LLM potentiality and awareness: A position paper from the perspective of trustworthy and responsible AI modeling. Authorea Prepr","DOI":"10.36227\/techrxiv.170905626.67078570\/v1"},{"key":"11389_CR175","doi-asserted-by":"publisher","DOI":"10.6028\/NIST.SP.1270","author":"R Schwartz","year":"2022","unstructured":"Schwartz R, Vassilev A, Greene K, Perine L, Burt A, Hall P (2022) Towards a standard for identifying and managing bias in artificial intelligence. Nat Inst Standards Technol Gaithersburg MD. https:\/\/doi.org\/10.6028\/NIST.SP.1270","journal-title":"Nat Inst Standards Technol Gaithersburg MD"},{"key":"11389_CR176","first-page":"9086","volume":"37","author":"L Schwinn","year":"2024","unstructured":"Schwinn L, Dobre D, Xhonneux S, Gidel G, G\u00fcnnemann S (2024) Soft prompt threats: attacking safety alignment and unlearning in open-source llms through the embedding space. Adv Neural Inf Process Syst 37:9086\u20139116","journal-title":"Adv Neural Inf Process Syst"},{"key":"11389_CR177","unstructured":"Shah MA, Sharma R, Dhamyal H, Olivier R, Shah A, Konan J, Alharthi D, Bukhari HT, Baali M, Deshmukh S, Kuhlmann M, Raj B, Singh R (2023) LoFT: Local proxy fine-tuning for improving transferability of adversarial attacks against large language model. arXiv prepr. arxiv:2310.04445v2 [cs.CL]"},{"key":"11389_CR178","doi-asserted-by":"crossref","unstructured":"Shaikh O, Zhang H, Held W, Bernstein M, Yang D (2022) On second thought, let\u2019s not think step by step! Bias and toxicity in zero-shot reasoning. arXiv prepr. arxiv:2212.08061","DOI":"10.18653\/v1\/2023.acl-long.244"},{"key":"11389_CR179","unstructured":"Shayegani E, Mamun MAA, Fu Y, Zaree P, Dong Y, Abu-Ghazaleh N (2023) Survey of vulnerabilities in large language models revealed by adversarial attacks. arXiv prepr. arxiv:2310.10844"},{"key":"11389_CR180","doi-asserted-by":"crossref","unstructured":"Shen X, Chen Z, Backes M, Shen Y, Zhang Y (2023) \" do anything now\": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. arXiv prepr. arxiv:2308.03825","DOI":"10.1145\/3658644.3670388"},{"key":"11389_CR181","unstructured":"Sheng Y, Cao S, Li D, Zhu B, Li Z, Zhuo D, Gonzalez JE, Stoica I (2023) Fairness in serving large language models. arXiv prepr. arxiv:2401.00588"},{"key":"11389_CR182","unstructured":"Sheppard B, Richter A, Cohen A, Smith EA, Kneese T, Pelletier C, Baldini I, Dong Y (2023) Subtle misogyny detection and mitigation: An expert-annotated dataset. arXiv prepr. arxiv:2311.09443"},{"key":"11389_CR183","unstructured":"Shi J, Liu Y, Zhou P, Sun L (2023) Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt. arXiv prepr. arxiv:2304.12298"},{"key":"11389_CR184","doi-asserted-by":"crossref","unstructured":"Shi W, Shea R, Chen S, Zhang C, Jia R, Yu Z (2022) Just fine-tune twice: Selective differential privacy for large language models. arXiv prepr. arxiv:2204.07667","DOI":"10.18653\/v1\/2022.emnlp-main.425"},{"key":"11389_CR185","doi-asserted-by":"crossref","unstructured":"Shuster K, Poff S, Chen M, Kiela D, Weston J (2021) Retrieval augmentation reduces hallucination in conversation. arXiv prepr. arxiv:2104.07567","DOI":"10.18653\/v1\/2021.findings-emnlp.320"},{"key":"11389_CR186","unstructured":"Shu M, Wang J, Zhu C, Geiping J, Xiao C, Goldstein T (2024) On the exploitability of instruction tuning. NeurIPS 36"},{"key":"11389_CR187","unstructured":"Simon N, Muise C (2022) TattleTale: Storytelling with planning and large language models. In: ICAPS Workshop Sched. Plan. Appl"},{"key":"11389_CR188","unstructured":"Singh S (2024) Flipkart Enhances AI Safety in E-Commerce: Implementing NVIDIA NeMo Guardrails. https:\/\/blog.flipkart.tech\/flipkart-enhances-ai-safety-in-e-commerce-implementing-nvidia-nemo-guardrails-cb2f293b29c0"},{"key":"11389_CR189","unstructured":"Singhal K, Tu T, Gottweis J, Sayres R, Wulczyn E, Hou L, Clark K, Pfohl S, Cole-Lewis H, Neal D, et\u00a0al (2023) Towards expert-level medical question answering with large language models. arXiv prepr. arxiv:2305.09617"},{"key":"11389_CR190","doi-asserted-by":"crossref","unstructured":"Song L, Shokri R, Mittal P (2019) Privacy risks of securing machine learning models against adversarial examples. In: Proc. 2019 ACM SIGSAC Conf. Comput. Commun. Secur., pp. 241\u2013257. Association for Computing Machinery, London, United Kingdom","DOI":"10.1145\/3319535.3354211"},{"key":"11389_CR191","unstructured":"Sun H, Pei J, Choi M, Jurgens D (2023) Aligning with whom? large language models have gender and racial biases in subjective nlp tasks. arXiv prepr. arxiv:2311.09730"},{"key":"11389_CR192","unstructured":"Tamirisa R, Bharathi B, Phan L, Zhou A, Gatti A, Suresh T, Lin M, Wang J, Wang R, Arel R, et\u00a0al (2024) Tamper-resistant safeguards for open-weight llms. arXiv preprint arXiv:2408.00761"},{"key":"11389_CR193","doi-asserted-by":"crossref","unstructured":"Tang X, Jin Q, Zhu K, Yuan T, Zhang Y, Zhou W, Qu M, Zhao Y, Tang J, Zhang Z, et\u00a0al (2024) Prioritizing safeguarding over autonomy: risks of LLM agents for science. arXiv prepr. arxiv:2402.04247","DOI":"10.1038\/s41467-025-63913-1"},{"key":"11389_CR194","unstructured":"Tao Y, Viberg O, Baker RS, Kizilcec RF (2023) Auditing and mitigating cultural bias in LLMs. arXiv prepr. arxiv:2311.14096"},{"key":"11389_CR195","unstructured":"team GA (2023) Tutorial of Guidance AI. https:\/\/guidance.readthedocs.io\/en\/latest\/tutorials.html"},{"key":"11389_CR196","unstructured":"Team GA (2024) use cases of Guardrails AI. https:\/\/hub.guardrailsai.com"},{"key":"11389_CR197","unstructured":"Team N (2024a) Building Blocks for Agentic AI. https:\/\/www.nvidia.com\/en-gb\/ai\/"},{"key":"11389_CR198","unstructured":"Team T (2023) Documentation of Guardrails AI. https:\/\/www.trulens.org"},{"key":"11389_CR199","unstructured":"Touvron H, Lavril T, Izacard G, Martinet X, Lachaux M-A, Lacroix T, Rozi\u00e8re B, Goyal N, Hambro E, Azhar F, et\u00a0al (2023) Llama: Open and efficient foundation language models. arXiv prepr. arxiv:2302.13971"},{"key":"11389_CR200","unstructured":"Trist EL, Bamforth KW (1957) Studies in the Quality of Life: Delivered by the Institute of Personnel Management in November 1957. Lecture Series"},{"key":"11389_CR201","doi-asserted-by":"crossref","unstructured":"Ungless EL, Rafferty A, Nag H, Ross B (2022) A Robust Bias Mitigation procedure based on the stereotype content model. arXiv prepr. arxiv:2210.14552","DOI":"10.18653\/v1\/2022.nlpcss-1.23"},{"issue":"11","key":"11389_CR202","doi-asserted-by":"publisher","first-page":"908","DOI":"10.1109\/32.730542","volume":"24","author":"A van Lamsweerde","year":"1998","unstructured":"van Lamsweerde A, Darimont R, Letier E (1998) Managing conflicts in goal-driven requirements engineering. IEEE Trans Softw Eng 24(11):908\u2013926. https:\/\/doi.org\/10.1109\/32.730542","journal-title":"IEEE Trans Softw Eng"},{"key":"11389_CR203","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (eds.) Adv. Neural Inf. Process. Syst. 30 (NeurIPS 2017), vol. 30. Curran Associates, Inc., ???"},{"key":"11389_CR204","unstructured":"Vega J, Chaudhary I, Xu C, Singh G (2023) Bypassing the safety training of open-source llms with priming attacks. arXiv preprint arXiv:2312.12321"},{"key":"11389_CR205","unstructured":"Vinyals O, Fortunato M, Jaitly N (2015) Pointer networks. In: Cortes, C., Lawrence, N., Lee, D., Sugiyama, M., Garnett, R. (eds.) Adv. Neural Inf. Process. Syst. 28 (NeurIPS 2015), vol. 28. Curran Associates, Inc., ???"},{"key":"11389_CR206","unstructured":"Vivien (2023) Better Steering LLM Agents with LMQL. https:\/\/vivien000.github.io\/blog\/journal\/better-steering-LLM-agents-with-LMQL.html"},{"key":"11389_CR207","first-page":"104471","volume":"37","author":"A Wachi","year":"2024","unstructured":"Wachi A, Tran T, Sato R, Tanabe T, Akimoto Y (2024) Stepwise alignment for constrained language model policy optimization. Adv Neural Inf Process Syst 37:104471\u2013104520","journal-title":"Adv Neural Inf Process Syst"},{"key":"11389_CR208","first-page":"31232","volume":"36","author":"B Wang","year":"2023","unstructured":"Wang B, Chen W, Pei H, Xie C, Kang M, Zhang C, Xu C, Xiong Z, Dutta R, Schaeffer R et al (2023) Decodingtrust: a comprehensive assessment of trustworthiness in gpt models. Adv Neural Inf Process Syst 36:31232\u201331339","journal-title":"Adv Neural Inf Process Syst"},{"key":"11389_CR209","unstructured":"Wang B, Chen W, Pei H, Xie C, Kang M, Zhang C, X, C, Xiong Z, Dutta R, Schaeffer R, Truong ST, Arora S, Mazeika M, Hendrycks D, Lin Z, Cheng Y, Koyejo S, Song D, Li B (2024) DecodingTrust: A comprehensive assessment of trustworthiness in GPT models. arXiv prepr. arxiv:2306.11698"},{"key":"11389_CR210","doi-asserted-by":"crossref","unstructured":"Wang Y, Liu X, Li Y, Chen M, Xiao C (2024) Adashield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting. In: European Conference on Computer Vision, pp. 77\u201394. Springer","DOI":"10.1007\/978-3-031-72661-3_5"},{"key":"11389_CR211","unstructured":"Wang H, Shu K (2023) Backdoor activation attack: Attack large language models using activation steering for safety-alignment. arXiv prepr. arxiv:2311.09433"},{"key":"11389_CR212","doi-asserted-by":"crossref","unstructured":"Wang X, Wang H, Yang D (2022) Measure and improve robustness in NLP models: A survey. arXiv prepr. arxiv:2112.08313v2 [cs.CL]","DOI":"10.18653\/v1\/2022.naacl-main.339"},{"key":"11389_CR213","unstructured":"Wang X, Wei J, Schuurmans D, Le QV, Chi EH, Narang S, Chowdhery A, Zhou D (2023) Self-consistency improves chain of thought reasoning in language models. In: 11th Int. Conf. Learn. Represent. (ICLR 2023)"},{"key":"11389_CR214","unstructured":"Wan S, Nikolaidis C, Song D, Molnar D, Crnkovich J, Grace J, Bhatt M, Chennabasappa S, Whitman S, Ding S, et\u00a0al (2024) Cyberseceval 3: Advancing the evaluation of cybersecurity risks and capabilities in large language models. arXiv preprint arXiv:2408.01605"},{"key":"11389_CR215","unstructured":"Webster M, Schmitt J (2024) LLM hallucinations: How to detect and prevent them with CI. CircleCI Blog ("},{"key":"11389_CR216","first-page":"24824","volume":"35","author":"J Wei","year":"2022","unstructured":"Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, Le QV, Zhou D et al (2022) Chain-of-thought prompting elicits reasoning in large language models. NeurIPS 35:24824\u201324837","journal-title":"NeurIPS"},{"key":"11389_CR217","first-page":"119","volume":"36","author":"A Wei","year":"2024","unstructured":"Wei A, Haghtalab N, Steinhardt J (2024) Jailbroken: how does llm safety training fail? NeurIPS 36:119","journal-title":"NeurIPS"},{"key":"11389_CR218","unstructured":"Wei J, Kim S, Jung H, Kim Y-H (2023) Leveraging large language models to power chatbots for collecting user self-reported data. arXiv prepr. arxiv:2301.05843"},{"key":"11389_CR219","unstructured":"Wei Z, Wang Y, Wang Y (2023) Jailbreak and guard aligned language models with only few in-context demonstrations. arXiv prepr. arxiv:2310.06387"},{"key":"11389_CR220","doi-asserted-by":"crossref","unstructured":"Welbl J, Glaese A, Uesato J, Dathathri S, Mellor J, Hendricks LA, Anderson K, Kohli P, Coppin B, Huang P-S (2021) Challenges in detoxifying language models. arXiv prepr. arxiv:2109.07445","DOI":"10.18653\/v1\/2021.findings-emnlp.210"},{"key":"11389_CR221","doi-asserted-by":"crossref","unstructured":"Welbl J, Glaese A, Uesato J, Dathathri S, Mellor J, Hendricks LA, Anderson K, Kohli P, Coppin B, Huang P-S (2021) Challenges in detoxifying language models. In: Moens, M.-F., Huang, X., Specia, L., Yih, S.W.-t. (eds.) Find. Assoc. Comput. Linguist.: EMNLP 2021, pp. 2447\u20132469. Association for Computational Linguistics, ???","DOI":"10.18653\/v1\/2021.findings-emnlp.210"},{"key":"11389_CR222","doi-asserted-by":"crossref","unstructured":"Weng F, Xu Y, Fu C, Wang W (2024) Mmj-bench: A comprehensive study on jailbreak attacks and defenses for multimodal large language models. arXiv preprint arXiv:2408.08464","DOI":"10.1609\/aaai.v39i26.34983"},{"key":"11389_CR223","doi-asserted-by":"crossref","unstructured":"Wen J, Ke P, Sun H, Zhang Z, Li C, Bai J, Huang M (2023) Unveiling the implicit toxicity in large language models. In: 2023 Conf. Empir. Methods Nat. Lang. Process. (EMNLP 2023)","DOI":"10.18653\/v1\/2023.emnlp-main.84"},{"key":"11389_CR224","first-page":"1","volume":"36","author":"A Xiang","year":"2022","unstructured":"Xiang A (2022) Being \u201cseen\u201d vs. \u201cmis-seen\u201d: tensions between privacy and fairness in computer vision. Harv J Law Technol 36:1","journal-title":"Harv J Law Technol"},{"key":"11389_CR225","unstructured":"Xiao Y, Jin Y, Bai Y, Wu Y, Yang X, Luo X, Yu W, Zhao X, Liu Y, Chen H, et\u00a0al (2023) Large language models can be good privacy protection learners. arXiv prepr. arxiv:2310.02469"},{"key":"11389_CR226","doi-asserted-by":"crossref","unstructured":"Xiao Y, Liang PP, Bhatt U, Neiswanger W, Salakhutdinov R, Morency L-P (2022) Uncertainty quantification with pre-trained language models: A large-scale empirical analysis. arXiv prepr. arxiv:2210.04714","DOI":"10.18653\/v1\/2022.findings-emnlp.538"},{"issue":"12","key":"11389_CR227","doi-asserted-by":"publisher","first-page":"1486","DOI":"10.1038\/s42256-023-00765-8","volume":"5","author":"Y Xie","year":"2023","unstructured":"Xie Y, Yi J, Shao J, Curl J, Lyu L, Chen Q, Xie X, Wu F (2023) Defending chatgpt against jailbreak attack via self-reminders. Nat Mach Intell 5(12):1486\u20131496","journal-title":"Nat Mach Intell"},{"key":"11389_CR228","doi-asserted-by":"publisher","unstructured":"Xie Z, Lukasiewicz T (2023) An empirical analysis of parameter-efficient methods for debiasing pre-trained language models. In: Rogers, A., Boyd-Graber, J.L., Okazaki, N. (eds.) Proc. 61st Annu. Meet. Assoc. Comput. Linguist., pp. 15730\u201315745. Association for Computational Linguistics, ???. https:\/\/doi.org\/10.18653\/V1\/2023.ACL-LONG.876","DOI":"10.18653\/V1\/2023.ACL-LONG.876"},{"key":"11389_CR229","doi-asserted-by":"crossref","unstructured":"Xu Z, Jiang F, Niu L, Jia J, Lin BY, Poovendran R (2024) SafeDecoding: Defending against jailbreak attacks via safety-aware decoding. arXiv prepr. arxiv:2402.08983","DOI":"10.18653\/v1\/2024.acl-long.303"},{"key":"11389_CR230","unstructured":"Yang Y, Dan S, Roth D, Lee I (2024) Benchmarking llm guardrails in handling multilingual toxicity. arXiv preprint arXiv:2410.22153"},{"key":"11389_CR231","doi-asserted-by":"crossref","unstructured":"Yang R, Fu M, Tantithamthavorn C, Arora C, Vandenhurk L, Chua J (2025) Ragva: Engineering retrieval augmented generation-based virtual assistants in practice. Journal of Systems and Software, 112436","DOI":"10.1016\/j.jss.2025.112436"},{"key":"11389_CR232","doi-asserted-by":"crossref","unstructured":"Yang J, Zhang X, Liang K, Liu Y (2023) Exploring the application of large language models in detecting and protecting personally identifiable information in archival data: a comprehensive study. In: IEEE Int. Conf. Big Data (BigData), pp. 2116\u20132123. IEEE, ???","DOI":"10.1109\/BigData59044.2023.10386949"},{"key":"11389_CR233","doi-asserted-by":"crossref","unstructured":"Yao H, Lou J, Ren K, Qin Z (2023) PromptCARE: Prompt copyright protection by watermark injection and verification. In: 2024 IEEE Symp. Secur. Priv. (SP 2024)","DOI":"10.1109\/SP54263.2024.00209"},{"key":"11389_CR234","unstructured":"Yeh K-C, Chi J-A, Lian D-C, Hsieh S-K (2023) Evaluating interfaced LLM bias. In: Proc. 35th Conf. Comput. Linguist. Speech Process. (ROCLING 2023), pp. 292\u2013299"},{"key":"11389_CR235","unstructured":"Ye W, Ou M, Li T, chen Y, Ma X, Yanggong Y, Wu S, Fu J, Chen G, Wang H, Zhao J (2023) Assessing hidden risks of LLMs: An empirical study on robustness, consistency, and credibility. arXiv prepr. arxiv:2305.10235v4 [cs.LG]"},{"key":"11389_CR236","unstructured":"Yi S, Liu Y, Sun Z, Cong T, He X, Song J, Xu K, Li Q (2024) Jailbreak attacks and defenses against large language models: A survey. arXiv preprint arXiv:2407.04295"},{"key":"11389_CR237","unstructured":"Yin X, Qi Y, Hu J, Chen Z, Dong Y, Zhao X, Huang X, Ruan W (2025) Taiji: Textual anchoring for immunizing jailbreak images in vision language models. arXiv preprint arXiv:2503.10872"},{"key":"11389_CR238","unstructured":"Yong ZX, Menghini C, Bach S (2023) Low-resource languages jailbreak GPT-4. In: Soc. Responsible Lang. Model. Res"},{"key":"11389_CR239","unstructured":"Yu X (2024) Create an AI Agent with Llama Guard in Anypoint Platform. https:\/\/medium.com\/@yuxiaojian\/create-an-ai-agent-with-llama-guard-in-anypoint-platform-a313b2c0b51f"},{"key":"11389_CR240","doi-asserted-by":"crossref","unstructured":"Yuan T, He Z, Dong L, Wang Y, Zhao R, Xia T, Xu L, Zhou B, Li F, Zhang Z, et\u00a0al (2024) R-judge: Benchmarking safety risk awareness for LLM agents. arXiv prepr. arxiv:2401.10019","DOI":"10.18653\/v1\/2024.findings-emnlp.79"},{"key":"11389_CR241","unstructured":"Yuan Y, Jiao W, Wang W, Huang J-t, He P, Shi S, Tu Z (2024) GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher. In: 12th Int. Conf. Learn. Represent. (ICLR 2024)"},{"key":"11389_CR242","unstructured":"Yu J, Lin X, Xing X (2023) Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts. arXiv prepr. arxiv:2309.10253"},{"key":"11389_CR243","unstructured":"Yu D, Naik S, Backurs A, Gopi S, Inan HA, Kamath G, Kulkarni J, Lee YT, Manoel A, Wutschitz L, Yekhanin S, Zhang H (2022) Differentially private fine-tuning of language models. In: 10th Int. Conf. Learn. Represent. (ICLR 2022). OpenReview.net"},{"key":"11389_CR244","doi-asserted-by":"publisher","unstructured":"Zampieri M, Malmasi S, Nakov P, Rosenthal S, Farra N, Kumar R (2019) Predicting the type and target of offensive posts in social media. In: Burstein, J., Doran, C., Solorio, T. (eds.) Proc. 2019 Conf. N. Am. Chapter Assoc. Comput. Linguist.: Hum. Lang. Technol., pp. 1415\u20131420. Association for Computational Linguistics, Minneapolis, Minnesota. https:\/\/doi.org\/10.18653\/v1\/N19-1144","DOI":"10.18653\/v1\/N19-1144"},{"issue":"21","key":"11389_CR245","doi-asserted-by":"publisher","first-page":"7640","DOI":"10.3390\/app10217640","volume":"10","author":"C Zeng","year":"2020","unstructured":"Zeng C, Li S, Li Q, Hu J, Hu J (2020) A survey on machine reading comprehension-tasks, evaluation metrics and benchmark datasets. Appl Sci 10(21):7640","journal-title":"Appl Sci"},{"key":"11389_CR246","doi-asserted-by":"crossref","unstructured":"Zhan Q, Fang R, Bindu R, Gupta A, Hashimoto T, Kang D (2023) Removing rlhf protections in gpt-4 via fine-tuning. arXiv prepr. arxiv:2311.05553","DOI":"10.18653\/v1\/2024.naacl-short.59"},{"key":"11389_CR247","doi-asserted-by":"crossref","unstructured":"Zhang W, Deng Y, Liu B, Pan SJ, Bing L (2023) Sentiment analysis in the era of large language models: A reality check. arXiv prepr. arxiv:2305.15005","DOI":"10.18653\/v1\/2024.findings-naacl.246"},{"key":"11389_CR248","unstructured":"Zhang H, Guo Z, Zhu H, Cao B, Lin L, Jia J, Chen J, Wu D (2023) On the safety of open-sourced large language models: Does alignment really prevent them from being misused? arXiv prepr. arxiv:2310.01581"},{"key":"11389_CR249","unstructured":"Zhang B, Shen X, Si WM, Sha Z, Chen Z, Salem A, Shen Y, Backes M, Zhang Y (2023) Comprehensive assessment of toxicity in ChatGPT. arXiv prepr. arxiv:2311.14685 [cs.CY]"},{"key":"11389_CR250","unstructured":"Zhang Z, Yang J, Ke P, Huang M (2023) Defending large language models against jailbreaking attacks through goal prioritization. arXiv prepr. arxiv:2311.09096"},{"key":"11389_CR251","unstructured":"Zhao J, Chen K, Yuan X, Qi Y, Zhang W, Yu N (2023) Silent guardian: Protecting text from malicious exploitation by large language models. arXiv prepr. arxiv:2312.09669"},{"key":"11389_CR252","doi-asserted-by":"crossref","unstructured":"Zhao S, Jia M, Tuan LA, Wen J (2024) Universal vulnerabilities in large language models: In-context learning backdoor attacks. arXiv prepr. arxiv:2401.05949","DOI":"10.18653\/v1\/2024.emnlp-main.642"},{"key":"11389_CR253","unstructured":"Zhou KZ, Sanfilippo MR (2023) Public perceptions of gender bias in large language models: Cases of chatgpt and ernie. arXiv prepr. arxiv:2309.09120"},{"key":"11389_CR254","unstructured":"Zhou A, Li B, Wang H (2024) Robust prompt optimization for defending language models against jailbreaking attacks. arXiv prepr. arxiv:2401.17263"},{"key":"11389_CR255","unstructured":"Zhou W, Wang X, Xiong L, Xia H, Gu Y, Chai M, Zhu F, Huang C, Dou S, Xi Z, et\u00a0al (2024) EasyJailbreak: A unified framework for jailbreaking large language models. arXiv prepr. arxiv:2403.12171"},{"key":"11389_CR256","unstructured":"Zhu K, Wang J, Zhou J, Wang Z, Chen H, Wang Y, Yang L, Ye W, Zhang Y, Gong NZ, Xie X (2023) PromptBench: Towards evaluating the robustness of large language models on adversarial prompts. arXiv prepr. arxiv:2306.04528v4 [cs.CL]"},{"key":"11389_CR257","unstructured":"Zhu S, Zhang R, An B, Wu G, Barrow J, Wang Z, Huang F, Nenkova A, Sun T (2023) Autodan: Automatic and interpretable adversarial attacks on large language models. arXiv prepr. arxiv:2310.15140"},{"key":"11389_CR258","unstructured":"Zou W, Geng R, Wang B, Jia J (2024) PoisonedRAG: Knowledge poisoning attacks to retrieval-augmented generation of large language models. arXiv prepr. arxiv:2402.07867"},{"key":"11389_CR259","unstructured":"Zou A, Wang Z, Carlini N, Nasr M, Kolter JZ, Fredrikson M (2023) Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11389-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-025-11389-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11389-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T00:38:28Z","timestamp":1764981508000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-025-11389-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,17]]},"references-count":259,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["11389"],"URL":"https:\/\/doi.org\/10.1007\/s10462-025-11389-2","relation":{},"ISSN":["1573-7462"],"issn-type":[{"value":"1573-7462","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,17]]},"assertion":[{"value":"21 October 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 August 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 October 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"382"}}