{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T16:03:32Z","timestamp":1784390612609,"version":"3.55.0"},"reference-count":151,"publisher":"Association for Computing Machinery (ACM)","issue":"7","funder":[{"DOI":"10.13039\/100006752","name":"U.S. Army Corps of Engineers","doi-asserted-by":"publisher","award":["W912HZ-23-2-0004"],"award-info":[{"award-number":["W912HZ-23-2-0004"]}],"id":[{"id":"10.13039\/100006752","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>Large Language Models (LLMs) are advancing rapidly and promising transformation across fields but pose challenges in oversight, ethics, and user trust. This review addresses trust issues like unintentional harms, opacity, vulnerability, misalignment with values, and environmental impact, all of which affect trust. Factors undermining trust include societal biases, opaque processes, misuse potential, and technology evolution challenges, especially in finance, healthcare, education, and policy. Recommended solutions include ethical oversight, industry accountability, regulation, and public involvement to reshape AI norms and incorporate ethics into development. A framework assesses trust in LLMs, analyzing trust dynamics and providing guidelines for responsible AI development. The review highlights limitations in building trustworthy AI, aiming to create a transparent and accountable ecosystem that maximizes benefits and minimizes risks, offering guidance for researchers, policymakers, and industry in fostering trust and ensuring responsible use of LLMs. We validate our frameworks through comprehensive experimental assessment across seven contemporary models, demonstrating substantial improvements in trustworthiness characteristics and identifying important disagreements with existing literature. Both theoretical foundations and empirical validation are provided in comprehensive supplementary materials.<\/jats:p>","DOI":"10.1145\/3777382","type":"journal-article","created":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T11:53:21Z","timestamp":1763380401000},"page":"1-43","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["Towards Trustworthy AI: A Review of Ethical and Robust Large Language Models"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8833-2274","authenticated-orcid":false,"given":"Md Meftahul","family":"Ferdaus","sequence":"first","affiliation":[{"name":"Department of Computer Science, University of New Orleans","place":["New Orleans, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-0932-2009","authenticated-orcid":false,"given":"Mahdi","family":"Abdelguerfi","sequence":"additional","affiliation":[{"name":"Computer Science, University of New Orleans","place":["New Orleans, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-3699-6762","authenticated-orcid":false,"given":"Elias","family":"Loup","sequence":"additional","affiliation":[{"name":"US Naval Research Laboratory Stennis Space Center","place":["Stennis Space Center, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2925-8856","authenticated-orcid":false,"given":"Kendall","family":"N. Niles","sequence":"additional","affiliation":[{"name":"US Army Corps of Engineers Mississippi Valley Division","place":["Vicksburg, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2782-4092","authenticated-orcid":false,"given":"Ken","family":"Pathak","sequence":"additional","affiliation":[{"name":"US Army Corps of Engineers, US Army Corps of Engineers Mississippi Valley Division","place":["Vicksburg, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4038-119X","authenticated-orcid":false,"given":"Steven","family":"Sloan","sequence":"additional","affiliation":[{"name":"US Army Corps of Engineers Mississippi Valley Division","place":["Vicksburg, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,1,8]]},"reference":[{"key":"e_1_3_3_2_2","volume-title":"Public Constitutional AI","author":"Abiri Anonymous","year":"2025","unstructured":"Anonymous Abiri. 2025. Public Constitutional AI. Technical Report. Georgia Law Review. Retrieved from, https:\/\/georgialawreview.org\/wp-content\/uploads\/2025\/05\/Abiri_Public-Constitutional-AI.pdf"},{"key":"e_1_3_3_3_2","doi-asserted-by":"crossref","unstructured":"Saleh Afroogh Ali Akbari Emmie Malone Mohammadali Kargar and Hananeh Alambeigi. 2024. Trust in AI: Progress challenges and future directions. Humanities and Social Sciences Communications 11 1568 (2024).","DOI":"10.1057\/s41599-024-04044-8"},{"key":"e_1_3_3_4_2","unstructured":"Mousa Al-kfairy. 2025. Strategic integration of generative AI in organizational settings: Applications challenges and adoption requirements. IEEE Engineering Management Review (2025) 1\u201314."},{"key":"e_1_3_3_5_2","first-page":"58","volume-title":"Informatics","author":"Al-Kfairy Mousa","year":"2024","unstructured":"Mousa Al-Kfairy, Dheya Mustafa, Nir Kshetri, Mazen Insiew, and Omar Alfandi. 2024. Ethical challenges and solutions of generative AI: An interdisciplinary perspective. In Informatics, Vol. 11. Multidisciplinary Digital Publishing Institute, 58."},{"key":"e_1_3_3_6_2","doi-asserted-by":"crossref","unstructured":"Mazen Aljohani Jingwei Hou Srinivas Kommu and Xin Wang. 2025. A comprehensive survey on the trustworthiness of large language models in healthcare. arXiv preprint arxiv:2502.15871 (2025).","DOI":"10.18653\/v1\/2025.findings-emnlp.356"},{"key":"e_1_3_3_7_2","doi-asserted-by":"crossref","unstructured":"Deafallah Alsadie. 2024. A comprehensive review of AI techniques for resource management in fog computing: Trends challenges and future directions. IEEE Access 12 (2024) 118007\u2013118059.","DOI":"10.1109\/ACCESS.2024.3447097"},{"key":"e_1_3_3_8_2","unstructured":"Afshine Amidi and Shervine Amidi. 2017. Deep Learning Cheatsheet. (2017). Retrieved from https:\/\/stanford.edu\/shervine\/teaching\/cs-230\/cheatsheet-deep-learning-tips-and-tricks"},{"key":"e_1_3_3_9_2","unstructured":"Anonymous. 2025. Agreeing to Disagree: Human-AI Collaboration in Ethical Decision-Making. SSRN Electronic Journal. (2025). Retrieved from https:\/\/papers.ssrn.com\/sol3\/papers.cfm?abstract_id=5262517"},{"key":"e_1_3_3_10_2","unstructured":"Anonymous. 2025. Where is morality on wheels? Decoding large language model (LLM) ethical decision-making in autonomous vehicles. Transportation Research Interdisciplinary Perspectives (2025). Retrieved from https:\/\/www.sciencedirect.com\/science\/article\/abs\/pii\/S2214367X25000572"},{"key":"e_1_3_3_11_2","unstructured":"Anthropic. 2025. Introducing Claude 4. Anthropic News. (May2025). Retrieved from https:\/\/www.anthropic.com\/news\/claude-4"},{"key":"e_1_3_3_12_2","unstructured":"Anthropic Research Team. 2025. Agentic Misalignment: How LLMs could be insider threats. Anthropic Research. (June2025). Retrieved from https:\/\/www.anthropic.com\/research\/agentic-misalignment"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3430984.3430987"},{"key":"e_1_3_3_14_2","unstructured":"Gregor Bachmann and Vaishnavh Nagarajan. 2024. The pitfalls of next-token prediction. arXiv preprint arxiv:2403.06963 (2024). https:\/\/arxiv.org\/abs\/2403.06963ICML 2024."},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467307"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","unstructured":"Yuntao Bai Andy Jones Kamal Ndousse Amanda Askell Anna Chen Nova DasSarma Dawn Drain Stanislav Fort Deep Ganguli Tom Henighan et\u00a0al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. ArXiv abs\/2204.05862 (2022). DOI:10.48550\/arXiv.2204.05862","DOI":"10.48550\/arXiv.2204.05862"},{"key":"e_1_3_3_17_2","unstructured":"Yuntao Bai Saurav Kadavath Sandipan Kundu Amanda Askell Jackson Kernion Andy Jones Andy Chen Anna Goldie Azalia Mirhoseini C McKinnon and et al.2022. Constitutional AI: Harmlessness from AI Feedback. ArXiv abs\/2212.08073 (2022)."},{"key":"e_1_3_3_18_2","unstructured":"Leonardo Berti Flavio Giorgi and Gjergji Kasneci. 2025. Emergent abilities in large language models: A survey. arXiv preprint arxiv:2503.05788 (2025)."},{"key":"e_1_3_3_19_2","unstructured":"Fernando Berzal. 2025. Differential privacy in machine learning: From symbolic AI to LLMs. arXiv preprint arxiv:2506.11687 (2025). https:\/\/arxiv.org\/abs\/2506.11687"},{"key":"e_1_3_3_20_2","volume-title":"Advances in Neural Information Processing Systems","author":"Bhardwaj Eshta","year":"2024","unstructured":"Eshta Bhardwaj, Harshit Gujral, Siyi Wu, Ciara Zogheib, Tegan Maharaj, and Christoph Becker. 2024. The state of data curation at NeurIPS: An assessment of dataset development practices in the datasets and benchmarks track. In Advances in Neural Information Processing Systems, Vol. 37."},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","unstructured":"Ashwin Bhat and Arijit Raychowdhury. 2023. Non-Uniform Interpolation in Integrated Gradients for Low-Latency Explainable-AI. (2023). DOI:10.48550\/arXiv.2302.11107arxiv:cs.AI\/2302.11107","DOI":"10.48550\/arXiv.2302.11107"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","unstructured":"David C. Brock and Burt Grad. 2022. Expert systems: Commercializing artificial intelligence. IEEE Annals of the History of Computing 44 (2022) 5\u20137. DOI:10.1109\/mahc.2022.3149612","DOI":"10.1109\/mahc.2022.3149612"},{"key":"e_1_3_3_23_2","unstructured":"Boxi Cao Keming Lu Xinyu Lu Jiawei Chen Mengjie Ren Hao Xiang Peilin Liu Yaojie Lu Ben He Xianpei Han Le Sun Hongyu Lin and Bowen Yu. 2024. Towards scalable automated alignment of LLMs: A survey. arXiv preprint arxiv:2406.01252 (2024). https:\/\/arxiv.org\/abs\/2406.01252"},{"key":"e_1_3_3_24_2","unstructured":"Stephen Casper Xander Davies Claudia Shi Thomas Krendl Gilbert J\u00e9r\u00e9my Scheurer Javier Rando Rachel Freedman Tomasz Korbak David Lindner Pedro Freire et\u00a0al. 2023. Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv preprint arxiv:2307.15217 (2023). https:\/\/arxiv.org\/abs\/2307.15217"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","unstructured":"S. Cass. 2016. What would marvin minsky read? Key works from the AI Titan\u2019s favorite authors. IEEE Spectrum 53 (2016) 22\u201322. DOI:10.1109\/MSPEC.2016.7439586","DOI":"10.1109\/MSPEC.2016.7439586"},{"key":"e_1_3_3_26_2","doi-asserted-by":"crossref","unstructured":"Shreyas Chaudhari Pranjal Aggarwal Vishvak Murahari Tanmay Rajpurohit Ashwin Kalyan Karthik Narasimhan Ameet Deshpande and Bruno Castro da Silva. 2024. RLHF deciphered: A critical analysis of reinforcement learning from human feedback for LLMs. Comput. Surveys 58 2 (2024) 1\u201337.","DOI":"10.1145\/3743127"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","unstructured":"Jiahao Chen Liang Wang Yifan Zhang and Xiaojun Liu. 2024. Towards a better understanding of evaluating trustworthiness in AI systems. Comput. Surveys 57 3 (2024) 1\u201342. Retrieved from 10.1145\/3721976","DOI":"10.1145\/3721976"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","unstructured":"Yu-Neng Chuang Guanchu Wang Chia yuan Chang Ruixiang Tang et\u00a0al. 2024. Large language models as faithful explainers. ArXiv abs\/2402.04678 (2024). DOI:10.48550\/arXiv.2402.04678","DOI":"10.48550\/arXiv.2402.04678"},{"key":"e_1_3_3_29_2","unstructured":"Constitutional AI Project. 2025. Constitutional AI: Tracking Anthropic\u2019s Revolutionary Framework. Constitutional AI Website. (2025). Retrieved from https:\/\/constitutional.ai\/"},{"key":"e_1_3_3_30_2","unstructured":"NVIDIA Corporation. 2017. Deep Learning: An Introductory Guide. (2017). Retrieved from https:\/\/research.nvidia.com\/publication\/2017-06_Deep-Learning-Introductory-Guide"},{"key":"e_1_3_3_31_2","volume-title":"Databricks AI Governance Framework (DAGF): A Comprehensive Approach to Responsible AI","author":"Inc. Databricks","year":"2025","unstructured":"Databricks Inc.2025. Databricks AI Governance Framework (DAGF): A Comprehensive Approach to Responsible AI. Technical Report. Databricks, San Francisco, CA. Retrieved from https:\/\/www.databricks.com\/ai-governance-frameworkAccessed: August 12, 2025."},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","unstructured":"Melis Dokumac\u0131. 2024. Legal frameworks for AI regulations. Human Computer Interaction (2024). DOI:10.62802\/ytst2927","DOI":"10.62802\/ytst2927"},{"key":"e_1_3_3_33_2","unstructured":"Sarah El-Sayed Canfer Akbulut Alexander McCroskery Ariel D Procaccia Vincent Conitzer and Iyad Rahwan. 2024. A mechanism-based approach to mitigating harms from persuasive generative AI. arXiv preprint arxiv:2404.15058 (2024)."},{"key":"e_1_3_3_34_2","unstructured":"European Parliament and Council of the European Union. 2024. Regulation (EU) 2024\/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act)... (July2024) 144 pages. Retrieved from https:\/\/eur-lex.europa.eu\/eli\/reg\/2024\/1689\/ojAccessed: August 12 2025."},{"key":"e_1_3_3_35_2","doi-asserted-by":"crossref","unstructured":"Oliver et\u00a0al. Faust. 2018. Deep learning for healthcare applications based on physiological signals: A review. Nature Scientific Reports 8 1 (2018) 1\u201313.","DOI":"10.1016\/j.cmpb.2018.04.005"},{"key":"e_1_3_3_36_2","unstructured":"Victor Gallego. 2024. Configurable safety tuning of language models with synthetic preference data. arXiv preprint arxiv:2404.00495 (2024)."},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3531820"},{"key":"e_1_3_3_38_2","doi-asserted-by":"crossref","unstructured":"Sai Rohit Gantla. 2025. Exploring mechanistic interpretability in large language models: Challenges approaches and insights. 2025 International Conference on Data Science and Engineering (2025). Retrieved from https:\/\/ieeexplore.ieee.org\/abstract\/document\/11011640\/","DOI":"10.1109\/ICDSAAI65575.2025.11011640"},{"key":"e_1_3_3_39_2","unstructured":"Satvik Golechha Maheep Chaudhary Joan Velja Alessandro Abate and Nandi Schoots. 2025. Studying cross-cluster modularity in neural networks. arXiv preprint arxiv:2502.02470 (2025)."},{"key":"e_1_3_3_40_2","unstructured":"Dongyu Gong and Hantao Zhang. 2024. Self-attention limits working memory capacity of transformer-based models. arXiv preprint arxiv:2409.10715 (2024). https:\/\/arxiv.org\/abs\/2409.10715"},{"key":"e_1_3_3_41_2","unstructured":"Charlie Griffin Louis Thomson Buck Shlegeris and Alessandro Abate. 2024. Games for AI control: Models of safety evaluations of AI deployment protocols. arXiv preprint arxiv:2409.07985 (2024)."},{"key":"e_1_3_3_42_2","unstructured":"Neel Guha Christie M Lawrence Lindsay A Gailmard Kit T Rodolfa Faiz Saremi Rishi Beaumavant Intikhab Deborah Mariano Raj Florentino Covellan Colleen Hemingberg et\u00a0al. 2024. AI regulation has its own alignment problem: The technical and institutional feasibility of disclosure registration licensing and auditing. George Washington Law Review 92 5 (2024) 1473."},{"key":"e_1_3_3_43_2","unstructured":"Daya Guo Dejian Yang Haowei Zhang Junxiao Song Ruoyu Zhang Runxin Xu Qihao Zhu Shirong Ma Peiyi Wang Xiao Bi et\u00a0al. 2025. DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. arXiv preprint arxiv:2501.12948 (2025)."},{"key":"e_1_3_3_44_2","unstructured":"Yufei Guo Muzhe Guo Juntao Su Zhou Yang Mengqiu Zhu Hongfei Li Mengyang Qiu and Shuo Shuo Liu. 2024. Bias in large language models: Origin evaluation and mitigation. arXiv preprint arxiv:2411.10915 (2024). https:\/\/arxiv.org\/abs\/2411.10915"},{"key":"e_1_3_3_45_2","unstructured":"Muhammad Usman Hadi Rizwan Qureshi Abbas Shah Muhammad Irfan Anas Zafar Muhammad Bilal Shaikh Naveed Akhtar Jia Wu Seyedali Mirjalili et\u00a0al. 2023. Large language models: A comprehensive survey of its applications challenges limitations and future prospects. Authorea Preprints 1 3 (2023) 1\u201326."},{"key":"e_1_3_3_46_2","doi-asserted-by":"crossref","unstructured":"Michael Hahn. 2020. Theoretical limitations of self-attention in neural sequence models. Transactions of the Association for Computational Linguistics 8 (2020) 156\u2013171. Retrieved from https:\/\/direct.mit.edu\/tacl\/article-abstract\/doi\/10.1162\/tacl_a_00306\/43545","DOI":"10.1162\/tacl_a_00306"},{"key":"e_1_3_3_47_2","unstructured":"Peter Henderson Tatsunori Hashimoto and Mark A Lemley. 2023. Where\u2019s the liability in harmful AI speech? Journal of Free Speech Law 3 (2023) 535\u2013594."},{"key":"e_1_3_3_48_2","volume-title":"AI Action Plan for Justice: Transforming the Justice System through Artificial Intelligence","author":"Justice HM Courts & Tribunals Service and Ministry of","year":"2025","unstructured":"HM Courts & Tribunals Service and Ministry of Justice. 2025. AI Action Plan for Justice: Transforming the Justice System through Artificial Intelligence. Technical Report. Ministry of Justice, United Kingdom, London, UK. Retrieved from https:\/\/www.gov.uk\/government\/publications\/ai-action-plan-justiceAccessed: August 12, 2025."},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","unstructured":"Yue Huang Lichao Sun Haoran Wang Siyuan Wu Qihui Zhang Yuan Li Chujie Gao Yixin Huang Wenhan Lyu Yixuan Zhang et\u00a0al. 2024. TrustLLM: Trustworthiness in large language models. ArXiv abs\/2401.05561 (2024). DOI:10.48550\/arXiv.2401.05561","DOI":"10.48550\/arXiv.2401.05561"},{"key":"e_1_3_3_50_2","volume-title":"AI Verify: Singapore\u2019s Framework for AI Governance and Testing","author":"Commission Infocomm Media Development Authority and Personal Data Protection","year":"2025","unstructured":"Infocomm Media Development Authority and Personal Data Protection Commission. 2025. AI Verify: Singapore\u2019s Framework for AI Governance and Testing. Technical Report. Government of Singapore, Singapore. Retrieved from https:\/\/www.aiverify.sg\/Accessed: August 12, 2025."},{"key":"e_1_3_3_51_2","unstructured":"International Organization for Standardization and International Electrotechnical Commission. 2023. ISO\/IEC 42001:2023 Information technology\u2014Artificial intelligence\u2014Management system. (December2023). Retrieved from https:\/\/www.iso.org\/standard\/81230.htmlAccessed: August 12 2025."},{"key":"e_1_3_3_52_2","doi-asserted-by":"crossref","unstructured":"Christine Jacob No\u00e9 Brasier Emanuele Laurenzi Sabina Heuss Stavroula-Georgia Mougiakakou Arzu C\u00f6ltekin and Marc K Peter. 2025. AI for IMPACTS framework for evaluating the long-term real-world impacts of AI-powered clinician tools: Systematic review and narrative synthesis. Journal of Medical Internet Research 27 (2025) e67485.","DOI":"10.2196\/67485"},{"key":"e_1_3_3_53_2","volume-title":"Proceedings of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Jain Sarthak","year":"2019","unstructured":"Sarthak Jain and Byron C. Wallace. 2019. Attention is not explanation. In Proceedings of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies."},{"key":"e_1_3_3_54_2","doi-asserted-by":"crossref","unstructured":"Hareem Javed Shaker El-Sappagh and Tamer Abuhmed. 2024. Robustness in deep learning models for medical diagnostics: Security and adversarial challenges towards robust AI applications. Artificial Intelligence Review 57 11 (2024) 1\u201355.","DOI":"10.1007\/s10462-024-11005-9"},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","unstructured":"Vitor Jeronymo Leandro Bonifacio Hugo Abonizio Marzieh Fadaee Roberto Lotufo Jakub Zavrel and Rodrigo Nogueira. 2023. InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval. (2023). DOI:10.48550\/arXiv.2301.01820arxiv:cs.CL\/2301.01820","DOI":"10.48550\/arXiv.2301.01820"},{"key":"e_1_3_3_56_2","unstructured":"Fengqing Jiang Zhangchen Xu Luyao Niu Zhen Xiang Bhaskar Ramasubramanian Bo Li and Radha Poovendran. 2024. ArtPrompt: ASCII art-based jailbreak attacks against aligned LLMs. arXiv preprint arxiv:2402.11753 (2024)."},{"key":"e_1_3_3_57_2","volume-title":"Artificial Intelligence and Human Rights: Parliamentary Inquiry Report","author":"Rights Joint Committee on Human","year":"2024","unstructured":"Joint Committee on Human Rights. 2024. Artificial Intelligence and Human Rights: Parliamentary Inquiry Report. Technical Report. UK Parliament, London, UK. Retrieved from https:\/\/committees.parliament.uk\/work\/1549\/artificial-intelligence-and-human-rights\/news\/Accessed: August 12, 2025."},{"key":"e_1_3_3_58_2","volume-title":"Proceedings of the 35th International Conference on Machine Learning","author":"Kim Been","year":"2018","unstructured":"Been Kim, Justin Gilmer, Fernanda Viegas, Ulf Erlingsson, and Martin Wattenberg. 2018. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV). In Proceedings of the 35th International Conference on Machine Learning."},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","unstructured":"John Kirchenbauer Jonas Geiping Yuxin Wen Jonathan Katz Ian Miers and Tom Goldstein. 2023. A Watermark for Large Language Models. (2023). DOI:10.48550\/arXiv.2301.10226arxiv:cs.CL\/2301.10226","DOI":"10.48550\/arXiv.2301.10226"},{"key":"e_1_3_3_60_2","unstructured":"Neeraja Kirtane V. Manushree and Aditya Kane. 2022. Efficient gender debiasing of pre-trained indic language models. ArXiv abs\/2209.03661 (2022)."},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","unstructured":"T. M. Kolade. 2024. Artificial intelligence and global security: Strengthening international cooperation and diplomatic relations. Archives of Current Research International (2024). DOI:10.9734\/acri\/2024\/v24i11945","DOI":"10.9734\/acri\/2024\/v24i11945"},{"key":"e_1_3_3_62_2","doi-asserted-by":"crossref","unstructured":"Ali Kore Elyar Abbasi Bavil Vallijah Subasri Moustafa Abdalla Benjamin Fine Elham Dolatabadi and Mohamed Abdalla. 2024. Empirical data drift detection experiments on real-world medical imaging data. Nature Communications 15 1 (2024) 1887.","DOI":"10.1038\/s41467-024-46142-w"},{"key":"e_1_3_3_63_2","doi-asserted-by":"crossref","unstructured":"Dominik Kowald Sebastian Scher Viktoria Pammer-Schindler Peter M\u00fcllner Kerstin Waxnegger Lea Demelius Angela Fessl Maximilian Toller Inti Gabriel Mendoza Estrada Ilija Simic et\u00a0al. 2024. Establishing and evaluating trustworthy AI: Overview and research challenges. arXiv preprint arxiv:2411.09973 (2024).","DOI":"10.3389\/fdata.2024.1467222"},{"key":"e_1_3_3_64_2","first-page":"1078","volume-title":"International Conference on Machine Learning","author":"Bras Ronan Le","year":"2020","unstructured":"Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew Peters, Ashish Sabharwal, and Yejin Choi. 2020. Adversarial filters of dataset biases. In International Conference on Machine Learning. PMLR, 1078\u20131088."},{"key":"e_1_3_3_65_2","first-page":"26874","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Lee Harrison","year":"2024","unstructured":"Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, and Sushant Prakash. 2024. RLAIF vs. RLHF: Scaling reinforcement learning from human feedback with AI feedback. In Proceedings of the 41st International Conference on Machine Learning. PMLR, 26874\u201326901."},{"key":"e_1_3_3_66_2","unstructured":"Seanie Lee Minsu Kim Lynn Cherif David Dobre Juho Lee Sung Ju Hwang Kenji Kawaguchi Gauthier Gidel Yoshua Bengio Nikolay Malkin and Moksh Jain. 2024. Learning diverse attacks on large language models for robust red-teaming and safety tuning. arXiv preprint arxiv:2405.18540 (2024)."},{"key":"e_1_3_3_67_2","unstructured":"Sung Une Lee Harsha Perera Yue Liu Boming Xia Qinghua Lu and Liming Zhu. 2024. Responsible AI question bank: A comprehensive tool for AI risk assessment. arXiv preprint arxiv:2408.11820 (2024)."},{"key":"e_1_3_3_68_2","doi-asserted-by":"publisher","unstructured":"Yuxuan Lei Jianxun Lian Jing Yao Xu Huang Defu Lian and Xing Xie. 2023. RecExplainer: Aligning large language models for recommendation model interpretability. ArXiv abs\/2311.10947 (2023). DOI:10.48550\/arXiv.2311.10947","DOI":"10.48550\/arXiv.2311.10947"},{"key":"e_1_3_3_69_2","unstructured":"Simon Lermen Charlie Rogers-Smith and Jeffrey Ladish. 2023. LoRA fine-tuning efficiently undoes safety training in llama 2-chat 70B. arXiv preprint arxiv:2310.20624 (2023)."},{"key":"e_1_3_3_70_2","unstructured":"Aaron J. Li Satyapriya Krishna and Himabindu Lakkaraju. 2024. More RLHF more trust? On the impact of preference alignment on trustworthiness. arXiv preprint arxiv:2404.18870 (2024)."},{"key":"e_1_3_3_71_2","volume-title":"International Conference on Learning Representations","author":"Li Binghui","year":"2025","unstructured":"Binghui Li and Yuanzhi Li. 2025. Adversarial training can provably improve robustness: Theoretical analysis of feature learning process under structured data. In International Conference on Learning Representations."},{"key":"e_1_3_3_72_2","doi-asserted-by":"publisher","unstructured":"Tianlin Li Xiaoyu Zhang Chao Du Tianyu Pang Qian Liu Qing Guo Chao Shen and Yang Liu. 2024. Your large language model is secretly a fairness proponent and you should prompt it like one. ArXiv abs\/2402.12150 (2024). DOI:10.48550\/arXiv.2402.12150","DOI":"10.48550\/arXiv.2402.12150"},{"key":"e_1_3_3_73_2","unstructured":"Yingji Li Mengnan Du Rui Song Xin Wang and Ying Wang. 2023. A survey on fairness in large language models. arXiv preprint arxiv:2308.10149 (2023)."},{"key":"e_1_3_3_74_2","doi-asserted-by":"publisher","unstructured":"Zuowei Li. 2024. AI ethics and transparency in operations management: How governance mechanisms can reduce data bias and privacy risks. Journal of Applied Economics and Policy Studies (2024). DOI:10.54254\/2977-5701\/13\/2024130","DOI":"10.54254\/2977-5701\/13\/2024130"},{"key":"e_1_3_3_75_2","doi-asserted-by":"crossref","unstructured":"Z. Li F. Liu W. Yang S. Peng and J. Zhou. 2021. A survey of convolutional neural networks: Analysis applications and prospects. IEEE Trans. Neural Netw. Learn. Syst. 33 12 (2021) 6999\u20137019.","DOI":"10.1109\/TNNLS.2021.3084827"},{"key":"e_1_3_3_76_2","doi-asserted-by":"publisher","unstructured":"David Lindner and Mennatallah El-Assady. 2022. Humans are not boltzmann distributions: Challenges and opportunities for modelling human feedback and interaction in reinforcement learning. ArXiv abs\/2206.13316 (2022). DOI:10.48550\/arXiv.2206.13316","DOI":"10.48550\/arXiv.2206.13316"},{"key":"e_1_3_3_77_2","unstructured":"Yang Liu Yuanshun Yao Jean-Francois Ton Xiaoying Zhang Ruocheng Guo Hao Cheng Yegor Klochkov Muhammad Faaiz Taufiq and Hang Li. 2023. Trustworthy LLMs: A survey and guideline for evaluating large language models\u2019 alignment. arXiv preprint arxiv:2308.05374 (2023)."},{"key":"e_1_3_3_78_2","doi-asserted-by":"crossref","unstructured":"Zhaoming Liu. 2024. Cultural bias in large language models: A comprehensive analysis and mitigation strategies. Journal of Transcultural Communication 3 2 (2024) 224\u2013244.","DOI":"10.1515\/jtc-2023-0019"},{"key":"e_1_3_3_79_2","unstructured":"Zhaowei Liu Congying Xia Wei He and Chengming Wang. 2024. Trustworthiness and self-awareness in large language models: An exploration through the think-solve-verify framework. Proceedings of the 2024 Joint International Conference on Computational Linguistics Language Resources and Evaluation (2024) 16758\u201316769."},{"key":"e_1_3_3_80_2","doi-asserted-by":"crossref","unstructured":"Brady Lund Zeynep Orhan Nishith Reddy Mannuru Ravi Varma Kumar Bevara Brett Porter Meka Kasi Vinaih and Padmapadanand Bhaskara. 2025. Standards frameworks and legislation for artificial intelligence (AI) transparency. AI and Ethics 5 (2025) 3639\u20133655. Retrieved from https:\/\/link.springer.com\/article\/10.1007\/s43681-025-00661-4","DOI":"10.1007\/s43681-025-00661-4"},{"key":"e_1_3_3_81_2","doi-asserted-by":"publisher","unstructured":"Xinyin Ma Xinchao Wang Gongfan Fang Yongliang Shen and Weiming Lu. 2022. Prompting to Distill: Boosting Data-Free Knowledge Distillation via Reinforced Prompt. (2022). DOI:10.48550\/arXiv.2205.07523arxiv:cs.CL\/2205.07523","DOI":"10.48550\/arXiv.2205.07523"},{"key":"e_1_3_3_82_2","doi-asserted-by":"crossref","unstructured":"David Manheim Sammy Martin Mark Bailey Mikhail Samin and Ross Greutzmacher. 2025. The necessity of AI audit standards boards. AI & Society (2025). Retrieved from https:\/\/link.springer.com\/article\/10.1007\/s00146-025-02320-yOpen access.","DOI":"10.1007\/s00146-025-02320-y"},{"key":"e_1_3_3_83_2","doi-asserted-by":"publisher","unstructured":"Niklas Maus Patrick Chao Eric Wong and Jacob R. Gardner. 2023. Adversarial Prompting for Black Box Foundation Models. (2023). DOI:10.48550\/arXiv.2302.04237arxiv:cs.CL\/2302.04237","DOI":"10.48550\/arXiv.2302.04237"},{"key":"e_1_3_3_84_2","doi-asserted-by":"publisher","unstructured":"J. McCarthy M. Minsky N. Rochester and C. Shannon. 2006. A proposal for the dartmouth summer research project on artificial intelligence august 31 1955. AI Mag. 27 (2006) 12\u201314. DOI:10.1609\/aimag.v27i4.1904","DOI":"10.1609\/aimag.v27i4.1904"},{"key":"e_1_3_3_85_2","volume-title":"Proceedings of the Digital Education Council Global Summit","author":"Singapore Ministry of Education","year":"2024","unstructured":"Ministry of Education Singapore. 2024. Digital education council global summit: Responsible AI in education. In Proceedings of the Digital Education Council Global Summit. Ministry of Education Singapore, Singapore. Retrieved from https:\/\/www.moe.gov.sg\/digital-education-summit-2024Accessed: August 12, 2025."},{"key":"e_1_3_3_86_2","doi-asserted-by":"crossref","unstructured":"Seyed Tohid Hosseini Mortaji and Mohammad Ebrahim Sadeghi. 2024. Assessing the reliability of artificial intelligence systems: Challenges metrics and future directions. International Journal of Innovation in Management Economics and Social Sciences 4 4 (2024) 1\u201315.","DOI":"10.59615\/ijimes.4.2.1"},{"key":"e_1_3_3_87_2","doi-asserted-by":"publisher","unstructured":"Nithesh Naik B. Hameed D. Shetty Dishant Swain M. Shah R. Paul Kaivalya Aggarwal Sufyan Ibrahim Vathsala Patil Komal Smriti Suyog Shetty Bhavan Prasad Rai P. Chlosta and B. Somani. 2022. Legal and ethical consideration in artificial intelligence in healthcare: Who takes responsibility? Frontiers in Surgery 9 (2022). DOI:10.3389\/fsurg.2022.862322","DOI":"10.3389\/fsurg.2022.862322"},{"key":"e_1_3_3_88_2","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3594074"},{"key":"e_1_3_3_89_2","volume-title":"AI Risk Management Framework (AI RMF 1.0): Generative AI Profile","author":"Technology National Institute of Standards and","year":"2024","unstructured":"National Institute of Standards and Technology. 2024. AI Risk Management Framework (AI RMF 1.0): Generative AI Profile. Technical Report NIST AI 600-1. U.S. Department of Commerce, Gaithersburg, MD. Retrieved from https:\/\/www.nist.gov\/itl\/ai-risk-management-frameworkAccessed: August 12, 2025."},{"key":"e_1_3_3_90_2","unstructured":"Joe O\u2019Brien Shaun Ee and Zoe Williams. 2023. Deployment corrections: An incident response framework for frontier AI models. arXiv preprint arxiv:2310.00328 (2023)."},{"key":"e_1_3_3_91_2","doi-asserted-by":"crossref","first-page":"2307","DOI":"10.1145\/3637528.3671883","volume-title":"Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Oosterhuis Harrie","year":"2024","unstructured":"Harrie Oosterhuis, Rolf Jagerman, Zhen Qin, Xuanhui Wang, and Michael Bendersky. 2024. Reliable confidence intervals for information retrieval evaluation using generative AI. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. ACM, 2307\u20132317."},{"key":"e_1_3_3_92_2","volume-title":"OECD AI Principles: Updated Guidance for Responsible AI","author":"Development Organisation for Economic Co-operation and","year":"2024","unstructured":"Organisation for Economic Co-operation and Development. 2024. OECD AI Principles: Updated Guidance for Responsible AI. Technical Report. OECD Publishing, Paris, France. Retrieved from https:\/\/www.oecd.org\/digital\/artificial-intelligence\/ai-principles\/Accessed: August 12, 2025."},{"key":"e_1_3_3_93_2","doi-asserted-by":"publisher","unstructured":"Celso Cancela Outeda. 2024. The EU\u2019s AI act: A framework for collaborative governance. Internet Things 27 (2024) 101291. DOI:10.1016\/j.iot.2024.101291","DOI":"10.1016\/j.iot.2024.101291"},{"key":"e_1_3_3_94_2","doi-asserted-by":"publisher","DOI":"10.1177\/1541931218621046"},{"key":"e_1_3_3_95_2","unstructured":"Yujian Peng Jiawei Wang Hao Yu and Amir Houmansadr. 2024. Data extraction attacks in retrieval-augmented generation via backdoors. arXiv preprint arxiv:2411.01705 (2024). https:\/\/arxiv.org\/abs\/2411.01705"},{"key":"e_1_3_3_96_2","doi-asserted-by":"crossref","unstructured":"F. Piccialli V. Di Somma F. Giampaolo S. Cuomo and G. Fortino. 2021. A survey on deep learning in medicine: Why how and when? Inf. Fusion 66 (2021) 111\u2013137.","DOI":"10.1016\/j.inffus.2020.09.006"},{"key":"e_1_3_3_97_2","doi-asserted-by":"crossref","unstructured":"Nineta Polemi Isabel Pra\u00e7a Kitty Kioskli and Adrien B\u00e9cue. 2024. Challenges and efforts in managing AI trustworthiness risks: A state of knowledge. Frontiers in Big Data 7 (2024) 1381163.","DOI":"10.3389\/fdata.2024.1381163"},{"key":"e_1_3_3_98_2","unstructured":"Aman Priyanshu Yash Maurya and Zuofei Hong. 2024. AI governance and accountability: An analysis of anthropic\u2019s claude. arXiv preprint arxiv:2407.01557 (2024)."},{"key":"e_1_3_3_99_2","doi-asserted-by":"publisher","unstructured":"V. Rajaraman. 2014. JohnMcCarthy\u2014father of artificial intelligence. Resonance 19 (2014) 198\u2013207. DOI:10.1007\/S12045-014-0027-9","DOI":"10.1007\/S12045-014-0027-9"},{"key":"e_1_3_3_100_2","doi-asserted-by":"crossref","unstructured":"Maribeth Rauh Nahema Marchal Arianna Manzini Lisa Anne Hendricks Philip Torr Abeba Birhane and Laura Weidinger. 2024. Gaps in the safety evaluation of generative AI. Proceedings of the 2024 AAAI\/ACM Conference on AI Ethics and Society 7 (2024) 1057\u20131069.","DOI":"10.1609\/aies.v7i1.31717"},{"key":"e_1_3_3_101_2","unstructured":"Anka Reuel Ben Bucknall Stephen Casper Tim Fist Lisa Soder Onni Aarne Lewis Hammond Lujain Ibrahim Alan Chan Peter Wills et\u00a0al. 2024. Open problems in technical AI governance. Transactions on Machine Learning Research (2024)."},{"key":"e_1_3_3_102_2","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939778"},{"key":"e_1_3_3_103_2","doi-asserted-by":"crossref","unstructured":"Lorenzo Rossi Michael Aerni Jie Zhang and Florian Tram\u00e8r. 2025. Membership inference attacks on sequence models. arXiv preprint arxiv:2506.05126 (2025). https:\/\/arxiv.org\/abs\/2506.05126Accepted to the 8th Deep Learning Security and Privacy Workshop (DLSP) workshop (best paper award).","DOI":"10.1109\/SPW67851.2025.00014"},{"key":"e_1_3_3_104_2","doi-asserted-by":"crossref","unstructured":"Paul R\u00f6ttger Fabio Pernisi Bertie Vidgen and Dirk Hovy. 2025. SafetyPrompts: A systematic review of open datasets for evaluating and improving large language model safety. Proceedings of the AAAI Conference on Artificial Intelligence 39 21 (2025) 23456\u201323464.","DOI":"10.1609\/aaai.v39i26.34975"},{"key":"e_1_3_3_105_2","doi-asserted-by":"crossref","unstructured":"Tahereh Saheb and Tayebeh Saheb. 2023. Topical review of artificial intelligence national policies: A mixed method analysis. Technology in Society 74 (2023) 102316.","DOI":"10.1016\/j.techsoc.2023.102316"},{"key":"e_1_3_3_106_2","doi-asserted-by":"crossref","unstructured":"Esra \u015eahin Nurten Nur Arslan and Dilek \u00d6zdemir. 2025. Unlocking the black box: An in-depth review on interpretability explainability and reliability in deep learning. Neural Computing and Applications 37 (2025) 1\u201342.","DOI":"10.1007\/s00521-024-10437-2"},{"key":"e_1_3_3_107_2","doi-asserted-by":"crossref","unstructured":"Sridharan Sankaran. 2025. Enhancing trust through standards: A comparative risk-impact framework for aligning ISO AI standards with global ethical and regulatory contexts. arXiv preprint arxiv:2504.16139 (2025).","DOI":"10.1109\/ACDSA65407.2025.11166403"},{"key":"e_1_3_3_108_2","doi-asserted-by":"crossref","unstructured":"I. H. Sarker. 2021. Machine learning: Algorithms real-world applications and research directions. SN Comput. Sci. 2 (2021) 1\u201321.","DOI":"10.1007\/s42979-021-00592-x"},{"key":"e_1_3_3_109_2","doi-asserted-by":"crossref","unstructured":"A. Saxe S. Nelli and C. Summerfield. 2021. If deep learning is the answer what is the question? Nat. Rev. Neurosci. 22 1 (2021) 55\u201367.","DOI":"10.1038\/s41583-020-00395-8"},{"key":"e_1_3_3_110_2","unstructured":"Sarah Schwettmann Tamar Rott Shaham Joanna Materzynska Neil Chowdhury Shuang Li Jacob Andreas David Bau and Antonio Torralba. 2023. FIND: A function description benchmark for evaluating interpretability methods. arXiv e-prints (2023) arXiv\u20132309."},{"key":"e_1_3_3_111_2","doi-asserted-by":"publisher","unstructured":"Erfan Shayegani Md. Abdullah Al Mamun Yu Fu Pedram Zaree et\u00a0al. 2023. Survey of vulnerabilities in large language models revealed by adversarial attacks. ArXiv abs\/2310.10844 (2023). DOI:10.48550\/arXiv.2310.10844","DOI":"10.48550\/arXiv.2310.10844"},{"key":"e_1_3_3_112_2","doi-asserted-by":"crossref","unstructured":"Kalkidan Bisrat Shiferaw Maike Roloff Irina Balaur Danielle Welter Andra Waagmeester and Dagmar Waltemath. 2025. Guidelines and standard frameworks for artificial intelligence in medicine: A systematic review. JAMIA Open 8 1 (2025) ooae155.","DOI":"10.1093\/jamiaopen\/ooae155"},{"key":"e_1_3_3_113_2","unstructured":"Chandan Singh Armin Askari Rich Caruana and Jianfeng Gao. 2022. Augmenting interpretable models with llms during training. arXiv preprint arxiv:2209.11799 (2022)."},{"key":"e_1_3_3_114_2","doi-asserted-by":"publisher","unstructured":"Taylor Sorensen Joshua Robinson Christopher Rytting Alexander Glenn Shaw Kyle Rogers Alexia Pauline Delorey Mahmoud Khalil Nancy Fulda and David Wingate. 2022. An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels. (2022). DOI:10.18653\/v1\/2022.acl-long.60arxiv:cs.CL\/2206.15076","DOI":"10.18653\/v1\/2022.acl-long.60"},{"key":"e_1_3_3_115_2","volume-title":"Global AI Governance Action Plan","author":"China State Council of the People\u2019s Republic of","year":"2025","unstructured":"State Council of the People\u2019s Republic of China. 2025. Global AI Governance Action Plan. Technical Report. Ministry of Science and Technology, People\u2019s Republic of China, Shanghai, China. Retrieved from https:\/\/www.most.gov.cn\/ai-governance-plan\/Accessed: August 12, 2025."},{"key":"e_1_3_3_116_2","doi-asserted-by":"publisher","DOI":"10.5555\/3305890.3306024"},{"key":"e_1_3_3_117_2","doi-asserted-by":"publisher","unstructured":"V. Sze Yu hsin Chen Tien-Ju Yang and J. Emer. 2017. Efficient processing of deep neural networks: A tutorial and survey. Proc. IEEE 105 (2017) 2295\u20132329. DOI:10.1109\/JPROC.2017.2761740","DOI":"10.1109\/JPROC.2017.2761740"},{"key":"e_1_3_3_118_2","unstructured":"Henrik Skaug S\u00e6tra and John Danaher. 2025. Resolving the battle of short-vs. long-term AI risks. AI and Ethics (2025). Retrieved from https:\/\/link.springer.com\/article\/10.1007\/s43681-023-00336-y"},{"key":"e_1_3_3_119_2","doi-asserted-by":"crossref","unstructured":"Ruixiang Tang Yu-Neng Chuang and Xia Hu. 2024. The science of detecting LLM-generated text. Commun. ACM 67 4 (2024) 50\u201359.","DOI":"10.1145\/3624725"},{"key":"e_1_3_3_120_2","doi-asserted-by":"crossref","unstructured":"Ruixiang Tang Dehan Kong Lo-li Huang and Hui Xue. 2023. Large language models can be lazy learners: Analyze shortcuts in in-context learning. Findings of the Association for Computational Linguistics: ACL 2023 (2023) 4645\u20134657.","DOI":"10.18653\/v1\/2023.findings-acl.284"},{"key":"e_1_3_3_121_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.284"},{"key":"e_1_3_3_122_2","volume-title":"Proceedings of the 7th International Conference on Learning Representations","author":"Tenney Ian","year":"2019","unstructured":"Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. What do you learn from context? Probing for sentence structure in contextualized word representations. In Proceedings of the 7th International Conference on Learning Representations."},{"key":"e_1_3_3_123_2","volume-title":"America\u2019s AI Action Plan","author":"House The White","year":"2025","unstructured":"The White House. 2025. America\u2019s AI Action Plan. Technical Report. Executive Office of the President of the United States. Retrieved from https:\/\/www.whitehouse.gov\/ai-action-plan\/Accessed: August 12, 2025."},{"key":"e_1_3_3_124_2","doi-asserted-by":"crossref","unstructured":"E. Tokpo Pieter Delobelle B. Berendt and T. Calders. 2023. How far can it go?: On intrinsic gender bias mitigation for text classification. ArXiv abs\/2301.12855 (2023).","DOI":"10.18653\/v1\/2023.eacl-main.248"},{"key":"e_1_3_3_125_2","unstructured":"Miles Turpin Julian Michael Ethan Perez and Samuel Bowman. 2024. Language models don\u2019t always say what they think: Unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems 36 (2024) 74952\u201374965."},{"key":"e_1_3_3_126_2","doi-asserted-by":"crossref","unstructured":"Oskar Van der Wal Jaap Jumelet K. Schulz and Willem H. Zuidema. 2022. The birth of bias: A case study on the evolution of gender bias in an English language model. ArXiv abs\/2207.10245 (2022).","DOI":"10.18653\/v1\/2022.gebnlp-1.8"},{"key":"e_1_3_3_127_2","doi-asserted-by":"publisher","DOI":"10.34190\/iccws.17.1.25"},{"key":"e_1_3_3_128_2","volume-title":"Proceedings of the ACL Workshop on BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Vig Jesse","year":"2019","unstructured":"Jesse Vig. 2019. BERTViz: A tool for visualizing multi-head self-attention in the BERT model. In Proceedings of the ACL Workshop on BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP."},{"key":"e_1_3_3_129_2","doi-asserted-by":"crossref","unstructured":"G. Vilone and L. Longo. 2021. Notions of explainability and evaluation approaches for explainable artificial intelligence. Inf. Fusion 76 (2021) 89\u2013106.","DOI":"10.1016\/j.inffus.2021.05.009"},{"key":"e_1_3_3_130_2","unstructured":"Boxin Wang Weixin Chen Hengzhi Pei Chulin Xie Mintong Kang Chenhui Zhang Chejian Xu Zidi Xiong Ritik Dutta Rylan Schaeffer et\u00a0al. 2023. DecodingTrust: A comprehensive assessment of trustworthiness in GPT models. arXiv preprint arxiv:2306.11698 (2023)."},{"key":"e_1_3_3_131_2","unstructured":"Chengdong Wang Yanshan Liu Bin Bi Daoguang Zhang Zhongzhou Li Yongbin Ma Yongdong He et\u00a0al. 2025. Safety in large reasoning models: A survey. arXiv preprint arxiv:2504.17704 (2025)."},{"key":"e_1_3_3_132_2","unstructured":"Kai Wang Yihao Zhang and Meng Sun. 2025. When thinking LLMs lie: Unveiling the strategic deception in representations of reasoning models. arXiv preprint arxiv:2506.04909 (2025)."},{"key":"e_1_3_3_133_2","unstructured":"Tianyang Wang Yunze Wang Jun Zhou Benji Peng Xinyuan Song Charles Zhang Xintian Sun Qian Niu Junyu Liu Silin Chen et\u00a0al. 2025. From aleatoric to epistemic: Exploring uncertainty quantification techniques in artificial intelligence. arXiv preprint arxiv:2501.03282 (2025)."},{"key":"e_1_3_3_134_2","unstructured":"Ze Wang Zekun Wu Jeremy Zhang Xin Guan Navya Jain Skylar Lu Saloni Gupta and Adriano Koshiyama. 2024. Bias amplification: Large language models as increasingly biased media. arXiv preprint arxiv:2410.15234 (2024)."},{"key":"e_1_3_3_135_2","unstructured":"Zhiyu Wang Jiahao Zhu Jiayi Li Fanzhi Mu Yanghua Zhang Chao Huang Yansong Zhang et\u00a0al. 2024. Automated red-teaming for large language models: A survey. arXiv preprint arxiv:2507.05538 (2024). https:\/\/arxiv.org\/abs\/2507.05538"},{"key":"e_1_3_3_136_2","unstructured":"Jason Wei Xuezhi Wang Dale Schuurmans Maarten Bosma Fei Xia Ed Chi Quoc V Le Denny Zhou et\u00a0al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (2022) 24824\u201324837."},{"key":"e_1_3_3_137_2","unstructured":"Laura Weidinger John Mellor Sebastian Riedel and Isabelle Augenstein. 2021. Ethical and social risks of harm from language models. arXiv preprint arxiv:2112.04359 (2021)."},{"key":"e_1_3_3_138_2","doi-asserted-by":"publisher","unstructured":"Yixuan Weng Minjun Zhu Shizhu He Kang Liu and Jun Zhao. 2022. Large Language Models are reasoners with Self-Verification. (2022). DOI:10.48550\/arXiv.2212.09561arxiv:cs.CL\/2212.09561","DOI":"10.48550\/arXiv.2212.09561"},{"key":"e_1_3_3_139_2","unstructured":"Xinyi Wu Yifei Wang Stefanie Jegelka and Ali Jadbabaie. 2025. On the emergence of position bias in transformers. arXiv preprint arxiv:2502.01951 (2025). https:\/\/arxiv.org\/abs\/2502.01951ICML 2025 camera-ready."},{"key":"e_1_3_3_140_2","unstructured":"Xinyi Wu Yifei Wang Stefanie Jegelka and Ali Jadbabaie. 2025. Unpacking the bias of large language models. MIT News. (June2025). Retrieved from https:\/\/news.mit.edu\/2025\/unpacking-large-language-model-bias-0617International Conference on Machine Learning."},{"key":"e_1_3_3_141_2","unstructured":"Zhongyuan Xie Jixi Guo Tao Yu and Shuai Li. 2024. Calibrating reasoning in language models with internal consistency. Advances in Neural Information Processing Systems 37 (2024). Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2024\/hash\/d037fd021c9aace128b8ce25001cdb6c-Abstract-Conference.html"},{"key":"e_1_3_3_142_2","unstructured":"Rui Xin Niloofar Mireshghallah Shiyi Shirley Li Mufan Duan Hyunwoo Kim et\u00a0al. 2025. A false sense of privacy: Evaluating textual data sanitization beyond surface-level privacy leakage. arXiv preprint arxiv:2504.21035 (2025). https:\/\/arxiv.org\/abs\/2504.21035"},{"key":"e_1_3_3_143_2","doi-asserted-by":"publisher","unstructured":"Jianhao Yuan Shuyang Sun Daniel Omeiza Bo Zhao et\u00a0al. 2024. RAG-driver: Generalisable driving explanations with retrieval-augmented in-context learning in multi-modal large language model. ArXiv abs\/2402.10828 (2024). DOI:10.48550\/arXiv.2402.10828","DOI":"10.48550\/arXiv.2402.10828"},{"key":"e_1_3_3_144_2","doi-asserted-by":"crossref","unstructured":"Esmat Zaidan and Imad Antoine Ibrahim. 2024. AI governance in a complex and rapidly changing regulatory landscape: A global perspective. Humanities and Social Sciences Communications 11 1121 (2024).","DOI":"10.1057\/s41599-024-03560-x"},{"key":"e_1_3_3_145_2","unstructured":"Xue Zhang. 2025. Constitution or collapse? Exploring constitutional AI with llama 3-8B. arXiv preprint arxiv:2504.04918 (2025)."},{"key":"e_1_3_3_146_2","unstructured":"Yuki Zhang Liang Wang Ming Chen Xiaojun Liu and Jie Zhou. 2024. Temporal consistency evaluation in large language models. arXiv preprint arxiv:2410.04640 (2024). https:\/\/arxiv.org\/abs\/2410.04640"},{"key":"e_1_3_3_147_2","unstructured":"Zhexin Zhang Leqi Lei Lindong Wu Rui Sun Yongkang Huang Chong Long Xiao Liu Xuanyu Lei Jie Tang and Minlie Huang. 2023. SafetyBench: Evaluating the safety of large language models. arXiv preprint arxiv:2309.07045 (2023)."},{"key":"e_1_3_3_148_2","doi-asserted-by":"crossref","unstructured":"Haiyan Zhao Hanjie Chen Fan Yang Ninghao Liu Huiqi Deng Hengyi Cai Shuaiqiang Wang Dawei Yin and Mengnan Du. 2024. Explainability for large language models: A survey. ACM Transactions on Intelligent Systems and Technology 15 2 (2024) 1\u201338.","DOI":"10.1145\/3639372"},{"key":"e_1_3_3_149_2","unstructured":"Kaiwen Zhou Chengzhi Liu Xuandong Zhao Shreedhar Jangam Jayanth Srinivasa Gaowen Liu Dawn Song and Xin Eric Wang. 2025. The hidden risks of large reasoning models: A safety assessment of R1. arXiv preprint arxiv:2502.12659 (2025). https:\/\/arxiv.org\/abs\/2502.12659"},{"key":"e_1_3_3_150_2","doi-asserted-by":"publisher","unstructured":"Yongchao Zhou Andrei Ioan Muresanu Ziwen Han Keiran Paster Silviu Pitis Harris Chan and Jimmy Ba. 2022. Large Language Models Are Human-Level Prompt Engineers. (2022). DOI:10.48550\/arXiv.2211.01910arxiv:cs.CL\/2211.01910","DOI":"10.48550\/arXiv.2211.01910"},{"key":"e_1_3_3_151_2","unstructured":"Zhaowei Zhu Jialu Wang Hao Cheng and Yang Liu. 2023. Unmasking and improving data credibility: A study with datasets for training harmless language models. arXiv preprint arxiv:2311.11202 (2023)."},{"key":"e_1_3_3_152_2","unstructured":"Andy Zou Zifan Wang Nicholas Carlini Milad Nasr J Zico Kolter and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arxiv:2307.15043 (2023)."}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3777382","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,8]],"date-time":"2026-01-08T15:56:38Z","timestamp":1767887798000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3777382"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,8]]},"references-count":151,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3777382"],"URL":"https:\/\/doi.org\/10.1145\/3777382","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,8]]},"assertion":[{"value":"2025-04-20","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-09","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}