{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T05:09:32Z","timestamp":1785906572317,"version":"3.56.0"},"reference-count":208,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2024,2,22]],"date-time":"2024-02-22T00:00:00Z","timestamp":1708560000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2024,4,30]]},"abstract":"<jats:p>Large language models (LLMs) have demonstrated impressive capabilities in natural language processing. However, their internal mechanisms are still unclear and this lack of transparency poses unwanted risks for downstream applications. Therefore, understanding and explaining these models is crucial for elucidating their behaviors, limitations, and social impacts. In this article, we introduce a taxonomy of explainability techniques and provide a structured overview of methods for explaining Transformer-based language models. We categorize techniques based on the training paradigms of LLMs: traditional fine-tuning-based paradigm and prompting-based paradigm. For each paradigm, we summarize the goals and dominant approaches for generating local explanations of individual predictions and global explanations of overall model knowledge. We also discuss metrics for evaluating generated explanations and discuss how explanations can be leveraged to debug models and improve performance. Lastly, we examine key challenges and emerging opportunities for explanation techniques in the era of LLMs in comparison to conventional deep learning models.<\/jats:p>","DOI":"10.1145\/3639372","type":"journal-article","created":{"date-parts":[[2024,1,2]],"date-time":"2024-01-02T21:59:20Z","timestamp":1704232760000},"page":"1-38","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":547,"title":["Explainability for Large Language Models: A Survey"],"prefix":"10.1145","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-5358-6895","authenticated-orcid":false,"given":"Haiyan","family":"Zhao","sequence":"first","affiliation":[{"name":"New Jersey Institute of Technology, Newark, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-5547-6634","authenticated-orcid":false,"given":"Hanjie","family":"Chen","sequence":"additional","affiliation":[{"name":"Johns Hopkins University, Baltimore, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3442-754X","authenticated-orcid":false,"given":"Fan","family":"Yang","sequence":"additional","affiliation":[{"name":"Wake Forest University, Winston-Salem, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9170-2424","authenticated-orcid":false,"given":"Ninghao","family":"Liu","sequence":"additional","affiliation":[{"name":"University of Georgia, Athens, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2946-6678","authenticated-orcid":false,"given":"Huiqi","family":"Deng","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7147-5666","authenticated-orcid":false,"given":"Hengyi","family":"Cai","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology, CAS, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9212-1947","authenticated-orcid":false,"given":"Shuaiqiang","family":"Wang","sequence":"additional","affiliation":[{"name":"Baidu Inc., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6813-680X","authenticated-orcid":false,"given":"Dawei","family":"Yin","sequence":"additional","affiliation":[{"name":"Baidu Inc., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1614-6069","authenticated-orcid":false,"given":"Mengnan","family":"Du","sequence":"additional","affiliation":[{"name":"New Jersey Institute of Technology, Newark, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,2,22]]},"reference":[{"key":"e_1_3_2_2_1","unstructured":"Ebtesam Almazrouei Hamza Alobeidli Abdulaziz Alshamsi Alessandro Cappelli Ruxandra Cojocaru Merouane Debbah Etienne Goffinet Daniel Heslow Julien Launay Quentin Malartic Badreddine Noune Baptiste Pannier and Guilherme Penedo. 2023. Falcon-40B: An open large language model with state-of-the-art performance. (2023). https:\/\/huggingface.co\/tiiuae\/falcon-40b"},{"key":"e_1_3_2_3_1","unstructured":"Anthropic. 2023. Decomposing Language Models Into Understandable Components? Retrieved from https:\/\/www.anthropic.com\/index\/decomposing-language-models-into-understandable-components. Accessed 24-11-2023."},{"key":"e_1_3_2_4_1","unstructured":"AnthropicAI. 2023. Introducing Claude. Retrieved from https:\/\/www.anthropic.com\/index\/introducing-claude"},{"key":"e_1_3_2_5_1","unstructured":"Omer Antverg and Yonatan Belinkov. 2022. On the Pitfalls of Analyzing Individual Neurons in Language Models. arXiv:2110.07483."},{"key":"e_1_3_2_6_1","doi-asserted-by":"crossref","unstructured":"Marianna Apidianaki and Aina Gar\u00ed Soler. 2021. ALL Dolphins Are Intelligent and SOME Are Friendly: Probing BERT for Nouns\u2019 Semantic Properties and their Prototypicality. arXiv:2110.06376.","DOI":"10.18653\/v1\/2021.blackboxnlp-1.7"},{"key":"e_1_3_2_7_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i7.16734"},{"key":"e_1_3_2_8_1","doi-asserted-by":"crossref","unstructured":"Pepa Atanasova Oana-Maria Camburu Christina Lioma Thomas Lukasiewicz Jakob Grue Simonsen and Isabelle Augenstein. 2023. Faithfulness Tests for Natural Language Explanations. arXiv:2305.18029.","DOI":"10.18653\/v1\/2023.acl-short.25"},{"key":"e_1_3_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467307"},{"key":"e_1_3_2_10_1","doi-asserted-by":"crossref","unstructured":"Hritik Bansal Karthik Gopalakrishnan Saket Dingliwal Sravan Bodapati Katrin Kirchhoff and Dan Roth. 2022. Rethinking the role of scale for in-context learning: An interpretability-based case study at 66 billion scale. arXiv:2212.09095.","DOI":"10.18653\/v1\/2023.acl-long.660"},{"key":"e_1_3_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482126"},{"key":"e_1_3_2_12_1","doi-asserted-by":"crossref","unstructured":"Jasmijn Bastings Sebastian Ebert Polina Zablotskaia Anders Sandholm and Katja Filippova. 2022. Will you find these shortcuts? A protocol for evaluating the faithfulness of input salience methods for text classification. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing 976\u2013991.","DOI":"10.18653\/v1\/2022.emnlp-main.64"},{"key":"e_1_3_2_13_1","unstructured":"Anthony Bau Yonatan Belinkov Hassan Sajjad Nadir Durrani Fahim Dalvi and James Glass. 2018. Identifying and Controlling Important Neurons in Neural Machine Translation. arXiv:1811.01157."},{"key":"e_1_3_2_14_1","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00422"},{"key":"e_1_3_2_15_1","first-page":"1","volume-title":"Proceedings of the 8th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Belinkov Yonatan","year":"2017","unstructured":"Yonatan Belinkov, Llu\u00eds M\u00e0rquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2017. Evaluating layers of representation in neural machine translation on part-of-speech and semantic tagging tasks. In Proceedings of the 8th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Asian Federation of Natural Language Processing, Taipei, Taiwan, 1\u201310. Retrieved from https:\/\/aclanthology.org\/I17-1001"},{"key":"e_1_3_2_16_1","unstructured":"Lukas Berglund Meg Tong Max Kaufmann Mikita Balesni Asa Cooper Stickland Tomasz Korbak and Owain Evans. 2023. The Reversal Curse: LLMs trained on \u201cA is B\u201d fail to learn \u201cB is A\u201d. arXiv:2309.12288."},{"key":"e_1_3_2_17_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.72"},{"key":"e_1_3_2_18_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.367"},{"key":"e_1_3_2_19_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-2003"},{"key":"e_1_3_2_20_1","unstructured":"Trenton Bricken Adly Templeton Joshua Batson Brian Chen Adam Jermyn Tom Conerly Nicholas L. Turner Cem Anil Carson Denison Amanda Askell Robert Lasenby Yifan Wu Shauna Kravec Nicholas Schiefer Tim Maxwell Nicholas Joseph Alex Tamkin Karina Nguyen Brayden McLean Josiah E. Burke Tristan Hume Shan Carter Tom Henighan and Chris Olah. 2023. Towards monosemanticity: Decomposing language models with dictionary learning. Retrieved from https:\/\/transformer-circuits.pub\/2023\/monosemantic-features\/index.html. Accessed 24-11-2023."},{"key":"e_1_3_2_21_1","article-title":"Language models are few-shot learners","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et\u00a0al. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems (NeurIPS) 33, 2020, 1877\u20131901.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_3_2_22_1","unstructured":"Gino Brunner Yang Liu Damian Pascual Oliver Richter Massimiliano Ciaramita and Roger Wattenhofer. 2019. On identifiability in transformers. arXiv:1908.04211."},{"key":"e_1_3_2_23_1","unstructured":"S\u00e9bastien Bubeck Varun Chandrasekaran Ronen Eldan Johannes Gehrke Eric Horvitz Ece Kamar Peter Lee Yin Tat Lee Yuanzhi Li Scott Lundberg Harsha Nori Hamid Palangi Marco Tulio Ribeiro and Yi Zhang. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv:2303.12712."},{"key":"e_1_3_2_24_1","unstructured":"Captum. 2022. Testing with Concept Activation Vectors (TCAV) on Sensitivity Classification Examples and a ConvNet Model Trained on IMDB DataSet. Retrieved from https:\/\/github.com\/pytorch\/captum\/blob\/master\/tutorials\/TCAV_NLP.ipynb. Accessed 24-11-2023."},{"key":"e_1_3_2_25_1","unstructured":"Nicholas Carlini Milad Nasr Christopher A. Choquette-Choo Matthew Jagielski Irena Gao Anas Awadalla Pang Wei Koh Daphne Ippolito Katherine Lee Florian Tramer and Ludwig Schmidt. 2023. Are aligned neural networks adversarially aligned? arXiv:2306.15447."},{"key":"e_1_3_2_26_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.bigscience-1.5"},{"key":"e_1_3_2_27_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.345"},{"key":"e_1_3_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00084"},{"key":"e_1_3_2_29_1","unstructured":"Angelica Chen Ravid Schwartz-Ziv Kyunghyun Cho Matthew L. Leavitt and Naomi Saphra. 2023c. Sudden drops in the loss: Syntax acquisition phase transitions and simplicity bias in MLMs. arXiv:2309.07311."},{"key":"e_1_3_2_30_1","unstructured":"Boli Chen Yao Fu Guangwei Xu Pengjun Xie Chuanqi Tan Mosha Chen and Liping Jing. 2021. Probing BERT in Hyperbolic Spaces. arXiv:2104.03869."},{"key":"e_1_3_2_31_1","article-title":"REV: Information-theoretic evaluation of free-text rationales","author":"Chen Hanjie","year":"2023","unstructured":"Hanjie Chen, Faeze Brahman, Xiang Ren, Yangfeng Ji, Yejin Choi, and Swabha Swayamdipta. 2023a. REV: Information-theoretic evaluation of free-text rationales. The 61th Annual Meeting of the Association for Computational Linguistics (ACL) (2023).","journal-title":"The 61th Annual Meeting of the Association for Computational Linguistics (ACL)"},{"key":"e_1_3_2_32_1","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-023-00657-x"},{"key":"e_1_3_2_33_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i10.21289"},{"key":"e_1_3_2_34_1","unstructured":"Yanda Chen Ruiqi Zhong Narutatsu Ri Chen Zhao He He Jacob Steinhardt Zhou Yu and Kathleen McKeown. 2023d. Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations. arXiv:2307.08678."},{"key":"e_1_3_2_35_1","unstructured":"Wei-Lin Chiang Zhuohan Li Zi Lin Ying Sheng Zhanghao Wu Hao Zhang Lianmin Zheng Siyuan Zhuang Yonghao Zhuang Joseph E. Gonzalez Ion Stoica and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. Retrieved from https:\/\/lmsys.org\/blog\/2023-03-30-vicuna\/. Accessed 21-08-2023."},{"key":"e_1_3_2_36_1","unstructured":"Aakanksha Chowdhery Sharan Narang Jacob Devlin Maarten Bosma Gaurav Mishra Adam Roberts Paul Barham Hyung Won Chung Charles Sutton Sebastian Gehrmann Parker Schuh Kensen Shi Sasha Tsvyashchenko Joshua Maynez Abhishek Rao Parker Barnes Yi Tay Noam Shazeer Vinodkumar Prabhakaran Emily Reif Nan Du Ben Hutchinson Reiner Pope James Bradbury Jacob Austin Michael Isard Guy Gur-Ari Pengcheng Yin Toju Duke Anselm Levskaya Sanjay Ghemawat Sunipa Dev Henryk Michalewski Xavier Garcia Vedant Misra Kevin Robinson Liam Fedus Denny Zhou Daphne Ippolito David Luan Hyeontaek Lim Barret Zoph Alexander Spiridonov Ryan Sepassi David Dohan Shivani Agrawal Mark Omernick Andrew M. Dai Thanumalayan Sankaranarayana Pillai Marie Pellat Aitor Lewkowycz Erica Moreira Rewon Child Oleksandr Polozov Katherine Lee Zongwei Zhou Xuezhi Wang Brennan Saeta Mark Diaz Orhan Firat Michele Catasta Jason Wei Kathy Meier-Hellstern Douglas Eck Jeff Dean Slav Petrov and Noah Fiedel. 2022. PaLM: Scaling Language Modeling with Pathways. Journal of Machine Learning Research 24 240 (2023) 1\u2013113."},{"key":"e_1_3_2_37_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.40"},{"key":"e_1_3_2_38_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W19-4828"},{"key":"e_1_3_2_39_1","article-title":"Electra: Pre-training text encoders as discriminators rather than generators","author":"Clark Kevin","year":"2020","unstructured":"Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. Electra: Pre-training text encoders as discriminators rather than generators. International Conference on Learning Representations (ICLR) (2020).","journal-title":"International Conference on Learning Representations (ICLR)"},{"key":"e_1_3_2_40_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33016309"},{"key":"e_1_3_2_41_1","unstructured":"Sanjoy Dasgupta Nave Frost and Michal Moshkovitz. 2022. Framework for evaluating faithfulness of local explanations. International Conference on Machine Learning (ICMR\u201922). PMLR 4794\u20134815."},{"key":"e_1_3_2_42_1","doi-asserted-by":"crossref","unstructured":"Joseph F. DeRose Jiayao Wang and Matthew Berger. 2020. Attention flows: Analyzing and comparing attention mechanisms in language models. IEEE Transactions on Visualization and Computer Graphics 27 2 (2020) 1160\u20131170.","DOI":"10.1109\/TVCG.2020.3028976"},{"key":"e_1_3_2_43_1","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Vol. 1 4171\u20134186."},{"key":"e_1_3_2_44_1","volume-title":"Proceedings of the North American Chapter of the Association for Computational Linguistics (NAACL \u201919)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the North American Chapter of the Association for Computational Linguistics (NAACL \u201919)."},{"key":"e_1_3_2_45_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.408"},{"key":"e_1_3_2_46_1","unstructured":"Finale Doshi-Velez and Been Kim. 2017. Towards a rigorous science of interpretable machine learning. stat 1050 (2017) 2."},{"key":"e_1_3_2_47_1","unstructured":"Timothy Dozat and Christopher D. Manning. 2016. Deep Biaffine Attention for Neural Dependency Parsing. arXiv: 1611.01734."},{"key":"e_1_3_2_48_1","article-title":"Shortcut learning of large language models in natural language understanding","author":"Du Mengnan","year":"2023","unstructured":"Mengnan Du, Fengxiang He, Na Zou, Dacheng Tao, and Xia Hu. 2023. Shortcut learning of large language models in natural language understanding. Communications of the ACM (CACM) 67, 1 (2023), 110\u2013120,","journal-title":"Communications of the ACM (CACM)"},{"key":"e_1_3_2_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3359786"},{"key":"e_1_3_2_50_1","unstructured":"Mengnan Du Ninghao Liu Fan Yang Shuiwang Ji and Xia Hu. 2019b. On attribution of recurrent neural network predictions via additive decomposition. The World Wide Web Conference (WWW\u201919) 383\u2013393."},{"key":"e_1_3_2_51_1","unstructured":"Mengnan Du Varun Manjunatha Rajiv Jain Ruchi Deshpande Franck Dernoncourt Jiuxiang Gu Tong Sun and Xia Hu. 2021. Towards interpreting and mitigating shortcut learning behavior of NLU models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 915\u2013929."},{"key":"e_1_3_2_52_1","doi-asserted-by":"crossref","unstructured":"Jinhao Duan Hao Cheng Shiqi Wang Chenan Wang Alex Zavalny Renjing Xu Bhavya Kailkhura and Kaidi Xu. 2023. Shifting attention to relevance: Towards the uncertainty estimation of large language models. arXiv:2307.01379.","DOI":"10.18653\/v1\/2024.acl-long.276"},{"key":"e_1_3_2_53_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.387"},{"key":"e_1_3_2_54_1","article-title":"A Mathematical Framework for Transformer Circuits \u2014 transformer-circuits.pub","author":"Elhage Nelson","year":"2021","unstructured":"Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021. A Mathematical Framework for Transformer Circuits \u2014 transformer-circuits.pub. Retrieved from https:\/\/transformer-circuits.pub\/2021\/framework\/index.html. [Accessed 27-11-2023].","journal-title":"R"},{"key":"e_1_3_2_55_1","doi-asserted-by":"crossref","unstructured":"Joseph Enguehard. 2023. Sequential Integrated Gradients: A Simple but Effective Method for Explaining Language Models. arXiv:2305.15853.","DOI":"10.18653\/v1\/2023.findings-acl.477"},{"key":"e_1_3_2_56_1","doi-asserted-by":"crossref","unstructured":"Kawin Ethayarajh and Dan Jurafsky. 2021. Attention Flows are Shapley Value Explanations. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing 2 (2021) 49\u201354.","DOI":"10.18653\/v1\/2021.acl-short.8"},{"key":"e_1_3_2_57_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1407"},{"key":"e_1_3_2_58_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.498"},{"key":"e_1_3_2_59_1","doi-asserted-by":"crossref","unstructured":"Mor Geva Avi Caciularu Kevin Wang and Yoav Goldberg. 2022. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing 30\u201345.","DOI":"10.18653\/v1\/2022.emnlp-main.3"},{"key":"e_1_3_2_60_1","doi-asserted-by":"crossref","unstructured":"Mor Geva Roei Schuster Jonathan Berant and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 5484\u20135495.","DOI":"10.18653\/v1\/2021.emnlp-main.446"},{"key":"e_1_3_2_61_1","first-page":"2242","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Ghorbani Amirata","year":"2019","unstructured":"Amirata Ghorbani and James Zou. 2019. Data shapley: Equitable valuation of data for machine learning. In Proceedings of the International Conference on Machine Learning. PMLR, 2242\u20132251."},{"key":"e_1_3_2_62_1","unstructured":"Olga Golovneva Moya Chen Spencer Poff Martin Corredor Luke Zettlemoyer Maryam Fazel-Zarandi and Asli Celikyilmaz. 2022. Roscoe: A suite of metrics for scoring step-by-step reasoning. arXiv:2212.07919."},{"key":"e_1_3_2_63_1","unstructured":"Roger Grosse Juhan Bae Cem Anil Nelson Elhage Alex Tamkin Amirhossein Tajdini Benoit Steiner Dustin Li Esin Durmus Ethan Perez Kamil\u0117 Luko\u0161i\u016bt\u0117 Karina Nguyen Nicholas Joseph Sam McCandlish Jared Kaplan and Samuel R. Bowman. 2023. Studying large language model generalization with influence functions. arXiv:2308.03296. Retrieved from https:\/\/arxiv.org\/abs\/2308.03296"},{"key":"e_1_3_2_64_1","unstructured":"Arnav Gudibande Eric Wallace Charlie Snell Xinyang Geng Hao Liu Pieter Abbeel Sergey Levine and Dawn Song. 2023. The false promise of imitating proprietary LLMs. arXiv:2305.15717."},{"key":"e_1_3_2_65_1","doi-asserted-by":"crossref","unstructured":"Han Guo Nazneen Fatema Rajani Peter Hase Mohit Bansal and Caiming Xiong. 2020. Fastif: Scalable influence functions for efficient model interpretation and debugging. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing 10333\u201310350.","DOI":"10.18653\/v1\/2021.emnlp-main.808"},{"key":"e_1_3_2_66_1","unstructured":"Wes Gurnee and Max Tegmark. 2023. Language models represent space and time. arXiv:2310.02207."},{"key":"e_1_3_2_67_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i14.17533"},{"key":"e_1_3_2_68_1","article-title":"Data cleansing for models trained with SGD","volume":"32","author":"Hara Satoshi","year":"2019","unstructured":"Satoshi Hara, Atsushi Nitanda, and Takanori Maehara. 2019. Data cleansing for models trained with SGD. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_69_1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"He Pengcheng","year":"2021","unstructured":"Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. DeBERTa: Decoding-enhanced BERT with disentangled attention. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_70_1","unstructured":"Danny Hernandez Tom Brown Tom Conerly Nova DasSarma Dawn Drain Sheer El-Showk Nelson Elhage Zac Hatfield-Dodds Tom Henighan Tristan Hume Scott Johnston Ben Mann Chris Olah Catherine Olsson Dario Amodei Nicholas Joseph Jared Kaplan and Sam McCandlish. 2022. Scaling Laws and Interpretability of Learning from Repeated Data. arXiv:2205.10487."},{"key":"e_1_3_2_71_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1275"},{"key":"e_1_3_2_72_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1419"},{"key":"e_1_3_2_73_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-demos.22"},{"key":"e_1_3_2_74_1","unstructured":"Yuheng Huang Jiayang Song Zhijie Wang Huaming Chen and Lei Ma. 2023. Look before you leap: An exploratory study of uncertainty measurement for large language models. arXiv:2307.10236."},{"key":"e_1_3_2_75_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.386"},{"key":"e_1_3_2_76_1","unstructured":"Sarthak Jain and Byron C. Wallace. 2019. Attention is not explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 1 (2019) 3543\u20133556."},{"key":"e_1_3_2_77_1","doi-asserted-by":"crossref","unstructured":"Theo Jaunet Corentin Kervadec Romain Vuillemot Grigory Antipov Moez Baccouche and Christian Wolf. 2021. VisQA: X-raying vision and language reasoning in transformers. IEEE Transactions on Visualization and Computer Graphics 28 1 (2021) 976\u2013986.","DOI":"10.1109\/TVCG.2021.3114683"},{"key":"e_1_3_2_78_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1356"},{"key":"e_1_3_2_79_1","doi-asserted-by":"crossref","unstructured":"Di Jin Zhijing Jin Joey Tianyi Zhou and Peter Szolovits. 2020. Is bert really robust? natural language attack on text classification and entailment. In Proceedings of the AAAI Conference on Artificial Intelligence 34 5 (2020) 8018\u20138025.","DOI":"10.1609\/aaai.v34i05.6311"},{"key":"e_1_3_2_80_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-emnlp.242"},{"key":"e_1_3_2_81_1","unstructured":"Nikhil Kandpal Haikang Deng Adam Roberts Eric Wallace and Colin Raffel. 2023. Large Language Models Struggle to Learn Long-Tail Knowledge. In International Conference on Machine Learning. PMLR 15696\u201315707."},{"key":"e_1_3_2_82_1","unstructured":"Cheongwoong Kang and Jaesik Choi. 2023. Impact of Co-occurrence on Factual Knowledge of Large Language Models. arXiv:2310.08256. Retrieved from https:\/\/arxiv.org\/abs\/2310.08256"},{"key":"e_1_3_2_83_1","unstructured":"Divyansh Kaushik Eduard Hovy and Zachary C. Lipton. 2020. Learning the difference that makes a difference with counterfactually-augmented data. arXiv preprint arXiv:1909.12434 (2019)."},{"key":"e_1_3_2_84_1","unstructured":"Been Kim Martin Wattenberg Justin Gilmer Carrie Cai James Wexler Fernanda Viegas and Rory Sayres. 2018. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV). In International Conference on Machine Learning PMLR 2668\u20132677."},{"key":"e_1_3_2_85_1","doi-asserted-by":"crossref","unstructured":"Pieter-Jan Kindermans Sara Hooker Julius Adebayo Maximilian Alber Kristof T. Sch\u00fctt Sven D\u00e4hne Dumitru Erhan and Been Kim. 2017. The (Un)reliability of saliency methods. Explainable AI: Interpreting Explaining and Visualizing Deep Learning Springer 267\u2013280.","DOI":"10.1007\/978-3-030-28954-6_14"},{"key":"e_1_3_2_86_1","first-page":"1885","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Koh Pang Wei","year":"2017","unstructured":"Pang Wei Koh and Percy Liang. 2017. Understanding black-box predictions via influence functions. In Proceedings of the International Conference on Machine Learning. PMLR, 1885\u20131894."},{"key":"e_1_3_2_87_1","first-page":"16","volume-title":"Proceedings of the EACL Hackashop on News Media Content Analysis and Automated Report Generation","author":"Kokalj Enja","year":"2021","unstructured":"Enja Kokalj, Bla\u017e \u0160krlj, Nada Lavra\u010d, Senja Pollak, and Marko Robnik-\u0160ikonja. 2021. BERT meets shapley: Extending SHAP explanations to transformer-based classifiers. In Proceedings of the EACL Hackashop on News Media Content Analysis and Automated Report Generation. Association for Computational Linguistics, Online, 16\u201321. Retrieved from https:\/\/aclanthology.org\/2021.hackashop-1.3"},{"key":"e_1_3_2_88_1","doi-asserted-by":"crossref","unstructured":"Olga Kovaleva Alexey Romanov Anna Rogers and Anna Rumshisky. 2019. Revealing the dark secrets of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) 4365\u20134374.","DOI":"10.18653\/v1\/D19-1445"},{"key":"e_1_3_2_89_1","unstructured":"Nicholas Kroeger Dan Ley Satyapriya Krishna Chirag Agarwal and Himabindu Lakkaraju. 2023. Are large language models Post Hoc Explainers? arXiv:2310.05797."},{"key":"e_1_3_2_90_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.450"},{"key":"e_1_3_2_91_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.eacl-main.234"},{"key":"e_1_3_2_92_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-emnlp.38"},{"key":"e_1_3_2_93_1","unstructured":"Tamera Lanham Anna Chen Ansh Radhakrishnan Benoit Steiner Carson Denison Danny Hernandez Dustin Li Esin Durmus Evan Hubinger Jackson Kernion Kamil\u0117 Luko\u0161i\u016bt\u0117 Karina Nguyen Newton Cheng Nicholas Joseph Nicholas Schiefer Oliver Rausch Robin Larson Sam McCandlish Sandipan Kundu Saurav Kadavath Shannon Yang Thomas Henighan Timothy Maxwell Timothy Telleen-Lawton Tristan Hume Zac Hatfield-Dodds Jared Kaplan Jan Brauner Samuel R. Bowman and Ethan Perez. 2023. Measuring faithfulness in chain-of-thought reasoning. arXiv:2307.13702. Retrieved from https:\/\/arxiv.org\/abs\/2307.13702"},{"key":"e_1_3_2_94_1","unstructured":"Dong-Ho Lee Akshen Kadakia Brihi Joshi Aaron Chan Ziyi Liu Kiran Narahari Takashi Shibuya Ryosuke Mitani Toshiyuki Sekiya Jay Pujara and Xiang Ren. 2022. XMD: An End-to-End Framework for Interactive Explanation-Based Debugging of NLP Models. arXiv:2210.16978. Retrieved from https:\/\/arxiv.org\/abs\/2210.16978"},{"key":"e_1_3_2_95_1","doi-asserted-by":"crossref","unstructured":"Bai Li Zining Zhu Guillaume Thomas Yang Xu and Frank Rudzicz. 2021b. How is BERT Surprised? Layerwise Detection of Linguistic Anomalies. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing 1 2021 4215\u20134228.","DOI":"10.18653\/v1\/2021.acl-long.325"},{"key":"e_1_3_2_96_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.400"},{"key":"e_1_3_2_97_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16-1082"},{"key":"e_1_3_2_98_1","doi-asserted-by":"crossref","unstructured":"Jiaoda Li Ryan Cotterell and Mrinmaya Sachan. 2022. Probing via Prompting. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 1144\u20131157.","DOI":"10.18653\/v1\/2022.naacl-main.84"},{"key":"e_1_3_2_99_1","unstructured":"Jiwei Li Will Monroe and Dan Jurafsky. 2017. Understanding Neural Networks through Representation Erasure. arXiv:1612.08220."},{"key":"e_1_3_2_100_1","unstructured":"Yingji Li Mengnan Du Rui Song Xin Wang and Ying Wang. 2023a. A survey on fairness in large language models. arXiv:2308.10149."},{"key":"e_1_3_2_101_1","unstructured":"Zongxia Li Paiheng Xu Fuxiao Liu and Hyemi Song. 2023b. Towards Understanding In-Context Learning with Contrastive Demonstrations and Saliency Maps. arXiv:2307.05052."},{"key":"e_1_3_2_102_1","unstructured":"Tom Lieberum Matthew Rahtz J\u00e1nos Kram\u00e1r Geoffrey Irving Rohin Shah and Vladimir Mikulik. 2023. Does circuit analysis interpretability scale? Evidence from multiple choice capabilities in chinchilla. arXiv:2307.09458."},{"key":"e_1_3_2_103_1","unstructured":"Yongjie Lin Yi Chern Tan and Robert Frank. 2019. Open sesame: Getting inside BERT\u2019s linguistic knowledge. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP 241\u2013253."},{"key":"e_1_3_2_104_1","unstructured":"Jiaxiang Liu Tianxiang Hu Yan Zhang Xiaotang Gai Yang Feng and Zuozhu Liu. 2023a. A ChatGPT Aided Explainable Framework for Zero-Shot Medical Image Diagnosis. arXiv:2307.01981."},{"key":"e_1_3_2_105_1","unstructured":"Ninghao Liu Yunsong Meng Xia Hu Tie Wang and Bo Long. 2020. Are interpretations fairly evaluated? A definition driven pipeline for post-hoc interpretability. arXiv:2009.07494."},{"key":"e_1_3_2_106_1","unstructured":"Nelson F. Liu Kevin Lin John Hewitt Ashwin Paranjape Michele Bevilacqua Fabio Petroni and Percy Liang. 2023b. Lost in the middle: How language models use long contexts. arXiv:2307.03172."},{"key":"e_1_3_2_107_1","unstructured":"Yibing Liu Haoliang Li Yangyang Guo Chenqi Kong Jing Li and Shiqi Wang. 2022. Rethinking attention-model explainability through faithfulness violation test. In International Conference on Machine Learning. PMLR 2022 13807\u201313824."},{"key":"e_1_3_2_108_1","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv:1907.11692."},{"key":"e_1_3_2_109_1","article-title":"A unified approach to interpreting model predictions","volume":"30","author":"Lundberg Scott M.","year":"2017","unstructured":"Scott M. Lundberg and Su-In Lee. 2017a. A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_110_1","volume-title":"Advances in Neural Information Processing Systems","author":"Lundberg Scott M.","year":"2017","unstructured":"Scott M. Lundberg and Su-In Lee. 2017b. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems, Vol. 30. Curran Associates, Inc. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2017\/hash\/8a20a8621978632d76c43dfd28b67767-Abstract.html"},{"key":"e_1_3_2_111_1","first-page":"14485","volume-title":"Proceedings of the 39th International Conference on Machine Learning","author":"Lundstrom Daniel D.","year":"2022","unstructured":"Daniel D. Lundstrom, Tianjian Huang, and Meisam Razaviyayn. 2022. A rigorous study of integrated gradients method and extensions to internal neuron attributions. In Proceedings of the 39th International Conference on Machine Learning. PMLR, 14485\u201314508. Retrieved from https:\/\/proceedings.mlr.press\/v162\/lundstrom22a.htmlISSN: 2640-3498."},{"key":"e_1_3_2_112_1","unstructured":"Siwen Luo Hamish Ivison Caren Han and Josiah Poon. 2022. Local Interpretations for Explainable Natural Language Processing: A Survey. arXiv:2103.11072."},{"key":"e_1_3_2_113_1","doi-asserted-by":"crossref","unstructured":"Satyapriya Krishna Jiaqi Ma Dylan Slack Asma Ghandeharioun Sameer Singh and Himabindu Lakkaraju. 2023. Post Hoc explanations of language models can improve language models. arXiv:2305.11426.","DOI":"10.21203\/rs.3.rs-3006112\/v1"},{"key":"e_1_3_2_114_1","unstructured":"Aman Madaan and Amir Yazdanbakhsh. 2022. Text and patterns: For effective chain of thought it takes two to tango. arXiv:2209.07686. Retrieved from https:\/\/arxiv.org\/abs\/2209.07686"},{"key":"e_1_3_2_115_1","unstructured":"Samuel Marks and Max Tegmark. 2023. The geometry of truth: Emergent linear structure in large language model representations of true\/false datasets. arXiv:2310.06824."},{"key":"e_1_3_2_116_1","unstructured":"Sammy Martin. 2023. Ten Levels of AI Alignment Difficulty. Retrieved from https:\/\/www.lesswrong.com\/posts\/EjgfreeibTXRx9Ham\/ten-levels-of-ai-alignment-difficulty. Accessed 21-08-2023."},{"key":"e_1_3_2_117_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1151"},{"key":"e_1_3_2_118_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i17.17745"},{"key":"e_1_3_2_119_1","doi-asserted-by":"crossref","unstructured":"Rowan Hall Maudslay and Ryan Cotterell. 2021. Do Syntactic Probes Probe Syntax? Experiments with Jabberwocky Probing. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 124\u2013131.","DOI":"10.18653\/v1\/2021.naacl-main.11"},{"key":"e_1_3_2_120_1","unstructured":"Thomas McGrath Matthew Rahtz Janos Kramar Vladimir Mikulik and Shane Legg. 2023. The hydra effect: Emergent self-repair in language model computations. arXiv:2307.15771."},{"key":"e_1_3_2_121_1","doi-asserted-by":"crossref","unstructured":"Nick McKenna Tianyi Li Liang Cheng Mohammad Javad Hosseini Mark Johnson and Mark Steedman. 2023. Sources of Hallucination by Large Language Models on Inference Tasks. arXiv:2305.14552.","DOI":"10.18653\/v1\/2023.findings-emnlp.182"},{"key":"e_1_3_2_122_1","unstructured":"Yusuf Mehdi. 2023. Reinventing Search with a New AI-powered Microsoft Bing and Edge Your Copilot for the Web. Retrieved from https:\/\/blogs.microsoft.com\/blog\/2023\/02\/07\/reinventing-search-with-a-new-ai-powered-microsoft-bing-and-edge-your-copilot-for-the-web\/. Accessed 21-08-2023."},{"key":"e_1_3_2_123_1","first-page":"17359","article-title":"Locating and editing factual associations in GPT","volume":"35","author":"Meng Kevin","year":"2022","unstructured":"Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. Advances in Neural Information Processing Systems 35 (2022), 17359\u201317372.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_124_1","unstructured":"Vivek Miglani Narine Kokhlikyan Bilal Alsallakh Miguel Martin and Orion Reblitz-Richardson. 2020. Investigating Saturation Effects in Integrated Gradients. arXiv:2010.12697."},{"key":"e_1_3_2_125_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.387"},{"key":"e_1_3_2_126_1","doi-asserted-by":"crossref","unstructured":"Hosein Mohebbi Ali Modarressi and Mohammad Taher Pilehvar. 2021. Exploring the role of BERT token representations to explain sentence probing results. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing 792\u2013806.","DOI":"10.18653\/v1\/2021.emnlp-main.61"},{"key":"e_1_3_2_127_1","doi-asserted-by":"crossref","unstructured":"Gr\u00e9goire Montavon Sebastian Bach Alexander Binder Wojciech Samek and Klaus-Robert M\u00fcller. 2015. Explaining nonlinear classification decisions with deep taylor decomposition. Pattern Recognition 65 (2017) 211\u2013222.","DOI":"10.1016\/j.patcog.2016.11.008"},{"key":"e_1_3_2_128_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-28954-6_10"},{"key":"e_1_3_2_129_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.eacl-main.243"},{"key":"e_1_3_2_130_1","unstructured":"Jesse Mu and Jacob Andreas. 2021. Compositional explanations of neurons. Advances in Neural Information Processing Systems 33 (2020) 17153\u201317163."},{"key":"e_1_3_2_131_1","unstructured":"Subhabrata Mukherjee Arindam Mitra Ganesh Jawahar Sahaj Agarwal Hamid Palangi and Ahmed Awadallah. 2023. Orca: Progressive learning from complex explanation traces of gpt-4. arXiv:2306.02707."},{"key":"e_1_3_2_132_1","unstructured":"Michael Neely Stefan F. Schouten Maurits J. R. Bleeker and Ana Lucic. 2021. Order in the court: Explainable ai methods prone to disagreement. arXiv:2105.03287."},{"key":"e_1_3_2_133_1","unstructured":"Maxwell Nye Anders Johan Andreassen Guy Gur-Ari Henryk Michalewski Jacob Austin David Bieber David Dohan Aitor Lewkowycz Maarten Bosma David Luan Charles Sutton and Augustus Odena. 2021. Show your work: Scratchpads for intermediate computation with language models. arXiv:2112.00114. Retrieved from https:\/\/arxiv.org\/abs\/2112.00114"},{"key":"e_1_3_2_134_1","article-title":"Zoom In: An Introduction to Circuits","author":"Olah Chris","year":"2020","unstructured":"Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. 2020a. Zoom In: An Introduction to Circuits. Retrieved from https:\/\/distill.pub\/2020\/circuits\/zoom-in\/. Accessed 24-11-2023.","journal-title":"R"},{"key":"e_1_3_2_135_1","article-title":"Naturally Occurring Equivariance in Neural Networks \u2014 distill.pub","author":"Olah Chris","year":"2020","unstructured":"Chris Olah, Nick Cammarata, Chelsea Voss, Ludwig Schubert, and Gabriel Goh. 2020b. Naturally Occurring Equivariance in Neural Networks \u2014 distill.pub. Retrieved from https:\/\/distill.pub\/2020\/circuits\/equivariance\/. Accessed 27-11-2023.","journal-title":"R"},{"key":"e_1_3_2_136_1","article-title":"In-context Learning and Induction Heads \u2014 transformer-circuits.pub","author":"Olsson Catherine","year":"2022","unstructured":"Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2022. In-context Learning and Induction Heads \u2014 transformer-circuits.pub. Retrieved from https:\/\/transformer-circuits.pub\/2022\/in-context-learning-and-induction-heads\/index.html. Accessed 27-11-2023.","journal-title":"R"},{"key":"e_1_3_2_137_1","unstructured":"OpenAI. 2023a. GPT-4 Technical Report. arXiv:2303.08774."},{"key":"e_1_3_2_138_1","volume-title":"Language Models Can Explain Neurons in Language Models","year":"2023","unstructured":"OpenAI. 2023b. Language Models Can Explain Neurons in Language Models. Retrieved from https:\/\/openai.com\/research\/language-models-can-explain-neurons-in-language-models?s=09"},{"key":"e_1_3_2_139_1","doi-asserted-by":"crossref","unstructured":"Cheonbok Park Inyoup Na Yongjang Jo Sungbok Shin Jaehyo Yoo Bum Chul Kwon Jian Zhao Hyungjong Noh Yeonsoo Lee and Jaegul Choo. 2019. SANVis: Visual analytics for understanding self-attention networks. IEEE Visualization Conference (VIS\u201919) IEEE 146\u2013150.","DOI":"10.1109\/VISUAL.2019.8933677"},{"key":"e_1_3_2_140_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1179"},{"key":"e_1_3_2_141_1","doi-asserted-by":"crossref","unstructured":"Fabio Petroni Tim Rockt\u00e4schel Patrick Lewis Anton Bakhtin Yuxiang Wu Alexander H. Miller and Sebastian Riedel. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) 2463\u20132473.","DOI":"10.18653\/v1\/D19-1250"},{"key":"e_1_3_2_142_1","article-title":"Weight Banding \u2014 distill.pub","author":"Petrov Michael","year":"2021","unstructured":"Michael Petrov, Chelsea Voss, Ludwig Schubert, Nick Cammarata, Gabriel Goh, and Chris Olah. 2021. Weight Banding \u2014 distill.pub. https:\/\/distill.pub\/2020\/circuits\/weight-banding\/. Accessed 27-11-2023.","journal-title":"https:\/\/distill.pub\/2020\/circuits\/weight-banding\/"},{"key":"e_1_3_2_143_1","doi-asserted-by":"crossref","unstructured":"Archiki Prasad Swarnadeep Saha Xiang Zhou and Mohit Bansal. 2023. ReCEval: Evaluating reasoning chains via correctness and informativeness. arXiv:2304.10703.","DOI":"10.18653\/v1\/2023.emnlp-main.622"},{"key":"e_1_3_2_144_1","first-page":"19920","article-title":"Estimating training data influence by tracing gradient descent","volume":"33","author":"Pruthi Garima","year":"2020","unstructured":"Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan. 2020. Estimating training data influence by tracing gradient descent. Advances in Neural Information Processing Systems 33 (2020), 19920\u201319930.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_145_1","unstructured":"Luyu Qiu Yi Yang Caleb Chen Cao Jing Liu Yueyuan Zheng Hilary Hei Ting Ngai Janet Hsiao and Lei Chen. 2021. Resisting Out-of-Distribution Data Problem in Perturbation of XAI. arXiv:2107.14000."},{"key":"e_1_3_2_146_1","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, and Gretchen Krueger Ilya. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_2_147_1","unstructured":"Ansh Radhakrishnan Karina Nguyen Anna Chen Carol Chen Carson Denison Danny Hernandez Esin Durmus Evan Hubinger Jackson Kernion Kamil\u0117 Luko\u0161i\u016bt\u0117 Newton Cheng Nicholas Joseph Nicholas Schiefer Oliver Rausch Sam McCandlish Sheer El Showk Tamera Lanham Tim Maxwell Venkatesa Chandrasekaran Zac Hatfield-Dodds Jared Kaplan Jan Brauner Samuel R. Bowman and Ethan Perez. 2023. Question decomposition improves the faithfulness of model-generated reasoning. arXiv:2307.11768. Retrieved from https:\/\/arxiv.org\/abs\/2307.11768"},{"key":"e_1_3_2_148_1","doi-asserted-by":"publisher","DOI":"10.5555\/3455716.3455856"},{"key":"e_1_3_2_149_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1487"},{"key":"e_1_3_2_150_1","first-page":"88","volume-title":"Proceedings of the 9th Joint Conference on Lexical and Computational Semantics","author":"Ravichander Abhilasha","year":"2020","unstructured":"Abhilasha Ravichander, Eduard Hovy, Kaheer Suleman, Adam Trischler, and Jackie Chi Kit Cheung. 2020. On the systematicity of probing contextualized word representations: The case of hypernymy in BERT. In Proceedings of the 9th Joint Conference on Lexical and Computational Semantics. Association for Computational Linguistics, Barcelona, Spain (Online), 88\u2013102. Retrieved from https:\/\/aclanthology.org\/2020.starsem-1.10"},{"key":"e_1_3_2_151_1","unstructured":"Ruiyang Ren Yuhao Wang Yingqi Qu Wayne Xin Zhao Jing Liu Hao Tian Hua Wu Ji-Rong Wen and Haifeng Wang. 2023. Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation. arXiv:2307.11019."},{"key":"e_1_3_2_152_1","doi-asserted-by":"crossref","unstructured":"Marco Tulio Ribeiro Sameer Singh and Carlos Guestrin. 2016. \u201cWhy should i trust you?\u201d: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 1135\u20131144.","DOI":"10.1145\/2939672.2939778"},{"key":"e_1_3_2_153_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00349"},{"key":"e_1_3_2_154_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-acl.336"},{"key":"e_1_3_2_155_1","doi-asserted-by":"crossref","unstructured":"Soumya Sanyal and Xiang Ren. 2021. Discretized integrated gradients for explaining language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing 10285\u201310299.","DOI":"10.18653\/v1\/2021.emnlp-main.805"},{"key":"e_1_3_2_156_1","article-title":"Interpretability Creationism","author":"Saphra Naomi","year":"2022","unstructured":"Naomi Saphra. 2022. Interpretability Creationism. Retrieved from https:\/\/nsaphra.github.io\/post\/creationism\/. [Accessed 22-10-2023].","journal-title":"R"},{"key":"e_1_3_2_157_1","doi-asserted-by":"crossref","unstructured":"Sofia Serrano and Noah A. Smith. 2019. Is attention interpretable? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics 2931\u20132951.","DOI":"10.18653\/v1\/P19-1282"},{"key":"e_1_3_2_158_1","doi-asserted-by":"crossref","unstructured":"Lloyd S. Shapley et\u00a0al. 1953. A value for n-person games. (1953).","DOI":"10.1515\/9781400881970-018"},{"key":"e_1_3_2_159_1","unstructured":"Yaozong Shen Lijie Wang Ying Chen Xinyan Xiao Jing Liu and Hua Wu. 2022. An Interpretability Evaluation Benchmark for Pre-trained Language Models. arXiv:2207.13948."},{"key":"e_1_3_2_160_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.71"},{"key":"e_1_3_2_161_1","unstructured":"Chandan Singh Aliyah R. Hsu Richard Antonello Shailee Jain Alexander G. Huth Bin Yu and Jianfeng Gao. 2023. Explaining black box text modules in natural language with language models. arXiv:2305.09863."},{"key":"e_1_3_2_162_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.384"},{"key":"e_1_3_2_163_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i10.21386"},{"key":"e_1_3_2_164_1","doi-asserted-by":"crossref","unstructured":"Hendrik Strobelt Sebastian Gehrmann Michael Behrisch Adam Perer Hanspeter Pfister and Alexander M. Rush. 2018. Seq2Seq-Vis: A visual debugging tool for sequence-to-sequence models. IEEE Transactions on Visualization and Computer Graphics 25 1 (2018) 353\u2013363.","DOI":"10.1109\/TVCG.2018.2865044"},{"key":"e_1_3_2_165_1","article-title":"Axiomatic attribution for deep networks","author":"Sundararajan Mukund","year":"2017","unstructured":"Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017. Axiomatic attribution for deep networks. International Conference on Machine Learning (ICML). PMLR, (2017), 3319\u20133328.","journal-title":"International Conference on Machine Learning (ICML)"},{"issue":"6","key":"e_1_3_2_166_1","first-page":"7","article-title":"Alpaca: A strong, replicable instruction-following model","volume":"3","author":"Taori Rohan","year":"2023","unstructured":"Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpaca: A strong, replicable instruction-following model. Stanford Center for Research on Foundation Models. 3, 6 (2023), 7. https:\/\/crfm.stanford.edu\/2023\/03\/13\/alpaca.html","journal-title":"Stanford Center for Research on Foundation Models"},{"key":"e_1_3_2_167_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1452"},{"key":"e_1_3_2_168_1","unstructured":"Ian Tenney Patrick Xia Berlin Chen Alex Wang Adam Poliak R. Thomas McCoy Najoung Kim Benjamin Van Durme Samuel R. Bowman Dipanjan Das et\u00a0al. 2019b. What do you learn from context? Probing for sentence structure in contextualized word representations. arXiv preprint arXiv:1905.06316 (2019)."},{"key":"e_1_3_2_169_1","doi-asserted-by":"crossref","unstructured":"James Thorne Andreas Vlachos Christos Christodoulopoulos and Arpit Mittal. 2019. Generating token-level explanations for natural language inference. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Volume 1 (Long and Short Papers) (2019) 963\u2013969.","DOI":"10.18653\/v1\/N19-1101"},{"key":"e_1_3_2_170_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.15"},{"key":"e_1_3_2_171_1","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar Aurelien Rodriguez Armand Joulin Edouard Grave and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_2_172_1","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale Dan Bikel Lukas Blecher Cristian Canton Ferrer Moya Chen Guillem Cucurull David Esiobu Jude Fernandes Jeremy Fu Wenyin Fu Brian Fuller Cynthia Gao Vedanuj Goswami Naman Goyal Anthony Hartshorn Saghar Hosseini Rui Hou Hakan Inan Marcin Kardas Viktor Kerkez Madian Khabsa Isabel Kloumann Artem Korenev Punit Singh Koura Marie-Anne Lachaux Thibaut Lavril Jenya Lee Diana Liskovich Yinghai Lu Yuning Mao Xavier Martinet Todor Mihaylov Pushkar Mishra Igor Molybog Yixin Nie Andrew Poulton Jeremy Reizenstein Rashi Rungta Kalyan Saladi Alan Schelten Ruan Silva Eric Michael Smith Ranjan Subramanian Xiaoqing Ellen Tan Binh Tang Ross Taylor Adina Williams Jian Xiang Kuan Puxin Xu Zheng Yan Iliyan Zarov Yuchen Zhang Angela Fan Melanie Kambadur Sharan Narang Aurelien Rodriguez Robert Stojnic Sergey Edunov and Thomas Scialom. 2023. LLAMa-2: Open foundation and fine-tuned chat models. (2023). Retrieved from https:\/\/ai.meta.com\/research\/publications\/llama-2-open-foundation-and-fine-tuned-chat-models\/"},{"key":"e_1_3_2_173_1","doi-asserted-by":"crossref","unstructured":"Marcos Treviso Alexis Ross Nuno M. Guerreiro and Andr\u00e9 F. T. Martins. 2023. CREST: A joint framework for rationalization and counterfactual text generation. (2023). arXiv preprint arXiv:2305.17075.","DOI":"10.18653\/v1\/2023.acl-long.842"},{"key":"e_1_3_2_174_1","unstructured":"Miles Turpin Julian Michael Ethan Perez and Samuel R. Bowman. 2023. Language models don\u2019t always say what they think: Unfaithful explanations in chain-of-thought prompting. arXiv:2305.04388."},{"key":"e_1_3_2_175_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3358028"},{"key":"e_1_3_2_176_1","article-title":"Residual networks behave like ensembles of relatively shallow networks","volume":"29","author":"Veit Andreas","year":"2016","unstructured":"Andreas Veit, Michael J. Wilber, and Serge Belongie. 2016. Residual networks behave like ensembles of relatively shallow networks. Advances in Neural Information Processing Systems 29 (2016).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_177_1","unstructured":"Jesse Vig. 2019. BertViz: A tool for visualizing multi-head self-attention in the BERT model. In ICLR Workshop: Debugging Machine Learning Models Vol. 23 353\u2013355."},{"key":"e_1_3_2_178_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.91"},{"key":"e_1_3_2_179_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1580"},{"key":"e_1_3_2_180_1","article-title":"Branch Specialization \u2014 distill.pub","author":"Voss Chelsea","year":"2021","unstructured":"Chelsea Voss, Gabriel Goh, Nick Cammarata, Michael Petrov, Ludwig Schubert, and Chris Olah. 2021. Branch Specialization \u2014 distill.pub. Retrieved from https:\/\/distill.pub\/2020\/circuits\/branch-specialization\/. [Accessed 26-11-2023].","journal-title":"R"},{"key":"e_1_3_2_181_1","doi-asserted-by":"crossref","unstructured":"Alex Wang Amanpreet Singh Julian Michael Felix Hill Omer Levy and Samuel R. Bowman. 2019. Glue: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP 353\u2013355.","DOI":"10.18653\/v1\/W18-5446"},{"key":"e_1_3_2_182_1","doi-asserted-by":"crossref","unstructured":"Boshi Wang Sewon Min Xiang Deng Jiaming Shen You Wu Luke Zettlemoyer and Huan Sun. 2022a. Towards understanding chain-of-thought prompting: An empirical study of what matters. arXiv:2212.10001.","DOI":"10.18653\/v1\/2023.acl-long.153"},{"key":"e_1_3_2_183_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-naacl.14"},{"key":"e_1_3_2_184_1","unstructured":"Jerry Wei Da Huang Yifeng Lu Denny Zhou and Quoc V. Le. 2023a. Simple Synthetic Data Reduces Sycophancy in Large Language Models. arXiv:2308.03958."},{"key":"e_1_3_2_185_1","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, Denny Zhou, et\u00a0al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (2022), 24824\u201324837.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_186_1","unstructured":"Jerry Wei Jason Wei Yi Tay Dustin Tran Albert Webson Yifeng Lu Xinyun Chen Hanxiao Liu Da Huang Denny Zhou and Tengyu Ma. 2023b. Larger language models do in-context learning differently. arXiv:2303.03846."},{"key":"e_1_3_2_187_1","unstructured":"Laura Weidinger John Mellor Maribeth Rauh Conor Griffin Jonathan Uesato Po-Sen Huang Myra Cheng Mia Glaese Borja Balle Atoosa Kasirzadeh Zac Kenton Sasha Brown Will Hawkins Tom Stepleton Courtney Biles Abeba Birhane Julia Haas Laura Rimell Lisa Hendricks Anne Isaac Legassick William Irving Sean Geoffrey and Iason Gabriel. 2021. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359."},{"key":"e_1_3_2_188_1","doi-asserted-by":"crossref","unstructured":"Sarah Wiegreffe and Yuval Pinter. 2019. Attention is not not Explanation. arXiv:1908.04626. Retrieved from https:\/\/arxiv.org\/abs\/1908.04626","DOI":"10.18653\/v1\/D19-1002"},{"key":"e_1_3_2_189_1","unstructured":"Skyler Wu Eric Meng Shen Charumathi Badrinath Jiaqi Ma and Himabindu Lakkaraju. 2023b. Analyzing chain-of-thought prompting in large language models via gradient-based feature attributions. arXiv preprint arXiv:2307.13339"},{"key":"e_1_3_2_190_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.523"},{"key":"e_1_3_2_191_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.173"},{"key":"e_1_3_2_192_1","unstructured":"Xuansheng Wu Wenlin Yao Jianshu Chen Xiaoman Pan Xiaoyang Wang Ninghao Liu and Dong Yu. 2023c. From language modeling to instruction following: Understanding the behavior shift in LLMs after instruction tuning. arXiv:2310.00492."},{"key":"e_1_3_2_193_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.383"},{"key":"e_1_3_2_194_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.blackboxnlp-1.24"},{"key":"e_1_3_2_195_1","unstructured":"Zhengxuan Wu and Desmond C. Ong. 2021. On Explaining Your Explanations of BERT: An Empirical Study with Sequence Classification. arXiv:2101.00196."},{"key":"e_1_3_2_196_1","unstructured":"Miao Xiong Zhiyuan Hu Xinyang Lu Yifei Li Jie Fu Junxian He and Bryan Hooi. 2023. Can LLMs express their uncertainty? An empirical evaluation of confidence elicitation in LLMs. arXiv:2306.13063. Retrieved from https:\/\/arxiv.org\/abs\/2306.13063"},{"key":"e_1_3_2_197_1","first-page":"30378","article-title":"The unreliability of explanations in few-shot prompting for textual reasoning","volume":"35","author":"Ye Xi","year":"2022","unstructured":"Xi Ye and Greg Durrett. 2022. The unreliability of explanations in few-shot prompting for textual reasoning. Advances in Neural Information Processing Systems 35 (2022), 30378\u201330392. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/hash\/c402501846f9fe03e2cac015b3f0e6b1-Abstract-Conference.html","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_198_1","doi-asserted-by":"crossref","unstructured":"Xi Ye and Greg Durrett. 2023. Explanation selection using unlabeled data for chain-of-thought prompting. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing 619\u2013637.","DOI":"10.18653\/v1\/2023.emnlp-main.41"},{"key":"e_1_3_2_199_1","unstructured":"Catherine Yeh Yida Chen Aoyu Wu Cynthia Chen Fernanda Vi\u00e9gas and Martin Wattenberg. 2023. AttentionViz: A Global View of Transformer Attention. arXiv:2305.03210."},{"key":"e_1_3_2_200_1","article-title":"Representer point selection for explaining deep neural networks","volume":"31","author":"Yeh Chih-Kuan","year":"2018","unstructured":"Chih-Kuan Yeh, Joon Kim, Ian En-Hsu Yen, and Pradeep K. Ravikumar. 2018. Representer point selection for explaining deep neural networks. Advances in Neural Information Processing Systems 31 (2018).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_201_1","doi-asserted-by":"crossref","unstructured":"Fan Yin Jesse Vig Philippe Laban Shafiq Joty Caiming Xiong and Chien-Sheng Jason Wu. 2023. Did you read the instructions? Rethinking the effectiveness of task definitions in instruction learning. arXiv:2306.01150.","DOI":"10.18653\/v1\/2023.acl-long.172"},{"key":"e_1_3_2_202_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-emnlp.255"},{"key":"e_1_3_2_203_1","unstructured":"Mert Yuksekgonul Maggie Wang and James Zou. 2023. Post-hoc concept bottleneck models. In ICLR 2022 Workshop on PAIR \\(^2\\) Struct: Privacy Accountability Interpretability Robustness Reasoning on Structured Data."},{"key":"e_1_3_2_204_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.blackboxnlp-1.24"},{"key":"e_1_3_2_205_1","unstructured":"Susan Zhang Stephen Roller Naman Goyal Mikel Artetxe Moya Chen Shuohui Chen Christopher Dewan Mona Diab Xian Li Xi Victoria Lin Todor Mihaylov Myle Ott and Sam Shleifer Kurt Shuster Daniel Simig Punit Singh Koura Anjali Sridhar Tianlu Wang and Luke Zettlemoyer. 2022a. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068."},{"key":"e_1_3_2_206_1","unstructured":"Yue Zhang Yafu Li Leyang Cui Deng Cai Lemao Liu Tingchen Fu Xinting Huang Enbo Zhao Yu Zhang Yulong Chen Longyue Wang Anh Tuan Luu Wei Bi Freda Shi and Shuming Shi. 2023. Siren\u2019s Song in the AI Ocean: A Survey on Hallucination in Large Language Models. arXiv:2309.01219."},{"key":"e_1_3_2_207_1","doi-asserted-by":"crossref","unstructured":"Zexuan Zhong Dan Friedman and Danqi Chen. 2021. Factual probing is [mask]: Learning vs. learning to recall. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 5017\u20135033.","DOI":"10.18653\/v1\/2021.naacl-main.398"},{"key":"e_1_3_2_208_1","unstructured":"Chunting Zhou Pengfei Liu Puxin Xu Srini Iyer Jiao Sun Yuning Mao Xuezhe Ma Avia Efrat Ping Yu Lili Yu Susan Zhang Gargi Ghosh Mike Lewis Luke Zettlemoyer and Omer Levy. 2023. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206."},{"key":"e_1_3_2_209_1","unstructured":"Andy Zou Long Phan Sarah Chen James Campbell Phillip Guo Richard Ren Alexander Pan Xuwang Yin Mantas Mazeika Ann-Kathrin Dombrowski Shashwat Goel Nathaniel Li Michael J. Byun and Zifan Wang Alex Mallen Steven Basart Sanmi Koyejo Dawn Song Matt Fredrikson J. Zico Kolter and Dan Hendrycks. 2023. Representation engineering: A top-down approach to AI transparency. arXiv:2310.01405. Retrieved from https:\/\/arxiv.org\/abs\/2310.01405"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639372","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3639372","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:54:11Z","timestamp":1750287251000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639372"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,22]]},"references-count":208,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,4,30]]}},"alternative-id":["10.1145\/3639372"],"URL":"https:\/\/doi.org\/10.1145\/3639372","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,22]]},"assertion":[{"value":"2023-09-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-30","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}