{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T18:08:23Z","timestamp":1775498903214,"version":"3.50.1"},"reference-count":57,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T00:00:00Z","timestamp":1775433600000},"content-version":"vor","delay-in-days":95,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Warning: This paper contains offensive text.<\/jats:p>\n                  <jats:p>We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (internals). Using this framework, we examine over 20 OLMo, Llama, and Mistral models, bridging behavioral and internal perspectives for toxicity for the first time. Our results show that LMs strongly encode information about the toxicity level of inputs and subsequent outputs, particularly in lower layers. Focusing on how unique LMs differ offers both correlative and causal evidence that they generate less toxic output when strongly encoding information about the input toxicity. We also highlight the heterogeneity of toxicity, as model behavior and internals vary across unique attributes such as Threat. Finally, four case studies analyzing detoxification, multi-prompt evaluations, model quantization, and pre-training dynamics underline the practical impact of aligned probing with further concrete insights. Our findings contribute to a more holistic understanding of LMs, both within and beyond the context of toxicity.<\/jats:p>\n                  <jats:p>alignedprobing.github.io<\/jats:p>","DOI":"10.1162\/tacl.a.613","type":"journal-article","created":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T16:49:57Z","timestamp":1775494197000},"page":"271-291","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":0,"title":["Aligned Probing: Relating Toxic Behavior and Model Internals"],"prefix":"10.1162","volume":"14","author":[{"given":"Andreas","family":"Waldis","sequence":"first","affiliation":[{"name":"Ubiquitous Knowledge Processing Lab (UKP Lab), Technical University of Darmstadt, Germany"},{"name":"Information Systems Research Lab, Lucerne University of Applied Sciences and Arts, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vagrant","family":"Gautam","sequence":"additional","affiliation":[{"name":"Spoken Language Systems, Saarland University, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anne","family":"Lauscher","sequence":"additional","affiliation":[{"name":"Data Science Group, University of Hamburg, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dietrich","family":"Klakow","sequence":"additional","affiliation":[{"name":"Spoken Language Systems, Saarland University, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Iryna","family":"Gurevych","sequence":"additional","affiliation":[{"name":"Ubiquitous Knowledge Processing Lab (UKP Lab), Technical University of Darmstadt, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2026,4,1]]},"reference":[{"key":"2026040612495084100_bib1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2405.15032","article-title":"Aya 23: Open weight releases to further multilingual progress","author":"Aryabumi","year":"2024","journal-title":"CoRR"},{"issue":"1","key":"2026040612495084100_bib2","doi-asserted-by":"publisher","first-page":"207","DOI":"10.1162\/coli_a_00422","article-title":"Probing classifiers: Promises, shortcomings, and advances","volume":"48","author":"Belinkov","year":"2022","journal-title":"Computational Linguistics"},{"key":"2026040612495084100_bib3","article-title":"Mechanistic interpretability for AI safety \u2013 A review","author":"Bereska","year":"2024","journal-title":"Transactions on Machine Learning Research"},{"key":"2026040612495084100_bib4","doi-asserted-by":"publisher","first-page":"275","DOI":"10.18653\/v1\/2024.woah-1.22","article-title":"Subjective isms? On the danger of conflating hate and offence in abusive language detection","volume-title":"Proceedings of the 8th Workshop on Online Abuse and Harms (WOAH 2024)","author":"Cercas Curry","year":"2024"},{"issue":"1","key":"2026040612495084100_bib5","doi-asserted-by":"publisher","first-page":"293","DOI":"10.1162\/coli_a_00492","article-title":"Language model behavior: A comprehensive survey","volume":"50","author":"Chang","year":"2024","journal-title":"Computational Linguistics"},{"key":"2026040612495084100_bib6","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i27.35011","article-title":"RTP-LX: can llms evaluate toxicity in multilingual scenarios?","author":"de Wynter","year":"2024","journal-title":"ArXiv preprint"},{"key":"2026040612495084100_bib7","doi-asserted-by":"publisher","first-page":"160","DOI":"10.1162\/tacl_a_00359","article-title":"Amnesic probing: Behavioral explanation with amnesic counterfactuals","volume":"9","author":"Elazar","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"issue":"3","key":"2026040612495084100_bib8","doi-asserted-by":"publisher","first-page":"1097","DOI":"10.1162\/coli_a_00524","article-title":"Bias and fairness in large language models: A survey","volume":"50","author":"Gallegos","year":"2024","journal-title":"Computational Linguistics"},{"key":"2026040612495084100_bib9","doi-asserted-by":"publisher","first-page":"3356","DOI":"10.18653\/v1\/2020.findings-emnlp.301","article-title":"RealToxicityPrompts: Evaluating neural toxic degeneration in language models","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Gehman","year":"2020"},{"key":"2026040612495084100_bib10","article-title":"The llama 3 herd of models","author":"Grattafiori","year":"2024"},{"key":"2026040612495084100_bib11","doi-asserted-by":"publisher","first-page":"15789","DOI":"10.18653\/v1\/2024.acl-long.841","article-title":"OLMo: Accelerating the science of language models","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Groeneveld","year":"2024"},{"key":"2026040612495084100_bib12","doi-asserted-by":"publisher","first-page":"3309","DOI":"10.18653\/v1\/2022.acl-long.234","article-title":"ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Hartvigsen","year":"2022"},{"key":"2026040612495084100_bib13","article-title":"We can\u2019t understand AI using our existing vocabulary","author":"Hewitt","year":"2025"},{"key":"2026040612495084100_bib14","doi-asserted-by":"publisher","first-page":"2733","DOI":"10.18653\/v1\/D19-1275","article-title":"Designing and interpreting probes with control tasks","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Hewitt","year":"2019"},{"key":"2026040612495084100_bib15","article-title":"The curious case of neural text degeneration","volume-title":"8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26\u201330, 2020","author":"Holtzman","year":"2020"},{"key":"2026040612495084100_bib16","doi-asserted-by":"publisher","first-page":"5040","DOI":"10.18653\/v1\/2023.emnlp-main.306","article-title":"Prompting is not a substitute for probability measurements in large language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Jennifer","year":"2023"},{"key":"2026040612495084100_bib17","article-title":"Polyglotoxicityprompts: Multilingual evaluation of neural toxic degeneration in large language models","author":"Jain","year":"2024","journal-title":"ArXiv preprint"},{"key":"2026040612495084100_bib18","article-title":"Mistral 7b","author":"Jiang","year":"2023","journal-title":"ArXiv preprint"},{"key":"2026040612495084100_bib19","article-title":"Stop anthropomorphizing intermediate tokens as reasoning\/thinking traces!","author":"Kambhampati","year":"2025","journal-title":"ArXiv preprint"},{"key":"2026040612495084100_bib20","doi-asserted-by":"publisher","first-page":"235","DOI":"10.18653\/v1\/S19-1026","article-title":"Probing what different NLP tasks teach machines about function word comprehension","volume-title":"Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM 2019)","author":"Kim","year":"2019"},{"key":"2026040612495084100_bib21","doi-asserted-by":"publisher","first-page":"3299","DOI":"10.18653\/v1\/2023.eacl-main.241","article-title":"Language generation models can cause harm: So what can we do about it? An actionable survey","volume-title":"Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics","author":"Kumar","year":"2023"},{"key":"2026040612495084100_bib22","doi-asserted-by":"publisher","first-page":"8818","DOI":"10.18653\/v1\/2022.acl-long.603","article-title":"Probing for the usage of grammatical number","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Lasri","year":"2022"},{"key":"2026040612495084100_bib23","doi-asserted-by":"publisher","first-page":"7901","DOI":"10.18653\/v1\/2022.emnlp-main.539","article-title":"SocioProbe: What, when, and where language models learn about sociodemographics","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Lauscher","year":"2022"},{"key":"2026040612495084100_bib24","article-title":"A mechanistic understanding of alignment algorithms: A case study on DPO and toxicity","volume-title":"Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21\u201327, 2024","author":"Lee","year":"2024"},{"key":"2026040612495084100_bib25","doi-asserted-by":"publisher","first-page":"13422","DOI":"10.18653\/v1\/2024.findings-emnlp.784","article-title":"Preference tuning for toxicity mitigation generalizes across languages","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2024","author":"Li","year":"2024"},{"key":"2026040612495084100_bib26","article-title":"Holistic evaluation of language models","volume":"2023","author":"Liang","year":"2023","journal-title":"Transactions on Machine Learning Research"},{"key":"2026040612495084100_bib27","doi-asserted-by":"publisher","first-page":"3245","DOI":"10.18653\/v1\/2024.naacl-long.179","article-title":"A pretrainer\u2019s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Longpre","year":"2024"},{"key":"2026040612495084100_bib28","article-title":"Decoupled weight decay regularization","volume-title":"7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6\u20139, 2019","author":"Loshchilov","year":"2019"},{"key":"2026040612495084100_bib29","doi-asserted-by":"publisher","first-page":"933","DOI":"10.1162\/tacl_a_00681","article-title":"State of what art? A call for multi-prompt LLM evaluation","volume":"12","author":"Mizrahi","year":"2024","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2026040612495084100_bib30","doi-asserted-by":"publisher","first-page":"3078","DOI":"10.18653\/v1\/2024.emnlp-main.181","article-title":"From insights to actions: The impact of interpretability and analysis research on NLP","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Mosbach","year":"2024"},{"key":"2026040612495084100_bib31","doi-asserted-by":"publisher","first-page":"68","DOI":"10.18653\/v1\/2020.blackboxnlp-1.7","article-title":"On the interplay between fine-tuning and sentence-level probing for linguistic knowledge in pre-trained transformers","volume-title":"Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"Mosbach","year":"2020"},{"key":"2026040612495084100_bib32","first-page":"3143","article-title":"Does BERT rediscover a classical NLP pipeline?","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Niu","year":"2022"},{"key":"2026040612495084100_bib33","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2312.12651","article-title":"Toxic bias: Perspective API misreads german as more toxic","author":"Nogara","year":"2023","journal-title":"CoRR"},{"key":"2026040612495084100_bib34","article-title":"2 olmo 2 furious","author":"OLMo","year":"2025"},{"key":"2026040612495084100_bib35","doi-asserted-by":"publisher","first-page":"4262","DOI":"10.18653\/v1\/2021.acl-long.329","article-title":"Probing toxic content in large pre-trained language models","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Ousidhoum","year":"2021"},{"key":"2026040612495084100_bib36","doi-asserted-by":"publisher","first-page":"107","DOI":"10.18653\/v1\/2023.c3nlp-1.11","article-title":"Toward disambiguating the definitions of abusive, offensive, toxic, and uncivil comments","volume-title":"Proceedings of the First Workshop on Cross-Cultural Considerations in NLP (C3NLP)","author":"Pachinger","year":"2023"},{"key":"2026040612495084100_bib37","doi-asserted-by":"publisher","first-page":"7595","DOI":"10.18653\/v1\/2023.emnlp-main.472","article-title":"On the challenges of using black-box apis for toxicity evaluation in research","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6\u201310, 2023","author":"Pozzobon","year":"2023"},{"key":"2026040612495084100_bib38","article-title":"Direct preference optimization: Your language model is secretly a reward model","volume-title":"Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10\u201316, 2023","author":"Rafailov","year":"2023"},{"key":"2026040612495084100_bib39","doi-asserted-by":"publisher","first-page":"3363","DOI":"10.18653\/v1\/2021.eacl-main.295","article-title":"Probing the probing paradigm: Does probing accuracy entail task relevance?","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Ravichander","year":"2021"},{"key":"2026040612495084100_bib40","doi-asserted-by":"publisher","first-page":"5884","DOI":"10.18653\/v1\/2022.naacl-main.431","article-title":"Annotators with attitudes: How annotator beliefs and identities bias toxic language detection","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Sap","year":"2022"},{"key":"2026040612495084100_bib41","doi-asserted-by":"publisher","first-page":"480","DOI":"10.18653\/v1\/2024.blackboxnlp-1.30","article-title":"Mechanistic?","volume-title":"Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP","author":"Saphra","year":"2024"},{"key":"2026040612495084100_bib42","article-title":"Quantifying language models\u2019 sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting","volume-title":"ICLR","author":"Sclar","year":"2024"},{"key":"2026040612495084100_bib43","doi-asserted-by":"publisher","DOI":"10.70777\/si.v2i6.15919","article-title":"The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity","author":"Shojaee","year":"2025"},{"key":"2026040612495084100_bib44","doi-asserted-by":"publisher","first-page":"24","DOI":"10.18653\/v1\/2023.mrl-1.3","article-title":"Counterfactually probing language identity in multilingual models","volume-title":"Proceedings of the 3rd Workshop on Multi-lingual Representation Learning (MRL)","author":"Srinivasan","year":"2023"},{"key":"2026040612495084100_bib45","doi-asserted-by":"publisher","first-page":"4593","DOI":"10.18653\/v1\/P19-1452","article-title":"BERT rediscovers the classical NLP pipeline","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Tenney","year":"2019"},{"key":"2026040612495084100_bib46","article-title":"What do you learn from context? Probing for sentence structure in contextualized word representations","volume-title":"7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6\u20139, 2019","author":"Tenney","year":"2019"},{"key":"2026040612495084100_bib47","article-title":"Llama 2: Open foundation and fine-tuned chat models","author":"Touvron","year":"2023","journal-title":"ArXiv preprint"},{"key":"2026040612495084100_bib48","doi-asserted-by":"publisher","first-page":"183","DOI":"10.18653\/v1\/2020.emnlp-main.14","article-title":"Information-theoretic probing with minimum description length","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Voita","year":"2020"},{"key":"2026040612495084100_bib49","doi-asserted-by":"publisher","first-page":"2197","DOI":"10.18653\/v1\/2024.findings-eacl.146","article-title":"Dive into the chasm: Probing the gap between in- and cross-topic generalization","volume-title":"Findings of the Association for Computational Linguistics: EACL 2024","author":"Waldis","year":"2024"},{"key":"2026040612495084100_bib50","doi-asserted-by":"publisher","first-page":"1616","DOI":"10.1162\/tacl_a_00718","article-title":"Holmes: A benchmark to assess the linguistic competence of language models","volume":"12","author":"Waldis","year":"2024","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2026040612495084100_bib51","doi-asserted-by":"publisher","first-page":"3093","DOI":"10.18653\/v1\/2024.acl-long.171","article-title":"Detoxifying large language models via knowledge editing","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wang","year":"2024"},{"key":"2026040612495084100_bib52","doi-asserted-by":"publisher","first-page":"377","DOI":"10.1162\/tacl_a_00321","article-title":"BLiMP: The benchmark of linguistic minimal pairs for English","volume":"8","author":"Warstadt","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2026040612495084100_bib53","doi-asserted-by":"publisher","first-page":"138","DOI":"10.18653\/v1\/W16-5618","article-title":"Are you a racist or am I seeing things? Annotator influence on hate speech detection on Twitter","volume-title":"Proceedings of the First Workshop on NLP and Computational Social Science","author":"Waseem","year":"2016"},{"key":"2026040612495084100_bib54","doi-asserted-by":"publisher","first-page":"78","DOI":"10.18653\/v1\/W17-3012","article-title":"Understanding abuse: A typology of abusive language detection subtasks","volume-title":"Proceedings of the First Workshop on Abusive Language Online","author":"Waseem","year":"2017"},{"key":"2026040612495084100_bib55","doi-asserted-by":"publisher","first-page":"1322","DOI":"10.18653\/v1\/2023.emnlp-main.84","article-title":"Unveiling the implicit toxicity in large language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Wen","year":"2023"},{"key":"2026040612495084100_bib56","unstructured":"Peter\n              West\n            , XimingLu, NouhaDziri, FaezeBrahman, LinjieLi, Jena D.Hwang, LiweiJiang, JillianFisher, AbhilashaRavichander, Khyathi RaghaviChandu, BenjaminNewman, Pang WeiKoh, AllysonEttinger, and YejinChoi. 2024. The generative AI paradox: \u201cwhat it can create, it may not understand\u201d. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\u201311, 2024. OpenReview.net."},{"key":"2026040612495084100_bib57","article-title":"Model merging in llms, mllms, and beyond: Methods, theories, applications and opportunities","author":"Yang","year":"2024","journal-title":"ArXiv preprint"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/TACL.a.613\/2590825\/tacl.a.613.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/TACL.a.613\/2590825\/tacl.a.613.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T16:50:01Z","timestamp":1775494201000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/TACL.a.613\/136155\/Aligned-Probing-Relating-Toxic-Behavior-and-Model"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026]]},"references-count":57,"URL":"https:\/\/doi.org\/10.1162\/tacl.a.613","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026]]},"published":{"date-parts":[[2026]]}}}