{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T16:10:41Z","timestamp":1784391041022,"version":"3.55.0"},"reference-count":59,"publisher":"MIT Press - Journals","issue":"1","license":[{"start":{"date-parts":[[2021,10,4]],"date-time":"2021-10-04T00:00:00Z","timestamp":1633305600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,4,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Probing classifiers have emerged as one of the prominent methodologies for interpreting and analyzing deep neural network models of natural language processing. The basic idea is simple\u2014a classifier is trained to predict some linguistic property from a model\u2019s representations\u2014and has been used to examine a wide variety of models and properties. However, recent studies have demonstrated various methodological limitations of this approach. This squib critically reviews the probing classifiers framework, highlighting their promises, shortcomings, and advances.<\/jats:p>","DOI":"10.1162\/coli_a_00422","type":"journal-article","created":{"date-parts":[[2021,10,5]],"date-time":"2021-10-05T00:13:56Z","timestamp":1633392836000},"page":"207-219","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":153,"title":["Probing Classifiers: Promises, Shortcomings, and Advances"],"prefix":"10.1162","volume":"48","author":[{"given":"Yonatan","family":"Belinkov","sequence":"first","affiliation":[{"name":"Technion \u2013 Israel Institute of Technology. belinkov@technion.ac.il"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2022,4,4]]},"reference":[{"key":"2022040614221404600_bib1","article-title":"Fine-grained analysis of sentence embeddings using auxiliary prediction tasks","volume":"abs\/1608.04207","author":"Adi","year":"2016","journal-title":"CoRR"},{"key":"2022040614221404600_bib2","article-title":"Fine-grained analysis of sentence embeddings using auxiliary prediction tasks","volume-title":"International Conference on Learning Representations (ICLR)","author":"Adi","year":"2017"},{"key":"2022040614221404600_bib3","article-title":"Understanding intermediate layers using linear classifier probes","author":"Alain","year":"2016","journal-title":"arXiv preprint arXiv:1610.01644v3"},{"key":"2022040614221404600_bib4","article-title":"Identifying and controlling important neurons in neural machine translation","volume-title":"International Conference on Learning Representations","author":"Bau","year":"2019"},{"key":"2022040614221404600_bib5","unstructured":"Belinkov, Yonatan\n          . 2018. On Internal Language Representations in Deep Learning: An Analysis of Machine Translation and Speech Recognition. Ph.D. thesis, Massachusetts Institute of Technology."},{"key":"2022040614221404600_bib6","doi-asserted-by":"publisher","first-page":"861","DOI":"10.18653\/v1\/P17-1080","article-title":"What do neural machine translation models learn about morphology?","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Belinkov","year":"2017"},{"issue":"1","key":"2022040614221404600_bib7","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1162\/coli_a_00367","article-title":"On the linguistic representational power of neural machine translation models","volume":"46","author":"Belinkov","year":"2020","journal-title":"Computational Linguistics"},{"key":"2022040614221404600_bib8","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18653\/v1\/2020.acl-tutorials.1","article-title":"Interpretability and analysis in neural NLP","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts","author":"Belinkov","year":"2020"},{"key":"2022040614221404600_bib9","first-page":"2441","article-title":"Analyzing hidden representations in end-to-end automatic speech recognition systems","volume-title":"Advances in Neural Information Processing Systems","author":"Belinkov","year":"2017"},{"key":"2022040614221404600_bib10","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1162\/tacl_a_00254","article-title":"Analysis methods in neural language processing: A survey","volume":"7","author":"Belinkov","year":"2019","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2022040614221404600_bib11","first-page":"1","article-title":"Evaluating layers of representation in neural machine translation on part-of-speech and semantic tagging tasks","volume-title":"Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Belinkov","year":"2017"},{"key":"2022040614221404600_bib12","doi-asserted-by":"publisher","first-page":"4487","DOI":"10.18653\/v1\/2020.acl-main.411","article-title":"DeFormer: Decomposing pre-trained transformers for faster question answering","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Cao","year":"2020"},{"key":"2022040614221404600_bib13","first-page":"960","article-title":"Low-complexity probing via finding subnetworks","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Cao","year":"2021"},{"key":"2022040614221404600_bib14","doi-asserted-by":"publisher","first-page":"2952","DOI":"10.18653\/v1\/P19-1283","article-title":"Correlating neural and symbolic representations of language","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Chrupa\u0142a","year":"2019"},{"key":"2022040614221404600_bib15","doi-asserted-by":"publisher","first-page":"4146","DOI":"10.18653\/v1\/2020.acl-main.381","article-title":"Analyzing analytical methods: The case of phonology in neural models of spoken language","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Chrupa\u0142a","year":"2020"},{"key":"2022040614221404600_bib16","first-page":"1396","article-title":"On the shortest arborescence of a directed graph","volume":"14","author":"Chu","year":"1965","journal-title":"Science Sinica"},{"key":"2022040614221404600_bib17","doi-asserted-by":"publisher","first-page":"276","DOI":"10.18653\/v1\/W19-4828","article-title":"What does BERT look at? An analysis of BERT\u2019s attention","volume-title":"Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Clark","year":"2019"},{"key":"2022040614221404600_bib18","doi-asserted-by":"publisher","first-page":"2126","DOI":"10.18653\/v1\/P18-1198","article-title":"What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Conneau","year":"2018"},{"key":"2022040614221404600_bib19","doi-asserted-by":"publisher","first-page":"4908","DOI":"10.18653\/v1\/2020.emnlp-main.398","article-title":"Analyzing redundancy in pretrained transformer models","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Dalvi","year":"2020"},{"key":"2022040614221404600_bib20","first-page":"447","article-title":"A survey of the state of explainable AI for natural language processing","volume-title":"Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing","author":"Danilevsky","year":"2020"},{"key":"2022040614221404600_bib21","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"issue":"4","key":"2022040614221404600_bib22","doi-asserted-by":"crossref","first-page":"233","DOI":"10.6028\/jres.071B.032","article-title":"Optimum branchings","volume":"71","author":"Edmonds","year":"1967","journal-title":"Journal of Research of the National Bureau of Standards B"},{"key":"2022040614221404600_bib23","doi-asserted-by":"publisher","first-page":"160","DOI":"10.1162\/tacl_a_00359","article-title":"Amnesic probing: Behavioral explanation with amnesic counterfactuals","volume":"9","author":"Elazar","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2022040614221404600_bib24","doi-asserted-by":"publisher","first-page":"134","DOI":"10.18653\/v1\/W16-2524","article-title":"Probing for semantic evidence of composition by means of simple classification tasks","volume-title":"Proceedings of the 1st Workshop on Evaluating Vector-Space Representations for NLP","author":"Ettinger","year":"2016"},{"issue":"2","key":"2022040614221404600_bib25","doi-asserted-by":"publisher","first-page":"333","DOI":"10.1162\/coli_a_00404","article-title":"CausaLM: Causal model explanation through counterfactual language models","volume":"47","author":"Feder","year":"2021","journal-title":"Computational Linguistics"},{"key":"2022040614221404600_bib26","doi-asserted-by":"publisher","first-page":"240","DOI":"10.18653\/v1\/W18-5426","article-title":"Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information","volume-title":"Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Giulianelli","year":"2018"},{"key":"2022040614221404600_bib27","doi-asserted-by":"publisher","first-page":"12","DOI":"10.18653\/v1\/D15-1002","article-title":"Distributional vectors encode referential attributes","volume-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing","author":"Gupta","year":"2015"},{"key":"2022040614221404600_bib28","doi-asserted-by":"publisher","first-page":"7389","DOI":"10.18653\/v1\/2020.acl-main.659","article-title":"A tale of a probe and a parser","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Hall Maudslay","year":"2020"},{"key":"2022040614221404600_bib29","doi-asserted-by":"publisher","first-page":"2733","DOI":"10.18653\/v1\/D19-1275","article-title":"Designing and interpreting probes with control tasks","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Hewitt","year":"2019"},{"key":"2022040614221404600_bib30","article-title":"Do attention heads in BERT track syntactic dependencies?","author":"Htut","year":"2019","journal-title":"arXiv preprint arXiv:1911.12246"},{"key":"2022040614221404600_bib31","doi-asserted-by":"publisher","first-page":"907","DOI":"10.1613\/jair.1.11196","article-title":"Visualisation and \u2019diagnostic classifiers\u2019 reveal how recurrent and recursive neural networks process hierarchical structure","volume":"61","author":"Hupkes","year":"2018","journal-title":"Journal of Artificial Intelligence Research"},{"key":"2022040614221404600_bib32","doi-asserted-by":"publisher","first-page":"2067","DOI":"10.18653\/v1\/D15-1246","article-title":"What\u2019s in an embedding? Analyzing word embeddings through multilingual evaluation","volume-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing","author":"K\u00f6hn","year":"2015"},{"key":"2022040614221404600_bib33","doi-asserted-by":"publisher","first-page":"4","DOI":"10.3389\/neuro.06.004.2008","article-title":"Representational similarity analysis\u2014connecting the branches of systems neuroscience","volume":"2","author":"Kriegeskorte","year":"2008","journal-title":"Frontiers in Systems Neuroscience"},{"key":"2022040614221404600_bib34","article-title":"Hierarchical multitask learning for CTC-based speech recognition","author":"Krishna","year":"2019","journal-title":"arXiv preprint arXiv:1807.06234"},{"key":"2022040614221404600_bib35","doi-asserted-by":"publisher","first-page":"11","DOI":"10.18653\/v1\/N19-1002","article-title":"The emergence of number and syntax units in LSTM language models","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Lakretz","year":"2019"},{"key":"2022040614221404600_bib36","doi-asserted-by":"publisher","first-page":"3637","DOI":"10.18653\/v1\/2020.coling-main.325","article-title":"Picking BERT\u2019s brain: Probing for linguistic dependencies in contextualized embeddings using representational similarity analysis","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Lepori","year":"2020"},{"key":"2022040614221404600_bib37","first-page":"1073","article-title":"Linguistic knowledge and transferability of contextual representations","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Liu","year":"2019"},{"key":"2022040614221404600_bib38","article-title":"Predicting inductive biases of pre-trained models","volume-title":"International Conference on Learning Representations","author":"Lovering","year":"2021"},{"key":"2022040614221404600_bib39","doi-asserted-by":"publisher","first-page":"263","DOI":"10.18653\/v1\/W19-4827","article-title":"From balustrades to Pierre Vinken: Looking for syntax in transformer self-attentions","volume-title":"Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Mare\u010dek","year":"2019"},{"key":"2022040614221404600_bib40","doi-asserted-by":"publisher","first-page":"6792","DOI":"10.18653\/v1\/2020.emnlp-main.552","article-title":"Asking without telling: Exploring latent ontologies in contextual representations","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Michael","year":"2020"},{"key":"2022040614221404600_bib41","doi-asserted-by":"publisher","first-page":"3138","DOI":"10.18653\/v1\/2020.emnlp-main.254","article-title":"Pareto probing: Trading off accuracy for complexity","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Pimentel","year":"2020"},{"key":"2022040614221404600_bib42","doi-asserted-by":"publisher","first-page":"4609","DOI":"10.18653\/v1\/2020.acl-main.420","article-title":"Information-theoretic probing for linguistic structure","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Pimentel","year":"2020"},{"key":"2022040614221404600_bib43","doi-asserted-by":"publisher","first-page":"287","DOI":"10.18653\/v1\/W18-5431","article-title":"An analysis of encoder representations in transformer-based machine translation","volume-title":"Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Raganato","year":"2018"},{"key":"2022040614221404600_bib44","first-page":"3363","article-title":"Probing the probing paradigm: Does probing accuracy entail task relevance?","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Ravichander","year":"2021"},{"key":"2022040614221404600_bib45","doi-asserted-by":"publisher","first-page":"842","DOI":"10.1162\/tacl_a_00349","article-title":"A primer in BERTology: What we know about how BERT works","volume":"8","author":"Rogers","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2022040614221404600_bib46","doi-asserted-by":"publisher","first-page":"1526","DOI":"10.18653\/v1\/D16-1159","article-title":"Does string-based neural MT learn source syntax?","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing","author":"Shi","year":"2016"},{"key":"2022040614221404600_bib47","doi-asserted-by":"publisher","first-page":"1393","DOI":"10.18653\/v1\/2020.findings-emnlp.125","article-title":"Investigating transferability in pretrained language models","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Tamkin","year":"2020"},{"key":"2022040614221404600_bib48","article-title":"What do you learn from context? Probing for sentence structure in contextualized word representations","volume-title":"International Conference on Learning Representations","author":"Tenney","year":"2019"},{"key":"2022040614221404600_bib49","doi-asserted-by":"publisher","first-page":"862","DOI":"10.18653\/v1\/2021.findings-acl.76","article-title":"What if this modified that? Syntactic interventions with counterfactual embeddings","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Tucker","year":"2021"},{"key":"2022040614221404600_bib50","first-page":"109","article-title":"Investigating \u2018aspect\u2019 in NMT and SMT: Translating the English simple past and present perfect","volume":"7","author":"Vanmassenhove","year":"2017","journal-title":"Computational Linguistics in the Netherlands Journal"},{"key":"2022040614221404600_bib51","article-title":"Diagnostic classifiers revealing how neural networks process hierarchical structure","volume-title":"CoCo@ NIPS","author":"Veldhoen","year":"2016"},{"key":"2022040614221404600_bib52","first-page":"12388","article-title":"Investigating gender bias in language models using causal mediation analysis","volume-title":"Advances in Neural Information Processing Systems","author":"Vig","year":"2020"},{"key":"2022040614221404600_bib53","doi-asserted-by":"publisher","first-page":"183","DOI":"10.18653\/v1\/2020.emnlp-main.14","article-title":"Information-theoretic probing with minimum description length","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Voita","year":"2020"},{"key":"2022040614221404600_bib54","doi-asserted-by":"publisher","first-page":"20","DOI":"10.18653\/v1\/2020.emnlp-tutorials.3","article-title":"Interpreting predictions of NLP models","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts","author":"Wallace","year":"2020"},{"key":"2022040614221404600_bib55","doi-asserted-by":"publisher","first-page":"4166","DOI":"10.18653\/v1\/2020.acl-main.383","article-title":"Perturbed masking: Parameter-free probing for analyzing and interpreting BERT","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Wu","year":"2020"},{"key":"2022040614221404600_bib56","doi-asserted-by":"publisher","first-page":"359","DOI":"10.18653\/v1\/W18-5448","article-title":"Language modeling teaches you more than translation does: Lessons learned through auxiliary syntactic task analysis","volume-title":"Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Zhang","year":"2018"},{"key":"2022040614221404600_bib57","doi-asserted-by":"publisher","first-page":"1112","DOI":"10.18653\/v1\/2021.acl-long.90","article-title":"When do you need billions of words of pretraining data?","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Zhang","year":"2021"},{"key":"2022040614221404600_bib58","doi-asserted-by":"publisher","first-page":"5070","DOI":"10.18653\/v1\/2021.naacl-main.401","article-title":"DirectProbe: Studying representations without classifiers","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Zhou","year":"2021"},{"key":"2022040614221404600_bib59","doi-asserted-by":"publisher","first-page":"9251","DOI":"10.18653\/v1\/2020.emnlp-main.744","article-title":"An information theoretic view on selecting linguistic probes","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Zhu","year":"2020"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/48\/1\/207\/2006605\/coli_a_00422.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/48\/1\/207\/2006605\/coli_a_00422.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,4,6]],"date-time":"2022-04-06T14:22:56Z","timestamp":1649254976000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/48\/1\/207\/107571\/Probing-Classifiers-Promises-Shortcomings-and"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022]]},"references-count":59,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,4,4]]},"published-print":{"date-parts":[[2022,4,4]]}},"URL":"https:\/\/doi.org\/10.1162\/coli_a_00422","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"value":"0891-2017","type":"print"},{"value":"1530-9312","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022]]},"published":{"date-parts":[[2022]]}}}