{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T02:12:02Z","timestamp":1784167922237,"version":"3.55.0"},"reference-count":193,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2020,10,6]],"date-time":"2020-10-06T00:00:00Z","timestamp":1601942400000},"content-version":"vor","delay-in-days":554,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The field of natural language processing has seen impressive progress in recent years, with neural network models replacing many of the traditional systems. A plethora of new models have been proposed, many of which are thought to be opaque compared to their feature-rich counterparts. This has led researchers to analyze, interpret, and evaluate neural networks in novel and more fine-grained ways. In this survey paper, we review analysis methods in neural language processing, categorize them according to prominent research trends, highlight existing limitations, and point to potential directions for future work.<\/jats:p>","DOI":"10.1162\/tacl_a_00254","type":"journal-article","created":{"date-parts":[[2019,4,2]],"date-time":"2019-04-02T13:04:45Z","timestamp":1554210285000},"page":"49-72","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":198,"title":["Analysis Methods in Neural Language Processing: A Survey"],"prefix":"10.1162","volume":"7","author":[{"given":"Yonatan","family":"Belinkov","sequence":"first","affiliation":[{"name":"MIT Computer Science and Artificial Intelligence Laboratory, United Stated."},{"name":"Harvard School of Engineering and Applied Sciences Cambridge, MA, USA. belinkov@mit.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"James","family":"Glass","sequence":"additional","affiliation":[{"name":"MIT Computer Science and Artificial Intelligence Laboratory, United Stated. glass@mit.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2019,4,1]]},"reference":[{"key":"2021060823284521400_bib1","doi-asserted-by":"crossref","unstructured":"Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017a. Analysis of sentence embedding models using prediction tasks in natural language processing. IBM Journal of Research and Development, 61(4):3\u20139.","DOI":"10.1147\/JRD.2017.2702858"},{"key":"2021060823284521400_bib2","unstructured":"Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017. Fine-Grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks. In International Conference on Learning Representations (ICLR)."},{"key":"2021060823284521400_bib3","doi-asserted-by":"crossref","unstructured":"Roee Aharoni and Yoav Goldberg. 2017. Morphological Inflection Generation with Hard Monotonic Attention. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2004\u20132015. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P17-1183"},{"key":"2021060823284521400_bib4","unstructured":"Wasi Uddin Ahmad, Xueying Bai, Zhechao Huang, Chao Jiang, Nanyun Peng, and Kai-Wei Chang. 2018. Multi-task Learning for Universal Sentence Embeddings: A Thorough Evaluation using Transfer and Auxiliary Tasks. arXiv preprint arXiv:1804.07911v2."},{"key":"2021060823284521400_bib5","doi-asserted-by":"crossref","unstructured":"Afra Alishahi, Marie Barking, and Grzegorz Chrupa\u0142a. 2017. Encoding of phonology in a recurrent neural model of grounded speech. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), pages 368\u2013378. Association for Computational Linguistics.","DOI":"10.18653\/v1\/K17-1037"},{"key":"2021060823284521400_bib6","doi-asserted-by":"crossref","unstructured":"David Alvarez-Melis and Tommi Jaakkola. 2017. A causal framework for explaining the predictions of black-box sequence-to-sequence models. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 412\u2013421. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D17-1042"},{"key":"2021060823284521400_bib7","doi-asserted-by":"crossref","unstructured":"Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018. Generating Natural Language Adversarial Examples. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2890\u20132896. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D18-1316"},{"key":"2021060823284521400_bib8","doi-asserted-by":"crossref","unstructured":"Leila Arras, Franziska Horn, Gr\u00e9goire Montavon, Klaus-Robert M\u00fcller, and Wojciech Samek. 2017a. \u201cWhat is relevant in a text document?\u201d: An interpretable machine learning approach. PLOS ONE, 12(8):1\u201323.","DOI":"10.1371\/journal.pone.0181142"},{"key":"2021060823284521400_bib9","doi-asserted-by":"crossref","unstructured":"Leila Arras, Gr\u00e9goire Montavon, Klaus-Robert M\u00fcller, and Wojciech Samek. 2017b. Explaining Recurrent Neural Network Predictions in Sentiment Analysis. In Proceedings of the 8th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, pages 159\u2013168. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W17-5221"},{"key":"2021060823284521400_bib10","doi-asserted-by":"crossref","unstructured":"Mikel Artetxe, Gorka Labaka, Inigo Lopez-Gazpio, and Eneko Agirre. 2018. Uncovering Divergent Linguistic Information in Word Embeddings with Lessons for Intrinsic and Extrinsic Evaluation. In Proceedings of the 22nd Conference on Computational Natural Language Learning, pages 282\u2013291. Association for Computational Linguistics.","DOI":"10.18653\/v1\/K18-1028"},{"key":"2021060823284521400_bib11","doi-asserted-by":"crossref","unstructured":"Malika Aubakirova and Mohit Bansal. 2016. Interpreting Neural Networks to Improve Politeness Comprehension. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2035\u20132041. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D16-1216"},{"key":"2021060823284521400_bib12","unstructured":"Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural Machine Translation by Jointly Learning to Align and Translate. arXiv preprint arXiv:1409.0473v7."},{"key":"2021060823284521400_bib13","unstructured":"Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2018. Identifying and Controlling Important Neurons in Neural Machine Translation. arXiv preprint arXiv:1811.01157v1."},{"key":"2021060823284521400_bib14","doi-asserted-by":"crossref","unstructured":"Rachel Bawden, Rico Sennrich, Alexandra Birch, and Barry Haddow. 2018. Evaluating Discourse Phenomena in Neural Machine Translation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1304\u20131313. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N18-1118"},{"key":"2021060823284521400_bib15","unstructured":"Yonatan Belinkov . 2018. On Internal Language Representations in Deep Learning: An Analysis of Machine Translation and Speech Recognition. Ph.D. thesis, Massachusetts Institute of Technology."},{"key":"2021060823284521400_bib16","unstructured":"Yonatan Belinkov and Yonatan Bisk. 2018. Synthetic and Natural Noise Both Break Neural Machine Translation. In International Conference on Learning Representations (ICLR)."},{"key":"2021060823284521400_bib17","doi-asserted-by":"crossref","unstructured":"Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2017a. What do Neural Machine Translation Models Learn about Morphology? In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 861\u2013872. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P17-1080"},{"key":"2021060823284521400_bib18","unstructured":"Yonatan Belinkov and James Glass. 2017, Analyzing Hidden Representations in End-to-End Automatic Speech Recognition Systems, I.Guyon, U. V.Luxburg, S.Bengio, H.Wallach, R.Fergus, S.Vishwanathan, and R.Garnett, editors, Advances in Neural Information Processing Systems 30, pages 2441\u20132451. Curran Associates, Inc."},{"key":"2021060823284521400_bib19","unstructured":"Yonatan Belinkov, Llu\u00eds M\u00e0rquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2017b. Evaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging Tasks. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1\u2013pages10. Asian Federation of Natural Language Processing."},{"key":"2021060823284521400_bib20","doi-asserted-by":"crossref","unstructured":"Jean-Philippe Bernardy . 2018. Can Recurrent Neural Networks Learn Nested Recursion?LiLT (Linguistic Issues in Language Technology), 16(1).","DOI":"10.33011\/lilt.v16i.1417"},{"key":"2021060823284521400_bib21","doi-asserted-by":"crossref","unstructured":"Arianna Bisazza and Clara Tump. 2018. The Lazy Encoder: A Fine-Grained Analysis of the Role of Morphology in Neural Machine Translation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2871\u20132876. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D18-1313"},{"key":"2021060823284521400_bib22","doi-asserted-by":"crossref","unstructured":"Terra Blevins, Omer Levy, and Luke Zettlemoyer. 2018. Deep RNNs Encode Soft Hierarchical Syntax. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 14\u201319. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-2003"},{"key":"2021060823284521400_bib23","doi-asserted-by":"crossref","unstructured":"Mikael Bod\u00e9n and Janet Wiles. 2002. On learning context-free and context-sensitive languages. IEEE Transactions on Neural Networks, 13(2): 491\u2013493.","DOI":"10.1109\/72.991436"},{"key":"2021060823284521400_bib24","doi-asserted-by":"crossref","unstructured":"Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632\u2013642. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D15-1075"},{"key":"2021060823284521400_bib25","unstructured":"Elia Bruni, Gemma Boleda, Marco Baroni, and Nam Khanh Tran. 2012. Distributional Semantics in Technicolor. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 136\u2013145. Association for Computational Linguistics."},{"key":"2021060823284521400_bib26","unstructured":"Gino Brunner, Yuyi Wang, Roger Wattenhofer, and Michael Weigelt. 2017. Natural Language Multitasking: Analyzing and Improving Syntactic Saliency of Hidden Representations. The 31st Annual Conference on Neural Information Processing (NIPS)\u2014Workshop on Learning Disentangled Features: From Perception to Control."},{"key":"2021060823284521400_bib27","doi-asserted-by":"crossref","unstructured":"Aljoscha Burchardt, Vivien Macketanz, Jon Dehdari, Georg Heigold, Jan-Thorsten Peter, and Philip Williams. 2017. A Linguistic Evaluation of Rule-Based, Phrase-Based, and Neural MT Engines. The Prague Bulletin of Mathematical Linguistics, 108(1):159\u2013170.","DOI":"10.1515\/pralin-2017-0017"},{"key":"2021060823284521400_bib28","doi-asserted-by":"crossref","unstructured":"Franck Burlot and Fran\u00e7ois Yvon. 2017. Evaluating the morphological competence of Machine Translation Systems. In Proceedings of the Second Conference on Machine Translation, pages 43\u201355. Association for Compu tational Linguistics.","DOI":"10.18653\/v1\/W17-4705"},{"key":"2021060823284521400_bib29","doi-asserted-by":"crossref","unstructured":"Mike Casey . 1996. The Dynamics of Discrete-Time Computation, with Application to Re current Neural Networks and Finite State Machine Extraction. Neural Computation, 8(6):1135\u20131178.","DOI":"10.1162\/neco.1996.8.6.1135"},{"key":"2021060823284521400_bib30","unstructured":"Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017. SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 1\u201314. Association for Computational Linguistics."},{"key":"2021060823284521400_bib31","doi-asserted-by":"crossref","unstructured":"Rahma Chaabouni, Ewan Dunbar, Neil Zeghidour, and Emmanuel Dupoux. 2017. Learning weakly supervised multimodal phoneme embeddings. In Interspeech 2017.","DOI":"10.21437\/Interspeech.2017-1689"},{"key":"2021060823284521400_bib32","doi-asserted-by":"crossref","unstructured":"Stephan K. Chalup and Alan D. Blair. 2003. Incremental Training of First Order Recurrent Neural Networks to Predict a Context-Sensitive Language. Neural Networks, 16(7):955\u2013972.","DOI":"10.1016\/S0893-6080(03)00054-6"},{"key":"2021060823284521400_bib33","unstructured":"Jonathan Chang, Sean Gerrish, Chong Wang, Jordan L. Boyd-graber, and David M. Blei. 2009, Reading Tea Leaves: How Humans Interpret Topic Models, Y.Bengio, D.Schuurmans, J. D.Lafferty, C. K. I.Williams, and A.Culotta, editors, Advances in Neural Information Processing Systems 22, pages 288\u2013296, Curran Associates, Inc.."},{"key":"2021060823284521400_bib34","doi-asserted-by":"crossref","unstructured":"Hongge Chen, Huan Zhang, Pin-Yu Chen, Jinfeng Yi, and Cho-Jui Hsieh. 2018a. Attacking visual language grounding with adversarial examples: A case study on neural image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2587\u20132597. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-1241"},{"key":"2021060823284521400_bib35","doi-asserted-by":"crossref","unstructured":"Xinchi Chen, Xipeng Qiu, Chenxi Zhu, Shiyu Wu , and XuanjingHuang. 2015. Sentence Modeling with Gated Recursive Neural Network. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 793\u2013798. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D15-1092"},{"key":"2021060823284521400_bib36","doi-asserted-by":"crossref","unstructured":"Yining Chen, Sorcha Gilroy, Andreas Maletti, Jonathan May, and Kevin Knight. 2018b. Recurrent Neural Networks as Weighted Language Recognizers. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2261\u20132271. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N18-1205"},{"key":"2021060823284521400_bib37","unstructured":"Minhao Cheng, Jinfeng Yi, Huan Zhang, Pin-Yu Chen, and Cho-Jui Hsieh. 2018. Seq2Sick: Evaluating the Robustness of Sequence-to-Sequence Models with Adversarial Examples. arXiv preprint arXiv:1803.01128v1."},{"key":"2021060823284521400_bib38","doi-asserted-by":"crossref","unstructured":"Grzegorz Chrupa\u0142a, Lieke Gelderloos, and Afra Alishahi. 2017. Representations of language in a model of visually grounded speech signal. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 613\u2013622. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P17-1057"},{"key":"2021060823284521400_bib39","doi-asserted-by":"crossref","unstructured":"Ond\u0159ej C\u00edfka and Ond\u0159ej Bojar. 2018. Are BLEU and Meaning Representation in Opposition? In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1362\u20131371. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-1126"},{"key":"2021060823284521400_bib40","doi-asserted-by":"crossref","unstructured":"Alexis Conneau, Germ\u00e1n Kruszewski, Guillaume Lample, Lo\u00efc Barrault, and Marco Baroni. 2018. What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2126\u20132136. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-1198"},{"key":"2021060823284521400_bib41","unstructured":"Robin Cooper, Dick Crouch, Jan van Eijck, Chris Fox, Josef van Genabith, Jan Jaspars, Hans Kamp, David Milward, Manfred Pinkal, Massimo Poesio, Steve Pulman, Ted Briscoe, Holger Maier, and Karsten Konrad. 1996, Using the framework. Technical report, The FraCaS Consortium."},{"key":"2021060823284521400_bib42","doi-asserted-by":"crossref","unstructured":"Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, D. Anthony Bau, and James Glass. 2019a, January. What Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP Models. In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence (AAAI).","DOI":"10.1609\/aaai.v33i01.33016309"},{"key":"2021060823284521400_bib43","unstructured":"Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, and Stephan Vogel. 2017. Understanding and Improving Morphological Learning in the Neural Machine Translation Decoder. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 142\u2013151. Asian Federation of Natural Language Processing."},{"key":"2021060823284521400_bib44","doi-asserted-by":"crossref","unstructured":"Fahim Dalvi, Avery Nortonsmith, D. Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, and James Glass. 2019b, January. NeuroX: A Toolkit for Analyzing Individual Neurons in Neural Networks. In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence (AAAI): Demonstrations Track.","DOI":"10.1609\/aaai.v33i01.33019851"},{"key":"2021060823284521400_bib45","unstructured":"Sreerupa Das, C. Lee Giles, and Guo-Zheng Sun. 1992. Learning Context-Free Grammars: Capabilities and Limitations of a Recurrent Neural Network with an External Stack Memory. In Proceedings of The Fourteenth Annual Conference of Cognitive Science Society. Indiana University, page 14."},{"key":"2021060823284521400_bib46","unstructured":"Ishita Dasgupta, Demi Guo, Andreas Stuhlm\u00fcller, Samuel J. Gershman, and Noah D. Goodman. 2018. Evaluating Compositionality in Sentence Embeddings. arXiv preprint arXiv:1802. 04302v2."},{"key":"2021060823284521400_bib47","doi-asserted-by":"crossref","unstructured":"Dhanush Dharmaretnam and Alona Fyshe. 2018. The Emergence of Semantics in Neural Network Representations of Visual Information. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 776\u2013780. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N18-2122"},{"key":"2021060823284521400_bib48","doi-asserted-by":"crossref","unstructured":"Yanzhuo Ding, Yang Liu, Huanbo Luan, and Maosong Sun. 2017. Visualizing and Understanding Neural Machine Translation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1150\u20131159. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P17-1106"},{"key":"2021060823284521400_bib49","unstructured":"Finale Doshi-Velez and Been Kim. 2017. Towards a Rigorous Science of Interpretable Machine Learning. In arXiv preprint arXiv: 1702.08608v2."},{"key":"2021060823284521400_bib50","doi-asserted-by":"crossref","unstructured":"Finale Doshi-Velez, Mason Kortz, Ryan Budish, Chris Bavitz, Sam Gershman, David O\u2019Brien, Stuart Shieber, James Waldo, David Weinberger, and Alexandra Wood. 2017. Accountability of AI Under the Law: The Role of Explanation. Privacy Law Scholars Conference.","DOI":"10.2139\/ssrn.3064761"},{"key":"2021060823284521400_bib51","doi-asserted-by":"crossref","unstructured":"Jennifer Drexler and James Glass. 2017. Analysis of Audio-Visual Features for Unsupervised Speech Recognition. In International Workshop on Grounding Language Understanding.","DOI":"10.21437\/GLU.2017-12"},{"key":"2021060823284521400_bib52","unstructured":"Javid Ebrahimi, Daniel Lowd, and Dejing Dou. 2018a. On Adversarial Examples for Character-Level Neural Machine Translation. In Proceedings of the 27th International Conference on Computational Linguistics, pages 653\u2013663. Association for Computational Linguistics."},{"key":"2021060823284521400_bib53","doi-asserted-by":"crossref","unstructured":"Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018b. HotFlip: White-Box Adversarial Examples for Text Classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 31\u201336. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-2006"},{"key":"2021060823284521400_bib54","doi-asserted-by":"crossref","unstructured":"Ali Elkahky, Kellie Webster, Daniel Andor, and Emily Pitler. 2018. A Challenge Set and Methods for Noun-Verb Ambiguity. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2562\u20132572. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D18-1277"},{"key":"2021060823284521400_bib55","doi-asserted-by":"crossref","unstructured":"Zied Elloumi, Laurent Besacier, Olivier Galibert, and Benjamin Lecouteux. 2018. Analyzing Learned Representations of a Deep ASR Performance Prediction Model. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 9\u201315. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W18-5402"},{"key":"2021060823284521400_bib56","doi-asserted-by":"crossref","unstructured":"Jeffrey L. Elman . 1989. Representation and Structure in Connectionist Models, University of California, San Diego, Center for Research in Language.","DOI":"10.21236\/ADA259504"},{"key":"2021060823284521400_bib57","doi-asserted-by":"crossref","unstructured":"Jeffrey L. Elman . 1990. Finding Structure in Time. Cognitive Science, 14(2):179\u2013211.","DOI":"10.1207\/s15516709cog1402_1"},{"key":"2021060823284521400_bib58","doi-asserted-by":"crossref","unstructured":"Jeffrey L. Elman . 1991. Distributed representations, simple recurrent networks, and grammatical structure. Machine Learning, 7(2\u20133): 195\u2013225.","DOI":"10.1007\/BF00114844"},{"key":"2021060823284521400_bib59","doi-asserted-by":"crossref","unstructured":"Allyson Ettinger, Ahmed Elgohary, and Philip Resnik. 2016. Probing for semantic evidence of composition by means of simple classification tasks. In Proceedings of the 1st Workshop on Evaluating Vector-Space Representations for NLP, pages 134\u2013139. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W16-2524"},{"key":"2021060823284521400_bib60","doi-asserted-by":"crossref","unstructured":"Manaal Faruqui, Yulia Tsvetkov, Pushpendre Rastogi, and Chris Dyer. 2016. Problems With Evaluation of Word Embeddings Using Word Similarity Tasks. In Proceedings of the 1st Workshop on Evaluating Vector Space Representations for NLP.","DOI":"10.18653\/v1\/W16-2506"},{"key":"2021060823284521400_bib61","doi-asserted-by":"crossref","unstructured":"Manaal Faruqui, Yulia Tsvetkov, Dani Yogatama, Chris Dyer, and Noah A. Smith. 2015. Sparse Overcomplete Word Vector Representations. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1491\u20131500. Association for Computational Linguistics.","DOI":"10.3115\/v1\/P15-1144"},{"key":"2021060823284521400_bib62","doi-asserted-by":"crossref","unstructured":"Shi Feng, Eric Wallace, Alvin Grissom II , MohitIyyer, PedroRodriguez, and JordanBoyd-Graber. 2018. Pathologies of Neural Models Make Interpretations Difficult. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3719\u20133728. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D18-1407"},{"key":"2021060823284521400_bib63","doi-asserted-by":"crossref","unstructured":"Lev Finkelstein, Evgeniy Gabrilovich, Yossi Matias, Ehud Rivlin, Zach Solan, Gadi Wolfman, and Eytan Ruppin. 2002. Placing Search in Context: The Concept Revisited. ACM Transactions on Information Systems, 20(1):116\u2013131.","DOI":"10.1145\/503104.503110"},{"key":"2021060823284521400_bib64","doi-asserted-by":"crossref","unstructured":"Robert Frank, Donald Mathis, and William Badecker. 2013. The Acquisition of Anaphora by Simple Recurrent Networks. Language Acquisition, 20(3):181\u2013227.","DOI":"10.1080\/10489223.2013.796950"},{"key":"2021060823284521400_bib65","unstructured":"Cynthia Freeman, Jonathan Merriman, Abhinav Aggarwal, Ian Beaver, and Abdullah Mueen. 2018. Paying Attention to Attention: Highlighting Influential Samples in Sequential Analysis. arXiv preprint arXiv:1808.02113v1."},{"key":"2021060823284521400_bib66","doi-asserted-by":"crossref","unstructured":"Alona Fyshe, Leila Wehbe, Partha P. Talukdar, Brian Murphy, and Tom M. Mitchell. 2015. A Compositional and Interpretable Semantic Space. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 32\u201341. Association for Computational Linguistics.","DOI":"10.3115\/v1\/N15-1004"},{"key":"2021060823284521400_bib67","doi-asserted-by":"crossref","unstructured":"David Gaddy, Mitchell Stern, and Dan Klein. 2018. What\u2019s Going On in Neural Constituency Parsers? An Analysis. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 999\u20131010. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N18-1091"},{"key":"2021060823284521400_bib68","doi-asserted-by":"crossref","unstructured":"J. Ganesh, Manish Gupta, and Vasudeva Varma. 2017. Interpretation of Semantic Tweet Representations. In Proceedings of the 2017 IEEE\/ACM International Conference on Advances in Social Networks Analysis and Mining 2017, ASONAM \u201917, pages 95\u2013102, New York, NY, USA. ACM.","DOI":"10.1145\/3110025.3110083"},{"key":"2021060823284521400_bib69","doi-asserted-by":"crossref","unstructured":"Ji Gao , JackLanchantin, Mary LouSoffa, and YanjunQi. 2018. Black-box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers. arXiv preprint arXiv: 1801.04354v5.","DOI":"10.1109\/SPW.2018.00016"},{"key":"2021060823284521400_bib70","unstructured":"Lieke Gelderloos and Grzegorz Chrupa\u0142a. 2016. From phonemes to images: Levels of representation in a recurrent neural model of visually-grounded language learning. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 1309\u20131319, Osaka, Japan, The COLING 2016 Organizing Committee."},{"key":"2021060823284521400_bib71","doi-asserted-by":"crossref","unstructured":"Felix A. Gers and J\u00fcrgen Schmidhuber. 2001. LSTM Recurrent Networks Learn Simple Context-Free and Context-Sensitive Languages. IEEE Transactions on Neural Networks, 12(6): 1333\u20131340.","DOI":"10.1109\/72.963769"},{"key":"2021060823284521400_bib72","doi-asserted-by":"crossref","unstructured":"Daniela Gerz, Ivan Vuli\u0107, Felix Hill, Roi Reichart, and Anna Korhonen. 2016. SimVerb-3500: A Large-Scale Evaluation Set of Verb Similarity. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2173\u20132182. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D16-1235"},{"key":"2021060823284521400_bib73","unstructured":"Hamidreza Ghader and Christof Monz. 2017. What does Attention in Neural Machine Translation Pay Attention to? In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 30\u201339. Asian Federation of Natural Language Processing."},{"key":"2021060823284521400_bib74","doi-asserted-by":"crossref","unstructured":"Reza Ghaeini, Xiaoli Fern, and Prasad Tadepalli. 2018. Interpreting Recurrent and Attention-Based Neural Models: A Case Study on Natural Language Inference. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4952\u20134957. Association for Computational Linguistics&gt;.","DOI":"10.18653\/v1\/D18-1537"},{"key":"2021060823284521400_bib75","doi-asserted-by":"crossref","unstructured":"Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema. 2018. Under the Hood: Using Diagnostic Classifiers to Investigate and Improve How Language Models Track Agreement Information. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 240\u2013248. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W18-5426"},{"key":"2021060823284521400_bib76","doi-asserted-by":"crossref","unstructured":"Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018. Breaking NLI Systems with Sentences that Require Simple Lexical Inferences. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 650\u2013655. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-2103"},{"key":"2021060823284521400_bib77","doi-asserted-by":"crossref","unstructured":"Fr\u00e9deric Godin, Kris Demuynck, Joni Dambre, Wesley De Neve, and Thomas Demeester. 2018. Explaining Character-Aware Neural Networks for Word-Level Prediction: Do They Discover Linguistic Rules? In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3275\u20133284. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D18-1365"},{"key":"2021060823284521400_bib78","doi-asserted-by":"crossref","unstructured":"Yoav Goldberg . 2017. Neural Network methods for Natural Language Processing, volume 10 of Synthesis Lectures on Human Language Technologies. Morgan & Claypool Publishers.","DOI":"10.2200\/S00762ED1V01Y201703HLT037"},{"key":"2021060823284521400_bib79","unstructured":"Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning, MIT Press. http:\/\/www.deepleaningbook.org."},{"key":"2021060823284521400_bib80","unstructured":"Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu , DavidWarde-Farley, SherjilOzair, AaronCourville, and YoshuaBengio. 2014. Generative Adversarial Nets. In Advances in Neural Information Processing Systems, pages 2672\u20132680."},{"key":"2021060823284521400_bib81","unstructured":"Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and Harnessing Adversarial Examples. In International Conference on Learning Representations (ICLR)."},{"key":"2021060823284521400_bib82","doi-asserted-by":"crossref","unstructured":"Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018. Colorless Green Recurrent Networks Dream Hierarchically. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1195\u20131205. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N18-1108"},{"key":"2021060823284521400_bib83","doi-asserted-by":"crossref","unstructured":"Abhijeet Gupta, Gemma Boleda, Marco Baroni, and Sebastian Pad\u00f3. 2015. Distributional vectors encode referential attributes. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 12\u201321. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D15-1002"},{"key":"2021060823284521400_bib84","doi-asserted-by":"crossref","unstructured":"Pankaj Gupta and Hinrich Sch\u00fctze. 2018. LISA: Explaining Recurrent Neural Network Judgments via Layer-wIse Semantic Accumulation and Example to Pattern Transformation. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 154\u2013164. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W18-5418"},{"key":"2021060823284521400_bib85","doi-asserted-by":"crossref","unstructured":"Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. Annotation Artifacts in Natural Language Inference Data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 107\u2013112. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N18-2017"},{"key":"2021060823284521400_bib86","doi-asserted-by":"crossref","unstructured":"Catherine L. Harris . 1990. Connectionism and Cognitive Linguistics. Connection Science, 2(1\u20132):7\u201333.","DOI":"10.1080\/09540099008915660"},{"key":"2021060823284521400_bib87","doi-asserted-by":"crossref","unstructured":"David Harwath and James Glass. 2017. Learning Word-Like Units from Joint Audio-Visual Analysis. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 506\u2013517. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P17-1047"},{"key":"2021060823284521400_bib88","unstructured":"Georg Heigold, G\u00fcnter Neumann, and Josef van Genabith. 2018. How Robust Are Character-Based Word Embeddings in Tagging and MT Against Wrod Scramlbing or Randdm Nouse? In Proceedings of the 13th Conference of The Association for Machine Translation in the Americas (Volume 1: Research Track), pages 68\u201379."},{"key":"2021060823284521400_bib89","doi-asserted-by":"crossref","unstructured":"Felix Hill, Roi Reichart, and Anna Korhonen. 2015. SimLex-999: Evaluating Semantic Models with (Genuine) Similarity Estimation. Computational Linguistics, 41(4):665\u2013695.","DOI":"10.1162\/COLI_a_00237"},{"key":"2021060823284521400_bib90","doi-asserted-by":"crossref","unstructured":"Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2018. Visualisation and \u201cdiagnostic classifiers\u201d reveal how recurrent and recursive neural networks process hierarchical structure. Journal of Artificial Intelligence Research, 61:907\u2013926.","DOI":"10.1613\/jair.1.11196"},{"key":"2021060823284521400_bib91","doi-asserted-by":"crossref","unstructured":"Pierre Isabelle, Colin Cherry, and George Foster. 2017. A Challenge Set Approach to Evaluating Machine Translation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2486\u20132496. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D17-1263"},{"key":"2021060823284521400_bib92","unstructured":"Pierre Isabelle and Roland Kuhn. 2018. A Challenge Set for French\u2013\u00bf English Machine Translation. arXiv preprint arXiv:1806.02725v2."},{"key":"2021060823284521400_bib93","unstructured":"Hitoshi Isahara . 1995. JEIDA\u2019s test-sets for quality evaluation of MT systems\u2014technical evaluation from the developer\u2019s point of view. In Proceedings of MT Summit V."},{"key":"2021060823284521400_bib94","doi-asserted-by":"crossref","unstructured":"Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018. Adversarial Example Generation with Syntactically Controlled Paraphrase Networks. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1875\u20131885. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N18-1170"},{"key":"2021060823284521400_bib95","doi-asserted-by":"crossref","unstructured":"Alon Jacovi, Oren Sar Shalom, and Yoav Goldberg. 2018. Understanding Convolutional Neural Networks for Text Classification. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 56\u201365. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W18-5408"},{"key":"2021060823284521400_bib96","doi-asserted-by":"crossref","unstructured":"Inigo Jauregi Unanue, Ehsan Zare Borzeshi, and Massimo Piccardi. 2018. A Shared Attention Mechanism for Interpretation of Neural Automatic Post-Editing Systems. In Proceedings of the 2nd Workshop on Neural Machine Translation and Generation, pages 11\u201317. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W18-2702"},{"key":"2021060823284521400_bib97","doi-asserted-by":"crossref","unstructured":"Robin Jia and Percy Liang. 2017. Adversarial examples for evaluating reading comprehension systems. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2021\u20132031. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D17-1215"},{"key":"2021060823284521400_bib98","unstructured":"Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016. Exploring the Limits of Language Modeling. arXiv preprint arXiv:1602.02410v2."},{"key":"2021060823284521400_bib99","doi-asserted-by":"crossref","unstructured":"Akos K\u00e1d\u00e1r, Grzegorz Chrupa\u0142a, and Afra Alishahi. 2017. Representation of Linguistic Form and Function in Recurrent Neural Networks. Computational Linguistics, 43(4):761\u2013780.","DOI":"10.1162\/COLI_a_00300"},{"key":"2021060823284521400_bib100","unstructured":"Andrej Karpathy, Justin Johnson, and Fei-Fei Li. 2015. Visualizing and Understanding Recurrent Networks. arXiv preprint arXiv:1506.02078v2."},{"key":"2021060823284521400_bib101","doi-asserted-by":"crossref","unstructured":"Urvashi Khandelwal, He He , PengQi, and DanJurafsky. 2018. Sharp Nearby, Fuzzy Far Away: How Neural Language Models Use Context. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 284\u2013294. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-1027"},{"key":"2021060823284521400_bib102","doi-asserted-by":"crossref","unstructured":"Margaret King and Kirsten Falkedal. 1990. Using Test Suites in Evaluation of Machine Translation Systems. In COLNG 1990 Volume 2: Papers Presented to the 13th International Conference on Computational Linguistics.","DOI":"10.3115\/997939.997976"},{"key":"2021060823284521400_bib103","doi-asserted-by":"crossref","unstructured":"Eliyahu Kiperwasser and Yoav Goldberg. 2016. Simple and Accurate Dependency Parsing Using Bidirectional LSTM Feature Representations. Transactions of the Association for Computational Linguistics, 4:313\u2013327.","DOI":"10.1162\/tacl_a_00101"},{"key":"2021060823284521400_bib104","unstructured":"Sungryong Koh, Jinee Maeng, Ji-Young Lee, Young-Sook Chae, and Key-Sun Choi. 2001. A test suite for evaluation of English-to-Korean machine translation systems. In MT Summit Conference."},{"key":"2021060823284521400_bib105","doi-asserted-by":"crossref","unstructured":"Arne K\u00f6hn . 2015. What\u2019s in an Embedding? Analyzing Word Embeddings through Multilingual Evaluation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 2067\u20132073, Lisbon, Portugal. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D15-1246"},{"key":"2021060823284521400_bib106","unstructured":"Volodymyr Kuleshov, Shantanu Thakoor, Tingfung Lau, and Stefano Ermon. 2018. Adversarial Examples for Natural Language Classification Problems."},{"key":"2021060823284521400_bib107","unstructured":"Brenden Lake and Marco Baroni. 2018. Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent Networks. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 2873\u20132882, Stockholmsm\u00e4ssan, Stockholm, Sweden. PMLR."},{"key":"2021060823284521400_bib108","unstructured":"Sabine Lehmann, Stephan Oepen, Sylvie Regnier-Prost, Klaus Netter, Veronika Lux, Judith Klein, Kirsten Falkedal, Frederik Fouvry, Dominique Estival, Eva Dauphin, Herve Compagnion, Judith Baur, Lorna Balkan, and Doug Arnold. 1996. TSNLP\u2014Test Suites for Natural Language Processing. In COLING 1996 Volume 2: The 16th International Conference on Computational Linguistics."},{"key":"2021060823284521400_bib109","doi-asserted-by":"crossref","unstructured":"Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2016. Rationalizing Neural Predictions. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 107\u2013117. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D16-1011"},{"key":"2021060823284521400_bib110","unstructured":"Ira Leviant and Roi Reichart. 2015. Separated by an Un-Common Language: Towards Judgment Language Informed Vector Space Modeling. arXiv preprint arXiv:1508.00106v5."},{"key":"2021060823284521400_bib111","unstructured":"Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016a. Visualizing and Understanding Neural Models in NLP. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 681\u2013691. Association for Computational Linguistics."},{"key":"2021060823284521400_bib112","unstructured":"Jiwei Li, Will Monroe, and Dan Jurafsky. 2016b. Understanding Neural Networks through Representation Erasure. arXiv preprint arXiv: 1612.08220v3."},{"key":"2021060823284521400_bib113","doi-asserted-by":"crossref","unstructured":"Bin Liang, Hongcheng Li, Miaoqiang Su , PanBian, XirongLi, and WenchangShi. 2018. Deep Text Classification Can Be Fooled. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI-18, pages 4208\u20134215. International Joint Conferences on Artificial Intelligence Organization.","DOI":"10.24963\/ijcai.2018\/585"},{"key":"2021060823284521400_bib114","doi-asserted-by":"crossref","unstructured":"Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016. Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies. Transactions of the Association for Computational Linguistics, 4:521\u2013535.","DOI":"10.1162\/tacl_a_00115"},{"key":"2021060823284521400_bib115","unstructured":"Zachary C. Lipton . 2016. The Mythos of Model Interpretability. In ICML Workshop on Human Interpretability of Machine Learning."},{"key":"2021060823284521400_bib116","unstructured":"Nelson F. Liu, Omer Levy, Roy Schwartz, Chenhao Tan, and Noah A. Smith. 2018. LSTMs Exploit Linguistic Attributes of Data. In Proceedings of The Third Workshop on Representation Learning for NLP, pages 180\u2013186. Association for Computational Linguistics."},{"key":"2021060823284521400_bib117","unstructured":"Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Song. 2017. Delving into Transferable Adversarial Examples and Black-Box Attacks. In International Conference on Learning Representations (ICLR)."},{"key":"2021060823284521400_bib118","unstructured":"Thang Luong, Richard Socher, and Christopher Manning. 2013. Better Word Representations with Recursive Neural Networks for Morphology. In Proceedings of the Seventeenth Conference on Computational Natural Language Learning, pages 104\u2013113. Association for Computational Linguistics."},{"key":"2021060823284521400_bib119","doi-asserted-by":"crossref","unstructured":"Jean Maillard and Stephen Clark. 2018. Latent Tree Learning with Differentiable Parsers: Shift-Reduce Parsing and Chart Parsing. In Proceedings of the Workshop on the Relevance of Linguistic Structure in Neural Architectures for NLP, pages 13\u201318. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W18-2903"},{"key":"2021060823284521400_bib120","doi-asserted-by":"crossref","unstructured":"Marco Marelli, Luisa Bentivogli, Marco Baroni, Raffaella Bernardi, Stefano Menini, and Roberto Zamparelli. 2014. SemEval-2014 Task 1: Evaluation of Compositional Distributional Semantic Models on Full Sentences through Semantic Relatedness and Textual Entailment. In Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014), pages 1\u20138. Association for Computational Linguistics.","DOI":"10.3115\/v1\/S14-2001"},{"key":"2021060823284521400_bib121","unstructured":"R. Thomas McCoy, Robert Frank, and Tal Linzen. 2018. Revisiting the poverty of the stimulus: Hierarchical generalization without a hierarchical bias in recurrent neural networks. In Proceedings of the 40th Annual Conference of the Cognitive Science Society."},{"key":"2021060823284521400_bib122","doi-asserted-by":"crossref","unstructured":"Risto Miikkulainen and Michael G. Dyer. 1991. Natural Language Processing with Modular Pdp Networks and Distributed Lexicon. Cognitive Science, 15(3):343\u2013399.","DOI":"10.1207\/s15516709cog1503_2"},{"key":"2021060823284521400_bib123","doi-asserted-by":"crossref","unstructured":"Tom\u00e1\u0161 Mikolov, Martin Karafi\u00e1t, Luk\u00e1\u0161 Burget, Jan \u010cernocky\u0300, and Sanjeev Khudanpur. 2010. Recurrent neural network based language model. In Eleventh Annual Conference of the International Speech Communication Association.","DOI":"10.1109\/ICASSP.2011.5947611"},{"key":"2021060823284521400_bib124","doi-asserted-by":"crossref","unstructured":"Yao Ming, Shaozu Cao, Ruixiang Zhang, Zhen Li, Yuanzhe Chen, Yangqiu Song, and Huamin Qu. 2017. Understanding Hidden Memories of Recurrent Neural Networks. In IEEE Conference on Visual Analytics Science and Technology (IEEE VAST 2017).","DOI":"10.1109\/VAST.2017.8585721"},{"key":"2021060823284521400_bib125","doi-asserted-by":"crossref","unstructured":"Gr\u00e9goire Montavon, Wojciech Samek, and Klaus-Robert M\u00fcller. 2018. Methods for interpreting and understanding deep neural networks. Digital Signal Processing, 73:1\u201315.","DOI":"10.1016\/j.dsp.2017.10.011"},{"key":"2021060823284521400_bib126","doi-asserted-by":"crossref","unstructured":"Pramod Kaushik Mudrakarta, Ankur Taly, Mukund Sundararajan, and Kedar Dhamdhere. 2018. Did the Model Understand the Question? In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1896\u20131906. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-1176"},{"key":"2021060823284521400_bib127","doi-asserted-by":"crossref","unstructured":"James Mullenbach, Sarah Wiegreffe, Jon Duke, Jimeng Sun, and Jacob Eisenstein. 2018. Explainable Prediction of Medical Codes from Clinical Text. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1101\u20131111. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N18-1100"},{"key":"2021060823284521400_bib128","unstructured":"W. James Murdoch, Peter J. Liu, and Bin Yu. 2018. Beyond Word Importance: Contextual Decomposition to Extract Interactions from LSTMs. In International Conference on Learning Representations."},{"key":"2021060823284521400_bib129","unstructured":"Brian Murphy, Partha Talukdar, and Tom Mitchell. 2012. Learning Effective and Interpretable Semantic Models Using Non-Negative Sparse Embedding. In Proceedings of COLING 2012, pages 1933\u20131950. The COLING 2012 Organizing Committee."},{"key":"2021060823284521400_bib130","doi-asserted-by":"crossref","unstructured":"Tasha Nagamine, Michael L. Seltzer, and Nima Mesgarani. 2015. Exploring How Deep Neural Networks Form Phonemic Categories. In Interspeech 2015.","DOI":"10.21437\/Interspeech.2015-422"},{"key":"2021060823284521400_bib131","doi-asserted-by":"crossref","unstructured":"Tasha Nagamine, Michael L. Seltzer, and Nima Mesgarani. 2016. On the Role of Nonlinear Transformations in Deep Neural Network Acoustic Models. In Interspeech 2016, pages 803\u2013807.","DOI":"10.21437\/Interspeech.2016-1406"},{"key":"2021060823284521400_bib132","unstructured":"Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018. Stress Test Evaluation for Natural Language Inference. In Proceedings of the 27th International Conference on Computational Linguistics, pages 2340\u20132353. Association for Computational Linguistics."},{"key":"2021060823284521400_bib133","doi-asserted-by":"crossref","unstructured":"Nina Narodytska and Shiva Kasiviswanathan. 2017. Simple Black-Box Adversarial Attacks on Deep Neural Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 1310\u20131318.","DOI":"10.1109\/CVPRW.2017.172"},{"key":"2021060823284521400_bib134","doi-asserted-by":"crossref","unstructured":"Lars Niklasson and Fredrik Lin\u00e5ker. 2000. Distributed representations for extended syntactic transformation. Connection Science, 12(3\u20134):299\u2013314.","DOI":"10.1080\/09540090010014070"},{"key":"2021060823284521400_bib135","doi-asserted-by":"crossref","unstructured":"Tong Niu and Mohit Bansal. 2018. Adversarial Over-Sensitivity and Over-Stability Strategies for Dialogue Models. In Proceedings of the 22nd Conference on Computational Natural Language Learning, pages 486\u2013496. Association for Computational Linguistics.","DOI":"10.18653\/v1\/K18-1047"},{"key":"2021060823284521400_bib136","unstructured":"Nicolas Papernot, Patrick McDaniel, and Ian Goodfellow. 2016. Transferability in Machine Learning: From Phenomena to Black-Box Attacks Using Adversarial Samples. arXiv preprint arXiv:1605.07277v1."},{"key":"2021060823284521400_bib137","doi-asserted-by":"crossref","unstructured":"Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, Somesh Jha, Z. Berkay Celik, and Ananthram Swami. 2017. Practical Black-Box Attacks Against Machine Learning. In Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security, ASIA CCS \u201917, pages 506\u2013519, New York, NY, USA, ACM.","DOI":"10.1145\/3052973.3053009"},{"key":"2021060823284521400_bib138","doi-asserted-by":"crossref","unstructured":"Nicolas Papernot, Patrick McDaniel, Ananthram Swami, and Richard Harang. 2016. Crafting Adversarial Input Sequences for Recurrent Neural Networks. In Military Communications Conference, MILCOM 2016, pages 49\u201354. IEEE.","DOI":"10.1109\/MILCOM.2016.7795300"},{"key":"2021060823284521400_bib139","doi-asserted-by":"crossref","unstructured":"Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. 2018. Multimodal Explanations: Justifying Decisions and Pointing to the Evidence. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR.2018.00915"},{"key":"2021060823284521400_bib140","doi-asserted-by":"crossref","unstructured":"Sungjoon Park, JinYeong Bak, and Alice Oh. 2017. Rotated Word Vector Representations and Their Interpretability. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 401\u2013411. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D17-1041"},{"key":"2021060823284521400_bib141","doi-asserted-by":"crossref","unstructured":"Matthew Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih. 2018. Dissecting Contextual Word Embeddings: Architecture and Representation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1499\u20131509. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D18-1179"},{"key":"2021060823284521400_bib142","doi-asserted-by":"crossref","unstructured":"Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, and Benjamin Van Durme. 2018a. Collecting Diverse Natural Language Inference Problems for Sentence Representation Evaluation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 67\u201381. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D18-1007"},{"key":"2021060823284521400_bib143","doi-asserted-by":"crossref","unstructured":"Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis Only Baselines in Natural Language Inference. In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics, pages 180\u2013191. Association for Computational Linguistics.","DOI":"10.18653\/v1\/S18-2023"},{"key":"2021060823284521400_bib144","doi-asserted-by":"crossref","unstructured":"Jordan B. Pollack . 1990. Recursive distributed representations. Artificial Intelligence, 46(1):77\u2013105.","DOI":"10.1016\/0004-3702(90)90005-K"},{"key":"2021060823284521400_bib145","doi-asserted-by":"crossref","unstructured":"Peng Qian, Xipeng Qiu, and Xuanjing Huang. 2016a. Analyzing Linguistic Knowledge in Sequential Model of Sentence. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 826\u2013835, Austin, Texas. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D16-1079"},{"key":"2021060823284521400_bib146","doi-asserted-by":"crossref","unstructured":"Peng Qian, Xipeng Qiu, and Xuanjing Huang. 2016b. Investigating Language Universal and Specific Properties in Word Embeddings. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1478\u20131488, Berlin, Germany. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P16-1140"},{"key":"2021060823284521400_bib147","doi-asserted-by":"crossref","unstructured":"Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018. Semantically Equivalent Adversarial Rules for Debugging NLP models. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 856\u2013865. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-1079"},{"key":"2021060823284521400_bib148","unstructured":"Mat\u012bss Rikters . 2018. Debugging Neural Machine Translations. arXiv preprint arXiv:1808. 02733v1."},{"key":"2021060823284521400_bib149","doi-asserted-by":"crossref","unstructured":"Annette Rios Gonzales, Laura Mascarell, and Rico Sennrich. 2017. Improving Word Sense Disambiguation in Neural Machine Translation with Sense Embeddings. In Proceedings of the Second Conference on Machine Translation, pages 11\u201319. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W17-4702"},{"key":"2021060823284521400_bib150","unstructured":"Tim Rockt\u00e4schel, Edward Grefenstette, Karl Moritz Hermann, Tom\u00e1\u0161 Ko\u010disky\u0300, and Phil Blunsom. 2016. Reasoning about Entailment with Neural Attention. In International Conference on Learning Representations (ICLR)."},{"key":"2021060823284521400_bib151","doi-asserted-by":"crossref","unstructured":"Andras Rozsa, Ethan M. Rudd, and Terrance E. Boult. 2016. Adversarial Diversity and Hard Positive Generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 25\u201332.","DOI":"10.1109\/CVPRW.2016.58"},{"key":"2021060823284521400_bib152","doi-asserted-by":"crossref","unstructured":"Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018. Gender Bias in Coreference Resolution. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 8\u201314. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N18-2002"},{"key":"2021060823284521400_bib153","unstructured":"D. E. Rumelhart and J. L. McClelland. 1986. Parallel Distributed Processing: Explorations in the Microstructure of Cognition. volume 2, chapter On Leaning the Past Tenses of English Verbs, pages 216\u2013271. MIT Press, Cambridge, MA, USA."},{"key":"2021060823284521400_bib154","unstructured":"Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015. A Neural Attention Model for Abstractive Sentence Summarization. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 379\u2013389. Association for Computational Linguistics."},{"key":"2021060823284521400_bib155","doi-asserted-by":"crossref","unstructured":"Keisuke Sakaguchi, Kevin Duh, Matt Post, and Benjamin Van Durme. 2017. Robsut Wrod Reocginiton via Semi-Character Recurrent Neural Network. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, February 4-9, 2017, San Francisco, California, USA., pages 3281\u20133287. AAAI Press.","DOI":"10.1609\/aaai.v31i1.10970"},{"key":"2021060823284521400_bib156","unstructured":"Suranjana Samanta and Sameep Mehta. 2017. Towards Crafting Text Adversarial Samples. arXiv preprint arXiv:1707.02812v1."},{"key":"2021060823284521400_bib157","doi-asserted-by":"crossref","unstructured":"Ivan Sanchez, Jeff Mitchell, and Sebastian Riedel. 2018. Behavior Analysis of NLI Models: Uncovering the Influence of Three Factors on Robustness. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1975\u20131985. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N18-1179"},{"key":"2021060823284521400_bib158","doi-asserted-by":"crossref","unstructured":"Motoki Sato, Jun Suzuki, Hiroyuki Shindo, and Yuji Matsumoto. 2018. Interpretable Adversarial Perturbation in Input Embedding Space for Text. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI-18, pages 4323\u20134330. International Joint Conferences on Artificial Intelligence Organization.","DOI":"10.24963\/ijcai.2018\/601"},{"key":"2021060823284521400_bib159","unstructured":"Lutfi Kerem Senel, Ihsan Utlu, Veysel Yucesoy, Aykut Koc, and Tolga Cukur. 2018. Semantic Structure and Interpretability of Word Embeddings. IEEE\/ACM Transactions on Audio, Speech, and Language Processing."},{"key":"2021060823284521400_bib160","doi-asserted-by":"crossref","unstructured":"Rico Sennrich . 2017. How Grammatical Is Character-Level Neural Machine Translation? Assessing MT Quality with Contrastive Translation Pairs. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 376\u2013382. Association for Computational Linguistics.","DOI":"10.18653\/v1\/E17-2060"},{"key":"2021060823284521400_bib161","unstructured":"Haoyue Shi, Jiayuan Mao, Tete Xiao, Yuning Jiang, and Jian Sun. 2018. Learning Visually- Grounded Semantics from Contrastive Adversarial Samples. In Proceedings of the 27th International Conference on Computational Linguistics, pages 3715\u20133727. Association for Computational Linguistics."},{"key":"2021060823284521400_bib162","doi-asserted-by":"crossref","unstructured":"Xing Shi, Kevin Knight, and Deniz Yuret. 2016a. Why Neural Translations are the Right Length. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2278\u20132282. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D16-1248"},{"key":"2021060823284521400_bib163","doi-asserted-by":"crossref","unstructured":"Xing Shi, Inkit Padhi, and Kevin Knight. 2016b. Does String-Based Neural MT Learn Source Syntax? In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1526\u20131534, Austin, Texas. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D16-1159"},{"key":"2021060823284521400_bib164","unstructured":"Chandan Singh, W. James Murdoch, and Bin Yu. 2018. Hierarchical interpretations for neural network predictions. arXiv preprint arXiv:1806.05337v1."},{"key":"2021060823284521400_bib165","doi-asserted-by":"crossref","unstructured":"Hendrik Strobelt, Sebastian Gehrmann, Michael Behrisch, Adam Perer, Hanspeter Pfister, and Alexander M. Rush. 2018a. Seq2Seq-Vis: A Visual Debugging Tool for Sequence- to-Sequence Models. arXiv preprint arXiv: 1804.09299v1.","DOI":"10.18653\/v1\/W18-5451"},{"key":"2021060823284521400_bib166","doi-asserted-by":"crossref","unstructured":"Hendrik Strobelt, Sebastian Gehrmann, Hanspeter Pfister, and Alexander M. Rush. 2018b. LSTMVis: A Tool for Visual Analysis of Hidden State Dynamics in Recurrent Neural Networks. IEEE Transactions on Visualization and Computer Graphics, 24(1):667\u2013676.","DOI":"10.1109\/TVCG.2017.2744158"},{"key":"2021060823284521400_bib167","unstructured":"Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017. Axiomatic Attribution for Deep Networks. In Proceedings of the 34th International Conference on Machine Learning, Volume 70 of Proceedings of Machine Learning Research, pages 3319\u20133328, International Convention Centre, Sydney, Australia. PMLR."},{"key":"2021060823284521400_bib168","unstructured":"Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to Sequence Learning with Neural Networks. In Advances in neural information processing systems, pages 3104\u20133112."},{"key":"2021060823284521400_bib169","unstructured":"Mirac Suzgun, Yonatan Belinkov, and Stuart M. Shieber. 2019. On Evaluating the Generalization of LSTM Models in Formal Languages. In Proceedings of the Society for Computation in Linguistics (SCiL)."},{"key":"2021060823284521400_bib170","unstructured":"Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2014. Intriguing properties of neural networks. In International Conference on Learning Representations (ICLR)."},{"key":"2021060823284521400_bib171","doi-asserted-by":"crossref","unstructured":"Gongbo Tang, Rico Sennrich, and Joakim Nivre. 2018. An Analysis of Attention Mechanisms: The Case of Word Sense Disambiguation in Neural Machine Translation. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 26\u201335. Association for Computational Linguistics.","DOI":"10.18653\/v1\/W18-6304"},{"key":"2021060823284521400_bib172","doi-asserted-by":"crossref","unstructured":"Yi Tay , Anh TuanLuu, and Siu CheungHui. 2018. CoupleNet: Paying Attention to Couples with Coupled Attention for Relationship Recommendation. In Proceedings of the Twelfth International AAAI Conference on Web and Social Media (ICWSM).","DOI":"10.1609\/icwsm.v12i1.15007"},{"key":"2021060823284521400_bib173","doi-asserted-by":"crossref","unstructured":"Ke Tran , AriannaBisazza, and ChristofMonz. 2018. The Importance of Being Recurrent for Modeling Hierarchical Structure. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4731\u20134736. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D18-1503"},{"key":"2021060823284521400_bib174","unstructured":"Eva Vanmassenhove, Jinhua Du , and AndyWay. 2017. Investigating \u201cAspect\u201d in NMT and SMT: Translating the English Simple Past and Present Perfect. Computational Linguistics in the Netherlands Journal, 7:109\u2013128."},{"key":"2021060823284521400_bib175","unstructured":"Sara Veldhoen, Dieuwke Hupkes, and Willem Zuidema. 2016. Diagnostic Classifiers: Revealing How Neural Networks Process Hierarchical Structure. In CEUR Workshop Proceedings."},{"key":"2021060823284521400_bib176","doi-asserted-by":"crossref","unstructured":"Elena Voita, Pavel Serdyukov, Rico Sennrich, and Ivan Titov. 2018. Context-Aware Neural Machine Translation Learns Anaphora Resolution. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1264\u20131274. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-1117"},{"key":"2021060823284521400_bib177","doi-asserted-by":"crossref","unstructured":"Ekaterina Vylomova, Trevor Cohn, Xuanli He, and Gholamreza Haffari. 2016. Word Representation Models for Morphologically Rich Languages in Neural Machine Translation. arXiv preprint arXiv:1606.04217v1.","DOI":"10.18653\/v1\/W17-4115"},{"key":"2021060823284521400_bib178","doi-asserted-by":"crossref","unstructured":"Alex Wang, Amapreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018a. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. arXiv preprint arXiv:1804.07461v1.","DOI":"10.18653\/v1\/W18-5446"},{"key":"2021060823284521400_bib179","doi-asserted-by":"crossref","unstructured":"Shuai Wang, Yanmin Qian, and Kai Yu. 2017a. What Does the Speaker Embedding Encode? In Interspeech 2017, pages 1497\u20131501.","DOI":"10.21437\/Interspeech.2017-1125"},{"key":"2021060823284521400_bib180","doi-asserted-by":"crossref","unstructured":"Xinyi Wang, Hieu Pham, Pengcheng Yin, and Graham Neubig. 2018b. A Tree-Based Decoder for Neural Machine Translation. In Conference on Empirical Methods in Natural Language Processing (EMNLP). Brussels, Belgium.","DOI":"10.18653\/v1\/D18-1509"},{"key":"2021060823284521400_bib181","doi-asserted-by":"crossref","unstructured":"Yu-Hsuan Wang, Cheng-Tao Chung, and Hung-yi Lee. 2017b. Gate Activation Signal Analysis for Gated Recurrent Neural Networks and Its Correlation with Phoneme Boundaries. In Interspeech 2017.","DOI":"10.21437\/Interspeech.2017-877"},{"key":"2021060823284521400_bib182","doi-asserted-by":"crossref","unstructured":"Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018. On the Practical Computational Power of Finite Precision RNNs for Language Recognition. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 740\u2013745. Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-2117"},{"key":"2021060823284521400_bib183","doi-asserted-by":"crossref","unstructured":"Adina Williams, Andrew Drozdov, and Samuel R. Bowman. 2018. Do latent tree learning models identify meaningful structure in sentences?Transactions of the Association for Computational Linguistics, 6:253\u2013267.","DOI":"10.1162\/tacl_a_00019"},{"key":"2021060823284521400_bib184","unstructured":"Zhizheng Wu and SimonKing. 2016. Investigating gated recurrent networks for speech synthesis. In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5140\u20135144. IEEE."},{"key":"2021060823284521400_bib185","unstructured":"Puyudi Yang, Jianbo Chen, Cho-Jui Hsieh, Jane-Ling Wang, and Michael I. Jordan. 2018. Greedy Attack and Gumbel Attack: Generating Adversarial Examples for Discrete Data. arXiv preprint arXiv:1805.12316v1."},{"key":"2021060823284521400_bib186","doi-asserted-by":"crossref","unstructured":"Wenpeng Yin, Hinrich Sch\u00fctze, Bing Xiang, and Bowen Zhou. 2016. ABCNN: Attention-Based Convolutional Neural Network for Modeling Sentence Pairs. Transactions of the Association for Computational Linguistics, 4:259\u2013272.","DOI":"10.1162\/tacl_a_00097"},{"key":"2021060823284521400_bib187","unstructured":"Xiaoyong Yuan, Pan He, Qile Zhu, and Xiaolin Li. 2017. Adversarial Examples: Attacks and Defenses for Deep Learning. arXiv preprint arXiv:1712.07107v3."},{"key":"2021060823284521400_bib188","unstructured":"Omar Zaidan, Jason Eisner, and Christine Piatko. 2007. Using \u201cAnnotator Rationales\u201d to Improve Machine Learning for Text Categorization. In Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Proceedings of the Main Conference, pages 260\u2013267. Association for Computational Linguistics."},{"key":"2021060823284521400_bib189","doi-asserted-by":"crossref","unstructured":"Quan-shi Zhang and Song-chun Zhu. 2018. Visual interpretability for deep learning: A survey. Frontiers of Information Technology & Electronic Engineering, 19(1):27\u201339.","DOI":"10.1631\/FITEE.1700808"},{"key":"2021060823284521400_bib190","doi-asserted-by":"crossref","unstructured":"Ye Zhang , IainMarshall, and Byron C.Wallace. 2016. Rationale-Augmented Convolutional Neural Networks for Text Classification. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 795\u2013804. Association for Computational Linguistics.","DOI":"10.18653\/v1\/D16-1076"},{"key":"2021060823284521400_bib191","doi-asserted-by":"crossref","unstructured":"Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018a. Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 15\u201320. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N18-2003"},{"key":"2021060823284521400_bib192","unstructured":"Junbo Zhao, Yoon Kim, Kelly Zhang, Alexander Rush, and Yann LeCun. 2018b. Adversarially Regularized Autoencoders. In Proceedings of the 35th International Conference on Machine Learning, Volume 80 of Proceedings of Machine Learning Research, pages 5902\u20135911, Stockholmsm\u00e4ssan, Stockholm, Sweden. PMLR."},{"key":"2021060823284521400_bib193","unstructured":"Zhengli Zhao, Dheeru Dua, and Sameer Singh. 2018c. Generating Natural Adversarial Examples. In International Conference on Learning Representations."}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00254\/1923061\/tacl_a_00254.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"http:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00254\/1923061\/tacl_a_00254.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,15]],"date-time":"2023-09-15T09:07:15Z","timestamp":1694768835000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00254\/43503\/Analysis-Methods-in-Neural-Language-Processing-A"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,4,1]]},"references-count":193,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00254","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2019,4]]},"published":{"date-parts":[[2019,4,1]]}}}