{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T16:43:30Z","timestamp":1783183410487,"version":"3.54.6"},"reference-count":67,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2022,11,29]],"date-time":"2022-11-29T00:00:00Z","timestamp":1669680000000},"content-version":"vor","delay-in-days":332,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,11,22]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Natural Language Inference (NLI) and Semantic Textual Similarity (STS) are widely used benchmark tasks for compositional evaluation of pre-trained language models. Despite growing interest in linguistic universals, most NLI\/STS studies have focused almost exclusively on English. In particular, there are no available multilingual NLI\/STS datasets in Japanese, which is typologically different from English and can shed light on the currently controversial behavior of language models in matters such as sensitivity to word order and case particles. Against this background, we introduce JSICK, a Japanese NLI\/STS dataset that was manually translated from the English dataset SICK. We also present a stress-test dataset for compositional inference, created by transforming syntactic structures of sentences in JSICK to investigate whether language models are sensitive to word order and case particles. We conduct baseline experiments on different pre-trained language models and compare the performance of multilingual models when applied to Japanese and other languages. The results of the stress-test experiments suggest that the current pre-trained language models are insensitive to word order and case marking.<\/jats:p>","DOI":"10.1162\/tacl_a_00518","type":"journal-article","created":{"date-parts":[[2022,11,29]],"date-time":"2022-11-29T18:32:16Z","timestamp":1669746736000},"page":"1266-1284","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":11,"title":["Compositional Evaluation on Japanese Textual Entailment and Similarity"],"prefix":"10.1162","volume":"10","author":[{"given":"Hitomi","family":"Yanaka","sequence":"first","affiliation":[{"name":"The University of Tokyo, Japan. hyanaka@is.s.u-tokyo.ac.jp"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Koji","family":"Mineshima","sequence":"additional","affiliation":[{"name":"Keio University, Japan. minesima@abelard.flet.keio.ac.jp"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2022,11,22]]},"reference":[{"key":"2022112918312854100_bib1","doi-asserted-by":"publisher","first-page":"497","DOI":"10.18653\/v1\/S16-1081","article-title":"SemEval-2016 task 1: Semantic textual similarity, monolingual and cross-lingual evaluation","volume-title":"Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016)","author":"Agirre","year":"2016"},{"key":"2022112918312854100_bib2","article-title":"FarsTail: A Persian natural language inference dataset","volume":"cs.CL\/2009.08820","author":"Amirkhani","year":"2020","journal-title":"CoRR"},{"key":"2022112918312854100_bib3","unstructured":"Masayuki\n              Asahara\n             and YujiMatsumoto. 2003. ipadic version 2.7.0 User\u2019s Manual. Nara Institute of Science and Technology."},{"key":"2022112918312854100_bib4","doi-asserted-by":"publisher","first-page":"632","DOI":"10.18653\/v1\/D15-1075","article-title":"A large annotated corpus for learning natural language inference","volume-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing","author":"Bowman","year":"2015"},{"key":"2022112918312854100_bib5","doi-asserted-by":"crossref","first-page":"1","DOI":"10.18653\/v1\/S17-2001","article-title":"SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation","volume-title":"Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017)","author":"Cer","year":"2017"},{"key":"2022112918312854100_bib6","doi-asserted-by":"publisher","first-page":"4536","DOI":"10.18653\/v1\/2020.emnlp-main.367","article-title":"Improving multilingual models with language-clustered vocabularies","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Chung","year":"2020"},{"key":"2022112918312854100_bib7","doi-asserted-by":"publisher","first-page":"8440","DOI":"10.18653\/v1\/2020.acl-main.747","article-title":"Unsupervised cross-lingual representation learning at scale","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Conneau","year":"2020"},{"key":"2022112918312854100_bib8","doi-asserted-by":"publisher","first-page":"2475","DOI":"10.18653\/v1\/D18-1269","article-title":"XNLI: Evaluating cross-lingual sentence representations","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Conneau","year":"2018"},{"key":"2022112918312854100_bib9","article-title":"FraCaS\u2013a framework for computational semantics","volume":"D6","author":"Cooper","year":"1994","journal-title":"Deliverable"},{"key":"2022112918312854100_bib10","doi-asserted-by":"publisher","first-page":"177","DOI":"10.1007\/11736790_9","article-title":"The pascal recognising textual entailment challenge","volume-title":"Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment","author":"Dagan","year":"2006"},{"key":"2022112918312854100_bib11","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"issue":"285","key":"2022112918312854100_bib12","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1093\/mind\/LXXII.285.1","article-title":"Compound thoughts","volume":"72","author":"Frege","year":"1963","journal-title":"Mind"},{"key":"2022112918312854100_bib13","first-page":"81","article-title":"Natural language inference with mixed effects","volume-title":"Proceedings of the Ninth Joint Conference on Lexical and Computational Semantics","author":"Gantt","year":"2020"},{"key":"2022112918312854100_bib14","doi-asserted-by":"publisher","first-page":"650","DOI":"10.18653\/v1\/P18-2103","article-title":"Breaking NLI systems with sentences that require simple lexical inferences","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Glockner","year":"2018"},{"key":"2022112918312854100_bib15","doi-asserted-by":"publisher","first-page":"1958","DOI":"10.18653\/v1\/2020.acl-main.177","article-title":"Probing linguistic systematicity","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Goodwin","year":"2020"},{"key":"2022112918312854100_bib16","doi-asserted-by":"publisher","first-page":"12946","DOI":"10.1609\/aaai.v35i14.17531","article-title":"BERT & family eat word salad: Experiments with text understanding","volume-title":"Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021","author":"Gupta","year":"2021"},{"key":"2022112918312854100_bib17","doi-asserted-by":"publisher","first-page":"422","DOI":"10.18653\/v1\/2020.findings-emnlp.39","article-title":"KorNLI and KorSTS: New benchmark datasets for Korean natural language understanding","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Ham","year":"2020"},{"key":"2022112918312854100_bib18","volume-title":"Definiteness and Indefiniteness. A Study in Reference and Grammaticality Prediction","author":"Hawkins","year":"1978"},{"key":"2022112918312854100_bib19","first-page":"6827","article-title":"Japanese realistic textual entailment corpus","volume-title":"Proceedings of the 12th Language Resources and Evaluation Conference","author":"Hayashibe","year":"2020"},{"key":"2022112918312854100_bib20","unstructured":"Irene\n              Heim\n            \n          . 1982. The Semantics of Definite and Indefinite Noun Phrases. Ph.D. thesis, UMass Amherst."},{"key":"2022112918312854100_bib21","doi-asserted-by":"publisher","first-page":"204","DOI":"10.18653\/v1\/2021.acl-short.27","article-title":"How effective is BERT without word ordering? implications for language understanding and data privacy","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)","author":"Hessel","year":"2021"},{"key":"2022112918312854100_bib22","volume-title":"Japanese: Descriptive Grammar","author":"Hinds","year":"1986"},{"key":"2022112918312854100_bib23","volume-title":"Logical Form Constraints and Configurational Structures in Japanese","author":"Hoji","year":"1985"},{"key":"2022112918312854100_bib24","doi-asserted-by":"publisher","first-page":"3512","DOI":"10.18653\/v1\/2020.findings-emnlp.314","article-title":"OCNLI: Original Chinese Natural Language Inference","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Hai","year":"2020"},{"key":"2022112918312854100_bib25","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1016\/B978-044481714-3\/50011-4","article-title":"Compositionality","volume-title":"Handbook of Logic and Language","author":"Janssen","year":"1997"},{"key":"2022112918312854100_bib26","doi-asserted-by":"publisher","first-page":"6282","DOI":"10.18653\/v1\/2020.acl-main.560","article-title":"The state and fate of linguistic diversity and inclusion in the NLP world","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Joshi","year":"2020"},{"key":"2022112918312854100_bib27","article-title":"Textual inference: Getting logic from humans","volume-title":"Proceedings of the 12th International Conference on Computational Semantics (IWCS) \u2014 Short papers","author":"Kalouli","year":"2017"},{"issue":"2","key":"2022112918312854100_bib28","doi-asserted-by":"publisher","first-page":"170","DOI":"10.2307\/411200","article-title":"The structure of a semantic theory","volume":"39","author":"Katz","year":"1963","journal-title":"Language"},{"key":"2022112918312854100_bib29","doi-asserted-by":"publisher","first-page":"58","DOI":"10.1007\/978-3-319-50953-2_5","article-title":"An inference problem set for evaluating semantic theories and semantic processing systems for Japanese","volume-title":"New Frontiers in Artificial Intelligence","author":"Ai","year":"2017"},{"key":"2022112918312854100_bib30","doi-asserted-by":"publisher","first-page":"5203","DOI":"10.18653\/v1\/2021.acl-long.405","article-title":"Lower perplexity is not always human-like","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Kuribayashi","year":"2021"},{"key":"2022112918312854100_bib31","first-page":"2479","article-title":"FlauBERT: Unsupervised language model pre-training for French","volume-title":"Proceedings of the 12th Language Resources and Evaluation Conference","author":"Le","year":"2020"},{"key":"2022112918312854100_bib32","article-title":"Tregex and tsurgeon: Tools for querying and manipulating tree data structures","volume-title":"Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC\u201906)","author":"Levy","year":"2006"},{"key":"2022112918312854100_bib33","doi-asserted-by":"publisher","first-page":"6008","DOI":"10.18653\/v1\/2020.emnlp-main.484","article-title":"XGLUE: A new benchmark dataset for cross-lingual pre-training, understanding and generation","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Liang","year":"2020"},{"key":"2022112918312854100_bib34","doi-asserted-by":"publisher","first-page":"5210","DOI":"10.18653\/v1\/2020.acl-main.465","article-title":"How can we accelerate progress towards human-like linguistic generalization?","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Linzen","year":"2020"},{"key":"2022112918312854100_bib35","article-title":"RoBERTa: A robustly optimized bert pretraining approach","volume":"cs.CL\/1907.11692","author":"Liu","year":"2019","journal-title":"CoRR"},{"key":"2022112918312854100_bib36","first-page":"216","article-title":"A SICK cure for the evaluation of compositional distributional semantic models","volume-title":"Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC\u201914)","author":"Marelli","year":"2014"},{"key":"2022112918312854100_bib37","doi-asserted-by":"publisher","first-page":"3428","DOI":"10.18653\/v1\/P19-1334","article-title":"Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"McCoy","year":"2019"},{"key":"2022112918312854100_bib38","doi-asserted-by":"publisher","first-page":"221","DOI":"10.1007\/978-94-010-2506-5_10","article-title":"The proper treatment of quantification in ordinary English","volume-title":"Approaches to Natural Language","author":"Montague","year":"1973"},{"key":"2022112918312854100_bib39","doi-asserted-by":"publisher","first-page":"2292","DOI":"10.18653\/v1\/D15-1276","article-title":"Morphological analysis for unsegmented languages using recurrent neural network language model","volume-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing","author":"Morita","year":"2015"},{"key":"2022112918312854100_bib40","first-page":"2340","article-title":"Stress test evaluation for natural language inference","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics","author":"Naik","year":"2018"},{"issue":"2","key":"2022112918312854100_bib41","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1023\/B:JEAL.0000019058.46668.c1","article-title":"Japanese plurals are exceptional","volume":"13","author":"Nakanishi","year":"2004","journal-title":"Journal of East Asian Linguistics"},{"key":"2022112918312854100_bib42","article-title":"KLUE: Korean language understanding evaluation","volume":"cs.CL\/2105.09680","author":"Park","year":"2021","journal-title":"CoRR"},{"key":"2022112918312854100_bib43","doi-asserted-by":"publisher","first-page":"677","DOI":"10.1162\/tacl_a_00293","article-title":"Inherent disagreements in human textual inferences","volume":"7","author":"Pavlick","year":"2019","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2022112918312854100_bib44","doi-asserted-by":"publisher","first-page":"1145","DOI":"10.18653\/v1\/2021.findings-acl.98","article-title":"Out of order: How important is the sequential order of words in a sentence in natural language understanding tasks?","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Pham","year":"2021"},{"key":"2022112918312854100_bib45","doi-asserted-by":"publisher","first-page":"3532","DOI":"10.18653\/v1\/N19-1356","article-title":"Studying the inductive biases of RNNs with synthetic variations of natural languages","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Ravfogel","year":"2019"},{"key":"2022112918312854100_bib46","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1007\/978-3-319-99722-3_31","article-title":"SICK-BR: A portuguese corpus for inference","volume-title":"Computational Processing of the Portuguese Language","author":"Real","year":"2018"},{"key":"2022112918312854100_bib47","doi-asserted-by":"publisher","first-page":"196","DOI":"10.18653\/v1\/K19-1019","article-title":"Diversify your datasets: Analyzing generalization via controlled variance in adversarial datasets","volume-title":"Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL)","author":"Rozen","year":"2019"},{"key":"2022112918312854100_bib48","doi-asserted-by":"publisher","first-page":"3118","DOI":"10.18653\/v1\/2021.acl-long.243","article-title":"How good is your tokenizer? On the monolingual performance of multilingual language models","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Rust","year":"2021"},{"key":"2022112918312854100_bib49","unstructured":"Mamoru\n              Saito\n            \n          . 1985. Some Asymmetries in Japanese and their Theoretical Implications. Ph.D. thesis, NA Cambridge."},{"key":"2022112918312854100_bib50","doi-asserted-by":"publisher","first-page":"5149","DOI":"10.1109\/ICASSP.2012.6289079","article-title":"Japanese and Korean voice search","volume-title":"2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Schuster","year":"2012"},{"key":"2022112918312854100_bib51","first-page":"173","article-title":"ALUE: Arabic language understanding evaluation","volume-title":"Proceedings of the Sixth Arabic Natural Language Processing Workshop","author":"Seelawi","year":"2021"},{"key":"2022112918312854100_bib52","doi-asserted-by":"publisher","first-page":"4717","DOI":"10.18653\/v1\/2020.emnlp-main.381","article-title":"RussianSuperGLUE: A Russian language understanding evaluation benchmark","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Shavrina","year":"2020"},{"key":"2022112918312854100_bib53","volume-title":"The Languages of Japan","author":"Shibatani","year":"1990"},{"key":"2022112918312854100_bib54","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.230","article-title":"Masked language modeling and the distributional hypothesis: Order word matters pre-training for little","volume":"cs.CL\/2104.06644","author":"Sinha","year":"2021","journal-title":"CoRR"},{"key":"2022112918312854100_bib55","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.569","article-title":"UnNatural Language Inference","volume-title":"Proceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP2021)","author":"Sinha","year":"2021"},{"key":"2022112918312854100_bib56","doi-asserted-by":"publisher","first-page":"54","DOI":"10.18653\/v1\/D18-2010","article-title":"Juman++: A morphological analysis toolkit for scriptio continua","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Tolmachev","year":"2018"},{"key":"2022112918312854100_bib57","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-5446","article-title":"GLUE: A multi-task benchmark and analysis platform for natural language understanding","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Wang","year":"2019"},{"key":"2022112918312854100_bib58","first-page":"454","article-title":"Examining the inductive bias of neural language models with artificial languages","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"White","year":"2021"},{"key":"2022112918312854100_bib59","doi-asserted-by":"publisher","first-page":"1474","DOI":"10.18653\/v1\/2021.eacl-main.126","article-title":"SICK-NL: A dataset for Dutch natural language inference","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Wijnholds","year":"2021"},{"key":"2022112918312854100_bib60","doi-asserted-by":"publisher","first-page":"1112","DOI":"10.18653\/v1\/N18-1101","article-title":"A broad-coverage challenge corpus for sentence understanding through inference","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Williams","year":"2018"},{"key":"2022112918312854100_bib61","first-page":"4762","article-title":"CLUE: A Chinese language understanding evaluation benchmark","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Liang","year":"2020"},{"key":"2022112918312854100_bib62","doi-asserted-by":"publisher","first-page":"920","DOI":"10.18653\/v1\/2021.eacl-main.78","article-title":"Exploring transitivity in neural NLI models through veridicality","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Yanaka","year":"2021"},{"key":"2022112918312854100_bib63","doi-asserted-by":"publisher","first-page":"3687","DOI":"10.18653\/v1\/D19-1382","article-title":"PAWS-X: A cross- lingual adversarial dataset for paraphrase identification","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Yang","year":"2019"},{"key":"2022112918312854100_bib64","doi-asserted-by":"publisher","first-page":"277","DOI":"10.18653\/v1\/P17-1026","article-title":"A* CCG parsing with a supertag and dependency factored model","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Yoshikawa","year":"2017"},{"key":"2022112918312854100_bib65","article-title":"Multilingualization of natural language inference datasets using machine translation (in Japanese)","volume-title":"Proceedings of the 244th Meeting of Natural Language Processing","author":"Yoshikoshi","year":"2020"},{"key":"2022112918312854100_bib66","article-title":"Bertscore: Evaluating text generation with bert","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Zhang","year":"2020"},{"key":"2022112918312854100_bib67","first-page":"1298","article-title":"PAWS: Paraphrase adversaries from word scrambling","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Zhang","year":"2019"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00518\/2060724\/tacl_a_00518.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00518\/2060724\/tacl_a_00518.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,11,29]],"date-time":"2022-11-29T18:32:34Z","timestamp":1669746754000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00518\/113850\/Compositional-Evaluation-on-Japanese-Textual"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022]]},"references-count":67,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00518","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022]]},"published":{"date-parts":[[2022]]}}}