{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T00:17:03Z","timestamp":1782173823855,"version":"3.54.5"},"reference-count":164,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2022,3,26]],"date-time":"2022-03-26T00:00:00Z","timestamp":1648252800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["61872446"],"award-info":[{"award-number":["61872446"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"NSF of Hunan Province","award":["2019JJ20024"],"award-info":[{"award-number":["2019JJ20024"]}]},{"name":"Science and Technology Innovation Program of Hunan Province","award":["2020RC4046"],"award-info":[{"award-number":["2020RC4046"]}]},{"name":"European Research Council","award":["H2020-ERC-2017-ADG 788506"],"award-info":[{"award-number":["H2020-ERC-2017-ADG 788506"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2023,3,31]]},"abstract":"<jats:p>How to transfer the semantic information in a sentence to a computable numerical embedding form is a fundamental problem in natural language processing. An informative universal sentence embedding can greatly promote subsequent natural language processing tasks. However, unlike universal word embeddings, a widely accepted general-purpose sentence embedding technique has not been developed. This survey summarizes the current universal sentence-embedding methods, categorizes them into four groups from a linguistic view, and ultimately analyzes their reported performance. Sentence embeddings trained from words in a bottom-up manner are observed to have different, nearly opposite, performance patterns in downstream tasks compared to those trained from logical relationships between sentences. By comparing differences of training schemes in and between groups, we analyze possible essential reasons for different performance patterns. We additionally collect incentive strategies handling sentences from other models and propose potentially inspiring future research directions.<\/jats:p>","DOI":"10.1145\/3482853","type":"journal-article","created":{"date-parts":[[2022,3,26]],"date-time":"2022-03-26T09:22:29Z","timestamp":1648286549000},"page":"1-42","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["A Brief Overview of Universal Sentence Representation Methods: A Linguistic View"],"prefix":"10.1145","volume":"55","author":[{"given":"Ruiqi","family":"Li","sequence":"first","affiliation":[{"name":"KU Leuven, Belgium and National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiang","family":"Zhao","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marie-Francine","family":"Moens","sequence":"additional","affiliation":[{"name":"KU Leuven, Celestijnenlaan, Heverlee, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,3,26]]},"reference":[{"key":"e_1_3_2_2_2","article-title":"Fine-grained analysis of sentence embeddings using auxiliary prediction tasks","volume":"1608","author":"Adi Yossi","year":"2016","unstructured":"Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2016. Fine-grained analysis of sentence embeddings using auxiliary prediction tasks. Corr abs\/1608.04207 (2016).","journal-title":"Corr"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S15-2045"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/S14-2010"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S16-1081"},{"key":"e_1_3_2_6_2","first-page":"385","volume-title":"Proceedings of the 6th International Workshop on Semantic Evaluation","author":"Agirre Eneko","year":"2012","unstructured":"Eneko Agirre, Daniel M. Cer, Mona T. Diab, and Aitor Gonzalez-Agirre. 2012. SemEval-2012 task 6: a pilot on semantic textual similarity. In Proceedings of the 6th International Workshop on Semantic Evaluation. 385\u2013393."},{"key":"e_1_3_2_7_2","first-page":"32","volume-title":"Proceedings of the 2nd Joint Conference on Lexical and Computational Semantics","author":"Agirre Eneko","year":"2013","unstructured":"Eneko Agirre, Daniel M. Cer, Mona T. Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013. *SEM 2013 shared task: Semantic textual similarity. In Proceedings of the 2nd Joint Conference on Lexical and Computational Semantics. 32\u201343."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00106"},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of 5th International Conference on Learning Representations","author":"Arora Sanjeev","year":"2017","unstructured":"Sanjeev Arora, Yingyu Liang, and Tengyu Ma. 2017. A simple but tough-to-beat baseline for sentence embeddings. In Proceedings of 5th International Conference on Learning Representations."},{"key":"e_1_3_2_10_2","article-title":"Layer normalization","volume":"1607","author":"Ba Lei Jimmy","year":"2016","unstructured":"Lei Jimmy Ba, Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer normalization. Corr abs\/1607.06450 (2016).","journal-title":"Corr"},{"key":"e_1_3_2_11_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations","author":"Bahdanau Dzmitry","year":"2015","unstructured":"Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In Proceedings of the 3rd International Conference on Learning Representations."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-2131"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.5555\/1622248.1622254"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00051"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1075"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jml.2006.07.005"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S17-2001"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-2029"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.2307\/1412159"},{"key":"e_1_3_2_20_2","volume-title":"Proceedings of 3rd International Conference on Learning Representations","author":"Chen Liang-Chieh","year":"2015","unstructured":"Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. 2015. Semantic image segmentation with deep convolutional nets and fully connected CRFs. In Proceedings of 3rd International Conference on Learning Representations."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2699184"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1017"},{"key":"e_1_3_2_23_2","first-page":"67","article-title":"Chinese students word-solving strategies in reading in English","author":"Chern Chiou-Lan","year":"1993","unstructured":"Chiou-Lan Chern. 1993. Chinese students word-solving strategies in reading in English. Sec. Lang. Read. Vocab. Learn. (1993), 67\u201385.","journal-title":"Sec. Lang. Read. Vocab. Learn."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1179"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/0346-251X(80)90003-2"},{"key":"e_1_3_2_26_2","first-page":"2807","volume-title":"Proceedings of the 26th International Conference on Computational Linguistics","author":"Collell Guillem","year":"2016","unstructured":"Guillem Collell and Marie-Francine Moens. 2016. Is an image worth more than a thousand words? On the fine-grain semantic differences between visual and linguistic representations. In Proceedings of the 26th International Conference on Computational Linguistics. 2807\u20132817."},{"key":"e_1_3_2_27_2","first-page":"4378","volume-title":"Proceedings of the 31st AAAI Conference on Artificial Intelligence","author":"Collell Guillem","year":"2017","unstructured":"Guillem Collell, Ted Zhang, and Marie-Francine Moens. 2017. Imagined visual representations as multimodal embeddings. In Proceedings of the 31st AAAI Conference on Artificial Intelligence. 4378\u20134384."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390177"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078186"},{"key":"e_1_3_2_30_2","volume-title":"Proceedings of the 11th International Conference on Language Resources and Evaluation","author":"Conneau Alexis","year":"2018","unstructured":"Alexis Conneau and Douwe Kiela. 2018. SentEval: An evaluation toolkit for universal sentence representations. In Proceedings of the 11th International Conference on Language Resources and Evaluation."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1070"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1198"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1023860521975"},{"key":"e_1_3_2_34_2","unstructured":"Andrew M. Dai Christopher Olah and Quoc V. Le. 2015. Document embedding with paragraph vectors. arxiv:1507.07998."},{"key":"e_1_3_2_35_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Volume 1 Jill Burstein Christy Doran and Thamar Solorio (Eds.). Association for Computational Linguistics 4171\u20134186."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.3115\/1220355.1220406"},{"key":"e_1_3_2_37_2","unstructured":"Jingfei Du Edouard Grave Beliz Gunel Vishrav Chaudhary Onur Celebi Michael Auli Ves Stoyanov and Alexis Conneau. 2020. Self-training improves pre-training for natural language understanding. arxiv:2010.02194."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1207\/s15516709cog1402_1"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1006"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00298"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313425"},{"key":"e_1_3_2_42_2","volume-title":"LibGuides: Literature Review: Transition Words","author":"Fresno California State University","year":"2018","unstructured":"California State University Fresno. 2018. LibGuides: Literature Review: Transition Words."},{"key":"e_1_3_2_43_2","first-page":"1","volume-title":"Forum Linguisticum","author":"Fries Peter H.","year":"1981","unstructured":"Peter H. Fries. 1981. On the status of theme in English: Arguments from discourse. In Forum Linguisticum, Vol. 6. Helmut Buske Verlag (Papers in Textlinguistics 45) Hamburg, 1\u201338."},{"key":"e_1_3_2_44_2","first-page":"758","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Ganitkevitch Juri","year":"2013","unstructured":"Juri Ganitkevitch, Benjamin Van Durme, and Chris Callison-Burch. 2013. PPDB: the paraphrase database. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 758\u2013764."},{"key":"e_1_3_2_45_2","first-page":"69","article-title":"Processes in language production","volume":"3","author":"Garrett Merrill F.","year":"1989","unstructured":"Merrill F. Garrett. 1989. Processes in language production. Ling.: Cambr. Surv. 3 (1989), 69\u201396.","journal-title":"Ling.: Cambr. Surv."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.5555\/3305381.3305510"},{"key":"e_1_3_2_47_2","unstructured":"Yoav Goldberg. 2019. Assessing BERT\u2019s syntactic abilities. arxiv:1901.05287."},{"key":"e_1_3_2_48_2","first-page":"347","article-title":"Learning task-dependent distributed representations by backpropagation through structure","volume":"1","author":"Goller Christoph","year":"1996","unstructured":"Christoph Goller and Andreas Kuchler. 1996. Learning task-dependent distributed representations by backpropagation through structure. Neural Netw. 1 (1996), 347\u2013352.","journal-title":"Neural Netw."},{"key":"e_1_3_2_49_2","first-page":"290","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence","author":"Guo Guibing","year":"2018","unstructured":"Guibing Guo, Songlin Zhai, Fajie Yuan, Yuan Liu, and Xingwei Wang. 2018. VSE-ens: Visual-semantic embeddings with efficient negative sampling. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence. 290\u2013297."},{"key":"e_1_3_2_50_2","first-page":"1024","volume-title":"Proceedings of the Advances in Neural Information Processing Systems Conference","author":"Hamilton William L.","year":"2017","unstructured":"William L. Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs. In Proceedings of the Advances in Neural Information Processing Systems Conference. 1024\u20131034."},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.3301938"},{"key":"e_1_3_2_52_2","article-title":"Efficient natural language response suggestion for smart reply","volume":"1705","author":"Henderson Matthew","year":"2017","unstructured":"Matthew Henderson, Rami Al-Rfou, Brian Strope, Yun-Hsuan Sung, L\u00e1szl\u00f3 Luk\u00e1cs, Ruiqi Guo, Sanjiv Kumar, Balint Miklos, and Ray Kurzweil. 2017. Efficient natural language response suggestion for smart reply. Corr abs\/1705.00652.","journal-title":"Corr"},{"key":"e_1_3_2_53_2","first-page":"4129","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Hewitt John","year":"2019","unstructured":"John Hewitt and Christopher D. Manning. 2019. A structural probe for finding syntax in word representations. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4129\u20134138."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16-1162"},{"key":"e_1_3_2_55_2","unstructured":"Geoffrey E. Hinton. 1984. Distributed representations. In Parallel Distributed Processing: Explorations in the Microstructure of Cognition: Foundations . MIT Press 77\u2013109."},{"key":"e_1_3_2_56_2","first-page":"473","volume-title":"Proceedings of the Advances in Neural Information Processing Systems Conference","author":"Hochreiter Sepp","year":"1996","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1996. LSTM can solve hard long time lag problems. In Proceedings of the Advances in Neural Information Processing Systems Conference. 473\u2013479."},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/1014052.1014073"},{"key":"e_1_3_2_59_2","unstructured":"Thomas Huckin Joel Bloch et\u00a0al. 1993. Strategies for inferring word-meanings in context: A cognitive model. In Second Language Reading And Vocabulary Learning . Ablex Publishing Corporation 153\u2013178."},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1162"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1356"},{"key":"e_1_3_2_62_2","article-title":"Discourse-based objectives for fast unsupervised sentence representation learning","volume":"1705","author":"Jernite Yacine","year":"2017","unstructured":"Yacine Jernite, Samuel R. Bowman, and David Sontag. 2017. Discourse-based objectives for fast unsupervised sentence representation learning. Corr abs\/1705.00557 (2017).","journal-title":"Corr"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1052"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-1062"},{"key":"e_1_3_2_65_2","first-page":"240","article-title":"note on regression and inheritance in the case of two parents","volume":"58","author":"Karl Pearson","year":"1895","unstructured":"Pearson Karl. 1895. note on regression and inheritance in the case of two parents. Proc. Roy. Societ. Lond. Series I 58 (1895), 240\u2013242.","journal-title":"Proc. Roy. Societ. Lond. Series I"},{"key":"e_1_3_2_66_2","unstructured":"Taeuk Kim Jihun Choi Daniel Edmiston and Sang-goo Lee. 2020. Are pre-trained language models aware of phrases? Simple but strong baselines for grammar induction. In Proceedings of the International Conference on Learning Representations . https:\/\/openreview.net\/pdf?id=H1xPR3NtPB."},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1181"},{"key":"e_1_3_2_68_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: a method for stochastic optimization. In Proceedings of the 3rd International Conference on Learning Representations."},{"key":"e_1_3_2_69_2","volume-title":"Proceedings of the 5th International Conference on Learning Representations","author":"Kipf Thomas N.","year":"2017","unstructured":"Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In Proceedings of the 5th International Conference on Learning Representations."},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1524"},{"key":"e_1_3_2_71_2","article-title":"Unifying visual-semantic embeddings with multimodal neural language models","volume":"1411","author":"Kiros Ryan","year":"2014","unstructured":"Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel. 2014. Unifying visual-semantic embeddings with multimodal neural language models. Corr abs\/1411.2539 (2014).","journal-title":"Corr"},{"key":"e_1_3_2_72_2","first-page":"3294","volume-title":"Proceedings of the Advances in Neural Information Processing Systems Conference","author":"Kiros Ryan","year":"2015","unstructured":"Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Skip-thought vectors. In Proceedings of the Advances in Neural Information Processing Systems Conference. 3294\u20133302."},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1007"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-2012"},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/146370.146380"},{"key":"e_1_3_2_76_2","article-title":"Electrophysiological correlates of processing causal relationships between sentences","author":"Kuperberg G. R.","year":"2011","unstructured":"G. R. Kuperberg, D. Caplan, M. Eddy, J. Cotton, and P. J. Holcomb. 2011. Electrophysiological correlates of processing causal relationships between sentences. J. Cogn. Neurosci. 23, 5 (2011), 1230\u20131246.","journal-title":"J. Cogn. Neurosci."},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.5555\/2886521.2886636"},{"key":"e_1_3_2_78_2","unstructured":"Zhenzhong Lan Mingda Chen Sebastian Goodman Kevin Gimpel Piyush Sharma and Radu Soricut. 2019. ALBERT: a lite BERT for self-supervised learning of language representations. arxiv:1909.11942."},{"key":"e_1_3_2_79_2","first-page":"1188","volume-title":"Proceedings of the 31st International Conference on Machine Learning","author":"Le Quoc V.","year":"2014","unstructured":"Quoc V. Le and Tomas Mikolov. 2014. Distributed representations of sentences and documents. In Proceedings of the 31st International Conference on Machine Learning. 1188\u20131196."},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.733"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1278"},{"key":"e_1_3_2_82_2","doi-asserted-by":"publisher","DOI":"10.3115\/1072228.1072378"},{"key":"e_1_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00640"},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_85_2","article-title":"A structured self-attentive sentence embedding","volume":"1703","author":"Lin Zhouhan","year":"2017","unstructured":"Zhouhan Lin, Minwei Feng, C\u00edcero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017. A structured self-attentive sentence embedding. Corr abs\/1703.03130 (2017).","journal-title":"Corr"},{"key":"e_1_3_2_86_2","first-page":"4881","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence, the 30th Innovative Applications of Artificial Intelligence, and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence","author":"Liu Tianyu","year":"2018","unstructured":"Tianyu Liu, Kexiang Wang, Lei Sha, Baobao Chang, and Zhifang Sui. 2018. Table-to-text generation by structure-aware Seq2seq learning. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence, the 30th Innovative Applications of Artificial Intelligence, and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence. 4881\u20134888."},{"key":"e_1_3_2_87_2","article-title":"RoBERTa: a robustly optimized BERT pretraining approach","volume":"1907","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: a robustly optimized BERT pretraining approach. Corr abs\/1907.11692.","journal-title":"Corr"},{"key":"e_1_3_2_88_2","article-title":"Learning natural language inference using bidirectional LSTM model and inner-attention","volume":"1605","author":"Liu Yang","year":"2016","unstructured":"Yang Liu, Chengjie Sun, Lei Lin, and Xiaolong Wang. 2016. Learning natural language inference using bidirectional LSTM model and inner-attention. Corr abs\/1605.09090.","journal-title":"Corr"},{"key":"e_1_3_2_89_2","volume-title":"Proceedings of the 6th International Conference on Learning Representations Conference","author":"Logeswaran Lajanugen","year":"2018","unstructured":"Lajanugen Logeswaran and Honglak Lee. 2018. An efficient framework for learning sentence representations. In Proceedings of the 6th International Conference on Learning Representations Conference."},{"key":"e_1_3_2_90_2","first-page":"216","volume-title":"Proceedings of the 9th International Conference on Language Resources and Evaluation","author":"Marelli Marco","year":"2014","unstructured":"Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014. A SICK cure for the evaluation of compositional distributional semantic models. In Proceedings of the 9th International Conference on Language Resources and Evaluation. 216\u2013223."},{"key":"e_1_3_2_91_2","first-page":"153","article-title":"Essai d\u2019une recherche statistique sur le texte du roman \u201cEugene Onegin\u201d illustrant la liaison des epreuve en chain (\u2018Example of a statistical investigation of the text of \u2018Eugene Onegin\u2019 illustrating the dependence between samples in chain\u201d)","volume":"7","author":"Markov A. A.","year":"1913","unstructured":"A. A. Markov. 1913. Essai d\u2019une recherche statistique sur le texte du roman \u201cEugene Onegin\u201d illustrant la liaison des epreuve en chain (\u2018Example of a statistical investigation of the text of \u2018Eugene Onegin\u2019 illustrating the dependence between samples in chain\u201d). Izvistia Imperatorskoi Akademii Nauk (Bulletin De l\u2019acad\u00e9mie imp\u00e9riale Des Sciences De st.-p\u00e9tersbourg) 7 (1913), 153\u2013162. English translation by Morris Halle, 1956.","journal-title":"Izvistia Imperatorskoi Akademii Nauk (Bulletin De l\u2019acad\u00e9mie imp\u00e9riale Des Sciences De st.-p\u00e9tersbourg)"},{"key":"e_1_3_2_92_2","doi-asserted-by":"publisher","DOI":"10.5220\/0006595904860492"},{"key":"e_1_3_2_93_2","volume-title":"Proceedings of the 1st International Conference on Learning Representations","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In Proceedings of the 1st International Conference on Learning Representations."},{"key":"e_1_3_2_94_2","unstructured":"Tom\u00e1s Mikolov Quoc V. Le and Ilya Sutskever. 2013. Exploiting similarities among languages for machine translation. Retrieved from http:\/\/arxiv.org\/abs\/1309.4168."},{"key":"e_1_3_2_95_2","first-page":"3111","volume-title":"Proceedings of the Advances in Neural Information Processing Systems Conference","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In Proceedings of the Advances in Neural Information Processing Systems Conference. 3111\u20133119."},{"key":"e_1_3_2_96_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.576"},{"key":"e_1_3_2_97_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00063"},{"key":"e_1_3_2_98_2","doi-asserted-by":"publisher","DOI":"10.1017\/S135132499800182X"},{"key":"e_1_3_2_99_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-2060"},{"key":"e_1_3_2_100_2","article-title":"DisSent: Sentence representation learning from explicit discourse relations","volume":"1710","author":"Nie Allen","year":"2017","unstructured":"Allen Nie, Erin D. Bennett, and Noah D. Goodman. 2017. DisSent: Sentence representation learning from explicit discourse relations. Corr abs\/1710.04334.","journal-title":"Corr"},{"key":"e_1_3_2_101_2","doi-asserted-by":"publisher","DOI":"10.1162\/jocn.2006.18.7.1098"},{"key":"e_1_3_2_102_2","volume-title":"Proceedings of the 42nd European Conference on IR Research: Advances in Information Retrieval","author":"Nikolaev Fedor","unstructured":"Fedor Nikolaev and Alexander Kotov. [n.d.]. Joint word and entity embeddings for entity retrieval from a knowledge graph. In Proceedings of the 42nd European Conference on IR Research: Advances in Information Retrieval."},{"key":"e_1_3_2_103_2","unstructured":"Martin Paczynski Tali Ditman Kana Okano and Gina R. Kuperberg. 2007. Drawing inferences during discourse comprehension: An ERP study. Retrieved on 10 Nov. 2021 from https:\/\/citeseerx.ist.psu.edu\/viewdoc\/download?doi=10.1.1.505.3466&rep=rep1&type=pdf."},{"key":"e_1_3_2_104_2","doi-asserted-by":"publisher","DOI":"10.3115\/1218955.1218990"},{"key":"e_1_3_2_105_2","doi-asserted-by":"publisher","DOI":"10.3115\/1219840.1219855"},{"key":"e_1_3_2_106_2","volume-title":"Proceedings of the 1st Conference on Argumentation","author":"Peldszus Andreas","year":"2015","unstructured":"Andreas Peldszus and Manfred Stede. 2015. An annotated corpus of argumentative microtexts. In Proceedings of the 1st Conference on Argumentation."},{"key":"e_1_3_2_107_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_108_2","article-title":"Evaluation of sentence embeddings in downstream and linguistic probing tasks","volume":"1806","author":"Perone Christian S.","year":"2018","unstructured":"Christian S. Perone, Roberto Silveira, and Thomas S. Paula. 2018. Evaluation of sentence embeddings in downstream and linguistic probing tasks. Corr abs\/1806.06259.","journal-title":"Corr"},{"key":"e_1_3_2_109_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1202"},{"key":"e_1_3_2_110_2","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(90)90005-K"},{"key":"e_1_3_2_111_2","doi-asserted-by":"publisher","DOI":"10.1145\/2036264.2036277"},{"key":"e_1_3_2_112_2","unstructured":"Xipeng Qiu Tianxiang Sun Yige Xu Yunfan Shao Ning Dai and Xuanjing Huang. 2020. Pre-trained models for natural language processing: a survey. arxiv:2003.08271."},{"key":"e_1_3_2_113_2","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever. 2018. Improving language understanding by generative pre-training. (2018). https:\/\/www.cs.ubc.ca\/amuham01\/LING530\/papers\/radford2018improving.pdf."},{"key":"e_1_3_2_114_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1264"},{"key":"e_1_3_2_115_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_2_116_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.365"},{"key":"e_1_3_2_117_2","doi-asserted-by":"publisher","DOI":"10.2307\/1771893"},{"key":"e_1_3_2_118_2","doi-asserted-by":"publisher","DOI":"10.1017\/S0272263199004039"},{"key":"e_1_3_2_119_2","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3269277"},{"key":"e_1_3_2_120_2","article-title":"Concatenated  \\( p \\) -mean word embeddings as universal cross-lingual sentence representations","volume":"1803","author":"R\u00fcckl\u00e9 Andreas","year":"2018","unstructured":"Andreas R\u00fcckl\u00e9, Steffen Eger, Maxime Peyrard, and Iryna Gurevych. 2018. Concatenated \\( p \\) -mean word embeddings as universal cross-lingual sentence representations. Corr abs\/1803.01400 (2018).","journal-title":"Corr"},{"key":"e_1_3_2_121_2","volume-title":"Artificial Intelligence: a Modern Approach (4th Edition)","author":"Russell Stuart J.","year":"2020","unstructured":"Stuart J. Russell and Peter Norvig. 2020. Artificial Intelligence: a Modern Approach (4th Edition). Pearson."},{"key":"e_1_3_2_122_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2012.6289079"},{"key":"e_1_3_2_123_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-2037"},{"key":"e_1_3_2_124_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-2619"},{"key":"e_1_3_2_125_2","volume-title":"Proceedings of the 11th International Conference on Language Resources and Evaluation","author":"Schwenk Holger","year":"2018","unstructured":"Holger Schwenk and Xian Li. 2018. A corpus for multilingual document classification in eight languages. In Proceedings of the 11th International Conference on Language Resources and Evaluation."},{"key":"e_1_3_2_126_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1099"},{"key":"e_1_3_2_127_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1162"},{"key":"e_1_3_2_128_2","first-page":"3715","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics","author":"Shi Haoyue","year":"2018","unstructured":"Haoyue Shi, Jiayuan Mao, Tete Xiao, Yuning Jiang, and Jian Sun. 2018. Learning visually-grounded semantics from contrastive adversarial samples. In Proceedings of the 27th International Conference on Computational Linguistics. 3715\u20133727."},{"key":"e_1_3_2_129_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1351"},{"key":"e_1_3_2_130_2","first-page":"801","volume-title":"Proceedings of the Advances in Neural Information Processing Systems Conference","author":"Socher Richard","year":"2011","unstructured":"Richard Socher, Eric H. Huang, Jeffrey Pennington, Andrew Y. Ng, and Christopher D. Manning. 2011. Dynamic pooling and unfolding recursive autoencoders for paraphrase detection. In Proceedings of the Advances in Neural Information Processing Systems Conference. 801\u2013809."},{"key":"e_1_3_2_131_2","first-page":"1631","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Socher Richard","year":"2013","unstructured":"Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. 1631\u20131642."},{"key":"e_1_3_2_132_2","doi-asserted-by":"crossref","first-page":"657","DOI":"10.1075\/slcs.7.36sor","volume-title":"Possibilities and Limitations of Pragmatics","author":"Sorensen Viggo","year":"1981","unstructured":"Viggo Sorensen. 1981. Coherence as a pragmatic concept. In Possibilities and Limitations of Pragmatics. John Benjamins, 657."},{"key":"e_1_3_2_133_2","doi-asserted-by":"publisher","DOI":"10.1162\/COLI_a_00295"},{"key":"e_1_3_2_134_2","volume-title":"Proceedings of the 6th International Conference on Learning Representations","author":"Subramanian Sandeep","year":"2018","unstructured":"Sandeep Subramanian, Adam Trischler, Yoshua Bengio, and Christopher J. Pal. 2018. Learning general purpose distributed sentence representations via large scale multi-task learning. In Proceedings of the 6th International Conference on Learning Representations."},{"key":"e_1_3_2_135_2","first-page":"3104","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems","author":"Sutskever Ilya","year":"2014","unstructured":"Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In Proceedings of the Annual Conference on Neural Information Processing Systems. 3104\u20133112."},{"key":"e_1_3_2_136_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-20161-5_37"},{"key":"e_1_3_2_137_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1150"},{"key":"e_1_3_2_138_2","unstructured":"Yi Tay Vinh Q. Tran Sebastian Ruder Jai Prakash Gupta Hyung Won Chung Dara Bahri Zhen Qin Simon Baumgartner Cong Yu and Donald Metzler. 2021. Charformer: Fast character transformers via gradient-based subword tokenization. Retrieved from https:\/\/arxiv.org\/abs\/2106.12672."},{"key":"e_1_3_2_139_2","volume-title":"Discourse and Dialogue","author":"Dijk Teun A. Van","year":"1985","unstructured":"Teun A. Van Dijk. 1985. Handbook of discourse analysis. In Discourse and Dialogue. Citeseer."},{"key":"e_1_3_2_140_2","first-page":"5998","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Annual Conference on Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_2_141_2","volume-title":"Proceedings of the 6th International Conference on Learning Representations","author":"Velickovic Petar","year":"2018","unstructured":"Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Li\u00f2, and Yoshua Bengio. 2018. Graph attention networks. In Proceedings of the 6th International Conference on Learning Representations."},{"key":"e_1_3_2_142_2","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390294"},{"key":"e_1_3_2_143_2","first-page":"2773","volume-title":"Proceedings of the Advances in Neural Information Processing Systems Conference","author":"Vinyals Oriol","year":"2015","unstructured":"Oriol Vinyals, Lukasz Kaiser, Terry Koo, Slav Petrov, Ilya Sutskever, and Geoffrey E. Hinton. 2015. Grammar as a foreign language. In Proceedings of the Advances in Neural Information Processing Systems Conference. 2773\u20132781."},{"key":"e_1_3_2_144_2","first-page":"56","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics","author":"Vulic Ivan","year":"2017","unstructured":"Ivan Vulic, Nikola Mrksic, Roi Reichart, Diarmuid \u00d3. S\u00e9aghdha, Steve J. Young, and Anna Korhonen. 2017. Morph-fitting: Fine-tuning word vector spaces with simple language-specific rules. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics. 56\u201368."},{"key":"e_1_3_2_145_2","volume-title":"Proceedings of the 7th International Conference on Learning Representations","author":"Wang Alex","year":"2019","unstructured":"Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 7th International Conference on Learning Representations."},{"key":"e_1_3_2_146_2","doi-asserted-by":"publisher","DOI":"10.1145\/2983323.2983755"},{"key":"e_1_3_2_147_2","volume-title":"Explorations in Applied Linguistics","author":"Widdowson Henry George","year":"1979","unstructured":"Henry George Widdowson. 1979. Explorations in Applied Linguistics. Vol. 1. Oxford University Press."},{"key":"e_1_3_2_148_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-005-7880-9"},{"key":"e_1_3_2_149_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00143"},{"key":"e_1_3_2_150_2","volume-title":"Proceedings of the 4th International Conference on Learning Representations","author":"Wieting John","year":"2016","unstructured":"John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2016. Towards universal paraphrastic sentence embeddings. In Proceedings of the 4th International Conference on Learning Representations."},{"key":"e_1_3_2_151_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1190"},{"key":"e_1_3_2_152_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1101"},{"key":"e_1_3_2_153_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1408"},{"key":"e_1_3_2_154_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S15-2001"},{"key":"e_1_3_2_155_2","unstructured":"Linting Xue Aditya Barua Noah Constant Rami Al-Rfou Sharan Narang Mihir Kale Adam Roberts and Colin Raffel. 2021. ByT5: Towards a token-free future with pre-trained byte-to-byte models. (2021). arXiv 2105.13626 ."},{"key":"e_1_3_2_156_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-demos.12"},{"key":"e_1_3_2_157_2","first-page":"5754","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems","author":"Yang Zhilin","year":"2019","unstructured":"Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019. XLNet: Generalized autoregressive pretraining for language understanding. In Proceedings of the Annual Conference on Neural Information Processing Systems. 5754\u20135764."},{"key":"e_1_3_2_158_2","volume-title":"Discourse Analysis","author":"Yule George","year":"1986","unstructured":"George Yule and Gillian R. Brown. 1986. Discourse Analysis. Cambridge University Press."},{"issue":"1","key":"e_1_3_2_159_2","first-page":"56","article-title":"A comparison of inferencing and meaning-guessing of new lexicon in context versus non-context vocabulary presentation","volume":"9","author":"Zaid M. A.","year":"2009","unstructured":"M. A. Zaid. 2009. A comparison of inferencing and meaning-guessing of new lexicon in context versus non-context vocabulary presentation. Read. Matrix: Int. Online J. 9, 1 (2009), 56\u201366.","journal-title":"Read. Matrix: Int. Online J."},{"key":"e_1_3_2_160_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1009"},{"key":"e_1_3_2_161_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.629"},{"key":"e_1_3_2_162_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1481"},{"key":"e_1_3_2_163_2","first-page":"1298","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies)","author":"Zhang Yuan","year":"2019","unstructured":"Yuan Zhang, Jason Baldridge, and Luheng He. 2019. PAWS: Paraphrase adversaries from word scrambling. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies). 1298\u20131308."},{"key":"e_1_3_2_164_2","doi-asserted-by":"publisher","DOI":"10.5555\/2832747.2832816"},{"key":"e_1_3_2_165_2","first-page":"19","volume-title":"Proceedings of IEEE International Conference on Computer Vision","author":"Zhu Yukun","year":"2015","unstructured":"Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of IEEE International Conference on Computer Vision. 19\u201327."}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3482853","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3482853","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:48:45Z","timestamp":1750193325000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3482853"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,26]]},"references-count":164,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,3,31]]}},"alternative-id":["10.1145\/3482853"],"URL":"https:\/\/doi.org\/10.1145\/3482853","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,3,26]]},"assertion":[{"value":"2019-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-08-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-03-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}