{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T21:47:43Z","timestamp":1771710463802,"version":"3.50.1"},"reference-count":104,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2020,10,30]],"date-time":"2020-10-30T00:00:00Z","timestamp":1604016000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"University of Rijeka","award":["uniri-drustv-18-20"],"award-info":[{"award-number":["uniri-drustv-18-20"]}]},{"name":"University of Rijeka","award":["uniri-drustv-18-38"],"award-info":[{"award-number":["uniri-drustv-18-38"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>In natural language processing, text needs to be transformed into a machine-readable representation before any processing. The quality of further natural language processing tasks greatly depends on the quality of those representations. In this survey, we systematize and analyze 50 neural models from the last decade. The models described are grouped by the architecture of neural networks as shallow, recurrent, recursive, convolutional, and attention models. Furthermore, we categorize these models by representation level, input level, model type, and model supervision. We focus on task-independent representation models, discuss their advantages and drawbacks, and subsequently identify the promising directions for future neural text representation models. We describe the evaluation datasets and tasks used in the papers that introduced the models and compare the models based on relevant evaluations. The quality of a representation model can be evaluated as its capability to generalize to multiple unrelated tasks. Benchmark standardization is visible amongst recent models and the number of different tasks models are evaluated on is increasing.<\/jats:p>","DOI":"10.3390\/info11110511","type":"journal-article","created":{"date-parts":[[2020,10,30]],"date-time":"2020-10-30T09:29:32Z","timestamp":1604050172000},"page":"511","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":35,"title":["Survey of Neural Text Representation Models"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6343-0938","authenticated-orcid":false,"given":"Karlo","family":"Babi\u0107","sequence":"first","affiliation":[{"name":"Center for Artificial Intelligence and Cybersecurity and Department of Informatics, University of Rijeka, 51000 Rijeka, Croatia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1900-5333","authenticated-orcid":false,"given":"Sanda","family":"Martin\u010di\u0107-Ip\u0161i\u0107","sequence":"additional","affiliation":[{"name":"Center for Artificial Intelligence and Cybersecurity and Department of Informatics, University of Rijeka, 51000 Rijeka, Croatia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9513-9467","authenticated-orcid":false,"given":"Ana","family":"Me\u0161trovi\u0107","sequence":"additional","affiliation":[{"name":"Center for Artificial Intelligence and Cybersecurity and Department of Informatics, University of Rijeka, 51000 Rijeka, Croatia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,10,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Manning, C.D., Raghavan, P., and Sch\u00fctze, H. (2008). Introduction to Information Retrieval, Cambridge University Press.","DOI":"10.1017\/CBO9780511809071"},{"key":"ref_2","first-page":"1","article-title":"Neural network methods for natural language processing","volume":"10","author":"Goldberg","year":"2017","journal-title":"Synth. Lect. Hum. Lang. Technol."},{"key":"ref_3","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford","year":"2019","journal-title":"OpenAI Blog"},{"key":"ref_4","unstructured":"Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep Learning, MIT Press."},{"key":"ref_5","unstructured":"Sutskever, I., Vinyals, O., and Le, Q.V. (2020, October 29). Sequence to Sequence Learning with Neural Networks. Available online: https:\/\/papers.nips.cc\/paper\/5346-sequence-to-sequence-learning-with-neural-networks.pdf."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"23253","DOI":"10.1109\/ACCESS.2017.2776930","article-title":"Deep Convolution Neural Networks for Twitter Sentiment Analysis","volume":"6","author":"Jianqiang","year":"2018","journal-title":"IEEE Access"},{"key":"ref_7","unstructured":"Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M.W. (2020). Realm: Retrieval-augmented language model pre-training. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Hailu, T.T., Yu, J., and Fantaye, T.G. (2020). A Framework for Word Embedding Based Automatic Text Summarization and Evaluation. Information, 11.","DOI":"10.3390\/info11020078"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Bodrunova, S.S., Orekhov, A.V., Blekanov, I.S., Lyudkevich, N.S., and Tarasov, N.A. (2020). Topic Detection Based on Sentence Embeddings and Agglomerative Clustering with Markov Moment. Future Internet, 12.","DOI":"10.3390\/fi12090144"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Martin\u010di\u0107-Ip\u0161i\u0107, S., Mili\u010di\u0107, T., and Todorovski, L. (2019). The Influence of Feature Representation of Text on the Performance of Document Classification. Appl. Sci., 9.","DOI":"10.3390\/app9040743"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1109\/MCI.2018.2840738","article-title":"Recent trends in deep learning based natural language processing","volume":"13","author":"Young","year":"2018","journal-title":"IEEE Comput. Intell. Mag."},{"key":"ref_12","unstructured":"Jing, K., Xu, J., and He, B. (2019). A Survey on Neural Network Language Models. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"743","DOI":"10.1613\/jair.1.11259","article-title":"From word to sense embeddings: A survey on vector representations of meaning","volume":"63","author":"Pilehvar","year":"2018","journal-title":"J. Artif. Intell. Res."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"569","DOI":"10.1613\/jair.1.11640","article-title":"A survey of cross-lingual word embedding models","volume":"65","author":"Ruder","year":"2019","journal-title":"J. Artif. Intell. Res."},{"key":"ref_15","unstructured":"A\u00dfenmacher, M., and Heumann, C. (2020). On the comparability of Pre-trained Language Models. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Finkelstein, L., Gabrilovich, E., Matias, Y., Rivlin, E., Solan, Z., Wolfman, G., and Ruppin, E. (2001, January 1\u20135). Placing search in context: The concept revisited. Proceedings of the 10th International Conference on World Wide Web, Hong Kong, China.","DOI":"10.1145\/371920.372094"},{"key":"ref_17","unstructured":"Marelli, M., Menini, S., Baroni, M., Bentivogli, L., Bernardi, R., and Zamparelli, R. (2014, January 26\u201331). A SICK Cure for the Evaluation of Compositional Distributional Semantic Models. Proceedings of the 9th Language Resources and Evaluation Conference, Reykjavik, Iceland. Available online: http:\/\/www.lrec-conf.org\/proceedings\/lrec2014\/pdf\/363_Paper.pdf."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Bowman, S.R., Angeli, G., Potts, C., and Manning, C.D. (2015, January 17\u201321). A large annotated corpus for learning natural language inference. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, Lisbon, Portugal.","DOI":"10.18653\/v1\/D15-1075"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Williams, A., Nangia, N., and Bowman, S. (2018, January 1\u20136). A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1 (Long Papers)), New Orleans, LA, USA.","DOI":"10.18653\/v1\/N18-1101"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Dolan, B., Quirk, C., and Brockett, C. (2004, January 23\u201327). Unsupervised construction of large paraphrase corpora: Exploiting massively parallel news sources. Proceedings of the 20th international conference on Computational Linguistics, Geneva, Switzerland.","DOI":"10.3115\/1220355.1220406"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Agirre, E., Banea, C., Cardie, C., Cer, D., Diab, M., Gonzalez-Agirre, A., Guo, W., Mihalcea, R., Rigau, G., and Wiebe, J. (2014, January 23\u201324). Semeval-2014 task 10: Multilingual semantic textual similarity. Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014), Dublin, Ireland.","DOI":"10.3115\/v1\/S14-2010"},{"key":"ref_22","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv."},{"key":"ref_23","unstructured":"Mikolov, T., Yih, W.T., and Zweig, G. (2013, January 9\u201314). Linguistic regularities in continuous space word representations. Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Atlanta, GA, USA."},{"key":"ref_24","unstructured":"Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C.D., Ng, A., and Potts, C. (2013, January 18\u201321). Recursive deep models for semantic compositionality over a sentiment treebank. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Seattle, WA, USA."},{"key":"ref_25","unstructured":"Maas, A.L., Daly, R.E., Pham, P.T., Huang, D., Ng, A.Y., and Potts, C. (2011, January 19\u201324). Learning word vectors for sentiment analysis. Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1, Portland, OR, USA."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Pang, B., and Lee, L. (2005, January 17). Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics. Association for Computational Linguistics, Ann Arbor, MI, USA.","DOI":"10.3115\/1219840.1219855"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Hu, M., and Liu, B. (2004, January 22). Mining and summarizing customer reviews. Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Seattle, WA, USA.","DOI":"10.1145\/1014052.1014073"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Rosenthal, S., Farra, N., and Nakov, P. (2017, January 3\u20134). SemEval-2017 Task 4: Sentiment Analysis in Twitter. Proceedings of the SemEval \u201917 11th International Workshop on Semantic Evaluation, Vancouver, BC, Canada.","DOI":"10.18653\/v1\/S17-2088"},{"key":"ref_29","unstructured":"Voorhees, E.M., and Harman, D. (2002, January 19\u201322). Overview of TREC 2002. Proceedings of the Eleventh Text REtrieval Conference, Gaithersburg, MD, USA. Available online: https:\/\/trec.nist.gov\/pubs\/trec11\/papers\/OVERVIEW.11.pdf."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. (2016). Squad: 100,000+ questions for machine comprehension of text. arXiv.","DOI":"10.18653\/v1\/D16-1264"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1007\/s10579-005-7880-9","article-title":"Annotating expressions of opinions and emotions in language","volume":"39","author":"Wiebe","year":"2005","journal-title":"Lang. Resour. Eval."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"453","DOI":"10.1162\/tacl_a_00276","article-title":"Natural questions: A benchmark for question answering research","volume":"7","author":"Kwiatkowski","year":"2019","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_33","unstructured":"Berant, J., Chou, A., Frostig, R., and Liang, P. (2013, January 18\u201321). Semantic parsing on freebase from question-answer pairs. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Seattle, WA, USA."},{"key":"ref_34","first-page":"313","article-title":"Building a large annotated corpus of English: The Penn Treebank","volume":"19","author":"Marcus","year":"1993","journal-title":"Comput. Linguist."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wang, J., Wang, Z., Zhang, D., and Yan, J. (2017, January 19\u201325). Combining Knowledge with Deep Convolutional Neural Networks for Short Text Classification. Proceedings of the International Joint Conference on Artificial Intelligence, Melbourne, Australia.","DOI":"10.24963\/ijcai.2017\/406"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S.R. (2018). Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv.","DOI":"10.18653\/v1\/W18-5446"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E. (2017, January 9\u201311). RACE: Large-scale ReAding Comprehension Dataset From Examinations. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Copenhagen, Denmark.","DOI":"10.18653\/v1\/D17-1082"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Pang, B., and Lee, L. (2004, January 21\u201326). A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. Proceedings of the 42nd annual meeting on Association for Computational Linguistics. Association for Computational Linguistics, Barcelona, Spain.","DOI":"10.3115\/1218955.1218990"},{"key":"ref_39","unstructured":"Jelinek, F. (1998). Statistical Methods for Speech Recognition, MIT Press."},{"key":"ref_40","unstructured":"Huang, E.H., Socher, R., Manning, C.D., and Ng, A.Y. (2012, January 8\u201314). Improving word representations via global context and multiple word prototypes. Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Long Papers-Volume 1, Jeju Island, Korea."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Peters, M.E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018). Deep contextualized word representations. arXiv.","DOI":"10.18653\/v1\/N18-1202"},{"key":"ref_42","unstructured":"Akbik, A., Blythe, D., and Vollgraf, R. (2018, January 21\u201325). Contextual string embeddings for sequence labeling. Proceedings of the 27th International Conference on Computational Linguistics, Santa Fe, NM, USA."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"112","DOI":"10.1016\/j.inffus.2019.06.009","article-title":"Convolution\u2013deconvolution word embedding: An end-to-end multi-prototype fusion embedding method for natural language processing","volume":"53","author":"Shuang","year":"2020","journal-title":"Inf. Fusion"},{"key":"ref_44","unstructured":"Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2020, October 29). Improving Language Understanding by Generative Pre-Training. Available online: https:\/\/s3-us-west-2.amazonaws.com\/openai-assets\/research-covers\/language-unsupervised\/language_understanding_paper.pdf."},{"key":"ref_45","unstructured":"Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R.R., and Le, Q.V. (2019). Xlnet: Generalized autoregressive pretraining for language understanding. Advances in Neural Information Processing Systems, MIT Press."},{"key":"ref_46","unstructured":"Botha, J., and Blunsom, P. (2014, January 21\u201326). Compositional morphology for word representations and language modelling. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_47","unstructured":"Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., and Dean, J. (2020, October 29). Distributed Representations of Words and Phrases and Their Compositionality. Available online: https:\/\/papers.nips.cc\/paper\/5021-distributed-representations-of-words-and-phrases-and-their-compositionality.pdf."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Kim, Y., Jernite, Y., Sontag, D., and Rush, A.M. (2016, January 12\u201317). Character-aware neural language models. Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA.","DOI":"10.1609\/aaai.v30i1.10362"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1162\/tacl_a_00051","article-title":"Enriching word vectors with subword information","volume":"5","author":"Bojanowski","year":"2017","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_50","first-page":"23","article-title":"A new algorithm for data compression","volume":"12","author":"Gage","year":"1994","journal-title":"C Users J."},{"key":"ref_51","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2020, October 29). Attention Is All You Need. Available online: https:\/\/papers.nips.cc\/paper\/7181-attention-is-all-you-need.pdf."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Kudo, T., and Richardson, J. (2018). Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv.","DOI":"10.18653\/v1\/D18-2012"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Schuster, M., and Nakajima, K. (2012, January 25\u201330). Japanese and korean voice search. Proceedings of the 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Kyoto, Japan.","DOI":"10.1109\/ICASSP.2012.6289079"},{"key":"ref_54","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1613\/jair.4992","article-title":"A primer on neural network models for natural language processing","volume":"57","author":"Goldberg","year":"2016","journal-title":"J. Artif. Intell. Res."},{"key":"ref_56","unstructured":"Shi, T., and Liu, Z. (2014). Linking GloVe with word2vec. arXiv."},{"key":"ref_57","unstructured":"Levy, O., and Goldberg, Y. (2014). Neural word embedding as implicit matrix factorization. Advances in Neural Information Processing Systems, Mit Press."},{"key":"ref_58","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1080\/00437956.1954.11659520","article-title":"Distributional structure","volume":"10","author":"Harris","year":"1954","journal-title":"Word"},{"key":"ref_59","unstructured":"Palmer, F. (1957). A Synopsis of Linguistic Theory 1930\u20131955. Studies in Linguistic Analysis, Philological Society."},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Levy, O., and Goldberg, Y. (2014, January 22\u201327). Dependency-based word embeddings. Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Baltimore, MD, USA.","DOI":"10.3115\/v1\/P14-2050"},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C. (2014, January 25\u201329). Glove: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_62","unstructured":"Le, Q., and Mikolov, T. (2014, January 21\u201326). Distributed representations of sentences and documents. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Cho, K., Van Merri\u00ebnboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_64","unstructured":"Kiros, R., Zhu, Y., Salakhutdinov, R.R., Zemel, R., Urtasun, R., Torralba, A., and Fidler, S. (2015). Skip-thought vectors. Advances in Neural Information Processing Systems, Mit Press."},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Lai, S., Xu, L., Liu, K., and Zhao, J. (2015, January 25\u201330). Recurrent convolutional neural networks for text classification. Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, Austin, TX, USA.","DOI":"10.1609\/aaai.v29i1.9513"},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Hill, F., Cho, K., and Korhonen, A. (2016). Learning distributed representations of sentences from unlabelled data. arXiv.","DOI":"10.18653\/v1\/N16-1162"},{"key":"ref_67","unstructured":"McCann, B., Bradbury, J., Xiong, C., and Socher, R. (2017). Learned in translation: Contextualized word vectors. Advances in Neural Information Processing Systems, Mit Press."},{"key":"ref_68","unstructured":"Subramanian, S., Trischler, A., Bengio, Y., and Pal, C.J. (2018). Learning general purpose distributed sentence representations via large scale multi-task learning. arXiv."},{"key":"ref_69","doi-asserted-by":"crossref","first-page":"597","DOI":"10.1162\/tacl_a_00288","article-title":"Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond","volume":"7","author":"Artetxe","year":"2019","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_70","unstructured":"Auli, M., Galley, M., Quirk, C., and Zweig, G. (2013, January 18\u201321). Joint language and translation modeling with recurrent neural networks. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Seattle, WA, USA."},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Conneau, A., Lample, G., Rinott, R., Williams, A., Bowman, S.R., Schwenk, H., and Stoyanov, V. (2018). XNLI: Evaluating cross-lingual sentence representations. arXiv.","DOI":"10.18653\/v1\/D18-1269"},{"key":"ref_72","unstructured":"Socher, R., Pennington, J., Huang, E.H., Ng, A.Y., and Manning, C.D. (2011, January 27\u201331). Semi-supervised recursive autoencoders for predicting sentiment distributions. Proceedings of the Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Edinburgh, Scotland, UK."},{"key":"ref_73","unstructured":"Socher, R., Huval, B., Manning, C.D., and Ng, A.Y. (2012, January 12\u201314). Semantic compositionality through recursive matrix-vector spaces. Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning. Association for Computational Linguistics, Jeju Island, Korea."},{"key":"ref_74","unstructured":"Zhao, H., Lu, Z., and Poupart, P. (2015, January 25\u201331). Self-adaptive hierarchical sentence model. Proceedings of the Twenty-Fourth International Joint Conference on Artificial Intelligence, Buenos Aires, Argentina."},{"key":"ref_75","doi-asserted-by":"crossref","unstructured":"Tai, K.S., Socher, R., and Manning, C.D. (2015). Improved semantic representations from tree-structured long short-term memory networks. arXiv.","DOI":"10.3115\/v1\/P15-1150"},{"key":"ref_76","unstructured":"Yogatama, D., Blunsom, P., Dyer, C., Grefenstette, E., and Ling, W. (2016). Learning to compose words into sentences with reinforcement learning. arXiv."},{"key":"ref_77","doi-asserted-by":"crossref","unstructured":"Choi, J., Yoo, K.M., and Lee, S.G. (2018, January 2\u20137). Learning to compose task-specific tree structures. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11975"},{"key":"ref_78","doi-asserted-by":"crossref","unstructured":"Drozdov, A., Verga, P., Yadav, M., Iyyer, M., and McCallum, A. (2019). Unsupervised latent tree induction with deep inside-outside recursive autoencoders. arXiv.","DOI":"10.18653\/v1\/N19-1116"},{"key":"ref_79","unstructured":"Nakagawa, T., Inui, K., and Kurohashi, S. (2010, January 2\u20134). Dependency tree-based sentiment classification using CRFs with hidden variables. Proceedings of the Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, Los Angeles, CA, USA."},{"key":"ref_80","doi-asserted-by":"crossref","unstructured":"Bowman, S.R., Gauthier, J., Rastogi, A., Gupta, R., Manning, C.D., and Potts, C. (2016). A fast unified model for parsing and sentence understanding. arXiv.","DOI":"10.18653\/v1\/P16-1139"},{"key":"ref_81","doi-asserted-by":"crossref","unstructured":"Collobert, R., and Weston, J. (2008, January 5\u20139). A unified architecture for natural language processing: Deep neural networks with multitask learning. Proceedings of the 25th International Conference on Machine Learning, Helsinki, Finland.","DOI":"10.1145\/1390156.1390177"},{"key":"ref_82","doi-asserted-by":"crossref","unstructured":"Kalchbrenner, N., Grefenstette, E., and Blunsom, P. (2014). A convolutional neural network for modelling sentences. arXiv.","DOI":"10.3115\/v1\/P14-1062"},{"key":"ref_83","unstructured":"Kalchbrenner, N., Espeholt, L., Simonyan, K., Oord, A.v.d., Graves, A., and Kavukcuoglu, K. (2016). Neural machine translation in linear time. arXiv."},{"key":"ref_84","doi-asserted-by":"crossref","unstructured":"Gan, Z., Pu, Y., Henao, R., Li, C., He, X., and Carin, L. (2016). Learning generic sentence representations using convolutional neural networks. arXiv.","DOI":"10.18653\/v1\/D17-1254"},{"key":"ref_85","unstructured":"Zaremba, W., Sutskever, I., and Vinyals, O. (2014). Recurrent neural network regularization. arXiv."},{"key":"ref_86","doi-asserted-by":"crossref","unstructured":"Chung, J., Cho, K., and Bengio, Y. (2016). A character-level decoder without explicit segmentation for neural machine translation. arXiv.","DOI":"10.18653\/v1\/P16-1160"},{"key":"ref_87","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv."},{"key":"ref_88","doi-asserted-by":"crossref","unstructured":"Yang, Z., Yang, D., Dyer, C., He, X., Smola, A., and Hovy, E. (2016, January 12\u201317). Hierarchical attention networks for document classification. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego, CA, USA.","DOI":"10.18653\/v1\/N16-1174"},{"key":"ref_89","unstructured":"Kim, Y., Denton, C., Hoang, L., and Rush, A.M. (2017). Structured attention networks. arXiv."},{"key":"ref_90","unstructured":"Lin, Z., Feng, M., Santos, C.N.d., Yu, M., Xiang, B., Zhou, B., and Bengio, Y. (2017). A structured self-attentive sentence embedding. arXiv."},{"key":"ref_91","doi-asserted-by":"crossref","unstructured":"Shen, T., Zhou, T., Long, G., Jiang, J., Pan, S., and Zhang, C. (2018, January 2\u20137). Disan: Directional self-attention network for rnn\/cnn-free language understanding. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11941"},{"key":"ref_92","unstructured":"Shen, T., Zhou, T., Long, G., Jiang, J., and Zhang, C. (2018). Bi-directional block self-attention for fast and memory-efficient sequence modeling. arXiv."},{"key":"ref_93","doi-asserted-by":"crossref","unstructured":"Shen, T., Zhou, T., Long, G., Jiang, J., Wang, S., and Zhang, C. (2018). Reinforced self-attention network: A hybrid of hard and soft attention for sequence modeling. arXiv.","DOI":"10.24963\/ijcai.2018\/604"},{"key":"ref_94","doi-asserted-by":"crossref","first-page":"63","DOI":"10.1162\/tacl_a_00005","article-title":"Learning structured text representations","volume":"6","author":"Liu","year":"2018","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_95","unstructured":"Lample, G., and Conneau, A. (2019). Cross-lingual language model pretraining. arXiv."},{"key":"ref_96","doi-asserted-by":"crossref","unstructured":"Guo, Q., Qiu, X., Liu, P., Shao, Y., Xue, X., and Zhang, Z. (2019). Star-transformer. arXiv.","DOI":"10.18653\/v1\/N19-1133"},{"key":"ref_97","doi-asserted-by":"crossref","unstructured":"Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q.V., and Salakhutdinov, R. (2019). Transformer-xl: Attentive language models beyond a fixed-length context. arXiv.","DOI":"10.18653\/v1\/P19-1285"},{"key":"ref_98","unstructured":"Song, K., Tan, X., Qin, T., Lu, J., and Liu, T.Y. (2019). Mass: Masked sequence to sequence pre-training for language generation. arXiv."},{"key":"ref_99","doi-asserted-by":"crossref","unstructured":"Reimers, N., and Gurevych, I. (2019). Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv.","DOI":"10.18653\/v1\/D19-1410"},{"key":"ref_100","unstructured":"Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. (2019). Albert: A lite bert for self-supervised learning of language representations. arXiv."},{"key":"ref_101","doi-asserted-by":"crossref","first-page":"64","DOI":"10.1162\/tacl_a_00300","article-title":"Spanbert: Improving pre-training by representing and predicting spans","volume":"8","author":"Joshi","year":"2020","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_102","unstructured":"Clark, K., Luong, M.T., Le, Q.V., and Manning, C.D. (2020). Electra: Pre-training text encoders as discriminators rather than generators. arXiv."},{"key":"ref_103","unstructured":"Conneau, A., Lample, G., Ranzato, M., Denoyer, L., and J\u00e9gou, H. (2017). Word translation without parallel data. arXiv."},{"key":"ref_104","unstructured":"Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., and Askell, A. (2020). Language models are few-shot learners. arXiv."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/11\/11\/511\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:27:05Z","timestamp":1760178425000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/11\/11\/511"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,30]]},"references-count":104,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2020,11]]}},"alternative-id":["info11110511"],"URL":"https:\/\/doi.org\/10.3390\/info11110511","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,10,30]]}}}