{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T14:09:13Z","timestamp":1781791753800,"version":"3.54.5"},"reference-count":67,"publisher":"Cambridge University Press (CUP)","issue":"4","license":[{"start":{"date-parts":[[2020,4,6]],"date-time":"2020-04-06T00:00:00Z","timestamp":1586131200000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2021,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The recent breakthroughs in deep neural architectures across multiple machine learning fields have led to the widespread use of deep neural models. These learners are often applied as black-box models that ignore or insufficiently utilize a wealth of preexisting semantic information. In this study, we focus on the text classification task, investigating methods for augmenting the input to deep neural networks (DNNs) with semantic information. We extract semantics for the words in the preprocessed text from the WordNet semantic graph, in the form of weighted concept terms that form a semantic frequency vector. Concepts are selected via a variety of semantic disambiguation techniques, including a basic, a part-of-speech-based, and a semantic embedding projection method. Additionally, we consider a weight propagation mechanism that exploits semantic relationships in the concept graph and conveys a spreading activation component. We enrich word2vec embeddings with the resulting semantic vector through concatenation or replacement and apply the semantically augmented word embeddings on the classification task via a DNN. Experimental results over established datasets demonstrate that our approach of semantic augmentation in the input space boosts classification performance significantly, with concatenation offering the best performance. We also note additional interesting findings produced by our approach regarding the behavior of term frequency - inverse document frequency normalization on semantic vectors, along with the radical dimensionality reduction potential with negligible performance loss.<\/jats:p>","DOI":"10.1017\/s1351324920000170","type":"journal-article","created":{"date-parts":[[2020,4,6]],"date-time":"2020-04-06T07:55:21Z","timestamp":1586159721000},"page":"391-425","update-policy":"https:\/\/doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":25,"title":["Text classification with semantically enriched word embeddings"],"prefix":"10.1017","volume":"27","author":[{"given":"N.","family":"Pittaras","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"G.","family":"Giannakopoulos","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"G.","family":"Papadakis","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"V.","family":"Karkaletsis","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"56","published-online":{"date-parts":[[2020,4,6]]},"reference":[{"key":"S1351324920000170_ref5","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1084"},{"key":"S1351324920000170_ref2","first-page":"1137","article-title":"A neural probabilistic language model","volume":"3","author":"Bengio","year":"2003","journal-title":"Journal of Machine Learning Research"},{"key":"S1351324920000170_ref6","doi-asserted-by":"publisher","DOI":"10.1561\/2200000016"},{"key":"S1351324920000170_ref34","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1987.1165125"},{"key":"S1351324920000170_ref19","unstructured":"Ganitkevitch, J. , Van Durme, B. and Callison-Burch, C. (2013). Ppdb: The paraphrase database. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. The Association for Computational Linguistics, pp. 758\u2013764."},{"key":"S1351324920000170_ref46","first-page":"3111","volume-title":"Advances in Neural Information Processing Systems","author":"Mikolov","year":"2013"},{"key":"S1351324920000170_ref58","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1170"},{"key":"S1351324920000170_ref22","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143892"},{"key":"S1351324920000170_ref21","doi-asserted-by":"crossref","unstructured":"Goikoetxea, J. , Agirre, E. and Soroa, A. (2016). Single or multiple? combining word representations independently learned from text and wordnet. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence. AAAI Press, pp. 2608\u20132614.","DOI":"10.1609\/aaai.v30i1.10321"},{"key":"S1351324920000170_ref25","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"S1351324920000170_ref14","first-page":"16","article-title":"Using wordnet for text categorization","volume":"5","author":"Elberrichi","year":"2008","journal-title":"International Arab Journal of Information Technology (IAJIT)"},{"key":"S1351324920000170_ref3","doi-asserted-by":"publisher","DOI":"10.1561\/2200000006"},{"key":"S1351324920000170_ref44","doi-asserted-by":"crossref","unstructured":"Mikolov, T. , Karafi\u00e1t, M. , Burget, L. , Cernock\u00fd, J. and Khudanpur, S. (2010). Recurrent neural network based language model. In INTERSPEECH. ISCA, pp. 1045\u20131048.","DOI":"10.1109\/ICASSP.2011.5947611"},{"key":"S1351324920000170_ref17","unstructured":"Fried, D. and Duh, K. (2014). Incorporating both distributional and relational semantics in word representations. CoRR, abs\/1412.5836."},{"key":"S1351324920000170_ref11","doi-asserted-by":"publisher","DOI":"10.1037\/0033-295X.82.6.407"},{"key":"S1351324920000170_ref54","doi-asserted-by":"publisher","DOI":"10.7763\/IJCCE.2014.V3.286"},{"key":"S1351324920000170_ref66","doi-asserted-by":"publisher","DOI":"10.1145\/2661829.2662038"},{"key":"S1351324920000170_ref45","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2011.5947611"},{"key":"S1351324920000170_ref28","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1010"},{"key":"S1351324920000170_ref37","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-72588-6_127"},{"key":"S1351324920000170_ref52","doi-asserted-by":"publisher","DOI":"10.3115\/1621474.1621480"},{"key":"S1351324920000170_ref24","unstructured":"Hinton, G.E. , McClelland, J.L. , Rumelhart, D.E. (1984) Distributed representations. In Parallel distributed processing: explorations in the microstructure of cognition, vol. 1. MIT Press, pp. 77\u2013109."},{"key":"S1351324920000170_ref61","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1.11640"},{"key":"S1351324920000170_ref9","unstructured":"Chollet, F. , et al. (2015). Keras. https:\/\/keras.io."},{"key":"S1351324920000170_ref31","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-04898-2_455"},{"key":"S1351324920000170_ref63","unstructured":"Scott, S. and Matwin, S. (1998). Text classification using wordnet hypernyms. In Usage of WordNet in Natural Language Processing Systems."},{"key":"S1351324920000170_ref12","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390177"},{"key":"S1351324920000170_ref15","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/N15-1184"},{"key":"S1351324920000170_ref47","doi-asserted-by":"publisher","DOI":"10.1145\/219717.219748"},{"key":"S1351324920000170_ref42","first-page":"121","volume-title":"Advances in Neural Information Processing Systems","author":"Mcauliffe","year":"2008"},{"key":"S1351324920000170_ref49","first-page":"246","article-title":"Hierarchical probabilistic neural network language model","volume":"5","author":"Morin","year":"2005","journal-title":"Aistats"},{"key":"S1351324920000170_ref13","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9"},{"key":"S1351324920000170_ref41","volume-title":"Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition","author":"Martin","year":"2009"},{"key":"S1351324920000170_ref39","unstructured":"Luong, T. , Socher, R. and Manning, C. (2013). Better word representations with recursive neural networks for morphology. In Proceedings of the 17th Conference on Computational Natural Language Learning (CoNLL). The Association for Computational Linguistics, pp. 104\u2013113."},{"key":"S1351324920000170_ref10","unstructured":"Chung, J. , Gulcehre, C. , Cho, K. and Bengio, Y. (2014). Empirical evaluation of gated recurrent neural networks on sequence modeling. In NIPS Deep Learning Workshop. MIT Press."},{"key":"S1351324920000170_ref62","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(88)90021-0"},{"key":"S1351324920000170_ref56","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"S1351324920000170_ref59","doi-asserted-by":"publisher","DOI":"10.1108\/00330330610681286"},{"key":"S1351324920000170_ref8","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1110"},{"key":"S1351324920000170_ref27","doi-asserted-by":"publisher","DOI":"10.3233\/HIS-2004-13-402"},{"key":"S1351324920000170_ref53","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2012.07.001"},{"key":"S1351324920000170_ref65","doi-asserted-by":"crossref","unstructured":"Vulic, I. and Mrk\u0161ic, N. (2018). Specialising word vectors for lexical entailment. In Proceedings of NAACL-HLT. The Association for Computational Lingustics, pp. 1134\u20131145.","DOI":"10.18653\/v1\/N18-1103"},{"key":"S1351324920000170_ref18","doi-asserted-by":"crossref","unstructured":"Fukunaga, K. (1990). Introduction to Statistical Pattern Recognition. New York: Academic Press.","DOI":"10.1016\/B978-0-08-047865-4.50007-7"},{"key":"S1351324920000170_ref4","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-662-44848-9_9"},{"key":"S1351324920000170_ref26","unstructured":"Huang, E.H. , Socher, R. , Manning, C.D. and Ng, A.Y. (2012). Improving word representations via global context and multiple word prototypes. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Long Papers-Volume 1. The Association for Computational Linguistics, pp. 873\u2013882."},{"key":"S1351324920000170_ref55","unstructured":"Parker, R. , Graff, D. , Kong, J. , Chen, K. and Maeda, K. (2011). English Gigaword Fifth Edition LDC2011T07. Philadelphia: Linguistic Data Consortium, https:\/\/catalog.ldc.upenn.edu\/LDC2011T07."},{"key":"S1351324920000170_ref7","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1189"},{"key":"S1351324920000170_ref32","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/E17-2068"},{"key":"S1351324920000170_ref36","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2017.7966016"},{"key":"S1351324920000170_ref51","doi-asserted-by":"publisher","DOI":"10.1145\/1459352.1459355"},{"key":"S1351324920000170_ref64","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24309-2_27"},{"key":"S1351324920000170_ref60","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1173"},{"key":"S1351324920000170_ref1","unstructured":"Abadi, M. , Barham, P. , Chen, J. , Chen, Z. , Davis, A. , Dean, J. , Devin, M. , Ghemawat, S. , Irving, G. , Isard, M. , et al. (2016). Tensorflow: a system for large-scale machine learning. In Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation. USENIX Association, pp. 265\u2013283."},{"key":"S1351324920000170_ref57","doi-asserted-by":"publisher","DOI":"10.1002\/meet.1450440226"},{"key":"S1351324920000170_ref48","doi-asserted-by":"publisher","DOI":"10.3115\/1075671.1075742"},{"key":"S1351324920000170_ref43","unstructured":"Mikolov, T. , Chen, K. , Corrado, G. and Dean, J. (2013a). Efficient estimation of word representations in vector space. In Proceedings of the First International Conference on Learning Representations."},{"key":"S1351324920000170_ref30","doi-asserted-by":"publisher","DOI":"10.1007\/BFb0026683"},{"key":"S1351324920000170_ref35","doi-asserted-by":"publisher","DOI":"10.1016\/B978-1-55860-377-6.50048-7"},{"key":"S1351324920000170_ref29","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-016-2401-x"},{"key":"S1351324920000170_ref67","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-2089"},{"key":"S1351324920000170_ref50","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00179"},{"key":"S1351324920000170_ref33","unstructured":"Jurgens, D.A. , Turney, P.D. , Mohammad, S.M. and Holyoak, K.J. (2012). Semeval-2012 task 2: Measuring degrees of relational similarity. In Proceedings of the First Joint Conference on Lexical and Computational Semantics-Volume 1: Proceedings of the Main Conference and the Shared Task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation. The Association for Computational Linguistics, pp. 356\u2013364."},{"key":"S1351324920000170_ref23","doi-asserted-by":"publisher","DOI":"10.1145\/2484028.2484140"},{"key":"S1351324920000170_ref38","doi-asserted-by":"publisher","DOI":"10.3115\/1118108.1118117"},{"key":"S1351324920000170_ref16","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1191"},{"key":"S1351324920000170_ref20","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1004"},{"key":"S1351324920000170_ref40","unstructured":"Mansuy, T.N. and Hilderman, R.J. (2006). Evaluating wordnet features in text classification models. In FLAIRS Conference. AAAI Press, pp. 568\u2013573."}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324920000170","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,10,20]],"date-time":"2022-10-20T17:57:26Z","timestamp":1666288646000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324920000170\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,4,6]]},"references-count":67,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2021,7]]}},"alternative-id":["S1351324920000170"],"URL":"https:\/\/doi.org\/10.1017\/s1351324920000170","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,4,6]]},"assertion":[{"value":"\u00a9 The Author(s) 2020. Published by Cambridge University Press","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}},{"value":"This is an Open Access article, distributed under the terms of the Creative Commons Attribution licence (http:\/\/creativecommons.org\/licenses\/by\/4.0\/), which permits unrestricted re-use, distribution, and reproduction in any medium, provided the original work is properly cited.","name":"license","label":"License","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}},{"value":"This content has been made available to all.","name":"free","label":"Free to read"}]}}