{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,2]],"date-time":"2026-03-02T22:28:53Z","timestamp":1772490533908,"version":"3.50.1"},"reference-count":53,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2019,5,12]],"date-time":"2019-05-12T00:00:00Z","timestamp":1557619200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key Research and Development Plan","award":["2016YFB0801204"],"award-info":[{"award-number":["2016YFB0801204"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Words have different meanings (i.e., senses) depending on the context. Disambiguating the correct sense is important and a challenging task for natural language processing. An intuitive way is to select the highest similarity between the context and sense definitions provided by a large lexical database of English, WordNet. In this database, nouns, verbs, adjectives, and adverbs are grouped into sets of cognitive synonyms interlinked through conceptual semantics and lexicon relations. Traditional unsupervised approaches compute similarity by counting overlapping words between the context and sense definitions which must match exactly. Similarity should compute based on how words are related rather than overlapping by representing the context and sense definitions on a vector space model and analyzing distributional semantic relationships among them using latent semantic analysis (LSA). When a corpus of text becomes more massive, LSA consumes much more memory and is not flexible to train a huge corpus of text. A word-embedding approach has an advantage in this issue. Word2vec is a popular word-embedding approach that represents words on a fix-sized vector space model through either the skip-gram or continuous bag-of-words (CBOW) model. Word2vec is also effectively capturing semantic and syntactic word similarities from a huge corpus of text better than LSA. Our method used Word2vec to construct a context sentence vector, and sense definition vectors then give each word sense a score using cosine similarity to compute the similarity between those sentence vectors. The sense definition also expanded with sense relations retrieved from WordNet. If the score is not higher than a specific threshold, the score will be combined with the probability of that sense distribution learned from a large sense-tagged corpus, SEMCOR. The possible answer senses can be obtained from high scores. Our method shows that the result (50.9% or 48.7% without the probability of sense distribution) is higher than the baselines (i.e., original, simplified, adapted and LSA Lesk) and outperforms many unsupervised systems participating in the SENSEVAL-3 English lexical sample task.<\/jats:p>","DOI":"10.3390\/fi11050114","type":"journal-article","created":{"date-parts":[[2019,5,13]],"date-time":"2019-05-13T05:35:39Z","timestamp":1557725739000},"page":"114","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":71,"title":["Word Sense Disambiguation Using Cosine Similarity Collaborates with Word2vec and WordNet"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1188-7212","authenticated-orcid":false,"given":"Korawit","family":"Orkphol","sequence":"first","affiliation":[{"name":"Information Security Research Center, College of Computer Science and Technology, Harbin Engineering University, Harbin 150001, China"}]},{"given":"Wu","family":"Yang","sequence":"additional","affiliation":[{"name":"Information Security Research Center, College of Computer Science and Technology, Harbin Engineering University, Harbin 150001, China"}]}],"member":"1968","published-online":{"date-parts":[[2019,5,12]]},"reference":[{"key":"ref_1","unstructured":"Zhong, Z., and Ng, H.T. (2010, January 13). It Makes Sense: A Wide-Coverage Word Sense Disambiguation System for Free Text. Proceedings of the ACL 2010 System Demonstrations, Uppsala, Sweden."},{"key":"ref_2","unstructured":"Lesk, M. (, January January). Automatic Sense Disambiguation Using Machine Readable Dictionaries: How to Tell a Pine Cone from an Ice Cream Cone. Proceedings of the 5th Annual International Conference on Systems Documentation, Toronto, ON, Canada."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Cowie, J., Guthrie, J., and Guthrie, L. (1992, January 23\u201328). Lexical Disambiguation Using Simulated Annealing. Proceedings of the 14th conference on Computational linguistics-Volume 1, Nantes, France.","DOI":"10.3115\/992066.992125"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1023\/A:1002693207386","article-title":"Framework and Results for English SENSEVAL","volume":"34","author":"Kilgarriff","year":"2000","journal-title":"Comput. Humanit."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Banerjee, S., and Pedersen, T. (2002, January 17\u201323). An Adapted Lesk Algorithm for Word Sense Disambiguation Using WordNet. Proceedings of the International Conference on Intelligent Text Processing and Computational Linguistics, Mexico City, Mexico.","DOI":"10.1007\/3-540-45715-1_11"},{"key":"ref_6","unstructured":"Basile, P., Caputo, A., and Semeraro, G. (2014, January 23\u201329). An Enhanced Lesk Word Sense Disambiguation Algorithm through a Distributional Semantic Model. Proceedings of the COLING 2014, 25th International Conference on Computational Linguistics: Technical Papers, Dublin, Ireland."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Yarowsky, D. (1995, January 26\u201330). Unsupervised Word Sense Disambiguation Rivaling Supervised Methods. Proceedings of the 33rd Annual Meeting on Association for Computational Linguistics, Cambridge, MA, USA.","DOI":"10.3115\/981658.981684"},{"key":"ref_8","unstructured":"Miller, G. (1998). WordNet: An Electronic Lexical Database, MIT press."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1080\/01638539809545028","article-title":"An Introduction to Latent Semantic Analysis","volume":"25","author":"Landauer","year":"1998","journal-title":"Discourse Process."},{"key":"ref_10","unstructured":"Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., and Dean, J. (2013, January 5\u20138). Distributed Representations of Words and Phrases and Their Compositionality. Proceedings of the Advances in Neural Information Processing Systems (NIPS 2013), Lake Tahoe, NV, USA."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Baroni, M., Dinu, G., and Kruszewski, G. (2014, January 23\u201325). Don\u2019t Count, Predict! A Systematic Comparison of Context-Counting vs. Context-Predicting Semantic Vectors. Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Baltimore, MD, USA.","DOI":"10.3115\/v1\/P14-1023"},{"key":"ref_12","unstructured":"Altszyler, E., Sigman, M., Ribeiro, S., and Slezak, D.F. (2016). Comparative Study of LSA vs Word2vec Embeddings in Small Corpora: A Case Study in Dreams Database. arXiv."},{"key":"ref_13","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. arXiv."},{"key":"ref_14","unstructured":"Mihalcea, R. (1998). Semcor Semantically Tagged Corpus, Unpublished manuscript."},{"key":"ref_15","unstructured":"Mihalcea, R., Chklovski, T., and Kilgarriff, A. (2004, January 25\u201326). The Senseval-3 English Lexical Sample Task. Proceedings of the SENSEVAL-3, the Third International Workshop on the Evaluation of Systems for the Semantic Analysis of Text, Barcelona, Spain."},{"key":"ref_16","unstructured":"Vasilescu, F., Langlais, P., and Lapalme, G. (2004, January 26\u201328). Evaluating Variants of the Lesk Approach for Disambiguating Words. Proceedings of the Fourth International Conference on Language Resources and Evaluation (Lrec\u201904), Lisbon, Portugal."},{"key":"ref_17","unstructured":"Edmonds, P., and Cotton, S. (2001, January 5\u20136). SENSEVAL-2: Overview. Proceedings of the Second International Workshop on Evaluating Word Sense Disambiguation Systems, SENSEVAL \u201901, Toulouse, France."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Mihalcea, R., Tarau, P., and Figa, E. (2004, January 23\u201327). PageRank on Semantic Networks, with Application to Word Sense Disambiguation. Proceedings of the 20th International Conference on Computational Linguistics, Geneva, Switzerland.","DOI":"10.3115\/1220355.1220517"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"217","DOI":"10.1016\/j.artint.2012.07.001","article-title":"BabelNet: The Automatic Construction, Evaluation and Application of a Wide-Coverage Multilingual Semantic Network","volume":"193","author":"Navigli","year":"2012","journal-title":"Artif. Intell."},{"key":"ref_20","unstructured":"Navigli, R., Jurgens, D., and Vannella, D. (2013, January 14\u201315). Semeval-2013 Task 12: Multilingual Word Sense Disambiguation. Proceedings of the Second Joint Conference on Lexical and Computational Semantics (* SEM), Volume 2 and the Seventh International Workshop on Semantic Evaluation (SemEval 2013), Atlanta, GA, USA."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"185","DOI":"10.1016\/0925-2312(93)90006-O","article-title":"Backpropagation and Stochastic Gradient Descent Method","volume":"5","author":"Amari","year":"1993","journal-title":"Neurocomputing"},{"key":"ref_22","first-page":"307","article-title":"Noise-Contrastive Estimation of Unnormalized Statistical Models, with Applications to Natural Image Statistics","volume":"13","author":"Gutmann","year":"2012","journal-title":"J. Mach. Learn. Res."},{"key":"ref_23","unstructured":"Mnih, A., and Whye Teh, Y. (July, January 26). A Fast and Simple Algorithm for Training Neural Probabilistic Language Models. Proceedings of the 29th International Conference on Machine Learning, ICML 2012, Edinburgh, UK."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Faruqui, M., Dodge, J., Kumar Jauhar, S., Dyer, C., Hovy, E., and Smith, N.A. (June, January 31). Retrofitting Word Vectors to Semantic Lexicons. Proceedings of the Human Language Technologies: The 2015 Annual Conference of the North American Chapter of the ACL, Denver, CO, USA.","DOI":"10.3115\/v1\/N15-1184"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Yu, L., Moritz Hermann, K., Blunsom, P., and Pulman, S. (2014, January 8\u201313). Deep Learning for Answer Sentence Selection. Proceedings of the Deep Learning and Representation Learning Workshop: NIPS-2014, Montr\u00e9al, QC, Canada.","DOI":"10.1201\/b17103-3"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Kenter, T., and de Rijke, M. (2015, January 18\u201323). Short Text Similarity with Word Embeddings. Proceedings of the 24th ACM International on Conference on Information and Knowledge Management, Melbourne, Australia.","DOI":"10.1145\/2806416.2806475"},{"key":"ref_27","first-page":"35","article-title":"Modern Information Retrieval: A Brief Overview","volume":"24","author":"Singhal","year":"2001","journal-title":"IEEE Data Eng. Bull."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Manning, C.D., Raghavan, P., and Sch\u00fctze, H. (2008). Introduction to Information Retrieval, Cambridge University Press.","DOI":"10.1017\/CBO9780511809071"},{"key":"ref_29","unstructured":"(2018, November 10). SENSEVAL 3 Home Page. Available online: http:\/\/web.eecs.umich.edu\/~mihalcea\/senseval\/senseval3\/data.html."},{"key":"ref_30","unstructured":"Resnik, P., and Yarowsky, D. A Perspective on Word Sense Disambiguation Methods and Their Evaluation. Tagging Text with Lexical Semantics: Why, What, and How?."},{"key":"ref_31","unstructured":"(2018, November 15). Proposal for Senseval Scoring Scheme. Available online: http:\/\/web.eecs.umich.edu\/~mihalcea\/senseval\/senseval3\/scoring\/scorescheme.txt."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"421","DOI":"10.1007\/s10579-010-9124-x","article-title":"Steven Bird, Ewan Klein and Edward Loper: Natural Language Processing with Python, Analyzing Text with the Natural Language Toolkit","volume":"44","author":"Wagner","year":"2010","journal-title":"Lang. Resour. Eval."},{"key":"ref_33","first-page":"2825","article-title":"Scikit-Learn: Machine Learning in Python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref_34","unstructured":"\u0158eh\u016f\u0159ek, R., and Sojka, P. (2010, January 22). Software Framework for Topic Modelling with Large Corpora. Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks, Valletta, Malta."},{"key":"ref_35","unstructured":"(2018, October 09). Google Code Archive: word2vec. Available online: https:\/\/code.google.com\/archive\/p\/word2vec\/."},{"key":"ref_36","unstructured":"(2018, October 09). Word Analogies Test for word2vec. Available online: https:\/\/raw.githubusercontent.com\/RaRe-Technologies\/gensim\/develop\/gensim\/test\/test_data\/questions-words.txt."},{"key":"ref_37","unstructured":"Han, L. (2018, October 21). UMBC Webbase Corpus. Available online: https:\/\/ebiquity.umbc.edu\/resource\/html\/id\/351."},{"key":"ref_38","unstructured":"Han, L., Kashyap, A.L., Finin, T., Mayfield, J., and Weese, J. (2013, January 13\u201314). UMBC_EBIQUITY-CORE: Semantic Textual Similarity Systems. Proceedings of the Second Joint Conference on Lexical and Computational Semantics, Atlanta, Georgia, USA."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Hilbe, J.M. (2009). Logistic Regression Models, Chapman and Hall\/CRC.","DOI":"10.1201\/9781420075779"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"239","DOI":"10.1023\/B:USER.0000028980.13669.44","article-title":"User Modelling for News Web Sites with Word Sense Based Techniques","volume":"14","author":"Magnini","year":"2004","journal-title":"User Model. User-Adapt. Interact."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"491","DOI":"10.13053\/cys-18-3-2043","article-title":"Soft Similarity and Soft Cosine Measure: Similarity of Features in Vector Space Model","volume":"18","author":"Sidorov","year":"2014","journal-title":"Computaci\u00f3n y Sistemas"},{"key":"ref_42","unstructured":"Le, Q., and Mikolov, T. (2014, January 22\u201324). Distributed Representations of Sentences and Documents. Proceedings of the 31st International Conference on Machine Learning, Beijing, China."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). Glove: Global Vectors for Word Representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Bojanowski, P., Grave, E., Joulin, A., and Mikolov, T. (2016). Enriching Word Vectors with Subword Information. arXiv.","DOI":"10.1162\/tacl_a_00051"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Yu, H., and Hatzivassiloglou, V. (2003, January 11\u201312). Towards Answering Opinion Questions: Separating Facts from Opinions and Identifying the Polarity of Opinion Sentences. Proceedings of the 2003 Conference on Empirical Methods in Natural Language Processing, Sapporo, Japan.","DOI":"10.3115\/1119355.1119372"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Kim, S.-M., and Hovy, E. (2004, January 23\u201327). Determining the Sentiment of Opinions. Proceedings of the 20th international conference on Computational Linguistics, Geneva, Switzerland.","DOI":"10.3115\/1220355.1220555"},{"key":"ref_47","unstructured":"Baccianella, S., Esuli, A., and Sebastiani, F. (2010, January 17\u201320). Sentiwordnet 3.0: An Enhanced Lexical Resource for Sentiment Analysis and Opinion Mining. Proceedings of the Seventh conference on International Language Resources and Evaluation (LREC\u201910), Valletta, Malta."},{"key":"ref_48","unstructured":"Musto, C., Semeraro, G., and Polignano, M. (2014, January 10). A Comparison of Lexicon-Based Approaches for Sentiment Analysis of Microblog Posts. Proceedings of the 8th International Workshop on Information Filtering and Retrieval, Pisa, Italy."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"89","DOI":"10.1007\/s00146-014-0549-4","article-title":"Social Media Analytics: A Survey of Techniques, Tools and Platforms","volume":"30","author":"Batrinca","year":"2015","journal-title":"AI & Soc."},{"key":"ref_50","first-page":"973","article-title":"Discovering Consumer Insight from Twitter via Sentiment Analysis","volume":"18","author":"Chamlertwat","year":"2012","journal-title":"J. UCS"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Hamzehei, A., Ebrahimi, M., Shafiee, E., Wong, R.K., and Chen, F. (July, January 27). Scalable Sentiment Analysis for Microblogs Based on Semantic Scoring. Proceedings of the 2015 IEEE International Conference on Services Computing, New York, NY, USA.","DOI":"10.1109\/SCC.2015.45"},{"key":"ref_52","unstructured":"Preo\u0163iuc-Pietro, D., Liu, Y., Hopkins, D., and Ungar, L. (August, January 30). Beyond Binary Labels: Political Ideology Prediction of Twitter Users. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vancouver, BC, Canada."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Liu, Y., Zhang, L., Nie, L., Chen, Y., and Rosenblum, D.S. (2016, January 12\u201317). Fortune Teller: Predicting Your Career Path. Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (AAAI\u201916), Phoenix, AZ, USA.","DOI":"10.1609\/aaai.v30i1.9969"}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/11\/5\/114\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:51:15Z","timestamp":1760187075000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/11\/5\/114"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,5,12]]},"references-count":53,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2019,5]]}},"alternative-id":["fi11050114"],"URL":"https:\/\/doi.org\/10.3390\/fi11050114","relation":{},"ISSN":["1999-5903"],"issn-type":[{"value":"1999-5903","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,5,12]]}}}