{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,9]],"date-time":"2026-03-09T23:20:27Z","timestamp":1773098427967,"version":"3.50.1"},"reference-count":55,"publisher":"MDPI AG","issue":"17","license":[{"start":{"date-parts":[[2019,8,28]],"date-time":"2019-08-28T00:00:00Z","timestamp":1566950400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"This research was funded by Jilin Provincial Science and Technology Department of China","award":["20170204002GX"],"award-info":[{"award-number":["20170204002GX"]}]},{"name":"Jilin Province Development and Reform Commission of China","award":["2015Y056"],"award-info":[{"award-number":["2015Y056"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Text representation is one of the key tasks in the field of natural language processing (NLP). Traditional feature extraction and weighting methods often use the bag-of-words (BoW) model, which may lead to a lack of semantic information as well as the problems of high dimensionality and high sparsity. At present, to solve these problems, a popular idea is to utilize deep learning methods. In this paper, feature weighting, word embedding, and topic models are combined to propose an unsupervised text representation method named the feature, probability, and word embedding method. The main idea is to use the word embedding technology Word2Vec to obtain the word vector, and then combine this with the feature weighted TF-IDF and the topic model LDA. Compared with traditional feature engineering, the proposed method not only increases the expressive ability of the vector space model, but also reduces the dimensions of the document vector. Besides this, it can be used to solve the problems of the insufficient information, high dimensions, and high sparsity of BoW. We use the proposed method for the task of text categorization and verify the validity of the method.<\/jats:p>","DOI":"10.3390\/s19173728","type":"journal-article","created":{"date-parts":[[2019,8,28]],"date-time":"2019-08-28T11:23:18Z","timestamp":1566991398000},"page":"3728","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":26,"title":["A Method of Short Text Representation Based on the Feature Probability Embedded Vector"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9018-2711","authenticated-orcid":false,"given":"Wanting","family":"Zhou","sequence":"first","affiliation":[{"name":"School of Information Science and Technology, Northeast Normal University, Changchun 130117, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hanbin","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Northeast Normal University, Changchun 130117, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongguang","family":"Sun","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Northeast Normal University, Changchun 130117, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tieli","family":"Sun","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Northeast Normal University, Changchun 130117, China"},{"name":"Department of General Computer, College of Humanities and Science of Northeast Normal University, Changchun 130117, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,8,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"794","DOI":"10.1109\/TFUZZ.2017.2690222","article-title":"Fuzzy bag-of-words model for document representation","volume":"26","author":"Zhao","year":"2018","journal-title":"IEEE Trans. Fuzzy Syst."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zhao, R., and Mao, K. (2014, January 14\u201314). Supervised adaptive-transfer PLSA for cross-domain text classification. Proceedings of the 2014 IEEE International Conference on Data Mining Workshop, Shenzhen, China.","DOI":"10.1109\/ICDMW.2014.163"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1561\/1500000011","article-title":"Opinion mining and sentiment analysis","volume":"2","author":"Pang","year":"2008","journal-title":"Found. Trends Inf. Retr."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"721","DOI":"10.1109\/TPAMI.2008.110","article-title":"Supervised and traditional term weighting methods for automatic text categorization","volume":"31","author":"Lan","year":"2009","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_5","first-page":"179","article-title":"Optimization mutual information text feature selection method based on word frequency","volume":"40","author":"Liu","year":"2014","journal-title":"Comput. Eng."},{"key":"ref_6","first-page":"3279","article-title":"Improved information gain text feature selection algorithm based on word frequency information","volume":"34","author":"Shi","year":"2014","journal-title":"J. Comput. Appl."},{"key":"ref_7","unstructured":"Mikolov, T., Sutskever, I., Chen, K., Corrado, G., and Dean, J. (2013). Distributed representations of words and phrases and their compositionality. Advances in Neural Information Processing Systems, NIPS, Inc."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Joulin, A., Grave, E., Bojanowski, P., and Mikolov, T. (2016). Bag of tricks for efficient text classification. arXiv.","DOI":"10.18653\/v1\/E17-2068"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C. (2014, January 25\u201329). Glove: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Kim, Y. (2014). Convolutional neural networks for sentence classification. arXiv.","DOI":"10.3115\/v1\/D14-1181"},{"key":"ref_11","unstructured":"Zhang, X., Zhao, J., and LeCun, Y. (2015). Character-level convolutional networks for text classification. Advances in Neural Information Processing Systems, Neural Information Processing Systems Foundation, Inc."},{"key":"ref_12","unstructured":"Liu, P., Qiu, X., and Huang, X. (2016). Recurrent neural network for text classification with multi-task learning. arXiv."},{"key":"ref_13","unstructured":"Henaff, M., Weston, J., Szlam, A., Bordes, A., and LeCun, Y. (2016). Tracking the world state with recurrent entity networks. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Yang, Z., Yang, D., Dyer, C., He, X., Smola, A., and Hovy, E. (2016, January 12\u201317). Hierarchical attention networks for document classification. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego, CA, USA.","DOI":"10.18653\/v1\/N16-1174"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1715780","DOI":"10.1155\/2016\/1715780","article-title":"A feature selection approach based on inter-class and intra-class relative contributions of terms","volume":"2016","author":"Zhou","year":"2016","journal-title":"Comput. Intell. Neurosci."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"245","DOI":"10.1016\/j.eswa.2016.09.009","article-title":"Turning from TF\u2013IDF to TF\u2013IGM for term weighting in text classification","volume":"66","author":"Chen","year":"2016","journal-title":"Expert Syst. Appl."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1186\/s13673-018-0135-8","article-title":"QER: A new feature selection method for sentiment analysis","volume":"8","author":"Parlar","year":"2018","journal-title":"Hum. Cent. Comput. Inf. Sci."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"75","DOI":"10.1007\/s13042-015-0347-4","article-title":"Sentimental feature selection for sentiment analysis of Chinese online reviews","volume":"9","author":"Zheng","year":"2018","journal-title":"Int. J. Mach. Learn. Cybern."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"377","DOI":"10.1016\/j.ins.2017.11.035","article-title":"Double regularization methods for robust feature selection and SVM classification via DC programming","volume":"429","author":"Maldonado","year":"2018","journal-title":"Inf. Sci."},{"key":"ref_20","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Artetxe, M., Labaka, G., Lopez-Gazpio, I., and Agirre, E. (2018). Uncovering divergent linguistic information in word embeddings with lessons for intrinsic and extrinsic evaluation. arXiv.","DOI":"10.18653\/v1\/K18-1028"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Peters, M.E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018). Deep contextualized word representations. arXiv.","DOI":"10.18653\/v1\/N18-1202"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"212","DOI":"10.1016\/j.joi.2016.01.006","article-title":"Selecting publication keywords for domain analysis in bibliometrics: A comparison of three methods","volume":"10","author":"Chen","year":"2016","journal-title":"J. Informetr."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1031","DOI":"10.1007\/s11192-017-2574-9","article-title":"A domain keyword analysis approach extending Term Frequency-Keyword Active Index with Google Word2Vec model","volume":"114","author":"Hu","year":"2018","journal-title":"Scientometrics"},{"key":"ref_25","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv."},{"key":"ref_26","unstructured":"Moody, C.E. (2016). Mixing dirichlet topic models and word embeddings to make lda2vec. arXiv."},{"key":"ref_27","unstructured":"Hinton, G.E. (1986, January 15\u201317). Learning distributed representations of concepts. Proceedings of the Eighth Annual Conference of the Cognitive Science Society, Amherst, MA, USA."},{"key":"ref_28","first-page":"1137","article-title":"A neural probabilistic language model","volume":"3","author":"Bengio","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"513","DOI":"10.1016\/0306-4573(88)90021-0","article-title":"Term-weighting approaches in automatic text retrieval","volume":"24","author":"Salton","year":"1988","journal-title":"Inf. Process. Manag."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"012023","DOI":"10.1088\/1757-899X\/261\/1\/012023","article-title":"Sentiments Analysis of Reviews Based on ARCNN Model","volume":"Volume 261","author":"Xu","year":"2017","journal-title":"IOP Conference Series: Materials Science and Engineering"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Ren, X., Zhang, L., Ye, W., Hua, H., and Zhang, S. (2018, January 4\u20137). Attention Enhanced Chinese Word Embeddings. Proceedings of the International Conference on Artificial Neural Networks, Rhodes, Greece.","DOI":"10.1007\/978-3-030-01418-6_16"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wang, J., and Li, C. (2011, January 25\u201327). An iterative voting method based on word density for text classification. Proceedings of the International Conference on Web Intelligence, Mining and Semantics, Sogndal, Norway.","DOI":"10.1145\/1988688.1988751"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Liu, C.Z., Sheng, Y.X., Wei, Z.Q., and Yang, Y.Q. (2018, January 24\u201327). Research of Text Classification Based on Improved TF\u2013IDF Algorithm. Proceedings of the 2018 IEEE International Conference of Intelligent Robotic and Control Engineering (IRCE), Lanzhou, China.","DOI":"10.1109\/IRCE.2018.8492945"},{"key":"ref_34","first-page":"993","article-title":"Latent dirichlet allocation","volume":"3","author":"Blei","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_35","unstructured":"Altszyler, E., Sigman, M., Ribeiro, S., and Slezak, D.F. (2016). Comparative study of LSA vs Word2vec embeddings in small corpora: A case study in dreams database. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Zhao, J., Lan, M., and Tian, J.F. (2015, January 4\u20135). Ecnu: Using traditional similarity measurements and word embedding for semantic textual similarity estimation. Proceedings of the 9th International Workshop on Semantic Evaluation, Denver, CO, USA.","DOI":"10.18653\/v1\/S15-2021"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Shi, M., Liu, J., Zhou, D., Tang, M., and Cao, B. (2017, January 25\u201330). WE-LDA: A word embeddings augmented LDA model for web services clustering. Proceedings of the 2017 IEEE International Conference on Web Services (ICWS), Honolulu, HI, USA.","DOI":"10.1109\/ICWS.2017.9"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Das, R., Zaheer, M., and Dyer, C. (2015, January 26\u201331). Gaussian lda for topic models with word embeddings. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Beijing, China.","DOI":"10.3115\/v1\/P15-1077"},{"key":"ref_39","unstructured":"Gregor, H. (2019, August 27). Parameter Estimation for Text Analysis. Available online: www.arbylon.net\/publications\/text-est2.pdf."},{"key":"ref_40","unstructured":"(2019, August 27). Extreme Multi-Label Classification Repository. Available online: http:\/\/manikvarma.org\/downloads\/XC\/XMLRepository.html."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Babbar, R., and Sch\u00f6lkopf, B. (2017, January 6\u201310). Dismec: Distributed sparse machines for extreme multi-label classification. Proceedings of the Tenth ACM International Conference on Web Search and Data Mining, Cambridge, UK.","DOI":"10.1145\/3018661.3018741"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Prabhu, Y., Kag, A., Harsola, S., Agrawal, R., and Varma, M. (2018, January 23\u201327). Parabel: Partitioned label trees for extreme classification with application to dynamic search advertising. Proceedings of the 2018 World Wide Web Conference, Lyon, France.","DOI":"10.1145\/3178876.3185998"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"1329","DOI":"10.1007\/s10994-019-05791-5","article-title":"Data scarcity, robustness and extreme multi-label classification","volume":"108","author":"Babbar","year":"2019","journal-title":"Mach. Learn."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Wang, H., Lu, Y., and Zhai, C.X. (2011, January 21\u201324). Latent aspect rating analysis without aspect keyword supervision. Proceedings of the 17th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Diego, CA, USA.","DOI":"10.1145\/2020408.2020505"},{"key":"ref_45","unstructured":"Qiu, X., Zhang, Q., and Huang, X. (2013, January 4\u20139). Fudannlp: A toolkit for chinese natural language processing. Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Sofia, Bulgaria."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Tan, S.B., Cheng, X.Q., Wang, Y.F., and Xu, H.B. (2009, January 6\u20139). Adapting naive bayes to domain adaptation for sentiment analysis. Proceedings of the European Conference on Information Retrieval, Toulouse, France.","DOI":"10.1007\/978-3-642-00958-7_31"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Zhang, H.-P., Liu, Q., Cheng, X.-Q., Zhang, H., and Yu, H.-K. (2003, January 11\u201312). Chinese lexical analysis using hierarchical hidden markov model. Proceedings of the Second SIGHAN Workshop on Chinese Language Processing, Sapporo, Japan.","DOI":"10.3115\/1119250.1119259"},{"key":"ref_48","unstructured":"(2010, January 12). Chinese Stop Words List. Available online: https:\/\/download.csdn.net\/download\/echo1004\/1987618."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"128","DOI":"10.1016\/S0019-9958(70)90081-1","article-title":"A generalized k-nearest neighbor rule","volume":"16","author":"Patrick","year":"1970","journal-title":"Inf. Control"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1145\/1961189.1961199","article-title":"LIBSVM: A library for support vector machines","volume":"2","author":"Chang","year":"2011","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"69","DOI":"10.1023\/A:1009982220290","article-title":"An evaluation of statistical approaches to text categorization","volume":"1","author":"Yang","year":"1999","journal-title":"Inf. Retr."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Martin\u010di\u0107-Ip\u0161i\u0107, S., and Mili\u010di\u0107, T. (2019). The Influence of Feature Representation of Text on the Performance of Document Classification. Appl. Sci., 9.","DOI":"10.3390\/app9040743"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Zhang, Q., Chen, H., and Huang, X. (2014, January 24\u201328). Chinese-English mixed text normalization. Proceedings of the 7th ACM International Conference on Web Search and Data Mining, New York, NY, USA.","DOI":"10.1145\/2556195.2556228"},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"9139","DOI":"10.1016\/j.eswa.2011.01.047","article-title":"Exploiting effective features for Chinese sentiment classification","volume":"38","author":"Zhai","year":"2011","journal-title":"Expert Syst. Appl."},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1016\/j.knosys.2015.06.015","article-title":"A survey on opinion mining and sentiment analysis: Tasks, approaches and applications","volume":"89","author":"Ravi","year":"2015","journal-title":"Knowl. Based Syst."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/17\/3728\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:14:46Z","timestamp":1760188486000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/17\/3728"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,8,28]]},"references-count":55,"journal-issue":{"issue":"17","published-online":{"date-parts":[[2019,9]]}},"alternative-id":["s19173728"],"URL":"https:\/\/doi.org\/10.3390\/s19173728","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,8,28]]}}}