{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:48:18Z","timestamp":1760237298158,"version":"build-2065373602"},"reference-count":30,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2020,3,30]],"date-time":"2020-03-30T00:00:00Z","timestamp":1585526400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>The rapid growth of Internet technologies has led to an enormous increase in the number of electronic documents used worldwide. To organize and manage big data for unstructured documents effectively and efficiently, text categorization has been employed in recent decades. To conduct text categorization tasks, documents are usually represented using the bag-of-words model, owing to its simplicity. In this representation for text classification, feature selection becomes an essential method because all terms in the vocabulary induce enormous feature space corresponding to the documents. In this paper, we propose a new feature selection method that considers term similarity to avoid the selection of redundant terms. Term similarity is measured using a general method such as mutual information, and serves as a second measure in feature selection in addition to term ranking. To consider balance of term ranking and term similarity for feature selection, we use a quadratic programming-based numerical optimization approach. Experimental results demonstrate that considering term similarity is effective and has higher accuracy than conventional methods.<\/jats:p>","DOI":"10.3390\/e22040395","type":"journal-article","created":{"date-parts":[[2020,3,31]],"date-time":"2020-03-31T13:27:19Z","timestamp":1585661239000},"page":"395","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Generalized Term Similarity for Feature Selection in Text Classification Using Quadratic Programming"],"prefix":"10.3390","volume":"22","author":[{"given":"Hyunki","family":"Lim","sequence":"first","affiliation":[{"name":"Image and Media Research Center, Korea Institute of Science and Technology, 5 Hwarang-Ro 14-gil, Seongbuk-Gu, Seoul 02792, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7124-1141","authenticated-orcid":false,"given":"Dae-Won","family":"Kim","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Chung-Ang University, 221 Heukseok-Dong, Dongjak-Gu, Seoul 06974, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,3,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1011441423217","article-title":"Text categorization based on regularized linear classification methods","volume":"4","author":"Zhang","year":"2001","journal-title":"Inf. Retr."},{"key":"ref_2","first-page":"60","article-title":"A survey of text mining techniques and applications","volume":"1","author":"Gupta","year":"2009","journal-title":"J. Emerg. Technol. Web Intell."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1016\/j.knosys.2016.02.017","article-title":"Two feature weighting approaches for naive Bayes text classifiers","volume":"100","author":"Zhang","year":"2016","journal-title":"Knowl.-Based Syst."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1656","DOI":"10.1109\/TKDE.2014.2373357","article-title":"Relevance feature discovery for text mining","volume":"27","author":"Li","year":"2015","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2508","DOI":"10.1109\/TKDE.2016.2563436","article-title":"Toward optimal feature selection in naive Bayes for text categorization","volume":"28","author":"Tang","year":"2016","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.patrec.2014.02.013","article-title":"t-Test feature selection approach based on term frequency for text categorization","volume":"45","author":"Wang","year":"2014","journal-title":"Pattern Recognit. Lett."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"518","DOI":"10.1016\/j.ins.2016.08.073","article-title":"Terms-based discriminative information space for robust text classification","volume":"372","author":"Junejo","year":"2016","journal-title":"Inf. Sci."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Dasgupta, A., Drineas, P., Harb, B., Josifovski, V., and Mahoney, M.W. (2007, January 12\u201315). Feature selection methods for text classification. Proceedings of the 13th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Jose, CA, USA.","DOI":"10.1145\/1281192.1281220"},{"key":"ref_9","first-page":"1289","article-title":"An extensive empirical study of feature selection metrics for text classification","volume":"3","author":"Forman","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"95","DOI":"10.1002\/nav.3800030109","article-title":"An algorithm for quadratic programming","volume":"3","author":"Frank","year":"1956","journal-title":"Nav. Res. Logist."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"382","DOI":"10.2307\/1909468","article-title":"The simplex method for quadratic programming","volume":"27","author":"Wolfe","year":"1959","journal-title":"Econometrica"},{"key":"ref_12","unstructured":"Lin, J., and Gunopulos, D. (2003, January 1\u20133). Dimensionality reduction by random projection and latent semantic indexing. Proceedings of the Text Mining Workshop, at The 3rd SIAM International Conference on Data Mining, San Francisco, CA, USA."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Bingham, E., and Mannila, H. (2001, January 26\u201329). Random projection in dimensionality reduction: applications to image and text data. Proceedings of the 7th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA.","DOI":"10.1145\/502512.502546"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"301","DOI":"10.1007\/s10044-003-0196-8","article-title":"Discriminative features for text document classification","volume":"6","author":"Torkkola","year":"2004","journal-title":"Pattern Anal. Appl."},{"key":"ref_15","unstructured":"Yang, Y., and Pedersen, J.O. (1997, January 8\u201312). A comparative study on feature selection in text categorization. Proceedings of the Fourteenth International Conference on Machine Learning (ICML), Tennessee, TN, USA."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"81","DOI":"10.1007\/BF00116251","article-title":"Induction of decision trees","volume":"1","author":"Quinlan","year":"1986","journal-title":"Mach. Learn."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1016\/j.eswa.2015.08.050","article-title":"An improved global feature selection scheme for text classification","volume":"43","author":"Uysal","year":"2016","journal-title":"Expert Syst. Appl."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1602","DOI":"10.1109\/TKDE.2016.2522427","article-title":"A Bayesian classification approach using class-specific features for text categorization","volume":"28","author":"Tang","year":"2016","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1016\/j.neucom.2015.01.031","article-title":"A two-stage Markov blanket based feature selection algorithm for text classification","volume":"157","author":"Javed","year":"2015","journal-title":"Neurocomputing"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"226","DOI":"10.1016\/j.knosys.2012.06.005","article-title":"A novel probabilistic feature selection method for text classification","volume":"36","author":"Uysal","year":"2012","journal-title":"Knowl.-Based Syst."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1016\/j.patrec.2017.02.004","article-title":"Optimization approach for feature selection in multi-label classification","volume":"89","author":"Lim","year":"2017","journal-title":"Pattern Recognit. Lett."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Lim, H., and Kim, D.W. (2016, January 4\u20138). Convex optimization approach for multi-label feature selection based on mutual information. Proceedings of the 23rd International Conference on Pattern Recognition (ICPR), Cancun, Mexico.","DOI":"10.1109\/ICPR.2016.7899851"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Boyd, S., and Vandenberghe, L. (2004). Convex Optimization, Cambridge University Press.","DOI":"10.1017\/CBO9780511804441"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"2039","DOI":"10.1162\/NECO_a_00770","article-title":"Indefinite proximity learning: A review","volume":"27","author":"Schleif","year":"2015","journal-title":"Neural Comput."},{"key":"ref_25","unstructured":"Gu, S., and Guo, Y. (2012, January 22\u201326). Learning SVM Classifiers with Indefinite Kernels. Proceedings of the 26th Conference on Artificial Intelligence (AAAI), Toronto, ON, Canada."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"123","DOI":"10.1137\/0907009","article-title":"Computing the minimum eigenvalue of a symmetric positive definite Toeplitz matrix","volume":"7","author":"Cybenko","year":"1986","journal-title":"SIAM J. Sci. Comput."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1007\/BF01587086","article-title":"An extension of Karmarkar\u2019s projective algorithm for convex quadratic programming","volume":"44","author":"Ye","year":"1989","journal-title":"Math. Program."},{"key":"ref_28","unstructured":"McCallum, A., and Nigam, K. (1998, January 26\u201330). A comparison of event models for naive bayes text classification. Proceedings of the 1998 15th Conference on Artificial Intelligence (AAAI) Workshop on Learning for Text Categorization, Madison, WI, USA."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"2793","DOI":"10.1016\/j.camwa.2011.07.045","article-title":"A two-stage feature selection method for text categorization","volume":"62","author":"Meng","year":"2011","journal-title":"Comput. Math. Appl."},{"key":"ref_30","first-page":"1491","article-title":"Quadratic programming feature selection","volume":"11","author":"Huerta","year":"2010","journal-title":"J. Mach. Learn. Res."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/22\/4\/395\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:13:36Z","timestamp":1760174016000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/22\/4\/395"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,30]]},"references-count":30,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2020,4]]}},"alternative-id":["e22040395"],"URL":"https:\/\/doi.org\/10.3390\/e22040395","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2020,3,30]]}}}