{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T06:40:59Z","timestamp":1777704059848,"version":"3.51.4"},"reference-count":40,"publisher":"SAGE Publications","issue":"4","license":[{"start":{"date-parts":[[2017,10,1]],"date-time":"2017-10-01T00:00:00Z","timestamp":1506816000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"published-print":{"date-parts":[[2017,10]]},"abstract":"<jats:p>\n                    In text classification field, many classifiers cannot deal with the features with large dimensions, thus it is very important to filter the redundant information from the original feature space efficiently and achieve the features with best qualities. On this basis, a new two-step based feature selection method is proposed in this paper. Firstly, some definitions (word semantic correlation, set semantic correlation, semantic correlative and semantic correlative set) are introduced, and the algorithm of generating the semantic correlative sets is given. Secondly, the process of the two-step based feature selection method is described: in the first step, a feature subset is obtained by using an optimal feature selection method, and a set of semantic correlative sets is generated by using the selected feature subset; in the second step, the redundant information of the selected features is filtered by using the generated semantic correlative sets. Finally, in order to avoid local optimum when searching the best threshold, the conception of memory recall position is introduced and an improved memory recall mechanism based fruit fly optimization algorithm is proposed. In the experiments, two typical classifiers: support vector machine and na\u00efve bayes are used on four datasets: Reuters50, SMSSPAS, WebKB and 20-Newsgroups, and the 10-cross validation is carried out when the measurements of F\n                    <jats:sub>1<\/jats:sub>\n                    and receiver operating curve are used. Experimental results show that the proposed method achieves higher accuracy than several representative traditional feature selection methods and runs faster than typical mutual information based feature selection methods, illustrating its effectiveness on achieving the best features in text classification filed.\n                  <\/jats:p>","DOI":"10.3233\/jifs-161541","type":"journal-article","created":{"date-parts":[[2017,9,22]],"date-time":"2017-09-22T10:59:38Z","timestamp":1506077978000},"page":"2059-2073","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":2,"title":["Two-step based feature selection method for filtering redundant information"],"prefix":"10.1177","volume":"33","author":[{"given":"Youwei","family":"Wang","sequence":"first","affiliation":[{"name":"School of Information, Central University of Finance and Economics, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lizhou","family":"Feng","sequence":"additional","affiliation":[{"name":"School of Science and Engineering, Tianjin University of Finance and Economics, Tianjin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Li","sequence":"additional","affiliation":[{"name":"School of Information, Central University of Finance and Economics, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2017,10]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2014.11.038"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2011.09.160"},{"key":"e_1_3_2_4_2","first-page":"412","volume-title":"Proceedings of the 14th International Conference on Machine Learning","author":"Yang Y.","year":"1997","unstructured":"YangY. and PedersenJ., A comparative study on feature set selection in text categorization, In FisherD.H. (Ed.), Proceedings of the 14th International Conference on Machine Learning, San Francisco, CA: Morgan Kaufmann, 1997, pp. 412\u2013420."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2006.04.001"},{"issue":"321","key":"e_1_3_2_6_2","first-page":"1","article-title":"Association and estimation in contingency tables","volume":"63","author":"Mosteller F.","year":"1986","unstructured":"MostellerF., Association and estimation in contingency tables, Journal of the American Statistical Association (American Statistical Association)63(321) (1986), 1\u201328.","journal-title":"Journal of the American Statistical Association (American Statistical Association)"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1002\/asi.21023"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2011.09.160"},{"issue":"1","key":"e_1_3_2_9_2","first-page":"1482","article-title":"Feature selection based on term frequency and T-test for text categorization [J]","volume":"45","author":"Wang D.","year":"2013","unstructured":"WangD., ZhangH., LiuR., et al., Feature selection based on term frequency and T-test for text categorization [J], Pattern Recognition Letters45(1) (2013), 1482\u20131486.","journal-title":"Pattern Recognition Letters"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.3233\/IFS-141240"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12530-015-9131-7"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2016.03.041"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.csda.2004.03.005"},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","unstructured":"KruskalJ.B. and WishM. Multidimensional scaling [M] Sage 1978.","DOI":"10.4135\/9781412985130"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2016.02.001"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2015.06.016"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2005.66"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/505282.505283"},{"key":"e_1_3_2_19_2","first-page":"412","article-title":"A comparative study on feature selection in text categorization [C]","author":"Yang Y.","year":"1997","unstructured":"YangY. and Pedersen.J.O., A comparative study on feature selection in text categorization [C], in: Proceedings of the Fourteenth International Conference on Machine Learning, 1997, pp. 412\u2013420.","journal-title":"in: Proceedings of the Fourteenth International Conference on Machine Learning"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2005.159"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2011.12.003"},{"issue":"5","key":"e_1_3_2_22_2","first-page":"2829","article-title":"Artificial intelligence: A modern approach[J]","volume":"263","author":"Norvig P.","year":"2003","unstructured":"NorvigP. and RussellS.J., Artificial intelligence: A modern approach[J], Applied Mechanics & Materials263(5) (2003), 2829\u20132833.","journal-title":"Applied Mechanics & Materials"},{"key":"e_1_3_2_23_2","first-page":"43","article-title":"A cluster based hybrid feature selection approach[C]","author":"Jaskowiak P.A.","year":"2015","unstructured":"JaskowiakP.A. and CampelloR.J.G.B., A cluster based hybrid feature selection approach[C], Brazilian Conference on Intelligent Systems IEEE, 2015, pp. 43\u201348.","journal-title":"Brazilian Conference on Intelligent Systems IEEE"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/72.298224"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2008.2005601"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2015.02.031"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2015.06.010"},{"key":"e_1_3_2_28_2","doi-asserted-by":"crossref","unstructured":"HuangJ. CaiY. and Xu.X. A hybrid genetic algorithm for feature selection wrapper based on mutual information [J] 28(13) (2007) 1825\u20131844.","DOI":"10.1016\/j.patrec.2007.05.011"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2011.12.005"},{"issue":"4","key":"e_1_3_2_30_2","first-page":"7794","article-title":"Renyi entropy, mutual information, and fluctuation properties of Fermi liquids [J]","volume":"86","author":"Swingle B.","year":"2010","unstructured":"SwingleB., Renyi entropy, mutual information, and fluctuation properties of Fermi liquids [J], Physical Review B Condensed Matter86(4) (2010), 7794\u20137794.","journal-title":"Physical Review B Condensed Matter"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2014.04.019"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2015.06.016"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2011.07.001"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2015.09.006"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2014.02.021"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.14257\/ijdta.2016.9.3.21"},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","unstructured":"PorterM.F. An algorithm for suffix stripping [M]. Readings in information retrieval. Morgan Kaufmann Publishers Inc. 1997 pp. 130\u2013137.","DOI":"10.1108\/eb046814"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13721-015-0078-1"},{"key":"e_1_3_2_39_2","first-page":"307","article-title":"A comparison of event models for naive Bayes spam filtering [C]","volume":"1","author":"McCallum A.","unstructured":"McCallumA. and NigamK., A comparison of event models for naive Bayes spam filtering [C], EACL \u201903 Proceedings of the Tenth Conference on European Chapter of the Association for Computational Linguistics, Volume 1, pp.\u00a0307\u2013314.","journal-title":"EACL \u201903 Proceedings of the Tenth Conference on European Chapter of the Association for Computational Linguistics"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00362-012-0443-4"},{"key":"e_1_3_2_41_2","unstructured":"CorderG.W. and ForemanD.I. Comparing Two Related Samples: The Wilcoxon Signed Ranks Test [M] Nonparametric Statistics for Non-Statisticians: A Step-by-Step Approach John Wiley & Sons Inc 2011 pp. 38\u201356."}],"container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/JIFS-161541","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.3233\/JIFS-161541","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/JIFS-161541","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:40:19Z","timestamp":1777455619000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.3233\/JIFS-161541"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,10]]},"references-count":40,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2017,10]]}},"alternative-id":["10.3233\/JIFS-161541"],"URL":"https:\/\/doi.org\/10.3233\/jifs-161541","relation":{},"ISSN":["1064-1246","1875-8967"],"issn-type":[{"value":"1064-1246","type":"print"},{"value":"1875-8967","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,10]]}}}