{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T14:45:21Z","timestamp":1740149121877,"version":"3.37.3"},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2022,4,5]],"date-time":"2022-04-05T00:00:00Z","timestamp":1649116800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,4,5]],"date-time":"2022-04-05T00:00:00Z","timestamp":1649116800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61972134"],"award-info":[{"award-number":["61972134"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["11601129"],"award-info":[{"award-number":["11601129"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Intell Syst"],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Attribute reduction is an important issue in rough set theory. However, the rough set theory-based attribute reduction algorithms need to be improved to deal with high-dimensional data. A distributed version of the attribute reduction algorithm is necessary to enable it to effectively handle big data. The partition of attribute space is an important research direction. In this paper, a distributed attribution reduction algorithm based on cosine similarity (DARCS) for high-dimensional data pre-processing under the Spark framework is proposed. First, to avoid the repeated calculation of similar attributes, the algorithm gathers similar attributes based on similarity measure to form multiple clusters. And then one attribute is selected randomly as a representative from each cluster to form a candidate attribute subset to participate in the subsequent reduction operation. At the same time, to improve computing efficiency, an improved method is introduced to calculate the attribute dependency in the divided sub-attribute space. Experiments on eight datasets show that, on the premise of avoiding critical information loss, the reduction ability and computing efficiency of DARCS have been improved by 0.32 to 39.61% and 31.32 to 93.79% respectively compared to the distributed version of attribute reduction algorithm based on a random partitioning of the attributes space.<\/jats:p>","DOI":"10.1007\/s44196-022-00076-7","type":"journal-article","created":{"date-parts":[[2022,4,5]],"date-time":"2022-04-05T13:10:16Z","timestamp":1649164216000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["A Distributed Attribute Reduction Algorithm for High-Dimensional Data under the Spark Framework"],"prefix":"10.1007","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5965-4661","authenticated-orcid":false,"given":"Zhengjiang","family":"Wu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiuyu","family":"Mei","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yaning","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tian","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junwei","family":"Luo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,4,5]]},"reference":[{"issue":"2","key":"76_CR1","doi-asserted-by":"publisher","first-page":"171","DOI":"10.1007\/s11036-013-0489-0","volume":"19","author":"M Chen","year":"2014","unstructured":"Chen, M., Mao, S., Liu, Y.: Big data: a survey. Mob. Netw. Appl. 19(2), 171\u2013209 (2014)","journal-title":"Mob. Netw. Appl."},{"key":"76_CR2","doi-asserted-by":"crossref","unstructured":"Li, T., Luo, C., Chen, H., Zhang, J.: Pickt: a solution for big data analysis. In: International Conference on Rough Sets and Knowledge Technology, pp. 15\u201325 (2015). Springer","DOI":"10.1007\/978-3-319-25754-9_2"},{"issue":"3","key":"76_CR3","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1007\/s00530-015-0494-1","volume":"23","author":"L Gao","year":"2017","unstructured":"Gao, L., Song, J., Liu, X., Shao, J., Liu, J., Shao, J.: Learning in high-dimensional multimedia data: the state of the art. Multimedia Syst. 23(3), 303\u2013313 (2017)","journal-title":"Multimedia Syst."},{"issue":"1","key":"76_CR4","first-page":"97","volume":"26","author":"X Wu","year":"2013","unstructured":"Wu, X., Zhu, X., Wu, G.-Q., Ding, W.: Data mining with big data. IEEE Trans. Knowl. Data Eng. 26(1), 97\u2013107 (2013)","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"76_CR5","doi-asserted-by":"crossref","unstructured":"Anderson, M., Cafarella, M.: Input selection for fast feature engineering. In: IEEE International Conference on Data Engineering, pp. 577\u2013588 (2016)","DOI":"10.1109\/ICDE.2016.7498272"},{"issue":"66\u201371","key":"76_CR6","first-page":"13","volume":"10","author":"L Van Der Maaten","year":"2009","unstructured":"Van Der Maaten, L., Postma, E., Van den Herik, J., et al.: Dimensionality reduction: a comparative. J Mach Learn Res 10(66\u201371), 13 (2009)","journal-title":"J Mach Learn Res"},{"issue":"6","key":"76_CR7","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.camwa.2020.12.002","volume":"40","author":"S Xu","year":"2021","unstructured":"Xu, S., Li, S., Liu, H., Garg, H., Jin, X., Zhao, J.: An understandable way to discover methods to model interval input-output samples. Comp. Appl. Math 40(6), 1\u201321 (2021)","journal-title":"Comp. Appl. Math"},{"issue":"5","key":"76_CR8","doi-asserted-by":"publisher","first-page":"341","DOI":"10.1007\/BF01001956","volume":"11","author":"Z Pawlak","year":"1982","unstructured":"Pawlak, Z.: Rough sets. Int. J. Comp. Inf. Sci. 11(5), 341\u2013356 (1982)","journal-title":"Int. J. Comp. Inf. Sci."},{"key":"76_CR9","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.knosys.2013.04.001","volume":"49","author":"Y-C Ko","year":"2013","unstructured":"Ko, Y.-C., Fujita, H., Tzeng, G.-H.: A fuzzy integral fusion approach in analyzing competitiveness patterns from wcy2010. Knowl-Based Syst. 49, 1\u20139 (2013)","journal-title":"Knowl-Based Syst."},{"issue":"1","key":"76_CR10","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1016\/j.ins.2006.06.003","volume":"177","author":"Z Pawlak","year":"2007","unstructured":"Pawlak, Z., Skowron, A.: Rudiments of rough sets. Inform. Sci. 177(1), 3\u201327 (2007)","journal-title":"Inform. Sci."},{"issue":"4","key":"76_CR11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s40314-021-01507-5","volume":"40","author":"H Garg","year":"2021","unstructured":"Garg, H., Rizk-Allah, R.M.: A novel approach for solving rough multi-objective transportation problem: development and prospects. Comp. Appl. Math 40(4), 1\u201324 (2021)","journal-title":"Comp. Appl. Math"},{"issue":"9\u201310","key":"76_CR12","doi-asserted-by":"publisher","first-page":"597","DOI":"10.1016\/j.artint.2010.04.018","volume":"174","author":"Y Qian","year":"2010","unstructured":"Qian, Y., Liang, J., Pedrycz, W., Dang, C.: Positive approximation: an accelerator for attribute reduction in rough set theory. Artif. Intell. 174(9\u201310), 597\u2013618 (2010)","journal-title":"Artif. Intell."},{"key":"76_CR13","doi-asserted-by":"publisher","first-page":"461","DOI":"10.1016\/j.ins.2016.09.018","volume":"373","author":"Y Zhang","year":"2016","unstructured":"Zhang, Y., Li, T., Luo, C., Zhang, J., Chen, H.: Incremental updating of rough approximations in interval-valued information systems under attribute generalization. Inform. Sci. 373, 461\u2013475 (2016)","journal-title":"Inform. Sci."},{"key":"76_CR14","doi-asserted-by":"publisher","first-page":"175","DOI":"10.1016\/j.ijar.2017.10.012","volume":"92","author":"MS Raza","year":"2018","unstructured":"Raza, M.S., Qamar, U.: Feature selection using rough set-based direct dependency calculation by avoiding the positive region. Int. J. Approx. Reason. 92, 175\u2013197 (2018)","journal-title":"Int. J. Approx. Reason."},{"issue":"1","key":"76_CR15","doi-asserted-by":"publisher","first-page":"1473","DOI":"10.2991\/ijcis.d.200915.004","volume":"13","author":"Y Gao","year":"2020","unstructured":"Gao, Y., Lv, C., Wu, Z.: Attribute reduction of boolean matrix in neighborhood rough set model. Int. J. Comput. Int. Sys. 13(1), 1473\u20131482 (2020)","journal-title":"Int. J. Comput. Int. Sys."},{"key":"76_CR16","doi-asserted-by":"publisher","first-page":"64","DOI":"10.1016\/j.ins.2020.05.010","volume":"535","author":"Y Chen","year":"2020","unstructured":"Chen, Y., Liu, K., Song, J., Fujita, H., Yang, X., Qian, Y.: Attribute group for attribute reduction. Inform. Sci. 535, 64\u201380 (2020)","journal-title":"Inform. Sci."},{"key":"76_CR17","doi-asserted-by":"publisher","first-page":"351","DOI":"10.1016\/j.ins.2016.09.012","volume":"373","author":"H Chen","year":"2016","unstructured":"Chen, H., Li, T., Cai, Y., Luo, C., Fujita, H.: Parallel attribute reduction in dominance-based neighborhood rough set. Inform. Sci. 373, 351\u2013368 (2016)","journal-title":"Inform. Sci."},{"key":"76_CR18","doi-asserted-by":"publisher","first-page":"671","DOI":"10.1016\/j.ins.2014.04.019","volume":"279","author":"J Qian","year":"2014","unstructured":"Qian, J., Miao, D., Zhang, Z., Yue, X.: Parallel attribute reduction algorithms using mapreduce. Inform. Sci. 279, 671\u2013690 (2014)","journal-title":"Inform. Sci."},{"key":"76_CR19","doi-asserted-by":"publisher","first-page":"18","DOI":"10.1016\/j.simpat.2016.01.010","volume":"64","author":"E-SM El-Alfy","year":"2016","unstructured":"El-Alfy, E.-S.M., Alshammari, M.A.: Towards scalable rough set based attribute subset selection for intrusion detection using parallel genetic algorithm in mapreduce. Simul. Model. Pract. Ther. 64, 18\u201329 (2016)","journal-title":"Simul. Model. Pract. Ther."},{"issue":"1","key":"76_CR20","doi-asserted-by":"publisher","first-page":"226","DOI":"10.1109\/TFUZZ.2017.2647966","volume":"26","author":"Q Hu","year":"2018","unstructured":"Hu, Q., Zhang, L., Zhou, Y., Pedrycz, W.: Large-scale multimodality attribute reduction with multi-kernel fuzzy rough sets. Trans. Fuz. Sys. 26(1), 226\u2013238 (2018)","journal-title":"Trans. Fuz. Sys."},{"issue":"11","key":"76_CR21","first-page":"6","volume":"43","author":"JB Xia","year":"2016","unstructured":"Xia, J.B., Wei, Z., Fu, K., Chen, Z.: Review of research and application on hadoop in cloud computing. Comput. Sci. 43(11), 6\u201311 (2016)","journal-title":"Comput. Sci."},{"key":"76_CR22","doi-asserted-by":"crossref","unstructured":"Shanahan, J.G., Dai, L.: Large scale distributed data science using apache spark. In: Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 2323\u20132324 (2015)","DOI":"10.1145\/2783258.2789993"},{"issue":"1","key":"76_CR23","doi-asserted-by":"publisher","first-page":"389","DOI":"10.1007\/s11192-016-1945-y","volume":"109","author":"IAT Hashem","year":"2016","unstructured":"Hashem, I.A.T., Anuar, N.B., Gani, A., Yaqoob, I., Xia, F., Khan, S.U.: Mapreduce: review and open challenges. Scientometrics 109(1), 389\u2013422 (2016)","journal-title":"Scientometrics"},{"issue":"2","key":"76_CR24","first-page":"393","volume":"21","author":"J Wang","year":"2020","unstructured":"Wang, J., Yang, Y., Wang, T., Sherratt, R.S., Zhang, J.: Big data service architecture: a survey. J. Int. Technol. 21(2), 393\u2013405 (2020)","journal-title":"J. Int. Technol."},{"issue":"10\u201310","key":"76_CR25","first-page":"95","volume":"10","author":"M Zaharia","year":"2010","unstructured":"Zaharia, M., Chowdhury, M., Franklin, M.J., Shenker, S., Stoica, I., et al.: Spark: cluster computing with working sets. HotCloud 10(10\u201310), 95 (2010)","journal-title":"HotCloud"},{"key":"76_CR26","unstructured":"Zhang, J., Li, T., Pan, Y.: Parallel large-scale attribute reduction on cloud systems. arXiv preprint arXiv:1610.01807 (2016)"},{"key":"76_CR27","doi-asserted-by":"crossref","unstructured":"Chen, M., Yuan, J., Li, L., Liu, D., Li, T.: A fast heuristic attribute reduction algorithm using spark. In: 2017 IEEE 37th International Conference on Distributed Computing Systems (ICDCS), pp. 2393\u20132398 (2017). IEEE","DOI":"10.1109\/ICDCS.2017.38"},{"issue":"9","key":"76_CR28","doi-asserted-by":"publisher","first-page":"1441","DOI":"10.1109\/TSMC.2017.2670926","volume":"48","author":"S Ram\u00edrez-Gallego","year":"2017","unstructured":"Ram\u00edrez-Gallego, S., Mouri\u00f1o-Tal\u00edn, H., Mart\u00ednez-Rego, D., Bol\u00f3n-Canedo, V., Ben\u00edtez, J.M., Alonso-Betanzos, A., Herrera, F.: An information theory-based feature selection framework for big data under apache spark. IEEE Trans. Syst. Man Cybern. 48(9), 1441\u20131453 (2017)","journal-title":"IEEE Trans. Syst. Man Cybern."},{"issue":"8","key":"76_CR29","doi-asserted-by":"publisher","first-page":"3321","DOI":"10.1007\/s10115-020-01467-y","volume":"62","author":"ZC Dagdia","year":"2020","unstructured":"Dagdia, Z.C., Zarges, C., Beck, G., Lebbah, M.: A scalable and effective rough set theory-based approach for big data pre-processing. Knowl. Inf. Syst. 62(8), 3321\u20133386 (2020)","journal-title":"Knowl. Inf. Syst."},{"key":"76_CR30","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1016\/j.knosys.2015.01.004","volume":"80","author":"Y Yao","year":"2015","unstructured":"Yao, Y.: The two sides of the theory of rough sets. Knowl.-Based Syst. 80, 67\u201377 (2015)","journal-title":"Knowl.-Based Syst."},{"issue":"1","key":"76_CR31","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1016\/j.ins.2006.06.007","volume":"177","author":"Z Pawlak","year":"2007","unstructured":"Pawlak, Z., Skowron, A.: Rough sets and boolean reasoning. Inform. Sci. 177(1), 41\u201373 (2007)","journal-title":"Inform. Sci."},{"key":"76_CR32","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2020.113400","volume":"154","author":"SP Patel","year":"2020","unstructured":"Patel, S.P., Upadhyay, S.H.: Euclidean distance based feature ranking and subset selection for bearing fault diagnosis. Expert Syst. Appl. 154, 113400 (2020)","journal-title":"Expert Syst. Appl."},{"key":"76_CR33","doi-asserted-by":"publisher","first-page":"39","DOI":"10.1016\/j.ins.2015.02.024","volume":"307","author":"P Xia","year":"2015","unstructured":"Xia, P., Zhang, L., Li, F.: Learning similarity with cosine similarity ensemble. Inform. Sci. 307, 39\u201352 (2015)","journal-title":"Inform. Sci."},{"key":"76_CR34","doi-asserted-by":"publisher","first-page":"11406114066","DOI":"10.1016\/j.eswa.2020.114066","volume":"166","author":"BI Kwak","year":"2021","unstructured":"Kwak, B.I., Han, M.L., Kim, H.K.: Cosine similarity based anomaly detection methodology for the can bus. Expert Syst. Appl. 166, 11406114066 (2021)","journal-title":"Expert Syst. Appl."},{"key":"76_CR35","doi-asserted-by":"publisher","first-page":"68","DOI":"10.1016\/j.patrec.2021.04.029","volume":"148","author":"J Chen","year":"2021","unstructured":"Chen, J., Guo, Z., Hu, J.: Ring-regularized cosine similarity learning for fine-grained face verification. Pattern Recogn. Lett. 148, 68\u201374 (2021)","journal-title":"Pattern Recogn. Lett."},{"key":"76_CR36","doi-asserted-by":"publisher","first-page":"101735","DOI":"10.1016\/j.artmed.2019.101735","volume":"101","author":"M Abdel-Basset","year":"2019","unstructured":"Abdel-Basset, M., Mohamed, M., Elhoseny, M., Chiclana, F., Zaied, A.E.-N.H., et al.: Cosine similarity measures of bipolar neutrosophic set for diagnosis of bipolar disorder diseases. Artif. Intell. Med. 101, 101735 (2019)","journal-title":"Artif. Intell. Med."},{"key":"76_CR37","doi-asserted-by":"publisher","first-page":"115224","DOI":"10.1016\/j.eswa.2021.115224","volume":"182","author":"A Hashemi","year":"2021","unstructured":"Hashemi, A., Dowlatshahi, M.B., Nezamabadi-pour, H.: Vmfs: a vikor-based multi-target feature selection. Expert Syst. Appl. 182, 115224 (2021)","journal-title":"Expert Syst. Appl."},{"key":"76_CR38","doi-asserted-by":"publisher","first-page":"106839","DOI":"10.1016\/j.csda.2019.106839","volume":"143","author":"A Bommert","year":"2020","unstructured":"Bommert, A., Sun, X., Bischl, B., Rahnenf\u00fchrer, J., Lang, M.: Benchmark for filter methods for feature selection in high-dimensional classification data. Comput. Stat. Data Anal. 143, 106839 (2020)","journal-title":"Comput. Stat. Data Anal."},{"issue":"2","key":"76_CR39","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1145\/2641190.2641198","volume":"15","author":"J Vanschoren","year":"2014","unstructured":"Vanschoren, J., Van Rijn, J.N., Bischl, B., Torgo, L.: Openml: networked science in machine learning. ACM SIGKDD Explor. Newslett. 15(2), 49\u201360 (2014)","journal-title":"ACM SIGKDD Explor. Newslett."},{"key":"76_CR40","unstructured":"Dua, D., Graff, C.: UCI Machine Learning Repository (2017). http:\/\/archive.ics.uci.edu\/ml"}],"container-title":["International Journal of Computational Intelligence Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44196-022-00076-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44196-022-00076-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44196-022-00076-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,4,5]],"date-time":"2022-04-05T13:44:47Z","timestamp":1649166287000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44196-022-00076-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,5]]},"references-count":40,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["76"],"URL":"https:\/\/doi.org\/10.1007\/s44196-022-00076-7","relation":{},"ISSN":["1875-6883"],"issn-type":[{"type":"electronic","value":"1875-6883"}],"subject":[],"published":{"date-parts":[[2022,4,5]]},"assertion":[{"value":"14 October 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 March 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 April 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that there is no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}],"article-number":"22"}}