{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T12:02:13Z","timestamp":1781784133033,"version":"3.54.5"},"reference-count":46,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2022,5,3]],"date-time":"2022-05-03T00:00:00Z","timestamp":1651536000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,5,3]],"date-time":"2022-05-03T00:00:00Z","timestamp":1651536000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Big Data"],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Dimension reduction is a preprocessing step in machine learning for eliminating undesirable features and increasing learning accuracy. In order to reduce the redundant features, there are data representation methods, each of which has its own advantages. On the other hand, big data with imbalanced classes is one of the most important issues in pattern recognition and machine learning. In this paper, a method is proposed in the form of a cost-sensitive optimization problem which implements the process of selecting and extracting the features simultaneously. The feature extraction phase is based on reducing error and maintaining geometric relationships between data by solving a manifold learning optimization problem. In the feature selection phase, the cost-sensitive optimization problem is adopted based on minimizing the upper limit of the generalization error. Finally, the optimization problem which is constituted from the above two problems is solved by adding a cost-sensitive term to create a balance between classes without manipulating the data. To evaluate the results of the feature reduction, the multi-class linear SVM classifier is used on the reduced data. The proposed method is compared with some other approaches on 21 datasets from the UCI learning repository, microarrays and high-dimensional datasets, as well as imbalanced datasets from the KEEL repository. The results indicate the significant efficiency of the proposed method compared to some similar approaches.<\/jats:p>","DOI":"10.1186\/s40537-022-00617-z","type":"journal-article","created":{"date-parts":[[2022,5,3]],"date-time":"2022-05-03T11:21:54Z","timestamp":1651576914000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":21,"title":["Improved cost-sensitive representation of data for solving the imbalanced big data classification problem"],"prefix":"10.1186","volume":"9","author":[{"given":"Mahboubeh","family":"Fattahi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8968-6744","authenticated-orcid":false,"given":"Mohammad Hossein","family":"Moattar","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yahya","family":"Forghani","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,5,3]]},"reference":[{"key":"617_CR1","doi-asserted-by":"publisher","first-page":"292","DOI":"10.1016\/j.compbiomed.2015.01.022","volume":"64","author":"S Rakkeitwinai","year":"2015","unstructured":"Rakkeitwinai S, et al. New feature selection for gene expression classification based on degree of class overlap in principal dimensions. Comput Biol Med. 2015;64:292\u20138.","journal-title":"Comput Biol Med"},{"issue":"17","key":"617_CR2","doi-asserted-by":"publisher","first-page":"2914","DOI":"10.1016\/j.neucom.2011.03.034","volume":"74","author":"MM Kabir","year":"2011","unstructured":"Kabir MM, Shahjahan M, Murase K. A new local search based hybrid genetic algorithm for feature selection. Neurocomputing. 2011;74(17):2914\u201328.","journal-title":"Neurocomputing"},{"issue":"4","key":"617_CR3","doi-asserted-by":"publisher","first-page":"2714","DOI":"10.1016\/j.eswa.2009.08.026","volume":"37","author":"SM Vieira","year":"2010","unstructured":"Vieira SM, Sousa JM, Runkler TA. Two cooperative ant colonies for feature selection using fuzzy models. Expert Syst Appl. 2010;37(4):2714\u201323.","journal-title":"Expert Syst Appl"},{"issue":"2","key":"617_CR4","doi-asserted-by":"publisher","first-page":"56","DOI":"10.38094\/jastt1224","volume":"1","author":"R Zebari","year":"2020","unstructured":"Zebari R, et al. A comprehensive review of dimensionality reduction techniques for feature selection and feature extraction. J Appl Sci Technol Trends. 2020;1(2):56\u201370.","journal-title":"J Appl Sci Technol Trends"},{"key":"617_CR5","doi-asserted-by":"publisher","DOI":"10.1155\/2018\/2879640","author":"Z Cheng","year":"2018","unstructured":"Cheng Z, Lu Z. A novel efficient feature dimensionality reduction method and its application in engineering. Complexity. 2018. https:\/\/doi.org\/10.1155\/2018\/2879640.","journal-title":"Complexity"},{"key":"617_CR6","doi-asserted-by":"crossref","unstructured":"Zebari DA, et al. A simultaneous approach for compression and encryption techniques using deoxyribonucleic acid. In: 2019 13th international conference on software, knowledge, information management and applications (SKIMA). IEEE; 2019.","DOI":"10.1109\/SKIMA47702.2019.8982392"},{"key":"617_CR7","doi-asserted-by":"publisher","first-page":"44","DOI":"10.1016\/j.inffus.2020.01.005","volume":"59","author":"S Ayesha","year":"2020","unstructured":"Ayesha S, Hanif MK, Talib R. Overview and comparative study of dimensionality reduction techniques for high dimensional data. Inf Fusion. 2020;59:44\u201358.","journal-title":"Inf Fusion"},{"issue":"5","key":"617_CR8","doi-asserted-by":"publisher","first-page":"571","DOI":"10.17706\/jcp.13.5.571-579","volume":"13","author":"N Abd-Alsabour","year":"2018","unstructured":"Abd-Alsabour N. On the role of dimensionality reduction. J Comput. 2018;13(5):571\u20139.","journal-title":"J Comput"},{"key":"617_CR9","doi-asserted-by":"crossref","unstructured":"Verleysen M, Fran\u00e7ois D. The curse of dimensionality in data mining and time series prediction. In: International work-conference on artificial neural networks. Springer; 2005.","DOI":"10.1007\/11494669_93"},{"key":"617_CR10","unstructured":"Peleg D, Meir R. A feature selection algorithm based on the global minimization of a generalization error bound. In: Advances in neural information processing systems. 2005."},{"issue":"4","key":"617_CR11","doi-asserted-by":"publisher","first-page":"44","DOI":"10.4018\/IJSI.2017100104","volume":"5","author":"MK Elhadad","year":"2017","unstructured":"Elhadad MK, Badran KM, Salama GI. A novel approach for ontology-based dimensionality reduction for web text document classification. Int J Softw Innov. 2017;5(4):44\u201358.","journal-title":"Int J Softw Innov"},{"key":"617_CR12","unstructured":"Luo W. Face recognition based on laplacian eigenmaps. In: 2011 International conference on computer science and service system (CSSS). IEEE; 2011."},{"key":"617_CR13","unstructured":"Abdullah A, et al. Sketching, embedding and dimensionality reduction in information theoretic spaces. In: Artificial intelligence and statistics. PMLR; 2016."},{"key":"617_CR14","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2019.105989","volume":"87","author":"Y Wang","year":"2020","unstructured":"Wang Y, Li T. Local feature selection based on artificial immune system for classification. Appl Soft Comput. 2020;87: 105989.","journal-title":"Appl Soft Comput"},{"key":"617_CR15","doi-asserted-by":"publisher","first-page":"154","DOI":"10.1016\/j.patcog.2018.01.012","volume":"78","author":"Y Zhao","year":"2018","unstructured":"Zhao Y, et al. Multi-view manifold learning with locality alignment. Pattern Recogn. 2018;78:154\u201366.","journal-title":"Pattern Recogn"},{"key":"617_CR16","doi-asserted-by":"crossref","unstructured":"Xu J, et al. Feature selection based on sparse imputation. In: The 2012 international joint conference on neural networks (IJCNN). IEEE; 2012.","DOI":"10.1109\/IJCNN.2012.6252639"},{"issue":"3","key":"617_CR17","doi-asserted-by":"publisher","first-page":"717","DOI":"10.1007\/s10489-019-01543-z","volume":"50","author":"SA Shahee","year":"2020","unstructured":"Shahee SA, Ananthakumar U. An effective distance based feature selection approach for imbalanced data. Appl Intell. 2020;50(3):717\u201345.","journal-title":"Appl Intell"},{"key":"617_CR18","doi-asserted-by":"publisher","first-page":"280","DOI":"10.1016\/j.patrec.2020.03.016","volume":"133","author":"H Chenxi","year":"2020","unstructured":"Chenxi H, et al. Sample imbalance disease classification model based on association rule feature selection. Pattern Recognit Lett. 2020;133:280\u20136.","journal-title":"Pattern Recognit Lett"},{"issue":"6","key":"617_CR19","doi-asserted-by":"publisher","first-page":"534","DOI":"10.1109\/TSE.2017.2731766","volume":"44","author":"KE Bennin","year":"2017","unstructured":"Bennin KE, et al. Mahakil: diversity based oversampling approach to alleviate the class imbalance issue in software defect prediction. IEEE Trans Softw Eng. 2017;44(6):534\u201350.","journal-title":"IEEE Trans Softw Eng"},{"key":"617_CR20","doi-asserted-by":"publisher","first-page":"59","DOI":"10.1016\/j.knosys.2018.01.002","volume":"145","author":"S Nakariyakul","year":"2018","unstructured":"Nakariyakul S. High-dimensional hybrid feature selection using interaction information-guided search. Knowl Based Syst. 2018;145:59\u201366.","journal-title":"Knowl Based Syst"},{"issue":"8","key":"617_CR21","doi-asserted-by":"publisher","first-page":"2656","DOI":"10.1016\/j.patcog.2015.02.025","volume":"48","author":"Z Zeng","year":"2015","unstructured":"Zeng Z, et al. A novel feature selection method considering feature interaction. Pattern Recogn. 2015;48(8):2656\u201366.","journal-title":"Pattern Recogn"},{"key":"617_CR22","doi-asserted-by":"crossref","unstructured":"Qi X, et al. WJMI: a new feature selection algorithm based on weighted joint mutual information. In: 2015 3rd international conference on mechatronics and industrial informatics (ICMII 2015). Atlantis Press; 2015.","DOI":"10.2991\/icmii-15.2015.108"},{"key":"617_CR23","unstructured":"Japkowicz N. The class imbalance problem: significance and strategies. In: Proc. of the Int\u2019l Conf. on artificial intelligence. 2000. Citeseer."},{"issue":"3","key":"617_CR24","doi-asserted-by":"publisher","first-page":"515","DOI":"10.1109\/TIT.1968.1054155","volume":"14","author":"P Hart","year":"1968","unstructured":"Hart P. The condensed nearest neighbor rule (corresp.). IEEE Trans Inf Theory. 1968;14(3):515\u20136.","journal-title":"IEEE Trans Inf Theory"},{"key":"617_CR25","first-page":"769","volume":"6","author":"I Tomek","year":"1976","unstructured":"Tomek I. Two modifications of CNN. IEEE Trans Syst Man Cybern. 1976;6:769\u201372.","journal-title":"IEEE Trans Syst Man Cybern."},{"issue":"3","key":"617_CR26","doi-asserted-by":"publisher","first-page":"5718","DOI":"10.1016\/j.eswa.2008.06.108","volume":"36","author":"S-J Yen","year":"2009","unstructured":"Yen S-J, Lee Y-S. Cluster-based under-sampling approaches for imbalanced data distributions. Expert Syst Appl. 2009;36(3):5718\u201327.","journal-title":"Expert Syst Appl"},{"issue":"3","key":"617_CR27","doi-asserted-by":"publisher","first-page":"275","DOI":"10.1162\/evco.2009.17.3.275","volume":"17","author":"S Garc\u00eda","year":"2009","unstructured":"Garc\u00eda S, Herrera F. Evolutionary undersampling for classification with imbalanced datasets: proposals and taxonomy. Evol Comput. 2009;17(3):275\u2013306.","journal-title":"Evol Comput"},{"key":"617_CR28","doi-asserted-by":"publisher","first-page":"321","DOI":"10.1613\/jair.953","volume":"16","author":"NV Chawla","year":"2002","unstructured":"Chawla NV, et al. SMOTE: synthetic minority over-sampling technique. J Artif Intell Res. 2002;16:321\u201357.","journal-title":"J Artif Intell Res"},{"key":"617_CR29","doi-asserted-by":"crossref","unstructured":"Han H, Wang W-Y, Mao B-H. Borderline-SMOTE: a new over-sampling method in imbalanced data sets learning. In: International conference on intelligent computing. Springer; 2005.","DOI":"10.1007\/11538059_91"},{"key":"617_CR30","doi-asserted-by":"crossref","unstructured":"Maciejewski T, Stefanowski J. Local neighbourhood extension of SMOTE for mining imbalanced data. In: 2011 IEEE symposium on computational intelligence and data mining (CIDM). IEEE; 2011.","DOI":"10.1109\/CIDM.2011.5949434"},{"key":"617_CR31","doi-asserted-by":"crossref","unstructured":"Bunkhumpornpat C, Sinapiromsaran K, Lursinsap C. Safe-level-smote: safe-level-synthetic minority over-sampling technique for handling the class imbalanced problem. In: Pacific-Asia conference on knowledge discovery and data mining. Springer; 2009.","DOI":"10.1007\/978-3-642-01307-2_43"},{"issue":"2","key":"617_CR32","doi-asserted-by":"publisher","first-page":"245","DOI":"10.1007\/s10115-011-0465-6","volume":"33","author":"E Ramentol","year":"2012","unstructured":"Ramentol E, et al. SMOTE-RS B*: a hybrid preprocessing approach based on oversampling and undersampling for high imbalanced data-sets using SMOTE and rough sets theory. Knowl Inf Syst. 2012;33(2):245\u201365.","journal-title":"Knowl Inf Syst"},{"key":"617_CR33","doi-asserted-by":"publisher","first-page":"45","DOI":"10.1016\/j.neucom.2016.10.053","volume":"224","author":"F Cheng","year":"2017","unstructured":"Cheng F, et al. Large cost-sensitive margin distribution machine for imbalanced data classification. Neurocomputing. 2017;224:45\u201357.","journal-title":"Neurocomputing"},{"key":"617_CR34","doi-asserted-by":"publisher","first-page":"70","DOI":"10.1016\/j.neucom.2016.09.120","volume":"261","author":"W Xiao","year":"2017","unstructured":"Xiao W, et al. Class-specific cost regulation extreme learning machine for imbalanced classification. Neurocomputing. 2017;261:70\u201382.","journal-title":"Neurocomputing"},{"key":"617_CR35","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2020.106020","volume":"200","author":"G Du","year":"2020","unstructured":"Du G, et al. Joint imbalanced classification and feature selection for hospital readmissions. Knowl Based Syst. 2020;200: 106020.","journal-title":"Knowl Based Syst"},{"key":"617_CR36","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2019.06.022","volume":"187","author":"BS Raghuwanshi","year":"2020","unstructured":"Raghuwanshi BS, Shukla S. SMOTE based class-specific extreme learning machine for imbalanced learning. Knowl Based Syst. 2020;187: 104814.","journal-title":"Knowl Based Syst"},{"key":"617_CR37","doi-asserted-by":"publisher","first-page":"214","DOI":"10.1016\/j.ins.2020.02.070","volume":"522","author":"H Yuan","year":"2020","unstructured":"Yuan H, et al. Low-rank matrix regression for image feature extraction and feature selection. Inf Sci. 2020;522:214\u201326.","journal-title":"Inf Sci"},{"key":"617_CR38","first-page":"424","volume":"25","author":"M Buvana","year":"2021","unstructured":"Buvana M, Muthumayil K, Jayasankar T. Content-based image retrieval based on hybrid feature extraction and feature selection technique pigeon inspired based optimization. Ann Roman Soc Cell Biol. 2021;25:424\u201343.","journal-title":"Ann Roman Soc Cell Biol"},{"key":"617_CR39","doi-asserted-by":"crossref","unstructured":"Wang Q. A hybrid sampling SVM approach to imbalanced data classification. In: Abstract and applied analysis. 2014. Hindawi.","DOI":"10.1155\/2014\/972786"},{"key":"617_CR40","doi-asserted-by":"crossref","unstructured":"Prachuabsupakij W. CLUS: a new hybrid sampling classification for imbalanced data. In: 2015 12th international joint conference on computer science and software engineering (JCSSE). IEEE; 2015.","DOI":"10.1109\/JCSSE.2015.7219810"},{"key":"617_CR41","doi-asserted-by":"publisher","first-page":"94","DOI":"10.1016\/j.asoc.2018.02.051","volume":"67","author":"S Maldonado","year":"2018","unstructured":"Maldonado S, L\u00f3pez J. Dealing with high-dimensional class-imbalanced datasets: embedded feature selection for SVM classification. Appl Soft Comput. 2018;67:94\u2013105.","journal-title":"Appl Soft Comput"},{"issue":"1","key":"617_CR42","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-021-00428-8","volume":"8","author":"M Roccetti","year":"2021","unstructured":"Roccetti M, et al. An alternative approach to dimension reduction for pareto distributed data: a case study. J Big Data. 2021;8(1):1\u201323.","journal-title":"J Big Data"},{"issue":"1","key":"617_CR43","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-020-00320-x","volume":"7","author":"S Thudumu","year":"2020","unstructured":"Thudumu S, et al. A comprehensive survey of anomaly detection techniques for high dimensional big data. J Big Data. 2020;7(1):1\u201330.","journal-title":"J Big Data"},{"issue":"1","key":"617_CR44","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-017-0093-4","volume":"4","author":"F Badaoui","year":"2017","unstructured":"Badaoui F, et al. Dimensionality reduction and class prediction algorithm with application to microarray Big Data. J Big Data. 2017;4(1):1\u201311.","journal-title":"J Big Data"},{"key":"617_CR45","doi-asserted-by":"publisher","first-page":"7940","DOI":"10.1109\/ACCESS.2016.2619719","volume":"4","author":"A Amin","year":"2016","unstructured":"Amin A, et al. Comparing oversampling techniques to handle the class imbalance problem: a customer churn prediction case study. IEEE Access. 2016;4:7940\u201357.","journal-title":"IEEE Access"},{"key":"617_CR46","doi-asserted-by":"crossref","unstructured":"Qi X, et al. WJMI: a new feature selection algorithm based on weighted joint mutual information. 2015.","DOI":"10.2991\/icmii-15.2015.108"}],"container-title":["Journal of Big Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-022-00617-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s40537-022-00617-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-022-00617-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,5,4]],"date-time":"2022-05-04T03:58:36Z","timestamp":1651636716000},"score":1,"resource":{"primary":{"URL":"https:\/\/journalofbigdata.springeropen.com\/articles\/10.1186\/s40537-022-00617-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,5,3]]},"references-count":46,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["617"],"URL":"https:\/\/doi.org\/10.1186\/s40537-022-00617-z","relation":{},"ISSN":["2196-1115"],"issn-type":[{"value":"2196-1115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,5,3]]},"assertion":[{"value":"22 November 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 April 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 May 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"This manuscript does not include studies involving human participants.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"This manuscript does not contain any individual person\u2019s data.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors approve that there is no conflict of interest in the manuscript.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"60"}}