{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T04:27:36Z","timestamp":1785472056272,"version":"3.56.0"},"reference-count":53,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2021,7,21]],"date-time":"2021-07-21T00:00:00Z","timestamp":1626825600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100014810","name":"Fondazione di Sardegna","doi-asserted-by":"publisher","award":["ADAM project - CUP: F74I19000900007"],"award-info":[{"award-number":["ADAM project - CUP: F74I19000900007"]}],"id":[{"id":"10.13039\/100014810","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Class imbalance and high dimensionality are two major issues in several real-life applications, e.g., in the fields of bioinformatics, text mining and image classification. However, while both issues have been extensively studied in the machine learning community, they have mostly been treated separately, and little research has been thus far conducted on which approaches might be best suited to deal with datasets that are class-imbalanced and high-dimensional at the same time (i.e., with a large number of features). This work attempts to give a contribution to this challenging research area by studying the effectiveness of hybrid learning strategies that involve the integration of feature selection techniques, to reduce the data dimensionality, with proper methods that cope with the adverse effects of class imbalance (in particular, data balancing and cost-sensitive methods are considered). Extensive experiments have been carried out across datasets from different domains, leveraging a well-known classifier, the Random Forest, which has proven to be effective in high-dimensional spaces and has also been successfully applied to imbalanced tasks. Our results give evidence of the benefits of such a hybrid approach, when compared to using only feature selection or imbalance learning methods alone.<\/jats:p>","DOI":"10.3390\/info12080286","type":"journal-article","created":{"date-parts":[[2021,7,21]],"date-time":"2021-07-21T11:53:23Z","timestamp":1626868403000},"page":"286","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":43,"title":["Learning from High-Dimensional and Class-Imbalanced Datasets Using Random Forests"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3983-6844","authenticated-orcid":false,"given":"Barbara","family":"Pes","sequence":"first","affiliation":[{"name":"Dipartimento di Matematica e Informatica, Universit\u00e0 di Cagliari, Via Ospedale 72, 09124 Cagliari, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,7,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"429","DOI":"10.3233\/IDA-2002-6504","article-title":"The class imbalance problem: A systematic study","volume":"6","author":"Japkowicz","year":"2002","journal-title":"Intell. Data Anal."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1263","DOI":"10.1109\/TKDE.2008.239","article-title":"Learning from imbalanced data","volume":"21","author":"He","year":"2009","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1007\/s13748-016-0094-0","article-title":"Learning from imbalanced data: Open challenges and future directions","volume":"5","author":"Krawczyk","year":"2016","journal-title":"Prog. Artif. Intell."},{"key":"ref_4","first-page":"31","article-title":"A Survey of Predictive Modeling on Imbalanced Domains","volume":"49","author":"Branco","year":"2016","journal-title":"ACM Comput. Surv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Blagus, R., and Lusa, L. (2010). Class prediction for high-dimensional class-imbalanced data. BMC Bioinform., 11.","DOI":"10.1186\/1471-2105-11-523"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"228","DOI":"10.1016\/j.ins.2014.07.015","article-title":"Feature selection for high-dimensional class-imbalanced data sets using Support Vector Machines","volume":"286","author":"Maldonado","year":"2014","journal-title":"Inf. Sci."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1016\/j.engappai.2016.10.008","article-title":"Feature selection for high dimensional imbalanced class data using harmony search","volume":"57","author":"Moayedikia","year":"2017","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_8","unstructured":"Shanab, A.A., and Khoshgoftaar, T.M. (2018, January 6\u20139). Is Gene Selection Enough for Imbalanced Bioinformatics Data?. Proceedings of the 2018 IEEE International Conference on Information Reuse and Integration for Data Science, Salt Lake City, UT, USA."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1765","DOI":"10.1007\/s13042-018-0853-2","article-title":"Research on classification method of high-dimensional class-imbalanced datasets based on SVM","volume":"10","author":"Zhang","year":"2019","journal-title":"Int. J. Mach. Learn. Cybern."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Fu, G.H., Wu, Y.J., Zong, M.J., and Pan, J. (2020). Hellinger distance-based stable sparse feature selection for high-dimensional class-imbalanced data. BMC Bioinform., 21.","DOI":"10.1186\/s12859-020-3411-3"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1010933404324","article-title":"Random forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learn."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"602","DOI":"10.1080\/21642583.2014.956265","article-title":"Random forests: From early developments to recent advancements","volume":"2","author":"Fawagreh","year":"2014","journal-title":"Syst. Sci. Control Eng."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1016\/j.inffus.2015.06.005","article-title":"Decision forest: Twenty years of research","volume":"27","author":"Rokach","year":"2016","journal-title":"Inf. Fusion"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Khoshgoftaar, T.M., Golawala, M., and Van Hulse, J. (2007, January 29\u201331). An Empirical Study of Learning from Imbalanced Data Using Random Forest. Proceedings of the 19th IEEE International Conference on Tools with Artificial Intelligence, Patras, Greece.","DOI":"10.1109\/ICTAI.2007.46"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"323","DOI":"10.1016\/j.ygeno.2012.04.003","article-title":"Random forests for genomic data analysis","volume":"99","author":"Chen","year":"2012","journal-title":"Genomics"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"220","DOI":"10.1016\/j.eswa.2016.12.035","article-title":"Learning from class-imbalanced data","volume":"73","author":"Haixiang","year":"2017","journal-title":"Expert Syst. Appl."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"863","DOI":"10.1613\/jair.1.11192","article-title":"SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year Anniversary","volume":"61","author":"Fernandez","year":"2018","journal-title":"J. Artif. Intell. Res."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"463","DOI":"10.1109\/TSMCC.2011.2161285","article-title":"A review on ensembles for the class imbalance problem: Bagging-, boosting-, and hybrid-based approaches","volume":"42","author":"Galar","year":"2012","journal-title":"IEEE Trans. Syst. Man Cybern. Part C Appl. Rev."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"80","DOI":"10.1145\/1007730.1007741","article-title":"Feature selection for text categorization on imbalanced data","volume":"6","author":"Zheng","year":"2004","journal-title":"ACM Sigkdd Explor. Newsl."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1388","DOI":"10.1109\/TKDE.2009.187","article-title":"Combating the small sample class imbalance problem using feature selection","volume":"22","author":"Wasikowski","year":"2010","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Bol\u00f3n-Canedo, V., S\u00e1nchez-Maro\u00f1o, N., and Alonso-Betanzos, A. (2015). Feature Selection for High-Dimensional Data, Artificial Intelligence: Foundations, Theory, and Algorithms, Springer.","DOI":"10.1007\/978-3-319-21858-8"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"4632","DOI":"10.1016\/j.eswa.2015.01.069","article-title":"Similarity of feature selection methods: An empirical study across data intensive classification tasks","volume":"42","author":"Pes","year":"2015","journal-title":"Expert Syst. Appl."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"106839","DOI":"10.1016\/j.csda.2019.106839","article-title":"Benchmark for filter methods for feature selection in high-dimensional classification data","volume":"143","author":"Bommert","year":"2020","journal-title":"Comput. Stat. Data Anal."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Cannas, L.M., Dess\u00ec, N., and Pes, B. (2010, January 13\u201316). A Filter-based Evolutionary Approach for Selecting Features in High-Dimensional Micro-array Data. Proceedings of the 6th International Conference on Intelligent Information Processing, Manchester, UK.","DOI":"10.1007\/978-3-642-16327-2_36"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Ahmed, N., Rafiq, J.I., and Islam, M.D.R. (2020). Enhanced Human Activity Recognition Based on Smartphone Sensor Data Using Hybrid Feature Selection Model. Sensors, 20.","DOI":"10.3390\/s20010317"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"78533","DOI":"10.1109\/ACCESS.2019.2922987","article-title":"A Survey on Hybrid Feature Selection Methods in Microarray Gene Expression Data for Cancer Classification","volume":"7","author":"Almugren","year":"2019","journal-title":"IEEE Access"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Dess\u00ec, N., and Pes, B. (2015). Stability in Biomarker Discovery: Does Ensemble Feature Selection Really Help?. Current Approaches in Applied Artificial Intelligence, Proceedings of the 28th International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems, IEA\/AIE 2015, Seoul, Korea, 10\u201312 June 2015, Springer. LNCS 9101.","DOI":"10.1007\/978-3-319-19066-2_19"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.inffus.2018.11.008","article-title":"Ensembles for feature selection: A review and future trends","volume":"52","year":"2019","journal-title":"Inf. Fusion"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"5951","DOI":"10.1007\/s00521-019-04082-3","article-title":"Ensemble feature selection for high-dimensional data: A stability analysis across multiple domains","volume":"32","author":"Pes","year":"2020","journal-title":"Neural Comput. Appl."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Haury, A.C., Gestraud, P., and Vert, J.P. (2011). The Influence of Feature Selection Methods on Accuracy, Stability and Interpretability of Molecular Signatures. PLoS ONE, 6.","DOI":"10.1371\/journal.pone.0028210"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.compbiomed.2015.08.010","article-title":"An Experimental Comparison of Feature Selection Methods on Two-Class Biomedical Datasets","volume":"66","author":"Gazda","year":"2015","journal-title":"Comput. Biol. Med."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Pes, B. (2017, January 21\u201323). Feature Selection for High-Dimensional Data: The Issue of Stability. Proceedings of the 2017 IEEE 26th International Conference on Enabling Technologies: Infrastructure for Collaborative Enterprises (WETICE), Poznan, Poland.","DOI":"10.1109\/WETICE.2017.28"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"395","DOI":"10.1007\/s10115-017-1140-3","article-title":"On the scalability of feature selection methods on high-dimensional data","volume":"56","year":"2018","journal-title":"Knowl. Inf. Syst."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Blagus, R., and Lusa, L. (2013). SMOTE for high-dimensional class-imbalanced data. BMC Bioinform., 14.","DOI":"10.1186\/1471-2105-14-106"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"13527","DOI":"10.1109\/ACCESS.2020.2966296","article-title":"Learning From High-Dimensional Biomedical Datasets: The Issue of Class Imbalance","volume":"8","author":"Pes","year":"2020","journal-title":"IEEE Access"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Sammut, C., and Webb, G.I. (2010). Cost-Sensitive Learning. Encyclopedia of Machine Learning, Springer.","DOI":"10.1007\/978-0-387-30164-8"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"113","DOI":"10.1016\/j.ins.2013.07.007","article-title":"An insight into classification with imbalanced data: Empirical results and current trends on using data intrinsic characteristics","volume":"250","author":"Palade","year":"2013","journal-title":"Inf. Sci."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.patrec.2021.01.008","article-title":"Large group activity security risk assessment and risk early warning based on random forest algorithm","volume":"144","author":"Chen","year":"2021","journal-title":"Pattern Recognit. Lett."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Figueroa, A., Peralta, B., and Nicolis, O. (2021). Coming to Grips with Age Prediction on Imbalanced Multimodal Community Question Answering Data. Information, 12.","DOI":"10.3390\/info12020048"},{"key":"ref_40","unstructured":"(2021, June 30). OpenML Datasets. Available online: https:\/\/www.openml.org\/search?type=data."},{"key":"ref_41","first-page":"78","article-title":"Microarray cancer feature selection: Review, challenges and research directions","volume":"1","author":"Hambali","year":"2020","journal-title":"Int. J. Cogn. Comput. Eng."},{"key":"ref_42","unstructured":"(2021, June 30). UCI Machine Learning Repository. Available online: https:\/\/archive.ics.uci.edu\/ml\/index.php."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"1118","DOI":"10.1109\/TKDE.2008.206","article-title":"Olex: Effective Rule Learning for Text Categorization","volume":"21","author":"Rullo","year":"2009","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"1757","DOI":"10.1016\/j.patcog.2004.03.009","article-title":"Learning multi-label scene classification","volume":"37","author":"Boutell","year":"2004","journal-title":"Pattern Recognit."},{"key":"ref_45","unstructured":"Witten, I.H., Frank, E., Hall, M.A., and Pal, C.J. (2016). Data Mining: Practical Machine Learning Tools and Techniques, Morgan Kaufmann."},{"key":"ref_46","unstructured":"(2021, June 30). Weka: Data Mining Software in Java. Available online: https:\/\/www.cs.waikato.ac.nz\/ml\/weka\/."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"239","DOI":"10.1023\/A:1024068626366","article-title":"Inference for the Generalization Error","volume":"52","author":"Nadeau","year":"2003","journal-title":"Mach. Learn."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"123","DOI":"10.1007\/978-1-4939-9442-7_6","article-title":"Feature Selection Applied to Microarray Data","volume":"Volume 1986","year":"2019","journal-title":"Microarray Bioinformatics"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Dess\u00ec, N., Milia, G., and Pes, B. (2013). Enhancing Random Forests Performance in Microarray Data Classification. Artificial Intelligence in Medicine, Proceedings of the 14th Conference on Artificial Intelligence in Medicine, AIME 2013, Murcia, Spain, 29 May\u20131 June 2013, Springer. LNCS 7885.","DOI":"10.1007\/978-3-642-38326-7_15"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Cilia, N.D., De Stefano, C., Fontanella, F., Raimondo, S., and Scotto di Freca, A. (2019). An Experimental Comparison of Feature-Selection and Classification Methods for Microarray Datasets. Information, 10.","DOI":"10.3390\/info10030109"},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"737","DOI":"10.1007\/s40745-019-00209-4","article-title":"On Regularisation Methods for Analysis of High Dimensional Data","volume":"6","author":"Sirimongkolkasem","year":"2019","journal-title":"Ann. Data. Sci."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Wu, S., Jiang, H., Shen, H., and Yang, Z. (2018). Gene Selection in Cancer Classification Using Sparse Logistic Regression with L1\/2 Regularization. Appl. Sci., 8.","DOI":"10.3390\/app8091569"},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"114","DOI":"10.1016\/j.jbi.2015.02.003","article-title":"Efficient and sparse feature selection for biomedical text classification via the elastic net: Application to ICU risk stratification from nursing notes","volume":"54","author":"Marafino","year":"2015","journal-title":"J. Biomed. Inform."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/12\/8\/286\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:32:35Z","timestamp":1760164355000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/12\/8\/286"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,21]]},"references-count":53,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2021,8]]}},"alternative-id":["info12080286"],"URL":"https:\/\/doi.org\/10.3390\/info12080286","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,7,21]]}}}