{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T20:22:16Z","timestamp":1781122936232,"version":"3.54.1"},"reference-count":39,"publisher":"Emerald","issue":"5","license":[{"start":{"date-parts":[[2021,5,14]],"date-time":"2021-05-14T00:00:00Z","timestamp":1620950400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.emerald.com\/insight\/site-policies"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["DTA"],"published-print":{"date-parts":[[2021,10,11]]},"abstract":"<jats:sec><jats:title content-type=\"abstract-subheading\">Purpose<\/jats:title><jats:p>Class imbalance learning, which exists in many domain problem datasets, is an important research topic in data mining and machine learning. One-class classification techniques, which aim to identify anomalies as the minority class from the normal data as the majority class, are one representative solution for class imbalanced datasets. Since one-class classifiers are trained using only normal data to create a decision boundary for later anomaly detection, the quality of the training set, i.e. the majority class, is one key factor that affects the performance of one-class classifiers.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-subheading\">Design\/methodology\/approach<\/jats:title><jats:p>In this paper, we focus on two data cleaning or preprocessing methods to address class imbalanced datasets. The first method examines whether performing instance selection to remove some noisy data from the majority class can improve the performance of one-class classifiers. The second method combines instance selection and missing value imputation, where the latter is used to handle incomplete datasets that contain missing values.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-subheading\">Findings<\/jats:title><jats:p>The experimental results are based on 44 class imbalanced datasets; three instance selection algorithms, including IB3, DROP3 and the GA, the CART decision tree for missing value imputation, and three one-class classifiers, which include OCSVM, IFOREST and LOF, show that if the instance selection algorithm is carefully chosen, performing this step could improve the quality of the training data, which makes one-class classifiers outperform the baselines without instance selection. Moreover, when class imbalanced datasets contain some missing values, combining missing value imputation and instance selection, regardless of which step is first performed, can maintain similar data quality as datasets without missing values.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-subheading\">Originality\/value<\/jats:title><jats:p>The novelty of this paper is to investigate the effect of performing instance selection on the performance of one-class classifiers, which has never been done before. Moreover, this study is the first attempt to consider the scenario of missing values that exist in the training set for training one-class classifiers. In this case, performing missing value imputation and instance selection with different orders are compared.<\/jats:p><\/jats:sec>","DOI":"10.1108\/dta-01-2021-0027","type":"journal-article","created":{"date-parts":[[2021,5,13]],"date-time":"2021-05-13T09:56:53Z","timestamp":1620899813000},"page":"771-787","source":"Crossref","is-referenced-by-count":5,"title":["Data cleaning issues in class imbalanced datasets: instance selection and missing values imputation for one-class classifiers"],"prefix":"10.1108","volume":"55","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5875-1355","authenticated-orcid":false,"given":"Zhenyuan","family":"Wang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5991-2253","authenticated-orcid":false,"given":"Chih-Fong","family":"Tsai","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5803-513X","authenticated-orcid":false,"given":"Wei-Chao","family":"Lin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"140","published-online":{"date-parts":[[2021,5,14]]},"reference":[{"issue":"1","key":"key2021100913305557800_ref001","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1007\/BF00153759","article-title":"Instance-based learning algorithms","volume":"6","year":"1991","journal-title":"Machine Learning"},{"key":"key2021100913305557800_ref002","doi-asserted-by":"crossref","first-page":"841","DOI":"10.1007\/s10115-019-01380-z","article-title":"Framework for extreme imbalance classification-SWIM\u2014sampling with the majority class","volume":"62","year":"2020","journal-title":"Knowledge and Information Systems"},{"issue":"2","key":"key2021100913305557800_ref003","article-title":"A survey of predictive modeling on imbalanced domains","volume":"49","year":"2016","journal-title":"ACM Computing Surveys"},{"issue":"2","key":"key2021100913305557800_ref004","doi-asserted-by":"crossref","first-page":"93","DOI":"10.1145\/335191.335388","article-title":"LOF: identifying density-based local outliers","volume":"29","year":"2000","journal-title":"SIGMOD Record"},{"issue":"6","key":"key2021100913305557800_ref005","doi-asserted-by":"crossref","first-page":"561","DOI":"10.1109\/TEVC.2003.819265","article-title":"Using evolutionary algorithms as instance selection for data reduction: an experimental study","volume":"7","year":"2003","journal-title":"IEEE Transactions on Evolutionary Computation"},{"issue":"3","key":"key2021100913305557800_ref006","first-page":"15:1","article-title":"Anomaly detection: a survey","volume":"41","year":"2009","journal-title":"ACM Computing Surveys"},{"key":"key2021100913305557800_ref007","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1016\/j.ins.2017.04.044","article-title":"Machine learning based mobile malware detection using highly imbalanced network traffic","volume":"433-434","year":"2018","journal-title":"Information Sciences"},{"key":"key2021100913305557800_ref008","doi-asserted-by":"crossref","first-page":"3685","DOI":"10.1007\/s00521-018-3747-z","article-title":"Imbalanced dataset-based echo state networks for anomaly detection","volume":"32","year":"2020","journal-title":"Neural Computing and Applications"},{"key":"key2021100913305557800_ref009","first-page":"1","article-title":"Statistical comparisons of classifiers over multiple data sets","volume":"7","year":"2006","journal-title":"Journal of Machine Learning Research"},{"key":"key2021100913305557800_ref010","doi-asserted-by":"crossref","first-page":"406","DOI":"10.1016\/j.patcog.2017.09.037","article-title":"A comparative evaluation of outlier detection algorithms: experiments and analyses","volume":"74","year":"2018","journal-title":"Pattern Recognition"},{"key":"key2021100913305557800_ref011","doi-asserted-by":"crossref","first-page":"103089","DOI":"10.1016\/j.jbi.2018.12.003","article-title":"A comprehensive data level analysis for cancer diagnosis on imbalanced data","volume":"90","year":"2019","journal-title":"Journal of Biomedical Informatics"},{"key":"key2021100913305557800_ref012","doi-asserted-by":"crossref","first-page":"463","DOI":"10.1109\/TSMCC.2011.2161285","article-title":"A review on ensembles for the class imbalance problem: bagging-, boosting-, and hybrid based approaches","volume":"42","year":"2012","journal-title":"IEEE Transactions on Systems, Man, and Cybernetics - Part C: Applications and Reviews"},{"issue":"3","key":"key2021100913305557800_ref013","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1109\/TPAMI.2011.142","article-title":"Prototype selection for nearest neighbor classification: taxonomy and empirical study","volume":"34","year":"2012","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"key2021100913305557800_ref014","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1007\/s00521-009-0295-6","article-title":"Pattern classification with missing data: a review","volume":"19","year":"2010","journal-title":"Neural Computing and Applications"},{"key":"key2021100913305557800_ref015","first-page":"192","article-title":"On the class imbalance problem","year":"2008"},{"key":"key2021100913305557800_ref016","doi-asserted-by":"crossref","first-page":"4711","DOI":"10.1007\/s00500-019-04501-6","article-title":"Ensemble learning via constraint projection and undersampling technique for class-imbalance problem","volume":"24","year":"2020","journal-title":"Soft Computing"},{"key":"key2021100913305557800_ref017","doi-asserted-by":"crossref","first-page":"7153","DOI":"10.1007\/s00521-018-3551-9","article-title":"A fuzzy twin support vector machine based on information entropy for class imbalance learning","volume":"31","year":"2019","journal-title":"Neural Computing and Applications"},{"key":"key2021100913305557800_ref018","doi-asserted-by":"crossref","first-page":"3687","DOI":"10.1007\/s13042-019-00953-2","article-title":"A Gaussian mixture model based combined resampling algorithm for classification of imbalanced credit data sets","volume":"10","year":"2019","journal-title":"International Journal of Machine Learning and Cybernetics"},{"issue":"2","key":"key2021100913305557800_ref019","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1023\/B:AIRE.0000045502.10941.a9","article-title":"A survey of outlier detection methodologies","volume":"22","year":"2004","journal-title":"Artificial Intelligence Review"},{"key":"key2021100913305557800_ref020","first-page":"1817479","article-title":"Outlier removal in model-based missing value imputation for medical datasets","volume":"2018","year":"2018","journal-title":"Journal of Healthcare Engineering"},{"issue":"3","key":"key2021100913305557800_ref021","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1017\/S026988891300043X","article-title":"One-class classification: taxonomy of study and review of techniques","volume":"29","year":"2014","journal-title":"The Knowledge Engineering Review"},{"key":"key2021100913305557800_ref501","doi-asserted-by":"crossref","first-page":"601","DOI":"10.1007\/s10115-018-1220-z","article-title":"Instance selection for one-class classification","volume":"59","year":"2019","journal-title":"Knowledge and Information Systems"},{"key":"key2021100913305557800_ref022","doi-asserted-by":"crossref","first-page":"1487","DOI":"10.1007\/s10462-019-09709-4","article-title":"Missing value imputation: a review and analysis of the literature (2006 \u2013 2017)","volume":"53","year":"2020","journal-title":"Artificial Intelligence Review"},{"key":"key2021100913305557800_ref023","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.jss.2015.04.038","article-title":"Learning to detect representative data for large scale instance selection","volume":"106","year":"2015","journal-title":"Journal of Systems and Software"},{"key":"key2021100913305557800_ref024","first-page":"413","article-title":"Isolation forest","year":"2008"},{"key":"key2021100913305557800_ref025","doi-asserted-by":"crossref","first-page":"113","DOI":"10.1016\/j.ins.2013.07.007","article-title":"An insight into classification with imbalanced data: empirical results and current trends on using data intrinsic characteristics","volume":"250","year":"2013","journal-title":"Information Sciences"},{"key":"key2021100913305557800_ref026","doi-asserted-by":"crossref","first-page":"133","DOI":"10.1007\/s10462-010-9165-y","article-title":"A review of instance selection methods","volume":"34","year":"2010","journal-title":"Artificial Intelligence Review"},{"key":"key2021100913305557800_ref027","doi-asserted-by":"crossref","first-page":"5951","DOI":"10.1007\/s00521-019-04082-3","article-title":"Ensemble feature selection for high-dimensional data: a stability analysis across multiple domains","volume":"32","year":"2020","journal-title":"Neural Computing and Applications"},{"key":"key2021100913305557800_ref028","doi-asserted-by":"crossref","first-page":"172","DOI":"10.1016\/j.neucom.2018.10.056","article-title":"Class imbalance learning using UnderBagging based kernelized extreme learning machine","volume":"329","year":"2019","journal-title":"Neurocomputing"},{"issue":"3","key":"key2021100913305557800_ref029","doi-asserted-by":"crossref","first-page":"457","DOI":"10.1080\/0952813X.2017.1409283","article-title":"Instance selection algorithm by ensemble margin","volume":"30","year":"2018","journal-title":"Journal of Experimental and Theoretical Artificial Intelligence"},{"issue":"2","key":"key2021100913305557800_ref030","first-page":"1253","article-title":"A comprehensive investigation of the role of imbalanced learning for software defect prediction","volume":"45","year":"2019","journal-title":"IEEE Transactions on Software Engineering"},{"key":"key2021100913305557800_ref031","doi-asserted-by":"crossref","first-page":"128","DOI":"10.1016\/j.inffus.2019.07.006","article-title":"Class-imbalanced dynamic financial distress prediction based on Adaboost-SVM ensemble combined with SMOTE and time weighting","volume":"54","year":"2020","journal-title":"Information Fusion"},{"key":"key2021100913305557800_ref032","doi-asserted-by":"crossref","first-page":"105601","DOI":"10.1016\/j.asoc.2019.105601","article-title":"Performance enhanced boosted SVM for imbalanced datasets","volume":"83","year":"2019","journal-title":"Applied Soft Computing"},{"key":"key2021100913305557800_ref033","doi-asserted-by":"crossref","first-page":"1191","DOI":"10.1016\/S0167-8655(99)00087-2","article-title":"Support vector domain description","volume":"20","year":"1999","journal-title":"Pattern Recognition Letters"},{"key":"key2021100913305557800_ref034","doi-asserted-by":"crossref","first-page":"106097","DOI":"10.1016\/j.knosys.2020.106097","article-title":"Ensemble feature selection in high dimension, low sample size datasets: parallel and serial combination approaches","volume":"203","year":"2020","journal-title":"Knowledge-Based Systems"},{"key":"key2021100913305557800_ref035","doi-asserted-by":"crossref","first-page":"47","DOI":"10.1016\/j.ins.2018.10.029","article-title":"Under-sampling class imbalanced datasets by combining clustering analysis and instance selection","volume":"477","year":"2019","journal-title":"Information Sciences"},{"issue":"10","key":"key2021100913305557800_ref036","doi-asserted-by":"crossref","first-page":"1388","DOI":"10.1109\/TKDE.2009.187","article-title":"Combating the small sample class imbalance problem using feature selection","volume":"22","year":"2010","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"issue":"3","key":"key2021100913305557800_ref037","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1023\/A:1007626913721","article-title":"Reduction techniques for instance-based learning algorithms","volume":"38","year":"2000","journal-title":"Machine Learning"},{"key":"key2021100913305557800_ref038","doi-asserted-by":"crossref","first-page":"13235","DOI":"10.1007\/s00500-019-03865-z","article-title":"Constraint nearest neighbor for instance selection","volume":"23","year":"2019","journal-title":"Soft Computing"}],"container-title":["Data Technologies and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/DTA-01-2021-0027\/full\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/DTA-01-2021-0027\/full\/html","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T23:14:54Z","timestamp":1753398894000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.emerald.com\/dta\/article\/55\/5\/771-787\/273982"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,14]]},"references-count":39,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2021,5,14]]},"published-print":{"date-parts":[[2021,10,11]]}},"alternative-id":["10.1108\/DTA-01-2021-0027"],"URL":"https:\/\/doi.org\/10.1108\/dta-01-2021-0027","relation":{},"ISSN":["2514-9288"],"issn-type":[{"value":"2514-9288","type":"print"}],"subject":[],"published":{"date-parts":[[2021,5,14]]}}}