{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T04:25:38Z","timestamp":1777695938283,"version":"3.51.4"},"reference-count":45,"publisher":"SAGE Publications","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IDA"],"published-print":{"date-parts":[[2020,12,18]]},"abstract":"<jats:p>Class imbalance is often a problem in various real-world datasets, where one class contains a small number of data and the other contains a large number of data. It is notably difficult to develop an effective model using traditional data mining and machine learning algorithms without using data preprocessing techniques to balance the dataset. Oversampling is often used as a pretreatment method for imbalanced datasets. Specifically, synthetic oversampling techniques focus on balancing the number of training instances between the majority class and the minority class by generating extra artificial minority class instances. However, the current oversampling techniques simply consider the imbalance of quantity and pay no attention to whether the distribution is balanced or not. Therefore, this paper proposes an entropy difference and kernel-based SMOTE (EDKS) which considers the imbalance degree of dataset from distribution by entropy difference and overcomes the limitation of SMOTE for nonlinear problems by oversampling in the feature space of support vector machine classifier. First, the EDKS method maps the input data into a feature space to increase the separability of the data. Then EDKS calculates the entropy difference in kernel space, determines the majority class and minority class, and finds the sparse regions in the minority class. Moreover, the proposed method balances the data distribution by synthesizing new instances and evaluating its retention capability. Our algorithm can effectively distinguish those datasets with the same imbalance ratio but different distribution. The experimental study evaluates and compares the performance of our method against state-of-the-art algorithms, and then demonstrates that the proposed approach is competitive with the state-of-art algorithms on multiple benchmark imbalanced datasets.<\/jats:p>","DOI":"10.3233\/ida-194761","type":"journal-article","created":{"date-parts":[[2020,12,22]],"date-time":"2020-12-22T15:04:11Z","timestamp":1608649451000},"page":"1239-1255","source":"Crossref","is-referenced-by-count":4,"title":["Entropy difference and kernel-based oversampling technique for imbalanced data learning"],"prefix":"10.1177","volume":"24","author":[{"given":"Xu","family":"Wu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Youlong","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lingyu","family":"Ren","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"key":"10.3233\/IDA-194761_ref1","doi-asserted-by":"crossref","first-page":"76","DOI":"10.1016\/j.ins.2017.10.017","article-title":"Imbalanced enterprise credit evaluation with DTESBD: Decision tree ensemble based on SMOTE and bagging with differentiated sampling rates","volume":"425","author":"Sun","year":"2018","journal-title":"Inf. Sci."},{"key":"10.3233\/IDA-194761_ref2","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1016\/j.knosys.2012.12.007","article-title":"Performance of corporate bankruptcy prediction models on imbalanced dataset: The effect of sampling methods","volume":"41","author":"Zhou","year":"2013","journal-title":"Knowl.-Based Syst."},{"issue":"6","key":"10.3233\/IDA-194761_ref3","doi-asserted-by":"crossref","first-page":"3027","DOI":"10.1016\/j.eswa.2013.10.033","article-title":"Detecting online auction shilling frauds using supervised learning","volume":"41","author":"Tsang","year":"2014","journal-title":"Expert Syst. Appl."},{"issue":"12","key":"10.3233\/IDA-194761_ref4","doi-asserted-by":"crossref","first-page":"4805","DOI":"10.1016\/j.eswa.2013.02.027","article-title":"Finding the needle: A risk-based ranking of product listings at online auction sites for non-delivery fraud prediction","volume":"40","author":"Almendra","year":"2013","journal-title":"Expert Syst. Appl."},{"key":"10.3233\/IDA-194761_ref5","first-page":"220","article-title":"Learning from class imbalanced data: Review of methods and applications","volume":"73","author":"Guo","year":"2016","journal-title":"Expert Syst. Appl."},{"issue":"8","key":"10.3233\/IDA-194761_ref6","doi-asserted-by":"crossref","first-page":"753","DOI":"10.1016\/j.knosys.2008.03.031","article-title":"A comparative study on rough set based class imbalance learning","volume":"21","author":"Liu","year":"2008","journal-title":"Knowl.-Based Syst."},{"issue":"9","key":"10.3233\/IDA-194761_ref7","doi-asserted-by":"crossref","first-page":"1263","DOI":"10.1109\/TKDE.2008.239","article-title":"Learning from imbalanced data","volume":"21","author":"He","year":"2009","journal-title":"IEEE Trans. Knowl. Data Eng."},{"issue":"1","key":"10.3233\/IDA-194761_ref8","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1007730.1007733","article-title":"Editorial: Special issue on learning from imbalanced data sets","volume":"6","author":"Chawla","year":"2004","journal-title":"ACM SIGKDD Explor. Newslett."},{"issue":"1","key":"10.3233\/IDA-194761_ref9","first-page":"25","article-title":"Handling imbalanced datasets: A review","volume":"30","author":"Kotsiantis","year":"2006","journal-title":"Science"},{"issue":"4","key":"10.3233\/IDA-194761_ref10","doi-asserted-by":"crossref","first-page":"463","DOI":"10.1109\/TSMCC.2011.2161285","article-title":"A review on ensembles for the class imbalance problem: Bagging-, boosting-, and hybrid-based approaches","volume":"42","author":"Galar","year":"2012","journal-title":"IEEE Trans. Syst. Man Cybern. Part C"},{"key":"10.3233\/IDA-194761_ref11","doi-asserted-by":"crossref","first-page":"142","DOI":"10.1016\/j.ins.2017.02.059","article-title":"Posterior probability based ensemble strategy using optimizing decision directed acyclic graph for multi-class classification","volume":"400-401","author":"Zhou","year":"2017","journal-title":"Inf. Sci."},{"issue":"6","key":"10.3233\/IDA-194761_ref12","doi-asserted-by":"crossref","first-page":"740","DOI":"10.1016\/j.knosys.2010.12.010","article-title":"A novel virtual sample generation method based on Gaussian distribution","volume":"24","author":"Yang","year":"2011","journal-title":"Knowl.-Based Syst."},{"key":"10.3233\/IDA-194761_ref13","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1016\/j.ins.2017.05.008","article-title":"Clustering-based undersampling in class imbalanced data","volume":"409-410","author":"Lin","year":"2017","journal-title":"Inf. Sci."},{"key":"10.3233\/IDA-194761_ref14","doi-asserted-by":"crossref","first-page":"184","DOI":"10.1016\/j.ins.2014.08.051","article-title":"SMOTE-IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering","volume":"291","author":"Sez","year":"2015","journal-title":"Inf. Sci."},{"key":"10.3233\/IDA-194761_ref15","unstructured":"H. He, Y. Bai, E.A. Garcia and S. Li, Adasyn: adaptive synthetic sampling approach for imbalanced learning, Neural Networks, 2008, IJCNN 2008, (IEEE World Congress on Computational Intelligence), in: IEEE International Joint Conference on Neural Networks, 2008, pp. 1322\u20131328."},{"issue":"9","key":"10.3233\/IDA-194761_ref16","doi-asserted-by":"crossref","first-page":"4065","DOI":"10.1109\/TNNLS.2017.2751612","article-title":"Classification of imbalanced data by oversampling in kernel space of support vector machines","volume":"29","author":"Mathew","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"10.3233\/IDA-194761_ref17","doi-asserted-by":"crossref","first-page":"5718","DOI":"10.1016\/j.eswa.2008.06.108","article-title":"Cluster-based undersampling approaches for imbalanced data distributions","volume":"36","author":"Yen","year":"2009","journal-title":"Expert Syst. Appl."},{"key":"10.3233\/IDA-194761_ref18","doi-asserted-by":"crossref","first-page":"88","DOI":"10.1016\/j.neucom.2017.03.011","article-title":"Fast-CBUS: A fast clustering-based undersampling method for addressing the class imbalance problem","volume":"243","author":"Ofek","year":"2017","journal-title":"Neurocomputing"},{"key":"10.3233\/IDA-194761_ref19","doi-asserted-by":"crossref","first-page":"100","DOI":"10.2307\/2346830","article-title":"A k-means clustering algorithm","volume":"28","author":"Hartigan","year":"1979","journal-title":"Appl. Stat."},{"key":"10.3233\/IDA-194761_ref20","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1613\/jair.953","article-title":"SMOTE: Synthetic minority over-sampling technique","volume":"16","author":"Chawla","year":"2002","journal-title":"J. Artif. Intell. Res."},{"key":"10.3233\/IDA-194761_ref21","doi-asserted-by":"crossref","unstructured":"H. Han, W.Y. Wang and B.H. Mao, Borderline-SMOTE: a new over-sampling method in imbalanced data sets learning, in: Proceedings of the 1st International Conference on Intelligent Computing, Hefei, China, 2005, 878\u2013887.","DOI":"10.1007\/11538059_91"},{"key":"10.3233\/IDA-194761_ref22","doi-asserted-by":"crossref","unstructured":"C. Bunkhumpornpat, K. Sinapiromsaran and C. Lursinsap, Safe-level-SMOTE: safe-level-synthetic minority oversampling technique for handling the class imbalanced problem, in: Proceedings of the 13th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining, 2009, pp. 475\u2013482.","DOI":"10.1007\/978-3-642-01307-2_43"},{"key":"10.3233\/IDA-194761_ref23","doi-asserted-by":"crossref","first-page":"664","DOI":"10.1007\/s10489-011-0287-y","article-title":"DBSMOTE: Density-based synthetic minority over-sampling technique","volume":"36","author":"Bunkhumpornpat","year":"2012","journal-title":"Appl. Intell."},{"issue":"2","key":"10.3233\/IDA-194761_ref24","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1109\/TKDE.2012.232","article-title":"MWMOTE-majority weighted minority oversampling technique for imbalanced data set learning","volume":"26","author":"Barua","year":"2014","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"10.3233\/IDA-194761_ref25","doi-asserted-by":"crossref","first-page":"863","DOI":"10.1613\/jair.1.11192","article-title":"SMOTE for learning from imbalanced data: Progress and challenges, marking the 15-year anniversary","volume":"61","author":"Fern\u00e1ndez","year":"2018","journal-title":"J. Artif. Intell. Res."},{"key":"10.3233\/IDA-194761_ref26","unstructured":"B. Wang and N. Japkowicz, Imbalanced data set learning with synthetic samples, in: Proc. IRIS Machine Learning Workshop, 2004, p. 19."},{"key":"10.3233\/IDA-194761_ref27","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1016\/j.patcog.2017.07.024","article-title":"Synthetic minority oversampling technique for multiclass imbalance problems","volume":"72","author":"Zhu","year":"2017","journal-title":"Pattern Recognit."},{"issue":"1","key":"10.3233\/IDA-194761_ref28","doi-asserted-by":"crossref","first-page":"92","DOI":"10.1007\/s10618-012-0295-5","article-title":"Training and assessing classification rules with imbalanced data","volume":"28","author":"Menardi","year":"2014","journal-title":"Data Min. Knowl. Discovery"},{"key":"10.3233\/IDA-194761_ref29","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1016\/j.knosys.2018.05.044","article-title":"Fuzzy rule-based oversampling technique for imbalanced and incomplete data learning","volume":"158","author":"Liu","year":"2018","journal-title":"Knowledge-Based Systems"},{"key":"10.3233\/IDA-194761_ref30","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1016\/j.inffus.2013.12.003","article-title":"RWO-sampling: A random walk over-sampling approach to imbalanced data classification","volume":"20","author":"Zhang","year":"2014","journal-title":"Inf. Fusion"},{"issue":"1","key":"10.3233\/IDA-194761_ref31","doi-asserted-by":"crossref","first-page":"222","DOI":"10.1109\/TKDE.2014.2324567","article-title":"Racog and wracog: Two probabilistic oversampling techniques","volume":"27","author":"Das","year":"2015","journal-title":"IEEE Trans. Knowl. Data Eng."},{"issue":"1","key":"10.3233\/IDA-194761_ref32","doi-asserted-by":"crossref","first-page":"238","DOI":"10.1109\/TKDE.2015.2458858","article-title":"To combat multi-class imbalanced problems by means of over-sampling techniques","volume":"28","author":"Abdi","year":"2016","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"10.3233\/IDA-194761_ref36","doi-asserted-by":"crossref","unstructured":"N. Cristianini, J. Kandola, A. Elisseeff and J. Shawe-Taylor, On kernel target alignment, in: Proc. Neural Inf. Process. Syst. (NIPS), Vancouver, BC, Canada, Dec. 2001, pp. 367\u2013373.","DOI":"10.7551\/mitpress\/1120.003.0052"},{"key":"10.3233\/IDA-194761_ref37","first-page":"2491","article-title":"SimpleMKL","volume":"9","author":"Rakotomamonjy","year":"2008","journal-title":"J. Mach. Learn. Res."},{"issue":"9","key":"10.3233\/IDA-194761_ref38","doi-asserted-by":"crossref","first-page":"1","DOI":"10.18637\/jss.v011.i09","article-title":"kernlab-An S4 package for kernel methods in R","volume":"11","author":"Karatzoglou","year":"2004","journal-title":"J. Stat. Softw."},{"key":"10.3233\/IDA-194761_ref39","doi-asserted-by":"crossref","unstructured":"Z.Q. Zeng and J. Gao, Improving SVM classification with imbalance data set, in: Proc. 16th Int. Conf. Neural Inf. Process., Bangkok, Thailand, Dec. 2009, pp. 389\u2013398.","DOI":"10.1007\/978-3-642-10677-4_44"},{"key":"10.3233\/IDA-194761_ref40","doi-asserted-by":"crossref","unstructured":"M. Perez-Ortiz, P.A. Gutierrez and C. Hervas-Martinez, Borderline Kernel Based Over-Sampling, in: Proc. Int. Conf. Hybrid Artif. Intell. Syst., Salamanca, Spain, Sep. 2013, pp. 472\u2013481.","DOI":"10.1007\/978-3-642-40846-5_47"},{"key":"10.3233\/IDA-194761_ref41","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.knosys.2014.09.002","article-title":"A proposal for evolutionary fuzzy systems using feature weighting: Dealing with overlapping in imbalanced datasets","volume":"73","author":"Alshomrani","year":"2015","journal-title":"Knowl.-Based Syst."},{"issue":"7","key":"10.3233\/IDA-194761_ref42","doi-asserted-by":"crossref","first-page":"1145","DOI":"10.1016\/S0031-3203(96)00142-2","article-title":"The use of the area under the ROC curve in the evaluation of machine learning algorithms","volume":"30","author":"Bradley","year":"1997","journal-title":"Pattern Recognit."},{"key":"10.3233\/IDA-194761_ref44","doi-asserted-by":"crossref","first-page":"100","DOI":"10.1016\/j.knosys.2016.10.018","article-title":"PBC4cip: A new contrast pattern-based classifier for class imbalance problems","volume":"115","author":"Loyola-Gonzlez","year":"2017","journal-title":"Knowl.-Based Syst."},{"key":"10.3233\/IDA-194761_ref46","unstructured":"J. Alcal\u00e1-Fdez, A. Fern\u00e1ndez, J. Luengo, J. Derrac and S. Garc\u00eda, Keel data-mining software tool: Data set repository, integration of algorithms and experimental analysis framework, Journal of Multiple-Valued Logic Soft Computing 17 (2011)."},{"issue":"2","key":"10.3233\/IDA-194761_ref47","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1023\/A:1009715923555","article-title":"A tutorial on support vector machines for pattern recognition","volume":"2","author":"Burges","year":"1998","journal-title":"Data Mining Knowl. Discovery"},{"issue":"12","key":"10.3233\/IDA-194761_ref48","doi-asserted-by":"crossref","first-page":"4263","DOI":"10.1109\/TCYB.2016.2606104","article-title":"A noise-filtered under-sampling scheme for imbalanced classification","volume":"47","author":"Kang","year":"2016","journal-title":"IEEE Transactions on Cybernetics"},{"key":"10.3233\/IDA-194761_ref49","doi-asserted-by":"crossref","first-page":"614","DOI":"10.1007\/s10489-015-0666-x","article-title":"Missing data imputation by K nearest neighbours based on grey relational structure and mutual information","volume":"43","author":"Pan","year":"2015","journal-title":"Appl Intell"},{"key":"10.3233\/IDA-194761_ref50","first-page":"1016","article-title":"Effective and efficient approach to classification with incomplete data","volume":"10","author":"Tran","year":"2018","journal-title":"Knowl.-Based Syst."}],"container-title":["Intelligent Data Analysis"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/IDA-194761","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:18:50Z","timestamp":1777454330000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/full\/10.3233\/IDA-194761"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,12,18]]},"references-count":45,"journal-issue":{"issue":"6"},"URL":"https:\/\/doi.org\/10.3233\/ida-194761","relation":{},"ISSN":["1088-467X","1571-4128"],"issn-type":[{"value":"1088-467X","type":"print"},{"value":"1571-4128","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,12,18]]}}}