{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,20]],"date-time":"2026-04-20T12:21:46Z","timestamp":1776687706928,"version":"3.51.2"},"reference-count":40,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2017]]},"DOI":"10.1587\/transinf.2016edp7204","type":"journal-article","created":{"date-parts":[[2017,1,31]],"date-time":"2017-01-31T17:30:07Z","timestamp":1485883807000},"page":"265-272","source":"Crossref","is-referenced-by-count":38,"title":["The Performance Stability of Defect Prediction Models with Class Imbalance: An Empirical Study"],"prefix":"10.1587","volume":"E100.D","author":[{"given":"Qiao","family":"YU","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, China University of Mining and Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shujuan","family":"JIANG","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, China University of Mining and Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yanmei","family":"ZHANG","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, China University of Mining and Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","doi-asserted-by":"crossref","unstructured":"[1] H. Wang, T.M. Khoshgoftaar, and A. Napolitano, \u201cAn empirical study on the stability of feature selection for imbalanced software engineering data,\u201d Proc. 11th Int. Conf. Mach. Learn. &amp; Appl., vol.1, pp.317-323, Boca Raton, USA, Dec. 2012.","DOI":"10.1109\/ICMLA.2012.60"},{"key":"2","doi-asserted-by":"crossref","unstructured":"[2] A. Sun, E.P. Lim, and Y. Liu, \u201cOn strategies for imbalanced text classification using SVM: A comparative study,\u201d Decis. Support Syst., vol.48, no.1, pp.191-201, 2009.","DOI":"10.1016\/j.dss.2009.07.011"},{"key":"3","doi-asserted-by":"crossref","unstructured":"[3] M.A. Mazurowski, P.A. Habas, J.M. Zurada, J.Y. Lo, J.A. Baker, and G.D. Tourassi, \u201cTraining neural network classifiers for medical decision making: The effects of imbalanced datasets on classification performance,\u201d Neural Networks, vol.21, no.2-3, pp.427-436, 2008.","DOI":"10.1016\/j.neunet.2007.12.031"},{"key":"4","doi-asserted-by":"crossref","unstructured":"[4] Y. Ma, G. Luo, and H. Chen, \u201cKernel based asymmetric learning for software defect prediction,\u201d IEICE Trans. Inf. &amp; Syst., vol.E95-D, no.1, pp.267-270, Jan. 2012.","DOI":"10.1587\/transinf.E95.D.267"},{"key":"5","doi-asserted-by":"crossref","unstructured":"[5] S. Wang and X. Yao, \u201cUsing class imbalance learning for software defect prediction,\u201d IEEE Trans. Reliab., vol.62, no.2, pp.434-443, 2013.","DOI":"10.1109\/TR.2013.2259203"},{"key":"6","doi-asserted-by":"crossref","unstructured":"[6] Z. Sun, Q. Song, X. Zhu, H. Sun, B. Xu, and Y. Zhou, \u201cA novel ensemble method for classifying imbalanced data,\u201d Pattern Recogn., vol.48, no.5, pp.1623-1637, 2015.","DOI":"10.1016\/j.patcog.2014.11.014"},{"key":"7","doi-asserted-by":"crossref","unstructured":"[7] L. Peng, H. Zhang, B. Yang, and Y. Chen, \u201cA new approach for imbalanced data classification based on data gravitation,\u201d Inform. Sciences, vol.288, pp.347-373, 2014.","DOI":"10.1016\/j.ins.2014.04.046"},{"key":"8","doi-asserted-by":"crossref","unstructured":"[8] V. L\u00f3pez, A. Fern\u00e1ndez, S. Garc\u00eda, V. Palade, and F. Herrera, \u201cAn insight into classification with imbalanced data: Empirical results and current trends on using data intrinsic characteristics,\u201d Inform. Sciences, vol.250, pp.113-141, 2013.","DOI":"10.1016\/j.ins.2013.07.007"},{"key":"9","doi-asserted-by":"crossref","unstructured":"[9] J. Forkman, \u201cEstimator and tests for common coefficients of variation in normal distributions,\u201d Commun. Statist.-Theory &amp; Methods, vol.38, no.2, pp.233-251, 2009.","DOI":"10.1080\/03610920802187448"},{"key":"10","unstructured":"[10] J.R. Quinlan, C4.5: Programs for machine learning, Morgan Kaufmann Publishers, San Francisco, USA, 1993."},{"key":"11","unstructured":"[11] G.H. John and P. Langley, \u201cEstimating continuous distributions in Bayesian classifiers,\u201d Proc. 11th Conf. Uncertainty Artif. Intell., pp.338-345, Montreal, Canada, Aug. 1995."},{"key":"12","doi-asserted-by":"crossref","unstructured":"[12] L. Breiman, \u201cRandom forests,\u201d Mach. Learn., vol.45, no.1, pp.5-32, 2001.","DOI":"10.1023\/A:1010933404324"},{"key":"13","doi-asserted-by":"crossref","unstructured":"[13] N. Japkowicz and S. Stephen, \u201cThe class imbalance problem: A systematic study,\u201d Intell. Data Anal., vol.6, no.5, pp.429-449, 2002.","DOI":"10.3233\/IDA-2002-6504"},{"key":"14","unstructured":"[14] T. Galinac Grbac, G. Mau\u0161a and B. Dalbelo-Ba\u0161ic, \u201cStability of software defect prediction in relation to levels of data imbalance,\u201d Proc. 2nd Workshop Softw. Qual. Anal. Monit. Improv. &amp; Appl., pp.1-10, Novi Sad, Serbia, Sept. 2013."},{"key":"15","doi-asserted-by":"crossref","unstructured":"[15] C. Seiffert, T.M. Khoshgoftaar, J. Van Hulse, and A. Folleco, \u201cAn empirical study of the classification performance of learners on imbalanced and noisy software quality data,\u201d Inform. Sciences, vol.259, pp.571-595, 2014.","DOI":"10.1016\/j.ins.2010.12.016"},{"key":"16","doi-asserted-by":"crossref","unstructured":"[16] D. Ryu, O. Choi, and J. Baik, \u201cValue-cognitive boosting with a support vector machine for cross-project defect prediction,\u201d Empir. Softw. Eng., vol.21, no.1, pp.43-71, 2016.","DOI":"10.1007\/s10664-014-9346-4"},{"key":"17","doi-asserted-by":"crossref","unstructured":"[17] M.A. Tahir, J. Kittler, and F. Yan, \u201cInverse random under sampling for class imbalance problem and its application to multi-label classification,\u201d Pattern Recogn., vol.45, no.10, pp.3738-3750, 2012.","DOI":"10.1016\/j.patcog.2012.03.014"},{"key":"18","doi-asserted-by":"crossref","unstructured":"[18] Y. Qian, Y. Liang, M. Li, G. Feng, and X. Shi, \u201cA resampling ensemble algorithm for classification of imbalance problems,\u201d Neurocomputing, vol.143, pp.57-67, 2014.","DOI":"10.1016\/j.neucom.2014.06.021"},{"key":"19","doi-asserted-by":"crossref","unstructured":"[19] \u00d6.F. Arar and K. Ayan, \u201cSoftware defect prediction using cost-sensitive neural network,\u201d Appl. Soft Comput., vol.33, pp.263-277, 2015.","DOI":"10.1016\/j.asoc.2015.04.045"},{"key":"20","doi-asserted-by":"crossref","unstructured":"[20] I.H. Laradji, M. Alshayeb, and L. Ghouti, \u201cSoftware defect prediction using ensemble learning on selected features,\u201d Inform. Softw. Technol., vol.58, pp.388-402, 2015.","DOI":"10.1016\/j.infsof.2014.07.005"},{"key":"21","doi-asserted-by":"crossref","unstructured":"[21] D. Rodriguez, I. Herraiz, R. Harrison, J. Dolado, and J.C. Riquelme, \u201cPreliminary comparison of techniques for dealing with imbalance in software defect prediction,\u201d Proc. 18th Int. Conf. Eval. Assess. Softw. Eng., pp.43:1-43:10, London, United Kingdom, May 2014.","DOI":"10.1145\/2601248.2601294"},{"key":"22","doi-asserted-by":"crossref","unstructured":"[22] M. Galar, A. Fern\u00e1ndez, E. Barrenechea, H. Bustince, and F. Herrera, \u201cA review on ensembles for the class imbalance problem: Bagging-, boosting-, and hybrid-based approaches,\u201d IEEE Trans. Syst. Man Cybern. Part C, vol.42, no.4, pp.463-484, 2012.","DOI":"10.1109\/TSMCC.2011.2161285"},{"key":"23","doi-asserted-by":"crossref","unstructured":"[23] C. Catal and B. Diri, \u201cInvestigating the effect of dataset size, metrics sets, and feature selection techniques on software fault prediction problem,\u201d Inform. Sciences, vol.179, no.8, pp.1040-1058, 2009.","DOI":"10.1016\/j.ins.2008.12.001"},{"key":"24","doi-asserted-by":"crossref","unstructured":"[24] Z. He, F. Shu, Y. Yang, M. Li, and Q. Wang, \u201cAn investigation on the feasibility of cross-project defect prediction,\u201d Automat. Softw. Eng., vol.19, no.2, pp.167-199, 2012.","DOI":"10.1007\/s10515-011-0090-3"},{"key":"25","doi-asserted-by":"crossref","unstructured":"[25] P. He, B. Li, X. Liu, J. Chen, and Y. Ma, \u201cAn empirical study on software defect prediction with a simplified metric set,\u201d Inform. Softw. Technol., vol.59, pp.170-190, 2015.","DOI":"10.1016\/j.infsof.2014.11.006"},{"key":"26","doi-asserted-by":"crossref","unstructured":"[26] J. Nam and S. Kim, \u201cCLAMI: Defect prediction on unlabeled datasets,\u201d Proc. 30th Int. Conf. Automat. Softw. Eng., pp.452-463, Lincoln, USA, Nov. 2015.","DOI":"10.1109\/ASE.2015.56"},{"key":"27","doi-asserted-by":"crossref","unstructured":"[27] L. Li and H. Leung, \u201cMining static code metrics for a robust prediction of software defect-proneness,\u201d Proc. 5th Int. Symp. Empir. Softw. Eng. &amp; Meas., pp.207-214, Banff, Canada, Sept. 2011.","DOI":"10.1109\/ESEM.2011.29"},{"key":"28","doi-asserted-by":"crossref","unstructured":"[28] K. Gao, T.M. Khoshgoftaar, H. Wang, and N. Seliya, \u201cChoosing software metrics for defect prediction: An investigation on feature selection techniques,\u201d Softw. Pract. Exper., vol.41, no.5, pp.579-606, 2011.","DOI":"10.1002\/spe.1043"},{"key":"29","doi-asserted-by":"crossref","unstructured":"[29] T.M. Khoshgoftaar, K. Gao, A. Napolitano, and R. Wald, \u201cA comparative study of iterative and non-iterative feature selection techniques for software defect prediction,\u201d Inform. Syst. Front., vol.16, no.5, pp.801-822, 2014.","DOI":"10.1007\/s10796-013-9430-0"},{"key":"30","doi-asserted-by":"crossref","unstructured":"[30] Y. Zhou, B. Xu, H. Leung, and L. Chen, \u201cAn in-depth study of the potentially confounding effect of class size in fault prediction,\u201d ACM Trans. Softw. Eng. Method., vol.23, no.1, pp.10:1-10:51, 2014.","DOI":"10.1145\/2556777"},{"key":"31","doi-asserted-by":"crossref","unstructured":"[31] D.W. Aha, D. Kibler, and M.K. Albert, \u201cInstance-based learning algorithms,\u201d Mach. Learn., vol.6, no.1, pp.37-66, 1991.","DOI":"10.1007\/BF00153759"},{"key":"32","doi-asserted-by":"crossref","unstructured":"[32] S. Le Cessie and J.C. Van Houwelingen, \u201cRidge estimators in logistic regression,\u201d Appl. Statist., vol.41, no.1, pp.191-201, 1992.","DOI":"10.2307\/2347628"},{"key":"33","doi-asserted-by":"crossref","unstructured":"[33] J.-G. Attali and G. Pag\u00e8s, \u201cApproximations of functions by a multilayer perceptron: A new approach,\u201d Neural Networks, vol.10, no.6, pp.1069-1081, 1997.","DOI":"10.1016\/S0893-6080(97)00010-5"},{"key":"34","doi-asserted-by":"crossref","unstructured":"[34] J. Huang and C.X. Ling, \u201cUsing AUC and accuracy in evaluating learning algorithms,\u201d IEEE Trans. Knowl. Data Eng., vol.17, no.3, pp.299-310, 2005.","DOI":"10.1109\/TKDE.2005.50"},{"key":"35","doi-asserted-by":"crossref","unstructured":"[35] Y. Jiang, J. Lin, B. Cukic, and T. Menzies, \u201cVariance analysis in software fault prediction models,\u201d Proc. 20th Int. Symp. Softw. Reliab. Eng., pp.99-108, Mysuru, India, Nov. 2009.","DOI":"10.1109\/ISSRE.2009.13"},{"key":"36","doi-asserted-by":"crossref","unstructured":"[36] E. Erturk and E.A. Sezer, \u201cA comparison of some soft computing methods for software fault prediction,\u201d Expert Syst. Appl., vol.42, no.4, pp.1872-1879, 2015.","DOI":"10.1016\/j.eswa.2014.10.025"},{"key":"37","doi-asserted-by":"crossref","unstructured":"[37] C. Tantithamthavorn, S. McIntosh, A.E. Hassan, and K. Matsumoto, \u201cAutomated parameter optimization of classification techniques for defect prediction models,\u201d Proc. 38th Int. Conf. Softw. Eng., pp.321-332, Austin, Texas, May 2016.","DOI":"10.1145\/2884781.2884857"},{"key":"38","doi-asserted-by":"crossref","unstructured":"[38] M. Shepperd, Q. Song, Z. Sun, and C. Mair, \u201cData quality: Some comments on the NASA software defect datasets,\u201d IEEE Trans. Softw. Eng., vol.39, no.9, pp.1208-1215, 2013.","DOI":"10.1109\/TSE.2013.11"},{"key":"39","doi-asserted-by":"crossref","unstructured":"[39] Y. Zhou, Y. Yang, B. Xu, H. Leung, and X. Zhou, \u201cSource code size estimation approaches for object-oriented systems from UML class diagrams: A comparative study,\u201d Inform. Softw. Technol., vol.56, no.2, pp.220-237, 2014.","DOI":"10.1016\/j.infsof.2013.09.003"},{"key":"40","doi-asserted-by":"crossref","unstructured":"[40] T.M. Khoshgoftaar, M. Golawala, and J. Van Hulse, \u201cAn empirical study of learning from imbalanced data using random forest,\u201d Proc. 19th Int. Conf. Tools with Artif. Intell., vol.2, pp.310-317, Patras, Greece, Oct. 2007.","DOI":"10.1109\/ICTAI.2007.46"}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E100.D\/2\/E100.D_2016EDP7204\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,9,18]],"date-time":"2019-09-18T00:26:26Z","timestamp":1568766386000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E100.D\/2\/E100.D_2016EDP7204\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017]]},"references-count":40,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2017]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2016edp7204","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"value":"0916-8532","type":"print"},{"value":"1745-1361","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017]]}}}