{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T07:34:06Z","timestamp":1784705646083,"version":"3.55.0"},"reference-count":37,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2021,1,21]],"date-time":"2021-01-21T00:00:00Z","timestamp":1611187200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Data"],"abstract":"<jats:p>It is recognized that the performance of any prediction model is a function of several factors. One of the most significant factors is the adopted preprocessing techniques. In other words, preprocessing is an essential process to generate an effective and efficient classification model. This paper investigates the impact of the most widely used preprocessing techniques, with respect to numerical features, on the performance of classification algorithms. The effect of combining various normalization techniques and handling missing values strategies is assessed on eighteen benchmark datasets using two well-known classification algorithms and adopting different performance evaluation metrics and statistical significance tests. According to the reported experimental results, the impact of the adopted preprocessing techniques varies from one classification algorithm to another. In addition, a statistically significant difference between the considered data preprocessing techniques is demonstrated.<\/jats:p>","DOI":"10.3390\/data6020011","type":"journal-article","created":{"date-parts":[[2021,1,22]],"date-time":"2021-01-22T03:31:30Z","timestamp":1611286290000},"page":"11","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":98,"title":["The Effect of Preprocessing Techniques, Applied to Numeric Features, on Classification Algorithms\u2019 Performance"],"prefix":"10.3390","volume":"6","author":[{"given":"Esra\u2019a","family":"Alshdaifat","sequence":"first","affiliation":[{"name":"Department of Computer Information System, Faculty of Prince Al-Hussein Bin Abdallah II For Information Technology, The Hashemite University, P.O. Box 330127, Zarqa 13133, Jordan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Doa\u2019a","family":"Alshdaifat","sequence":"additional","affiliation":[{"name":"Department of Computer Information System, Faculty of Prince Al-Hussein Bin Abdallah II For Information Technology, The Hashemite University, P.O. Box 330127, Zarqa 13133, Jordan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ayoub","family":"Alsarhan","sequence":"additional","affiliation":[{"name":"Department of Computer Information System, Faculty of Prince Al-Hussein Bin Abdallah II For Information Technology, The Hashemite University, P.O. Box 330127, Zarqa 13133, Jordan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fairouz","family":"Hussein","sequence":"additional","affiliation":[{"name":"Department of Computer Information System, Faculty of Prince Al-Hussein Bin Abdallah II For Information Technology, The Hashemite University, P.O. Box 330127, Zarqa 13133, Jordan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9700-4862","authenticated-orcid":false,"given":"Subhieh Moh\u2019d Faraj S.","family":"El-Salhi","sequence":"additional","affiliation":[{"name":"Department of Computer Information System, Faculty of Prince Al-Hussein Bin Abdallah II For Information Technology, The Hashemite University, P.O. Box 330127, Zarqa 13133, Jordan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,1,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Kuhn, M., and Johnson, K. (2013). Data Pre-processing. Applied Predictive Modeling, Springer.","DOI":"10.1007\/978-1-4614-6849-3"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"781","DOI":"10.1016\/j.ejor.2005.07.023","article-title":"The impact of preprocessing on data mining: An evaluation of classifier sensitivity in direct marketing","volume":"173","author":"Crone","year":"2006","journal-title":"Eur. J. Oper. Res."},{"key":"ref_3","first-page":"11","article-title":"Investigations on Impact of Feature Normalization Techniques on Classifier\u2019s Performance in Breast Tumor Classification","volume":"116","author":"KumarSingh","year":"2015","journal-title":"Int. J. Comput. Appl."},{"key":"ref_4","first-page":"27","article-title":"Assessment of Normalization Techniques on the Accuracy of Hyperspectral Data Clustering","volume":"XLII-4\/W4","author":"Babadi","year":"2017","journal-title":"ISPRS Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci."},{"key":"ref_5","unstructured":"Jiawei, H., Micheline, K., and Jian, P. (2011). Data Mining: Concepts and Techniques, Morgan Kaufmann."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Jayalskshmi, T., and Santhakumaran, A. (2010, January 12\u201313). Impact of Preprocessing for Diagnosis of Diabetes Mellitus Using Artificial Neural Networks. Proceedings of the 2010 Second International Conference on Machine Learning and Computing, Bangalore, India.","DOI":"10.1109\/ICMLC.2010.65"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"e4584","DOI":"10.7717\/peerj.4584","article-title":"Empirical evaluation of data normalization methods for molecular classification","volume":"6","author":"Huang","year":"2018","journal-title":"PeerJ"},{"key":"ref_8","unstructured":"Rozenstein, O., Paz Kagan, T., Salbach, C., and Karnieli, A. (2014). Comparing the Effect of Preprocessing Transformations on Methods of Land-Use Classification Derived From Spectral Soil Measurements. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., 1\u201312."},{"key":"ref_9","first-page":"311","article-title":"Effect of Missing Values on Data Classification","volume":"4","author":"Baitharu","year":"2013","journal-title":"J. Emerg. Trends Eng. Appl. Sci. (JETEAS)"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"822","DOI":"10.1016\/j.joca.2012.03.005","article-title":"Consequences of handling missing data for treatment response in osteoarthritis: A simulation study","volume":"20","author":"Olsen","year":"2012","journal-title":"Osteoarthr. Cartil."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Palumbo, F., Montanari, A., and Vichi, M. (2017). Missing Data Imputation and Its Effect on the Accuracy of Classification, Springer International Publishing. Data Science.","DOI":"10.1007\/978-3-319-55723-6"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1214\/aoms\/1177731944","article-title":"A Comparison of Alternative Tests of Significance for the Problem of m Rankings","volume":"11","author":"Friedman","year":"1940","journal-title":"Ann. Math. Stat."},{"key":"ref_13","unstructured":"Nemenyi, P. (1963). Distribution-Free Multiple Comparisons, Princeton University."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1016\/j.bdr.2019.04.001","article-title":"Anomaly Detection and Repair for Accurate Predictions in Geo-distributed Big Data","volume":"16","author":"Corizzo","year":"2019","journal-title":"Big Data Res."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Banks, D., McMorris, F.R., Arabie, P., and Gaul, W. (2004). The Treatment of Missing Values and its Effect on Classifier Accuracy. Classification, Clustering, and Data Mining Applications, Springer.","DOI":"10.1007\/978-3-642-17103-1"},{"key":"ref_16","first-page":"1623","article-title":"Handling Missing Values when Applying Classification Models","volume":"8","author":"Provost","year":"2007","journal-title":"J. Mach. Learn. Res."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1007\/s00521-009-0295-6","article-title":"Pattern Classification with Missing Data: A Review","volume":"19","year":"2010","journal-title":"Neural Comput. Appl."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Osborne, J. (2013). Best Practices in Data Cleaning: A Complete Guide to Everything You Need to Do before and after Collecting Your Data, SAGE.","DOI":"10.4135\/9781452269948"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"5621","DOI":"10.1016\/j.eswa.2015.02.050","article-title":"Hybrid prediction model with missing value imputation for medical data","volume":"42","author":"Purwar","year":"2015","journal-title":"Expert Syst. Appl."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"77","DOI":"10.1007\/s10115-011-0424-2","article-title":"On the choice of the best imputation methods for missing values considering three groups of classification methods","volume":"32","author":"Luengo","year":"2012","journal-title":"Knowl. Inf. Syst."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"105351","DOI":"10.1109\/ACCESS.2020.2999960","article-title":"Data Repair Without Prior Knowledge Using Deep Convolutional Neural Networks","volume":"8","author":"Qie","year":"2020","journal-title":"IEEE Access"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"799","DOI":"10.1080\/014311697218764","article-title":"An evaluation of some factors affecting the accuracy of classification by an artificial neural network","volume":"18","author":"Foody","year":"1997","journal-title":"Int. J. Remote Sens."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"26","DOI":"10.5539\/mas.v6n10p26","article-title":"Ranking Normalization Methods for Improving the Accuracy of SVM Algorithm by DEA Method","volume":"6","author":"Eftekhary","year":"2012","journal-title":"Mod. Appl. Sci."},{"key":"ref_24","unstructured":"Wohlrab, L., and F\u00fcrnkranz, J. (2009). A Comparison of Strategies for Handling Missing Values in Rule Learning, Knowledge Engineering Group, Technische Universit\u00e4t Darmstadt. Technical Report."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1082","DOI":"10.1007\/s11704-016-5203-5","article-title":"Impact of preprocessing on medical data classification","volume":"10","author":"Almuhaideb","year":"2016","journal-title":"Front. Comput. Sci."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1515\/jaiscr-2018-0002","article-title":"Classifiers accuracy improvement based on missing data imputation","volume":"8","author":"Jordanov","year":"2018","journal-title":"J. Artif. Intell. Soft Comput. Res."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"525","DOI":"10.3102\/00346543074004525","article-title":"Missing Data in Educational Research: A Review of Reporting Practices and Suggestions for Improvement","volume":"74","author":"Peugh","year":"2004","journal-title":"Rev. Educ. Res."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Aleryani, A., Wang, W., and De La Iglesia, B. (2018). Dealing with Missing Data and Uncertainty in the Context of Data Mining. Hybrid Artificial Intelligent Systems, Springer.","DOI":"10.1007\/978-3-319-92639-1_24"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Kim, T., Ko, W., and Kim, J. (2019). Analysis and Impact Evaluation of Missing Data Imputation in Day-ahead PV Generation Forecasting. Appl. Sci., 9.","DOI":"10.3390\/app9010204"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"653","DOI":"10.1016\/j.apenergy.2017.01.063","article-title":"Short-term wind speed forecasting by spectral analysis from long-term observations with missing values","volume":"191","author":"Filik","year":"2017","journal-title":"Appl. Energy"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"701","DOI":"10.1016\/j.ins.2020.08.003","article-title":"Multi-aspect renewable energy forecasting","volume":"546","author":"Corizzo","year":"2021","journal-title":"Inf. Sci."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"538","DOI":"10.1109\/TSTE.2017.2747765","article-title":"Short-Term Spatio-Temporal Forecasting of Photovoltaic Power Production","volume":"9","author":"Agoua","year":"2018","journal-title":"IEEE Trans. Sustain. Energy"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"698","DOI":"10.1007\/s10618-018-0605-7","article-title":"Spatial autocorrelation and entropy for renewable energy forecasting","volume":"33","author":"Ceci","year":"2019","journal-title":"Data Min. Knowl. Discov."},{"key":"ref_34","unstructured":"Lichman, M. (2019, June 01). UCI Machine Learning Repository. Available online: http:\/\/archive.ics.uci.edu\/ml."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"299","DOI":"10.1109\/TKDE.2005.50","article-title":"Using AUC and Accuracy in Evaluating Learning Algorithms","volume":"17","author":"Huang","year":"2005","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1145\/1656274.1656278","article-title":"The WEKA Data Mining Software: An Update","volume":"11","author":"Hall","year":"2009","journal-title":"SIGKDD Explor. Newsl."},{"key":"ref_37","first-page":"1","article-title":"Statistical Comparisons of Classifiers over Multiple Data Sets","volume":"7","year":"2006","journal-title":"J. Mach. Learn. Res."}],"container-title":["Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2306-5729\/6\/2\/11\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:13:47Z","timestamp":1760159627000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2306-5729\/6\/2\/11"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,1,21]]},"references-count":37,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2021,2]]}},"alternative-id":["data6020011"],"URL":"https:\/\/doi.org\/10.3390\/data6020011","relation":{},"ISSN":["2306-5729"],"issn-type":[{"value":"2306-5729","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,1,21]]}}}