{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T19:36:29Z","timestamp":1784921789228,"version":"3.55.0"},"reference-count":30,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2019,2,28]],"date-time":"2019-02-28T00:00:00Z","timestamp":1551312000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>The kNN (k-nearest neighbors) classification algorithm is one of the most widely used non-parametric classification methods, however it is limited due to memory consumption related to the size of the dataset, which makes them impractical to apply to large volumes of data. Variations of this method have been proposed, such as condensed KNN which divides the training dataset into clusters to be classified, other variations reduce the input dataset in order to apply the algorithm. This paper presents a variation of the kNN algorithm, of the type structure less NN, to work with categorical data. Categorical data, due to their nature, can be compressed in order to decrease the memory requirements at the time of executing the classification. The method proposes a previous phase of compression of the data to then apply the algorithm on the compressed data. This allows us to maintain the whole dataset in memory which leads to a considerable reduction of the amount of memory required. Experiments and tests carried out on known datasets show the reduction in the volume of information stored in memory and maintain the accuracy of the classification. They also show a slight decrease in processing time because the information is decompressed in real time (on-the-fly) while the algorithm is running.<\/jats:p>","DOI":"10.3390\/e21030234","type":"journal-article","created":{"date-parts":[[2019,3,1]],"date-time":"2019-03-01T03:33:00Z","timestamp":1551411180000},"page":"234","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":58,"title":["Compressed kNN: K-Nearest Neighbors with Data Compression"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3564-8929","authenticated-orcid":false,"given":"Jaime","family":"Salvador\u2013Meneses","sequence":"first","affiliation":[{"name":"Facultad de Ingenier\u00eda, Ciencias F\u00edsicas y Matem\u00e1tica, Universidad Central del Ecuador, Quito 170129, Ecuador"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6808-4648","authenticated-orcid":false,"given":"Zoila","family":"Ruiz\u2013Chavez","sequence":"additional","affiliation":[{"name":"Facultad de Ingenier\u00eda, Ciencias F\u00edsicas y Matem\u00e1tica, Universidad Central del Ecuador, Quito 170129, Ecuador"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7798-3055","authenticated-orcid":false,"given":"Jose","family":"Garcia\u2013Rodriguez","sequence":"additional","affiliation":[{"name":"Computer Technology Department, University of Alicante, 03080 Alicante, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,2,28]]},"reference":[{"key":"ref_1","first-page":"447","article-title":"Compression, Clustering and Pattern Discovery in Very High Dimensional Discrete-Attribute Datasets","volume":"17","author":"Grama","year":"2005","journal-title":"Techniques"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"64","DOI":"10.1016\/j.patrec.2018.04.015","article-title":"A Label Compression Method for Online Multi-Label Classification","volume":"111","author":"Ahmadi","year":"2018","journal-title":"Pattern Recognit. Lett."},{"key":"ref_3","first-page":"1","article-title":"A Survey of Clustering Techniques","volume":"7","author":"Rai","year":"2010","journal-title":"Int. J. Comput. Appl."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1016\/j.dam.2004.04.004","article-title":"Discrete models for data imputation","volume":"144","author":"Bruni","year":"2004","journal-title":"Discret. Appl. Math."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Duan, Z., and Wang, L. (2017). K-dependence Bayesian classifier ensemble. Entropy, 19.","DOI":"10.3390\/e19120651"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Jim\u00e9nez, F., Mart\u00ednez, C., Miralles-Pechu\u00e1n, L., S\u00e1nchez, G., and Sciavicco, G. (2018). Multi-Objective Evolutionary Rule-Based Classification with Categorical Data. Entropy, 20.","DOI":"10.3390\/e20090684"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"621","DOI":"10.2165\/00002018-200730070-00010","article-title":"Principles of Data Mining","volume":"30","author":"Hand","year":"2007","journal-title":"Drug Saf."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Guo, G., Wang, H., Bell, D., Bi, Y., and Greer, K. (2003). KNN Model-Based Approach in Classification. On the Move to Meaningful Internet Systems 2003: CoopIS, DOA, and ODBASE, Springer.","DOI":"10.1007\/978-3-540-39964-3_62"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Ouyang, J., Luo, H., Wang, Z., Tian, J., Liu, C., and Sheng, K. (2010, January 8\u201310). FPGA implementation of GZIP compression and decompression for IDC services. Proceedings of the 2010 International Conference on Field-Programmable Technology, FPT\u201910, Beijing, China.","DOI":"10.1109\/FPT.2010.5681489"},{"key":"ref_10","first-page":"302","article-title":"Survey of Nearest Neighbor techniques","volume":"8","author":"Bhatia","year":"2010","journal-title":"Int. J. Comput. Sci. Inf. Sec."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1483","DOI":"10.1016\/j.neucom.2008.11.026","article-title":"K nearest neighbours with mutual information for simultaneous classification and missing data imputation","volume":"72","author":"Verleysen","year":"2009","journal-title":"Neurocomputing"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"105","DOI":"10.1016\/j.artmed.2010.05.002","article-title":"Missing data imputation using statistical and machine learning methods in a real breast cancer problem","volume":"50","author":"Jerez","year":"2010","journal-title":"Artif. Intell. Med."},{"key":"ref_13","first-page":"44","article-title":"Comparison Classifier of Condensed KNN and K-Nearest Neighborhood Error Rate Method","volume":"2","author":"James","year":"2012","journal-title":"Comput. Sci. Technol. Int. J."},{"key":"ref_14","first-page":"622","article-title":"Stochastic Neighbor Compression","volume":"32","author":"Kusner","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_15","first-page":"1331","article-title":"ProtoNN: Compressed and Accurate kNN for Resource-scarce Devices","volume":"70","author":"Gupta","year":"2017","journal-title":"Icml2017"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2047","DOI":"10.1109\/TNNLS.2015.2451151","article-title":"Space Structure and Clustering of Categorical Data","volume":"27","author":"Qian","year":"2016","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Boriah, S., Chandola, V., and Kumar, V. (2008, January 24\u201326). Similarity Measures for Categorical Data: A Comparative Evaluation. Proceedings of the 2008 SIAM International Conference on Data Mining, Atlanta, GA, USA.","DOI":"10.1137\/1.9781611972788.22"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Alamuri, M., Surampudi, B.R., and Negi, A. (2014, January 6\u201311). A survey of distance\/similarity measures for categorical data. Proceedings of the International Joint Conference on Neural Networks, BeiJing, China.","DOI":"10.1109\/IJCNN.2014.6889941"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"519","DOI":"10.1080\/713827181","article-title":"An analysis of four missing data treatment methods for supervised learning","volume":"17","author":"Batista","year":"2003","journal-title":"Appl. Artif. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"125","DOI":"10.1016\/j.compbiomed.2015.02.006","article-title":"Missing data imputation on the 5-year survival prediction of breast cancer patients with unknown discrete values","volume":"59","author":"Abreu","year":"2015","journal-title":"Comput. Biol. Med."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"367","DOI":"10.15623\/ijret.2014.0310059","article-title":"Parallel KNN on GPU Architecture Using OpenCL","volume":"3","author":"Nikam","year":"2014","journal-title":"Int. J. Res. Eng. Technol."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Salvador-Meneses, J., Ruiz-Chavez, Z., and Garcia-Rodriguez, J. (2018, January 18\u201320). Low Level Big Data Compression. Proceedings of the 10th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management, Seville, Spain.","DOI":"10.5220\/0007228003530358"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"811","DOI":"10.24201\/edu.v31i3.15","article-title":"El formato Redatam","volume":"31","year":"2016","journal-title":"Estud. Demogr. Urbanos"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Salvador-Meneses, J., Ruiz-Chavez, Z., and Garcia-Rodriguez, J. (2018, January 18\u201320). Low Level Big Data Processing. Proceedings of the 10th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management, Seville, Spain.","DOI":"10.5220\/0007227103470352"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"2630","DOI":"10.1098\/rspa.2011.0704","article-title":"Statistical approach to normalization of feature vectors and clustering of mixed datasets","volume":"468","author":"Pham","year":"2012","journal-title":"Proc. R. Soc. A"},{"key":"ref_26","first-page":"236","article-title":"Breast Cancer Diagnosis on Three Different Datasets Using Multi-Classifiers Gouda","volume":"1","author":"Salama","year":"2012","journal-title":"Int. J. Comput. Inf. Technol."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1109\/LCA.2015.2434872","article-title":"Fast Bulk Bitwise and and or in DRAM","volume":"14","author":"Seshadri","year":"2015","journal-title":"IEEE Comput. Archit. Lett."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Yin, H., Camacho, D., Novais, P., and Tall\u00f3n-Ballesteros, A.J. (2018). Categorical Big Data Processing. Intelligent Data Engineering and Automated Learning\u2014IDEAL 2018, Springer International Publishing.","DOI":"10.1007\/978-3-030-03493-1"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Beygelzimer, A., Kakade, S., and Langford, J. (2006, January 25\u201329). Cover trees for nearest neighbor. Proceedings of the 23rd International Conference on Machine Learning\u2014ICML \u201906, Pittsburgh, PA, USA.","DOI":"10.1145\/1143844.1143857"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Yin, H., Camacho, D., Novais, P., and Tall\u00f3n-Ballesteros, A.J. (2018). Machine Learning Methods Based Preprocessing to Improve Categorical Data Classification. Intelligent Data Engineering and Automated Learning\u2014IDEAL 2018, Springer International Publishing.","DOI":"10.1007\/978-3-030-03493-1"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/21\/3\/234\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:35:25Z","timestamp":1760186125000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/21\/3\/234"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,2,28]]},"references-count":30,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2019,3]]}},"alternative-id":["e21030234"],"URL":"https:\/\/doi.org\/10.3390\/e21030234","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,2,28]]}}}