{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,16]],"date-time":"2026-01-16T00:07:30Z","timestamp":1768522050617,"version":"3.49.0"},"reference-count":43,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2025,3,7]],"date-time":"2025-03-07T00:00:00Z","timestamp":1741305600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,3,7]],"date-time":"2025-03-07T00:00:00Z","timestamp":1741305600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001807","name":"Funda\u00e7\u00e3o de Amparo \u00e0 Pesquisa do Estado de S\u00e3o Paulo","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100001807","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002322","name":"Coordena\u00e7\u00e3o de Aperfei\u00e7oamento de Pessoal de N\u00edvel Superior","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100002322","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003593","name":"Conselho Nacional de Desenvolvimento Cient\u00edfico e Tecnol\u00f3gico","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100003593","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100008047","name":"Carnegie Mellon University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100008047","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Data Min Knowl Disc"],"published-print":{"date-parts":[[2025,5]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>How to spot outliers in a large, unlabeled dataset with both numerical and categorical attributes? How to do it in a fast and scalable way? Outlier detection has many applications; it is covered therefore by an extensive literature. The distance-based detectors are the most popular ones. However, they still have two major drawbacks: (a) the intensive neighborhood search that takes hours or even days to complete in large data, and; (b) the inability to process categorical attributes. This paper tackles both problems by presenting <jats:sc>HySortOD<\/jats:sc>: a new, fast and scalable detector for numerical and categorical data. Our main focus is the analysis of datasets with many instances, and a low-to-moderate number of attributes. We studied dozens of real, benchmark datasets with up to <jats:bold>one million instances<\/jats:bold>; <jats:sc>HySortOD<\/jats:sc> outperformed <jats:bold>nine competitors<\/jats:bold> from the state of the art in runtime, being up to <jats:bold>six orders of magnitude faster<\/jats:bold> in large data, while maintaining high accuracy. Finally, we also performed an extensive experimental evaluation that confirms the ability of our method to obtain high-quality results from both real and synthetic datasets with categorical attributes.<\/jats:p>","DOI":"10.1007\/s10618-024-01084-1","type":"journal-article","created":{"date-parts":[[2025,3,7]],"date-time":"2025-03-07T03:17:13Z","timestamp":1741317433000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Efficient outlier detection in numerical and categorical data"],"prefix":"10.1007","volume":"39","author":[{"given":"Eug\u00eanio F.","family":"Cabral","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Braulio V.","family":"S\u00e1nchez Vinces","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guilherme D. F.","family":"Silva","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"J\u00f6rg","family":"Sander","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6795-3004","authenticated-orcid":false,"given":"Robson L. F.","family":"Cordeiro","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,3,7]]},"reference":[{"key":"1084_CR1","doi-asserted-by":"crossref","unstructured":"Jauhri A, McDanel B (2015) Connor Chris (2015) Outlier Detection for Large Scale Manufacturing Processes. In: IEEE Big Data. ( pp. 2771\u20132774)","DOI":"10.1109\/BigData.2015.7364079"},{"key":"1084_CR2","doi-asserted-by":"crossref","unstructured":"Anbarasi MS, Dhivya S (2017) Fraud Detection Using Outlier Predictor in Health Insurance data. In: ICICES, p. 6","DOI":"10.1109\/ICICES.2017.8070750"},{"issue":"7","key":"1084_CR3","first-page":"229","volume":"118","author":"D Tripathi","year":"2018","unstructured":"Tripathi D, Lone T, Sharma Y, Dwivedi S (2018) Credit Card fraud detection using local outlier factor. IJPAM 118(7):229\u2013234","journal-title":"IJPAM"},{"key":"1084_CR4","doi-asserted-by":"publisher","first-page":"568","DOI":"10.1016\/j.chb.2017.04.001","volume":"73","author":"PV Bindu","year":"2017","unstructured":"Bindu PV, Thilagam P, Ahuja D (2017) Discovering suspicious behavior in multilayer social networks. Comput Human Behav 73:568\u2013582","journal-title":"Comput Human Behav"},{"key":"1084_CR5","doi-asserted-by":"crossref","unstructured":"Jabez J, Muthukumar B (2015) Intrusion Detection System (ids): Anomaly Detection Using Outlier Detection Approach. In: ICICC, pp. 338\u2013346","DOI":"10.1016\/j.procs.2015.04.191"},{"key":"1084_CR6","doi-asserted-by":"publisher","first-page":"193","DOI":"10.1007\/s10462-012-9370-y","volume":"43","author":"N Shahid","year":"2015","unstructured":"Shahid N, Naqvi IH, Qaisar SB (2015) Characteristics and classification of outlier detection techniques for wireless sensor networks in harsh environments: survey. Artif Intell Rev 43:193\u2013228","journal-title":"Artif Intell Rev"},{"issue":"4","key":"1084_CR7","first-page":"891","volume":"30","author":"G Campos","year":"2016","unstructured":"Campos G, Zimek A, Sander J, Campello R, Micenkov\u00e1 B, Schubert E, Assent I, Houle ME (2016) On the evaluation of unsupervised outlier detection: measures, datasets, and an empirical study. DMKD 30(4):891\u2013927","journal-title":"DMKD"},{"key":"1084_CR8","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-47578-3","volume-title":"Outlier Analysis","author":"CC Aggarwal","year":"2017","unstructured":"Aggarwal CC (2017) Outlier Analysis. Springer, Switzerland"},{"issue":"1","key":"1084_CR9","doi-asserted-by":"publisher","first-page":"27","DOI":"10.1214\/aoms\/1177729885","volume":"21","author":"FE Grubbs","year":"1950","unstructured":"Grubbs FE (1950) Sample criteria for testing outlying observations. Ann Math Statis 21(1):27\u201358","journal-title":"Ann Math Statis"},{"key":"1084_CR10","unstructured":"Knorr EM, Ng RT (1998) Algorithms for mining distance-based outliers in large datasets. In: Proceedings of the 24rd International Conference on Very Large Data Bases. Morgan Kaufmann Publishers Inc: San Francisco (pp. 392\u2013403)"},{"issue":"2","key":"1084_CR11","doi-asserted-by":"publisher","first-page":"427","DOI":"10.1145\/335191.335437","volume":"29","author":"S Ramaswamy","year":"2000","unstructured":"Ramaswamy S, Rastogi R, Shim K (2000) Efficient algorithms for mining outliers from large data sets. SIGMOD 29(2):427\u2013438","journal-title":"SIGMOD"},{"issue":"2","key":"1084_CR12","doi-asserted-by":"publisher","first-page":"93","DOI":"10.1145\/335191.335388","volume":"29","author":"MM Breunig","year":"2000","unstructured":"Breunig MM, Kriegel H-P, Ng RT, Sander J (2000) LOF: Identifying density-based local outliers. SIGMOD 29(2):93\u2013104","journal-title":"SIGMOD"},{"key":"1084_CR13","doi-asserted-by":"crossref","unstructured":"Angiulli F, Pizzuti C (2002) Fast Outlier Detection in High Dimensional Spaces. In: European Conference on PKDD, pp. 15\u201327 (2002)","DOI":"10.1007\/3-540-45681-3_2"},{"key":"1084_CR14","doi-asserted-by":"crossref","unstructured":"Papadimitriou S, Kitagawa H, Gibbons PB, Faloutsos C (2003) LOCI Fast Outlier Detection. In: ICDE, pp. 315\u2013326","DOI":"10.1007\/978-3-540-45072-6_12"},{"key":"1084_CR15","doi-asserted-by":"crossref","unstructured":"Hautam\u00e4ki V, K\u00e4rkk\u00e4inen I, Fr\u00e4nti P (2004) Outlier Detection Using k-Nearest Neighbour Graph. In: ICPR, pp. 430\u2013433. IEEE","DOI":"10.1109\/ICPR.2004.1334558"},{"key":"1084_CR16","doi-asserted-by":"crossref","unstructured":"Amagata D, Onizuka M, Hara T. Fast and exact outlier detection in metric spaces: a proximity graph-based approach. InProceedings of the 2021 International Conference on Management of Data 2021 Jun 9 (pp. 36-48)","DOI":"10.1145\/3448016.3452782"},{"issue":"4","key":"1084_CR17","doi-asserted-by":"publisher","first-page":"797","DOI":"10.1007\/s00778-022-00729-1","volume":"31","author":"D Amagata","year":"2022","unstructured":"Amagata D, Onizuka M, Hara T (2022) Fast, exact, and parallel-friendly outlier detection algorithms with proximity graph in metric spaces. VLDB J. 31(4):797\u2013821","journal-title":"VLDB J."},{"issue":"1\u20132","key":"1084_CR18","doi-asserted-by":"publisher","first-page":"1469","DOI":"10.14778\/1920841.1921021","volume":"3","author":"GH Orair","year":"2010","unstructured":"Orair GH, Teixeira CHC, Meira W, Wang Y, Parthasarathy S (2010) Distance-based outlier detection: consolidation and renewed bearing. VLDB Endow 3(1\u20132):1469\u20131480","journal-title":"VLDB Endow"},{"issue":"4","key":"1084_CR19","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1371\/journal.pone.0152173","volume":"11","author":"M Goldstein","year":"2016","unstructured":"Goldstein M, Uchida S (2016) A comparative evaluation of unsupervised anomaly detection algorithms for multivariate data. PLoS ONE 11(4):1\u201331","journal-title":"PLoS ONE"},{"key":"1084_CR20","doi-asserted-by":"crossref","unstructured":"Kirner E, Schubert E, Zimek A (2017) Good and Bad Neighborhood Approximations for Outlier Detection Ensembles. In: SISAP, pp. 173\u2013187","DOI":"10.1007\/978-3-319-68474-1_12"},{"key":"1084_CR21","doi-asserted-by":"publisher","unstructured":"Akoglu L, Tong H, Vreeken J, Faloutsos C (2012) Fast and reliable anomaly detection in categorical data. In: Proceedings of the 21st ACM International Conference on Information and Knowledge Management. CIKM \u201912, pp. 415\u2013424. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/2396761.2396816. https:\/\/doi.org\/10.1145\/2396761.2396816","DOI":"10.1145\/2396761.2396816"},{"key":"1084_CR22","doi-asserted-by":"publisher","unstructured":"Tang G, Bailey J, Pei J, Dong G (2013) Mining multidimensional contextual outliers from categorical relational data. In: Proceedings of the 25th International Conference on Scientific and Statistical Database Management. SSDBM. Association for Computing Machinery, New York, NY, USA (2013). https:\/\/doi.org\/10.1145\/2484838.2484883. https:\/\/doi.org\/10.1145\/2484838.2484883","DOI":"10.1145\/2484838.2484883"},{"key":"1084_CR23","volume-title":"Digital Design and Computer Architecture","author":"DM Harris","year":"2012","unstructured":"Harris DM, Harris SL (2012) Digital Design and Computer Architecture. Morgan Kaufmann, Burlington"},{"issue":"1","key":"1084_CR24","doi-asserted-by":"publisher","first-page":"256","DOI":"10.1109\/TVCG.2017.2744685","volume":"24","author":"L Wilkinson","year":"2018","unstructured":"Wilkinson L (2018) Visualizing big data outliers through distributed aggregation. IEEE Trans. Vis. Comput. Graph. 24(1):256\u2013266. https:\/\/doi.org\/10.1109\/TVCG.2017.2744685","journal-title":"IEEE Trans. Vis. Comput. Graph."},{"key":"1084_CR25","doi-asserted-by":"publisher","unstructured":"Liu FT, Ting KM, Zhou Z (2008) Isolation forest. In: Proceedings of the 8th IEEE International Conference on Data Mining (ICDM 2008), December 15-19, 2008, Pisa, Italy, pp. 413\u2013422. IEEE Computer Society. https:\/\/doi.org\/10.1109\/ICDM.2008.17. https:\/\/doi.org\/10.1109\/ICDM.2008.17","DOI":"10.1109\/ICDM.2008.17"},{"issue":"1","key":"1084_CR26","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1145\/2133360.2133363","volume":"6","author":"FT Liu","year":"2012","unstructured":"Liu FT, Ting KM, Zhou Z (2012) Isolation-based anomaly detection. ACM Trans. Knowl. Discov. Data 6(1):3\u20131339. https:\/\/doi.org\/10.1145\/2133360.2133363","journal-title":"ACM Trans. Knowl. Discov. Data"},{"issue":"1","key":"1084_CR27","first-page":"190","volume":"28","author":"E Schubert","year":"2014","unstructured":"Schubert E, Zimek A, Kriegel HP (2014) Local outlier detection reconsidered: a generalized view on locality with applications to spatial, video, and network outlier detection. DMKD 28(1):190\u2013237","journal-title":"DMKD"},{"key":"1084_CR28","doi-asserted-by":"crossref","unstructured":"Fraideinberze AC, Rodrigues JF, Cordeiro RLF (2016) Effective and Unsupervised Fractal-based Feature Selection for Very Large Datasets: removing linear and non-linear attribute correlations. In: ICDM Workshops, pp. 615\u2013622 (2016)","DOI":"10.1109\/ICDMW.2016.0093"},{"issue":"2","key":"1084_CR29","first-page":"244","volume":"14","author":"C Traina Junior","year":"2002","unstructured":"Traina Junior C, Traina AJM, Faloutsos C, Seeger B (2002) Fast indexing and visualization of metric data sets using slim-trees. TKDE 14(2):244\u2013260","journal-title":"TKDE"},{"issue":"1","key":"1084_CR30","first-page":"59","volume":"24","author":"D Mo","year":"2012","unstructured":"Mo D, Huang SH (2012) Fractal-based intrinsic dimension estimation and its application in dimensionality reduction. TKDE 24(1):59\u201371","journal-title":"TKDE"},{"issue":"12","key":"1084_CR31","doi-asserted-by":"publisher","first-page":"1089","DOI":"10.14778\/2994509.2994526","volume":"9","author":"L Tran","year":"2016","unstructured":"Tran L, Fan L, Shahabi C (2016) Distance-based outlier detection in data streams. VLDB Endow. 9(12):1089\u20131100","journal-title":"VLDB Endow."},{"key":"1084_CR32","doi-asserted-by":"publisher","first-page":"1303","DOI":"10.14778\/3342263.3342269","volume":"12","author":"S Yoon","year":"2018","unstructured":"Yoon S, Lee J-G, Byung BS (2018) NETS: Extremely fast outlier detection from a data stream via set-based processing. Proceed VLDB Endow. 12:1303\u20131315","journal-title":"Proceed VLDB Endow."},{"key":"1084_CR33","doi-asserted-by":"publisher","unstructured":"Cabral EF, Cordeiro RLF (2020) Fast and scalable outlier detection with sorted hypercubes. In: Proceedings of the 29th ACM International Conference on Information & Knowledge Management. CIKM \u201920, pp. 95\u2013104. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3340531.3412033","DOI":"10.1145\/3340531.3412033"},{"key":"1084_CR34","doi-asserted-by":"publisher","unstructured":"Faloutsos C, Kamel I (1994) Beyond uniformity and independence: Analysis of r-trees using the concept of fractal dimension. In: Proceedings of the Thirteenth ACM SIGACT-SIGMOD-SIGART Symposium on Principles of Database Systems. PODS \u201994, pp. 4\u201313. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/182591.182593. https:\/\/doi.org\/10.1145\/182591.182593","DOI":"10.1145\/182591.182593"},{"key":"1084_CR35","doi-asserted-by":"publisher","DOI":"10.1515\/9783110198164","volume-title":"An Introduction to Abstract Algebra","author":"DJS Robinson","year":"2003","unstructured":"Robinson DJS (2003) An Introduction to Abstract Algebra, 2nd edn. Walter de Gruyter, Berlin","edition":"2"},{"key":"1084_CR36","unstructured":"Marcus B (2022) Lecture Notes, Math 342: Algebra and Coding Theory. https:\/\/personal.math.ubc.ca\/~marcus\/Math342\/Math342_Lectures3-4.pdf. Day of access: 17 Nov"},{"issue":"4","key":"1084_CR37","doi-asserted-by":"publisher","first-page":"483","DOI":"10.1007\/s00778-005-0178-0","volume":"16","author":"C Traina Junior","year":"2007","unstructured":"Traina Junior C, Santos Filho RF, Traina AJM, Vieira MR, Faloutsos C (2007) The omni-family of all-purpose access methods: a simple and effective way to make similarity search more efficient. VLDB J. 16(4):483\u2013505. https:\/\/doi.org\/10.1007\/s00778-005-0178-0","journal-title":"VLDB J."},{"key":"1084_CR38","doi-asserted-by":"publisher","DOI":"10.1063\/1.2810323","volume-title":"Fractals, Chaos, Power Laws. Minutes from an Infinite Paradise","author":"MR Schroeder","year":"1991","unstructured":"Schroeder MR (1991) Fractals, Chaos, Power Laws. Minutes from an Infinite Paradise. W H Freeman, New York"},{"issue":"2","key":"1084_CR39","doi-asserted-by":"publisher","first-page":"416","DOI":"10.1214\/aos\/1176346150","volume":"11","author":"J Rissanen","year":"1983","unstructured":"Rissanen J (1983) A universal prior for integers and estimation by minimum description length. Ann Statis 11(2):416\u2013431. https:\/\/doi.org\/10.1214\/aos\/1176346150","journal-title":"Ann Statis"},{"key":"1084_CR40","doi-asserted-by":"crossref","unstructured":"Chakrabarti D, Papadimitriou S, Modha DS, Faloutsos C (2004) Fully automatic cross-associations. In: KDD, pp. 79\u201388. ACM","DOI":"10.1145\/1014052.1014064"},{"key":"1084_CR41","doi-asserted-by":"crossref","unstructured":"Achtert E, Hettab A, Kriegel HP, Schubert E, Zimek A. Spatial outlier detection: Data, algorithms, visualizations. InAdvances in Spatial and Temporal Databases: 12th International Symposium, SSTD 2011, Minneapolis, MN, USA, August 24-26, 2011, Proceedings 12 2011 (pp. 512-516). Springer: Berlin","DOI":"10.1007\/978-3-642-22922-0_41"},{"key":"1084_CR42","doi-asserted-by":"crossref","unstructured":"Xia C, Lu H, Ooi BC, Hu J. Gorder: an efficient method for knn join processing. InProceedings of the Thirtieth international conference on Very large data bases-Volume 30 2004 Aug 31 (pp. 756-767).","DOI":"10.1016\/B978-012088469-8\/50067-X"},{"issue":"4","key":"1084_CR43","doi-asserted-by":"publisher","first-page":"561","DOI":"10.1007\/s00778-012-0305-7","volume":"22","author":"DV Kalashnikov","year":"2013","unstructured":"Kalashnikov DV (2013) Super-EGO: fast multi-dimensional similarity join. VLDB J 22(4):561\u2013585","journal-title":"VLDB J"}],"container-title":["Data Mining and Knowledge Discovery"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10618-024-01084-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10618-024-01084-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10618-024-01084-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,26]],"date-time":"2025-04-26T01:35:21Z","timestamp":1745631321000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10618-024-01084-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,7]]},"references-count":43,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,5]]}},"alternative-id":["1084"],"URL":"https:\/\/doi.org\/10.1007\/s10618-024-01084-1","relation":{},"ISSN":["1384-5810","1573-756X"],"issn-type":[{"value":"1384-5810","type":"print"},{"value":"1573-756X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,7]]},"assertion":[{"value":"22 September 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 November 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 March 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"18"}}