{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T07:37:18Z","timestamp":1774251438858,"version":"3.50.1"},"reference-count":46,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T00:00:00Z","timestamp":1753833600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T00:00:00Z","timestamp":1753833600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100007195","name":"Universit\u00e0 degli Studi di Napoli Federico II","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100007195","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Classif"],"published-print":{"date-parts":[[2026,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Clustering is one of the most ubiquitous unsupervised learning tasks, with applications to a wide variety of domains. Genomics data analysis makes no exception, yet genomics datasets combine features of different natures (continuous and categorical) and are, therefore, of mixed type. The hierarchical clustering of a set of genomics data observations requires pairwise distances or dissimilarities, and it returns a sequence of nested clustering partitions represented as a tree graph. The choice of the dissimilarity or distance measure affects the obtained sequence of cluster partitions. Furthermore, selecting the reference partition out of the nested sequence is up to the user: this is done by setting a threshold, that is, by cutting the tree-based graph horizontally. A permutation test-based procedure has been proposed in the literature to select the final partition based on more than a single threshold (non-horizontal cut). This paper introduces a novel top-down implementation of such permutation test-based procedure to identify the final partition out of a hierarchy of solutions. Different approaches for distance computations are considered to extend the procedure\u2019s applicability to the mixed data case.<\/jats:p>","DOI":"10.1007\/s00357-025-09516-3","type":"journal-article","created":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T09:49:10Z","timestamp":1753868950000},"page":"41-65","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Automatic Dendrogram Slicing for Mixed-Type Data Clustering"],"prefix":"10.1007","volume":"43","author":[{"given":"Lucio","family":"Palazzo","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alfonso","family":"Iodice D\u2019Enza","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Domenico","family":"Vistocco","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9027-5053","authenticated-orcid":false,"given":"Francesco","family":"Palumbo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,7,30]]},"reference":[{"issue":"2","key":"9516_CR1","doi-asserted-by":"publisher","first-page":"503","DOI":"10.1016\/j.datak.2007.03.016","volume":"63","author":"A Ahmad","year":"2007","unstructured":"Ahmad, A., & Dey, L. (2007). A k-mean clustering algorithm for mixed numeric and categorical data. Data & Knowledge Engineering, 63(2), 503\u2013527.","journal-title":"Data & Knowledge Engineering"},{"key":"9516_CR2","doi-asserted-by":"publisher","first-page":"31883","DOI":"10.1109\/ACCESS.2019.2903568","volume":"7","author":"A Ahmad","year":"2019","unstructured":"Ahmad, A., & Khan, S. S. (2019). Survey of state-of-the-art mixed data clustering algorithms. Ieee Access, 7, 31883\u201331902.","journal-title":"Ieee Access"},{"key":"9516_CR3","doi-asserted-by":"crossref","unstructured":"Anagnostou, P., Tasoulis, S., Plagianakos, V., et\u00a0al. (2022). Hipart: Hierarchical divisive clustering toolbox. arXiv:2209.08680","DOI":"10.21105\/joss.05024"},{"key":"9516_CR4","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s11704-019-9059-3","volume":"15","author":"P Bhattacharjee","year":"2021","unstructured":"Bhattacharjee, P., & Mitra, P. (2021). A survey of density based clustering algorithms. Frontiers of Computer Science, 15, 1\u201327.","journal-title":"Frontiers of Computer Science"},{"key":"9516_CR5","doi-asserted-by":"publisher","first-page":"325","DOI":"10.1023\/A:1009740529316","volume":"2","author":"D Boley","year":"1998","unstructured":"Boley, D. (1998). Principal direction divisive partitioning. Data Mining And Knowledge Discovery, 2, 325\u2013344.","journal-title":"Data Mining And Knowledge Discovery"},{"key":"9516_CR6","doi-asserted-by":"publisher","first-page":"52","DOI":"10.1016\/j.csda.2012.12.008","volume":"71","author":"C Bouveyron","year":"2014","unstructured":"Bouveyron, C., & Brunet-Saumard, C. (2014). Model-based clustering of high-dimensional data: A review. Computational Statistics & Data Analysis, 71, 52\u201378.","journal-title":"Computational Statistics & Data Analysis"},{"key":"9516_CR7","doi-asserted-by":"publisher","first-page":"285","DOI":"10.1007\/s00357-015-9179-x","volume":"32","author":"D Bruzzese","year":"2015","unstructured":"Bruzzese, D., & Vistocco, D. (2015). Despota: Dendrogram slicing through a pemutation test approach. Journal of Classification, 32, 285\u2013304.","journal-title":"Journal of Classification"},{"key":"9516_CR8","doi-asserted-by":"crossref","unstructured":"Chavent, M., Kuentz-Simonet, V., Labenne, A., et\u00a0al. (2014). Multivariate analysis of mixed data: The r package pcamixdata. arXiv:1411.4911","DOI":"10.32614\/CRAN.package.PCAmixdata"},{"key":"9516_CR9","doi-asserted-by":"crossref","unstructured":"Chen, J. W., & Dhahbi, J. (2021). Lung adenocarcinoma and lung squamous cell carcinoma cancer classification, biomarker identification, and gene expression analysis using overlapping feature selection methods. Scientific Reports, 11(1), 13323.","DOI":"10.1038\/s41598-021-92725-8"},{"issue":"4","key":"9516_CR10","doi-asserted-by":"publisher","first-page":"1295","DOI":"10.1007\/s13042-023-01968-6","volume":"15","author":"K Chu","year":"2024","unstructured":"Chu, K., Zhang, M., Xun, Y., et al. (2024). A hybrid similarity measure-based clustering approach for mixed attribute data. International Journal of Machine Learning and Cybernetics, 15(4), 1295\u20131311.","journal-title":"International Journal of Machine Learning and Cybernetics"},{"key":"9516_CR11","first-page":"231","volume":"2","author":"J De Leeuw","year":"1980","unstructured":"De Leeuw, J., & Van Rijckevorsel, J. (1980). Homals and Princals\u2014Some generalizations of principal components analysis. Data Analysis And Informatics, 2, 231\u2013242.","journal-title":"Data Analysis And Informatics"},{"key":"9516_CR12","doi-asserted-by":"crossref","unstructured":"Ellen, J. G., Jacob, E., Nikolaou, N., et al. (2023). Autoencoder-based multimodal prediction of non-small cell lung cancer survival. Scientific Reports, 13(1), 15761.","DOI":"10.1038\/s41598-023-42365-x"},{"issue":"1","key":"9516_CR13","doi-asserted-by":"publisher","first-page":"80","DOI":"10.1111\/insr.12274","volume":"87","author":"AH Foss","year":"2019","unstructured":"Foss, A. H., Markatou, M., & Ray, B. (2019). Distance metrics and clustering methods for mixed-type data. International Statistical Review, 87(1), 80\u2013109.","journal-title":"International Statistical Review"},{"key":"9516_CR14","unstructured":"Gao, LL., Bien, J., Witten, D. (2022). Selective inference for hierarchical clustering. Journal of the American Statistical Association pp. 1\u201311"},{"key":"9516_CR15","unstructured":"Good, P. (2013). Permutation tests: a practical guide to resampling methods for testing hypotheses. Springer Science & Business Media"},{"issue":"4","key":"9516_CR16","doi-asserted-by":"publisher","first-page":"857","DOI":"10.2307\/2528823","volume":"27","author":"J Gower","year":"1971","unstructured":"Gower, J. (1971). A general coefficient of similarity and some of its properties. Biometrics, 27(4), 857\u2013871.","journal-title":"Biometrics"},{"issue":"1","key":"9516_CR17","doi-asserted-by":"publisher","first-page":"100","DOI":"10.1038\/s43586-022-00184-w","volume":"2","author":"M Greenacre","year":"2022","unstructured":"Greenacre, M., Groenen, P. J., Hastie, T., et al. (2022). Principal component analysis. Nature Reviews Methods Primers, 2(1), 100.","journal-title":"Nature Reviews Methods Primers"},{"key":"9516_CR18","doi-asserted-by":"crossref","unstructured":"Hill, M., Smith, A. (1976). Principal component analysis of taxonomic data with multi-state discrete characters. Taxon pp. 249\u2013255","DOI":"10.2307\/1219449"},{"key":"9516_CR19","doi-asserted-by":"publisher","unstructured":"Horst, AM., Hill, AP., Gorman, KB. (2020). palmerpenguins: Palmer Archipelago (Antarctica) penguin data. https:\/\/doi.org\/10.5281\/zenodo.3960218, https:\/\/allisonhorst.github.io\/palmerpenguins\/, r package version 0.1.0","DOI":"10.5281\/zenodo.3960218"},{"key":"9516_CR20","unstructured":"Huang, Z. (1997). Clustering large data sets with mixed numeric and categorical values,\" proceedings of 1st pacific-asia conference on knowledge discovery and data mining"},{"key":"9516_CR21","doi-asserted-by":"publisher","DOI":"10.1201\/b21874","volume-title":"Exploratory multivariate analysis by example using R","author":"F Husson","year":"2017","unstructured":"Husson, F., L\u00ea, S., & Pag\u00e8s, J. (2017). Exploratory multivariate analysis by example using R. CRC Press."},{"issue":"1","key":"9516_CR22","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1023\/A:1021394316112","volume":"25","author":"Y Jung","year":"2003","unstructured":"Jung, Y., Park, H., Du, D. Z., et al. (2003). A decision criterion for the optimal number of clusters in hierarchical clustering. Journal of Global Optimization, 25(1), 91\u2013111.","journal-title":"Journal of Global Optimization"},{"issue":"2","key":"9516_CR23","doi-asserted-by":"publisher","first-page":"197","DOI":"10.1007\/BF02294458","volume":"56","author":"HA Kiers","year":"1991","unstructured":"Kiers, H. A. (1991). Simple structure in component analysis techniques for mixtures of qualitative and quantitative variables. Psychometrika, 56(2), 197\u2013212.","journal-title":"Psychometrika"},{"issue":"2","key":"9516_CR24","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1109\/TIT.1982.1056489","volume":"28","author":"S Lloyd","year":"1982","unstructured":"Lloyd, S. (1982). Least squares quantization in PCM. IEEE Transactions On Information Theory, 28(2), 129\u2013137.","journal-title":"IEEE Transactions On Information Theory"},{"key":"9516_CR25","unstructured":"MacQueen, J., et\u00a0al. (1967). Some methods for classification and analysis of multivariate observations. In: Proceedings of the fifth Berkeley symposium on mathematical statistics and probability, Oakland, CA, USA, pp. 281\u2013297"},{"key":"9516_CR26","doi-asserted-by":"crossref","unstructured":"McMurdie, PJ., Holmes, S. (2012). Phyloseq: A bioconductor package for handling and analysis of high-throughput phylogenetic sequence data. In: Biocomputing 2012. World Scientific, p. 235\u2013246","DOI":"10.1142\/9789814366496_0023"},{"key":"9516_CR27","doi-asserted-by":"publisher","first-page":"331","DOI":"10.1007\/s00357-016-9211-9","volume":"33","author":"PD McNicholas","year":"2016","unstructured":"McNicholas, P. D. (2016). Model-based clustering. Journal of Classification, 33, 331\u2013373.","journal-title":"Journal of Classification"},{"key":"9516_CR28","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1007\/BF02294245","volume":"50","author":"GW Milligan","year":"1985","unstructured":"Milligan, G. W., & Cooper, M. C. (1985). An examination of procedures for determining the number of clusters in a data set. Psychometrika, 50, 159\u2013179.","journal-title":"Psychometrika"},{"key":"9516_CR29","doi-asserted-by":"crossref","unstructured":"Mousavi, E., & Sehhati, M. (2023). A generalized multi-aspect distance metric for mixed-type data clustering. Pattern Recognition, 138, Article 109353.","DOI":"10.1016\/j.patcog.2023.109353"},{"issue":"4","key":"9516_CR30","first-page":"93","volume":"52","author":"J Pag\u00e8s","year":"2004","unstructured":"Pag\u00e8s, J. (2004). Analyse factorielle de donnees mixtes: principe et exemple d\u2019application. Revue de statistique appliqu\u00e9e, 52(4), 93\u2013111.","journal-title":"Revue de statistique appliqu\u00e9e"},{"key":"9516_CR31","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa, F., Varoquaux, G., Gramfort, A., et al. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825\u20132830.","journal-title":"Journal of Machine Learning Research"},{"issue":"8","key":"9516_CR32","doi-asserted-by":"publisher","first-page":"8219","DOI":"10.1007\/s10462-022-10366-3","volume":"56","author":"X Ran","year":"2023","unstructured":"Ran, X., Xi, Y., Lu, Y., et al. (2023). Comprehensive survey on hierarchical clustering algorithms and the recent developments. Artificial Intelligence Review, 56(8), 8219\u20138264.","journal-title":"Artificial Intelligence Review"},{"key":"9516_CR33","doi-asserted-by":"publisher","first-page":"251","DOI":"10.1007\/978-3-319-23528-8_16","volume-title":"Machine Learning and Knowledge Discovery in Databases","author":"M Ring","year":"2015","unstructured":"Ring, M., Otto, F., Becker, M., et al. (2015). Condist: A context-driven categorical distance measure. In A. Appice, P. Rodrigues, V. Santos Costa, et al. (Eds.), Machine Learning and Knowledge Discovery in Databases (pp. 251\u2013266). Cham: Springer International Publishing."},{"issue":"2","key":"9516_CR34","doi-asserted-by":"publisher","first-page":"85","DOI":"10.1038\/nrg3868","volume":"16","author":"MD Ritchie","year":"2015","unstructured":"Ritchie, M. D., Holzinger, E. R., Li, R., et al. (2015). Methods of integrating data to uncover genotype-phenotype interactions. Nature Reviews Genetics, 16(2), 85\u201397.","journal-title":"Nature Reviews Genetics"},{"key":"9516_CR35","doi-asserted-by":"crossref","unstructured":"Ross, B. (2014). Mutual information between discrete and continuous data sets. PloS one, 9(2), Article e87357.","DOI":"10.1371\/journal.pone.0087357"},{"key":"9516_CR36","doi-asserted-by":"publisher","first-page":"53","DOI":"10.1016\/0377-0427(87)90125-7","volume":"20","author":"PJ Rousseeuw","year":"1987","unstructured":"Rousseeuw, P. J. (1987). Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. Journal Of Computational And Applied Mathematics, 20, 53\u201365.","journal-title":"Journal Of Computational And Applied Mathematics"},{"key":"9516_CR37","doi-asserted-by":"publisher","first-page":"345","DOI":"10.1007\/s00357-018-9259-9","volume":"35","author":"M Roux","year":"2018","unstructured":"Roux, M. (2018). A comparative study of divisive and agglomerative hierarchical clustering algorithms. Journal of Classification, 35, 345\u2013366.","journal-title":"Journal of Classification"},{"issue":"2","key":"9516_CR38","doi-asserted-by":"publisher","first-page":"129","DOI":"10.2307\/3316064","volume":"31","author":"SK Sahu","year":"2003","unstructured":"Sahu, S. K., Dey, D. K., & Branco, M. D. (2003). A new class of multivariate skew distributions with applications to bayesian regression models. Canadian Journal of Statistics, 31(2), 129\u2013150.","journal-title":"Canadian Journal of Statistics"},{"key":"9516_CR39","doi-asserted-by":"crossref","unstructured":"Savaresi, SM., Boley, DL. (2001). On the performance of bisecting k-means and pddp. In: Proceedings of the 2001 SIAM International Conference on Data Mining, SIAM, pp. 1\u201314","DOI":"10.1137\/1.9781611972719.5"},{"key":"9516_CR40","volume-title":"Density estimation for statistics and data analysis","author":"BW Silverman","year":"1998","unstructured":"Silverman, B. W. (1998). Density estimation for statistics and data analysis. Routledge."},{"issue":"10","key":"9516_CR41","doi-asserted-by":"publisher","first-page":"3391","DOI":"10.1016\/j.patcog.2010.05.025","volume":"43","author":"SK Tasoulis","year":"2010","unstructured":"Tasoulis, S. K., Tasoulis, D. K., & Plagianakos, V. P. (2010). Enhancing principal direction divisive clustering. Pattern Recognition, 43(10), 3391\u20133411.","journal-title":"Pattern Recognition"},{"key":"9516_CR42","doi-asserted-by":"crossref","unstructured":"Van de Velden, M., Iodice D\u2019Enza, A., & Markos, A. (2019). Distance-based clustering of mixed data. Wiley Interdisciplinary Reviews: Computational Statistics, 11(3), Article e1456.","DOI":"10.1002\/wics.1456"},{"key":"9516_CR43","doi-asserted-by":"crossref","unstructured":"Van de Velden, M., Iodice D\u2019Enza, A., Markos, A., et al. (2024). A general framework for implementing distances for categorical variables. Pattern Recognition, 153, Article 110547.","DOI":"10.1016\/j.patcog.2024.110547"},{"key":"9516_CR44","doi-asserted-by":"crossref","unstructured":"Van\u00a0de Velden, M., Iodice\u00a0D\u2019Enza, A., Markos, A., et\u00a0al. (2024b). Unbiased mixed variables distance. arXiv:2411.00429","DOI":"10.2139\/ssrn.5010828"},{"issue":"1","key":"9516_CR45","doi-asserted-by":"publisher","first-page":"2477","DOI":"10.1038\/s41598-023-29719-1","volume":"13","author":"Y Wang","year":"2023","unstructured":"Wang, Y., Xiao, X., & Li, Y. (2023). Construction and validation of a cuproptosis-related lncrna signature for the prediction of the prognosis of luad and lusc. Scientific Reports, 13(1), 2477.","journal-title":"Scientific Reports"},{"key":"9516_CR46","doi-asserted-by":"publisher","unstructured":"Zeimpekis, D., Gallopoulos, E. (2008). Principal Direction Divisive Partitioning with Kernels and k-Means Steering, Springer London, London, pp. 45\u201364.https:\/\/doi.org\/10.1007\/978-1-84800-046-9_3,","DOI":"10.1007\/978-1-84800-046-9_3"}],"container-title":["Journal of Classification"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00357-025-09516-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00357-025-09516-3","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00357-025-09516-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T06:43:21Z","timestamp":1774248201000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00357-025-09516-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,30]]},"references-count":46,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,4]]}},"alternative-id":["9516"],"URL":"https:\/\/doi.org\/10.1007\/s00357-025-09516-3","relation":{},"ISSN":["0176-4268","1432-1343"],"issn-type":[{"value":"0176-4268","type":"print"},{"value":"1432-1343","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,30]]},"assertion":[{"value":"16 July 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 July 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}