{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,20]],"date-time":"2026-03-20T17:18:02Z","timestamp":1774027082795,"version":"3.50.1"},"reference-count":43,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,7,4]],"date-time":"2024-07-04T00:00:00Z","timestamp":1720051200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,7,4]],"date-time":"2024-07-04T00:00:00Z","timestamp":1720051200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Karlsruher Institut f\u00fcr Technologie (KIT)"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Classif"],"published-print":{"date-parts":[[2025,3]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Cluster analysis aims to find meaningful groups, called clusters, in data. The objects within a cluster should be similar to each other and dissimilar to objects from other clusters. The fundamental question arising is whether found clusters are \u201cvalid clusters\u201d or not. Existing cluster validity indices are computation-intensive, make assumptions about the underlying cluster structure, or cannot detect the absence of clusters. Thus, we present a new cluster validation framework to assess the validity of a clustering and determine the underlying number of clusters <jats:inline-formula>\n              <jats:alternatives>\n                <jats:tex-math>$$k^*$$<\/jats:tex-math>\n                <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:msup>\n                    <mml:mi>k<\/mml:mi>\n                    <mml:mo>\u2217<\/mml:mo>\n                  <\/mml:msup>\n                <\/mml:math>\n              <\/jats:alternatives>\n            <\/jats:inline-formula>. Within the framework, we introduce a new merge criterion analyzing the data in a one-dimensional projection, which maximizes the ratio of between-cluster- variance to within-cluster-variance in the clusters. Nonetheless, other local methods can be applied as a merge criterion within the framework. Experiments on synthetic and real-world data sets show promising results for both the overall framework and the introduced merge criterion.<\/jats:p>","DOI":"10.1007\/s00357-024-09481-3","type":"journal-article","created":{"date-parts":[[2024,7,4]],"date-time":"2024-07-04T13:01:39Z","timestamp":1720098099000},"page":"54-71","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Cluster Validation Based on Fisher\u2019s Linear Discriminant Analysis"],"prefix":"10.1007","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2934-7406","authenticated-orcid":false,"given":"Fabian","family":"K\u00e4chele","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nora","family":"Schneider","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,7,4]]},"reference":[{"issue":"1","key":"9481_CR1","doi-asserted-by":"publisher","first-page":"243","DOI":"10.1016\/j.patcog.2012.07.021","volume":"46","author":"O Arbelaitz","year":"2013","unstructured":"Arbelaitz, O., Gurrutxaga, I., Muguerza, J., P\u00e9rez, J. M., & Perona, I. (2013). An extensive comparative study of cluster validity indices. Pattern Recognition, 46(1), 243\u2013256.","journal-title":"Pattern Recognition"},{"issue":"2","key":"9481_CR2","doi-asserted-by":"publisher","first-page":"61","DOI":"10.1016\/0031-3203(82)90002-4","volume":"15","author":"TA Bailey","year":"1982","unstructured":"Bailey, T. A., & Dubes, R. (1982). Cluster validity profiles. Pattern Recognition, 15(2), 61\u201383.","journal-title":"Pattern Recognition"},{"issue":"349","key":"9481_CR3","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1080\/01621459.1975.10480256","volume":"70","author":"FB Baker","year":"1975","unstructured":"Baker, F. B., & Hubert, L. J. (1975). Measuring the power of hierarchical cluster analysis. Journal of the American Statistical Association, 70(349), 31\u201338.","journal-title":"Journal of the American Statistical Association"},{"issue":"1","key":"9481_CR4","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1080\/03610927408827101","volume":"3","author":"T Cali\u0144ski","year":"1974","unstructured":"Cali\u0144ski, T., & Harabasz, J. (1974). A dendrite method for cluster analysis. Communications in Statistics-theory and Methods, 3(1), 1\u201327.","journal-title":"Communications in Statistics-theory and Methods"},{"key":"9481_CR5","doi-asserted-by":"publisher","first-page":"7","DOI":"10.1007\/s00357-012-9098-z","volume":"29","author":"J Cerdeira","year":"2012","unstructured":"Cerdeira, J., Martins, M., & Silva, P. (2012). A combinatorial approach to assess the separability of clusters. Journal of Classification, 29, 7\u201322.","journal-title":"Journal of Classification"},{"key":"9481_CR6","doi-asserted-by":"crossref","unstructured":"Dangl, R., &\u00a0Leisch, F. (2019). Effects of resampling in determining the number of clusters in a data set. Journal of Classification 37.","DOI":"10.1007\/s00357-019-09328-2"},{"key":"9481_CR7","doi-asserted-by":"crossref","unstructured":"Davies, D.\u00a0L., & Bouldin, D.\u00a0W. (1979). A cluster separation measure. IEEE Transactions on Pattern Analysis and Machine Intelligence PAMI-1(2), 224\u2013227.","DOI":"10.1109\/TPAMI.1979.4766909"},{"issue":"2","key":"9481_CR8","doi-asserted-by":"publisher","first-page":"271","DOI":"10.1111\/rssb.12310","volume":"81","author":"A Delaigle","year":"2019","unstructured":"Delaigle, A., Hall, P., & Pham, T. (2019). Clustering functional data into groups by using projections. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 81(2), 271\u2013304.","journal-title":"Journal of the Royal Statistical Society: Series B (Statistical Methodology)"},{"key":"9481_CR9","unstructured":"Dua, D., &\u00a0Graff, C. (2017). UCI machine learning repository."},{"issue":"6","key":"9481_CR10","doi-asserted-by":"publisher","first-page":"645","DOI":"10.1016\/0031-3203(87)90034-3","volume":"20","author":"RC Dubes","year":"1987","unstructured":"Dubes, R. C. (1987). How many clusters are best? - An experiment. Pattern Recognition, 20(6), 645\u2013663.","journal-title":"Pattern Recognition"},{"issue":"2","key":"9481_CR11","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1111\/j.1469-1809.1936.tb02137.x","volume":"7","author":"RA Fisher","year":"1936","unstructured":"Fisher, R. A. (1936). The use of multiple measurements in taxonomic problems. Annals of eugenics, 7(2), 179\u2013188.","journal-title":"Annals of eugenics"},{"issue":"1","key":"9481_CR12","doi-asserted-by":"publisher","first-page":"162","DOI":"10.1080\/10618600.2019.1647846","volume":"29","author":"W Fu","year":"2020","unstructured":"Fu, W., & Perry, P. O. (2020). Estimating the number of clusters using cross-validation. Journal of Computational and Graphical Statistics, 29(1), 162\u2013173.","journal-title":"Journal of Computational and Graphical Statistics"},{"issue":"2","key":"9481_CR13","doi-asserted-by":"publisher","first-page":"263","DOI":"10.1016\/0022-5193(83)90340-5","volume":"101","author":"MA Gates","year":"1983","unstructured":"Gates, M. A., & Hansell, R. I. C. (1983). On the distinctness of clusters. Journal of Theoretical Biology, 101(2), 263\u2013273.","journal-title":"Journal of Theoretical Biology"},{"issue":"526","key":"9481_CR14","doi-asserted-by":"publisher","first-page":"893","DOI":"10.1080\/01621459.2018.1458618","volume":"114","author":"J Geng","year":"2019","unstructured":"Geng, J., Bhattacharya, A., & Pati, D. (2019). Probabilistic community detection with unknown number of communities. Journal of the American Statistical Association, 114(526), 893\u2013905.","journal-title":"Journal of the American Statistical Association"},{"key":"9481_CR15","doi-asserted-by":"publisher","first-page":"22","DOI":"10.1007\/978-4-431-65950-1_2","volume-title":"Data science, classification, and related methods, Tokyo","author":"AD Gordon","year":"1998","unstructured":"Gordon, A. D. (1998). Cluster validation. In C. Hayashi, K. Yajima, H.-H. Bock, N. Ohsumi, Y. Tanaka, & Y. Baba (Eds.), Data science, classification, and related methods, Tokyo (pp. 22\u201339). Springer Japan."},{"key":"9481_CR16","doi-asserted-by":"crossref","unstructured":"Halkidi, M.,\u00a0Batistakis, Y., &\u00a0Vazirgiannis, M. (2001). On clustering validation techniques. Journal of Intelligent Information Systems 17.","DOI":"10.1023\/A:1012801612483"},{"key":"9481_CR17","doi-asserted-by":"crossref","unstructured":"Handl, J.,\u00a0Knowles, J., & Kell, D.\u00a0B. (2005). Computational cluster validation in post-genomic data analysis. Bioinformatics 21(15), 3201\u20133212.","DOI":"10.1093\/bioinformatics\/bti517"},{"key":"9481_CR18","doi-asserted-by":"crossref","unstructured":"Hastie, T.,\u00a0Tibshirani, R., &\u00a0Friedman, J. (2009). The elements of statistical learning: data mining, inference and prediction (2 ed.). Springer.","DOI":"10.1007\/978-0-387-84858-7"},{"key":"9481_CR19","doi-asserted-by":"crossref","unstructured":"Hennig, C. (2015). What are the true clusters? Pattern Recognition Letters 64, 53\u201362. Philosophical Aspects of Pattern Recognition.","DOI":"10.1016\/j.patrec.2015.04.009"},{"key":"9481_CR20","doi-asserted-by":"crossref","unstructured":"Hennig, C. (2022). An empirical comparison and characterisation of nine popular clustering methods. Advances in Data Analysis and Classification.","DOI":"10.1007\/s11634-021-00478-z"},{"key":"9481_CR21","doi-asserted-by":"publisher","DOI":"10.1201\/b19706","volume-title":"Handbook of cluster analysis (1th","author":"C Hennig","year":"2015","unstructured":"Hennig, C., Meila, M., Murtagh, F., & Rocci, R. (2015). Handbook of cluster analysis (1th (edition). New York: Chapman and Hall\/CRC.","edition":"edition"},{"key":"9481_CR22","doi-asserted-by":"crossref","unstructured":"Ingrassia, S., &\u00a0Punzo, A. (2020). Cluster validation for mixtures of regressions via the total sum of squares decomposition. Journal of Classification 37(2), 526\u2013547.","DOI":"10.1007\/s00357-019-09326-4"},{"issue":"3","key":"9481_CR23","doi-asserted-by":"publisher","first-page":"547","DOI":"10.1198\/106186005X59586","volume":"14","author":"J Li","year":"2005","unstructured":"Li, J. (2005). Clustering based on a multilayer mixture model. Journal of Computational and Graphical Statistics, 14(3), 547\u2013568.","journal-title":"Journal of Computational and Graphical Statistics"},{"issue":"483","key":"9481_CR24","doi-asserted-by":"publisher","first-page":"1281","DOI":"10.1198\/016214508000000454","volume":"103","author":"Y Liu","year":"2008","unstructured":"Liu, Y., Hayes, D. N., Nobel, A., & Marron, J. S. (2008). Statistical significance of clustering for high-dimension, low-sample size data. Journal of the American Statistical Association, 103(483), 1281\u20131293.","journal-title":"Journal of the American Statistical Association"},{"issue":"1","key":"9481_CR25","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1080\/10618600.2014.978007","volume":"25","author":"V Melnykov","year":"2016","unstructured":"Melnykov, V. (2016). Merging mixture components for clustering through pairwise overlap. Journal of Computational and Graphical Statistics, 25(1), 66\u201390.","journal-title":"Journal of Computational and Graphical Statistics"},{"key":"9481_CR26","doi-asserted-by":"publisher","first-page":"97","DOI":"10.1007\/s00357-019-09314-8","volume":"37","author":"V Melnykov","year":"2020","unstructured":"Melnykov, V., & Michael, S. (2020). Clustering large datasets by merging k-means solutions. Journal of Classification, 37, 97\u2013123.","journal-title":"Journal of Classification"},{"issue":"2","key":"9481_CR27","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1007\/BF02294245","volume":"50","author":"G Milligan","year":"1985","unstructured":"Milligan, G., & Cooper, M. (1985). An examination of procedures for determining the number of clusters in a data set. Psychometrika, 50(2), 159\u2013179.","journal-title":"Psychometrika"},{"key":"9481_CR28","doi-asserted-by":"publisher","first-page":"583","DOI":"10.3233\/IDA-2007-11602","volume":"11","author":"MGH Omran","year":"2011","unstructured":"Omran, M. G. H., Engelbrecht, A. P., & Salman, A. (2011). An overview of clustering methods. Intelligent Data Analysis, 11, 583\u2013605.","journal-title":"Intelligent Data Analysis"},{"issue":"405","key":"9481_CR29","doi-asserted-by":"publisher","first-page":"184","DOI":"10.1080\/01621459.1989.10478754","volume":"84","author":"R Peck","year":"1989","unstructured":"Peck, R., Fisher, L., & Ness, J. V. (1989). Approximate confidence intervals for the number of clusters. Journal of the American Statistical Association, 84(405), 184\u2013191.","journal-title":"Journal of the American Statistical Association"},{"issue":"456","key":"9481_CR30","doi-asserted-by":"publisher","first-page":"1433","DOI":"10.1198\/016214501753382345","volume":"96","author":"D Pe\u00f1a","year":"2001","unstructured":"Pe\u00f1a, D., & Prieto, F. J. (2001). Cluster identification using projections. Journal of the American Statistical Association, 96(456), 1433\u20131445.","journal-title":"Journal of the American Statistical Association"},{"issue":"336","key":"9481_CR31","doi-asserted-by":"publisher","first-page":"846","DOI":"10.1080\/01621459.1971.10482356","volume":"66","author":"WM Rand","year":"1971","unstructured":"Rand, W. M. (1971). Objective criteria for the evaluation of clustering methods. Journal of the American Statistical Association, 66(336), 846\u2013850.","journal-title":"Journal of the American Statistical Association"},{"issue":"1","key":"9481_CR32","first-page":"27","volume":"5","author":"E Rend\u00f3n","year":"2011","unstructured":"Rend\u00f3n, E., Abundez, I., Arizmendi, A., & Quiroz, E. M. (2011). Internal versus external cluster validation indexes. International Journal of computers and communications, 5(1), 27\u201334.","journal-title":"International Journal of computers and communications"},{"key":"9481_CR33","doi-asserted-by":"crossref","unstructured":"Rossbroich, J.,\u00a0Durieux, J., & Wilderjans, T.\u00a0F. (2022). Model selection strategies for determining the optimal number of overlapping clusters in additive overlapping partitional clustering. Journal of Classification.","DOI":"10.1007\/s00357-021-09409-1"},{"key":"9481_CR34","unstructured":"Rousseeuw, P.\u00a0J., &\u00a0Kaufman, L. (1990). Finding groups in data: An introduction to cluster analysis. John Wiley & Sons."},{"key":"9481_CR35","doi-asserted-by":"crossref","unstructured":"Salvador, S., &\u00a0Chan, P. (2004). Determining the number of clusters\/segments in hierarchical clustering\/segmentation algorithms. In 16th IEEE international conference on tools with artificial intelligence, pp. 576\u2013584. IEEE.","DOI":"10.1109\/ICTAI.2004.50"},{"issue":"2","key":"9481_CR36","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1007\/BF02312508","volume":"9","author":"P Sneath","year":"1977","unstructured":"Sneath, P. (1977). A method for testing the distinctness of clusters: A test of the disjunction of two clusters in Euclidean space as measured by their overlap. Journal of the International Association for Mathematical Geology, 9(2), 123\u2013143.","journal-title":"Journal of the International Association for Mathematical Geology"},{"issue":"463","key":"9481_CR37","doi-asserted-by":"publisher","first-page":"750","DOI":"10.1198\/016214503000000666","volume":"98","author":"CA Sugar","year":"2003","unstructured":"Sugar, C. A., & James, G. M. (2003). Finding the number of clusters in a dataset: An information-theoretic approach. Journal of the American Statistical Association, 98(463), 750\u2013763.","journal-title":"Journal of the American Statistical Association"},{"issue":"2","key":"9481_CR38","doi-asserted-by":"publisher","first-page":"411","DOI":"10.1111\/1467-9868.00293","volume":"63","author":"R Tibshirani","year":"2001","unstructured":"Tibshirani, R., Walther, G., & Hastie, T. (2001). Estimating the number of clusters in a data set via the gap statistic. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 63(2), 411\u2013423.","journal-title":"Journal of the Royal Statistical Society: Series B (Statistical Methodology)"},{"issue":"3","key":"9481_CR39","doi-asserted-by":"publisher","DOI":"10.1002\/widm.1444","volume":"12","author":"T Ullmann","year":"2022","unstructured":"Ullmann, T., Hennig, C., & Boulesteix, A.-L. (2022). Validation of cluster analysis results on validation data: A systematic framework. WIREs Data Mining and Knowledge Discovery, 12(3), e1444.","journal-title":"WIREs Data Mining and Knowledge Discovery"},{"issue":"3","key":"9481_CR40","first-page":"235","volume":"2","author":"U von Luxburg","year":"2010","unstructured":"von Luxburg, U. (2010). Clustering stability: An overview. Foundations and Trends in Machine Learning, 2(3), 235\u2013274.","journal-title":"Foundations and Trends in Machine Learning"},{"key":"9481_CR41","doi-asserted-by":"crossref","unstructured":"Wierzcho\u0144, S.\u00a0T. (2018). Modern algorithms of cluster analysis. Springer International Publishing.","DOI":"10.1007\/978-3-319-69308-8"},{"key":"9481_CR42","doi-asserted-by":"publisher","first-page":"1033","DOI":"10.1038\/nmeth.3583","volume":"12","author":"C Wiwie","year":"2015","unstructured":"Wiwie, C., Baumbach, J., & R\u00f6ttger, R. (2015). Comparing the performance of biomedical clustering methods. Nature Methods, 12, 1033\u20131038.","journal-title":"Nature Methods"},{"issue":"3","key":"9481_CR43","doi-asserted-by":"publisher","first-page":"645","DOI":"10.1109\/TNN.2005.845141","volume":"16","author":"R Xu","year":"2005","unstructured":"Xu, R., & Wunsch, D. (2005). Survey of clustering algorithms. IEEE Transactions on Neural Networks, 16(3), 645\u2013678.","journal-title":"IEEE Transactions on Neural Networks"}],"container-title":["Journal of Classification"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00357-024-09481-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00357-024-09481-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00357-024-09481-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,3,19]],"date-time":"2025-03-19T08:10:55Z","timestamp":1742371855000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00357-024-09481-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,4]]},"references-count":43,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,3]]}},"alternative-id":["9481"],"URL":"https:\/\/doi.org\/10.1007\/s00357-024-09481-3","relation":{},"ISSN":["0176-4268","1432-1343"],"issn-type":[{"value":"0176-4268","type":"print"},{"value":"1432-1343","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,4]]},"assertion":[{"value":"10 June 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 July 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors comply with all ethical standards. No research involving Human Participants and\/or Animals was conducted.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical Approval"}},{"value":"The authors declare no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of Interest"}}]}}