{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T12:34:02Z","timestamp":1779194042027,"version":"3.51.4"},"reference-count":25,"publisher":"Oxford University Press (OUP)","issue":"13","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2007,7,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Biologists often employ clustering techniques in the explorative phase of microarray data analysis to discover relevant biological groupings. Given the availability of numerous clustering algorithms in the machine-learning literature, an user might want to select one that performs the best for his\/her data set or application. While various validation measures have been proposed over the years to judge the quality of clusters produced by a given clustering algorithm including their biological relevance, unfortunately, a given clustering algorithm can perform poorly under one validation measure while outperforming many other algorithms under another validation measure. A manual synthesis of results from multiple validation measures is nearly impossible in practice, especially, when a large number of clustering algorithms are to be compared using several measures. An automated and objective way of reconciling the rankings is needed.<\/jats:p>\n               <jats:p>Results: Using a Monte Carlo cross-entropy algorithm, we successfully combine the ranks of a set of clustering algorithms under consideration via a weighted aggregation that optimizes a distance criterion. The proposed weighted rank aggregation allows for a far more objective and automated assessment of clustering results than a simple visual inspection. We illustrate our procedure using one simulated as well as three real gene expression data sets from various platforms where we rank a total of eleven clustering algorithms using a combined examination of 10 different validation measures. The aggregate rankings were found for a given number of clusters k and also for an entire range of k.<\/jats:p>\n               <jats:p>Availability: R code for all validation measures and rank aggregation is available from the authors upon request.<\/jats:p>\n               <jats:p>Contact: somnath.datta@louisville.edu<\/jats:p>\n               <jats:p>Supplementary information: Supplementary information are available at http:\/\/www.somnathdatta.org\/Supp\/RankCluster\/supp.htm.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btm158","type":"journal-article","created":{"date-parts":[[2007,5,6]],"date-time":"2007-05-06T00:28:57Z","timestamp":1178411337000},"page":"1607-1615","source":"Crossref","is-referenced-by-count":244,"title":["Weighted rank aggregation of cluster validation measures: a Monte Carlo cross-entropy approach"],"prefix":"10.1093","volume":"23","author":[{"given":"Vasyl","family":"Pihur","sequence":"first","affiliation":[{"name":"Department of Bioinformatics and Biostatistics, University of Louisville, Louisville, KY 40202, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Susmita","family":"Datta","sequence":"additional","affiliation":[{"name":"Department of Bioinformatics and Biostatistics, University of Louisville, Louisville, KY 40202, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Somnath","family":"Datta","sequence":"additional","affiliation":[{"name":"Department of Bioinformatics and Biostatistics, University of Louisville, Louisville, KY 40202, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2007,5,5]]},"reference":[{"key":"2023062708444254400_B1","doi-asserted-by":"crossref","first-page":"R499","DOI":"10.1186\/bcr899","article-title":"Transcriptomic changes in human breast cancer progression as determined by serial analysis of gene expression","volume":"6","author":"Abba","year":"2004","journal-title":"Breast Cancer Res"},{"key":"2023062708444254400_B2","doi-asserted-by":"crossref","first-page":"803","DOI":"10.2307\/2532201","article-title":"Model-based Gaussian and non-Gaussian clustering","volume":"49","author":"Banfield","year":"1993","journal-title":"Biometrics"},{"key":"2023062708444254400_B3","doi-asserted-by":"crossref","first-page":"699","DOI":"10.1126\/science.282.5389.699","article-title":"The transcriptional program of sporulation in budding yeast","volume":"282","author":"Chu","year":"1998","journal-title":"Science"},{"key":"2023062708444254400_B4","doi-asserted-by":"crossref","first-page":"459","DOI":"10.1093\/bioinformatics\/btg025","article-title":"Comparisons and validation of statistical clustering techniques for microarray gene expression data","volume":"19","author":"Datta","year":"2003","journal-title":"Bioinformatics"},{"key":"2023062708444254400_B5","doi-asserted-by":"crossref","first-page":"397","DOI":"10.1186\/1471-2105-7-397","article-title":"Methods for evaluating clustering algorithms for gene expression data using a reference set of functional classes","volume":"7","author":"Datta","year":"2006","journal-title":"BMC Bioinformatics"},{"key":"2023062708444254400_B6","doi-asserted-by":"crossref","first-page":"19","DOI":"10.1007\/s10479-005-5724-z","article-title":"A tutorial on the Cross-Entropy method","volume":"134","author":"De Boer","year":"2005","journal-title":"Ann. Oper. Res"},{"key":"2023062708444254400_B7","doi-asserted-by":"crossref","first-page":"95","DOI":"10.1080\/01969727408546059","article-title":"Well separated clusters and fuzzy partitions","volume":"4","author":"Dunn","year":"1974","journal-title":"J. Cybern"},{"key":"2023062708444254400_B8","doi-asserted-by":"crossref","first-page":"134","DOI":"10.1137\/S0895480102412856","article-title":"Comparing top k lists","volume":"17","author":"Fagin","year":"2003","journal-title":"SIAM J. Discrete Math"},{"key":"2023062708444254400_B9","first-page":"1081","article-title":"Evolutionary multiobjective clustering","author":"Handl","year":"2004"},{"key":"2023062708444254400_B10","first-page":"547","article-title":"Exploiting the trade-off \u2013 the benefits of multiple objectives in data clustering","author":"Handl","year":"2005"},{"key":"2023062708444254400_B11","doi-asserted-by":"crossref","first-page":"3201","DOI":"10.1093\/bioinformatics\/bti517","article-title":"Computational cluster validation in post-genomic data analysis","volume":"21","author":"Handl","year":"2005","journal-title":"Bioinformatics"},{"key":"2023062708444254400_B12","doi-asserted-by":"crossref","first-page":"100","DOI":"10.2307\/2346830","article-title":"A k-means clustering algorithm","volume":"28","author":"Hartigan","year":"1979","journal-title":"Appl. Stat"},{"key":"2023062708444254400_B13","doi-asserted-by":"crossref","first-page":"126","DOI":"10.1093\/bioinformatics\/17.2.126","article-title":"A hierarchical unsupervised growing neural network for clustering gene expression patterns","volume":"17","author":"Herrero","year":"2001","journal-title":"Bioinformatics"},{"key":"2023062708444254400_B14","volume-title":"Fitting Groups in Data. An Introduction to Cluster Analysis","author":"Kaufman","year":"1990"},{"key":"2023062708444254400_B15","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-97966-8","volume-title":"Self-Organizing Maps","author":"Kohonen","year":"1997","edition":"2nd edn"},{"key":"2023062708444254400_B16","first-page":"424","article-title":"Multiobjective data clustering","author":"Law","year":"2004"},{"key":"2023062708444254400_B17","article-title":"Rank aggregation of putative microRNA targets with Cross-Entropy Monte Carlo methods","author":"Lin","year":"2006"},{"key":"2023062708444254400_B18","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1016\/0377-0427(87)90125-7","article-title":"Silhouettes: a graphical aid to the interpretation and validation of cluster analysis","volume":"20","author":"Rousseeuw","year":"1987","journal-title":"J. Comput. Appl. Math"},{"key":"2023062708444254400_B19","doi-asserted-by":"crossref","first-page":"89","DOI":"10.1016\/S0377-2217(96)00385-2","article-title":"Optimization of computer simulation models with rare events","volume":"99","author":"Rubinstein","year":"1997","journal-title":"Eur. J. Oper. Res"},{"key":"2023062708444254400_B20","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1023\/A:1010091220143","article-title":"The simulated Entropy method for combinatorial and continuous optimization","volume":"2","author":"Rubinstein","year":"1999","journal-title":"Methodol. Comput. Appl. Probab"},{"key":"2023062708444254400_B21","doi-asserted-by":"crossref","first-page":"304","DOI":"10.1007\/978-1-4757-6594-6_14","article-title":"Combinatorial optimization Cross-Entropy, Ants, and rare events","volume-title":"Stochastic Optimization: Algorithms and Applications","author":"Rubinstein","year":"2001"},{"key":"2023062708444254400_B22","volume-title":"The Cross-Entropy Method. A Unified Approach to Combinatorial Optimization, Monte-Carlo Simulation and Machine Learning","author":"Rubinstein","year":"2004"},{"key":"2023062708444254400_B23","volume-title":"Numerical Taxonomy","author":"Sneath","year":"1973"},{"key":"2023062708444254400_B24","first-page":"309","article-title":"Validating clustering for gene expression data","volume-title":"Bioinformatics","author":"Yeung","year":"2001"},{"key":"2023062708444254400_B25","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1186\/jbiol16","article-title":"The functional landscape of mouse gene expression","volume":"3","author":"Zhang","year":"2004","journal-title":"J. Biol"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/13\/1607\/50716060\/bioinformatics_23_13_1607.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/13\/1607\/50716060\/bioinformatics_23_13_1607.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,27]],"date-time":"2023-06-27T08:46:22Z","timestamp":1687855582000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/23\/13\/1607\/223480"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,5,5]]},"references-count":25,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2007,7,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btm158","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2007,7]]},"published":{"date-parts":[[2007,5,5]]}}}