{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T08:16:12Z","timestamp":1768292172595,"version":"3.49.0"},"reference-count":42,"publisher":"Oxford University Press (OUP)","issue":"14","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2006,7,15]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: The clustering of gene profiles across some experimental conditions of interest contributes significantly to the elucidation of unknown gene function, the validation of gene discoveries and the interpretation of biological processes. However, this clustering problem is not straightforward as the profiles of the genes are not all independently distributed and the expression levels may have been obtained from an experimental design involving replicated arrays. Ignoring the dependence between the gene profiles and the structure of the replicated data can result in important sources of variability in the experiments being overlooked in the analysis, with the consequent possibility of misleading inferences being made. We propose a random-effects model that provides a unified approach to the clustering of genes with correlated expression levels measured in a wide variety of experimental situations. Our model is an extension of the normal mixture model to account for the correlations between the gene profiles and to enable covariate information to be incorporated into the clustering process. Hence the model is applicable to longitudinal studies with or without replication, for example, time-course experiments by using time as a covariate, and to cross-sectional experiments by using categorical covariates to represent the different experimental classes.<\/jats:p><jats:p>Results: We show that our random-effects model can be fitted by maximum likelihood via the EM algorithm for which the E(expectation)and M(maximization) steps can be implemented in closed form. Hence our model can be fitted deterministically without the need for time-consuming Monte Carlo approximations. The effectiveness of our model-based procedure for the clustering of correlated gene profiles is demonstrated on three real datasets, representing typical microarray experimental designs, covering time-course, repeated-measurement and cross-sectional data. In these examples, relevant clusters of the genes are obtained, which are supported by existing gene-function annotation. A synthetic dataset is considered too.<\/jats:p><jats:p>Availability: A Fortran program blue called EMMIX-WIRE (EM-based MIXture analysis WIth Random Effects) is available on request from the corresponding author.<\/jats:p><jats:p>Contact: \u00a0gjm@maths.uq.edu.au<\/jats:p><jats:p>Supplementary information: \u00a0. Colour versions of Figures 1 and 2 are available as Supplementary material on Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btl165","type":"journal-article","created":{"date-parts":[[2006,5,5]],"date-time":"2006-05-05T00:19:22Z","timestamp":1146788362000},"page":"1745-1752","source":"Crossref","is-referenced-by-count":125,"title":["A Mixture model with random-effects components for clustering correlated gene-expression profiles"],"prefix":"10.1093","volume":"22","author":[{"given":"S. K.","family":"Ng","sequence":"first","affiliation":[{"name":"Department of Mathematics, University of Queensland 1 \u00a0 1 \u00a0 \u00a0 Brisbane, QLD 4072, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"G. J.","family":"McLachlan","sequence":"additional","affiliation":[{"name":"Department of Mathematics, University of Queensland 1 \u00a0 1 \u00a0 \u00a0 Brisbane, QLD 4072, Australia"},{"name":"Institute for Molecular Bioscience, University of Queensland 2 \u00a0 2 \u00a0 \u00a0 Brisbane, QLD 4072, Australia"},{"name":"ARC Centre for Complex Systems, University of Queensland 3 \u00a0 3 \u00a0 \u00a0 Brisbane, QLD 4072, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"K.","family":"Wang","sequence":"additional","affiliation":[{"name":"ARC Centre for Complex Systems, University of Queensland 3 \u00a0 3 \u00a0 \u00a0 Brisbane, QLD 4072, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"L.","family":"Ben-Tovim Jones","sequence":"additional","affiliation":[{"name":"Institute for Molecular Bioscience, University of Queensland 2 \u00a0 2 \u00a0 \u00a0 Brisbane, QLD 4072, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"S.-W.","family":"Ng","sequence":"additional","affiliation":[{"name":"Laboratory of Gynecologic Oncology, Department of Obstetrics, Gynecology and Reproductive Biology 4 \u00a0 4 \u00a0 \u00a0 Brigham and Women's Hospital, Boston, MA 02115, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2006,5,3]]},"reference":[{"key":"2023012408542571200_b1","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1038\/75556","article-title":"Gene Ontology: tool for the unification of biology","volume":"25","author":"Ashburner","year":"2000","journal-title":"Nat. Genet."},{"key":"2023012408542571200_b2","first-page":"206","article-title":"A variational Bayesian framework for graphical models","volume-title":"Advances in Neural Information Processing Systems 12","author":"Attias","year":"2000"},{"key":"2023012408542571200_b3","doi-asserted-by":"crossref","first-page":"281","DOI":"10.1089\/106652799318274","article-title":"Clustering gene expression patterns","volume":"6","author":"Ben-Dor","year":"1999","journal-title":"J. Comput. Biol."},{"key":"2023012408542571200_b4","doi-asserted-by":"crossref","first-page":"163","DOI":"10.1007\/0-387-23077-7_13","article-title":"Use of microarray data via model-based classification in the study and prediction of survival from lung cancer","volume-title":"Methods of Microarray Data Analysis IV","author":"Ben-Tovim Jones","year":"2005"},{"key":"2023012408542571200_b5","article-title":"Statistical approaches to analysing microarray data representing periodic biological processes: a case study using the yeast cell cycle","author":"Booth","year":"2004"},{"key":"2023012408542571200_b6","doi-asserted-by":"crossref","first-page":"331","DOI":"10.1093\/bib\/6.4.331","article-title":"Unsupervised pattern recognition: an introduction to the whys and wherefores of clustering microarray data","volume":"6","author":"Boutros","year":"2005","journal-title":"Brief Bioinform"},{"key":"2023012408542571200_b7","doi-asserted-by":"crossref","first-page":"243","DOI":"10.1191\/1471082X05st096oa","article-title":"Mixture of linear mixed models for clustering gene expression profiles from repeated microarray experiments","volume":"5","author":"Celeux","year":"2005","journal-title":"Stat. Model."},{"key":"2023012408542571200_b8","doi-asserted-by":"crossref","first-page":"687","DOI":"10.1081\/BIP-200025659","article-title":"A knowledge-based clustering algorithm driven by gene ontology","volume":"14","author":"Cheng","year":"2004","journal-title":"J. Biopharm. Stat."},{"key":"2023012408542571200_b9","first-page":"511","article-title":"How well do we understand the clusters in microarray data?","volume":"2","author":"Clare","year":"2002","journal-title":"In Silico Biol."},{"key":"2023012408542571200_b10","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1111\/j.2517-6161.1977.tb01600.x","article-title":"Maximum likelihood from incomplete data via the EM algorithm (with discussion)","volume":"39","author":"Dempster","year":"1977","journal-title":"J. R. Stat. Soc. B"},{"key":"2023012408542571200_b11","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4899-4541-9","volume-title":"An Introduction to the Bootstrap","author":"Efron","year":"1993"},{"key":"2023012408542571200_b12","doi-asserted-by":"crossref","first-page":"578","DOI":"10.1093\/comjnl\/41.8.578","article-title":"How many clusters? Which clustering method? Answers via model-based cluster analysis","volume":"41","author":"Fraley","year":"1998","journal-title":"Comp J."},{"key":"2023012408542571200_b13","doi-asserted-by":"crossref","first-page":"275","DOI":"10.1093\/bioinformatics\/18.2.275","article-title":"Mixture modelling of gene expression data from microarray experiments","volume":"18","author":"Ghosh","year":"2002","journal-title":"Bioinformatics"},{"key":"2023012408542571200_b14","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1186\/1297-9686-36-1-3","article-title":"Mixture model for inferring susceptibility to mastitis in diary cattle: a procedure for likelihood-based inference","volume":"36","author":"Gianola","year":"2004","journal-title":"Genet. Sel. Evol."},{"key":"2023012408542571200_b15","doi-asserted-by":"crossref","first-page":"1574","DOI":"10.1101\/gr.397002","article-title":"Judging the quality of gene expression-based clustering methods using gene annotation","volume":"12","author":"Gibbons","year":"2002","journal-title":"Genome Res."},{"key":"2023012408542571200_b16","volume-title":"Multilevel Statistical Models","author":"Goldstein","year":"1995","edition":"(2nd edn)"},{"key":"2023012408542571200_b17","doi-asserted-by":"crossref","first-page":"193","DOI":"10.1007\/BF01908075","article-title":"Comparing partitions","volume":"2","author":"Hubert","year":"1985","journal-title":"J. Classif."},{"key":"2023012408542571200_b18","doi-asserted-by":"crossref","first-page":"929","DOI":"10.1126\/science.292.5518.929","article-title":"Integrated genomic and proteomic analyses of a systemically perturbed metabolic network","volume":"292","author":"Ideker","year":"2001","journal-title":"Science"},{"issue":"Issue 1","key":"2023012408542571200_b19","article-title":"A new type of stochastic dependence revealed in gene expression data","volume":"5","author":"Klebanov","year":"2006","journal-title":"Stat. Appl. Genetics Mol. Biol."},{"key":"2023012408542571200_b20","doi-asserted-by":"crossref","first-page":"9834","DOI":"10.1073\/pnas.97.18.9834","article-title":"Importance of replication in microarray gene expression studies: statistical methods and evidence from repetitive cDNA hybridizations","volume":"97","author":"Lee","year":"2000","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012408542571200_b21","doi-asserted-by":"crossref","first-page":"474","DOI":"10.1093\/bioinformatics\/btg014","article-title":"Clustering of time-course gene expression data using a mixed-effects model with B-splines","volume":"19","author":"Luan","year":"2003","journal-title":"Bioinformatics"},{"key":"2023012408542571200_b22","volume-title":"Generalized, Linear, and Mixed Models","author":"McCulloch","year":"2001"},{"key":"2023012408542571200_b23","doi-asserted-by":"crossref","first-page":"318","DOI":"10.2307\/2347790","article-title":"On bootstrapping the likelihood ratio test statistic for the number of components in a normal mixture","volume":"36","author":"McLachlan","year":"1987","journal-title":"Appl. Stat."},{"key":"2023012408542571200_b24","doi-asserted-by":"crossref","DOI":"10.1002\/0471725293","volume-title":"Discriminant Analysis and Statistical Pattern Recognition","author":"McLachlan","year":"1992"},{"key":"2023012408542571200_b25","volume-title":"Mixture Models: Inference and Applications to Clustering","author":"McLachlan","year":"1988"},{"key":"2023012408542571200_b26","doi-asserted-by":"crossref","first-page":"413","DOI":"10.1093\/bioinformatics\/18.3.413","article-title":"A mixture model-based approach to the clustering of microarray expression data","volume":"18","author":"McLachlan","year":"2002","journal-title":"Bioinformatics"},{"key":"2023012408542571200_b27","doi-asserted-by":"crossref","DOI":"10.1002\/047172842X","volume-title":"Analyzing Microarray Gene Expression Data","author":"McLachlan","year":"2004"},{"key":"2023012408542571200_b28","doi-asserted-by":"crossref","first-page":"90","DOI":"10.1016\/j.jmva.2004.02.002","article-title":"On a resampling approach for tests on the number of clusters with mixture model-based clustering of tissue samples","volume":"90","author":"McLachlan","year":"2004","journal-title":"J. Multivar. Anal."},{"key":"2023012408542571200_b29","doi-asserted-by":"crossref","volume-title":"Finite Mixture Models","author":"McLachlan","DOI":"10.1002\/0471721182"},{"key":"2023012408542571200_b30","doi-asserted-by":"crossref","DOI":"10.18637\/jss.v004.i02","article-title":"The EMMIX software for the fitting of mixtures of normal and t-components","volume":"4","author":"McLachlan","year":"1999","journal-title":"J. Stat. Software"},{"key":"2023012408542571200_b31","doi-asserted-by":"crossref","first-page":"1194","DOI":"10.1093\/bioinformatics\/18.9.1194","article-title":"Bayesian infinite mixture model based clustering of gene expression profiles","volume":"18","author":"Medvedovic","year":"2002","journal-title":"Bioinformatics"},{"key":"2023012408542571200_b32","doi-asserted-by":"crossref","first-page":"R21","DOI":"10.1186\/gb-2003-4-3-r21","article-title":"Identification of expressed genes linked to malignancy of human colorectal carcinoma by parametric clustering of quantitative expression data","volume":"4","author":"Muro","year":"2003","journal-title":"Genome Biol."},{"key":"2023012408542571200_b33","first-page":"137","article-title":"The EM algorithm","volume-title":"Handbook of Computational Statistics Vol. 1","author":"Ng","year":"2004"},{"key":"2023012408542571200_b34","doi-asserted-by":"crossref","first-page":"2652","DOI":"10.3168\/jds.S0022-0302(05)72942-8","article-title":"A Bayesian threshold-normal mixture model for analysis of a continuous mastitis-related trait","volume":"88","author":"\u00d8deg\u00e5rd","year":"2005","journal-title":"J. Dairy Sci."},{"key":"2023012408542571200_b35","doi-asserted-by":"crossref","first-page":"795","DOI":"10.1093\/bioinformatics\/btl011","article-title":"Incorporating gene functions as priors in model-based clustering of microarray gene expression data","volume":"22","author":"Pan","year":"2006","journal-title":"Bioinformatics"},{"key":"2023012408542571200_b36","doi-asserted-by":"crossref","DOI":"10.1186\/gb-2002-3-2-research0009","article-title":"Model-based cluster analysis of microarray gene-expression data","volume":"3","author":"Pan","year":"2002","journal-title":"Genome Biol."},{"key":"2023012408542571200_b37","doi-asserted-by":"crossref","first-page":"1620","DOI":"10.1093\/bioinformatics\/btg227","article-title":"The effect of replication on gene expression microarray experiments","volume":"19","author":"Pavlidis","year":"2003","journal-title":"Bioinformatics"},{"key":"2023012408542571200_b38","doi-asserted-by":"crossref","first-page":"461","DOI":"10.1214\/aos\/1176344136","article-title":"Estimating the dimension of a model","volume":"6","author":"Schwarz","year":"1978","journal-title":"Ann. Stat."},{"key":"2023012408542571200_b39","doi-asserted-by":"crossref","first-page":"3273","DOI":"10.1091\/mbc.9.12.3273","article-title":"Comprehensive identification of cell cycle-regulated genes of the yeast Saccharomyces cerevisiae by microarray hybridization","volume":"9","author":"Spellman","year":"1998","journal-title":"Mol. Biol. Cell"},{"key":"2023012408542571200_b40","doi-asserted-by":"crossref","first-page":"12837","DOI":"10.1073\/pnas.0504609102","article-title":"Significance analysis of time course microarray experiments","volume":"102","author":"Storey","year":"2005","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012408542571200_b41","doi-asserted-by":"crossref","first-page":"977","DOI":"10.1093\/bioinformatics\/17.10.977","article-title":"Model-based clustering and data transformations for gene expression data","volume":"17","author":"Yeung","year":"2001","journal-title":"Bioinformatics"},{"key":"2023012408542571200_b42","doi-asserted-by":"crossref","first-page":"R34","DOI":"10.1186\/gb-2003-4-5-r34","article-title":"Clustering gene-expression data with repeated measurements","volume":"4","author":"Yeung","year":"2003","journal-title":"Genome Biol."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/14\/1745\/48841337\/bioinformatics_22_14_1745.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/14\/1745\/48841337\/bioinformatics_22_14_1745.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,8]],"date-time":"2025-01-08T20:28:33Z","timestamp":1736368113000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/22\/14\/1745\/227125"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,5,3]]},"references-count":42,"journal-issue":{"issue":"14","published-print":{"date-parts":[[2006,7,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btl165","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2006,7,15]]},"published":{"date-parts":[[2006,5,3]]}}}