{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T12:24:47Z","timestamp":1767961487986,"version":"3.49.0"},"reference-count":26,"publisher":"Oxford University Press (OUP)","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2006,2,15]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Differential gene expression detection and sample classification using microarray data have received much research interest recently. Owing to the large number of genes p and small number of samples n (p \u226b n), microarray data analysis poses big challenges for statistical analysis. An obvious problem owing to the \u2018large p small n\u2019 is over-fitting. Just by chance, we are likely to find some non-differentially expressed genes that can classify the samples very well. The idea of shrinkage is to regularize the model parameters to reduce the effects of noise and produce reliable inferences. Shrinkage has been successfully applied in the microarray data analysis. The SAM statistics proposed by Tusher et al. and the \u2018nearest shrunken centroid\u2019 proposed by Tibshirani et al. are ad hoc shrinkage methods. Both methods are simple, intuitive and prove to be useful in empirical studies.<\/jats:p><jats:p>Recently Wu proposed the penalized t\/F-statistics with shrinkage by formally using the \u21121 penalized linear regression models for two-class microarray data, showing good performance. In this paper we systematically discussed the use of penalized regression models for analyzing microarray data. We generalize the two-class penalized t\/F-statistics proposed by Wu to multi-class microarray data. We formally derive the ad hoc shrunken centroid used by Tibshirani et al. using the \u21121 penalized regression models. And we show that the penalized linear regression models provide a rigorous and unified statistical framework for sample classification and differential gene expression detection.<\/jats:p><jats:p>Availability: For the computer programs, detailed analysis results and R functions for the proposed methods, please see<\/jats:p><jats:p>Contact: \u00a0baolin@biostat.umn.edu<\/jats:p><jats:p>Supplementary information: \u00a0<\/jats:p>","DOI":"10.1093\/bioinformatics\/bti827","type":"journal-article","created":{"date-parts":[[2005,12,14]],"date-time":"2005-12-14T02:28:47Z","timestamp":1134527327000},"page":"472-476","source":"Crossref","is-referenced-by-count":25,"title":["Differential gene expression detection and sample classification using penalized linear regression models"],"prefix":"10.1093","volume":"22","author":[{"given":"Baolin","family":"Wu","sequence":"first","affiliation":[{"name":"Division of Biostatistics, School of Public Health, University of Minnesota \u00a0 A460 Mayo Building, MMC 303, Minneapolis, MN 55455, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2005,12,13]]},"reference":[{"key":"2023012408502729300_b1","doi-asserted-by":"crossref","first-page":"503","DOI":"10.1038\/35000501","article-title":"Distinct types of diffuse large B-cell lymphoma identified by gene expression profiling","volume":"403","author":"Alizadeh","year":"2000","journal-title":"Nature"},{"key":"2023012408502729300_b2","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1111\/j.2517-6161.1995.tb02031.x","article-title":"Controlling the false discovery rate: a practical and powerful approach to multiple testing","volume":"57","author":"Benjamini","year":"1995","journal-title":"J. R. Stat. Soc. B."},{"key":"2023012408502729300_b3","doi-asserted-by":"crossref","first-page":"185","DOI":"10.1093\/bioinformatics\/19.2.185","article-title":"A comparison of normalization methods for high density oligonucleotide array data based on variance and bias","volume":"19","author":"Bolstad","year":"2003","journal-title":"Bioinformatics"},{"key":"2023012408502729300_b4","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1010933404324","article-title":"Random forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learning"},{"key":"2023012408502729300_b5","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1023\/A:1009715923555","article-title":"A tutorial on support vector machines for pattern recognition","volume":"2","author":"Burges","year":"1998","journal-title":"Data Min. Knowl. Disc."},{"key":"2023012408502729300_b6","article-title":"Libsvm : a library for support vector machines","author":"Chang","year":"2001"},{"key":"2023012408502729300_b7","doi-asserted-by":"crossref","first-page":"457","DOI":"10.1038\/ng1296-457","article-title":"Use of a cDNA microarray to analyse gene expression patterns in human cancer","volume":"14","author":"DeRisi","year":"1996","journal-title":"Nat. Genet."},{"key":"2023012408502729300_b8","doi-asserted-by":"crossref","first-page":"425","DOI":"10.1093\/biomet\/81.3.425","article-title":"Ideal spatial adaptation by wavelet shrinkage","volume":"81","author":"Donoho","year":"1994","journal-title":"Biometrika"},{"key":"2023012408502729300_b9","first-page":"111","article-title":"Statistical methods for identifying differentially expressed genes in replicated cDNA microarray experiments","volume":"12","author":"Dudoit","year":"2002","journal-title":"Stat. Sinica"},{"key":"2023012408502729300_b10","doi-asserted-by":"crossref","first-page":"407","DOI":"10.1214\/009053604000000067","article-title":"Least angle regression","volume":"32","author":"Efron","year":"2004","journal-title":"Ann. Stat."},{"key":"2023012408502729300_b11","doi-asserted-by":"crossref","first-page":"531","DOI":"10.1126\/science.286.5439.531","article-title":"Molecular classification of cancer: class discovery and class prediction by gene expression monitoring","volume":"286","author":"Golub","year":"1999","journal-title":"Science"},{"key":"2023012408502729300_b12","doi-asserted-by":"crossref","first-page":"673","DOI":"10.1038\/89044","article-title":"Classification and diagnostic prediction of cancers using gene expression profiling and artificial neural networks","volume":"7","author":"Khan","year":"2001","journal-title":"Nat. Med."},{"key":"2023012408502729300_b13","volume-title":"Applied Linear Regression Models","author":"Kutner","year":"2004","edition":"4th edn."},{"key":"2023012408502729300_b14","doi-asserted-by":"crossref","first-page":"1675","DOI":"10.1038\/nbt1296-1675","article-title":"Expression monitoring by hybridization to high-density oligonucleotide arrays","volume":"14","author":"Lockhart","year":"1996","journal-title":"Nat. Biotechnol."},{"key":"2023012408502729300_b15","doi-asserted-by":"crossref","DOI":"10.1007\/b97411","volume-title":"The Analysis of Gene Expression Data: Methods and Software.","author":"Parmigiani","year":"2003"},{"key":"2023012408502729300_b16","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/415436a","article-title":"Prediction of central nervous system embryonal tumor outcome based on gene expression","volume":"415","author":"Pomeroy","year":"2002","journal-title":"Nature"},{"key":"2023012408502729300_b17","doi-asserted-by":"crossref","DOI":"10.1201\/9780203011232","volume-title":"Statistical Analysis of Gene Expression Microarray Data","author":"Speed","year":"2003"},{"key":"2023012408502729300_b18","first-page":"260","article-title":"The optimal discovery procedure II: applications to comparative microarray experiments","volume-title":"UW Biostatistics Working Paper Series","author":"Storey","year":"2005"},{"key":"2023012408502729300_b19","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1111\/j.2517-6161.1996.tb02080.x","article-title":"Regression shrinkage and selection via the lasso","volume":"58","author":"Tibshirani","year":"1996","journal-title":"J. R. Statistical Soc. B, Methodological"},{"key":"2023012408502729300_b20","doi-asserted-by":"crossref","first-page":"6567","DOI":"10.1073\/pnas.082099299","article-title":"Diagnosis of multiple cancer types by shrunken centroids of gene expression","volume":"99","author":"Tibshirani","year":"2002","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"2023012408502729300_b21","doi-asserted-by":"crossref","first-page":"104","DOI":"10.1214\/ss\/1056397488","article-title":"Class prediction by nearest shrunken centroids, with application to DNA microarrays","volume":"18","author":"Tibshirani","year":"2003","journal-title":"Stat. Sci."},{"key":"2023012408502729300_b22","doi-asserted-by":"crossref","first-page":"5116","DOI":"10.1073\/pnas.091062498","article-title":"Significance analysis of microarrays applied to the ionizing radiation response","volume":"98","author":"Tusher","year":"2001","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"2023012408502729300_b23","doi-asserted-by":"crossref","first-page":"11462","DOI":"10.1073\/pnas.201162998","article-title":"Predicting the clinical status of human breast cancer by using gene expression profiles","volume":"98","author":"West","year":"2001","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"2023012408502729300_b24","doi-asserted-by":"crossref","first-page":"1565","DOI":"10.1093\/bioinformatics\/bti217","article-title":"Differential gene expression detection using penalized linear regression models: the improved SAM statistics","volume":"21","author":"Wu","year":"2005","journal-title":"Bioinformatics"},{"key":"2023012408502729300_b25","doi-asserted-by":"crossref","first-page":"1636","DOI":"10.1093\/bioinformatics\/btg210","article-title":"Comparison of statistical methods for classification of ovarian cancer using mass spectrometry data","volume":"19","author":"Wu","year":"2003","journal-title":"Bioinformatics"},{"key":"2023012408502729300_b26","doi-asserted-by":"crossref","first-page":"e15","DOI":"10.1093\/nar\/30.4.e15","article-title":"Normalization for cDNA microarray data: a robust composite method addressing single and multiple slide systematic variation","volume":"30","author":"Yang","year":"2002","journal-title":"Nucleic Acids Res."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/4\/472\/48838871\/bioinformatics_22_4_472.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/4\/472\/48838871\/bioinformatics_22_4_472.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,6]],"date-time":"2025-01-06T13:46:33Z","timestamp":1736171193000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/22\/4\/472\/184122"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,12,13]]},"references-count":26,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2006,2,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bti827","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2006,2,15]]},"published":{"date-parts":[[2005,12,13]]}}}