{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T21:12:10Z","timestamp":1774473130235,"version":"3.50.1"},"reference-count":15,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2010,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec><jats:title>Background<\/jats:title><jats:p>Before conducting a microarray experiment, one important issue that needs to be determined is the number of arrays required in order to have adequate power to identify differentially expressed genes. This paper discusses some crucial issues in the problem formulation, parameter specifications, and approaches that are commonly proposed for sample size estimation in microarray experiments. Common methods for sample size estimation are formulated as the minimum sample size necessary to achieve a specified sensitivity (proportion of detected truly differentially expressed genes)<jats:italic>on average<\/jats:italic>at a specified false discovery rate (FDR) level and specified expected proportion (<jats:italic>\u03c0<\/jats:italic><jats:sub>1<\/jats:sub>) of the true differentially expression genes in the array. Unfortunately, the probability of detecting the specified sensitivity in such a formulation can be low. We formulate the sample size problem as the number of arrays needed to achieve a specified sensitivity with<jats:italic>95% probability<\/jats:italic>at the specified significance level. A permutation method using a small pilot dataset to estimate sample size is proposed. This method accounts for correlation and effect size heterogeneity among genes.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>A sample size estimate based on the common formulation, to achieve the desired sensitivity on average, can be calculated using a univariate method without taking the correlation among genes into consideration. This formulation of sample size problem is inadequate because the probability of detecting the specified sensitivity can be lower than 50%. On the other hand, the needed sample size calculated by the proposed permutation method will ensure detecting at least the desired sensitivity with 95% probability. The method is shown to perform well for a real example dataset using a small pilot dataset with 4-6 samples per group.<\/jats:p><\/jats:sec><jats:sec><jats:title>Conclusions<\/jats:title><jats:p>We recommend that the sample size problem should be formulated to detect a specified proportion of differentially expressed genes with 95% probability. This formulation ensures finding the desired proportion of true positives with high probability. The proposed permutation method takes the correlation structure and effect size heterogeneity into consideration and works well using only a small pilot dataset.<\/jats:p><\/jats:sec>","DOI":"10.1186\/1471-2105-11-48","type":"journal-article","created":{"date-parts":[[2010,1,26]],"date-time":"2010-01-26T07:16:11Z","timestamp":1264490171000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":52,"title":["Power and sample size estimation in microarray studies"],"prefix":"10.1186","volume":"11","author":[{"given":"Wei-Jiun","family":"Lin","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huey-Miin","family":"Hsueh","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"James J","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2010,1,25]]},"reference":[{"key":"3505_CR1","doi-asserted-by":"publisher","first-page":"24","DOI":"10.1152\/physiolgenomics.00037.2003","volume":"16","author":"MCK Yang","year":"2003","unstructured":"Yang MCK, Yang JJ, McIndoe RA, et al.: Microarray experimental design: power and sample size considerations. Physiol Genomics 2003, 16: 24\u201328. 10.1152\/physiolgenomics.00037.2003","journal-title":"Physiol Genomics"},{"key":"3505_CR2","doi-asserted-by":"publisher","first-page":"714","DOI":"10.1089\/cmb.2004.11.714","volume":"11","author":"SJ Wang","year":"2004","unstructured":"Wang SJ, Chen JJ: Sample size for identifying differentially expressed genes in microarray experiments. J Comput Biol 2004, 11: 714\u2013726. 10.1089\/cmb.2004.11.714","journal-title":"J Comput Biol"},{"key":"3505_CR3","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1093\/biostatistics\/kxh026","volume":"6","author":"S-H Jung","year":"2005","unstructured":"Jung S-H, Bang H, Young S: Sample size calculation for multiple testing in microarray data analysis. Biostatistics 2005, 6: 157\u2013169. 10.1093\/biostatistics\/kxh026","journal-title":"Biostatistics"},{"key":"3505_CR4","doi-asserted-by":"publisher","first-page":"S3097","DOI":"10.1093\/bioinformatics\/bti456","volume":"21","author":"S-H Jung","year":"2005","unstructured":"Jung S-H: Sample size for FDR-control in microarray data analysis. Bioinformatics 2005, 21: S3097\u20133104. 10.1093\/bioinformatics\/bti456","journal-title":"Bioinformatics"},{"key":"3505_CR5","doi-asserted-by":"publisher","first-page":"4263","DOI":"10.1093\/bioinformatics\/bti699","volume":"21","author":"S Pounds","year":"2005","unstructured":"Pounds S, Cheng C: Sample size determination for the false discovery rate. Bioinformatics 2005, 21: 4263\u20134267. 10.1093\/bioinformatics\/bti699","journal-title":"Bioinformatics"},{"key":"3505_CR6","doi-asserted-by":"publisher","first-page":"2267","DOI":"10.1002\/sim.2119","volume":"24","author":"SS Li","year":"2005","unstructured":"Li SS, Bigler J, Lampe JW, Potter JD, Feng Z: FDR-controlling testing procedures and sample size determination for microarrays. Statist Med 2005, 24: 2267\u20132280. 10.1002\/sim.2119","journal-title":"Statist Med"},{"key":"3505_CR7","doi-asserted-by":"publisher","first-page":"1502","DOI":"10.1093\/bioinformatics\/bti162","volume":"21","author":"C-A Tsai","year":"2005","unstructured":"Tsai C-A, Wang S-J, Chen D-T, et al.: Sample size for gene expression microarray experiments. Bioinformatics 2005, 21: 1502\u20131508. 10.1093\/bioinformatics\/bti162","journal-title":"Bioinformatics"},{"key":"3505_CR8","doi-asserted-by":"publisher","first-page":"4219","DOI":"10.1002\/sim.2862","volume":"26","author":"Y Shao","year":"2007","unstructured":"Shao Y, Tseng C-H: Sample size calculation with dependence adjustment for FDR-control in microarray studies. Statist Med 2007, 26: 4219\u20134237. 10.1002\/sim.2862","journal-title":"Statist Med"},{"key":"3505_CR9","doi-asserted-by":"publisher","first-page":"3543","DOI":"10.1002\/sim.1335","volume":"21","author":"M-L Lee","year":"2002","unstructured":"Lee M-L, Whitmore G: Power and sample size for DNA microarray studies. Statist Med 2002, 21: 3543\u201370. 10.1002\/sim.1335","journal-title":"Statist Med"},{"key":"3505_CR10","doi-asserted-by":"publisher","first-page":"106","DOI":"10.1186\/1471-2105-7-106","volume":"7","author":"R Tibshirani","year":"2006","unstructured":"Tibshirani R: A simple method for assessing sample sizes in microarray experiments. BMC Bioinformatics 2006, 7: 106. 10.1186\/1471-2105-7-106","journal-title":"BMC Bioinformatics"},{"key":"3505_CR11","doi-asserted-by":"publisher","first-page":"27","DOI":"10.1093\/biostatistics\/kxh015","volume":"6","author":"K Dobbin","year":"2005","unstructured":"Dobbin K, Simon R: Sample size determination in microarray experiments for class comparison and prognostic classification. Biostatistics 2005, 6: 27\u201338. 10.1093\/biostatistics\/kxh015","journal-title":"Biostatistics"},{"key":"3505_CR12","volume-title":"Statistical Methods for Meta-Analysis","author":"LV Hedges","year":"1985","unstructured":"Hedges LV, Olkin I: Statistical Methods for Meta-Analysis. Academic Press; 1985."},{"key":"3505_CR13","doi-asserted-by":"publisher","first-page":"479","DOI":"10.1111\/1467-9868.00346","volume":"64","author":"JD Storey","year":"2002","unstructured":"Storey JD: A direct approach to false discovery rates. Journal of the Royal Statistical Society, Series B 2002, 64: 479\u2013498. 10.1111\/1467-9868.00346","journal-title":"Journal of the Royal Statistical Society, Series B"},{"key":"3505_CR14","doi-asserted-by":"publisher","first-page":"6745","DOI":"10.1073\/pnas.96.12.6745","volume":"96","author":"U Alon","year":"1999","unstructured":"Alon U, Barkai N, Notterman DA, et al.: Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays. Proc Natl Acad Sci 1999, 96: 6745\u20136750. 10.1073\/pnas.96.12.6745","journal-title":"Proc Natl Acad Sci"},{"key":"3505_CR15","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1111\/j.2517-6161.1995.tb02031.x","volume":"57","author":"Y Benjamini","year":"1995","unstructured":"Benjamini Y, Hochberg y: Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society, Series B 1995, 57: 289\u2013300.","journal-title":"Journal of the Royal Statistical Society, Series B"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-11-48.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,16]],"date-time":"2025-02-16T23:38:47Z","timestamp":1739749127000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-11-48"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2010,1,25]]},"references-count":15,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2010,12]]}},"alternative-id":["3505"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-11-48","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2010,1,25]]},"assertion":[{"value":"18 August 2009","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 January 2010","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 January 2010","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"48"}}