{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T10:56:11Z","timestamp":1781088971102,"version":"3.54.1"},"reference-count":40,"publisher":"Oxford University Press (OUP)","issue":"21","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2009,11,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: Genome-wide association studies (GWAS) interrogate common genetic variation across the entire human genome in an unbiased manner and hold promise in identifying genetic variants with moderate or weak effect sizes. However, conventional testing procedures, which are mostly P-value based, ignore the dependency and therefore suffer from loss of efficiency. The goal of this article is to exploit the dependency information among adjacent single nucleotide polymorphisms (SNPs) to improve the screening efficiency in GWAS.<\/jats:p><jats:p>Results: We propose to model the linear block dependency in the SNP data using hidden Markov models (HMMs). A compound decision\u2013theoretic framework for testing HMM-dependent hypotheses is developed. We propose a powerful data-driven procedure [pooled local index of significance (PLIS)] that controls the false discovery rate (FDR) at the nominal level. PLIS is shown to be optimal in the sense that it has the smallest false negative rate (FNR) among all valid FDR procedures. By re-ranking significance for all SNPs with dependency considered, PLIS gains higher power than conventional P-value based methods. Simulation results demonstrate that PLIS dominates conventional FDR procedures in detecting disease-associated SNPs. Our method is applied to analysis of the SNP data from a GWAS of type 1 diabetes. Compared with the Benjamini\u2013Hochberg (BH) procedure, PLIS yields more accurate results and has better reproducibility of findings.<\/jats:p><jats:p>Conclusion: The genomic rankings based on our procedure are substantially different from the rankings based on the P-values. By integrating information from adjacent locations, the PLIS rankings benefit from the increased signal-to-noise ratio, hence our procedure often has higher statistical power and better reproducibility. It provides a promising direction in large-scale GWAS.<\/jats:p><jats:p>Availability: An R package PLIS has been developed to implement the PLIS procedure. Source codes are available upon request and will be available on CRAN (http:\/\/cran.r-project.org\/).<\/jats:p><jats:p>Contact: \u00a0zhiwei@njit.edu<\/jats:p><jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btp476","type":"journal-article","created":{"date-parts":[[2009,8,5]],"date-time":"2009-08-05T00:13:27Z","timestamp":1249431207000},"page":"2802-2808","source":"Crossref","is-referenced-by-count":49,"title":["Multiple testing in genome-wide association studies via hidden Markov models"],"prefix":"10.1093","volume":"25","author":[{"given":"Zhi","family":"Wei","sequence":"first","affiliation":[{"name":"1 Department of Computer Science, New Jersey Institute of Technology, Newark, NJ 07102, 2 Department of Statistics, North Carolina State University, Raleigh, NC 27695, 3 Center for Applied Genomics, The Children's Hospital of Philadelphia and 4 Division of Genetics, Department of Pediatrics, The Children's Hospital of Philadelphia, Philadelphia, PA 19104, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenguang","family":"Sun","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science, New Jersey Institute of Technology, Newark, NJ 07102, 2 Department of Statistics, North Carolina State University, Raleigh, NC 27695, 3 Center for Applied Genomics, The Children's Hospital of Philadelphia and 4 Division of Genetics, Department of Pediatrics, The Children's Hospital of Philadelphia, Philadelphia, PA 19104, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kai","family":"Wang","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science, New Jersey Institute of Technology, Newark, NJ 07102, 2 Department of Statistics, North Carolina State University, Raleigh, NC 27695, 3 Center for Applied Genomics, The Children's Hospital of Philadelphia and 4 Division of Genetics, Department of Pediatrics, The Children's Hospital of Philadelphia, Philadelphia, PA 19104, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hakon","family":"Hakonarson","sequence":"additional","affiliation":[{"name":"1 Department of Computer Science, New Jersey Institute of Technology, Newark, NJ 07102, 2 Department of Statistics, North Carolina State University, Raleigh, NC 27695, 3 Center for Applied Genomics, The Children's Hospital of Philadelphia and 4 Division of Genetics, Department of Pediatrics, The Children's Hospital of Philadelphia, Philadelphia, PA 19104, USA"},{"name":"1 Department of Computer Science, New Jersey Institute of Technology, Newark, NJ 07102, 2 Department of Statistics, North Carolina State University, Raleigh, NC 27695, 3 Center for Applied Genomics, The Children's Hospital of Philadelphia and 4 Division of Genetics, Department of Pediatrics, The Children's Hospital of Philadelphia, Philadelphia, PA 19104, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2009,5,4]]},"reference":[{"key":"2023013112193337100_B1","doi-asserted-by":"crossref","first-page":"703","DOI":"10.1038\/ng.381","article-title":"Genome-wide association study and meta-analysis find that over 40 loci affect risk of type 1 diabetes","volume":"41","author":"Barrett","year":"2009","journal-title":"Nat. Genet."},{"key":"2023013112193337100_B2","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1111\/j.2517-6161.1995.tb02031.x","article-title":"Controlling the false discovery rate: a practical and powerful approach to multiple testing","volume":"57","author":"Benjamini","year":"1995","journal-title":"J. R. Stat. Soc. B"},{"key":"2023013112193337100_B3","doi-asserted-by":"crossref","first-page":"60","DOI":"10.3102\/10769986025001060","article-title":"On the adaptive control of the false discovery rate in multiple testing with independent statistics","volume":"25","author":"Benjamini","year":"2000","journal-title":"J. Educ. Behav. Stat."},{"key":"2023013112193337100_B4","doi-asserted-by":"crossref","first-page":"1165","DOI":"10.1214\/aos\/1013699998","article-title":"The control of the false discovery rate in multiple testing under dependency","volume":"29","author":"Benjamini","year":"2001","journal-title":"Ann. Stat."},{"key":"2023013112193337100_B5","doi-asserted-by":"crossref","first-page":"1158","DOI":"10.1086\/522036","article-title":"So many correlated tests, so little time! Rapid adjustment of P values for multiple correlated tests","volume":"81","author":"Conneely","year":"2007","journal-title":"Am. J. Hum. Genet."},{"key":"2023013112193337100_B6","first-page":"111","article-title":"Statistical methods for identifying differentially expressed genes in replicated cDNA microarray experiments","volume":"12","author":"Dudoit","year":"2002","journal-title":"Stat. Sin."},{"key":"2023013112193337100_B7","doi-asserted-by":"crossref","first-page":"1151","DOI":"10.1198\/016214501753382129","article-title":"Empirical Bayes analysis of a microarray experiment","volume":"96","author":"Efron","year":"2001","journal-title":"J. Am. Stat. Assoc."},{"key":"2023013112193337100_B8","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1198\/016214504000000089","article-title":"Large-scale simultaneous hypothesis testing: the choice of a null hypothesis","volume":"99","author":"Efron","year":"2004","journal-title":"J. Am. Stat. Assoc."},{"key":"2023013112193337100_B9","doi-asserted-by":"crossref","first-page":"93","DOI":"10.1198\/016214506000001211","article-title":"Correlation and large-scale simultaneous testing","volume":"102","author":"Efron","year":"2007","journal-title":"J. Am. Stat. Assoc."},{"key":"2023013112193337100_B10","first-page":"197","article-title":"Simultaneous inference: when should hypothesis testing problems be combined?","volume":"1","author":"Efron","year":"2008","journal-title":"Ann. Appl. Stat."},{"key":"2023013112193337100_B11","doi-asserted-by":"crossref","first-page":"1518","DOI":"10.1109\/TIT.2002.1003838","article-title":"Hidden Markov processes","volume":"48","author":"Ephraim","year":"2002","journal-title":"IEEE Trans. Inf. Theory"},{"key":"2023013112193337100_B12","doi-asserted-by":"crossref","first-page":"275","DOI":"10.1111\/j.1467-9469.2006.00530.x","article-title":"Some results on the control of the false discovery rate under dependence","volume":"34","author":"Farcomeni","year":"2007","journal-title":"Scand. J. Stat."},{"key":"2023013112193337100_B13","volume-title":"Statistical Methods for Research Workers","author":"Fisher","year":"1958","edition":"13th"},{"key":"2023013112193337100_B14","doi-asserted-by":"crossref","first-page":"499","DOI":"10.1111\/1467-9868.00347","article-title":"Operating characteristic and extensions of the false discovery rate procedure","volume":"64","author":"Genovese","year":"2002","journal-title":"J. R. Stat. Soc. B"},{"key":"2023013112193337100_B15","doi-asserted-by":"crossref","first-page":"290","DOI":"10.2337\/db08-1022","article-title":"Follow up analysis of genome-wide association data identifies novel loci for type 1 diabetes","volume":"58","author":"Grant","year":"2008","journal-title":"Diabetes"},{"key":"2023013112193337100_B16","doi-asserted-by":"crossref","first-page":"13","DOI":"10.2202\/1544-6115.1360","article-title":"Adaptive choice of the number of bootstrap samples in large scale multiple testing","volume":"7","author":"Guo","year":"2008","journal-title":"Stat. Appl. Genet. Mol. Biol."},{"key":"2023013112193337100_B17","doi-asserted-by":"crossref","first-page":"591","DOI":"10.1038\/nature06010","article-title":"A genome-wide association study identifies KIAA0350 as a type 1 diabetes gene","volume":"448","author":"Hakonarson","year":"2007","journal-title":"Nature"},{"key":"2023013112193337100_B18","doi-asserted-by":"crossref","first-page":"R116","DOI":"10.1093\/hmg\/ddn246","article-title":"Autoimmune diseases: insights from genome-wide association studies","volume":"17","author":"Lettre","year":"2008","journal-title":"Hum. Mol. Genet."},{"key":"2023013112193337100_B19","doi-asserted-by":"crossref","first-page":"1141","DOI":"10.1080\/01621459.1996.10476984","article-title":"A smooth nonparametric estimate of a mixing distribution using mixtures of Gaussians","volume":"91","author":"Magder","year":"1996","journal-title":"J. Am. Stat. Assoc."},{"key":"2023013112193337100_B20","doi-asserted-by":"crossref","first-page":"373","DOI":"10.1214\/009053605000000741","article-title":"Estimating the proportion of false null hypotheses among a large number of independently tested hypotheses","volume":"34","author":"Meinshausen","year":"2006","journal-title":"Ann. Stat."},{"key":"2023013112193337100_B21","doi-asserted-by":"crossref","first-page":"3492","DOI":"10.1086\/324109","article-title":"Controlling the false-discovery rate in astrophysical data analysis","volume":"122","author":"Miller","year":"2001","journal-title":"Astronom. J."},{"key":"2023013112193337100_B22","doi-asserted-by":"crossref","first-page":"765","DOI":"10.1086\/383251","article-title":"A simple correction for multiple testing for single-nucleotide polymorphisms in linkage disequilibrium with each other","volume":"74","author":"Nyholt","year":"2004","journal-title":"Am. J. Hum. Genet."},{"key":"2023013112193337100_B23","doi-asserted-by":"crossref","first-page":"411","DOI":"10.1111\/j.1467-9868.2005.00509.x","article-title":"Variance of the number of false discoveries","volume":"67","author":"Owen","year":"2005","journal-title":"J. R. Stat. Soc. B"},{"key":"2023013112193337100_B24","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1007\/s10142-003-0085-7","article-title":"A mixture model approach to detecting differentially expressed genes with microarray data","volume":"3","author":"Pan","year":"2003","journal-title":"Funct. Integr. Genomics"},{"key":"2023013112193337100_B25","doi-asserted-by":"crossref","DOI":"10.2202\/1544-6115.1157","article-title":"Correlation between gene expression levels and limitations of the empirical Bayes methodology for finding differentially expressed genes","volume":"4","author":"Qiu","year":"2005","journal-title":"Stat. Appl. Genet. Mol. Biol."},{"key":"2023013112193337100_B26","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1109\/5.18626","article-title":"A tutorial on hidden Markov models and selected applications in speech recognition","volume":"77","author":"Rabiner","year":"1989","journal-title":"Proc. IEEE"},{"key":"2023013112193337100_B27","doi-asserted-by":"crossref","first-page":"829","DOI":"10.1093\/genetics\/164.2.829","article-title":"False discovery rate in linkage and association genome screens for complex disorders","volume":"164","author":"Sabatti","year":"2003","journal-title":"Genetics"},{"key":"2023013112193337100_B28","doi-asserted-by":"crossref","first-page":"35","DOI":"10.1038\/ng.271","article-title":"Genomewide association analysis of metabolic phenotypes in a birth cohort from a founder population","volume":"41","author":"Sabatti","year":"2009","journal-title":"Nat. Genet."},{"key":"2023013112193337100_B29","doi-asserted-by":"crossref","first-page":"394","DOI":"10.1214\/009053605000000778","article-title":"False discovery and false nondiscovery rates in single-step multiple testing procedures","volume":"34","author":"Sarkar","year":"2006","journal-title":"Ann. Stat."},{"key":"2023013112193337100_B30","doi-asserted-by":"crossref","first-page":"153","DOI":"10.1214\/07-AOAS133","article-title":"False discovery rate analysis of brain diffusion direction maps","volume":"2","author":"Schwartzman","year":"2008","journal-title":"Ann. Appl. Stat."},{"key":"2023013112193337100_B31","doi-asserted-by":"crossref","first-page":"479","DOI":"10.1111\/1467-9868.00346","article-title":"A direct approach to false discovery rates","volume":"64","author":"Storey","year":"2002","journal-title":"J. R. Stat. Soc. B"},{"key":"2023013112193337100_B32","doi-asserted-by":"crossref","first-page":"9440","DOI":"10.1073\/pnas.1530509100","article-title":"Statistical significance for genome-wide studies","volume":"100","author":"Storey","year":"2003","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023013112193337100_B33","doi-asserted-by":"crossref","first-page":"393","DOI":"10.1111\/j.1467-9868.2008.00694.x","article-title":"Large-scale multiple testing under dependence","volume":"71","author":"Sun","year":"2009","journal-title":"J. R. Stat. Soc. B"},{"key":"2023013112193337100_B34","doi-asserted-by":"crossref","first-page":"857","DOI":"10.1038\/ng2068","article-title":"Robust associations of four new chromosome regions from genome-wide analyses of type 1 diabetes","volume":"39","author":"Todd","year":"2007","journal-title":"Nat. Genet."},{"key":"2023013112193337100_B35","doi-asserted-by":"crossref","first-page":"5116","DOI":"10.1073\/pnas.091062498","article-title":"Significance analysis of microarrays applied to the ionizing radiation response","volume":"98","author":"Tusher","year":"2001","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023013112193337100_B36","article-title":"Multiple testing. Part III. Augmentation procedures for control of the generalized family-wise error rate and tail probabilities for the proportion of false positives","author":"van der Laan","year":"2004","journal-title":"U.C. Berkeley Division of Biostatistics Working Paper Series, Working Paper 141."},{"key":"2023013112193337100_B37","doi-asserted-by":"crossref","first-page":"1278","DOI":"10.1086\/522374","article-title":"Pathway based approaches for analysis of genome-wide association studies","volume":"81","author":"Wang","year":"2007","journal-title":"Am. J. Hum. Genet."},{"key":"2023013112193337100_B38","doi-asserted-by":"crossref","first-page":"1537","DOI":"10.1093\/bioinformatics\/btm129","article-title":"A Markov random field model for network-based analysis of genomic data","volume":"23","author":"Wei","year":"2007","journal-title":"Bioinformatics"},{"key":"2023013112193337100_B39","doi-asserted-by":"crossref","first-page":"408","DOI":"10.1214\/07--AOAS145","article-title":"A hidden spatial-temporal Markov random field model for network-based analysis of time course gene expression data","volume":"2","author":"Wei","year":"2008","journal-title":"Ann. Appl. Stat."},{"key":"2023013112193337100_B40","first-page":"364","article-title":"On false discovery control under dependence","volume":"36","author":"Wu","year":"2009","journal-title":"Ann. Stat."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/25\/21\/2802\/48998003\/bioinformatics_25_21_2802.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/25\/21\/2802\/48998003\/bioinformatics_25_21_2802.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,11]],"date-time":"2025-02-11T16:44:22Z","timestamp":1739292262000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/25\/21\/2802\/226040"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,5,4]]},"references-count":40,"journal-issue":{"issue":"21","published-print":{"date-parts":[[2009,11,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btp476","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2009,11,1]]},"published":{"date-parts":[[2009,5,4]]}}}