{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,20]],"date-time":"2026-01-20T14:11:50Z","timestamp":1768918310737,"version":"3.49.0"},"reference-count":51,"publisher":"Oxford University Press (OUP)","issue":"20","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2012,10,15]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: There is growing momentum to develop statistical learning (SL) methods as an alternative to conventional genome-wide association studies (GWAS). Methods such as random forests (RF) and gradient boosting machine (GBM) result in variable importance measures that indicate how well each single-nucleotide polymorphism (SNP) predicts the phenotype. For RF, it has been shown that variable importance measures are systematically affected by minor allele frequency (MAF) and linkage disequilibrium (LD). To establish RF and GBM as viable alternatives for analyzing genome-wide data, it is necessary to address this potential bias and show that SL methods do not significantly under-perform conventional GWAS methods.<\/jats:p><jats:p>Results: Both LD and MAF have a significant impact on the variable importance measures commonly used in RF and GBM. Dividing SNPs into overlapping subsets with approximate linkage equilibrium and applying SL methods to each subset successfully reduces the impact of LD. A welcome side effect of this approach is a dramatic reduction in parallel computing time, increasing the feasibility of applying SL methods to large datasets. The created subsets also facilitate a potential correction for the effect of MAF using pseudocovariates. Simulations using simulated SNPs embedded in empirical data\u2014assessing varying effect sizes, minor allele frequencies and LD patterns\u2014suggest that the sensitivity to detect effects is often improved by subsetting and does not significantly under-perform the Armitage trend test, even under ideal conditions for the trend test.<\/jats:p><jats:p>Availability: Code for the LD subsetting algorithm and pseudocovariate correction is available at http:\/\/www.nd.edu\/\u223cglubke\/code.html.<\/jats:p><jats:p>Contact: \u00a0glubke@nd.edu<\/jats:p><jats:p>Supplementary information: Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/bts483","type":"journal-article","created":{"date-parts":[[2012,7,31]],"date-time":"2012-07-31T06:55:55Z","timestamp":1343717755000},"page":"2615-2623","source":"Crossref","is-referenced-by-count":16,"title":["An integrated approach to reduce the impact of minor allele frequency and linkage disequilibrium on variable importance measures for genome-wide data"],"prefix":"10.1093","volume":"28","author":[{"given":"Raymond","family":"Walters","sequence":"first","affiliation":[{"name":"1 Department of Psychology, University of Notre Dame, Notre Dame, IN 46556, USA and 2Biological Psychology, VU University Amsterdam, Van der Boechorststraat 1, 1081 BT Amsterdam, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Charles","family":"Laurin","sequence":"additional","affiliation":[{"name":"1 Department of Psychology, University of Notre Dame, Notre Dame, IN 46556, USA and 2Biological Psychology, VU University Amsterdam, Van der Boechorststraat 1, 1081 BT Amsterdam, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gitta H.","family":"Lubke","sequence":"additional","affiliation":[{"name":"1 Department of Psychology, University of Notre Dame, Notre Dame, IN 46556, USA and 2Biological Psychology, VU University Amsterdam, Van der Boechorststraat 1, 1081 BT Amsterdam, The Netherlands"},{"name":"1 Department of Psychology, University of Notre Dame, Notre Dame, IN 46556, USA and 2Biological Psychology, VU University Amsterdam, Van der Boechorststraat 1, 1081 BT Amsterdam, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2012,7,30]]},"reference":[{"key":"2023012513140813900_bts483-B1","doi-asserted-by":"crossref","first-page":"375","DOI":"10.2307\/3001775","article-title":"Tests for linear trends in proportions and frequencies","volume":"11","author":"Armitage","year":"1955","journal-title":"Biometrics"},{"key":"2023012513140813900_bts483-B2","doi-asserted-by":"crossref","first-page":"879","DOI":"10.1002\/gepi.20543","article-title":"SNP selection in genome-wide and candidate gene studies via penalized logistic regression","volume":"34","author":"Ayers","year":"2010","journal-title":"Genet. Epidemiol."},{"key":"2023012513140813900_bts483-B3","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1093\/bioinformatics\/bth457","article-title":"Haploview: analysis and visualization of LD and haplotype maps","volume":"21","author":"Barrett","year":"2005","journal-title":"Bioinformatics"},{"key":"2023012513140813900_bts483-B4","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1016\/S0167-7152(02)00323-1","article-title":"A new bivariate binomial distribution","volume":"60","author":"Biswas","year":"2002","journal-title":"Stat. Probab. Lett."},{"key":"2023012513140813900_bts483-B5","doi-asserted-by":"crossref","first-page":"292","DOI":"10.1093\/bib\/bbr053","article-title":"Random forest Gini importance favours SNPs with large minor allele frequency: impact, sources and recommendations","volume":"13","author":"Boulesteix","year":"2011","journal-title":"Brief. Bioinform"},{"key":"2023012513140813900_bts483-B6","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1010933404324","article-title":"Random forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learn."},{"key":"2023012513140813900_B0","unstructured":"Breiman L (2002). Manual on setting up, using, and understanding random forests v3.1. http:\/\/oz.berkeley.edu\/users\/breiman\/Using_random_forests_V3.1.pdf"},{"key":"2023012513140813900_bts483-B7","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1002\/gepi.20041","article-title":"Identifying SNPs predictive of phenotype using random forests","volume":"28","author":"Bureau","year":"2005","journal-title":"Genet. Epidemiol."},{"key":"2023012513140813900_bts483-B8","doi-asserted-by":"crossref","DOI":"10.1145\/1143844.1143865","article-title":"An empirical comparison of supervised learning algorithms.","author":"Caruana","year":"2006"},{"key":"2023012513140813900_bts483-B9","doi-asserted-by":"crossref","first-page":"404","DOI":"10.1093\/biomet\/26.4.404","article-title":"The use of confidence or fiducial limits illustrated in the case of the binomial","volume":"26","author":"Clopper","year":"1934","journal-title":"Biometrika"},{"key":"2023012513140813900_bts483-B10","doi-asserted-by":"crossref","first-page":"241","DOI":"10.1038\/nrg2554","article-title":"Human genetic variation and its contribution to complex traits","volume":"10","author":"Frazer","year":"2009","journal-title":"Nat. Rev. Genet."},{"issue":"5","key":"2023012513140813900_bts483-B11","doi-asserted-by":"crossref","first-page":"1189","DOI":"10.1214\/aos\/1013203451","article-title":"Greedy function approximation: a gradient boosting machine","volume":"29","author":"Friedman","year":"2001","journal-title":"Ann. Statist."},{"issue":"200","key":"2023012513140813900_bts483-B12","doi-asserted-by":"crossref","first-page":"675","DOI":"10.1080\/01621459.1937.10503522","article-title":"The use of ranks to avoid the assumption of normality implicit in the analysis of variance","volume":"32","author":"Friedman","year":"1937","journal-title":"J. Am. Stat. Ass."},{"key":"2023012513140813900_bts483-B13","doi-asserted-by":"crossref","first-page":"360","DOI":"10.1111\/j.1469-1809.2009.00511.x","article-title":"Evaluating the ability of tree-based methods and logistic regression for the detection of SNP-SNP interaction","volume":"73","author":"Garcia-Magarinos","year":"2009","journal-title":"Ann. Hum. Genet."},{"key":"2023012513140813900_bts483-B14","doi-asserted-by":"crossref","first-page":"49","DOI":"10.1186\/1471-2156-11-49","article-title":"An application of random forests to a genome-wide association dataset: methodological considerations & new findings","volume":"11","author":"Goldstein","year":"2010","journal-title":"BMC Genet."},{"key":"2023012513140813900_bts483-B15","doi-asserted-by":"crossref","DOI":"10.1007\/978-0-387-84858-7","volume-title":"The Elements of Statistical Learning: Data Mining, Inference, and Prediction","author":"Hastie","year":"2009","edition":"2nd"},{"issue":"1","key":"2023012513140813900_bts483-B16","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1093\/bioinformatics\/btq600","article-title":"A variable selection method for genome-wide association studies","volume":"27","author":"He","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012513140813900_bts483-B17","doi-asserted-by":"crossref","first-page":"583","DOI":"10.1080\/01621459.1952.10483441","article-title":"Use of ranks in one-criterion variance analysis","volume":"47","author":"Kruskal","year":"1952","journal-title":"J. Am. Stat. Ass."},{"key":"2023012513140813900_bts483-B18","doi-asserted-by":"crossref","first-page":"516","DOI":"10.1093\/bioinformatics\/btq688","article-title":"The Bayesian lasso for genome-wide association studies","volume":"27","author":"Li","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012513140813900_bts483-B19","doi-asserted-by":"crossref","first-page":"i222","DOI":"10.1093\/bioinformatics\/btr227","article-title":"Detecting epistatic effects in association studies at a genomic level based on an ensemble approach","volume":"27","author":"Li","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012513140813900_bts483-B20","first-page":"18","article-title":"Classification and regression by randomForest","volume":"2","author":"Liaw","year":"2002","journal-title":"R News"},{"key":"2023012513140813900_bts483-B21","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1186\/1471-2156-5-32","article-title":"Screening large-scale association study data: exploiting interactions using random forests","volume":"5","author":"Lunetta","year":"2004","journal-title":"BMC Genet."},{"key":"2023012513140813900_bts483-B22","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1038\/456018a","article-title":"Personal genomes: the case of the missing heritability","volume":"456","author":"Maher","year":"2008","journal-title":"Nature"},{"key":"2023012513140813900_bts483-B23","doi-asserted-by":"crossref","first-page":"747","DOI":"10.1038\/nature08494","article-title":"Finding the missing heritability of complex diseases","volume":"461","author":"Manolio","year":"2009","journal-title":"Nature"},{"key":"2023012513140813900_bts483-B24","doi-asserted-by":"crossref","first-page":"356","DOI":"10.1038\/nrg2344","article-title":"Genome-wide association studies for complex traits: consensus, uncertainty and challenges","volume":"9","author":"McCarthy","year":"2008","journal-title":"Nat. Rev. Genet."},{"key":"2023012513140813900_bts483-B25","doi-asserted-by":"crossref","first-page":"750","DOI":"10.1016\/j.ajhg.2009.10.009","article-title":"Common variants in the Trichohyalin gene are associated with straight hair in Europeans","volume":"85","author":"Medland","year":"2009","journal-title":"Am. J. Hum. Genet."},{"key":"2023012513140813900_bts483-B26","doi-asserted-by":"crossref","first-page":"78","DOI":"10.1186\/1471-2105-10-78","article-title":"Performance of random forest when SNPs are in linkage disequilibrium","volume":"10","author":"Meng","year":"2009","journal-title":"BMC Bioinformatics"},{"key":"2023012513140813900_bts483-B27","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1159\/000073735","article-title":"The ubiquitous nature of epistasis in determining susceptibility to common human diseases","volume":"56","author":"Moore","year":"2003","journal-title":"Hum. Hered."},{"key":"2023012513140813900_bts483-B28","doi-asserted-by":"crossref","first-page":"1884","DOI":"10.1093\/bioinformatics\/btp331","article-title":"Predictor correlation impacts machine learning algorithms: implications for genomic studies","volume":"25","author":"Nicodemus","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012513140813900_bts483-B29","doi-asserted-by":"crossref","first-page":"71","DOI":"10.1186\/1471-2156-9-71","article-title":"Application of two machine learning algorithms to genetic association studies in the presence of covariates","volume":"9","author":"Nonyane","year":"2008","journal-title":"BMC Genet."},{"key":"2023012513140813900_bts483-B30","doi-asserted-by":"crossref","first-page":"S11","DOI":"10.1186\/1753-6561-5-S3-S11","article-title":"A comparison of random forests, boosting and support vector machines for genomic selection","volume":"5","author":"Ogutu","year":"2011","journal-title":"BMC Proc."},{"key":"2023012513140813900_bts483-B31","doi-asserted-by":"crossref","first-page":"570","DOI":"10.1038\/ng.610","article-title":"Estimation of effect size distribution from genome-wide association studies and implications for future discoveries","volume":"42","author":"Park","year":"2010","journal-title":"Nat. Genet."},{"key":"2023012513140813900_bts483-B32","doi-asserted-by":"crossref","first-page":"18026","DOI":"10.1073\/pnas.1114759108","article-title":"Distribution of allele frequencies and effect sizes and their interrelationships for common genetic susceptibility variants","volume":"108","author":"Park","year":"2011","journal-title":"Proc. Natl Acad. Sci. U.S.A."},{"key":"2023012513140813900_bts483-B33","volume-title":"R: A Language and Environment for Statistical Computing","author":"R Development Core Team","year":"2011"},{"key":"2023012513140813900_bts483-B34","volume-title":"GBM: Generalized boosted regression models","author":"Ridgeway","year":"2010"},{"key":"2023012513140813900_bts483-B35","doi-asserted-by":"crossref","first-page":"138","DOI":"10.1086\/321276","article-title":"Multifactor-dimensionality reduction reveals high-order interactions among estrogen-metabolism genes in sporadic breast cancer","volume":"69","author":"Ritchie","year":"2001","journal-title":"Am. J. Hum. Genet."},{"key":"2023012513140813900_bts483-B36","doi-asserted-by":"crossref","first-page":"e62","DOI":"10.1093\/nar\/gkr064","article-title":"Ranking causal variants and associated regions in genome-wide association studies by the support vector machine and random forest","volume":"39","author":"Roshan","year":"2011","journal-title":"Nucleic Acids Res."},{"key":"2023012513140813900_bts483-B37","doi-asserted-by":"crossref","first-page":"611","DOI":"10.1198\/106186008X344522","article-title":"A bias correction algorithm for the Gini variable importance measure in classification trees","volume":"17","author":"Sandri","year":"2008","journal-title":"J. Comp. Graph. Stat."},{"key":"2023012513140813900_bts483-B38","doi-asserted-by":"crossref","first-page":"393","DOI":"10.1007\/s11222-009-9132-0","article-title":"Analysis and correction of bias in total decrease in node impurity measures for tree-based algorithms","volume":"20","author":"Sandri","year":"2010","journal-title":"Stat Comput"},{"key":"2023012513140813900_bts483-B39","doi-asserted-by":"crossref","first-page":"1353","DOI":"10.1093\/bioinformatics\/bts163","article-title":"Matrix eQTL: ultra fast eQTL analysis via large matrix operations","volume":"28","author":"Shabalin","year":"2012","journal-title":"Bioinformatics"},{"key":"2023012513140813900_bts483-B40","first-page":"447","article-title":"Uncovering the total heritability explained by all true susceptibility variants in a genome-wide association study","volume":"35","author":"So","year":"2011","journal-title":"Genet. Epidemiol."},{"key":"2023012513140813900_bts483-B41","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1186\/1471-2105-8-25","article-title":"Bias in random forest variable importance measures: illustrations, sources and a solution","volume":"8","author":"Strobl","year":"2007","journal-title":"BMC Bioinform."},{"key":"2023012513140813900_bts483-B42","doi-asserted-by":"crossref","first-page":"307","DOI":"10.1186\/1471-2105-9-307","article-title":"Conditional variable importance for random forests","volume":"9","author":"Strobl","year":"2008","journal-title":"BMC Bioinform."},{"key":"2023012513140813900_bts483-B43","doi-asserted-by":"crossref","first-page":"S51","DOI":"10.1002\/gepi.20473","article-title":"Machine learning in genome-wide association studies","volume":"33","author":"Szymczak","year":"2009","journal-title":"Genet. Epidemiol."},{"key":"2023012513140813900_bts483-B44","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1111\/j.2517-6161.1996.tb02080.x","article-title":"Regression shrinkage and selection via the lasso","volume":"58","author":"Tibshirani","year":"1996","journal-title":"J. R. Stat. Soc. Ser. B."},{"key":"2023012513140813900_bts483-B45","doi-asserted-by":"crossref","first-page":"S69","DOI":"10.1186\/1753-6561-3-S7-S69","article-title":"Detecting significant single-nucleotide polymorphisms in a rheumatoid arthritis study using random forests","volume":"3","author":"Wang","year":"2009","journal-title":"BMC Proc."},{"key":"2023012513140813900_bts483-B46","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1038\/nrg1522","article-title":"Genome-wide association studies: theoretical and practical concerns","volume":"6","author":"Wang","year":"2005","journal-title":"Nat. Rev. Genet."},{"key":"2023012513140813900_bts483-B47","doi-asserted-by":"crossref","first-page":"2936","DOI":"10.1093\/bioinformatics\/btr512","article-title":"An empirical comparison of several recent epistatic interaction detection methods","volume":"27","author":"Wang","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012513140813900_bts483-B48","doi-asserted-by":"crossref","first-page":"275","DOI":"10.1002\/gepi.20459","article-title":"Screen and clean: a tool for identifying interactions in genome-wide association studies","volume":"34","author":"Wu","year":"2010","journal-title":"Genet. Epidemiol."},{"key":"2023012513140813900_bts483-B49","doi-asserted-by":"crossref","first-page":"565","DOI":"10.1038\/ng.608","article-title":"Common SNPs explain a large proportion of the heritability for human height","volume":"42","author":"Yang","year":"2010","journal-title":"Nat. Genet."},{"key":"2023012513140813900_bts483-B50","doi-asserted-by":"crossref","DOI":"10.1002\/9783527633654","volume-title":"A Statistical Approach to Genetic Epidemiology: Concepts and Applications","author":"Ziegler","year":"2010","edition":"2nd"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/28\/20\/2615\/48874596\/bioinformatics_28_20_2615.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/28\/20\/2615\/48874596\/bioinformatics_28_20_2615.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,4,27]],"date-time":"2024-04-27T23:34:13Z","timestamp":1714260853000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/28\/20\/2615\/201922"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,7,30]]},"references-count":51,"journal-issue":{"issue":"20","published-print":{"date-parts":[[2012,10,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bts483","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2012,10,15]]},"published":{"date-parts":[[2012,7,30]]}}}