{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T09:19:32Z","timestamp":1778059172495,"version":"3.51.4"},"reference-count":28,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T00:00:00Z","timestamp":1777680000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T00:00:00Z","timestamp":1778025600000},"content-version":"vor","delay-in-days":4,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"German Science Foundation","award":["DFG-Einzelf\u00f6rderung HO 6422\/1-3"],"award-info":[{"award-number":["DFG-Einzelf\u00f6rderung HO 6422\/1-3"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>\n                      <jats:bold>Background<\/jats:bold>\n                    <\/jats:title>\n                    <jats:p>Identifying relevant biomarkers is critical in clinical research and precision medicine, particularly when analysing high-dimensional data. Random forests (RFs) are promising for such settings due to their flexibility, ease of use, and their ability to handle data sets with more variables than samples. RFs assess the importance of each variable in predicting the outcome using variable importance (VIMP) scores. However, since the distribution of VIMP scores is intricate, standard statistical testing and multiple testing adjustments for variable selection are challenging.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>\n                      <jats:bold>Methods<\/jats:bold>\n                    <\/jats:title>\n                    <jats:p>\n                      We propose shadowVIMP, a novel method for multiple testing-controlled variable selection, based on an approach similar to permutation testing. It generates permuted counterparts for each variable and compares their VIMPs with those of the original variables over multiple iterations to calculate\n                      <jats:italic>p<\/jats:italic>\n                      -values. Unlike conventional permutation testing, shadowVIMP preserves the correlation structure between variables, mitigating biases caused by the over-selection of correlated variables in RFs. We evaluated shadowVIMP against three competing RF variable selection approaches using simulation designs previously employed in studies considering VIMPs and variable selection for RFs. These designs included high- and low-dimensional data, as well as correlated and categorical variables. For illustration, we also applied the method to a real-world example on Alzheimer\u2019s disease.\n                    <\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>\n                      <jats:bold>Conclusions<\/jats:bold>\n                    <\/jats:title>\n                    <jats:p>Our results showed that, compared to competing approaches, shadowVIMP offers advantages in high-dimensional settings, improving sensitivity while enabling multiple testing-adjusted results. Additionally, it demonstrated robustness against VIMP biases induced by correlated and categorical variables when using permutation-based VIMP. The method can be used to annotate standard VIMP plots, visually presenting selected variable sets based on different types of multiple testing adjustments and significance levels. Overall, shadowVIMP is a promising approach for providing multiple testing-adjusted variable selection while explicitly addressing known biases of RF\u2019s permutation-based VIMP measure. The shadowVIMP method is implemented in an R package , which is available on CRAN.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1186\/s12859-026-06412-4","type":"journal-article","created":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T07:06:57Z","timestamp":1777878417000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["ShadowVIMP: permutation-based multiple testing-controlled variable selection"],"prefix":"10.1186","volume":"27","author":[{"given":"Tim","family":"M\u00fcller","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Roman","family":"Hornung","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Silke","family":"Szymczak","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hannes","family":"Buchner","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,5,2]]},"reference":[{"key":"6412_CR1","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman L. Random forests. Mach Learn. 2001;45:5\u201332.","journal-title":"Mach Learn"},{"key":"6412_CR2","doi-asserted-by":"publisher","first-page":"885","DOI":"10.1007\/s11634-016-0276-4","volume":"12","author":"S Janitza","year":"2018","unstructured":"Janitza S, Celik E, Boulesteix A-L. A computationally fast variable importance test for random forests for high-dimensional data. Adv Data Anal Classif. 2018;12:885\u2013915.","journal-title":"Adv Data Anal Classif"},{"issue":"1","key":"6412_CR3","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/1471-2105-8-25","volume":"8","author":"C Strobl","year":"2007","unstructured":"Strobl C, et al. Bias in random forest variable importance measures: illustrations, sources and a solution. BMC Bioinf. 2007;8(1):1\u201321.","journal-title":"BMC Bioinf"},{"issue":"2","key":"6412_CR4","doi-asserted-by":"publisher","first-page":"492","DOI":"10.1093\/bib\/bbx124","volume":"20","author":"F Degenhardt","year":"2019","unstructured":"Degenhardt F, Seifert S, Szymczak S. Evaluation of variable selection methods for random forests and omics data sets. Br Bioinform. 2019;20(2):492\u2013503.","journal-title":"Br Bioinform"},{"key":"6412_CR5","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s12859-018-2264-5","volume":"19","author":"R Couronn\u00e9","year":"2018","unstructured":"Couronn\u00e9 R, Probst P, Boulesteix A-L. Random forest versus logistic regression: a large-scale benchmark experiment. BMC Bioinf. 2018;19:1\u201314.","journal-title":"BMC Bioinf"},{"issue":"1","key":"6412_CR6","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18637\/jss.v077.i01","volume":"77","author":"MN Wright","year":"2017","unstructured":"Wright MN, Ziegler A. Ranger: a fast implementation of random forests for high dimensional data in C++ and R. J Stat Softw. 2017;77(1):1\u201317. https:\/\/doi.org\/10.18637\/jss.v077.i01.","journal-title":"J Stat Softw"},{"key":"6412_CR7","unstructured":"Boulesteix A-L. Diagnostic Accuracy Studies: Lecture C1: Introduction to prediction modelling. 2022."},{"issue":"4","key":"6412_CR8","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0018850","volume":"6","author":"R Craig-Schapiro","year":"2011","unstructured":"Craig-Schapiro R, et al. Multiplexed immunoassay panel identifies novel CSF biomarkers for Alzheimer\u2019s disease diagnosis and prognosis. PLoS ONE. 2011;6(4):e18850.","journal-title":"PLoS ONE"},{"key":"6412_CR9","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4614-6849-3","volume-title":"Applied predictive modeling","author":"M Kuhn","year":"2013","unstructured":"Kuhn M, Johnson K, et al. Applied predictive modeling, vol. 26. Berlin: Springer; 2013."},{"key":"6412_CR10","unstructured":"Breiman L, Random forests. http:\/\/www.stat.berkeley.edu\/users\/breiman\/RandomForests\/cc_home.htm. 2008."},{"key":"6412_CR11","doi-asserted-by":"crossref","unstructured":"Strobl C, Zeileis A. Danger: high power!-exploring the statistical properties of a test for random forest variable importance. In: BMC Bioinformatics 2008;9","DOI":"10.1186\/1471-2105-9-307"},{"key":"6412_CR12","doi-asserted-by":"publisher","first-page":"50","DOI":"10.1016\/j.csda.2012.09.020","volume":"60","author":"A Hapfelmeier","year":"2013","unstructured":"Hapfelmeier A, Ulm K. A new variable selection approach using random forests. Comput Stat Data Anal. 2013;60:50\u201369.","journal-title":"Comput Stat Data Anal"},{"key":"6412_CR13","doi-asserted-by":"publisher","DOI":"10.1016\/j.csda.2022.107689","volume":"181","author":"A Hapfelmeier","year":"2023","unstructured":"Hapfelmeier A, Hornung R, Haller B. Efficient permutation testing of variable importance measures by the example of random forests. Comput Stat Data Anal. 2023;181:107689.","journal-title":"Comput Stat Data Anal"},{"key":"6412_CR14","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18637\/jss.v036.i11","volume":"36","author":"MB Kursa","year":"2010","unstructured":"Kursa MB, Rudnicki WR. Feature selection with the Boruta package. In J Stat Softw. 2010;36:1\u201313.","journal-title":"In J Stat Softw"},{"key":"6412_CR15","doi-asserted-by":"publisher","first-page":"93","DOI":"10.1016\/j.eswa.2019.05.028","volume":"134","author":"JL Speiser","year":"2019","unstructured":"Speiser JL, et al. A comparison of random forest variable selection methods for classification prediction modeling. Expert Syst Appl. 2019;134:93\u2013101.","journal-title":"Expert Syst Appl"},{"key":"6412_CR16","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/1471-2105-11-110","volume":"11","author":"KK Nicodemus","year":"2010","unstructured":"Nicodemus KK, et al. The behaviour of random forest permutation-based variable importance measures under predictor correlation. BMC Bioinf. 2010;11:1\u201313.","journal-title":"BMC Bioinf"},{"key":"6412_CR17","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-20933-9","volume-title":"A primer of permutation statistical methods","author":"KJ Berry","year":"2019","unstructured":"Berry KJ, Johnston JE, Mielke PW. A primer of permutation statistical methods. Berlin: Springer; 2019."},{"issue":"4","key":"6412_CR18","doi-asserted-by":"publisher","first-page":"323","DOI":"10.1037\/a0016973","volume":"14","author":"C Strobl","year":"2009","unstructured":"Strobl C, Malley J, Tutz G. An introduction to recursive partitioning: rationale, application, and characteristics of classification and regression trees, bagging, and random forests. Psychol Methods. 2009;14(4):323.","journal-title":"Psychol Methods"},{"key":"6412_CR19","unstructured":"Hapfelmeier A, Hornung R. rfvimptest: sequential permutation testing of random forest variable importance measures. R package version 0.1.4. 2025. https:\/\/CRAN.R-project.org\/package=rfvimptest."},{"key":"6412_CR20","doi-asserted-by":"crossref","unstructured":"Phipson B, Smyth GK. Permutation P-values should never be zero: calculating exact P-values when permutations are randomly drawn. In: statistical applications in genetics and molecular biology 9(1) (2010).","DOI":"10.2202\/1544-6115.1585"},{"issue":"3","key":"6412_CR21","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1002956","volume":"9","author":"Z Chen","year":"2013","unstructured":"Chen Z, Zhang W. Integrative analysis using module-guided random forests reveals correlated genetic factors related to mouse weight. PLoS Comput Biol. 2013;9(3):e1002956.","journal-title":"PLoS Comput Biol"},{"key":"6412_CR22","unstructured":"Szymczak S. Github: Pomona (R). https:\/\/github.com\/silkeszy\/Pomona. Accessed: 12\u201305-2023. 2023."},{"issue":"1","key":"6412_CR23","first-page":"1","volume":"19","author":"JH Friedman","year":"1991","unstructured":"Friedman JH. Multivariate adaptive regression splines. Ann Stat. 1991;19(1):1\u201367.","journal-title":"Ann Stat"},{"key":"6412_CR24","unstructured":"Leisch F, Dimitriadou E. mlbench: Machine learning benchmark problems. R package version 2.1-3. 2021."},{"key":"6412_CR25","unstructured":"M\u00fcller T, Hornung R, Miluch O. shadowVIMP-publication: electronic appendix. https:\/\/github.com\/mueller-staburo\/shadowVIMP-Publication. 2025."},{"issue":"4","key":"6412_CR26","doi-asserted-by":"publisher","first-page":"649","DOI":"10.1111\/rssb.12274","volume":"80","author":"L Lei","year":"2018","unstructured":"Lei L, Fithian W. AdaPT: an interactive procedure for multiple testing with side information. J R Stat Soc Ser B Stat Methodol. 2018;80(4):649\u201379.","journal-title":"J R Stat Soc Ser B Stat Methodol"},{"issue":"27","key":"6412_CR27","doi-asserted-by":"publisher","first-page":"5039","DOI":"10.1002\/sim.9900","volume":"42","author":"Y Luo","year":"2023","unstructured":"Luo Y, Guo X. Inference on tree-structured subgroups with subgroup size and subgroup effect relationship in clinical trials. Stat Med. 2023;42(27):5039\u201353.","journal-title":"Stat Med"},{"key":"6412_CR28","doi-asserted-by":"crossref","unstructured":"Mueller T, Miluch O. shadowVIMP: covariate selection based on VIMP permutation - like testing. R package version 1.0.2. 2025. 10.32614\/CRAN.package.shadowVIMP. https:\/\/CRAN.R-project.org\/package=shadowVIMP.","DOI":"10.32614\/CRAN.package.shadowVIMP"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-026-06412-4","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-026-06412-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-026-06412-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T08:47:41Z","timestamp":1778057261000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1186\/s12859-026-06412-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,2]]},"references-count":28,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,12]]}},"alternative-id":["6412"],"URL":"https:\/\/doi.org\/10.1186\/s12859-026-06412-4","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,2]]},"assertion":[{"value":"7 August 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 February 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 May 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"96"}}