{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T15:29:58Z","timestamp":1784820598177,"version":"3.55.0"},"reference-count":35,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2022,6,20]],"date-time":"2022-06-20T00:00:00Z","timestamp":1655683200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100002428","name":"Austrian Science Fund (FWF)","doi-asserted-by":"publisher","award":["I4739-B"],"award-info":[{"award-number":["I4739-B"]}],"id":[{"id":"10.13039\/501100002428","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>There is an increasing interest in machine learning (ML) algorithms for predicting patient outcomes, as these methods are designed to automatically discover complex data patterns. For example, the random forest (RF) algorithm is designed to identify relevant predictor variables out of a large set of candidates. In addition, researchers may also use external information for variable selection to improve model interpretability and variable selection accuracy, thereby prediction quality. However, it is unclear to which extent, if at all, RF and ML methods may benefit from external information. In this paper, we examine the usefulness of external information from prior variable selection studies that used traditional statistical modeling approaches such as the Lasso, or suboptimal methods such as univariate selection. We conducted a plasmode simulation study based on subsampling a data set from a pharmacoepidemiologic study with nearly 200,000 individuals, two binary outcomes and 1152 candidate predictor (mainly sparse binary) variables. When the scope of candidate predictors was reduced based on external knowledge RF models achieved better calibration, that is, better agreement of predictions and observed outcome rates. However, prediction quality measured by cross-entropy, AUROC or the Brier score did not improve. We recommend appraising the methodological quality of studies that serve as an external information source for future prediction model development.<\/jats:p>","DOI":"10.3390\/e24060847","type":"journal-article","created":{"date-parts":[[2022,6,21]],"date-time":"2022-06-21T01:43:27Z","timestamp":1655775807000},"page":"847","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Using Background Knowledge from Preceding Studies for Building a Random Forest Prediction Model: A Plasmode Simulation Study"],"prefix":"10.3390","volume":"24","author":[{"given":"Lorena","family":"Hafermann","sequence":"first","affiliation":[{"name":"Institute of Biometry and Clinical Epidemiology, Charit\u00e9\u2013Universit\u00e4tsmedizin Berlin, Corporate Member of Freie Universit\u00e4t Berlin and Humboldt-Universit\u00e4t zu Berlin, Charit\u00e9platz 1, 10117 Berlin, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5072-5347","authenticated-orcid":false,"given":"Nadja","family":"Klein","sequence":"additional","affiliation":[{"name":"Chair of Statistics and Data Science, School of Business and Economics, Humboldt-Universit\u00e4t zu Berlin, Unter den Linden 6, 10099 Berlin, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Geraldine","family":"Rauch","sequence":"additional","affiliation":[{"name":"Institute of Biometry and Clinical Epidemiology, Charit\u00e9\u2013Universit\u00e4tsmedizin Berlin, Corporate Member of Freie Universit\u00e4t Berlin and Humboldt-Universit\u00e4t zu Berlin, Charit\u00e9platz 1, 10117 Berlin, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4821-9928","authenticated-orcid":false,"given":"Michael","family":"Kammer","sequence":"additional","affiliation":[{"name":"Section for Clinical Biometrics, Center for Medical Statistics, Informatics and Intelligent Systems, Medical University of Vienna, Spitalgasse 23, 1090 Vienna, Austria"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1147-8491","authenticated-orcid":false,"given":"Georg","family":"Heinze","sequence":"additional","affiliation":[{"name":"Section for Clinical Biometrics, Center for Medical Statistics, Informatics and Intelligent Systems, Medical University of Vienna, Spitalgasse 23, 1090 Vienna, Austria"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,6,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1080\/09332480.2019.1579578","article-title":"A second chance to get causal inference right: A classification of data science tasks","volume":"32","author":"Hsu","year":"2019","journal-title":"Chance"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1214\/10-STS330","article-title":"To Explain or to Predict?","volume":"25","author":"Shmueli","year":"2010","journal-title":"Stat. Sci."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"e5595","DOI":"10.1136\/bmj.e5595","article-title":"Prognosis research strategy (PROGRESS) 1: A framework for researching clinical outcomes","volume":"346","author":"Hemingway","year":"2013","journal-title":"BMJ"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1214\/ss\/1009213726","article-title":"Statistical modelling: The two cultures","volume":"16","author":"Breiman","year":"2001","journal-title":"Stat. Sci."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Hastie, T., Tibshirani, R., and Friedman, J. (2009). The Elements of Statistical Learning, Springer. [2nd ed.].","DOI":"10.1007\/978-0-387-84858-7"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1111\/j.2517-6161.1996.tb02080.x","article-title":"Regression Shrinkage and Selection via the Lasso","volume":"58","author":"Tibshirani","year":"1996","journal-title":"J. R. Stat. Soc. Ser. B"},{"key":"ref_7","first-page":"559","article-title":"Boosting for high-dimensional linear models","volume":"34","year":"2006","journal-title":"Ann. Stat."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Harrell, F.E. (2015). Regression Modelling Strategies, Springer. [2nd ed.].","DOI":"10.1007\/978-3-319-19425-7"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Royston, P., and Sauerbrei, W. (2008). Multivariable Model-Building: A Pragmatic Approach to Regression Analysis Based on Fractional Polynomials for Modelling Continuous Variables, John Wiley & Sons Ltd.. [1st ed.].","DOI":"10.1002\/9780470770771"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"431","DOI":"10.1002\/bimj.201700067","article-title":"Variable selection\u2014A review and recommendation for the practicing statistician","volume":"60","author":"Heinze","year":"2018","journal-title":"Biom. J."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1186\/s41512-020-00074-3","article-title":"State of the art in selection of variables and functional forms in multivariable analysis\u2014Outstanding issues","volume":"4","author":"Sauerbrei","year":"2020","journal-title":"Diagn. Progn. Res."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"van der Ploeg, T., Austin, P.C., and Steyerberg, E.W. (2014). Modern modelling techniques are data hungry: A simulation study for predicting dichotomous endpoints. BMC Med. Res. Methodol., 14.","DOI":"10.1186\/1471-2288-14-137"},{"key":"ref_13","first-page":"1","article-title":"Weighted Lasso with Data Integration","volume":"10","author":"Bergerson","year":"2011","journal-title":"Stat. Appl. Genet. Mol. Biol."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1","DOI":"10.18637\/jss.v077.i01","article-title":"Ranger: A Fast Implementation of Random Forests for High Dimensional Data in C++ and R","volume":"77","author":"Wright","year":"2017","journal-title":"J. Stat. Softw."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"907","DOI":"10.1016\/0895-4356(96)00025-X","article-title":"Inappropriate Use of Bivariable Analysis to Screen Risk Factors for Use in Multivariable Analysis","volume":"49","author":"Sun","year":"1996","journal-title":"J. Clin. Epidemiol."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"6","DOI":"10.1111\/tri.12895","article-title":"Five myths about variable selection","volume":"30","author":"Heinze","year":"2017","journal-title":"Transpl. Int."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1010933404324","article-title":"Random forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learn."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"74","DOI":"10.3414\/ME00-01-0052","article-title":"Probability machines: Consistent probability estimation using nonparametric learning machines","volume":"51","author":"Malley","year":"2012","journal-title":"Methods Inf. Med."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Strobl, C., Boulesteix, A.L., Zeileis, A., and Hothorn, T. (2007). Bias in random forest variable importance measures: Illustrations, sources and a solution. BMC Bioinform., 8.","DOI":"10.1186\/1471-2105-8-25"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"128","DOI":"10.1097\/EDE.0b013e3181c30fb2","article-title":"Assessing the performance of prediction models: A framework for traditional and novel measures","volume":"21","author":"Steyerberg","year":"2010","journal-title":"Epidemiology"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"5964","DOI":"10.1038\/s41598-020-62318-y","article-title":"Comparative effectiveness of branded vs. generic versions of antihypertensive, lipid-lowering and hypoglycemic substances: A population-wide cohort study","volume":"10","author":"Tian","year":"2020","journal-title":"Sci. Rep."},{"key":"ref_22","unstructured":"WHO Collaborating Centre for Drug Statistics Methodology (2011). Guidelines for ATC Classification and DDD Assignment 2012, Norwegian Institute of Public Health."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"2074","DOI":"10.1002\/sim.8086","article-title":"Using simulation studies to evaluate statistical methods","volume":"38","author":"Morris","year":"2019","journal-title":"Stat. Med."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1","DOI":"10.18637\/jss.v033.i01","article-title":"Regularization Paths for Generalized Linear Models via Coordinate Descent","volume":"33","author":"Friedman","year":"2010","journal-title":"J. Stat. Softw."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"3166","DOI":"10.1177\/0962280220921415","article-title":"Regression shrinkage methods for clinical prediction models do not guarantee improved performance: Simulation study","volume":"29","author":"Steyerberg","year":"2020","journal-title":"Stat. Methods Med. Res."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Van Calster, B., McLernon, D.J., van Smeden, M., Wynants, L., Steyerberg, E.W., and on behalf of Topic Group \u2018Evaluating Diagnostic Tests and Prediction Models\u2019 of the STRATOS Initiative (2019). Calibration: The Achilles heel of predictive analytics. BMC Med., 17.","DOI":"10.1186\/s12916-019-1466-7"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"m1328","DOI":"10.1136\/bmj.m1328","article-title":"Prediction models for diagnosis and prognosis of covid-19: Systematic review and critical appraisal","volume":"369","author":"Wynants","year":"2020","journal-title":"BMJ"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"126","DOI":"10.1016\/j.jclinepi.2022.01.025","article-title":"Prediction models for living organ transplantation are poorly developed, reported, and validated: A systematic review","volume":"145","author":"Haller","year":"2022","journal-title":"J. Clin. Epidemiol."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"g7594","DOI":"10.1136\/bmj.g7594","article-title":"Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): The TRIPOD statement","volume":"350","author":"Collins","year":"2015","journal-title":"BMJ"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"W1","DOI":"10.7326\/M18-1377","article-title":"PROBAST: A tool to assess risk of bias and applicability of prediction model studies: Explanation and elaboration","volume":"170","author":"Moons","year":"2019","journal-title":"Ann. Intern. Med."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Hafermann, L., Becher, H., Herrmann, C., Klein, N., Heinze, G., and Rauch, G. (2021). Statistical model building: Background \u201cknowledge\u201d based on inappropriate preselection causes misspecification. BMC Med. Res. Methodol, 21.","DOI":"10.1186\/s12874-021-01373-z"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"301","DOI":"10.1111\/j.1467-9868.2005.00503.x","article-title":"Regularization and Variable Selection via the Elastic Net","volume":"67","author":"Zou","year":"2005","journal-title":"J. R. Stat. Soc. B."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"193","DOI":"10.1007\/s40258-014-0143-4","article-title":"Potential Savings in Prescription Drug Costs for Hypertension, Hyperlipidemia, and Diabetes Mellitus by Equivalent Drug Substitution in Austria: A Nationwide Cohort Study","volume":"13","author":"Heinze","year":"2014","journal-title":"Appl. Health Econ. Health Policy"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"90","DOI":"10.1002\/pds.3898","article-title":"Prevalence and determinants of unintended double medication of antihypertensive, lipid-lowering, and hypoglycemic drugs in Austria: A nationwide cohort study","volume":"25","author":"Heinze","year":"2015","journal-title":"Pharmacoepidemiol. Drug Saf."},{"key":"ref_35","unstructured":"Jandeck, L.M. (2014). Populationsweite Utilisationsuntersuchung in den chronischen Krankheitsbildern Hypertonie, Hyperlipid\u00e4mie und Typ 2 Diabetes Mellitus. [Inaugural Dissertation, Ruhr-Universit\u00e4t Bochum]."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/6\/847\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:35:32Z","timestamp":1760139332000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/6\/847"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,20]]},"references-count":35,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2022,6]]}},"alternative-id":["e24060847"],"URL":"https:\/\/doi.org\/10.3390\/e24060847","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,20]]}}}