{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T16:14:13Z","timestamp":1785946453601,"version":"3.56.0"},"reference-count":30,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2008,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>Cancer diagnosis and clinical outcome prediction are among the most important emerging applications of gene expression microarray technology with several molecular signatures on their way toward clinical deployment. Use of the most accurate classification algorithms available for microarray gene expression data is a critical ingredient in order to develop the best possible molecular signatures for patient care. As suggested by a large body of literature to date, support vector machines can be considered \"best of class\" algorithms for classification of such data. Recent work, however, suggests that random forest classifiers may outperform support vector machines in this domain.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>In the present paper we identify methodological biases of prior work comparing random forests and support vector machines and conduct a new rigorous evaluation of the two algorithms that corrects these limitations. Our experiments use 22 diagnostic and prognostic datasets and show that support vector machines outperform random forests, often by a large margin. Our data also underlines the importance of sound research design in benchmarking and comparison of bioinformatics algorithms.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>We found that both on average and in the majority of microarray datasets, random forests are outperformed by support vector machines both in the settings when no gene selection is performed and when several popular gene selection methods are used.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-9-319","type":"journal-article","created":{"date-parts":[[2008,7,22]],"date-time":"2008-07-22T18:14:07Z","timestamp":1216750447000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":541,"title":["A comprehensive comparison of random forests and support vector machines for microarray-based cancer classification"],"prefix":"10.1186","volume":"9","author":[{"given":"Alexander","family":"Statnikov","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lily","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Constantin F","family":"Aliferis","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2008,7,22]]},"reference":[{"key":"2304_CR1","doi-asserted-by":"publisher","first-page":"631","DOI":"10.1093\/bioinformatics\/bti033","volume":"21","author":"A Statnikov","year":"2005","unstructured":"Statnikov A, Aliferis CF, Tsamardinos I, Hardin D, Levy S: A comprehensive evaluation of multicategory classification methods for microarray gene expression cancer diagnosis. Bioinformatics 2005, 21: 631\u2013643.","journal-title":"Bioinformatics"},{"key":"2304_CR2","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman L: Random forests. Machine Learning 2001, 45: 5\u201332.","journal-title":"Machine Learning"},{"key":"2304_CR3","doi-asserted-by":"publisher","first-page":"1636","DOI":"10.1093\/bioinformatics\/btg210","volume":"19","author":"B Wu","year":"2003","unstructured":"Wu B, Abbott T, Fishman D, McMurray W, Mor G, Stone K, Ward D, Williams K, Zhao H: Comparison of statistical methods for classification of ovarian cancer using mass spectrometry data. Bioinformatics 2003, 19: 1636\u20131643.","journal-title":"Bioinformatics"},{"key":"2304_CR4","doi-asserted-by":"publisher","first-page":"869","DOI":"10.1016\/j.csda.2004.03.017","volume":"48","author":"JW Lee","year":"2005","unstructured":"Lee JW, Lee JB, Park M, Song SH: An extensive comparison of recent classification tools applied to microarray data. Computational Statistics & Data Analysis 2005, 48: 869\u2013885.","journal-title":"Computational Statistics & Data Analysis"},{"key":"2304_CR5","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1186\/1471-2105-7-3","volume":"7","author":"R Diaz-Uriarte","year":"2006","unstructured":"Diaz-Uriarte R, Alvarez de Andres S: Gene selection and classification of microarray data using random forest. BMC Bioinformatics 2006, 7: 3.","journal-title":"BMC Bioinformatics"},{"key":"2304_CR6","doi-asserted-by":"publisher","first-page":"706","DOI":"10.1137\/S0036144502411986","volume":"45","author":"R Rifkin","year":"2003","unstructured":"Rifkin R, Mukherjee S, Tamayo P, Ramaswamy S, Yeang CH, Angelo M, Reich M, Poggio T, Lander ES, Golub TR, Mesirov JP: An analytical method for multi-class molecular cancer classification. SIAM Reviews 2003, 45: 706\u2013723.","journal-title":"SIAM Reviews"},{"key":"2304_CR7","volume-title":"Proceedings of the First SIAM International Conference on Data Mining","author":"V Sindhwani","year":"2001","unstructured":"Sindhwani V, Bhattacharyya P, Rakshit S: Information Theoretic Feature Crediting in Multiclass Support Vector Machines. Proceedings of the First SIAM International Conference on Data Mining 2001."},{"key":"2304_CR8","doi-asserted-by":"publisher","first-page":"361","DOI":"10.1002\/(SICI)1097-0258(19960229)15:4<361::AID-SIM168>3.0.CO;2-4","volume":"15","author":"FE Harrell Jr.","year":"1996","unstructured":"Harrell FE Jr., Lee KL, Mark DB: Multivariable prognostic models: issues in developing models, evaluating assumptions and adequacy, and measuring and reducing errors. Stat Med 1996, 15: 361\u2013387.","journal-title":"Stat Med"},{"key":"2304_CR9","volume-title":"Proceedings of the Eighteenth International Joint Conference of Artificial Intelligence (IJCAI)","author":"CX Ling","year":"2003","unstructured":"Ling CX, Huang J, Zhang H: AUC: a statistically consistent and more discriminating measure than accuracy. Proceedings of the Eighteenth International Joint Conference of Artificial Intelligence (IJCAI) 2003."},{"key":"2304_CR10","volume-title":"Technical Report, HPL-2003-4, HP Laboratories","author":"T Fawcett","year":"2003","unstructured":"Fawcett T: ROC Graphs: Notes and Practical Considerations for Researchers. Technical Report, HPL-2003\u20134, HP Laboratories 2003."},{"key":"2304_CR11","first-page":"548","volume":"92","author":"B Efron","year":"1997","unstructured":"Efron B, Tibshirani R: Improvements on cross-validation: the .632+ bootstrap method. Journal of the American Statistical Association 1997, 92: 548\u2013560.","journal-title":"Journal of the American Statistical Association"},{"key":"2304_CR12","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-21606-5","volume-title":"The elements of statistical learning: data mining, inference, and prediction","author":"T Hastie","year":"2001","unstructured":"Hastie T, Tibshirani R, Friedman JH Springer series in statistics. In The elements of statistical learning: data mining, inference, and prediction. New York, Springer; 2001."},{"key":"2304_CR13","doi-asserted-by":"publisher","first-page":"278","DOI":"10.1186\/1471-2164-7-278","volume":"7","author":"AM Glas","year":"2006","unstructured":"Glas AM, Floore A, Delahaye LJ, Witteveen AT, Pover RC, Bakx N, Lahti-Domenici JS, Bruinsma TJ, Warmoes MO, Bernards R, Wessels LF, van't Veer LJ: Converting a breast cancer microarray signature into a high-throughput diagnostic test. BMC Genomics 2006, 7: 278.","journal-title":"BMC Genomics"},{"key":"2304_CR14","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1023\/A:1022936519097","volume":"17","author":"B Hammer","year":"2003","unstructured":"Hammer B, Gersmann K: A Note on the Universal Approximation Capability of Support Vector Machines. Neural Processing Letters 2003, 17: 43\u201353.","journal-title":"Neural Processing Letters"},{"key":"2304_CR15","doi-asserted-by":"publisher","first-page":"77","DOI":"10.1198\/016214502753479248","volume":"97","author":"S Dudoit","year":"2002","unstructured":"Dudoit S, Fridlyand J, Speed TP: Comparison of discrimination methods for the classification of tumors using gene expression data. Journal of the American Statistical Association 2002, 97: 77\u201388.","journal-title":"Journal of the American Statistical Association"},{"key":"2304_CR16","doi-asserted-by":"publisher","first-page":"147","DOI":"10.1093\/jnci\/djk018","volume":"99","author":"A Dupuy","year":"2007","unstructured":"Dupuy A, Simon RM: Critical review of published microarray studies for cancer outcome and guidelines on statistical analysis and reporting. J Natl Cancer Inst 2007, 99: 147\u2013157.","journal-title":"J Natl Cancer Inst"},{"key":"2304_CR17","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/3-540-45014-9_1","volume-title":"Proceedings of the First International Workshop on Multiple Classifier Systems","author":"TG Dietterich","year":"2000","unstructured":"Dietterich TG: Ensemble methods in machine learning. In Proceedings of the First International Workshop on Multiple Classifier Systems. New York, NY, Springer-Verlag; 2000:1\u201315."},{"key":"2304_CR18","volume-title":"Technical Report, Center for Bioinformatics & Molecular Biostatistics, University of California, San Francisco","author":"MR Segal","year":"2004","unstructured":"Segal MR: Machine Learning Benchmarks and Random Forest Regression. Technical Report, Center for Bioinformatics & Molecular Biostatistics, University of California, San Francisco 2004."},{"key":"2304_CR19","doi-asserted-by":"publisher","first-page":"491","DOI":"10.1016\/j.ijmedinf.2005.05.002","volume":"74","author":"A Statnikov","year":"2005","unstructured":"Statnikov A, Tsamardinos I, Dosbayev Y, Aliferis CF: GEMS: a system for automated cancer diagnosis and biomarker discovery from microarray gene expression data. Int J Med Inform 2005, 74: 491\u2013503.","journal-title":"Int J Med Inform"},{"key":"2304_CR20","volume-title":"Error estimation and model selection","author":"T Scheffer","year":"1999","unstructured":"Scheffer T: Error estimation and model selection. Ph.D.Thesis, Technischen Universit\u00e4t Berlin, School of Computer Science; 1999."},{"key":"2304_CR21","doi-asserted-by":"publisher","first-page":"906","DOI":"10.1093\/bioinformatics\/16.10.906","volume":"16","author":"TS Furey","year":"2000","unstructured":"Furey TS, Cristianini N, Duffy N, Bednarski DW, Schummer M, Haussler D: Support vector machine classification and validation of cancer tissue samples using microarray expression data. Bioinformatics 2000, 16: 906\u2013914.","journal-title":"Bioinformatics"},{"key":"2304_CR22","volume-title":"Statistical learning theory","author":"VN Vapnik","year":"1998","unstructured":"Vapnik VN Adaptive and learning systems for signal processing, communications, and control. In Statistical learning theory. New York, Wiley; 1998."},{"key":"2304_CR23","first-page":"1918","volume":"6","author":"RE Fan","year":"2005","unstructured":"Fan RE, Chen PH, Lin CJ: Working set selection using second order information for training support vector machines. Journal of Machine Learning Research 2005, 6: 1918.","journal-title":"Journal of Machine Learning Research"},{"key":"2304_CR24","first-page":"18","volume":"2","author":"A Liaw","year":"2002","unstructured":"Liaw A, Wiener M: Classification and regression by randomForest. R News 2002, 2: 18\u201322.","journal-title":"R News"},{"key":"2304_CR25","volume-title":"Manual on setting up, using, and understanding Random Forests v4.0","author":"L Breiman","year":"2003","unstructured":"Breiman L: Manual on setting up, using, and understanding Random Forests v4.0.2003. [ftp:\/\/ftp.stat.berkeley.edu\/pub\/users\/breiman\/]"},{"key":"2304_CR26","doi-asserted-by":"publisher","first-page":"389","DOI":"10.1023\/A:1012487302797","volume":"46","author":"I Guyon","year":"2002","unstructured":"Guyon I, Weston J, Barnhill S, Vapnik V: Gene selection for cancer classification using support vector machines. Machine Learning 2002, 46: 389\u2013422.","journal-title":"Machine Learning"},{"key":"2304_CR27","doi-asserted-by":"publisher","first-page":"1685","DOI":"10.1016\/j.patrec.2006.03.013","volume":"27","author":"X Chen","year":"2006","unstructured":"Chen X, Zeng X, van Alphen D: Multi-class feature selection for texture classification. Pattern Recognition Letters 2006, 27: 1685\u20131691.","journal-title":"Pattern Recognition Letters"},{"key":"2304_CR28","doi-asserted-by":"publisher","first-page":"531","DOI":"10.1126\/science.286.5439.531","volume":"286","author":"TR Golub","year":"1999","unstructured":"Golub TR, Slonim DK, Tamayo P, Huard C, Gaasenbeek M, Mesirov JP, Coller H, Loh ML, Downing JR, Caligiuri MA, Bloomfield CD, Lander ES: Molecular classification of cancer: class discovery and class prediction by gene expression monitoring. Science 1999, 286: 531\u2013537.","journal-title":"Science"},{"key":"2304_CR29","doi-asserted-by":"publisher","first-page":"1331","DOI":"10.1109\/IJCNN.2004.1380138","volume":"2","author":"J Menke","year":"2004","unstructured":"Menke J, Martinez TR: Using permutations instead of student's t distribution for p-values in paired-difference algorithm comparisons. Proceedings of 2004 IEEE International Joint Conference on Neural Networks 2004, 2: 1331\u20131335.","journal-title":"Proceedings of 2004 IEEE International Joint Conference on Neural Networks"},{"key":"2304_CR30","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4757-3235-1","volume-title":"Permutation tests: a practical guide to resampling methods for testing hypotheses","author":"PI Good","year":"2000","unstructured":"Good PI Springer series in statistics. In Permutation tests: a practical guide to resampling methods for testing hypotheses. 2nd edition. New York, Springer; 2000.","edition":"2nd"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-9-319.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T03:19:02Z","timestamp":1630466342000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-9-319"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,7,22]]},"references-count":30,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2008,12]]}},"alternative-id":["2304"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-9-319","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2008,7,22]]},"assertion":[{"value":"24 January 2008","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 July 2008","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 July 2008","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"319"}}