{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T22:59:25Z","timestamp":1784329165493,"version":"3.55.0"},"reference-count":24,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2011,7,29]],"date-time":"2011-07-29T00:00:00Z","timestamp":1311897600000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Med Inform Decis Mak"],"published-print":{"date-parts":[[2011,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>We present a method utilizing Healthcare Cost and Utilization Project (HCUP) dataset for predicting disease risk of individuals based on their medical diagnosis history. The presented methodology may be incorporated in a variety of applications such as risk management, tailored health communication and decision support systems in healthcare.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Methods<\/jats:title>\n            <jats:p>We employed the National Inpatient Sample (NIS) data, which is publicly available through Healthcare Cost and Utilization Project (HCUP), to train random forest classifiers for disease prediction. Since the HCUP data is highly imbalanced, we employed an ensemble learning approach based on repeated random sub-sampling. This technique divides the training data into multiple sub-samples, while ensuring that each sub-sample is fully balanced. We compared the performance of support vector machine (SVM), bagging, boosting and RF to predict the risk of eight chronic diseases.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>We predicted eight disease categories. Overall, the RF ensemble learning method outperformed SVM, bagging and boosting in terms of the area under the receiver operating characteristic (ROC) curve (AUC). In addition, RF has the advantage of computing the importance of each variable in the classification process.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusions<\/jats:title>\n            <jats:p>In combining repeated random sub-sampling with RF, we were able to overcome the class imbalance problem and achieve promising results. Using the national HCUP data set, we predicted eight disease categories with an average AUC of 88.79%.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1472-6947-11-51","type":"journal-article","created":{"date-parts":[[2011,7,29]],"date-time":"2011-07-29T18:18:55Z","timestamp":1311963535000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":576,"title":["Predicting disease risks from highly imbalanced data using random forest"],"prefix":"10.1186","volume":"11","author":[{"given":"Mohammed","family":"Khalilia","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sounak","family":"Chakraborty","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mihail","family":"Popescu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2011,7,29]]},"reference":[{"issue":"1","key":"428_CR1","doi-asserted-by":"publisher","first-page":"16","DOI":"10.1186\/1472-6947-10-16","volume":"10","author":"W Yu","year":"2010","unstructured":"Yu W: Application of support vector machine modeling for prediction of common diseases: the case of diabetes and pre-diabetes. BMC Medical Informatics and Decision Making. 2010, 10 (1): 16-10.1186\/1472-6947-10-16.","journal-title":"BMC Medical Informatics and Decision Making"},{"issue":"6","key":"428_CR2","doi-asserted-by":"publisher","first-page":"270","DOI":"10.1177\/106286069901400607","volume":"14","author":"P Hebert","year":"1999","unstructured":"Hebert P: Identifying persons with diabetes using Medicare claims data. American Journal of Medical Quality. 1999, 14 (6): 270-10.1177\/106286069901400607.","journal-title":"American Journal of Medical Quality"},{"key":"428_CR3","volume-title":"Medical Underwriting for Life Insurance","author":"V Fuster","year":"2008","unstructured":"Fuster V: Medical Underwriting for Life Insurance. 2008, McGraw-Hill's AccessMedicine"},{"key":"428_CR4","volume-title":"Machine Learning and Cybernetics, 2005. Proceedings of 2005 International Conference on","author":"T Yi","year":"2005","unstructured":"Yi T, Guo-Ji Z: The application of machine learning algorithm in underwriting process. Machine Learning and Cybernetics, 2005. Proceedings of 2005 International Conference on. 2005"},{"issue":"5","key":"428_CR5","doi-asserted-by":"publisher","first-page":"427","DOI":"10.1080\/10410230802342176","volume":"23","author":"E Cohen","year":"2008","unstructured":"Cohen E: Cancer coverage in general-audience and black newspapers. Health Communication. 2008, 23 (5): 427-435. 10.1080\/10410230802342176.","journal-title":"Health Communication"},{"key":"428_CR6","volume-title":"Overview of the Nationwide Inpatient Sample (NIS)","author":"HCUP Project","year":"2009","unstructured":"HCUP Project: Overview of the Nationwide Inpatient Sample (NIS). 2009, [http:\/\/www.hcup-us.ahrq.gov\/nisoverview.jsp]"},{"key":"428_CR7","volume-title":"Bioinformatics and Biomedicine, 2007. BIBM 2007. IEEE International Conference on","author":"ST Moturu","year":"2007","unstructured":"Moturu ST, Johnson WG, Huan L: Predicting Future High-Cost Patients: A Real-World Risk Modeling Application. Bioinformatics and Biomedicine, 2007. BIBM 2007. IEEE International Conference on. 2007"},{"key":"428_CR8","first-page":"769","volume-title":"Predicting individual disease risk based on medical history","author":"DA Davis","year":"2008","unstructured":"Davis DA, Chawla NV, Blumm N, Christakis N, Barab\u00e1si AL: Proceeding of the 17th ACM conference on Information and knowledge management. Predicting individual disease risk based on medical history. 2008, 769-778."},{"key":"428_CR9","volume-title":"BioInformatics and BioEngineering, 2008. BIBE 2008. 8th IEEE International Conference on","author":"DH Mantzaris","year":"2008","unstructured":"Mantzaris DH, Anastassopoulos GC, Lymberopoulos DK: Medical disease prediction using Artificial Neural Networks. BioInformatics and BioEngineering, 2008. BIBE 2008. 8th IEEE International Conference on. 2008"},{"key":"428_CR10","doi-asserted-by":"publisher","first-page":"242","DOI":"10.1109\/IJCBS.2009.23","volume-title":"Bioinformatics, Systems Biology and Intelligent Computing, 2009. IJCBS '09. International Joint Conference on","author":"W Zhang","year":"2009","unstructured":"Zhang W: A Comparative Study of Ensemble Learning Approaches in the Classification of Breast Cancer Metastasis. Bioinformatics, Systems Biology and Intelligent Computing, 2009. IJCBS '09. International Joint Conference on. 2009, 242-245."},{"key":"428_CR11","first-page":"183","volume-title":"A Smart Home Application to Eldercare: Current Status and Lessons Learned, Technology and Health Care","author":"M Skubic","year":"2009","unstructured":"Skubic M, Alexander G, Popescu M, Rantz M, Keller J: A Smart Home Application to Eldercare: Current Status and Lessons Learned, Technology and Health Care. 2009, 17 (3): 183-201."},{"key":"428_CR12","volume-title":"Proceedings of the AAAI'2000 Workshop on Imbalanced Data Sets","author":"F Provost","year":"2000","unstructured":"Provost F: Machine learning from imbalanced data sets 101. Proceedings of the AAAI'2000 Workshop on Imbalanced Data Sets. 2000"},{"issue":"5","key":"428_CR13","doi-asserted-by":"crossref","first-page":"429","DOI":"10.3233\/IDA-2002-6504","volume":"6","author":"N Japkowicz","year":"2002","unstructured":"Japkowicz N, Stephen S: The class imbalance problem: A systematic study. Intelligent Data Analysis. 2002, 6 (5): 429-449.","journal-title":"Intelligent Data Analysis"},{"key":"428_CR14","first-page":"725","volume-title":"Proceedings of the National Conference on Artificial Intelligence","author":"JR Quinlan","year":"1996","unstructured":"Quinlan JR: Bagging, boosting, and C4. 5. Proceedings of the National Conference on Artificial Intelligence. 1996, 725-730."},{"key":"428_CR15","volume-title":"Classification and regression trees","author":"L Breiman","year":"1984","unstructured":"Breiman L: Classification and regression trees. 1984, Wadsworth. Inc., Belmont, CA, 358:"},{"issue":"1","key":"428_CR16","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman L: Random forests. Machine learning. 2001, 45 (1): 5-32. 10.1023\/A:1010933404324.","journal-title":"Machine learning"},{"key":"428_CR17","volume-title":"Using random forest to learn imbalanced data","author":"C Chen","year":"2004","unstructured":"Chen C, Liaw A, Breiman L: Using random forest to learn imbalanced data. 2004, University of California, Berkeley"},{"key":"428_CR18","volume-title":"Manual-Setting Up, Using, and Understanding Random Forests V4. 0","author":"L Breiman","year":"2003","unstructured":"Breiman L, others: Manual-Setting Up, Using, and Understanding Random Forests V4. 0. 2003, [ftp:\/\/ftpstat.berkeley.edu\/pub\/users\/breiman]"},{"key":"428_CR19","doi-asserted-by":"crossref","first-page":"605","DOI":"10.1007\/978-0-387-84858-7_16","volume-title":"The elements of statistical learning: data mining, inference and prediction","author":"T Hastie","year":"2009","unstructured":"Hastie T: The elements of statistical learning: data mining, inference and prediction. 2009, 605-622."},{"key":"428_CR20","unstructured":"Bjoern M: A comparison of random forest and its Gini importance with standard chemometric methods for the feature selection and classification of spectral data. BMC Bioinformatics. 10:"},{"issue":"4","key":"428_CR21","first-page":"319","volume":"3","author":"J Mingers","year":"1989","unstructured":"Mingers J: An empirical comparison of selection measures for decision-tree induction. Machine learning. 1989, 3 (4): 319-342.","journal-title":"Machine learning"},{"key":"428_CR22","doi-asserted-by":"publisher","first-page":"1145","DOI":"10.1016\/S0031-3203(96)00142-2","volume":"30","author":"AP Bradley","year":"1997","unstructured":"Bradley AP: The use of the area under the ROC curve in the evaluation of machine learning algorithms. Pattern Recognition. 1997, 30: 1145-1159. 10.1016\/S0031-3203(96)00142-2.","journal-title":"Pattern Recognition"},{"issue":"1","key":"428_CR23","doi-asserted-by":"publisher","first-page":"150","DOI":"10.1021\/ci060164k","volume":"47","author":"D Palmer","year":"2007","unstructured":"Palmer D: Random forest models to predict aqueous solubility. J Chem Inf Model. 2007, 47 (1): 150-158. 10.1021\/ci060164k.","journal-title":"J Chem Inf Model"},{"key":"428_CR24","unstructured":"Liaw A, Wiener M: Classification and Regression by randomForest."}],"container-title":["BMC Medical Informatics and Decision Making"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/1472-6947-11-51.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/1472-6947-11-51\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/1472-6947-11-51","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1472-6947-11-51.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T15:40:32Z","timestamp":1630510832000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcmedinformdecismak.biomedcentral.com\/articles\/10.1186\/1472-6947-11-51"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,7,29]]},"references-count":24,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2011,12]]}},"alternative-id":["428"],"URL":"https:\/\/doi.org\/10.1186\/1472-6947-11-51","relation":{},"ISSN":["1472-6947"],"issn-type":[{"value":"1472-6947","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,7,29]]},"assertion":[{"value":"28 September 2010","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 July 2011","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 July 2011","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"51"}}