{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T14:17:59Z","timestamp":1787062679499,"version":"3.56.0"},"reference-count":49,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,4,11]],"date-time":"2023-04-11T00:00:00Z","timestamp":1681171200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,4,11]],"date-time":"2023-04-11T00:00:00Z","timestamp":1681171200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Big Data"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Using the wrong metrics to gauge classification of highly imbalanced Big Data may hide important information in experimental results. However, we find that analysis of metrics for performance evaluation and what they can hide or reveal is rarely covered in related works. Therefore, we address that gap by analyzing multiple popular performance metrics on three Big Data classification tasks. To the best of our knowledge, we are the first to utilize three new Medicare insurance claims datasets which became publicly available in 2021. These datasets are all highly imbalanced. Furthermore, the datasets are comprised of completely different data. We evaluate the performance of five ensemble learners in the Machine Learning task of Medicare fraud detection. Random Undersampling (RUS) is applied to induce five class ratios. The classifiers are evaluated with both the Area Under the Receiver Operating Characteristic Curve (AUC), and Area Under the Precision Recall Curve (AUPRC) metrics. We show that AUPRC provides a better insight into classification performance. Our findings reveal that the AUC metric hides the performance impact of RUS. However, classification results in terms of AUPRC show RUS has a detrimental effect. We show that, for highly imbalanced Big Data, the AUC metric fails to capture information about precision scores and false positive counts that the AUPRC metric reveals. Our contribution is to show AUPRC is a more effective metric for evaluating the performance of classifiers when working with highly imbalanced Big Data.<\/jats:p>","DOI":"10.1186\/s40537-023-00724-5","type":"journal-article","created":{"date-parts":[[2023,4,11]],"date-time":"2023-04-11T11:04:47Z","timestamp":1681211087000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":102,"title":["Evaluating classifier performance with highly imbalanced Big Data"],"prefix":"10.1186","volume":"10","author":[{"given":"John T.","family":"Hancock","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Taghi M.","family":"Khoshgoftaar","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Justin M.","family":"Johnson","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,4,11]]},"reference":[{"key":"724_CR1","first-page":"10","volume":"3","author":"M Bekkar","year":"2013","unstructured":"Bekkar M, Djemaa HK, Alitouche TA. Evaluation measures for models assessment over imbalanced data sets. J Inf Eng Appl. 2013;3:10.","journal-title":"J Inf Eng Appl"},{"key":"724_CR2","first-page":"451","volume-title":"Joint European Conference on Machine Learning and Knowledge Discovery in Databases","author":"K Boyd","year":"2013","unstructured":"Boyd K, Eng KH, Page CD. Area under the precision-recall curve: point estimates and confidence intervals. In: Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer: New York; 2013. p. 451\u201366."},{"key":"724_CR3","unstructured":"The Centers for Medicare and Medicaid Services: Medicare Durable Medical Equipment, Devices & Supplies \u2013 by Referring Provider and Service. https:\/\/data.cms.gov\/provider-summary-by-type-of-service\/medicare-durable-medical-equipment-devices-supplies\/medicare-durable-medical-equipment-devices-supplies-by-referring-provider-and-service 2021."},{"key":"724_CR4","unstructured":"The Centers for Medicare and Medicaid Services: Medicare Physician & Other Practitioners \u2013 by Provider and Service. https:\/\/data.cms.gov\/provider-summary-by-type-of-service\/medicare-physician-other-practitioners\/medicare-physician-other-practitioners-by-provider-and-service. 2021."},{"key":"724_CR5","unstructured":"The Centers for Medicare and Medicaid Services: Medicare Part D Prescribers \u2013 by Provider and Drug. https:\/\/data.cms.gov\/provider-summary-by-type-of-service\/medicare-part-d-prescribers\/medicare-part-d-prescribers-by-provider-and-drug 2021."},{"key":"724_CR6","doi-asserted-by":"crossref","unstructured":"De\u00a0Mauro A, Greco M, Grimaldi M. A formal definition of big data based on its essential features. Library Review 2016.","DOI":"10.1108\/LR-06-2015-0061"},{"key":"724_CR7","unstructured":"Civil Division, U.S. Department of Justice: Fraud Statistics, Overview. https:\/\/www.justice.gov\/opa\/press-release\/file\/1354316\/download 2020."},{"key":"724_CR8","unstructured":"Centers for Medicare and Medicaid Services: 2019 Estimated Improper Payment Rates for Centers for Medicare & Medicaid Services (CMS) Programs. https:\/\/www.cms.gov\/newsroom\/fact-sheets\/2019-estimated-improper-payment-rates-centers-medicare-medicaid-services-cms-programs 2019."},{"key":"724_CR9","doi-asserted-by":"crossref","unstructured":"Bauder RA, Khoshgoftaar TM, Hasanin T. Data sampling approaches with severely imbalanced big data for medicare fraud detection. In: 2018 IEEE 30th International Conference on Tools with Artificial Intelligence (ICTAI), 2018;137\u2013142. IEEE","DOI":"10.1109\/ICTAI.2018.00030"},{"issue":"1","key":"724_CR10","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-021-00460-8","volume":"8","author":"R Zuech","year":"2021","unstructured":"Zuech R, Hancock JT, Khoshgoftaar TM. Detecting web attacks using random undersampling and ensemble learners. J Big Data. 2021;8(1):1\u201320.","journal-title":"J Big Data"},{"key":"724_CR11","first-page":"8","volume":"31","author":"L Prokhorenkova","year":"2018","unstructured":"Prokhorenkova L, Gusev G, Vorobev A, Dorogush AV, Gulin A. Catboost: unbiased boosting with categorical features. Adva Neural Inf Process Syst. 2018;31:8.","journal-title":"Adva Neural Inf Process Syst"},{"key":"724_CR12","doi-asserted-by":"publisher","unstructured":"Chen T, Guestrin C. Xgboost: A scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining - KDD \u201916 2016. https:\/\/doi.org\/10.1145\/2939672.2939785.","DOI":"10.1145\/2939672.2939785"},{"key":"724_CR13","first-page":"3146","volume":"30","author":"G Ke","year":"2017","unstructured":"Ke G, Meng Q, Finley T, Wang T, Chen W, Ma W, Ye Q, Liu T-Y. Lightgbm: A highly efficient gradient boosting decision tree. Adva Neural Inf Process Syst. 2017;30:3146\u201354.","journal-title":"Adva Neural Inf Process Syst"},{"issue":"1","key":"724_CR14","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman L. Random forests. Mach Learn. 2001;45(1):5\u201332.","journal-title":"Mach Learn"},{"issue":"1","key":"724_CR15","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1007\/s10994-006-6226-1","volume":"63","author":"P Geurts","year":"2006","unstructured":"Geurts P, Ernst D, Wehenkel L. Extremely randomized trees. Mach Learn. 2006;63(1):3\u201342.","journal-title":"Mach Learn"},{"key":"724_CR16","first-page":"878","volume-title":"International Conference on Intelligent Computing","author":"H Han","year":"2005","unstructured":"Han H, Wang W-Y, Mao B-H. Borderline-smote: a new over-sampling method in imbalanced data sets learning. In: International Conference on Intelligent Computing. Springer: Berlin; 2005. p. 878\u201387."},{"key":"724_CR17","unstructured":"He H, Bai Y, Garcia EA, Li S. Adasyn: Adaptive synthetic sampling approach for imbalanced learning. In: 2008 IEEE International Joint Conference on Neural Networks (IEEE World Congress on Computational Intelligence), 2008;1322\u20131328. IEEE"},{"issue":"11","key":"724_CR18","doi-asserted-by":"publisher","first-page":"56","DOI":"10.1145\/2934664","volume":"59","author":"M Zaharia","year":"2016","unstructured":"Zaharia M, Xin RS, Wendell P, Das T, Armbrust M, Dave A, Meng X, Rosen J, Venkataraman S, Franklin MJ, et al. Apache spark: a unified engine for big data processing. Commun ACM. 2016;59(11):56\u201365.","journal-title":"Commun ACM"},{"issue":"1","key":"724_CR19","first-page":"191","volume":"41","author":"S Le Cessie","year":"1992","unstructured":"Le Cessie S, Van Houwelingen JC. Ridge estimators in logistic regression. J R Stat Soc. 1992;41(1):191\u2013201.","journal-title":"J R Stat Soc"},{"issue":"1","key":"724_CR20","first-page":"1235","volume":"17","author":"X Meng","year":"2016","unstructured":"Meng X, Bradley J, Yavuz B, Sparks E, Venkataraman S, Liu D, Freeman J, Tsai D, Amde M, Owen S, et al. Mllib: Machine learning in apache spark. J Mach Learn Res. 2016;17(1):1235\u201341.","journal-title":"J Mach Learn Res"},{"issue":"1","key":"724_CR21","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-019-0274-4","volume":"6","author":"T Hasanin","year":"2019","unstructured":"Hasanin T, Khoshgoftaar TM, Leevy JL, Bauder RA. Severely imbalanced big data challenges: investigating data sampling approaches. J Big Data. 2019;6(1):1\u201325.","journal-title":"J Big Data"},{"issue":"2","key":"724_CR22","doi-asserted-by":"publisher","first-page":"215","DOI":"10.1007\/s13748-019-00172-4","volume":"8","author":"LI Kuncheva","year":"2019","unstructured":"Kuncheva LI, Arnaiz-Gonzalez A, D\u00edez-Pastor J-F, Gunn IA. Instance selection improves geometric mean accuracy: a study on imbalanced data classification. Prog Artif Intell. 2019;8(2):215\u201328.","journal-title":"Prog Artif Intell"},{"issue":"9","key":"724_CR23","doi-asserted-by":"publisher","first-page":"1263","DOI":"10.1109\/TKDE.2008.239","volume":"21","author":"H He","year":"2009","unstructured":"He H, Garcia EA. Learning from imbalanced data. IEEE Trans Knowl Data Eng. 2009;21(9):1263\u201384.","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"724_CR24","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2020.106598","volume":"212","author":"WC Sleeman IV","year":"2021","unstructured":"Sleeman WC IV, Krawczyk B. Multi-class imbalanced big data classification on spark. Knowl-Based Syst. 2021;212: 106598.","journal-title":"Knowl-Based Syst"},{"key":"724_CR25","doi-asserted-by":"crossref","unstructured":"Calvert CL, Khoshgoftaar TM. Threshold based optimization of performance metrics with severely imbalanced big security data. In: 2019 IEEE 31st International Conference on Tools with Artificial Intelligence (ICTAI), 2019. p. 1328\u201334.","DOI":"10.1109\/ICTAI.2019.00184"},{"issue":"1","key":"724_CR26","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-019-0225-0","volume":"6","author":"JM Johnson","year":"2019","unstructured":"Johnson JM, Khoshgoftaar TM. Medicare fraud detection using neural networks. J Big Data. 2019;6(1):1\u201335.","journal-title":"J Big Data"},{"issue":"1","key":"724_CR27","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-020-00305-w","volume":"7","author":"JT Hancock","year":"2020","unstructured":"Hancock JT, Khoshgoftaar TM. Survey on categorical data for neural networks. J Big Data. 2020;7(1):1\u201341.","journal-title":"J Big Data"},{"issue":"5","key":"724_CR28","doi-asserted-by":"publisher","first-page":"1113","DOI":"10.1007\/s10796-020-10022-7","volume":"22","author":"JM Johnson","year":"2020","unstructured":"Johnson JM, Khoshgoftaar TM. The effects of data sampling with deep learning and highly imbalanced big data. Inf Syst Front. 2020;22(5):1113\u201331.","journal-title":"Inf Syst Front"},{"key":"724_CR29","unstructured":"Apache Software Foundation: Hadoop. https:\/\/hadoop.apache.org."},{"issue":"3","key":"724_CR30","doi-asserted-by":"publisher","first-page":"0118432","DOI":"10.1371\/journal.pone.0118432","volume":"10","author":"T Saito","year":"2015","unstructured":"Saito T, Rehmsmeier M. The precision-recall plot is more informative than the roc plot when evaluating binary classifiers on imbalanced datasets. PloS one. 2015;10(3):0118432.","journal-title":"PloS one"},{"issue":"2","key":"724_CR31","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1007\/BF00058655","volume":"24","author":"L Breiman","year":"1996","unstructured":"Breiman L. Bagging predictors. Machine learning. 1996;24(2):123\u201340.","journal-title":"Machine learning"},{"key":"724_CR32","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1201\/9780429246593","volume-title":"An Introduction to the Bootstrap","author":"B Efron","year":"1994","unstructured":"Efron B, Tibshirani RJ. An Introduction to the Bootstrap. Boca Raton: CRC Press; 1994. p. 5\u20136."},{"key":"724_CR33","doi-asserted-by":"crossref","unstructured":"Hasanin T, Khoshgoftaar TM. The effects of random undersampling with simulated class imbalance for big data. In: 2018 IEEE International Conference on Information Reuse and Integration (IRI), 2018. p. 70\u20139.","DOI":"10.1109\/IRI.2018.00018"},{"key":"724_CR34","first-page":"1189","volume":"34","author":"JH Friedman","year":"2001","unstructured":"Friedman JH. Greedy function approximation: a gradient boosting machine. Ann Stat. 2001;34:1189\u2013232.","journal-title":"Ann Stat"},{"issue":"1","key":"724_CR35","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-020-00369-8","volume":"7","author":"JT Hancock","year":"2020","unstructured":"Hancock JT, Khoshgoftaar TM. Catboost for big data: an interdisciplinary review. J Big Data. 2020;7(1):1\u201345.","journal-title":"J Big Data"},{"key":"724_CR36","doi-asserted-by":"crossref","unstructured":"Hancock JT, Khoshgoftaar TM. Leveraging lightgbm for categorical big data. In: 2021 IEEE Seventh International Conference on Big Data Computing Service and Applications (BigDataService), 2021. p. 149\u2013154.","DOI":"10.1109\/BigDataService52369.2021.00024"},{"key":"724_CR37","unstructured":"LEIE: Office of Inspector General Leie Downloadable Databases. https:\/\/oig.hhs.gov\/exclusions\/index.asp."},{"key":"724_CR38","doi-asserted-by":"crossref","unstructured":"Bauder RA, Khoshgoftaar TM. A novel method for fraudulent medicare claims detection from expected payment deviations (application paper). In: 2016 IEEE 17th International Conference on Information Reuse and Integration (IRI), 2016. p. 11\u201319.","DOI":"10.1109\/IRI.2016.11"},{"key":"724_CR39","unstructured":"The Centers for Medicare and Medicaid Services: Medicare Durable Medical Equipment, Devices & Supplies \u2013 by Referring Provider and Service Data Dictionary. https:\/\/data.cms.gov\/resources\/medicare-durable-medical-equipment-devices-supplies-by-referring-provider-and-service-data-dictionary 2021."},{"key":"724_CR40","unstructured":"The Centers for Medicare and Medicaid Services: Medicare Physician & Other Practitioners \u2013 by Provider and Service Data Dictionary. https:\/\/data.cms.gov\/resources\/medicare-physician-other-practitioners-by-provider-and-service-data-dictionary. 2021."},{"key":"724_CR41","unstructured":"The Centers for Medicare and Medicaid Services: Medicare Part D Prescribers \u2013 by Provider and Drug Data Dictionary. https:\/\/data.cms.gov\/resources\/medicare-part-d-prescribers-by-provider-and-drug-data-dictionary 2021."},{"key":"724_CR42","unstructured":"Van\u00a0Rossum G, Drake FL. Python\/C Api Manual-Python 3. CreateSpace 2009."},{"key":"724_CR43","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, et al. Scikit-learn: Machine learning in python. J Mach Learn Res. 2011;12:2825\u201330.","journal-title":"J Mach Learn Res"},{"key":"724_CR44","unstructured":"McGinnis W. Category Encoders. https:\/\/contrib.scikit-learn.org\/category_encoders\/."},{"issue":"4","key":"724_CR45","doi-asserted-by":"publisher","first-page":"276","DOI":"10.1007\/s42979-021-00656-y","volume":"2","author":"JM Johnson","year":"2021","unstructured":"Johnson JM, Khoshgoftaar TM. Medical provider embeddings for healthcare fraud detection. SN Computer Sci. 2021;2(4):276.","journal-title":"SN Computer Sci"},{"key":"724_CR46","unstructured":"XGBoost Parameters. XGBoost Developers. https:\/\/xgboost.readthedocs.io\/en\/stable\/parameter.html Accessed 9 Jul 2022."},{"key":"724_CR47","unstructured":"Parameters. Yandex Corporation. https:\/\/catboost.ai\/en\/docs\/references\/training-parameters\/common. Accessed 9 Jul 2022."},{"key":"724_CR48","doi-asserted-by":"publisher","DOI":"10.4135\/9781412983327","volume-title":"Analysis of Variance","author":"GR Iversen","year":"1987","unstructured":"Iversen GR, Norpoth H. Analysis of Variance, vol. 1. Newbury Park: Sage; 1987."},{"key":"724_CR49","doi-asserted-by":"publisher","first-page":"99","DOI":"10.2307\/3001913","volume":"56","author":"JW Tukey","year":"1949","unstructured":"Tukey JW. Comparing individual means in the analysis of variance. Biometrics. 1949;56:99\u2013114.","journal-title":"Biometrics"}],"container-title":["Journal of Big Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-023-00724-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s40537-023-00724-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-023-00724-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,4,11]],"date-time":"2023-04-11T11:05:22Z","timestamp":1681211122000},"score":1,"resource":{"primary":{"URL":"https:\/\/journalofbigdata.springeropen.com\/articles\/10.1186\/s40537-023-00724-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4,11]]},"references-count":49,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["724"],"URL":"https:\/\/doi.org\/10.1186\/s40537-023-00724-5","relation":{},"ISSN":["2196-1115"],"issn-type":[{"value":"2196-1115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,4,11]]},"assertion":[{"value":"25 September 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 March 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 April 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"42"}}