{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T02:43:01Z","timestamp":1787020981908,"version":"3.56.0"},"reference-count":52,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,10,9]],"date-time":"2023-10-09T00:00:00Z","timestamp":1696809600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,10,9]],"date-time":"2023-10-09T00:00:00Z","timestamp":1696809600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Big Data"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>As a means of building explainable machine learning models for Big Data, we apply a novel ensemble supervised feature selection technique. The technique is applied to publicly available insurance claims data from the United States public health insurance program, Medicare. We approach Medicare insurance fraud detection as a supervised machine learning task of anomaly detection through the classification of highly imbalanced Big Data. Our objectives for feature selection are to increase efficiency in model training, and to develop more explainable machine learning models for fraud detection. Using two Big Data datasets derived from two different sources of insurance claims data, we demonstrate how our feature selection technique reduces the dimensionality of the datasets by approximately 87.5% without compromising performance. Moreover, the reduction in dimensionality results in machine learning models that are easier to explain, and less prone to overfitting. Therefore, our primary contribution of the exposition of our novel feature selection technique leads to a further contribution to the application domain of automated Medicare insurance fraud detection. We utilize our feature selection technique to provide an explanation of our fraud detection models in terms of the definitions of the selected features. The ensemble supervised feature selection technique we present is flexible in that any collection of machine learning algorithms that maintain a list of feature importance values may be used. Therefore, researchers may easily employ variations of the technique we present.<\/jats:p>","DOI":"10.1186\/s40537-023-00821-5","type":"journal-article","created":{"date-parts":[[2023,10,9]],"date-time":"2023-10-09T06:14:02Z","timestamp":1696832042000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":35,"title":["Explainable machine learning models for Medicare fraud detection"],"prefix":"10.1186","volume":"10","author":[{"given":"John T.","family":"Hancock","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Richard A.","family":"Bauder","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huanjing","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Taghi M.","family":"Khoshgoftaar","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,10,9]]},"reference":[{"key":"821_CR1","unstructured":"Zuech R, Khoshgoftaar TM. A survey on feature selection for intrusion detection. In: Proceedings of the 21st issat international conference on reliability and quality in design; 2015. p. 150\u20135."},{"key":"821_CR2","unstructured":"Centers for medicare and medicaid services: about CMS; 2023. https:\/\/www.cms.gov\/About-CMS\/About-CMS."},{"key":"821_CR3","unstructured":"Civil Division, U.S. Department of Justice: fraud statistics, overview; 2020. https:\/\/www.justice.gov\/opa\/press-release\/file\/1354316\/download."},{"key":"821_CR4","unstructured":"Centers for Medicare and Medicaid Services: 2019 estimated improper payment rates for centers for medicare & medicaid services (CMS) programs; 2019. https:\/\/www.cms.gov\/newsroom\/fact-sheets\/2019-estimated-improper-payment-rates-centers-medicare-medicaid-services-cms-programs."},{"key":"821_CR5","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1007\/s10742-016-0154-8","volume":"17","author":"R Bauder","year":"2017","unstructured":"Bauder R, Khoshgoftaar TM, Seliya N. A survey on the state of healthcare upcoding fraud analysis and detection. Health Serv Outcomes Res Methodol. 2017;17:31\u201355.","journal-title":"Health Serv Outcomes Res Methodol"},{"key":"821_CR6","doi-asserted-by":"crossref","unstructured":"Mayaki MZA, Riveill M. Multiple inputs neural networks for fraud detection. In: 2022 international conference on machine learning, control, and robotics (MLCR). New York: IEEE; 2022. p. 8\u201313.","DOI":"10.1109\/MLCR57210.2022.00011"},{"key":"821_CR7","unstructured":"LEIE: office of inspector general Leie downloadable databases. https:\/\/oig.hhs.gov\/exclusions\/index.asp."},{"key":"821_CR8","doi-asserted-by":"crossref","unstructured":"Salekshahrezaee Z, Leevy JL, Khoshgoftaar TM. A class-imbalanced study with feature extraction via pca and convolutional autoencoder. In: 2022 IEEE 23rd international conference on information reuse and integration for data science (IRI). New York: IEEE; 2022. p. 63\u20138.","DOI":"10.1109\/IRI54793.2022.00026"},{"key":"821_CR9","doi-asserted-by":"crossref","unstructured":"Boyd K, Eng KH, Page CD. Area under the precision-recall curve: point estimates and confidence intervals. In: Joint European conference on machine learning and knowledge discovery in databases. Berlin: Springer; 2013. p. 451\u201366.","DOI":"10.1007\/978-3-642-40994-3_29"},{"key":"821_CR10","doi-asserted-by":"crossref","unstructured":"Waspada I, Bahtiar N, Wirawan PW, Awan BDA. Performance analysis of isolation forest algorithm in fraud detection of credit card transactions. Khazanah Informatika: Jurnal Ilmu Komputer dan Informatika 2020;6(2):165\u201375.","DOI":"10.23917\/khif.v6i2.10520"},{"key":"821_CR11","unstructured":"Kaggle: credit card fraud detection dataset; 2016. https:\/\/www.kaggle.com\/mlg-ulb\/creditcardfraud."},{"key":"821_CR12","doi-asserted-by":"crossref","unstructured":"Wang H, Khoshgoftaar TM, Napolitano A. A comparative study of ensemble feature selection techniques for software defect prediction. In: 2010 ninth international conference on machine learning and applications. New York: IEEE; 2010. p. 135\u201340.","DOI":"10.1109\/ICMLA.2010.27"},{"issue":"3","key":"821_CR13","first-page":"3343","volume":"12","author":"C Sailaja","year":"2021","unstructured":"Sailaja C, Teja GSSK, Mahesh G, Reddy PRS. Detection of fraudulent medicare providers using decision tree and logistic regression models. J Cardiovasc Dis Res. 2021;12(3):3343\u201352.","journal-title":"J Cardiovasc Dis Res"},{"key":"821_CR14","doi-asserted-by":"crossref","unstructured":"Bekkar M, Djemaa HK, Alitouche TA. Evaluation measures for models assessment over imbalanced data sets. J Inf Eng Appl. 2013;3(10):27\u201338.","DOI":"10.5121\/ijdkp.2013.3402"},{"issue":"3","key":"821_CR15","doi-asserted-by":"publisher","first-page":"96","DOI":"10.14445\/22315381\/IJETT-V69I3P216","volume":"69","author":"RY Gupta","year":"2021","unstructured":"Gupta RY, Mudigonda SS, Baruah PK. A comparative study of using various machine learning and deep learning-based fraud detection models for universal health coverage schemes. Int J Eng Trends Technol. 2021;69(3):96\u2013102.","journal-title":"Int J Eng Trends Technol"},{"issue":"1","key":"821_CR16","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-018-0138-3","volume":"5","author":"M Herland","year":"2018","unstructured":"Herland M, Khoshgoftaar TM, Bauder RA. Big data fraud detection using multiple medicare data sources. J Big Data. 2018;5(1):1\u201321.","journal-title":"J Big Data"},{"key":"821_CR17","unstructured":"The centers for medicare and medicaid services: medicare physician & other practitioners\u2014by provider and service; 2021. https:\/\/data.cms.gov\/provider-summary-by-type-of-service\/medicare-physician-other-practitioners\/medicare-physician-other-practitioners-by-provider-and-service."},{"key":"821_CR18","unstructured":"The Centers for Medicare and Medicaid Services: medicare part D prescribers\u2014by provider and drug; 2021. https:\/\/data.cms.gov\/provider-summary-by-type-of-service\/medicare-part-d-prescribers\/medicare-part-d-prescribers-by-provider-and-drug."},{"key":"821_CR19","unstructured":"The Centers for Medicare and Medicaid Services: medicare durable medical equipment, devices & supplies\u2014by referring provider and service; 2021. https:\/\/data.cms.gov\/provider-summary-by-type-of-service\/medicare-durable-medical-equipment-devices-supplies\/medicare-durable-medical-equipment-devices-supplies-by-referring-provider-and-service."},{"issue":"4","key":"821_CR20","doi-asserted-by":"publisher","first-page":"389","DOI":"10.1007\/s42979-023-01809-x","volume":"4","author":"JM Johnson","year":"2023","unstructured":"Johnson JM, Khoshgoftaar TM. Data-centric ai for healthcare fraud detection. SN Comput Sci. 2023;4(4):389.","journal-title":"SN Comput Sci"},{"key":"821_CR21","unstructured":"The Centers for Medicare and Medicaid Services: medicare physician & other practitioners\u2014by provider data dictionary; 2021. https:\/\/data.cms.gov\/resources\/medicare-physician-other-practitioners-by-provider-data-dictionary."},{"key":"821_CR22","unstructured":"The Centers for Medicare and Medicaid Services: medicare physician & other practitioners\u2014by provider; 2021. https:\/\/data.cms.gov\/provider-summary-by-type-of-service\/medicare-physician-other-practitioners\/medicare-physician-other-practitioners-by-provider."},{"key":"821_CR23","unstructured":"The Centers for Medicare and Medicaid Services: medicare part D prescribers\u2014by provider and drug data dictionary. https:\/\/data.cms.gov\/resources\/medicare-part-d-prescribers-by-provider-and-drug-data-dictionary 2021."},{"key":"821_CR24","unstructured":"The Centers for Medicare and Medicaid Services: medicare part D prescribers\u2014by provider data dictionary; 2020. https:\/\/data.cms.gov\/resources\/medicare-part-d-prescribers-by-provider-data-dictionary."},{"key":"821_CR25","unstructured":"The Centers for Medicare and Medicaid Services: medicare part D prescribers\u2014by provider; 2021. https:\/\/data.cms.gov\/provider-summary-by-type-of-service\/medicare-part-d-prescribers\/medicare-part-d-prescribers-by-provider."},{"key":"821_CR26","unstructured":"The Centers for Medicare and Medicaid Services: medicare physician & other practitioners\u2014by provider and service data dictionary; 2021. https:\/\/data.cms.gov\/resources\/medicare-physician-other-practitioners-by-provider-and-service-data-dictionary."},{"key":"821_CR27","doi-asserted-by":"crossref","unstructured":"Bauder RA, Khoshgoftaar TM. A novel method for fraudulent medicare claims detection from expected payment deviations (application paper). In: 2016 IEEE 17th international conference on information reuse and integration (IRI). New York: IEEE; 2016. p. 11\u20139.","DOI":"10.1109\/IRI.2016.11"},{"key":"821_CR28","doi-asserted-by":"crossref","unstructured":"Chen T, Guestrin C. Xgboost: a scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining\u2014KDD \u201916; 2016.","DOI":"10.1145\/2939672.2939785"},{"key":"821_CR29","first-page":"3146","volume":"30","author":"G Ke","year":"2017","unstructured":"Ke G, Meng Q, Finley T, Wang T, Chen W, Ma W, Ye Q, Liu T-Y. Lightgbm: a highly efficient gradient boosting decision tree. Adv Neural Inf Process Syst. 2017;30:3146\u201354.","journal-title":"Adv Neural Inf Process Syst"},{"issue":"1","key":"821_CR30","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1007\/s10994-006-6226-1","volume":"63","author":"P Geurts","year":"2006","unstructured":"Geurts P, Ernst D, Wehenkel L. Extremely randomized trees. Mach Learn. 2006;63(1):3\u201342.","journal-title":"Mach Learn"},{"issue":"1","key":"821_CR31","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman L. Random forests. Mach Learn. 2001;45(1):5\u201332.","journal-title":"Mach Learn"},{"key":"821_CR32","unstructured":"Prokhorenkova L, Gusev G, Vorobev A, Dorogush AV, Gulin A. Catboost: unbiased boosting with categorical features. In: Advances in neural information processing systems.  2018. Vol. 31, p. 2\u201311."},{"issue":"1","key":"821_CR33","first-page":"191","volume":"41","author":"S Le Cessie","year":"1992","unstructured":"Le Cessie S, Van Houwelingen JC. Ridge estimators in logistic regression. J R Stat Soc Ser C (Appl Stat). 1992;41(1):191\u2013201.","journal-title":"J R Stat Soc Ser C (Appl Stat)"},{"key":"821_CR34","volume-title":"Classification and regression trees","author":"L Breiman","year":"1984","unstructured":"Breiman L, Friedman J, Stone CJ, Olshen RA. Classification and regression trees. Taylor & Francis; 1984."},{"key":"821_CR35","doi-asserted-by":"publisher","first-page":"1189","DOI":"10.1214\/aos\/1013203451","volume":"29","author":"JH Friedman","year":"2001","unstructured":"Friedman JH. Greedy function approximation: a gradient boosting machine. Ann Stat. 2001;29:1189\u2013232.","journal-title":"Ann Stat"},{"issue":"4","key":"821_CR36","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s42979-021-00655-z","volume":"2","author":"JT Hancock","year":"2021","unstructured":"Hancock JT, Khoshgoftaar TM. Gradient boosted decision tree algorithms for Medicare fraud detection. SN Comput Sci. 2021;2(4):1\u201312.","journal-title":"SN Comput Sci"},{"key":"821_CR37","doi-asserted-by":"crossref","unstructured":"Leevy JL, Hancock JT, Zuech R, Khoshgoftaar TM. Detecting cybersecurity attacks using different network features with lightgbm and xgboost learners. In: 2020 IEEE second international conference on cognitive machine intelligence (CogMI). New York: IEEE; 2020. p. 190\u20137.","DOI":"10.1109\/CogMI50398.2020.00032"},{"issue":"1","key":"821_CR38","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-020-00369-8","volume":"7","author":"JT Hancock","year":"2020","unstructured":"Hancock JT, Khoshgoftaar TM. Catboost for big data: an interdisciplinary review. J big data. 2020;7(1):1\u201345.","journal-title":"J big data"},{"issue":"2","key":"821_CR39","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1007\/BF00058655","volume":"24","author":"L Breiman","year":"1996","unstructured":"Breiman L. Bagging predictors. Mach Learn. 1996;24(2):123\u201340.","journal-title":"Mach Learn"},{"key":"821_CR40","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1201\/9780429246593","volume-title":"An introduction to the bootstrap","author":"B Efron","year":"1994","unstructured":"Efron B, Tibshirani RJ. An introduction to the bootstrap. Boca Raton: CRC Press; 1994. p. 5\u20136."},{"key":"821_CR41","doi-asserted-by":"crossref","unstructured":"Hancock JT, Khoshgoftaar TM, Johnson JM. A comparative approach to threshold optimization for classifying imbalanced data. In: The international conference on collaboration and internet computing (CIC). New York: IEEE; 2022.","DOI":"10.1109\/CIC56439.2022.00028"},{"key":"821_CR42","doi-asserted-by":"crossref","unstructured":"Gu Q, Cai Z, Zhu L, Huang B. Data mining on imbalanced data sets. In: 2008 international conference on advanced computer theory and engineering. New York: IEEE; 2008. p. 1020\u20131024.","DOI":"10.1109\/ICACTE.2008.26"},{"issue":"2","key":"821_CR43","doi-asserted-by":"publisher","first-page":"215","DOI":"10.1007\/s13748-019-00172-4","volume":"8","author":"LI Kuncheva","year":"2019","unstructured":"Kuncheva LI, Arnaiz-Gonzalez A, D\u00edez-Pastor J-F, Gunn IA. Instance selection improves geometric mean accuracy: a study on imbalanced data classification. Progr Artif Intell. 2019;8(2):215\u201328.","journal-title":"Progr Artif Intell"},{"issue":"1","key":"821_CR44","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s12864-019-6413-7","volume":"21","author":"D Chicco","year":"2020","unstructured":"Chicco D, Jurman G. The advantages of the Matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation. BMC Genom. 2020;21(1):1\u201313.","journal-title":"BMC Genom"},{"key":"821_CR45","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-84858-7","volume-title":"The elements of statistical learning: data mining, inference, and prediction","author":"T Hastie","year":"2009","unstructured":"Hastie T, Tibshirani R, Friedman JH, Friedman JH. The elements of statistical learning: data mining, inference, and prediction, vol. 2. Heidelberg: Springer; 2009."},{"issue":"3","key":"821_CR46","doi-asserted-by":"publisher","first-page":"379","DOI":"10.1002\/j.1538-7305.1948.tb01338.x","volume":"27","author":"CE Shannon","year":"1948","unstructured":"Shannon CE. A mathematical theory of communication. Bell Syst Tech J. 1948;27(3):379\u2013423.","journal-title":"Bell Syst Tech J"},{"key":"821_CR47","doi-asserted-by":"publisher","DOI":"10.4135\/9781412983327","volume-title":"Analysis of variance","author":"GR Iversen","year":"1987","unstructured":"Iversen GR, Norpoth H. Analysis of variance, vol. 1. Newbury Park: Sage; 1987."},{"key":"821_CR48","doi-asserted-by":"publisher","first-page":"99","DOI":"10.2307\/3001913","volume":"5","author":"JW Tukey","year":"1949","unstructured":"Tukey JW. Comparing individual means in the analysis of variance. Biometrics. 1949;5:99\u2013114.","journal-title":"Biometrics"},{"key":"821_CR49","volume-title":"Data mining: practical machine learning tools and techniques. The Morgan Kaufmann series in data management systems","author":"IH Witten","year":"2011","unstructured":"Witten IH, Frank E, Hall MA. Data mining: practical machine learning tools and techniques. The Morgan Kaufmann series in data management systems. Pittsburgh: Elsevier Science; 2011."},{"key":"821_CR50","volume-title":"Python 3 reference manual createspace","author":"G Van Rossum","year":"2009","unstructured":"Van Rossum G, Drake F. Python 3 reference manual createspace. Scotts Valley; 2009."},{"key":"821_CR51","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V. Scikit-learn: machine learning in python. J Mach Learn Res. 2011;12:2825\u201330.","journal-title":"J Mach Learn Res"},{"key":"821_CR52","doi-asserted-by":"crossref","unstructured":"Calvert CL, Khoshgoftaar TM. Threshold based optimization of performance metrics with severely imbalanced big security data. In: 2019 IEEE 31st international conference on tools with artificial intelligence (ICTAI). New York: IEEE; 2019. p. 1328\u201334.","DOI":"10.1109\/ICTAI.2019.00184"}],"container-title":["Journal of Big Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-023-00821-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s40537-023-00821-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-023-00821-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,18]],"date-time":"2023-11-18T18:05:11Z","timestamp":1700330711000},"score":1,"resource":{"primary":{"URL":"https:\/\/journalofbigdata.springeropen.com\/articles\/10.1186\/s40537-023-00821-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,9]]},"references-count":52,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["821"],"URL":"https:\/\/doi.org\/10.1186\/s40537-023-00821-5","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-3076353\/v1","asserted-by":"object"}]},"ISSN":["2196-1115"],"issn-type":[{"value":"2196-1115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,9]]},"assertion":[{"value":"17 June 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 August 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 October 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"154"}}