{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T13:23:01Z","timestamp":1782739381797,"version":"3.54.5"},"reference-count":44,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,10,12]],"date-time":"2023-10-12T00:00:00Z","timestamp":1697068800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,10,12]],"date-time":"2023-10-12T00:00:00Z","timestamp":1697068800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Big Data"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Research into machine learning methods for fraud detection is of paramount importance, largely due to the substantial financial implications associated with fraudulent activities. Our investigation is centered around the Credit Card Fraud Dataset and the Medicare Part D dataset, both of which are highly imbalanced. The Credit Card Fraud Detection Dataset is large data and contains actual transactional content, which makes it an ideal benchmark for credit card fraud detection. The Medicare Part D dataset is big data, providing researchers the opportunity to examine national trends and patterns related to prescription drug usage and expenditures. This paper presents a detailed comparison of One-Class Classification (OCC) and binary classification algorithms, utilizing eight distinct classifiers. OCC is a more appealing option, since collecting a second label for binary classification can be very expensive and not possible to obtain within a reasonable time frame. We evaluate our models based on two key metrics: the Area Under the Precision-Recall Curve (AUPRC)) and the Area Under the Receiver Operating Characteristic Curve (AUC). Our results show that binary classification consistently outperforms OCC in detecting fraud within both datasets. In addition, we found that CatBoost is the most performant among the classifiers tested. Moreover, we contribute novel results by being the first to publish a performance comparison of OCC and binary classification specifically for fraud detection in the Credit Card Fraud and Medicare Part D datasets.<\/jats:p>","DOI":"10.1186\/s40537-023-00825-1","type":"journal-article","created":{"date-parts":[[2023,10,12]],"date-time":"2023-10-12T13:02:14Z","timestamp":1697115734000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":22,"title":["Investigating the effectiveness of one-class and binary classification for fraud detection"],"prefix":"10.1186","volume":"10","author":[{"given":"Joffrey L.","family":"Leevy","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"John","family":"Hancock","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Taghi M.","family":"Khoshgoftaar","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Azadeh","family":"Abdollah Zadeh","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,10,12]]},"reference":[{"key":"825_CR1","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-021-00447-5","volume":"8","author":"Z Salekshahrezaee","year":"2021","unstructured":"Salekshahrezaee Z, Leevy JL, Khoshgoftaar TM. A reconstruction error-based framework for label noise detection. J Big Data. 2021;8:1\u201316.","journal-title":"J Big Data"},{"key":"825_CR2","doi-asserted-by":"crossref","unstructured":"Bauder RA, Khoshgoftaar TM, Hasanin T. Data sampling approaches with severely imbalanced big data for medicare fraud detection. In: 2018 IEEE 30th International Conference on Tools with Artificial Intelligence (ICTAI), pp. 137\u2013142 2018;. IEEE","DOI":"10.1109\/ICTAI.2018.00030"},{"issue":"1","key":"825_CR3","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-020-00301-0","volume":"7","author":"T Hasanin","year":"2020","unstructured":"Hasanin T, Khoshgoftaar TM, Leevy JL, Bauder RA. Investigating class rarity in big data. J Big Data. 2020;7(1):1\u201317.","journal-title":"J Big Data"},{"issue":"9","key":"825_CR4","doi-asserted-by":"publisher","first-page":"1263","DOI":"10.1109\/TKDE.2008.239","volume":"21","author":"H He","year":"2009","unstructured":"He H, Garcia EA. Learning from imbalanced data. IEEE Trans Knowl Data Eng. 2009;21(9):1263\u201384.","journal-title":"IEEE Trans Knowl Data Eng"},{"issue":"1","key":"825_CR5","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-021-00514-x","volume":"8","author":"N Seliya","year":"2021","unstructured":"Seliya N, Abdollah Zadeh A, Khoshgoftaar TM. A literature review on one-class classification and its potential applications in big data. J Big Data. 2021;8(1):1\u201331.","journal-title":"J Big Data"},{"key":"825_CR6","unstructured":"Kaggle: Credit Card Fraud Detection. https:\/\/www.kaggle.com\/mlg-ulb\/creditcardfraud (2018)."},{"issue":"4","key":"825_CR7","doi-asserted-by":"publisher","first-page":"389","DOI":"10.1007\/s42979-023-01809-x","volume":"4","author":"JM Johnson","year":"2023","unstructured":"Johnson JM, Khoshgoftaar TM. Data-centric ai for healthcare fraud detection. SN Comp Sci. 2023;4(4):389.","journal-title":"SN Comp Sci"},{"key":"825_CR8","unstructured":"of Enterprise\u00a0Data, C.O., Analytics: Medicare Fee-For Service Provider Utilization & Payment Data Part D prescriber public use file: a methodological overview. https:\/\/www.cms.gov\/Research-Statistics-Data-and-Systems\/Statistics-Trends-and-Reports\/Medicare-Provider-Charge-Data\/Downloads\/Prescriber_Methods.pdf."},{"issue":"1","key":"825_CR9","doi-asserted-by":"publisher","first-page":"29","DOI":"10.1186\/s40537-018-0138-3","volume":"5","author":"M Herland","year":"2018","unstructured":"Herland M, Khoshgoftaar TM, Bauder RA. Big data fraud detection using multiple medicare data sources. J Big Data. 2018;5(1):29.","journal-title":"J Big Data"},{"key":"825_CR10","doi-asserted-by":"crossref","unstructured":"Hancock J, Khoshgoftaar TM. Medicare fraud detection using catboost. In: 2020 IEEE 21st International Conference on Information Reuse and Integration for Data Science (IRI), pp. 97\u2013103 2020;. IEEE Computer Society","DOI":"10.1109\/IRI49571.2020.00022"},{"key":"825_CR11","doi-asserted-by":"crossref","unstructured":"Hancock J, Khoshgoftaar TM, Johnson JM. The effects of random undersampling for big data medicare fraud detection. In: 2022 IEEE International Conference on Service-Oriented System Engineering (SOSE), pp. 141\u2013146 2022;. IEEE.","DOI":"10.1109\/SOSE55356.2022.00023"},{"key":"825_CR12","doi-asserted-by":"crossref","unstructured":"Kumar MS, Soundarya V, Kavitha S, Keerthika E, Aswini E. Credit card fraud detection using random forest algorithm. In: 2019 3rd International Conference on Computing and Communications Technologies (ICCCT), pp. 149\u2013153 2019;. IEEE","DOI":"10.1109\/ICCCT2.2019.8824930"},{"key":"825_CR13","doi-asserted-by":"crossref","unstructured":"Hancock J, Khoshgoftaar TM. Performance of catboost and xgboost in medicare fraud detection. In: 19th IEEE International Conference On Machine Learning And Applications (ICMLA) 2020;. IEEE.","DOI":"10.1109\/ICMLA51294.2020.00095"},{"key":"825_CR14","doi-asserted-by":"crossref","unstructured":"Alenzi HZ, Aljehane NO. Fraud detection in credit cards using logistic regression. International Journal of Advanced Computer Science and Applications. 2020. 11(12).","DOI":"10.14569\/IJACSA.2020.0111265"},{"key":"825_CR15","doi-asserted-by":"crossref","unstructured":"Najafabadi MM, Khoshgoftaar TM, Calvert C, Kemp C. A text mining approach for anomaly detection in application layer ddos attacks. In: The Thirtieth International Flairs Conference 2017.","DOI":"10.1109\/IRI.2017.44"},{"issue":"15","key":"825_CR16","doi-asserted-by":"publisher","first-page":"17073","DOI":"10.1007\/s10489-021-02671-1","volume":"52","author":"T Hayashi","year":"2022","unstructured":"Hayashi T, Fujita H. One-class ensemble classifier for data imbalance problems. Appl Intell. 2022;52(15):17073\u201389.","journal-title":"Appl Intell"},{"issue":"1","key":"825_CR17","doi-asserted-by":"publisher","first-page":"118","DOI":"10.1186\/s40537-023-00794-5","volume":"10","author":"JL Leevy","year":"2023","unstructured":"Leevy JL, Hancock J, Khoshgoftaar TM. Comparative analysis of binary and one-class classification techniques for credit card fraud data. J Big Data. 2023;10(1):118.","journal-title":"J Big Data"},{"issue":"1","key":"825_CR18","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-023-00724-5","volume":"10","author":"JT Hancock","year":"2023","unstructured":"Hancock JT, Khoshgoftaar TM, Johnson JM. Evaluating classifier performance with highly imbalanced big data. J Big Data. 2023;10(1):1\u201331.","journal-title":"J Big Data"},{"key":"825_CR19","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.eswa.2021.114750","volume":"175","author":"Z Li","year":"2021","unstructured":"Li Z, Huang M, Liu G, Jiang C. A hybrid method with dynamic weighted entropy for handling the problem of class imbalance with overlap in credit card fraud detection. Expert Syst Appl. 2021;175:1\u201310.","journal-title":"Expert Syst Appl"},{"key":"825_CR20","doi-asserted-by":"crossref","unstructured":"Jeragh M, AlSulaimi M. Combining auto encoders and one class support vectors machine for fraudulant credit card transactions detection. In: 2018 Second World Conference on Smart Trends in Systems, Security and Sustainability (WorldS4), pp. 178\u2013184 2018;. IEEE.","DOI":"10.1109\/WorldS4.2018.8611624"},{"key":"825_CR21","first-page":"42","volume":"4","author":"A Chandorkar","year":"2022","unstructured":"Chandorkar A. Credit card fraud detection using machine learning. Int Res J Moderniz Eng Technol Sci. 2022;4:42\u201350.","journal-title":"Int Res J Moderniz Eng Technol Sci"},{"key":"825_CR22","doi-asserted-by":"publisher","first-page":"1","DOI":"10.14445\/22312803\/IJCTT-V69I8P101","volume":"69","author":"H Bodepudi","year":"2021","unstructured":"Bodepudi H. Credit card fraud detection using unsupervised machine learning algorithms. Int J Comput Trends Technol. 2021;69:1\u201313.","journal-title":"Int J Comput Trends Technol"},{"issue":"2","key":"825_CR23","first-page":"394","volume":"6","author":"S Ounacer","year":"2018","unstructured":"Ounacer S, El Bour HA, Oubrahim Y, Ghoumari MY, Azzouazi M. Using isolation forest in anomaly detection: the case of credit card transactions. Periodic Eng Nat Sci. 2018;6(2):394\u2013400.","journal-title":"Periodic Eng Nat Sci"},{"key":"825_CR24","doi-asserted-by":"crossref","unstructured":"Hancock J, Khoshgoftaar TM, Johnson JM. Informative evaluation metrics for highly imbalanced big data classification. In: 2022 21st IEEE International Conference on Machine Learning and Applications (ICMLA) 2022; IEEE.","DOI":"10.1109\/ICMLA55696.2022.00224"},{"key":"825_CR25","doi-asserted-by":"crossref","unstructured":"Raza M, Qayyum U. Classical and deep learning classifiers for anomaly detection. In: 2019 16th International Bhurban Conference on Applied Sciences and Technology (IBCAST), pp. 614\u2013618 2019; IEEE.","DOI":"10.1109\/IBCAST.2019.8667245"},{"key":"825_CR26","doi-asserted-by":"crossref","unstructured":"Wu T-Y, Wang Y-T. Locally interpretable one-class anomaly detection for credit card fraud detection. In: 2021 International Conference on Technologies and Applications of Artificial Intelligence (TAAI), pp. 25\u201330 2021;. IEEE.","DOI":"10.1109\/TAAI54685.2021.00014"},{"key":"825_CR27","doi-asserted-by":"crossref","unstructured":"Salekshahrezaee Z, Leevy JL, Khoshgoftaar TM. Feature extraction for class imbalance using a convolutional autoencoder and data sampling. In: 2021 IEEE 33rd International Conference on Tools with Artificial Intelligence (ICTAI), pp. 217\u2013223 2021; IEEE.","DOI":"10.1109\/ICTAI52525.2021.00037"},{"key":"825_CR28","unstructured":"The Centers for Medicare and Medicaid Services: Medicare Part D Prescribers \u2013 by Provider and Drug. https:\/\/data.cms.gov\/provider-summary-by-type-of-service\/medicare-part-d-prescribers\/medicare-part-d-prescribers-by-provider-and-drug (2021)."},{"key":"825_CR29","unstructured":"The Centers for Medicare and Medicaid Services: Medicare Part D Prescribers - by Provider. https:\/\/data.cms.gov\/provider-summary-by-type-of-service\/medicare-part-d-prescribers\/medicare-part-d-prescribers-by-provider (2021)."},{"issue":"1","key":"825_CR30","doi-asserted-by":"publisher","first-page":"109","DOI":"10.1097\/ALN.0000000000001897","volume":"128","author":"GF Chamoun","year":"2018","unstructured":"Chamoun GF, Li L, Chamoun NG, Saini V, Sessler DI. Comparison of an updated risk stratification index to hierarchical condition categories. Anesthesiology. 2018;128(1):109\u201316.","journal-title":"Anesthesiology"},{"key":"825_CR31","unstructured":"OIG: Office of Inspector General Exclusion Authorities US Department of Health and Human Services. https:\/\/oig.hhs.gov\/."},{"key":"825_CR32","doi-asserted-by":"publisher","first-page":"3571","DOI":"10.1016\/j.matpr.2021.11.635","volume":"56","author":"JS Kushwah","year":"2022","unstructured":"Kushwah JS, Kumar A, Patel S, Soni R, Gawande A, Gupta S. Comparative study of regressor and classifier with decision tree using modern tools. Mat Today Proc. 2022;56:3571\u20136.","journal-title":"Mat Today Proc"},{"issue":"1","key":"825_CR33","first-page":"41","volume":"11","author":"SM Basha","year":"2018","unstructured":"Basha SM, Rajput DS, Vandhan V. Impact of gradient ascent and boosting algorithm in classification. Int J Intell Eng Syst (IJIES). 2018;11(1):41\u20139.","journal-title":"Int J Intell Eng Syst (IJIES)"},{"key":"825_CR34","unstructured":"Prokhorenkova L, Gusev G, Vorobev A, Dorogush AV, Gulin A. Catboost: unbiased boosting with categorical features. In: Advances in Neural Information Processing Systems, pp. 6638\u20136648 2018."},{"issue":"3","key":"825_CR35","doi-asserted-by":"publisher","first-page":"876","DOI":"10.1287\/moor.2016.0831","volume":"42","author":"A Gupta","year":"2017","unstructured":"Gupta A, Nagarajan V, Ravi R. Approximation algorithms for optimal decision trees and adaptive tsp problems. Mathemat Operat Res. 2017;42(3):876\u201396.","journal-title":"Mathemat Operat Res"},{"key":"825_CR36","doi-asserted-by":"publisher","first-page":"205","DOI":"10.1016\/j.inffus.2020.07.007","volume":"64","author":"S Gonz\u00e1lez","year":"2020","unstructured":"Gonz\u00e1lez S, Garc\u00eda S, Del Ser J, Rokach L, Herrera F. A practical tutorial on bagging and boosting based ensembles for machine learning: Algorithms, software tools, performance study, practical perspectives and opportunities. Inform Fusion. 2020;64:205\u201337.","journal-title":"Inform Fusion"},{"key":"825_CR37","doi-asserted-by":"publisher","first-page":"191","DOI":"10.1007\/s10994-008-5092-4","volume":"74","author":"R Kassab","year":"2009","unstructured":"Kassab R, Alexandre F. Incremental data-driven learning of a novelty detection model for one-class classification with application to high-dimensional noisy data. Mach Learn. 2009;74:191\u2013234.","journal-title":"Mach Learn"},{"key":"825_CR38","first-page":"11821","volume":"34","author":"G Sriramanan","year":"2021","unstructured":"Sriramanan G, Addepalli S, Baburaj A, et al. Towards efficient and effective adversarial training. Adv Neural Inform Proc Syst. 2021;34:11821\u201333.","journal-title":"Adv Neural Inform Proc Syst"},{"issue":"11","key":"825_CR39","doi-asserted-by":"publisher","first-page":"139","DOI":"10.1145\/3422622","volume":"63","author":"I Goodfellow","year":"2020","unstructured":"Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y. Generative adversarial networks. Commun ACM. 2020;63(11):139\u201344.","journal-title":"Commun ACM"},{"key":"825_CR40","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, et al. Scikit-learn: machine learning in python. J Mach Learning Res. 2011;12:2825\u201330.","journal-title":"J Mach Learning Res"},{"key":"825_CR41","doi-asserted-by":"crossref","unstructured":"Seliya N, Khoshgoftaar TM, Van\u00a0Hulse J. A study on the relationships of classifier performance metrics. In: Tools with Artificial Intelligence, 2009. ICTAI\u201909. 21st International Conference On, pp. 59\u201366 2009;. IEEE.","DOI":"10.1109\/ICTAI.2009.25"},{"key":"825_CR42","doi-asserted-by":"crossref","unstructured":"Davis J, Goadrich M. The relationship between precision-recall and roc curves. In: Proceedings of the 23rd International Conference on Machine Learning, pp. 233\u2013240 2006.","DOI":"10.1145\/1143844.1143874"},{"issue":"3","key":"825_CR43","first-page":"61","volume":"10","author":"J Platt","year":"1999","unstructured":"Platt J, et al. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Adv Large Margin Classif. 1999;10(3):61\u201374.","journal-title":"Adv Large Margin Classif"},{"key":"825_CR44","doi-asserted-by":"crossref","unstructured":"Zadrozny B, Elkan C. Transforming classifier scores into accurate multiclass probability estimates. In: Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 694\u2013699 2002.","DOI":"10.1145\/775047.775151"}],"container-title":["Journal of Big Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-023-00825-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s40537-023-00825-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-023-00825-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,18]],"date-time":"2023-11-18T15:06:29Z","timestamp":1700319989000},"score":1,"resource":{"primary":{"URL":"https:\/\/journalofbigdata.springeropen.com\/articles\/10.1186\/s40537-023-00825-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,12]]},"references-count":44,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["825"],"URL":"https:\/\/doi.org\/10.1186\/s40537-023-00825-1","relation":{},"ISSN":["2196-1115"],"issn-type":[{"value":"2196-1115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,12]]},"assertion":[{"value":"30 June 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 September 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 October 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"157"}}