{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T12:00:58Z","timestamp":1783512058909,"version":"3.55.0"},"reference-count":50,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,5,27]],"date-time":"2021-05-27T00:00:00Z","timestamp":1622073600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,5,27]],"date-time":"2021-05-27T00:00:00Z","timestamp":1622073600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Big Data"],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Class imbalance is an important consideration for cybersecurity and machine learning. We explore classification performance in detecting web attacks in the recent CSE-CIC-IDS2018 dataset. This study considers a total of eight random undersampling (RUS) ratios: no sampling, 999:1, 99:1, 95:5, 9:1, 3:1, 65:35, and 1:1. Additionally, seven different classifiers are employed: Decision Tree (DT), Random Forest (RF), CatBoost (CB), LightGBM (LGB), XGBoost (XGB), Naive Bayes (NB), and Logistic Regression (LR). For classification performance metrics, Area Under the Receiver Operating Characteristic Curve (AUC) and Area Under the Precision-Recall Curve (AUPRC) are both utilized to answer the following three research questions. The first question asks: \u201cAre various random undersampling ratios statistically different from each other in detecting web attacks?\u201d The second question asks: \u201cAre different classifiers statistically different from each other in detecting web attacks?\u201d And, our third question asks: \u201cIs the interaction between different classifiers and random undersampling ratios significant for detecting web attacks?\u201d Based on our experiments, the answers to all three research questions is \u201cYes\u201d. To the best of our knowledge, we are the first to apply random undersampling techniques to web attacks from the CSE-CIC-IDS2018 dataset while exploring various sampling ratios.<\/jats:p>","DOI":"10.1186\/s40537-021-00460-8","type":"journal-article","created":{"date-parts":[[2021,5,27]],"date-time":"2021-05-27T19:02:52Z","timestamp":1622142172000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":73,"title":["Detecting web attacks using random undersampling and ensemble learners"],"prefix":"10.1186","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5526-1094","authenticated-orcid":false,"given":"Richard","family":"Zuech","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"John","family":"Hancock","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Taghi M.","family":"Khoshgoftaar","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,5,27]]},"reference":[{"key":"460_CR1","unstructured":"Young J. US Ecommerce Sales Grow 14.9% in 2019. Accessed: 2020-11-28. https:\/\/www.digitalcommerce360.com\/article\/us-ecommerce-sales\/"},{"issue":"11","key":"460_CR2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s42452-020-03559-4","volume":"2","author":"P Radanliev","year":"2020","unstructured":"Radanliev P, De Roure D, Walton R, Van Kleek M, Montalvo RM, Santos O, Burnap P, Anthi E, et al. Artificial intelligence and machine learning in dynamic cyber risk analytics at the edge. SN Applied Sciences. 2020;2(11):1\u20138.","journal-title":"SN Applied Sciences"},{"issue":"1","key":"460_CR3","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-021-00426-w","volume":"8","author":"JL Leevy","year":"2021","unstructured":"Leevy JL, Hancock J, Zuech R, Khoshgoftaar TM. Detecting cybersecurity attacks across different network features and learners. Journal of Big Data. 2021;8(1):1\u201329.","journal-title":"Journal of Big Data"},{"key":"460_CR4","doi-asserted-by":"crossref","unstructured":"Wald R, Villanustre F, Khoshgoftaar TM, Zuech R, Robinson J, Muharemagic E. Using feature selection and classification to build effective and efficient firewalls. In: Proceedings of the 2014 IEEE 15th International Conference on Information Reuse and Integration (IEEE IRI 2014), 2014;pp. 850\u2013854 . IEEE","DOI":"10.1109\/IRI.2014.7051979"},{"issue":"01","key":"460_CR5","doi-asserted-by":"publisher","first-page":"1650001","DOI":"10.1142\/S0218539316500017","volume":"23","author":"MM Najafabadi","year":"2016","unstructured":"Najafabadi MM, Khoshgoftaar TM, Seliya N. Evaluating feature selection methods for network intrusion detection with kyoto data. International Journal of Reliability, Quality and Safety Engineering. 2016;23(01):1650001.","journal-title":"International Journal of Reliability, Quality and Safety Engineering"},{"key":"460_CR6","unstructured":"Amit I, Matherly J, Hewlett W, Xu Z., Meshi Y, Weinberger Y. Machine learning in cyber-security-problems, challenges and data sets. arXiv preprint arXiv:1812.07858 2018."},{"key":"460_CR7","doi-asserted-by":"crossref","unstructured":"Sharafaldin I, Lashkari AH, Ghorbani AA. Toward generating a new intrusion detection dataset and intrusion traffic characterization. In: ICISSP, 2018;pp. 108\u2013116.","DOI":"10.5220\/0006639801080116"},{"key":"460_CR8","unstructured":"CICIDS2017 Dataset. Accessed: 2020-08-28. https:\/\/www.unb.ca\/cic\/datasets\/ids-2017.html"},{"key":"460_CR9","unstructured":"CSE-CIC-IDS2018 Dataset. Accessed: 2020-08-28. https:\/\/www.unb.ca\/cic\/datasets\/ids-2018.html"},{"issue":"1","key":"460_CR10","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-019-0278-0","volume":"7","author":"JL Leevy","year":"2020","unstructured":"Leevy JL, Khoshgoftaar TM. A survey and analysis of intrusion detection models based on cse-cic-ids2018 big data. J Big Data. 2020;7(1):1\u201319.","journal-title":"J Big Data"},{"issue":"4","key":"460_CR11","first-page":"1","volume":"9","author":"RB Basnet","year":"2019","unstructured":"Basnet RB, Shash R, Johnson C, Walgren L, Doleck T. Towards detecting and classifying network intrusion traffic using deep learning frameworks. J Internet Serv Inf Secur. 2019;9(4):1\u201317.","journal-title":"J Internet Serv Inf Secur"},{"key":"460_CR12","doi-asserted-by":"crossref","unstructured":"Atefinia R, Ahmadi M. Network intrusion detection using multi-architectural modular deep neural network. J Supercomput 2020;1\u201323.","DOI":"10.1007\/s11227-020-03410-y"},{"key":"460_CR13","doi-asserted-by":"publisher","first-page":"101851","DOI":"10.1016\/j.cose.2020.101851","volume":"95","author":"X Li","year":"2020","unstructured":"Li X, Chen W, Zhang Q, Wu L. Building auto-encoder intrusion detection system based on random forest feature selection. Comput Secur. 2020;95:101851.","journal-title":"Comput Secur"},{"key":"460_CR14","first-page":"102564","volume":"54","author":"L D'hooge","year":"2020","unstructured":"D'hooge L, Wauters T, Volckaert B, De Turck F. Inter-dataset generalization strength of supervised machine learning methods for intrusion detection. J Inf Secur Appl. 2020;54:102564.","journal-title":"J Inf Secur Appl"},{"key":"460_CR15","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1214\/09-SS054","volume":"4","author":"S Arlot","year":"2010","unstructured":"Arlot S, Celisse A, et al. A survey of cross-validation procedures for model selection. Stat Surv. 2010;4:40\u201379.","journal-title":"Stat Surv"},{"issue":"1","key":"460_CR16","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1145\/1882471.1882479","volume":"12","author":"G Forman","year":"2010","unstructured":"Forman G, Scholz M. Apples-to-apples in cross-validation studies: pitfalls in classifier performance measurement. Acm Sigkdd Explor Newslett. 2010;12(1):49\u201357.","journal-title":"Acm Sigkdd Explor Newslett"},{"key":"460_CR17","unstructured":"Kohavi R et al. A study of cross-validation and bootstrap for accuracy estimation and model selection. In: Ijcai. Montreal, Canada; 1995. vol. 14, p. 1137\u20131145"},{"key":"460_CR18","unstructured":"Scikit-learn website. https:\/\/scikit-learn.org\/stable\/. Accessed 30 Jan 2021."},{"issue":"6","key":"460_CR19","first-page":"275","volume":"18","author":"AJ Myles","year":"2004","unstructured":"Myles AJ, Feudale RN, Liu Y, Woody NA, Brown SD. An introduction to decision tree modeling. J Chemom J Chemom Soc. 2004;18(6):275\u201385.","journal-title":"J Chemom J Chemom Soc"},{"issue":"1","key":"460_CR20","doi-asserted-by":"publisher","first-page":"77","DOI":"10.1023\/B:AMAI.0000018580.96245.c6","volume":"41","author":"LE Raileanu","year":"2004","unstructured":"Raileanu LE, Stoffel K. Theoretical comparison between the gini index and information gain criteria. Ann Math Artif Intell. 2004;41(1):77\u201393.","journal-title":"Ann Math Artif Intell"},{"key":"460_CR21","doi-asserted-by":"crossref","unstructured":"Khoshgoftaar TM, Golawala M, Van Hulse J. An empirical study of learning from imbalanced data using random forest. In: 19th IEEE international conference on tools with artificial intelligence (ICTAI 2007). IEEE; 2007. vol. 2, p. 310\u2013317","DOI":"10.1109\/ICTAI.2007.46"},{"issue":"1","key":"460_CR22","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman L. Random forests. Mach Learn. 2001;45(1):5\u201332.","journal-title":"Mach Learn"},{"issue":"2","key":"460_CR23","first-page":"123","volume":"24","author":"L Breiman","year":"1996","unstructured":"Breiman L. Bagging predictors. Mach Learn. 1996;24(2):123\u201340.","journal-title":"Mach Learn"},{"issue":"1","key":"460_CR24","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-019-0278-0","volume":"7","author":"JT Hancock","year":"2020","unstructured":"Hancock JT, Khoshgoftaar TM. Catboost for big data: an interdisciplinary review. J Big Data. 2020;7(1):1\u201345.","journal-title":"J Big Data"},{"key":"460_CR25","unstructured":"Prokhorenkova L, Gusev G, Vorobev A, Dorogush AV, Gulin A. Catboost: unbiased boosting with categorical features. In: Advances in neural information processing systems, 2018. pp. 6638\u20136648."},{"key":"460_CR26","unstructured":"LightGBM GitHub website. https:\/\/github.com\/microsoft\/LightGBM. Accessed 28 Aug 2020."},{"key":"460_CR27","doi-asserted-by":"publisher","first-page":"21","DOI":"10.3389\/fnbot.2013.00021","volume":"7","author":"A Natekin","year":"2013","unstructured":"Natekin A, Knoll A. Gradient boosting machines, a tutorial. Front Neurorobot. 2013;7:21.","journal-title":"Front Neurorobot"},{"key":"460_CR28","unstructured":"Ke G, Meng Q, Finley T, Wang T, Chen W, Ma W, Ye Q, Liu T-Y. Lightgbm: A highly efficient gradient boosting decision tree. In: Advances in neural information processing systems, 2017. pp. 3146\u20133154."},{"key":"460_CR29","doi-asserted-by":"crossref","unstructured":"Chen T, Guestrin C. Xgboost: A scalable tree boosting system. In: Proceedings of the 22nd Acm Sigkdd international conference on knowledge discovery and data mining, 2016. pp. 785\u2013794.","DOI":"10.1145\/2939672.2939785"},{"key":"460_CR30","unstructured":"Guo C, Berkhahn F. Entity embeddings of categorical variables. arXiv preprint arXiv:1604.06737 2016."},{"key":"460_CR31","unstructured":"Naive Bayes scikit-learn documentation. https:\/\/scikit-learn.org\/stable\/modules\/naive_bayes.html. Accessed 28 Aug 2020."},{"key":"460_CR32","volume-title":"Bayes theory","author":"JA Hartigan","year":"2012","unstructured":"Hartigan JA. Bayes theory. Berlin\/Heidelberg: Springer; 2012."},{"key":"460_CR33","unstructured":"sklearn.linear\\_model.LogisticRegression scikit-learn documentation. https:\/\/scikit-learn.org\/stable\/modules\/generated\/sklearn.linear_model.LogisticRegression.html. Accessed 28 Aug 2020"},{"key":"460_CR34","volume-title":"Introduction to linear regression analysis","author":"DC Montgomery","year":"2012","unstructured":"Montgomery DC, Peck EA, Vining GG. Introduction to linear regression analysis, vol. 821. Hoboken, NJ: Wiley; 2012."},{"issue":"1","key":"460_CR35","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1002\/isaf.1460","volume":"27","author":"S Lahmiri","year":"2020","unstructured":"Lahmiri S, Bekiros S, Giakoumelou A, Bezzina F. Performance assessment of ensemble learning systems in financial data classification. Intell Syst Account Financ Manag. 2020;27(1):3\u20139.","journal-title":"Intell Syst Account Financ Manag"},{"key":"460_CR36","unstructured":"Kaggle competitions website. https:\/\/www.kaggle.com\/competitions. Accessed 30 Jan 2021."},{"issue":"7","key":"460_CR37","doi-asserted-by":"publisher","first-page":"1145","DOI":"10.1016\/S0031-3203(96)00142-2","volume":"30","author":"AP Bradley","year":"1997","unstructured":"Bradley AP. The use of the area under the roc curve in the evaluation of machine learning algorithms. Pattern Recognit. 1997;30(7):1145\u201359.","journal-title":"Pattern Recognit"},{"issue":"6","key":"460_CR38","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/cc3000","volume":"8","author":"V Bewick","year":"2004","unstructured":"Bewick V, Cheek L, Ball J. Statistics review 13: receiver operating characteristic curves. Crit Care. 2004;8(6):1\u20135.","journal-title":"Crit Care"},{"issue":"7","key":"460_CR39","doi-asserted-by":"publisher","first-page":"928","DOI":"10.1161\/CIRCULATIONAHA.106.672402","volume":"115","author":"NR Cook","year":"2007","unstructured":"Cook NR. Use and misuse of the receiver operating characteristic curve in risk prediction. Circulation. 2007;115(7):928\u201335.","journal-title":"Circulation"},{"key":"460_CR40","doi-asserted-by":"crossref","unstructured":"Davis J, Goadrich M. The relationship between precision-recall and roc curves. In: Proceedings of the 23rd international conference on machine learning; 2006. pp. 233\u2013240.","DOI":"10.1145\/1143844.1143874"},{"key":"460_CR41","doi-asserted-by":"crossref","unstructured":"Boyd K, Eng KH, Page CD. Area under the precision-recall curve: point estimates and confidence intervals. In: Joint European conference on machine learning and knowledge discovery in databases. Springer; 2013. p. 451\u2013466","DOI":"10.1007\/978-3-642-40994-3_29"},{"issue":"1","key":"460_CR42","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1186\/s40537-019-0274-4","volume":"6","author":"T Hasanin","year":"2019","unstructured":"Hasanin T, Khoshgoftaar TM, Leevy JL, Bauder RA. Severely imbalanced big data challenges: investigating data sampling approaches. J Big Data. 2019;6(1):107.","journal-title":"J Big Data"},{"key":"460_CR43","doi-asserted-by":"crossref","unstructured":"Van Hulse J, Khoshgoftaar TM, Napolitano A. Experimental perspectives on learning from imbalanced data. In: Proceedings of the 24th international conference on machine learning; 2007. pp. 935\u2013942.","DOI":"10.1145\/1273496.1273614"},{"issue":"1","key":"460_CR44","doi-asserted-by":"publisher","first-page":"9","DOI":"10.1007\/s13755-018-0051-3","volume":"6","author":"RA Bauder","year":"2018","unstructured":"Bauder RA, Khoshgoftaar TM. The effects of varying class distribution on learner behavior for medicare fraud detection with imbalanced big data. Health Inf Sci Syst. 2018;6(1):9.","journal-title":"Health Inf Sci Syst"},{"issue":"1","key":"460_CR45","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1186\/s40537-019-0230-3","volume":"6","author":"CL Calvert","year":"2019","unstructured":"Calvert CL, Khoshgoftaar TM. Impact of class distribution on the detection of slow http dos attacks using big data. J Big Data. 2019;6(1):67.","journal-title":"J Big Data"},{"key":"460_CR46","doi-asserted-by":"crossref","unstructured":"Hasanin T, Khoshgoftaar TM, Bauder RA. Impact of data sampling with severely imbalanced big data. Reuse Intell Syst 2020;1.","DOI":"10.1201\/9781003034971-1"},{"key":"460_CR47","volume-title":"Experimental designs using ANOVA","author":"BG Tabachnick","year":"2007","unstructured":"Tabachnick BG, Fidell LS. Experimental designs using ANOVA. Belmont, CA: Thomson\/Brooks\/Cole; 2007."},{"key":"460_CR48","doi-asserted-by":"crossref","unstructured":"Tukey JW. Comparing individual means in the analysis of variance. Biometrics. 1949;99\u2013114.","DOI":"10.2307\/3001913"},{"key":"460_CR49","doi-asserted-by":"crossref","unstructured":"Calvert CL, Khoshgoftaar TM. Threshold based optimization of performance metrics with severely imbalanced big security data. In: 2019 IEEE 31st international conference on tools with artificial intelligence (ICTAI). IEEE; 2019. p. 1328\u20131334","DOI":"10.1109\/ICTAI.2019.00184"},{"issue":"1","key":"460_CR50","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-020-00301-0","volume":"7","author":"T Hasanin","year":"2020","unstructured":"Hasanin T, Khoshgoftaar TM, Leevy JL, Bauder RA. Investigating class rarity in big data. J Big Data. 2020;7(1):1\u201317.","journal-title":"J Big Data"}],"container-title":["Journal of Big Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-021-00460-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s40537-021-00460-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-021-00460-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,5,27]],"date-time":"2021-05-27T19:11:08Z","timestamp":1622142668000},"score":1,"resource":{"primary":{"URL":"https:\/\/journalofbigdata.springeropen.com\/articles\/10.1186\/s40537-021-00460-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,27]]},"references-count":50,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["460"],"URL":"https:\/\/doi.org\/10.1186\/s40537-021-00460-8","relation":{},"ISSN":["2196-1115"],"issn-type":[{"value":"2196-1115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,5,27]]},"assertion":[{"value":"28 January 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 May 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 May 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"75"}}