{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T15:07:53Z","timestamp":1787324873332,"version":"3.56.0"},"reference-count":45,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2022,12,15]],"date-time":"2022-12-15T00:00:00Z","timestamp":1671062400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,12,15]],"date-time":"2022-12-15T00:00:00Z","timestamp":1671062400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Big Data"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We propose a novel feature popularity framework, and introduce this new framework to the cybersecurity domain. Feature popularity has not yet been used in machine learning or data mining, and we implement it with three web attacks from the CSE-CIC-IDS2018 dataset: Brute Force, SQL Injection, and XSS web attacks. Feature popularity is based upon ensemble Feature Selection Techniques (FSTs) and allows us to more easily understand common and important features between different cyberattacks. Three filter-based and four supervised learning-based FSTs are used to generate feature subsets for each of our three different web attack datasets, and then our feature popularity frameworks are applied. Classification performance for feature popularity is mostly similar as compared to when \u201call features\u201d are evaluated (with feature popularity subsets having better performance in 5 out of 15 experiments). Our feature popularity technique effectively builds an ensemble of ensembles by first building an ensemble of FSTs for each dataset, and then building another ensemble across a dataset agreement dimension. The Jaccard similarity is also employed with our feature popularity framework in order to better identify which attack classes should (or should not) be grouped together when applying feature popularity. The four most popular features across all three web attacks from this experiment are: Flow_Bytes_s, Flow_IAT_Max, Fwd_IAT_Std, and Fwd_IAT_Total. When only using these four features as input to our models, classification performance is not seriously degraded. This feature popularity framework granted us new and previously unseen insights into the web attack detection process with CSE-CIC-IDS2018 big data, even though we had intensely studied it previously. We realized these four particular features cannot properly identify our three web attacks, as they operate mainly from the time dimension and NetFlow features from layers 3 and 4 of the OSI model. Conversely, our three web attacks operate in the application layer (7) of the OSI model and should not leave signatures in these four features. Feature popularity produces easier to explain models which provide domain experts better visibility into the problem, and can also reduce the complexity of implementing models in real-world systems.<\/jats:p>","DOI":"10.1186\/s40537-022-00661-9","type":"journal-article","created":{"date-parts":[[2022,12,15]],"date-time":"2022-12-15T04:29:07Z","timestamp":1671078547000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["A new feature popularity framework for detecting cyberattacks using popular features"],"prefix":"10.1186","volume":"9","author":[{"given":"Richard","family":"Zuech","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"John","family":"Hancock","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Taghi M.","family":"Khoshgoftaar","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,12,15]]},"reference":[{"key":"661_CR1","unstructured":"Young J. US ecommerce sales grow 14.9% in 2019. 2020. https:\/\/www.digitalcommerce360.com\/article\/us-ecommerce-sales\/. Accessed 28 Nov 2020"},{"key":"661_CR2","doi-asserted-by":"crossref","unstructured":"Saeys Y, Abeel T, Van\u00a0de Peer Y Robust feature selection using ensemble feature selection techniques. In: Joint European conference on machine learning and knowledge discovery in databases. Springer; 2008. p. 313\u2013325","DOI":"10.1007\/978-3-540-87481-2_21"},{"key":"661_CR3","doi-asserted-by":"publisher","first-page":"124","DOI":"10.1016\/j.knosys.2016.11.017","volume":"118","author":"B Seijo-Pardo","year":"2017","unstructured":"Seijo-Pardo B, Porto-D\u00edaz I, Bol\u00f3n-Canedo V, Alonso-Betanzos A. Ensemble feature selection: homogeneous and heterogeneous approaches. Knowl Based Syst. 2017;118:124\u201339.","journal-title":"Knowl Based Syst"},{"issue":"1","key":"661_CR4","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1007\/s10115-006-0040-8","volume":"12","author":"A Kalousis","year":"2007","unstructured":"Kalousis A, Prados J, Hilario M. Stability of feature selection algorithms: a study on high-dimensional spaces. Knowl Inf Syst. 2007;12(1):95\u2013116.","journal-title":"Knowl Inf Syst"},{"key":"661_CR5","doi-asserted-by":"crossref","unstructured":"Zuech R, Hancock J, Khoshgoftaar TM. Feature popularity between different web attacks with supervised feature selection rankers. In: 2021 20th IEEE international conference on machine learning and applications (ICMLA). IEEE; 2021. p. 30\u201337","DOI":"10.1109\/ICMLA52953.2021.00013"},{"key":"661_CR6","doi-asserted-by":"crossref","unstructured":"Sharafaldin I, Lashkari AH, Ghorbani AA. Toward generating a new intrusion detection dataset and intrusion traffic characterization. In: ICISSP. 2018. p. 108\u2013116","DOI":"10.5220\/0006639801080116"},{"key":"661_CR7","unstructured":"CICIDS2017 Dataset. 2020. https:\/\/www.unb.ca\/cic\/datasets\/ids-2017.html. Accessed 28 Aug 2020"},{"issue":"1","key":"661_CR8","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-018-0151-6","volume":"5","author":"JL Leevy","year":"2018","unstructured":"Leevy JL, Khoshgoftaar TM, Bauder RA, Seliya N. A survey on addressing high-class imbalance in big data. J Big Data. 2018;5(1):1\u201330.","journal-title":"J Big Data"},{"key":"661_CR9","first-page":"194","volume":"2","author":"RC Soltysik","year":"2013","unstructured":"Soltysik RC, Yarnold PR. Megaoda large sample and big data time trials: separating the chaff. Optim Data Anal. 2013;2:194\u20137.","journal-title":"Optim Data Anal"},{"issue":"2","key":"661_CR10","doi-asserted-by":"publisher","first-page":"423","DOI":"10.2308\/acch-51068","volume":"29","author":"M Cao","year":"2015","unstructured":"Cao M, Chychyla R, Stewart T. Big data analytics in financial statement audits. Account Horiz. 2015;29(2):423\u20139.","journal-title":"Account Horiz"},{"key":"661_CR11","unstructured":"OWASP Top Ten webpage. 2020. https:\/\/owasp.org\/www-project-top-ten\/. Accessed 10 Aug 2021"},{"key":"661_CR12","doi-asserted-by":"crossref","unstructured":"Sarhan M, Layeghy S, Portmann M. An explainable machine learning-based network intrusion detection system for enabling generalisability in securing iot networks. arXiv preprint arXiv:2104.07183 2021","DOI":"10.21203\/rs.3.rs-2035633\/v1"},{"issue":"1","key":"661_CR13","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-021-00426-w","volume":"8","author":"JL Leevy","year":"2021","unstructured":"Leevy JL, Hancock J, Zuech R, Khoshgoftaar TM. Detecting cybersecurity attacks across different network features and learners. J Big Data. 2021;8(1):1\u201329.","journal-title":"J Big Data"},{"key":"661_CR14","doi-asserted-by":"crossref","unstructured":"Fitni QRS, Ramli K. Implementation of ensemble learning and feature selection for performance improvements in anomaly-based intrusion detection systems. In: 2020 IEEE international conference on industry 4.0, artificial intelligence, and communications technology (IAICT). IEEE; 2020. p. 118\u2013124.","DOI":"10.1109\/IAICT50021.2020.9172014"},{"key":"661_CR15","doi-asserted-by":"publisher","first-page":"107120","DOI":"10.1016\/j.knosys.2021.107120","volume":"226","author":"M Beechey","year":"2021","unstructured":"Beechey M, Kyriakopoulos KG, Lambotharan S. Evidential classification and feature selection for cyber-threat hunting. Knowl Based Syst. 2021;226:107120.","journal-title":"Knowl Based Syst"},{"key":"661_CR16","doi-asserted-by":"crossref","unstructured":"Hua Y. An efficient traffic classification scheme using embedded feature selection and lightgbm. In: 2020 information communication technologies conference (ICTC). IEEE; 2020. p. 125\u2013130.","DOI":"10.1109\/ICTC49638.2020.9123302"},{"key":"661_CR17","doi-asserted-by":"publisher","DOI":"10.1016\/j.comnet.2020.107315","volume":"177","author":"H Zhang","year":"2020","unstructured":"Zhang H, Huang L, Wu CQ, Li Z. An effective convolutional neural network based on smote and gaussian mixture model for intrusion detection in imbalanced dataset. Comput Netw. 2020;177: 107315.","journal-title":"Comput Netw"},{"key":"661_CR18","unstructured":"CSE-CIC-IDS2018 Dataset. 2020. https:\/\/www.unb.ca\/cic\/datasets\/ids-2018.html. Accessed 28 Aug 2020"},{"key":"661_CR19","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1214\/09-SS054","volume":"4","author":"S Arlot","year":"2010","unstructured":"Arlot S, Celisse A, et al. A survey of cross-validation procedures for model selection. Stat Surv. 2010;4:40\u201379.","journal-title":"Stat Surv"},{"key":"661_CR20","unstructured":"Kohavi R, et al. A study of cross-validation and bootstrap for accuracy estimation and model selection. In: Ijcai. Montreal, Canada; 1995. p. 14, 1137\u20131145 ."},{"issue":"1","key":"661_CR21","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s40537-020-00387-6","volume":"8","author":"R Zuech","year":"2021","unstructured":"Zuech R, Hancock J, Khoshgoftaar TM. Investigating rarity in web attacks with ensemble learners. J Big Data. 2021;8(1):1\u201327.","journal-title":"J Big Data"},{"key":"661_CR22","unstructured":"Scikit-learn website. 2020. https:\/\/scikit-learn.org\/stable\/. Accessed 30 Jan 2021"},{"issue":"6","key":"661_CR23","first-page":"275","volume":"18","author":"AJ Myles","year":"2004","unstructured":"Myles AJ, Feudale RN, Liu Y, Woody NA, Brown SD. An introduction to decision tree modeling. J Chemom J Chemom Soc. 2004;18(6):275\u201385.","journal-title":"J Chemom J Chemom Soc"},{"issue":"1","key":"661_CR24","doi-asserted-by":"publisher","first-page":"77","DOI":"10.1023\/B:AMAI.0000018580.96245.c6","volume":"41","author":"LE Raileanu","year":"2004","unstructured":"Raileanu LE, Stoffel K. Theoretical comparison between the gini index and information gain criteria. Ann Math Artif Intell. 2004;41(1):77\u201393.","journal-title":"Ann Math Artif Intell"},{"issue":"1","key":"661_CR25","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman L. Random forests. Mach Learn. 2001;45(1):5\u201332.","journal-title":"Mach Learn"},{"issue":"2","key":"661_CR26","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1007\/BF00058655","volume":"24","author":"L Breiman","year":"1996","unstructured":"Breiman L. Bagging predictors. Mach Learn. 1996;24(2):123\u201340.","journal-title":"Mach Learn"},{"key":"661_CR27","unstructured":"CatBoost home page. 2020. https:\/\/catboost.ai\/. Accessed 28 Aug 2020"},{"key":"661_CR28","unstructured":"Prokhorenkova L, Gusev G, Vorobev A, Dorogush AV, Gulin A. Catboost: unbiased boosting with categorical features. In: Advances in neural information processing systems. 2018. p. 6638\u20136648"},{"key":"661_CR29","unstructured":"LightGBM GitHub website. 2020. https:\/\/github.com\/microsoft\/LightGBM. Accessed 28 Aug 2020"},{"key":"661_CR30","doi-asserted-by":"publisher","first-page":"21","DOI":"10.3389\/fnbot.2013.00021","volume":"7","author":"A Natekin","year":"2013","unstructured":"Natekin A, Knoll A. Gradient boosting machines, a tutorial. Front Neurorobot. 2013;7:21.","journal-title":"Front Neurorobot"},{"key":"661_CR31","unstructured":"Ke G, Meng Q, Finley T, Wang T, Chen W, Ma W, Ye Q, Liu T-Y. Lightgbm: a highly efficient gradient boosting decision tree. In: Advances in neural information processing systems. 2017. p. 3146\u20133154"},{"key":"661_CR32","doi-asserted-by":"crossref","unstructured":"Chen T, Guestrin C. Xgboost: a scalable tree boosting system. In: Proceedings of the 22nd Acm Sigkdd international conference on knowledge discovery and data mining. 2016. p. 785\u2013794","DOI":"10.1145\/2939672.2939785"},{"key":"661_CR33","unstructured":"Guo C, Berkhahn F. Entity embeddings of categorical variables. 2016. arXiv preprint arXiv:1604.06737"},{"key":"661_CR34","unstructured":"Scikit-learn Documentation\u2014Feature Selection. 2020. https:\/\/scikit-learn.org\/stable\/modules\/feature_selection.html. Accessed 16 Aug 2021"},{"key":"661_CR35","doi-asserted-by":"crossref","unstructured":"Zien A, Kr\u00e4mer N, Sonnenburg S, R\u00e4tsch G. The feature importance ranking measure. In: Joint European conference on machine learning and knowledge discovery in databases. Springer; 2009. p. 694\u2013709 .","DOI":"10.1007\/978-3-642-04174-7_45"},{"key":"661_CR36","unstructured":"Scikit-learn Documentation - chi2 Feature Selection. 2020. https:\/\/scikit-learn.org\/stable\/modules\/generated\/sklearn-.feature_selection.chi2.html. Accessed 16 Aug 2021"},{"issue":"1","key":"661_CR37","doi-asserted-by":"publisher","first-page":"175","DOI":"10.1007\/s00521-013-1368-0","volume":"24","author":"JR Vergara","year":"2014","unstructured":"Vergara JR, Est\u00e9vez PA. A review of feature selection methods based on mutual information. Neural Comput Appl. 2014;24(1):175\u201386.","journal-title":"Neural Comput Appl"},{"key":"661_CR38","unstructured":"Mohammad AH. Comparing two feature selections methods information gain and gain ratio on three different classification algorithms using arabic dataset. J Theor Appl Inf Technol 2018; 96(6)"},{"key":"661_CR39","unstructured":"info_gain Pypi project. 2020. https:\/\/pypi.org\/project\/info-gain\/. Accessed 16 Aug 2021."},{"issue":"7","key":"661_CR40","doi-asserted-by":"publisher","first-page":"1145","DOI":"10.1016\/S0031-3203(96)00142-2","volume":"30","author":"AP Bradley","year":"1997","unstructured":"Bradley AP. The use of the area under the roc curve in the evaluation of machine learning algorithms. Pattern Recogn. 1997;30(7):1145\u201359.","journal-title":"Pattern Recogn"},{"issue":"12","key":"661_CR41","doi-asserted-by":"publisher","first-page":"1334","DOI":"10.1109\/PROC.1983.12775","volume":"71","author":"JD Day","year":"1983","unstructured":"Day JD, Zimmermann H. The OSI reference model. Proc IEEE. 1983;71(12):1334\u201340.","journal-title":"Proc IEEE"},{"key":"661_CR42","doi-asserted-by":"crossref","unstructured":"Lashkari AH, Draper-Gil G, Mamun MSI, Ghorbani AA. Characterization of tor traffic using time based features. In: ICISSp, 2017. p. 253\u2013262","DOI":"10.5220\/0005740704070414"},{"key":"661_CR43","doi-asserted-by":"crossref","unstructured":"Draper-Gil G, Lashkari AH, Mamun MSI, Ghorbani AA. Characterization of encrypted and vpn traffic using time-related. In: Proceedings of the 2nd international conference on information systems security and privacy (ICISSP). 2016. p. 407\u2013414","DOI":"10.5220\/0005740704070414"},{"key":"661_CR44","unstructured":"OWASP A2:2017-Broken Authentication. 2020. https:\/\/owasp.org\/www-project-top-ten\/2017\/A2_2017-Broken-_Authentication. Accessed 10 Aug 2021"},{"key":"661_CR45","unstructured":"OWASP A10:2017-Insufficient Logging & Monitoring. 2020. https:\/\/owasp.org\/www-project-top-ten\/2017\/A10_2017--Insufficient_Logging%2526Monitoring. Accessed 10 Aug 2021"}],"container-title":["Journal of Big Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-022-00661-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s40537-022-00661-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-022-00661-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,15]],"date-time":"2022-12-15T16:19:04Z","timestamp":1671121144000},"score":1,"resource":{"primary":{"URL":"https:\/\/journalofbigdata.springeropen.com\/articles\/10.1186\/s40537-022-00661-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,15]]},"references-count":45,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,12]]}},"alternative-id":["661"],"URL":"https:\/\/doi.org\/10.1186\/s40537-022-00661-9","relation":{},"ISSN":["2196-1115"],"issn-type":[{"value":"2196-1115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,15]]},"assertion":[{"value":"21 March 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 October 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 December 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"119"}}