{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T03:56:44Z","timestamp":1784606204733,"version":"3.55.0"},"reference-count":47,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2021,5,27]],"date-time":"2021-05-27T00:00:00Z","timestamp":1622073600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,5,27]],"date-time":"2021-05-27T00:00:00Z","timestamp":1622073600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100005416","name":"Norges Forskningsr\u00e5d","doi-asserted-by":"publisher","award":["247678"],"award-info":[{"award-number":["247678"]}],"id":[{"id":"10.13039\/501100005416","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100016999","name":"Western Norway University Of Applied Sciences","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100016999","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Software Qual J"],"published-print":{"date-parts":[[2021,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Security remains under-addressed in many organisations, illustrated by the number of large-scale software security breaches. Preventing breaches can begin during software development if attention is paid to security during the software\u2019s design and implementation. One approach to security assurance during software development is to examine communications between developers as a means of studying the security concerns of the project. Prior research has investigated models for classifying project communication messages (e.g., issues or commits) as security related or not. A known problem is that these models are project-specific, limiting their use by other projects or organisations. We investigate whether we can build a generic classification model that can generalise across projects. We define a set of security keywords by extracting them from relevant security sources, dividing them into four categories: asset, attack\/threat, control\/mitigation, and implicit. Using different combinations of these categories and including them in the training dataset, we built a classification model and evaluated it on industrial, open-source, and research-based datasets containing over 45 different products. Our model based on harvested security keywords as a feature set shows average recall from 55 to 86%, minimum recall from 43 to 71% and maximum recall from 60 to 100%. An average f-score between 3.4 and 88%, an average g-measure of at least 66% across all the dataset, and an average AUC of ROC from 69 to 89%. In addition, models that use externally sourced features outperformed models that use project-specific features on average by a margin of 26\u201344% in recall, 22\u201350% in g-measure, 0.4\u201328% in f-score, and 15\u201319% in AUC of ROC. Further, our results outperform a state-of-the-art prediction model for security bug reports in all cases. We find using sound statistical and effect size tests that (1) using harvested security keywords as features to train a text classification model improve classification models and generalise to other projects significantly. (2) Including features in the training dataset before model construction improve classification models significantly. (3) Different security categories represent predictors for different projects. Finally, we introduce new and promising approaches to construct models that can generalise across different independent projects.<\/jats:p>","DOI":"10.1007\/s11219-020-09546-7","type":"journal-article","created":{"date-parts":[[2021,5,27]],"date-time":"2021-05-27T00:02:26Z","timestamp":1622073746000},"page":"509-553","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["An improved text classification modelling approach to identify security messages in heterogeneous projects"],"prefix":"10.1007","volume":"29","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0027-4522","authenticated-orcid":false,"given":"Tosin Daniel","family":"Oyetoyan","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Patrick","family":"Morrison","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,5,27]]},"reference":[{"key":"9546_CR1","doi-asserted-by":"crossref","unstructured":"Anvik, J., Hiew, L., & Murphy, G. C. (2006). Who should fix this bug? In: Proceedings of the 28th international conference on Software engineering, ACM, pp. 361\u2013370.","DOI":"10.1145\/1134285.1134336"},{"key":"9546_CR2","doi-asserted-by":"crossref","unstructured":"Bozorgi, M., Saul, L. K., Savage, S. & Voelker, G. M. (2010). Beyond heuristics: learning to classify vulnerabilities and predict exploits. In Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining, pp. 105\u2013114.","DOI":"10.1145\/1835804.1835821"},{"issue":"11","key":"9546_CR3","doi-asserted-by":"publisher","first-page":"943","DOI":"10.1109\/32.177364","volume":"18","author":"R Chillarege","year":"1992","unstructured":"Chillarege, R., Bhandari, I. S., Chaar, J. K., Halliday, M. J., Moebus, D. S., Ray, B. K., & Wong, M. Y. (1992). Orthogonal defect classification-a concept for in-process measurements. IEEE Transactions on Software Engineering, 18(11), 943\u2013956.","journal-title":"IEEE Transactions on Software Engineering"},{"key":"9546_CR4","unstructured":"Christey, S., & Martin, B. (n.d.). Buying into the bias: why vulnerability statistics suck, BlackHat, Las Vegas, USA, Tech. Rep 1."},{"key":"9546_CR5","doi-asserted-by":"publisher","unstructured":"Cleland-Huang, J., Settimi, R., Zou, X., & Solc, P. (2006). The detection and classification of non-functional requirements with application to early aspects. In Requirements Engineering, 14th IEEE International Conference, Minneapolis, Minnesota, pp. 39\u201348. https:\/\/doi.org\/10.1109\/RE.2006.65","DOI":"10.1109\/RE.2006.65"},{"key":"9546_CR6","unstructured":"Cois, C. A., & Kazman, R. (2015). Natural language processing to quantify security effort in the software development lifecycle. In SEKE, pp. 716\u2013 721."},{"key":"9546_CR7","doi-asserted-by":"crossref","unstructured":"Dai, X., et al. (2017). From social media to public health surveillance: word embedding based clustering method for twitter classification. In SoutheastCon pp.1\u20137.","DOI":"10.1109\/SECON.2017.7925400"},{"key":"9546_CR8","doi-asserted-by":"crossref","unstructured":"Debole, F., & Sebastiani, F. (2004). Supervised term weighting for automated text categorization. In Text mining and its applications, Springer, pp. 81\u201397.","DOI":"10.1007\/978-3-540-45219-5_7"},{"key":"9546_CR9","unstructured":"Demsar, J. (2006). Statistical comparisons of classifiers over multiple data sets. Journal of Machine learning research, 7, 1\u201330."},{"key":"9546_CR10","volume-title":"Evaluating and mitigating software supply chain security risks","author":"RJ Ellison","year":"2010","unstructured":"Ellison, R. J., Goodenough, J. B., Weinstock, C. B., & Woody, C. (2010). Evaluating and mitigating software supply chain security risks. Tech. rep.: CARNEGIE-MELLON UNIV PITTSBURGH PA SOFTWARE ENGINEERING INST."},{"key":"9546_CR11","unstructured":"Feinerer, I. (2013). Introduction to the tm Package Text Mining in R. Accessible enligne: http:\/\/cran.r-project.org\/web\/packages\/tm\/vignettes\/tm.pdf"},{"key":"9546_CR12","unstructured":"Forman, G. (2003). An extensive empirical study of feature selection metrics for text classification, Journal of machine learning research, 3, 1289\u20131305."},{"key":"9546_CR13","doi-asserted-by":"crossref","unstructured":"Gegick, M., Rotella, P., & Xie, T. (2010). Identifying security bug reports via text mining: An industrial case study. In Mining software repositories (MSR), 2010 7th IEEE working conference on, IEEE, pp. 11\u201320.","DOI":"10.1109\/MSR.2010.5463340"},{"issue":"9","key":"9546_CR14","doi-asserted-by":"publisher","first-page":"1263","DOI":"10.1109\/TKDE.2008.239","volume":"21","author":"H He","year":"2009","unstructured":"He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on knowledge and data engineering, 21(9), 1263\u20131284.","journal-title":"IEEE Transactions on knowledge and data engineering"},{"issue":"6","key":"9546_CR15","doi-asserted-by":"publisher","first-page":"1125","DOI":"10.1007\/s10664-012-9209-9","volume":"18","author":"A Hindle","year":"2013","unstructured":"Hindle, A., Ernst, N. A., Godfrey, M. W., & Mylopoulos, J. (2013). Auto- mated topic naming. Empirical Softw. Engg., 18(6), 1125\u20131155. https:\/\/doi.org\/10.1007\/s10664-012-9209-9.doi:10.1007\/s10664-012-9209-9","journal-title":"Empirical Softw. Engg."},{"key":"9546_CR16","doi-asserted-by":"crossref","unstructured":"Joachims, T. (1998). Text categorization with support vector machines: learning with many relevant features, Machine learning: ECML-98, 137\u2013142.","DOI":"10.1007\/BFb0026683"},{"issue":"11\u201312","key":"9546_CR17","doi-asserted-by":"publisher","first-page":"1073","DOI":"10.1016\/j.infsof.2007.02.015","volume":"49","author":"VB Kampenes","year":"2007","unstructured":"Kampenes, V. B., Dyb\u00e5, T., Hannay, J. E., & Sj\u00f8berg, D. I. K. (2007). A systematic review of effect size in software engineering experiments. Information and Software Technology, 49(11\u201312), 1073\u20131086.","journal-title":"Information and Software Technology"},{"key":"9546_CR18","unstructured":"Le, Q., & Mikolov, T. (2014). Distributed representations of sentences and documents. In Proceedings of the 31st International Conference on Machine Learning (ICML-14), pp. 1188\u20131196."},{"key":"9546_CR19","unstructured":"Louppe, G. (2014). Understanding random forests: from theory to practice. arXiv preprint arXiv: 1407.7502."},{"issue":"2\u20134","key":"9546_CR20","doi-asserted-by":"publisher","first-page":"100","DOI":"10.1017\/CBO9780511809071.007","volume":"100","author":"CD Manning","year":"2008","unstructured":"Manning, C. D., Raghavan, P., & Sch\u00fctze, H. (2008). Scoring, term weighting and the vector space model. Introduction to information retrieval, 100(2\u20134), 100\u2013123.","journal-title":"Introduction to information retrieval"},{"key":"9546_CR21","doi-asserted-by":"crossref","unstructured":"Massacci, F., & Nguyen, V. H. (2010). Which is the right source for vulnerability studies?: an empirical analysis on Mozilla Firefox. In Proceedings of the 6th International Workshop on Security Measurements and Metrics, ACM, p. 4.","DOI":"10.1145\/1853919.1853925"},{"key":"9546_CR22","doi-asserted-by":"crossref","unstructured":"Morrison, P., Oyetoyan, T. D., & Williams, L. (2018b). Identifying security issues in software development: are keywords enough? In Proceedings of the 40th International Conference on Software Engineering: Companion Proceeedings, pp. 426\u2013427.","DOI":"10.1145\/3183440.3195040"},{"key":"9546_CR23","doi-asserted-by":"crossref","unstructured":"Morrison, P. J., Pandita, R., Xiao, X., Chillarege, R., & Williams, L. (2018a). Are vulnerabilities discovered and resolved like other defects? Empirical Software Engineering, 23(3), 1383\u20131421.","DOI":"10.1007\/s10664-017-9541-1"},{"issue":"2","key":"9546_CR24","doi-asserted-by":"publisher","first-page":"103","DOI":"10.1023\/A:1007692713085","volume":"39","author":"K Nigam","year":"2000","unstructured":"Nigam, K., McCallum, A. K., Thrun, S., & Mitchell, T. (2000). Text classification from labeled and unlabeled documents using em. Machine learning, 39(2), 103\u2013134.","journal-title":"Machine learning"},{"key":"9546_CR25","doi-asserted-by":"crossref","unstructured":"Ohira, M., Kashiwa, Y., Yamatani, Y., Yoshiyuki, H., Maeda, Y., Lim- settho, N., Fujino, K., Hata, H., Ihara, A., & Matsumoto, K. (2015). A dataset of high impact bugs: manually-classified issue reports. In Mining Soft- ware Repositories (MSR), 2015 IEEE\/ACM 12th Working Conference on, IEEE, pp. 518\u2013521.","DOI":"10.1109\/MSR.2015.78"},{"key":"9546_CR26","unstructured":"Peters, F., Tun, T., Yu, Y., & Nuseibeh, B. (2017). Text filtering and ranking for security bug report prediction. IEEE Transactions on Software Engineering."},{"key":"9546_CR27","doi-asserted-by":"crossref","unstructured":"Pletea, D., Vasilescu, B., & Serebrenik, A. (2014). Security and emotion: sentiment analysis of security discussions on github. In Proceedings of the 11th working conference on mining software repositories, ACM, pp. 348\u2013351.","DOI":"10.1145\/2597073.2597117"},{"key":"9546_CR28","unstructured":"Ponemon-Institute, IBM-Security. (2017). Cost of data breach study: Global overview benchmark research sponsored by ibm security independently conducted by ponemon institute llc, Ponemon Institute Research Report."},{"issue":"1","key":"9546_CR29","first-page":"37","volume":"2","author":"D Powers","year":"2011","unstructured":"Powers, D. (2011). Evaluation: from precision, recall and F-Factor to ROC, informedness, markedness & correlation. J. Mach. Learn. Technol, 2(1), 37\u201363.","journal-title":"J. Mach. Learn. Technol"},{"key":"9546_CR30","unstructured":"R Development Core Team. (2008). R: A Language and Environment for Statistical Computing, R Foundation for Statistical Computing, Vienna, Austria, ISBN 3\u2013900051\u201307\u20130. http:\/\/www.R-project.org"},{"key":"9546_CR31","doi-asserted-by":"publisher","unstructured":"Ray, B., Hellendoorn, V., Godhane, S., Tu, Z., Bacchelli, A., & Devanbu, P. (2006). On the \u201dnaturalness\u201d of buggy code. In Proceedings of the 38th In- ternational Conference on Software Engineering, ICSE \u201916, ACM, New York, NY, USA, pp. 428\u2013439. https:\/\/doi.org\/10.1145\/2884781.2884848","DOI":"10.1145\/2884781.2884848"},{"key":"9546_CR32","doi-asserted-by":"crossref","unstructured":"Ray, B., Posnett, D., Filkov, V., & Devanbu, P. (2014). A large scale study of pro- gramming languages and code quality in github. In Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Soft- ware Engineering, ACM, pp. 155\u2013165.","DOI":"10.1145\/2635868.2635922"},{"key":"9546_CR33","doi-asserted-by":"crossref","unstructured":"Riaz, M., King, J., Slankas, J., & Williams, L. (2014). Hidden in plain sight: auto- matically identifying security requirements from natural language arti- facts. In Proc. 22nd RE, IEEE, pp. 183\u2013192.","DOI":"10.1109\/RE.2014.6912260"},{"issue":"5","key":"9546_CR34","doi-asserted-by":"publisher","first-page":"513","DOI":"10.1016\/0306-4573(88)90021-0","volume":"24","author":"G Salton","year":"1988","unstructured":"Salton, G., & Buckley, C. (1988). Term-weighting approaches in automatic text retrieval. Information processing & management, 24(5), 513\u2013523.","journal-title":"Information processing & management"},{"issue":"11","key":"9546_CR35","doi-asserted-by":"publisher","first-page":"1022","DOI":"10.1145\/182.358466","volume":"26","author":"G Salton","year":"1983","unstructured":"Salton, G., Fox, E. A., & Wu, H. (1983). Extended Boolean information retrieval. Communications of the ACM, 26(11), 1022\u20131036.","journal-title":"Communications of the ACM"},{"key":"9546_CR36","unstructured":"Salton, G., & McGill, M. J. (n.d.). Introduction to modern information retrieval."},{"issue":"10","key":"9546_CR37","doi-asserted-by":"publisher","first-page":"993","DOI":"10.1109\/TSE.2014.2340398","volume":"40","author":"R Scandariato","year":"2014","unstructured":"Scandariato, R., Walden, J., Hovsepyan, A., & Joosen, W. (2014). Predicting vul- nerable software components via text mining. IEEE Transactions on Software Engineering, 40(10), 993\u20131006.","journal-title":"IEEE Transactions on Software Engineering"},{"issue":"1","key":"9546_CR38","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/505282.505283","volume":"34","author":"F Sebastiani","year":"2002","unstructured":"Sebastiani, F. (2002). Machine learning in automated text categorization. ACM computing surveys (CSUR), 34(1), 1\u201347.","journal-title":"ACM computing surveys (CSUR)"},{"issue":"1","key":"9546_CR39","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1108\/eb026526","volume":"28","author":"K Sparck Jones","year":"1972","unstructured":"Sparck Jones, K. (1972). A statistical interpretation of term specificity and its application in retrieval. Journal of Documentation, 28(1), 11\u201321.","journal-title":"Journal of Documentation"},{"key":"9546_CR40","unstructured":"Tyo, J. P. (2016). Empirical analysis and automated classification of security bug reports."},{"key":"9546_CR41","unstructured":"Unterkalmsteiner, M., Abrahamsson, P., Wang, X., Nguyen-Duc, A., Shah, S., Bajwa, S. S., Baltes, G. H., Conboy, K., Cullina, E., Dennehy D., et al. (n.d.). Software startups\u2013a research agenda, e-Informatica Software En- gineering Journal 10 (1)."},{"key":"9546_CR42","doi-asserted-by":"crossref","unstructured":"Wijayasekara, D., Manic, M., & McQueen, M. (2014). Vulnerability identification and classification via text mining bug databases. In Industrial Electronics Society, IECON 2014\u201340th Annual Conference of the IEEE, IEEE, pp. 3612\u20133618.","DOI":"10.1109\/IECON.2014.7049035"},{"key":"9546_CR43","volume-title":"Data mining: practical machine learning tools and techniques","author":"IH Witten","year":"2005","unstructured":"Witten, I. H., & Frank, E. (2005). Data mining: practical machine learning tools and techniques (2nd ed.). San Francisco: Morgan Kaufmann.","edition":"2"},{"issue":"3","key":"9546_CR44","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/1361684.1361686","volume":"26","author":"HC Wu","year":"2008","unstructured":"Wu, H. C., Luk, R. W. P., Wong, K. F., & Kwok, K. L. (2008). Interpreting TF-IDF term weights as making relevance decisions. ACM Trans Inf Systems (TOIS), 26(3), 1\u201337.","journal-title":"ACM Trans Inf Systems (TOIS)"},{"key":"9546_CR45","unstructured":"Xia, T., Krishna, R., Chen, J., Mathew, G., Shen, X., & Menzies, T. (2018). Hyperparameter optimization for effort estimation.\u00a0arXiv preprint arXiv: 1805.00336."},{"key":"9546_CR46","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.infsof.2017.07.003","volume":"92","author":"M Yan","year":"2017","unstructured":"Yan, M., Zhang, X., Liu, C., Xu, L., Yang, M., & Yang, D. (2017). Automated change- prone class prediction on unlabeled dataset using unsupervised method. Information and Software Technology, 92, 1\u201316.","journal-title":"Information and Software Technology"},{"key":"9546_CR47","doi-asserted-by":"crossref","unstructured":"Zaman, S., Adams, B., & Hassan, A. E. (2011). Security versus performance bugs: a case study on firefox. In Proceedings of the 8th working conference on mining software repositories, ACM, pp. 93\u2013102.","DOI":"10.1145\/1985441.1985457"}],"container-title":["Software Quality Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11219-020-09546-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11219-020-09546-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11219-020-09546-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,6,5]],"date-time":"2021-06-05T01:12:34Z","timestamp":1622855554000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11219-020-09546-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,27]]},"references-count":47,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,6]]}},"alternative-id":["9546"],"URL":"https:\/\/doi.org\/10.1007\/s11219-020-09546-7","relation":{},"ISSN":["0963-9314","1573-1367"],"issn-type":[{"value":"0963-9314","type":"print"},{"value":"1573-1367","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,5,27]]},"assertion":[{"value":"26 December 2020","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 May 2021","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}