{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,20]],"date-time":"2026-02-20T23:57:51Z","timestamp":1771631871733,"version":"3.50.1"},"reference-count":34,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2022,10,1]],"date-time":"2022-10-01T00:00:00Z","timestamp":1664582400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,10,1]],"date-time":"2022-10-01T00:00:00Z","timestamp":1664582400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100005389","name":"Universit\u00e0 degli Studi dell'Insubria","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100005389","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Empir Software Eng"],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Context<\/jats:title>\n                <jats:p>The F-measure has been widely used as a performance metric when selecting binary classifiers for prediction, but it has also been widely criticized, especially given the availability of alternatives such as <jats:italic>\u03d5<\/jats:italic> (also known as Matthews Correlation Coefficient).<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Objectives<\/jats:title>\n                <jats:p>Our goals are to (1) investigate possible issues related to the F-measure in depth and show how <jats:italic>\u03d5<\/jats:italic> can address them, and (2) explore the relationships between the F-measure and <jats:italic>\u03d5<\/jats:italic>.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Method<\/jats:title>\n                <jats:p>Based on the definitions of <jats:italic>\u03d5<\/jats:italic> and the F-measure, we derive a few mathematical properties of these two performance metrics and of the relationships between them. To demonstrate the practical effects of these mathematical properties, we illustrate the outcomes of an empirical study involving 70 Empirical Software Engineering datasets and 837 classifiers.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>We show that <jats:italic>\u03d5<\/jats:italic> can be defined as a function of <jats:italic>Precision<\/jats:italic> and <jats:italic>Recall<\/jats:italic>, which are the only two performance metrics used to define the F-measure, and the rate of actually positive software modules in a dataset. Also, <jats:italic>\u03d5<\/jats:italic> can be expressed as a function of the F-measure and the rates of actual and estimated positive software modules. We derive the minimum and maximum value of <jats:italic>\u03d5<\/jats:italic> for any given value of the F-measure, and the conditions under which both the F-measure and <jats:italic>\u03d5<\/jats:italic> rank two classifiers in the same order.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusions<\/jats:title>\n                <jats:p>Our results show that <jats:italic>\u03d5<\/jats:italic> is a sensible and useful metric for assessing the performance of binary classifiers. We also recommend that the F-measure should not be used by itself to assess the performance of a classifier, but that the rate of positives should always be specified as well, at least to assess if and to what extent a classifier performs better than random classification. The mathematical relationships described here can also be used to re-interpret the conclusions of previously published papers that relied mainly on the F-measure as a performance metric.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1007\/s10664-022-10199-2","type":"journal-article","created":{"date-parts":[[2022,10,1]],"date-time":"2022-10-01T08:02:32Z","timestamp":1664611352000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["Comparing \u03d5 and the F-measure as performance metrics for software-related classifications"],"prefix":"10.1007","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5226-4337","authenticated-orcid":false,"given":"Luigi","family":"Lavazza","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sandro","family":"Morasca","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,10,1]]},"reference":[{"key":"10199_CR1","unstructured":"The SEACRAFT repository of empirical software engineering data. https:\/\/zenodo.org\/communities\/seacraft (2017)"},{"key":"10199_CR2","doi-asserted-by":"crossref","unstructured":"Bowes D, Hall T, Gray D (2012) Comparing the performance of fault prediction models which report multiple performance measures: recomputing the confusion matrix. In: Proceedings of the 8th international conference on predictive models in software engineering, pp 109\u2013118","DOI":"10.1145\/2365324.2365338"},{"issue":"2","key":"10199_CR3","doi-asserted-by":"publisher","first-page":"525","DOI":"10.1007\/s11219-016-9353-3","volume":"26","author":"D Bowes","year":"2018","unstructured":"Bowes D, Hall T, Petri\u0107 J (2018) Software defect prediction: do different classifiers find the same defects?. Softw Qual J 26(2):525\u2013552","journal-title":"Softw Qual J"},{"key":"10199_CR4","unstructured":"Cauchy A (1821) Cours d\u2019analyse de l\u2019\u00e9cole royale polyt\u00e9chnique, Vol. I. Analyse analyse. International Centre for Mechanical Sciences. Debure"},{"issue":"1","key":"10199_CR5","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s12864-019-6413-7","volume":"21","author":"D Chicco","year":"2020","unstructured":"Chicco D, Jurman G (2020) The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC genomics 21(1):1\u201313","journal-title":"BMC genomics"},{"issue":"1","key":"10199_CR6","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1177\/001316446002000104","volume":"20","author":"J Cohen","year":"1960","unstructured":"Cohen J (1960) A coefficient of agreement for nominal scales. Educ Psychol Meas 20(1):37\u201346. https:\/\/doi.org\/10.1177\/001316446002000104","journal-title":"Educ Psychol Meas"},{"key":"10199_CR7","volume-title":"Statistical power analysis for the behavioral sciences lawrence earlbaum associates","author":"J Cohen","year":"1988","unstructured":"Cohen J (1988) Statistical power analysis for the behavioral sciences lawrence earlbaum associates. Routledge, New York"},{"issue":"9","key":"10199_CR8","doi-asserted-by":"publisher","first-page":"e0222916","DOI":"10.1371\/journal.pone.0222916","volume":"14","author":"R Delgado","year":"2019","unstructured":"Delgado R, Tibau XA (2019) Why Cohen\u2019s Kappa should be avoided as performance measure in classification. PloS one 14(9):e0222916","journal-title":"PloS one"},{"key":"10199_CR9","doi-asserted-by":"publisher","first-page":"66647","DOI":"10.1109\/ACCESS.2020.2985780","volume":"8","author":"J Deng","year":"2020","unstructured":"Deng J, Lu L, Qiu S, Ou Y (2020) A suitable AST node granularity and multi-kernel transfer convolutional neural network for cross-project defect prediction. IEEE Access 8:66647\u201366661","journal-title":"IEEE Access"},{"issue":"9","key":"10199_CR10","doi-asserted-by":"publisher","first-page":"1057","DOI":"10.3390\/e22091057","volume":"22","author":"E Dias Canedo","year":"2020","unstructured":"Dias Canedo E, Cordeiro Mendes B (2020) Software requirements classification using machine learning algorithms. Entropy 22(9):1057","journal-title":"Entropy"},{"key":"10199_CR11","doi-asserted-by":"crossref","unstructured":"Gray D, Bowes D, Davey N, Sun Y, Christianson B (2011) The misuse of the NASA metrics data program data sets for automated software defect prediction. In: 15th annual conference on evaluation & assessment in software engineering (EASE 2011), pp 96\u2013103","DOI":"10.1049\/ic.2011.0012"},{"issue":"6","key":"10199_CR12","doi-asserted-by":"publisher","first-page":"1276","DOI":"10.1109\/TSE.2011.103","volume":"38","author":"T Hall","year":"2011","unstructured":"Hall T, Beecham S, Bowes D, Gray D, Counsell S (2011) A systematic literature review on fault prediction performance in software engineering. IEEE Trans Softw Eng 38(6):1276\u20131304","journal-title":"IEEE Trans Softw Eng"},{"key":"10199_CR13","first-page":"2813","volume":"13","author":"J Hern\u00e1ndez-Orallo","year":"2012","unstructured":"Hern\u00e1ndez-Orallo J., Flach PA, Ferri C (2012) A unified view of performance metrics: translating threshold choice into expected classification loss. J Mach Learn Res 13:2813\u20132869. http:\/\/dl.acm.org\/citation.cfm?id=2503332","journal-title":"J Mach Learn Res"},{"key":"10199_CR14","doi-asserted-by":"crossref","unstructured":"Jureczko M, Madeyski L (2010) Towards identifying software project clusters with regard to defect prediction. In: Proceedings of the 6th international conference on predictive models in software engineering, pp 1\u201310","DOI":"10.1145\/1868328.1868342"},{"issue":"3","key":"10199_CR15","doi-asserted-by":"publisher","first-page":"419","DOI":"10.1177\/09622802211060515","volume":"31","author":"L Lavazza","year":"2022","unstructured":"Lavazza L, Morasca S (2022) Considerations on the region of interest in the ROC space. Stat Methods Med Res 31(3):419\u2013437","journal-title":"Stat Methods Med Res"},{"issue":"2","key":"10199_CR16","doi-asserted-by":"publisher","first-page":"201","DOI":"10.1007\/s10515-011-0092-1","volume":"19","author":"M Li","year":"2012","unstructured":"Li M, Zhang H, Wu R, Zhou ZH (2012) Sample-based software defect prediction with active and semi-supervised learning. Autom Softw Eng 19 (2):201\u2013230","journal-title":"Autom Softw Eng"},{"key":"10199_CR17","doi-asserted-by":"publisher","first-page":"216","DOI":"10.1016\/j.patcog.2019.02.023","volume":"91","author":"A Luque","year":"2019","unstructured":"Luque A, Carrasco A, Mart\u00edn A, de Las Heras A (2019) The impact of class imbalance in classification performance metrics based on the binary confusion matrix. Pattern Recogn 91:216\u2013231","journal-title":"Pattern Recogn"},{"issue":"2","key":"10199_CR18","doi-asserted-by":"publisher","first-page":"442","DOI":"10.1016\/0005-2795(75)90109-9","volume":"405","author":"BW Matthews","year":"1975","unstructured":"Matthews BW (1975) Comparison of the predicted and observed secondary structure of t4 phage lysozyme. Biochimica et Biophysica Acta (BBA)-Protein Structure 405(2):442\u2013451","journal-title":"Biochimica et Biophysica Acta (BBA)-Protein Structure"},{"key":"10199_CR19","doi-asserted-by":"crossref","unstructured":"Menzies T, Di Stefano JS (2004) How good is your blind spot sampling policy. In: Eighth IEEE international symposium on high assurance systems engineering, 2004. Proceedings. IEEE, pp 129\u2013138","DOI":"10.1109\/HASE.2004.1281737"},{"key":"10199_CR20","doi-asserted-by":"crossref","unstructured":"Morasca S, Lavazza L (2016) Slope-based fault-proneness thresholds for software engineering measures. In: Proceedings of the 20th international conference on evaluation and assessment in software engineering, pp 1\u201310","DOI":"10.1145\/2915970.2915997"},{"key":"10199_CR21","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1016\/j.infsof.2017.03.005","volume":"89","author":"S Morasca","year":"2017","unstructured":"Morasca S, Lavazza L (2017) Risk-averse slope-based thresholds: Definition and empirical evaluation. Information & Software Technology 89:37\u201363. https:\/\/doi.org\/10.1016\/j.infsof.2017.03.005","journal-title":"Information & Software Technology"},{"issue":"5","key":"10199_CR22","doi-asserted-by":"publisher","first-page":"3977","DOI":"10.1007\/s10664-020-09861-4","volume":"25","author":"S Morasca","year":"2020","unstructured":"Morasca S, Lavazza L (2020) On the assessment of software defect prediction models via ROC curves. Empir Softw Eng 25(5):3977\u20134019","journal-title":"Empir Softw Eng"},{"issue":"1","key":"10199_CR23","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1140\/epjds\/s13688-020-00253-8","volume":"9","author":"F Pierri","year":"2020","unstructured":"Pierri F, Piccardi C, Ceri S (2020) A multi-layer approach to disinformation detection in us and italian news spreading on twitter. EPJ Data Science 9(1):35","journal-title":"EPJ Data Science"},{"key":"10199_CR24","unstructured":"Powers DM (2011) Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation"},{"key":"10199_CR25","doi-asserted-by":"publisher","first-page":"100172","DOI":"10.1109\/ACCESS.2020.2997939","volume":"8","author":"GF Scaranti","year":"2020","unstructured":"Scaranti GF, Carvalho LF, Barbon S, Proen\u00e7a ML (2020) Artificial immune systems and fuzzy logic to detect flooding attacks in software-defined networks. IEEE Access 8:100172\u2013100184","journal-title":"IEEE Access"},{"key":"10199_CR26","doi-asserted-by":"crossref","unstructured":"Serafini P (1985) Mathematics of multi objective optimization. International Centre for Mechanical Sciences. Springer","DOI":"10.1007\/978-3-7091-2822-0"},{"key":"10199_CR27","unstructured":"Singh PK, Agarwal D, Gupta A (2015) A systematic review on software defect prediction. In: 2015 2nd international conference on computing for sustainable global development (INDIACom). IEEE, pp 1793\u20131797"},{"issue":"4","key":"10199_CR28","doi-asserted-by":"publisher","first-page":"427","DOI":"10.1016\/j.ipm.2009.03.002","volume":"45","author":"M Sokolova","year":"2009","unstructured":"Sokolova M, Lapalme G (2009) A systematic analysis of performance measures for classification tasks. Information processing & management 45(4):427\u2013437","journal-title":"Information processing & management"},{"key":"10199_CR29","doi-asserted-by":"crossref","unstructured":"Sonbol R, Rebdawi G, Ghneim N (2020) Towards a semantic representation for functional software requirements. In: 2020 IEEE seventh international workshop on artificial intelligence for requirements engineering (AIRE). IEEE, pp 1\u20138","DOI":"10.1109\/AIRE51212.2020.00007"},{"issue":"12","key":"10199_CR30","doi-asserted-by":"publisher","first-page":"1253","DOI":"10.1109\/TSE.2018.2836442","volume":"45","author":"Q Song","year":"2019","unstructured":"Song Q, Guo Y, Shepperd M (2019) A comprehensive investigation of the role of imbalanced learning for software defect prediction. IEEE Trans. Software Eng. 45(12):1253\u20131269","journal-title":"IEEE Trans. Software Eng."},{"key":"10199_CR31","unstructured":"van Rijsbergen CJ (1979) Information retrieval. Butterworth"},{"key":"10199_CR32","doi-asserted-by":"crossref","unstructured":"Yao J, Shepperd M (2020) Assessing software defection prediction performance: Why using the Matthews correlation coefficient matters. In: Proceedings of the evaluation and assessment in software engineering, pp 120\u2013129","DOI":"10.1145\/3383219.3383232"},{"key":"10199_CR33","doi-asserted-by":"publisher","first-page":"106664","DOI":"10.1016\/j.infsof.2021.106664","volume":"139","author":"J Yao","year":"2021","unstructured":"Yao J, Shepperd M (2021) The impact of using biased performance metrics on software defect prediction research. Inf Softw Technol 139:106664","journal-title":"Inf Softw Technol"},{"issue":"6","key":"10199_CR34","doi-asserted-by":"publisher","first-page":"3186","DOI":"10.1007\/s10664-017-9516-2","volume":"22","author":"F Zhang","year":"2017","unstructured":"Zhang F, Keivanloo I, Zou Y (2017) Data transformation in cross-project defect prediction. Empir Softw Eng 22(6):3186\u20133218","journal-title":"Empir Softw Eng"}],"container-title":["Empirical Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10664-022-10199-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10664-022-10199-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10664-022-10199-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,11,21]],"date-time":"2022-11-21T02:17:13Z","timestamp":1668997033000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10664-022-10199-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,1]]},"references-count":34,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["10199"],"URL":"https:\/\/doi.org\/10.1007\/s10664-022-10199-2","relation":{},"ISSN":["1382-3256","1573-7616"],"issn-type":[{"value":"1382-3256","type":"print"},{"value":"1573-7616","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,1]]},"assertion":[{"value":"2 July 2022","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 October 2022","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no conflict of interest to declare.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"<!--Emphasis Type='Bold' removed-->Conflict of interest"}}],"article-number":"185"}}