{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T22:28:57Z","timestamp":1787351337349,"version":"build-2736575974"},"reference-count":50,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,6,26]],"date-time":"2024-06-26T00:00:00Z","timestamp":1719360000000},"content-version":"vor","delay-in-days":177,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,6,25]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Classification systems are evaluated in a countless number of papers. However, we find that evaluation practice is often nebulous. Frequently, metrics are selected without arguments, and blurry terminology invites misconceptions. For instance, many works use so-called \u2018macro\u2019 metrics to rank systems (e.g., \u2018macro F1\u2019) but do not clearly specify what they would expect from such a \u2018macro\u2019 metric. This is problematic, since picking a metric can affect research findings and thus any clarity in the process should be maximized. Starting from the intuitive concepts of bias and prevalence, we perform an analysis of common evaluation metrics. The analysis helps us understand the metrics\u2019 underlying properties, and how they align with expectations as found expressed in papers. Then we reflect on the practical situation in the field, and survey evaluation practice in recent shared tasks. We find that metric selection is often not supported with convincing arguments, an issue that can make a system ranking seem arbitrary. Our work aims at providing overview and guidance for more informed and transparent metric selection, fostering meaningful evaluation.<\/jats:p>","DOI":"10.1162\/tacl_a_00675","type":"journal-article","created":{"date-parts":[[2024,6,26]],"date-time":"2024-06-26T18:29:51Z","timestamp":1719426591000},"page":"820-836","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":137,"title":["A Closer Look at Classification Evaluation Metrics and a Critical\n                    Reflection of Common Evaluation Practice"],"prefix":"10.1162","volume":"12","author":[{"given":"Juri","family":"Opitz","sequence":"first","affiliation":[{"name":"University of Zurich, Switzerland. opitz.sci@gmail.com"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,6,25]]},"reference":[{"key":"2024081413395243800_bib1","doi-asserted-by":"publisher","first-page":"3938","DOI":"10.18653\/v1\/2020.acl-main.363","article-title":"An effectiveness metric for ordinal\n                        classification: Formal properties and experimental results","volume-title":"Proceedings of the 58th Annual Meeting of the Association for\n                        Computational Linguistics","author":"Amigo","year":"2020"},{"key":"2024081413395243800_bib2","doi-asserted-by":"publisher","first-page":"24","DOI":"10.18653\/v1\/S18-1003","article-title":"SemEval 2018 task 2: Multilingual emoji\n                        prediction","volume-title":"Proceedings of The 12th International\n                        Workshop on Semantic Evaluation","author":"Barbieri","year":"2018"},{"key":"2024081413395243800_bib3","doi-asserted-by":"publisher","first-page":"747","DOI":"10.18653\/v1\/S17-2126","article-title":"DataStories at SemEval- 2017 task 4: Deep\n                        LSTM with attention for message-level and topic-based sentiment\n                        analysis","volume-title":"Proceedings of the 11th International\n                        Workshop on Semantic Evaluation (SemEval-2017)","author":"Baziotis","year":"2017"},{"key":"2024081413395243800_bib4","first-page":"12","article-title":"Detecting spammers on\n                        twitter","volume-title":"Collaboration, electronic messaging,\n                        anti-abuse and spam conference (CEAS)","author":"Benevenuto","year":"2010"},{"issue":"1","key":"2024081413395243800_bib5","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s12864-019-6413-7","article-title":"The advantages of the matthews correlation\n                        coefficient (mcc) over f1 score and accuracy in binary classification\n                        evaluation","volume":"21","author":"Chicco","year":"2020","journal-title":"BMC Genomics"},{"key":"2024081413395243800_bib6","doi-asserted-by":"publisher","first-page":"573","DOI":"10.18653\/v1\/S17-2094","article-title":"BB_twtr at SemEval-2017 task 4: Twitter\n                        sentiment analysis with CNNs and LSTMs","volume-title":"Proceedings of the 11th International Workshop on Semantic\n                        Evaluation (SemEval-2017)","author":"Cliche","year":"2017"},{"issue":"1","key":"2024081413395243800_bib7","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1177\/001316446002000104","article-title":"A coefficient of agreement for nominal\n                        scales","volume":"20","author":"Cohen","year":"1960","journal-title":"Educational and Psychological\n                        Measurement"},{"issue":"9","key":"2024081413395243800_bib8","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1371\/journal.pone.0222916","article-title":"Why cohen\u2019s kappa should be avoided as\n                        performance measure in classification","volume":"14","author":"Delgado","year":"2019","journal-title":"PLOS\n                        ONE"},{"key":"2024081413395243800_bib9","doi-asserted-by":"publisher","first-page":"70","DOI":"10.18653\/v1\/2021.semeval-1.7","article-title":"SemEval-2021 task 6: Detection of\n                        persuasion techniques in texts and images","volume-title":"Proceedings of the 15th International Workshop on Semantic\n                        Evaluation (SemEval-2021)","author":"Dimitrov","year":"2021"},{"key":"2024081413395243800_bib10","doi-asserted-by":"publisher","first-page":"8189","DOI":"10.18653\/v1\/2020.emnlp-main.657","article-title":"Discriminatively-tuned generative\n                        classifiers for robust natural language inference","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural\n                        Language Processing (EMNLP)","author":"Ding","year":"2020"},{"key":"2024081413395243800_bib11","doi-asserted-by":"publisher","first-page":"221","DOI":"10.18653\/v1\/2023.semeval-1.31","article-title":"Epicurus at SemEval-2023 task 4: Improving\n                        prediction of human values behind arguments by leveraging their\n                        definitions","volume-title":"Proceedings of the 17th\n                        International Workshop on Semantic Evaluation (SemEval-2023)","author":"Fang","year":"2023"},{"issue":"8","key":"2024081413395243800_bib12","doi-asserted-by":"publisher","first-page":"861","DOI":"10.1016\/j.patrec.2005.10.010","article-title":"An introduction to ROC\n                        analysis","volume":"27","author":"Fawcett","year":"2006","journal-title":"Pattern Recognition Letters"},{"key":"2024081413395243800_bib13","article-title":"Precision-recall-gain curves: Pr analysis done\n                        right","volume-title":"Advances in Neural Information Processing\n                        Systems","author":"Flach","year":"2015"},{"key":"2024081413395243800_bib14","article-title":"Interpretable multi-dataset evaluation for\n                        named entity recognition","author":"Jinlan","year":"2020","journal-title":"arXiv preprint\n                        arXiv:2011.06854"},{"key":"2024081413395243800_bib15","doi-asserted-by":"publisher","DOI":"10.5244\/C.31.42","article-title":"Generative openmax for multi-class open\n                        set classification","volume-title":"British Machine Vision\n                        Conference","author":"Ge","year":"2017"},{"issue":"5\u20136","key":"2024081413395243800_bib16","doi-asserted-by":"publisher","first-page":"367","DOI":"10.1016\/j.compbiolchem.2004.09.006","article-title":"Comparing two k-category assignments by a\n                        k-category correlation coefficient","volume":"28","author":"Gorodkin","year":"2004","journal-title":"Computational\n                        Biology and Chemistry"},{"key":"2024081413395243800_bib17","article-title":"Metrics for multi-class classification: an\n                        overview","author":"Grandini","year":"2020","journal-title":"arXiv preprint\n                    arXiv:2008.05756"},{"key":"2024081413395243800_bib18","doi-asserted-by":"publisher","first-page":"700","DOI":"10.18653\/v1\/S17-2116","article-title":"Senti17 at SemEval-2017 task 4: Ten\n                        convolutional neural network voters for tweet polarity\n                        classification","volume-title":"Proceedings of the 11th\n                        International Workshop on Semantic Evaluation (SemEval-2017)","author":"Hamdan","year":"2017"},{"key":"2024081413395243800_bib19","doi-asserted-by":"publisher","first-page":"3905","DOI":"10.18653\/v1\/2022.naacl-main.287","article-title":"TRUE: Re-evaluating factual consistency\n                        evaluation","volume-title":"Proceedings of the 2022 Conference of\n                        the North American Chapter of the Association for Computational Linguistics:\n                        Human Language Technologies","author":"Or","year":"2022"},{"key":"2024081413395243800_bib20","doi-asserted-by":"publisher","first-page":"694","DOI":"10.18653\/v1\/S17-2115","article-title":"SiTAKA at SemEval-2017 task 4: Sentiment\n                        analysis in Twitter based on a rich set of features","volume-title":"Proceedings of the 11th International Workshop on Semantic\n                        Evaluation (SemEval-2017)","author":"Jabreel","year":"2017"},{"key":"2024081413395243800_bib21","article-title":"Practical text classification with large\n                        pre-trained language models","author":"Kant","year":"2018","journal-title":"arXiv preprint\n                        arXiv:1812.01207"},{"key":"2024081413395243800_bib22","first-page":"22443","article-title":"Achieving forgetting prevention and knowledge transfer in\n                        continual learning","volume-title":"Advances in Neural\n                        Information Processing Systems","author":"Ke","year":"2021"},{"key":"2024081413395243800_bib23","doi-asserted-by":"publisher","first-page":"675","DOI":"10.18653\/v1\/S17-2112","article-title":"Tweester at SemEval-2017 task 4: Fusion of\n                        semantic- affective and pairwise classification models for sentiment\n                        analysis in Twitter","volume-title":"Proceedings of the 11th\n                        International Workshop on Semantic Evaluation (SemEval-2017)","author":"Kolovou","year":"2017"},{"key":"2024081413395243800_bib24","doi-asserted-by":"publisher","first-page":"820","DOI":"10.1007\/s10618-014-0382-x","article-title":"Evaluation measures for hierarchical\n                        classification: A unified view and novel approaches","volume":"29","author":"Kosmopoulos","year":"2015","journal-title":"Data Mining and Knowledge Discovery"},{"key":"2024081413395243800_bib25","doi-asserted-by":"publisher","first-page":"216","DOI":"10.1016\/j.patcog.2019.02.023","article-title":"The impact of class imbalance in\n                        classification performance metrics based on the binary confusion\n                        matrix","volume":"91","author":"Luque","year":"2019","journal-title":"Pattern Recognition"},{"key":"2024081413395243800_bib26","volume-title":"An introduction to information\n                    retrieval","author":"Manning","year":"2009"},{"key":"2024081413395243800_bib27","doi-asserted-by":"publisher","first-page":"771","DOI":"10.18653\/v1\/S17-2130","article-title":"INGEOTEC at SemEval 2017 task 4: A B4MSA\n                        ensemble based on genetic programming for Twitter sentiment\n                        analysis","volume-title":"Proceedings of the 11th International\n                        Workshop on Semantic Evaluation (SemEval-2017)","author":"Miranda-Jim\u00e9nez","year":"2017"},{"key":"2024081413395243800_bib28","first-page":"5000","article-title":"Cooking up a neural-based model for recipe\n                        classification","volume-title":"Proceedings of the Twelfth\n                        Language Resources and Evaluation Conference","author":"Mohammadi","year":"2020"},{"key":"2024081413395243800_bib29","doi-asserted-by":"publisher","first-page":"777","DOI":"10.3389\/fpsyg.2017.00777","article-title":"An overview of interrater agreement on\n                        likert scales for researchers and practitioners","volume":"8","author":"O\u2019Neill","year":"2017","journal-title":"Frontiers in psychology"},{"key":"2024081413395243800_bib30","article-title":"Macro f1 and macro f1","author":"Opitz","year":"2019","journal-title":"arXiv preprint arXiv:1911.03347"},{"key":"2024081413395243800_bib31","doi-asserted-by":"publisher","first-page":"529","DOI":"10.13140\/RG.2.1.3754.1926","article-title":"Recall and precision versus the\n                        bookmaker","volume-title":"Cognitive Science - COGSCI","author":"Powers","year":"2003"},{"issue":"1","key":"2024081413395243800_bib32","first-page":"37","article-title":"Evaluation: From precision, recall and\n                        f-measure to roc, informedness, markedness &\n                        correlation","volume":"2","author":"Powers","year":"2011","journal-title":"Journal of Machine Learning\n                        Technologies"},{"key":"2024081413395243800_bib33","first-page":"345","article-title":"The problem with kappa","volume-title":"Proceedings of the 13th Conference of the European Chapter of the\n                        Association for Computational Linguistics","author":"Powers","year":"2012"},{"key":"2024081413395243800_bib34","article-title":"What the f-measure doesn\u2019t measure:\n                        Features, flaws, fallacies and fixes","author":"Powers","year":"2015","journal-title":"arXiv preprint\n                        arXiv:1503.06410"},{"key":"2024081413395243800_bib35","doi-asserted-by":"publisher","first-page":"4902","DOI":"10.18653\/v1\/2020.acl-main.442","article-title":"Beyond accuracy: Behavioral testing of NLP\n                        models with CheckList","volume-title":"Proceedings of the 58th\n                        Annual Meeting of the Association for Computational Linguistics","author":"Ribeiro","year":"2020"},{"key":"2024081413395243800_bib36","first-page":"6859","article-title":"Transferring confluent knowledge to\n                        argument mining","volume-title":"Proceedings of the 29th\n                        International Conference on Computational Linguistics","author":"Rodrigues","year":"2022"},{"key":"2024081413395243800_bib37","doi-asserted-by":"publisher","first-page":"502","DOI":"10.18653\/v1\/S17-2088","article-title":"Semeval-2017 task 4: Sentiment analysis in\n                        twitter","volume-title":"Proceedings of the 11th international\n                        workshop on semantic evaluation (SemEval-2017)","author":"Rosenthal","year":"2017"},{"key":"2024081413395243800_bib38","article-title":"NEATCLasS 2023: The 2nd workshop on novel\n                        evaluation approaches for text classification systems","author":"Ross","year":"2023","journal-title":"Workshop Proceedings of the 17th International AAAI Conference on\n                        Web and Social Media"},{"key":"2024081413395243800_bib39","doi-asserted-by":"publisher","first-page":"760","DOI":"10.18653\/v1\/S17-2128","article-title":"LIA at SemEval-2017 task 4: An ensemble of\n                        neural networks for sentiment classification","volume-title":"Proceedings of the 11th International Workshop on Semantic\n                        Evaluation (SemEval-2017)","author":"Rouvier","year":"2017"},{"key":"2024081413395243800_bib40","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1145\/2808194.2809449","article-title":"An axiomatically derived measure for the\n                        evaluation of classification algorithms","volume-title":"Proceedings of the 2015 International Conference on The Theory of\n                        Information Retrieval","author":"Sebastiani","year":"2015"},{"issue":"4","key":"2024081413395243800_bib41","doi-asserted-by":"publisher","first-page":"427","DOI":"10.1016\/j.ipm.2009.03.002","article-title":"A systematic analysis of performance\n                        measures for classification tasks","volume":"45","author":"Sokolova","year":"2009","journal-title":"Information\n                        Processing & Management"},{"issue":"3","key":"2024081413395243800_bib42","doi-asserted-by":"publisher","first-page":"619","DOI":"10.1162\/COLI_a_00295","article-title":"Parsing argumentation structures in\n                        persuasive essays","volume":"43","author":"Stab","year":"2017","journal-title":"Computational\n                        Linguistics"},{"issue":"1","key":"2024081413395243800_bib43","doi-asserted-by":"publisher","first-page":"168","DOI":"10.1016\/j.aci.2018.08.003","article-title":"Classification assessment\n                        methods","volume":"17","author":"Tharwat","year":"2020","journal-title":"Applied Computing and Informatics"},{"key":"2024081413395243800_bib44","doi-asserted-by":"publisher","first-page":"39","DOI":"10.18653\/v1\/S18-1005","article-title":"SemEval-2018 task 3: Irony detection in\n                        English tweets","volume-title":"Proceedings of The 12th\n                        International Workshop on Semantic Evaluation","author":"Van Hee","year":"2018"},{"key":"2024081413395243800_bib45","doi-asserted-by":"publisher","first-page":"5347","DOI":"10.18653\/v1\/2023.findings-emnlp.355","article-title":"Don\u2019t waste a single annotation: Improving\n                        single-label classifiers through soft labels","volume-title":"Findings of the Association for Computational Linguistics: EMNLP\n                        2023","author":"Ben","year":"2023"},{"key":"2024081413395243800_bib46","doi-asserted-by":"publisher","first-page":"978","DOI":"10.18653\/v1\/2020.coling-main.85","article-title":"Financial sentiment analysis: An\n                        investigation into common mistakes and silver bullets","volume-title":"Proceedings of the 28th International Conference on Computational\n                        Linguistics","author":"Xing","year":"2020"},{"key":"2024081413395243800_bib47","doi-asserted-by":"publisher","first-page":"65","DOI":"10.18653\/v1\/W19-8608","article-title":"Towards coherent and engaging spoken\n                        dialog response generation using automatic conversation\n                        evaluators","volume-title":"Proceedings of the 12th International\n                        Conference on Natural Language Generation","author":"Yi","year":"2019"},{"key":"2024081413395243800_bib48","doi-asserted-by":"publisher","first-page":"621","DOI":"10.18653\/v1\/S17-2102","article-title":"NNEMBs at SemEval-2017 task 4: Neural Twitter\n                        sentiment classification: a simple ensemble method with different\n                        embeddings","volume-title":"Proceedings of the 11th International\n                        Workshop on Semantic Evaluation (SemEval-2017)","author":"Yin","year":"2017"},{"key":"2024081413395243800_bib49","doi-asserted-by":"publisher","first-page":"645","DOI":"10.1145\/2187980.2188169","article-title":"Enhancing naive bayes with various\n                        smoothing methods for short text classification","volume-title":"Proceedings of the 21st International Conference on World Wide\n                        Web","author":"Yuan","year":"2012"},{"key":"2024081413395243800_bib50","doi-asserted-by":"publisher","first-page":"75","DOI":"10.18653\/v1\/S19-2010","article-title":"SemEval-2019 task 6: Identifying and\n                        categorizing offensive language in social media\n                    (OffensEval)","volume-title":"Proceedings of the 13th International\n                        Workshop on Semantic Evaluation","author":"Zampieri","year":"2019"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00675\/2465598\/tacl_a_00675.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00675\/2465598\/tacl_a_00675.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,14]],"date-time":"2024-08-14T13:40:18Z","timestamp":1723642818000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00675\/122720\/A-Closer-Look-at-Classification-Evaluation-Metrics"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":50,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00675","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}