{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,25]],"date-time":"2026-07-25T16:09:58Z","timestamp":1784995798530,"version":"3.55.0"},"reference-count":113,"publisher":"MIT Press","issue":"1","license":[{"start":{"date-parts":[[2024,9,19]],"date-time":"2024-09-19T00:00:00Z","timestamp":1726704000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,3,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Recent machine translation (MT) metrics calibrate their effectiveness by correlating with human judgment. However, these results are often obtained by averaging predictions across large test sets without any insights into the strengths and weaknesses of these metrics across different error types. Challenge sets are used to probe specific dimensions of metric behavior but there are very few such datasets and they either focus on a limited number of phenomena or a limited number of language pairs. We introduce ACES, a contrastive challenge set spanning 146 language pairs, aimed at discovering whether metrics can identify 68 translation accuracy errors. These phenomena range from basic alterations at the word\/character level to more intricate errors based on discourse and real-world knowledge. We conducted a large-scale study by benchmarking ACES on 47 metrics submitted to the WMT 2022 and WMT 2023 metrics shared tasks. We also measure their sensitivity to a range of linguistic phenomena. We further investigate claims that large language models (LLMs) are effective as MT evaluators, addressing the limitations of previous studies by using a dataset that covers a range of linguistic phenomena and language pairs and includes both low- and medium-resource languages. Our results demonstrate that different metric families struggle with different phenomena and that LLM-based methods are unreliable. We expose a number of major flaws with existing methods: Most metrics ignore the source sentence; metrics tend to prefer surface level overlap; and over-reliance on language-agnostic representations leads to confusion when the target language is similar to the source language. To further encourage detailed evaluation beyond singular scores, we expand ACES to include error span annotations, denoted as SPAN-ACES, and we use this dataset to evaluate span-based error metrics, showing that these metrics also need considerable improvement. Based on our observations, we provide a set of recommendations for building better MT metrics, including focusing on error labels instead of scores, ensembling, designing metrics to explicitly focus on the source sentence, focusing on semantic content rather than relying on the lexical overlap, and choosing the right pre-trained model for obtaining representations.<\/jats:p>","DOI":"10.1162\/coli_a_00537","type":"journal-article","created":{"date-parts":[[2024,9,19]],"date-time":"2024-09-19T15:59:59Z","timestamp":1726761599000},"page":"73-137","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":8,"title":["Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets"],"prefix":"10.1162","volume":"51","author":[{"given":"Nikita","family":"Moghe","sequence":"first","affiliation":[{"name":"University of Edinburgh, School of Informatics. nikitamoghe29@gmail.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Arnisa","family":"Fazla","sequence":"additional","affiliation":[{"name":"University of Zurich, Department of Computational Linguistics. arnisa.fazla@uzh.ch"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chantal","family":"Amrhein","sequence":"additional","affiliation":[{"name":"Supertext. chantal@supertext.ch"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tom","family":"Kocmi","sequence":"additional","affiliation":[{"name":"Microsoft. tom.kocmi@microsoft.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mark","family":"Steedman","sequence":"additional","affiliation":[{"name":"University of Edinburgh, School of Informatics. m.steedman@ed.ac.uk"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alexandra","family":"Birch","sequence":"additional","affiliation":[{"name":"University of Edinburgh, School of Informatics. a.birch@ed.ac.uk"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rico","family":"Sennrich","sequence":"additional","affiliation":[{"name":"University of Zurich, Department of Computational Linguistics. sennrich@cl.uzh.ch"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Liane","family":"Guillou","sequence":"additional","affiliation":[{"name":"University of Edinburgh, School of Informatics. liane.guillou@ed.ac.uk"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2025,3,15]]},"reference":[{"key":"2025032113511448100_bib1","first-page":"469","article-title":"Robust MT evaluation with sentence-level multilingual augmentation","volume-title":"Proceedings of the Seventh Conference on Machine Translation (WMT)","author":"Alves","year":"2022"},{"key":"2025032113511448100_bib2","doi-asserted-by":"publisher","first-page":"479","DOI":"10.18653\/v1\/2023.wmt-1.57","article-title":"ACES: Translation accuracy challenge sets for evaluating machine translation metrics","volume-title":"Proceedings of the Seventh Conference on Machine Translation (WMT)","author":"Amrhein","year":"2022"},{"key":"2025032113511448100_bib3","doi-asserted-by":"publisher","first-page":"693","DOI":"10.18653\/v1\/2023.wmt-1.57","article-title":"ACES: Translation accuracy challenge sets at WMT 2023","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"Amrhein","year":"2023"},{"key":"2025032113511448100_bib4","first-page":"1125","article-title":"Identifying weaknesses in machine translation metrics through minimum Bayes risk decoding: A case study for COMET","volume-title":"2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing","author":"Amrhein","year":"2022"},{"key":"2025032113511448100_bib5","first-page":"514","article-title":"Linguistically motivated evaluation of machine translation metrics based on a challenge set","volume-title":"Proceedings of the Seventh Conference on Machine Translation (WMT)","author":"Avramidis","year":"2022"},{"key":"2025032113511448100_bib6","first-page":"243","article-title":"Fine-grained evaluation of quality estimation for machine translation based on a linguistically motivated test suite","volume-title":"Proceedings of the AMTA 2018 Workshop on Translation Quality Estimation and Automatic Post-Editing","author":"Avramidis","year":"2018"},{"key":"2025032113511448100_bib7","doi-asserted-by":"publisher","first-page":"713","DOI":"10.18653\/v1\/2023.wmt-1.58","article-title":"Challenging the state-of-the-art machine translation metrics from a linguistic perspective","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"Avramidis","year":"2023"},{"key":"2025032113511448100_bib8","first-page":"887","article-title":"ParBLEU: Augmenting metrics with automatic paraphrases for the WMT\u2019(20 metrics shared task","volume-title":"Proceedings of the Fifth Conference on Machine Translation","author":"Bawden","year":"2020"},{"key":"2025032113511448100_bib9","doi-asserted-by":"publisher","first-page":"257","DOI":"10.18653\/v1\/D16-1025","article-title":"Neural versus phrase-based machine translation quality: A case study","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing","author":"Bentivogli","year":"2016"},{"key":"2025032113511448100_bib10","doi-asserted-by":"publisher","first-page":"131","DOI":"10.18653\/v1\/W16-2301","article-title":"Findings of the 2016 conference on machine translation","volume-title":"Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers","author":"Bojar","year":"2016"},{"key":"2025032113511448100_bib11","doi-asserted-by":"publisher","first-page":"45","DOI":"10.3384\/ecp197006","article-title":"Experiments on automatic error detection and correction for Uruguayan learners of English","volume-title":"Proceedings of the 12th Workshop on NLP for Computer Assisted Language Learning","author":"Brown","year":"2023"},{"key":"2025032113511448100_bib12","first-page":"249","article-title":"Re-evaluating the role of Bleu in machine translation research","volume-title":"11th Conference of the European Chapter of the Association for Computational Linguistics","author":"Callison-Burch","year":"2006"},{"key":"2025032113511448100_bib13","doi-asserted-by":"publisher","first-page":"4331","DOI":"10.18653\/v1\/2022.acl-long.298","article-title":"DiBiMT: A novel benchmark for measuring Word Sense Disambiguation biases in Machine Translation","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Campolungo","year":"2022"},{"key":"2025032113511448100_bib14","first-page":"19","article-title":"Extracting training data from large language models","volume-title":"USENIX Security Symposium","author":"Carlini","year":"2020"},{"key":"2025032113511448100_bib15","doi-asserted-by":"publisher","first-page":"109","DOI":"10.1515\/pralin-2017-0013","article-title":"Is neural machine translation the new state of the art?","volume":"108","author":"Castilho","year":"2017","journal-title":"The Prague Bulletin of Mathematical Linguistics"},{"key":"2025032113511448100_bib16","first-page":"530","article-title":"Exploring robustness of machine translation metrics: A study of twenty-two automatic metrics in the WMT22 metric task","volume-title":"Proceedings of the Seventh Conference on Machine Translation (WMT)","author":"Chen","year":"2022"},{"key":"2025032113511448100_bib17","article-title":"INSTRUCTEVAL: Towards holistic evaluation of instruction-tuned large language models","volume":"arXiv:2306.04757","author":"Chia","year":"2023","journal-title":"Computing Research Repository"},{"key":"2025032113511448100_bib18","doi-asserted-by":"publisher","first-page":"15607","DOI":"10.18653\/v1\/2023.acl-long.870","article-title":"Can large language models be an alternative to human evaluations?","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chiang","year":"2023"},{"key":"2025032113511448100_bib19","article-title":"Scaling instruction-finetuned language models","volume":"arXiv:2210.11416","author":"Chung","year":"2022","journal-title":"Computing Research Repository"},{"key":"2025032113511448100_bib20","doi-asserted-by":"publisher","first-page":"2475","DOI":"10.18653\/v1\/D18-1269","article-title":"XNLI: Evaluating cross-lingual sentence representations","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Conneau","year":"2018"},{"key":"2025032113511448100_bib21","doi-asserted-by":"publisher","first-page":"36","DOI":"10.18653\/v1\/2023.acl-long.3","article-title":"Detecting and mitigating hallucinations in machine translation: Model internal workings alone do well, sentence similarity even better","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Dale","year":"2023"},{"key":"2025032113511448100_bib22","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2025032113511448100_bib23","doi-asserted-by":"publisher","first-page":"736","DOI":"10.18653\/v1\/2023.wmt-1.60","article-title":"Embed_llama: Using LLM embeddings for the metrics shared task","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"Dr\u00e9ano","year":"2023"},{"key":"2025032113511448100_bib24","doi-asserted-by":"publisher","first-page":"728","DOI":"10.18653\/v1\/2023.wmt-1.59","article-title":"Tokengram_F, a fast and accurate token-based chrf ++ derivative","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"Dr\u00e9ano","year":"2023"},{"key":"2025032113511448100_bib25","first-page":"70293","article-title":"Faith and fate: Limits of transformers on compositionality","volume-title":"Thirty-seventh Conference on Neural Information Processing Systems","author":"Dziri","year":"2023"},{"key":"2025032113511448100_bib26","doi-asserted-by":"publisher","first-page":"744","DOI":"10.18653\/v1\/2023.wmt-1.61","article-title":"eBLEU: Unexpectedly good machine translation evaluation using simple word embeddings","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"ElNokrashy","year":"2023"},{"key":"2025032113511448100_bib27","doi-asserted-by":"publisher","first-page":"8517","DOI":"10.18653\/v1\/2021.emnlp-main.670","article-title":"Wino-X: Multilingual Winograd schemas for commonsense reasoning and coreference resolution","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Emelin","year":"2021"},{"issue":"107","key":"2025032113511448100_bib28","first-page":"1","article-title":"Beyond English-centric multilingual machine translation","volume":"22","author":"Fan","year":"2021","journal-title":"Journal of Machine Learning Research"},{"key":"2025032113511448100_bib29","doi-asserted-by":"publisher","first-page":"1066","DOI":"10.18653\/v1\/2023.wmt-1.100","article-title":"The devil is in the errors: Leveraging large language models for fine-grained machine translation evaluation","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"Fernandes","year":"2023"},{"key":"2025032113511448100_bib30","doi-asserted-by":"publisher","first-page":"1460","DOI":"10.1162\/tacl_a_00437","article-title":"Experts, errors, and context: A large-scale study of human evaluation for machine translation","volume":"9","author":"Freitag","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2025032113511448100_bib31","doi-asserted-by":"publisher","first-page":"578","DOI":"10.18653\/v1\/2023.wmt-1.51","article-title":"Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"Freitag","year":"2023"},{"key":"2025032113511448100_bib32","first-page":"46","article-title":"Results of WMT22 metrics shared task: Stop using BLEU \u2013 neural metrics are better and more robust","volume-title":"Proceedings of the Seventh Conference on Machine Translation (WMT)","author":"Freitag","year":"2022"},{"key":"2025032113511448100_bib33","first-page":"733","article-title":"Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain","volume-title":"Proceedings of the Sixth Conference on Machine Translation","author":"Freitag","year":"2021"},{"key":"2025032113511448100_bib34","doi-asserted-by":"publisher","first-page":"749","DOI":"10.18653\/v1\/2023.wmt-1.62","article-title":"Cometoid: Distilling strong reference-based machine translation metrics into even stronger quality estimation metrics","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"Gowda","year":"2023"},{"key":"2025032113511448100_bib35","doi-asserted-by":"publisher","first-page":"522","DOI":"10.1162\/tacl_a_00474","article-title":"The Flores-101 evaluation benchmark for low-resource and multilingual machine translation","volume":"10","author":"Goyal","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2025032113511448100_bib36","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00683","article-title":"xCOMET: Transparent machine translation evaluation through fine-grained error detection","volume":"arXiv:2310.10482","author":"Guerreiro","year":"2023","journal-title":"Computing Research Repository"},{"key":"2025032113511448100_bib37","first-page":"636","article-title":"PROTEST: A test suite for evaluating pronouns in machine translation","volume-title":"Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC\u201916)","author":"Guillou","year":"2016"},{"key":"2025032113511448100_bib38","doi-asserted-by":"publisher","first-page":"570","DOI":"10.18653\/v1\/W18-6435","article-title":"A pronoun test suite evaluation of the English\u2013German MT systems at WMT 2018","volume-title":"Proceedings of the Third Conference on Machine Translation: Shared Task Papers","author":"Guillou","year":"2018"},{"key":"2025032113511448100_bib39","first-page":"507","article-title":"A fine-grained analysis of BERTScore","volume-title":"Proceedings of the Sixth Conference on Machine Translation","author":"Hanna","year":"2021"},{"key":"2025032113511448100_bib40","doi-asserted-by":"publisher","first-page":"2486","DOI":"10.18653\/v1\/D17-1263","article-title":"A challenge set approach to evaluating machine translation","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing","author":"Isabelle","year":"2017"},{"issue":"12","key":"2025032113511448100_bib41","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3571730","article-title":"Survey of hallucination in natural language generation","volume":"55","author":"Ji","year":"2023","journal-title":"ACM Computing Surveys"},{"key":"2025032113511448100_bib42","doi-asserted-by":"publisher","first-page":"2021","DOI":"10.18653\/v1\/D17-1215","article-title":"Adversarial examples for evaluating reading comprehension systems","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing","author":"Jia","year":"2017"},{"key":"2025032113511448100_bib43","article-title":"Tigerscore: Towards building explainable metric for all text generation tasks","volume":"arxiv:2310.00752","author":"Jiang","year":"2023","journal-title":"Computing Research Repository"},{"key":"2025032113511448100_bib44","doi-asserted-by":"publisher","first-page":"754","DOI":"10.18653\/v1\/2023.wmt-1.63","article-title":"Metricx-23: The Google submission to the WMT 2023 metrics shared task","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"Juraska","year":"2023"},{"key":"2025032113511448100_bib45","doi-asserted-by":"publisher","first-page":"9540","DOI":"10.18653\/v1\/2022.emnlp-main.649","article-title":"DEMETR: Diagnosing evaluation metrics for translation","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Karpinska","year":"2022"},{"key":"2025032113511448100_bib46","doi-asserted-by":"publisher","first-page":"252","DOI":"10.18653\/v1\/N18-1023","article-title":"Looking beyond the surface: A challenge set for reading comprehension over multiple sentences","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Khashabi","year":"2018"},{"key":"2025032113511448100_bib47","doi-asserted-by":"publisher","DOI":"10.3115\/997939.997976","article-title":"Using test suites in evaluation of machine translation systems","volume-title":"COLING 1990 Volume 2: Papers presented to the 13th International Conference on Computational Linguistics","author":"King","year":"1990"},{"key":"2025032113511448100_bib48","doi-asserted-by":"publisher","first-page":"766","DOI":"10.18653\/v1\/2023.wmt-1.64","article-title":"GEMBA-MQM: Detecting translation quality error spans with GPT-4","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"Kocmi","year":"2023"},{"key":"2025032113511448100_bib49","first-page":"193","article-title":"Large language models are state-of-the-art evaluators of translation quality","volume-title":"Proceedings of the 24th Annual Conference of the European Association for Machine Translation","author":"Kocmi","year":"2023"},{"key":"2025032113511448100_bib50","first-page":"478","article-title":"To ship or not to ship: An extensive evaluation of automatic metrics for machine translation","volume-title":"Proceedings of the Sixth Conference on Machine Translation","author":"Kocmi","year":"2021"},{"key":"2025032113511448100_bib51","first-page":"541","article-title":"MS-COMET: More and Better Human Judgements Improve Metric Performance","volume-title":"Proceedings of the Seventh Conference on Machine Translation","author":"Kocmi","year":"2022"},{"key":"2025032113511448100_bib52","first-page":"79","article-title":"Europarl: A parallel corpus for statistical machine translation","volume-title":"Proceedings of Machine Translation Summit X: Papers","author":"Koehn","year":"2005"},{"key":"2025032113511448100_bib53","doi-asserted-by":"publisher","first-page":"102","DOI":"10.3115\/1654650.1654666","article-title":"Manual and automatic evaluation of machine translation between European languages","volume-title":"Proceedings on the Workshop on Statistical Machine Translation","author":"Koehn","year":"2006"},{"key":"2025032113511448100_bib54","doi-asserted-by":"publisher","first-page":"66","DOI":"10.18653\/v1\/D18-2012","article-title":"SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Kudo","year":"2018"},{"key":"2025032113511448100_bib55","doi-asserted-by":"publisher","first-page":"407","DOI":"10.26615\/978-954-452-049-6_054","article-title":"Improving discourse relation projection to build discourse annotated corpora","volume-title":"Proceedings of the International Conference Recent Advances in Natural Language Processing, RANLP 2017","author":"Laali","year":"2017"},{"key":"2025032113511448100_bib56","article-title":"ParCorFull: A parallel corpus annotated with full coreference","volume-title":"Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)","author":"Lapshinova-Koltunski","year":"2018"},{"issue":"75","key":"2025032113511448100_bib57","first-page":"1","article-title":"Towards explainable evaluation metrics for machine translation","volume":"25","author":"Leiter","year":"2024","journal-title":"Journal of Machine Learning Research"},{"key":"2025032113511448100_bib58","article-title":"PRD: Peer rank and discussion improve large language model based evaluations","author":"Li","year":"2023","journal-title":"Computing Research Repository"},{"key":"2025032113511448100_bib59","doi-asserted-by":"publisher","first-page":"27","DOI":"10.18653\/v1\/W17-5404","article-title":"BIBI system description: Building with CNNs and breaking with deep reinforcement learning","volume-title":"Proceedings of the First Workshop on Building Linguistically Generalizable NLP Systems","author":"Li","year":"2017"},{"key":"2025032113511448100_bib60","article-title":"Partial Could Be Better Than Whole: HW-TSC 2022 Submission for the Metrics Shared Task","volume-title":"Proceedings of the Seventh Conference on Machine Translation","author":"Liu","year":"2022"},{"key":"2025032113511448100_bib61","doi-asserted-by":"publisher","first-page":"507","DOI":"10.18653\/v1\/W19-5358","article-title":"YiSi - a unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources","volume-title":"Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1)","author":"Lo","year":"2019"},{"key":"2025032113511448100_bib62","doi-asserted-by":"publisher","first-page":"774","DOI":"10.18653\/v1\/2023.wmt-1.65","article-title":"Metric score landscape challenge (MSLC23): Understanding metrics\u2019 performance on a wider landscape of translation quality","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"Lo","year":"2023"},{"key":"2025032113511448100_bib63","doi-asserted-by":"publisher","first-page":"455","DOI":"10.5565\/rev\/tradumatica.77","article-title":"Multidimensional quality metrics (MQM): A framework for declaring and describing translation quality metrics","volume":"0","author":"Lommel","year":"2014","journal-title":"Tradum\u00e0tica: Tecnologies de la Traducci\u00f3"},{"key":"2025032113511448100_bib64","doi-asserted-by":"publisher","DOI":"10.20944\/preprints202303.0255.v1","article-title":"Error analysis prompting enables human-like translation evaluation in large language models: A case study on ChatGPT","volume":"arXiv:2303.13809","author":"Lu","year":"2023","journal-title":"Computing Research Repository"},{"key":"2025032113511448100_bib65","first-page":"432","article-title":"Linguistically motivated evaluation of the 2022 state-of-the-art machine translation systems for three language directions","volume-title":"Proceedings of the Seventh Conference on Machine Translation (WMT)","author":"Macketanz","year":"2022"},{"key":"2025032113511448100_bib66","doi-asserted-by":"publisher","first-page":"33","DOI":"10.18653\/v1\/W17-5405","article-title":"Breaking NLP: Using morphosyntax, semantics, pragmatics and world knowledge to fool sentiment analysis systems","volume-title":"Proceedings of the First Workshop on Building Linguistically Generalizable NLP Systems","author":"Mahler","year":"2017"},{"key":"2025032113511448100_bib67","first-page":"358","article-title":"Non-entailed subsequences as a challenge for natural language inference","author":"McCoy","year":"2019","journal-title":"Proceedings of the Society for Computation in Linguistics (SCiL)"},{"key":"2025032113511448100_bib68","doi-asserted-by":"publisher","first-page":"13060","DOI":"10.18653\/v1\/2023.acl-long.730","article-title":"Extrinsic evaluation of machine translation metrics","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Moghe","year":"2023"},{"key":"2025032113511448100_bib69","doi-asserted-by":"publisher","first-page":"798","DOI":"10.18653\/v1\/2023.wmt-1.66","article-title":"MEE4 and XLsim: IIIT HYD\u2019s submissions\u2019 for WMT23 metrics shared task","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"Mukherjee","year":"2023"},{"key":"2025032113511448100_bib70","article-title":"No language left behind: Scaling human-centered machine translation","volume":"arXiv:2207.04672","author":"Team","year":"2022","journal-title":"Computing Research Repository"},{"key":"2025032113511448100_bib71","doi-asserted-by":"publisher","first-page":"311","DOI":"10.3115\/1073083.1073135","article-title":"Bleu: A method for automatic evaluation of machine translation","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni","year":"2002"},{"key":"2025032113511448100_bib72","first-page":"569","article-title":"MaTESe: Machine translation evaluation as a sequence tagging problem","volume-title":"Proceedings of the Seventh Conference on Machine Translation (WMT)","author":"Perrella","year":"2022"},{"key":"2025032113511448100_bib73","first-page":"569","article-title":"Machine translation evaluation as a sequence tagging problem","volume-title":"Proceedings of the Seventh Conference on Machine Translation","author":"Perrella","year":"2022"},{"key":"2025032113511448100_bib74","doi-asserted-by":"publisher","first-page":"612","DOI":"10.18653\/v1\/W17-4770","article-title":"chrF++: words helping character n-grams","volume-title":"Proceedings of the Second Conference on Machine Translation","author":"Popovi\u0107","year":"2017"},{"key":"2025032113511448100_bib75","article-title":"Challenge test sets for MT evaluation","volume-title":"Proceedings of Machine Translation Summit XVII: Tutorial Abstracts","author":"Popovi\u0107","year":"2019"},{"key":"2025032113511448100_bib76","doi-asserted-by":"publisher","first-page":"470","DOI":"10.18653\/v1\/W19-5354","article-title":"The MuCoW test suite at WMT 2019: Automatically harvested multilingual contrastive word sense disambiguation test sets for machine translation","volume-title":"Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1)","author":"Raganato","year":"2019"},{"key":"2025032113511448100_bib77","doi-asserted-by":"publisher","first-page":"2976","DOI":"10.18653\/v1\/2021.eacl-main.259","article-title":"NoiseQA: Challenge set evaluation for user-centric question answering","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Ravichander","year":"2021"},{"key":"2025032113511448100_bib78","article-title":"COMET-22: Unbabel-IST 2022 submission for the metrics shared task","volume-title":"Proceedings of the Seventh Conference on Machine Translation","author":"Rei","year":"2022"},{"key":"2025032113511448100_bib79","doi-asserted-by":"publisher","first-page":"1089","DOI":"10.18653\/v1\/2023.acl-short.94","article-title":"The inside story: Towards better understanding of machine translation neural evaluation metrics","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Rei","year":"2023"},{"key":"2025032113511448100_bib80","doi-asserted-by":"publisher","first-page":"2685","DOI":"10.18653\/v1\/2020.emnlp-main.213","article-title":"COMET: A neural framework for MT evaluation","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Rei","year":"2020"},{"key":"2025032113511448100_bib81","doi-asserted-by":"publisher","first-page":"813","DOI":"10.3115\/1699571.1699619","article-title":"Unbounded dependency recovery for parser evaluation","volume-title":"Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing","author":"Rimell","year":"2009"},{"key":"2025032113511448100_bib82","doi-asserted-by":"publisher","first-page":"588","DOI":"10.18653\/v1\/W18-6437","article-title":"The word sense disambiguation test suite at WMT18","volume-title":"Proceedings of the Third Conference on Machine Translation: Shared Task Papers","author":"Rios","year":"2018"},{"key":"2025032113511448100_bib83","first-page":"7","article-title":"FANCY: A diagnostic data-set for NLI models","volume-title":"Proceedings of the Eighth Italian Conference on Computational Linguistics (CLiC-it)","author":"Rocchietti","year":"2021"},{"key":"2025032113511448100_bib84","doi-asserted-by":"publisher","first-page":"74","DOI":"10.18653\/v1\/W17-1609","article-title":"Social bias in elicited natural language inferences","volume-title":"Proceedings of the First ACL Workshop on Ethics in Natural Language Processing","author":"Rudinger","year":"2017"},{"key":"2025032113511448100_bib85","article-title":"BLOOM: A 176B-parameter open-access multilingual language model","volume":"arxiv:2211.05100","author":"Scao","year":"2022","journal-title":"Computing Research Repository"},{"key":"2025032113511448100_bib86","first-page":"921","article-title":"Learning to evaluate translation beyond English: BLEURT submissions to the WMT metrics 2020 shared task","volume-title":"Proceedings of the Fifth Conference on Machine Translation","author":"Sellam","year":"2020"},{"key":"2025032113511448100_bib87","doi-asserted-by":"publisher","first-page":"1715","DOI":"10.18653\/v1\/P16-1162","article-title":"Neural machine translation of rare words with subword units","volume-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Sennrich","year":"2016"},{"key":"2025032113511448100_bib88","doi-asserted-by":"publisher","first-page":"2888","DOI":"10.18653\/v1\/2021.emnlp-main.230","article-title":"Masked language modeling and the distributional hypothesis: Order word matters pre-training for little","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Sinha","year":"2021"},{"key":"2025032113511448100_bib89","article-title":"Adversarial evaluation for models of natural language","volume":"arXiv:1207.0245","author":"Smith","year":"2012","journal-title":"Computing Research Repository"},{"key":"2025032113511448100_bib90","first-page":"2204","article-title":"Benchmarking the performance of machine translation evaluation metrics with Chinese multiword expressions","volume-title":"Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)","author":"Song","year":"2024"},{"key":"2025032113511448100_bib91","doi-asserted-by":"publisher","first-page":"8776","DOI":"10.18653\/v1\/2023.emnlp-main.543","article-title":"Evaluation metrics in the era of GPT-4: Reliably evaluating large language models on sequence to sequence tasks","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Sottana","year":"2023"},{"key":"2025032113511448100_bib92","doi-asserted-by":"publisher","first-page":"76","DOI":"10.18653\/v1\/W19-5303","article-title":"Findings of the WMT 2020 shared task on machine translation robustness","volume-title":"Proceedings of the Fifth Conference on Machine Translation","author":"Specia","year":"2020"},{"key":"2025032113511448100_bib93","doi-asserted-by":"publisher","first-page":"61","DOI":"10.18653\/v1\/W17-5410","article-title":"Breaking sentiment analysis of movie reviews","volume-title":"Proceedings of the First Workshop on Building Linguistically Generalizable NLP Systems","author":"Stali\u016bnait\u0117","year":"2017"},{"key":"2025032113511448100_bib94","doi-asserted-by":"publisher","first-page":"1679","DOI":"10.18653\/v1\/P19-1164","article-title":"Evaluating gender bias in machine translation","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Stanovsky","year":"2019"},{"key":"2025032113511448100_bib95","doi-asserted-by":"publisher","first-page":"6596","DOI":"10.18653\/v1\/2024.naacl-long.367","article-title":"Not all metrics are guilty: Improving NLG evaluation by diversifying references","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Tang","year":"2024"},{"key":"2025032113511448100_bib96","first-page":"646","article-title":"CrossQE: HW-TSC 2022 submission for the quality estimation shared task","volume-title":"Proceedings of the Seventh Conference on Machine Translation (WMT)","author":"Tao","year":"2022"},{"key":"2025032113511448100_bib97","article-title":"Stanford alpaca: An instruction-following llama model","author":"Taori","year":"2023"},{"key":"2025032113511448100_bib98","doi-asserted-by":"publisher","first-page":"1063","DOI":"10.18653\/v1\/E17-1100","article-title":"A multifaceted evaluation of neural versus phrase-based machine translation for 9 language directions","volume-title":"Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers","author":"Toral","year":"2017"},{"key":"2025032113511448100_bib99","article-title":"Llama 2: Open foundation and fine-tuned chat models","author":"Touvron","year":"2023","journal-title":"Computing Research Repository"},{"key":"2025032113511448100_bib100","doi-asserted-by":"publisher","first-page":"10246","DOI":"10.18653\/v1\/2021.emnlp-main.803","article-title":"Contrastive conditioning for assessing disambiguation in MT: A case study of distilled bias","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Vamvas","year":"2021"},{"key":"2025032113511448100_bib101","doi-asserted-by":"publisher","first-page":"490","DOI":"10.18653\/v1\/2022.acl-short.53","article-title":"As little as possible, as much as necessary: Detecting over- and undertranslations with contrastive conditioning","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Vamvas","year":"2022"},{"issue":"11","key":"2025032113511448100_bib102","doi-asserted-by":"publisher","first-page":"1515","DOI":"10.1080\/1369118X.2020.1776370","article-title":"Understanding the societal impacts of machine translation: A critical review of the literature on medical and legal use cases","volume":"24","author":"Vieira","year":"2021","journal-title":"Information, Communication & Society"},{"key":"2025032113511448100_bib103","article-title":"Alibaba-Translate China\u2019s submission for WMT2022 Metrics Shared Task","volume-title":"Proceedings of the Seventh Conference on Machine Translation","author":"Wan","year":"2022"},{"key":"2025032113511448100_bib104","doi-asserted-by":"publisher","first-page":"8117","DOI":"10.18653\/v1\/2022.acl-long.558","article-title":"UniTE: Unified translation evaluation","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wan","year":"2022"},{"key":"2025032113511448100_bib105","article-title":"The generative AI paradox: What it can create, it may not understand","volume-title":"The Twelfth International Conference on Learning Representations","author":"West","year":"2024"},{"key":"2025032113511448100_bib106","doi-asserted-by":"publisher","first-page":"833","DOI":"10.18653\/v1\/D19-1077","article-title":"Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Wu","year":"2019"},{"key":"2025032113511448100_bib107","doi-asserted-by":"publisher","first-page":"820","DOI":"10.18653\/v1\/2023.wmt-1.70","article-title":"Empowering a metric with LLM-assisted named entity annotation: HW-TSC\u2019s submission to the WMT23 metrics shared task","volume-title":"Proceedings of the Eighth Conference on Machine Translation","author":"Wu","year":"2023"},{"key":"2025032113511448100_bib108","doi-asserted-by":"publisher","first-page":"5967","DOI":"10.18653\/v1\/2023.emnlp-main.365","article-title":"INSTRUCTSCORE: Towards explainable text generation evaluation with automatic feedback","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Xu","year":"2023"},{"key":"2025032113511448100_bib109","doi-asserted-by":"publisher","first-page":"3687","DOI":"10.18653\/v1\/D19-1382","article-title":"PAWS-X: A cross-lingual adversarial dataset for paraphrase identification","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Yang","year":"2019"},{"key":"2025032113511448100_bib110","first-page":"69","article-title":"Findings of the WMT 2022 shared task on quality estimation","volume-title":"Proceedings of the Seventh Conference on Machine Translation (WMT)","author":"Zerva","year":"2022"},{"key":"2025032113511448100_bib111","article-title":"BERTScore: Evaluating text generation with BERT","volume-title":"8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26\u201330, 2020","author":"Zhang","year":"2020"},{"key":"2025032113511448100_bib112","doi-asserted-by":"publisher","first-page":"15","DOI":"10.18653\/v1\/N18-2003","article-title":"Gender bias in coreference resolution: Evaluation and debiasing methods","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers)","author":"Zhao","year":"2018"},{"key":"2025032113511448100_bib113","doi-asserted-by":"publisher","first-page":"33","DOI":"10.18653\/v1\/2021.mwe-1.5","article-title":"PIE: A parallel idiomatic expression corpus for idiomatic sentence generation and paraphrasing","volume-title":"Proceedings of the 17th Workshop on Multiword Expressions (MWE 2021)","author":"Zhou","year":"2021"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/51\/1\/73\/2481445\/coli_a_00537.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/51\/1\/73\/2481445\/coli_a_00537.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,3,21]],"date-time":"2025-03-21T18:19:21Z","timestamp":1742581161000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/51\/1\/73\/124465\/Machine-Translation-Meta-Evaluation-through"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025]]},"references-count":113,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,3,15]]},"published-print":{"date-parts":[[2025,3,15]]}},"URL":"https:\/\/doi.org\/10.1162\/coli_a_00537","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"value":"0891-2017","type":"print"},{"value":"1530-9312","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2025]]},"published":{"date-parts":[[2025]]}}}