{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T23:35:24Z","timestamp":1785972924968,"version":"3.56.0"},"reference-count":35,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2021,12,23]],"date-time":"2021-12-23T00:00:00Z","timestamp":1640217600000},"content-version":"vor","delay-in-days":356,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,12,17]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Human evaluation of modern high-quality machine translation systems is a difficult problem, and there is increasing evidence that inadequate evaluation procedures can lead to erroneous conclusions. While there has been considerable research on human evaluation, the field still lacks a commonly accepted standard procedure. As a step toward this goal, we propose an evaluation methodology grounded in explicit error analysis, based on the Multidimensional Quality Metrics (MQM) framework. We carry out the largest MQM research study to date, scoring the outputs of top systems from the WMT 2020 shared task in two language pairs using annotations provided by professional translators with access to full document context. We analyze the resulting data extensively, finding among other results a substantially different ranking of evaluated systems from the one established by the WMT crowd workers, exhibiting a clear preference for human over machine output. Surprisingly, we also find that automatic metrics based on pre-trained embeddings can outperform human crowd workers. We make our corpus publicly available for further research.<\/jats:p>","DOI":"10.1162\/tacl_a_00437","type":"journal-article","created":{"date-parts":[[2021,12,24]],"date-time":"2021-12-24T05:57:54Z","timestamp":1640325474000},"page":"1460-1474","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":127,"title":["Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation"],"prefix":"10.1162","volume":"9","author":[{"given":"Markus","family":"Freitag","sequence":"first","affiliation":[{"name":"Google Research. freitag@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"George","family":"Foster","sequence":"additional","affiliation":[{"name":"Google Research. fosterg@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David","family":"Grangier","sequence":"additional","affiliation":[{"name":"Google Research. grangier@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Viresh","family":"Ratnakar","sequence":"additional","affiliation":[{"name":"Google Research. vratnakar@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qijun","family":"Tan","sequence":"additional","affiliation":[{"name":"Google Research. qijuntan@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wolfgang","family":"Macherey","sequence":"additional","affiliation":[{"name":"Google Research. wmach@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2021,12,17]]},"reference":[{"key":"2021122316152622700_bib1","volume-title":"Language and Machines: Computers in Translation and Linguistics; a Report","author":"ALPAC","year":"1966"},{"key":"2021122316152622700_bib2","first-page":"1127","article-title":"Involving Language Professionals in the Evaluation of Machine Translation","volume-title":"Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC\u201912)","author":"Avramidis","year":"2012"},{"key":"2021122316152622700_bib3","first-page":"1","article-title":"Findings of the 2020 Conference on Machine Translation (WMT20)","volume-title":"Proceedings of the Fifth Conference on Machine Translation","author":"Barrault","year":"2020"},{"key":"2021122316152622700_bib4","article-title":"Machine Translation Human Evaluation: An investigation of evaluation based on Post-Editing and its relation with Direct Assessment","volume-title":"International Workshop on Spoken Language Translation","author":"Bentivogli","year":"2018"},{"key":"2021122316152622700_bib5","doi-asserted-by":"crossref","first-page":"169","DOI":"10.18653\/v1\/W17-4717","article-title":"Findings of the 2017 Conference on Machine Translation (WMT17)","volume-title":"Second Conference on Machine Translation","author":"Bojar","year":"2017"},{"key":"2021122316152622700_bib6","doi-asserted-by":"publisher","first-page":"131","DOI":"10.18653\/v1\/W16-2301","article-title":"Findings of the 2016 Conference on Machine Translation","volume-title":"Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers","author":"Bojar","year":"2016"},{"key":"2021122316152622700_bib7","article-title":"Proceedings of the Third Workshop on Statistical Machine Translation","volume-title":"Proceedings of the Third Workshop on Statistical Machine Translation","author":"Callison-Burch","year":"2008"},{"key":"2021122316152622700_bib8","article-title":"A Comparative Quality Evaluation of PBSMT and NMT using Professional Translators","author":"Castilho","year":"2017","journal-title":"AAMT"},{"key":"2021122316152622700_bib9","first-page":"215","article-title":"What\u2019s the difference between professional human and machine translation? A blind multi-language study on domain-specific MT","volume-title":"Proceedings of the 22nd Annual Conference of the European Association for Machine Translation","author":"Fischer","year":"2020"},{"key":"2021122316152622700_bib10","unstructured":"Marina Fomicheva . 2017. The Role of Human Reference Translation in Machine Translation Evaluation. Ph.D. thesis, Universitat Pompeu Fabra."},{"key":"2021122316152622700_bib11","doi-asserted-by":"crossref","first-page":"192","DOI":"10.18653\/v1\/W18-6320","article-title":"Exploring gap filling as a cheaper alternative to reading comprehension questionnaires when evaluating machine translation for gisting","volume-title":"Proceedings of the Third Conference on Machine Translation: Research Papers","author":"Forcada","year":"2018"},{"key":"2021122316152622700_bib12","first-page":"34","article-title":"APE at scale and its implications on MT evaluation biases","volume-title":"Proceedings of the Fourth Conference on Machine Translation","author":"Freitag","year":"2019"},{"key":"2021122316152622700_bib13","doi-asserted-by":"crossref","first-page":"61","DOI":"10.18653\/v1\/2020.emnlp-main.5","article-title":"BLEU might be guilty but references are not innocent","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Freitag","year":"2020"},{"key":"2021122316152622700_bib14","first-page":"33","article-title":"Continuous measurement scales in human evaluation of machine translation","volume-title":"Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse","author":"Graham","year":"2013"},{"issue":"1","key":"2021122316152622700_bib15","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1017\/S1351324915000339","article-title":"Can machine translation systems be evaluated by the crowd alone?","volume":"23","author":"Graham","year":"2017","journal-title":"Natural Language Engineering"},{"key":"2021122316152622700_bib16","doi-asserted-by":"crossref","first-page":"72","DOI":"10.18653\/v1\/2020.emnlp-main.6","article-title":"Translationese in machine translation evaluation","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Graham","year":"2020"},{"key":"2021122316152622700_bib17","article-title":"Achieving Human Parity on Automatic Chinese to English News Translation","author":"Hassan","year":"2018","journal-title":"arXiv preprint arXiv:1803.05567"},{"issue":"3","key":"2021122316152622700_bib18","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1007\/s10590-018-9214-x","article-title":"Quantitative fine- grained human evaluation of machine translation systems: A case study on english to croatian","volume":"32","author":"Klubi\u010dka","year":"2018","journal-title":"Machine Translation"},{"key":"2021122316152622700_bib19","doi-asserted-by":"crossref","first-page":"102","DOI":"10.3115\/1654650.1654666","article-title":"Manual and automatic evaluation of machine translation between european languages","volume-title":"Proceedings on the Workshop on Statistical Machine Translation","author":"Koehn","year":"2006"},{"key":"2021122316152622700_bib20","first-page":"1318","article-title":"Translationese and its dialects","volume-title":"Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies - Volume 1","author":"Koppel","year":"2011"},{"key":"2021122316152622700_bib21","doi-asserted-by":"crossref","first-page":"653","DOI":"10.1613\/jair.1.11371","article-title":"A set of recommendations for assessing human\u2013machine parity in language translation","volume":"67","author":"L\u00e4ubli","year":"2020","journal-title":"Journal of Artificial Intelligence Research"},{"key":"2021122316152622700_bib22","doi-asserted-by":"crossref","first-page":"4791","DOI":"10.18653\/v1\/D18-1512","article-title":"Has machine translation achieved human parity? A case for document-level evaluation","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"L\u00e4ubli","year":"2018"},{"key":"2021122316152622700_bib23","doi-asserted-by":"crossref","first-page":"455","DOI":"10.5565\/rev\/tradumatica.77","article-title":"Multidimensional quality metrics (MQM): A framework for declaring and describing translation quality metrics","author":"Lommel","year":"2014","journal-title":"Tradum\u00e0tica"},{"key":"2021122316152622700_bib24","first-page":"688","article-title":"Results of the WMT20 metrics shared task","volume-title":"Proceedings of the Fifth Conference on Machine Translation","author":"Mathur","year":"2020"},{"key":"2021122316152622700_bib25","doi-asserted-by":"publisher","first-page":"571","DOI":"10.1162\/tacl_a_00040","article-title":"Comparing Bayesian models of annotation","volume":"6","author":"Paun","year":"2018","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2021122316152622700_bib26","doi-asserted-by":"publisher","first-page":"392","DOI":"10.18653\/v1\/W15-3049","article-title":"chrF: Character n-gram F-score for automatic MT evaluation","volume-title":"Proceedings of the Tenth Workshop on Statistical Machine Translation","author":"Popovi\u0107","year":"2015"},{"key":"2021122316152622700_bib27","doi-asserted-by":"crossref","first-page":"5059","DOI":"10.18653\/v1\/2020.coling-main.444","article-title":"Informative manual evaluation of machine translation output","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Popovi\u0107","year":"2020"},{"key":"2021122316152622700_bib28","doi-asserted-by":"crossref","first-page":"2685","DOI":"10.18653\/v1\/2020.emnlp-main.213","article-title":"COMET: A neural framework for MT evaluation","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Rei","year":"2020"},{"key":"2021122316152622700_bib29","first-page":"3652","article-title":"A reading comprehension corpus for machine translation evaluation","volume-title":"Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC\u201916)","author":"Scarton","year":"2016"},{"key":"2021122316152622700_bib30","doi-asserted-by":"crossref","first-page":"158","DOI":"10.18653\/v1\/2020.inlg-1.22","article-title":"A gold standard methodology for evaluating accuracy in data-to-text systems","volume-title":"Proceedings of the 13th International Conference on Natural Language Generation","author":"Thomson","year":"2020"},{"key":"2021122316152622700_bib31","first-page":"185","article-title":"Reassessing claims of human parity and super-human performance in machine translation at WMT 2019","volume-title":"Proceedings of the 22nd Annual Conference of the European Association for Machine Translation","author":"Toral","year":"2020"},{"key":"2021122316152622700_bib32","doi-asserted-by":"crossref","first-page":"113","DOI":"10.18653\/v1\/W18-6312","article-title":"Attaining the unattainable? Reassessing claims of human parity in neural machine translation","volume-title":"Proceedings of the Third Conference on Machine Translation: Research Papers","author":"Toral","year":"2018"},{"key":"2021122316152622700_bib33","doi-asserted-by":"crossref","first-page":"96","DOI":"10.3115\/1626355.1626368","article-title":"Human evaluation of machine translation through binary system comparisons","volume-title":"Proceedings of the Second Workshop on Statistical Machine Translation","author":"Vilar","year":"2007"},{"key":"2021122316152622700_bib34","article-title":"The ARPA MT evaluation methodologies: Evolution, lessons, and future approaches","volume-title":"Proceedings of the First Conference of the Association for Machine Translation in the Americas","author":"White","year":"1994"},{"key":"2021122316152622700_bib35","doi-asserted-by":"crossref","first-page":"73","DOI":"10.18653\/v1\/W19-5208","article-title":"The effect of translationese in machine translation test sets","volume-title":"Proceedings of the Fourth Conference on Machine Translation (Volume 1: Research Papers)","author":"Zhang","year":"2019"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00437\/1979261\/tacl_a_00437.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00437\/1979261\/tacl_a_00437.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,15]],"date-time":"2024-09-15T04:36:57Z","timestamp":1726375017000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00437\/108866\/Experts-Errors-and-Context-A-Large-Scale-Study-of"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021]]},"references-count":35,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00437","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2021]]},"published":{"date-parts":[[2021]]}}}