{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T06:26:32Z","timestamp":1783405592985,"version":"3.54.6"},"reference-count":62,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,7,16]],"date-time":"2024-07-16T00:00:00Z","timestamp":1721088000000},"content-version":"vor","delay-in-days":197,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,7,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Metrics are the foundation for automatic evaluation in grammatical error correction (GEC), with their evaluation of the metrics (meta-evaluation) relying on their correlation with human judgments. However, conventional meta-evaluations in English GEC encounter several challenges, including biases caused by inconsistencies in evaluation granularity and an outdated setup using classical systems. These problems can lead to misinterpretation of metrics and potentially hinder the applicability of GEC techniques. To address these issues, this paper proposes SEEDA, a new dataset for GEC meta-evaluation. SEEDA consists of corrections with human ratings along two different granularities: edit-based and sentence-based, covering 12 state-of-the-art systems including large language models, and two human corrections with different focuses. The results of improved correlations by aligning the granularity in the sentence-level meta-evaluation suggest that edit-based metrics may have been underestimated in existing studies. Furthermore, correlations of most metrics decrease when changing from classical to neural systems, indicating that traditional metrics are relatively poor at evaluating fluently corrected sentences with many edits.<\/jats:p>","DOI":"10.1162\/tacl_a_00676","type":"journal-article","created":{"date-parts":[[2024,7,16]],"date-time":"2024-07-16T18:36:17Z","timestamp":1721154977000},"page":"837-855","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":8,"title":["Revisiting Meta-evaluation for Grammatical Error Correction"],"prefix":"10.1162","volume":"12","author":[{"given":"Masamune","family":"Kobayashi","sequence":"first","affiliation":[{"name":"Tokyo Metropolitan University, Japan. kobayashi-masamune@ed.tmu.ac.jp"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Masato","family":"Mita","sequence":"additional","affiliation":[{"name":"CyberAgent Inc., Japan"},{"name":"Tokyo Metropolitan University, Japan. mita_masato@cyberagent.co.jp"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mamoru","family":"Komachi","sequence":"additional","affiliation":[{"name":"Hitotsubashi University, Japan. mamoru.komachi@r.hit-u.ac.jp"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,7,15]]},"reference":[{"key":"2024071618354371800_bib1","first-page":"343","article-title":"Reference-based metrics can be replaced with reference-less metrics in evaluating grammatical error correction systems","volume-title":"Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers)","author":"Asano","year":"2017"},{"key":"2024071618354371800_bib2","doi-asserted-by":"publisher","first-page":"4260","DOI":"10.18653\/v1\/D19-1435","article-title":"Parallel iterative edit models for local sequence transduction","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Awasthi","year":"2019"},{"key":"2024071618354371800_bib3","doi-asserted-by":"publisher","first-page":"1","DOI":"10.3115\/v1\/W14-3302","article-title":"Findings of the 2013 Workshop on Statistical Machine Translation","volume-title":"Proceedings of the Eighth Workshop on Statistical Machine Translation","author":"Bojar","year":"2013"},{"key":"2024071618354371800_bib4","doi-asserted-by":"publisher","first-page":"52","DOI":"10.18653\/v1\/W19-4406","article-title":"The BEA-2019 shared task on grammatical error correction","volume-title":"Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications","author":"Bryant","year":"2019"},{"key":"2024071618354371800_bib5","doi-asserted-by":"publisher","first-page":"793","DOI":"10.18653\/v1\/P17-1074","article-title":"Automatic annotation and evaluation of error types for grammatical error correction","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Bryant","year":"2017"},{"key":"2024071618354371800_bib6","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1162\/coli_a_00478","article-title":"Grammatical error correction: A survey of the state of the art","author":"Bryant","year":"2023","journal-title":"Computational Linguistics"},{"key":"2024071618354371800_bib7","doi-asserted-by":"publisher","first-page":"15607","DOI":"10.18653\/v1\/2023.acl-long.870","article-title":"Can large language models be an alternative to human evaluations?","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chiang","year":"2023"},{"key":"2024071618354371800_bib8","doi-asserted-by":"publisher","first-page":"2528","DOI":"10.18653\/v1\/D18-1274","article-title":"Neural quality estimation of grammatical error correction","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Chollampatt","year":"2018"},{"key":"2024071618354371800_bib9","first-page":"2730","article-title":"A reassessment of reference-based grammatical error correction metrics","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics","author":"Chollampatt","year":"2018"},{"key":"2024071618354371800_bib10","doi-asserted-by":"publisher","first-page":"1372","DOI":"10.18653\/v1\/P18-1127","article-title":"Automatic metric validation for grammatical error correction","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Choshen","year":"2018"},{"key":"2024071618354371800_bib11","doi-asserted-by":"publisher","first-page":"124","DOI":"10.18653\/v1\/N18-2020","article-title":"Reference-less measure of faithfulness for grammatical error correction","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers)","author":"Choshen","year":"2018"},{"issue":"1","key":"2024071618354371800_bib12","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1177\/001316446002000104","article-title":"A coefficient of agreement for nominal scales","volume":"20","author":"Cohen","year":"1960","journal-title":"Educational and Psychological Measurement"},{"key":"2024071618354371800_bib13","article-title":"Analyzing the performance of GPT-3.5 and GPT-4 in grammatical error correction","author":"Coyne","year":"2023"},{"key":"2024071618354371800_bib14","first-page":"568","article-title":"Better evaluation for grammatical error correction","volume-title":"Proceedings of the 2012 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Dahlmeier","year":"2012"},{"key":"2024071618354371800_bib15","first-page":"54","article-title":"HOO 2012: A report on the preposition and determiner error correction shared task","volume-title":"Proceedings of the Seventh Workshop on Building Educational Applications Using NLP","author":"Dale","year":"2012"},{"key":"2024071618354371800_bib16","first-page":"242","article-title":"Helping our own: The HOO 2011 pilot shared task","volume-title":"Proceedings of the 13th European Workshop on Natural Language Generation","author":"Dale","year":"2011"},{"key":"2024071618354371800_bib17","doi-asserted-by":"publisher","first-page":"1132","DOI":"10.1162\/tacl_a_00417","article-title":"A statistical analysis of summarization evaluation metrics using resampling methods","volume":"9","author":"Deutsch","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024071618354371800_bib18","doi-asserted-by":"publisher","first-page":"4171","DOI":"10.18653\/v1\/N19-1423","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2024071618354371800_bib19","doi-asserted-by":"publisher","first-page":"3614","DOI":"10.18653\/v1\/2023.findings-acl.223","article-title":"TransGEC: Improving grammatical error correction with translationese","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Fang","year":"2023"},{"key":"2024071618354371800_bib20","article-title":"Appraise: An open-source toolkit for manual phrase-based evaluation of translations","volume-title":"Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC\u201910)","author":"Federmann","year":"2010"},{"key":"2024071618354371800_bib21","doi-asserted-by":"publisher","first-page":"578","DOI":"10.3115\/v1\/N15-1060","article-title":"Towards a standard evaluation method for grammatical error detection and correction","volume-title":"Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Felice","year":"2015"},{"key":"2024071618354371800_bib22","first-page":"46","article-title":"Results of WMT22 metrics shared task: Stop using BLEU \u2013 neural metrics are better and more robust","volume-title":"Proceedings of the Seventh Conference on Machine Translation (WMT)","author":"Freitag","year":"2022"},{"key":"2024071618354371800_bib23","first-page":"733","article-title":"Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain","volume-title":"Proceedings of the Sixth Conference on Machine Translation","author":"Freitag","year":"2021"},{"key":"2024071618354371800_bib24","doi-asserted-by":"publisher","first-page":"6891","DOI":"10.18653\/v1\/2022.emnlp-main.463","article-title":"Revisiting grammatical error correction evaluation and beyond","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Gong","year":"2022"},{"key":"2024071618354371800_bib25","doi-asserted-by":"publisher","first-page":"2085","DOI":"10.18653\/v1\/2020.coling-main.188","article-title":"Taking the correction difficulty into account in grammatical error correction evaluation","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Gotou","year":"2020"},{"key":"2024071618354371800_bib26","doi-asserted-by":"publisher","first-page":"461","DOI":"10.18653\/v1\/D15-1052","article-title":"Human evaluation of grammatical error correction systems","volume-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing","author":"Grundkiewicz","year":"2015"},{"key":"2024071618354371800_bib27","doi-asserted-by":"publisher","first-page":"252","DOI":"10.18653\/v1\/W19-4427","article-title":"Neural grammatical error correction systems with unsupervised pre-training on synthetic data","volume-title":"Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications","author":"Grundkiewicz","year":"2019"},{"key":"2024071618354371800_bib28","doi-asserted-by":"publisher","first-page":"3009","DOI":"10.18653\/v1\/2021.emnlp-main.239","article-title":"Is this the end of the gold standard? A straightforward reference-less grammatical error correction metric","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Md","year":"2021"},{"key":"2024071618354371800_bib29","doi-asserted-by":"publisher","first-page":"25","DOI":"10.3115\/v1\/W14-1703","article-title":"The AMU system in the CoNLL-2014 shared task: Grammatical error correction by data-intensive and feature-rich statistical machine translation","volume-title":"Proceedings of the Eighteenth Conference on Computational Natural Language Learning: Shared Task","author":"Junczys-Dowmunt","year":"2014"},{"key":"2024071618354371800_bib30","doi-asserted-by":"publisher","first-page":"4248","DOI":"10.18653\/v1\/2020.acl-main.391","article-title":"Encoder-decoder models can benefit from pre-trained masked language models in grammatical error correction","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Kaneko","year":"2020"},{"key":"2024071618354371800_bib31","doi-asserted-by":"publisher","first-page":"1236","DOI":"10.18653\/v1\/D19-1119","article-title":"An empirical study of incorporating pseudo data into grammatical error correction","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Kiyono","year":"2019"},{"key":"2024071618354371800_bib32","first-page":"193","article-title":"Large language models are state-of-the-art evaluators of translation quality","volume-title":"Proceedings of the 24th Annual Conference of the European Association for Machine Translation","author":"Kocmi","year":"2023"},{"key":"2024071618354371800_bib33","first-page":"707","article-title":"Binary codes capable of correcting deletions, insertions and reversals","volume":"10","author":"Levenshtein","year":"1966","journal-title":"Soviet Physics Doklady"},{"key":"2024071618354371800_bib34","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.703","article-title":"BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension","volume":"abs\/ 1910.13461","author":"Lewis","year":"2019","journal-title":"CoRR"},{"key":"2024071618354371800_bib35","doi-asserted-by":"publisher","first-page":"6878","DOI":"10.18653\/v1\/2023.acl-long.380","article-title":"TemplateGEC: Improving grammatical error correction with detection template","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Li","year":"2023"},{"key":"2024071618354371800_bib36","article-title":"Benchmarking generation and evaluation capabilities of large language models for instruction controllable summarization","author":"Liu","year":"2023"},{"key":"2024071618354371800_bib37","doi-asserted-by":"publisher","first-page":"5441","DOI":"10.18653\/v1\/2021.naacl-main.429","article-title":"Neural quality estimation with multiple hypotheses for grammatical error correction","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Liu","year":"2021"},{"key":"2024071618354371800_bib38","doi-asserted-by":"publisher","DOI":"10.5565\/rev\/tradumatica.77","article-title":"Multidimensional Quality Metrics (MQM): A framework for declaring and describing translation quality metrics","author":"Lommel","year":"2014"},{"key":"2024071618354371800_bib39","first-page":"3578","article-title":"IMPARA: Impact-based metric for GEC using parallel data","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Maeda","year":"2022"},{"key":"2024071618354371800_bib40","doi-asserted-by":"publisher","first-page":"4984","DOI":"10.18653\/v1\/2020.acl-main.448","article-title":"Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Mathur","year":"2020"},{"key":"2024071618354371800_bib41","doi-asserted-by":"publisher","first-page":"9194","DOI":"10.18653\/v1\/2023.acl-long.511","article-title":"Benchmarking large language model capabilities for conditional generation","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Maynez","year":"2023"},{"key":"2024071618354371800_bib42","first-page":"894","article-title":"Evaluating performance of grammatical error detection to maximize learning effect","volume-title":"Coling 2010: Posters","author":"Nagata","year":"2010"},{"key":"2024071618354371800_bib43","unstructured":"Hiroki\n              Nakayama\n            , TakahiroKubo, JunyaKamura, YasufumiTaniguchi, and XuLiang. 2018. doccano: Text annotation tool for human. Software available from https:\/\/github.com\/doccano\/doccano."},{"key":"2024071618354371800_bib44","doi-asserted-by":"publisher","first-page":"551","DOI":"10.1162\/tacl_a_00282","article-title":"Enabling robust grammatical error correction in new domains: Data sets, metrics, and analyses","volume":"7","author":"Napoles","year":"2019","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024071618354371800_bib45","doi-asserted-by":"publisher","first-page":"588","DOI":"10.3115\/v1\/P15-2097","article-title":"Ground truth for grammatical error correction metrics","volume-title":"Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)","author":"Napoles","year":"2015"},{"key":"2024071618354371800_bib46","article-title":"GLEU without tuning","author":"Napoles","year":"2016","journal-title":"CoRR"},{"key":"2024071618354371800_bib47","doi-asserted-by":"publisher","first-page":"2109","DOI":"10.18653\/v1\/D16-1228","article-title":"There\u2019s no comparison: Reference-less evaluation metrics in grammatical error correction","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing","author":"Napoles","year":"2016"},{"key":"2024071618354371800_bib48","doi-asserted-by":"publisher","first-page":"229","DOI":"10.18653\/v1\/E17-2037","article-title":"JFLEG: A fluency corpus and benchmark for grammatical error correction","volume-title":"Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers","author":"Napoles","year":"2017"},{"key":"2024071618354371800_bib49","doi-asserted-by":"publisher","first-page":"1","DOI":"10.3115\/v1\/W14-1701","article-title":"The CoNLL-2014 shared task on grammatical error correction","volume-title":"Proceedings of the Eighteenth Conference on Computational Natural Language Learning: Shared Task","author":"Ng","year":"2014"},{"key":"2024071618354371800_bib50","first-page":"1","article-title":"The CoNLL-2013 shared task on grammatical error correction","volume-title":"Proceedings of the Seventeenth Conference on Computational Natural Language Learning: Shared Task","author":"Ng","year":"2013"},{"key":"2024071618354371800_bib51","doi-asserted-by":"publisher","first-page":"163","DOI":"10.18653\/v1\/2020.bea-1.16","article-title":"GECToR \u2013 grammatical error correction: Tag, not rewrite","volume-title":"Proceedings of the Fifteenth Workshop on Innovative Use of NLP for Building Educational Applications","author":"Omelianchuk","year":"2020"},{"key":"2024071618354371800_bib52","doi-asserted-by":"publisher","first-page":"311","DOI":"10.3115\/1073083.1073135","article-title":"BLEU: A method for automatic evaluation of machine translation","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni","year":"2002"},{"key":"2024071618354371800_bib53","article-title":"Language models are unsupervised multitask learners","author":"Radford","year":"2019"},{"key":"2024071618354371800_bib54","doi-asserted-by":"publisher","first-page":"702","DOI":"10.18653\/v1\/2021.acl-short.89","article-title":"A simple recipe for multilingual grammatical error correction","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)","author":"Rothe","year":"2021"},{"key":"2024071618354371800_bib55","doi-asserted-by":"publisher","first-page":"34","DOI":"10.3115\/v1\/W14-1704","article-title":"The Illinois-Columbia system in the CoNLL-2014 shared task","volume-title":"Proceedings of the Eighteenth Conference on Computational Natural Language Learning: Shared Task","author":"Rozovskaya","year":"2014"},{"key":"2024071618354371800_bib56","doi-asserted-by":"publisher","first-page":"169","DOI":"10.1162\/tacl_a_00091","article-title":"Reassessing the goals of grammatical error correction: Fluency instead of grammaticality","volume":"4","author":"Sakaguchi","year":"2016","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024071618354371800_bib57","doi-asserted-by":"publisher","first-page":"1","DOI":"10.3115\/v1\/W14-3301","article-title":"Efficient elicitation of annotations for human evaluation of machine translation","volume-title":"Proceedings of the Ninth Workshop on Statistical Machine Translation","author":"Sakaguchi","year":"2014"},{"key":"2024071618354371800_bib58","doi-asserted-by":"publisher","first-page":"3842","DOI":"10.18653\/v1\/2022.acl-long.266","article-title":"Ensembling and knowledge distilling of large sequence taggers for grammatical error correction","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Tarnavskyi","year":"2022"},{"key":"2024071618354371800_bib59","first-page":"5998","article-title":"Attention is all you need","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani","year":"2017"},{"key":"2024071618354371800_bib60","doi-asserted-by":"publisher","first-page":"7752","DOI":"10.18653\/v1\/2021.emnlp-main.611","article-title":"LM-critic: Language models for unsupervised grammatical error correction","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Yasunaga","year":"2021"},{"key":"2024071618354371800_bib61","doi-asserted-by":"publisher","first-page":"6516","DOI":"10.18653\/v1\/2020.coling-main.573","article-title":"SOME: Reference-less sub-metrics optimized for manual evaluations of grammatical error correction","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Yoshimura","year":"2020"},{"key":"2024071618354371800_bib62","article-title":"BERTScore: Evaluating text generation with BERT","author":"Zhang","year":"2019","journal-title":"CoRR"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00676\/2461947\/tacl_a_00676.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00676\/2461947\/tacl_a_00676.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,16]],"date-time":"2024-07-16T18:36:28Z","timestamp":1721154988000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00676\/123651\/Revisiting-Meta-evaluation-for-Grammatical-Error"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":62,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00676","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}