{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T23:21:54Z","timestamp":1778628114646,"version":"3.51.4"},"reference-count":47,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,4,17]],"date-time":"2024-04-17T00:00:00Z","timestamp":1713312000000},"content-version":"vor","delay-in-days":107,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,4,16]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Pretrained character-level and byte-level language models have been shown to be competitive with popular subword models across a range of Natural Language Processing tasks. However, there has been little research on their effectiveness for neural machine translation (NMT), particularly within the popular pretrain-then-finetune paradigm. This work performs an extensive comparison across multiple languages and experimental conditions of character- and subword-level pretrained models (ByT5 and mT5, respectively) on NMT. We show the effectiveness of character-level modeling in translation, particularly in cases where fine-tuning data is limited. In our analysis, we show how character models\u2019 gains in translation quality are reflected in better translations of orthographically similar words and rare words. While evaluating the importance of source texts in driving model predictions, we highlight word-level patterns within ByT5, suggesting an ability to modulate word-level and character-level information during generation. We conclude by assessing the efficiency tradeoff of byte models, suggesting their usage in non-time-critical scenarios to boost translation quality.<\/jats:p>","DOI":"10.1162\/tacl_a_00651","type":"journal-article","created":{"date-parts":[[2024,4,17]],"date-time":"2024-04-17T20:20:59Z","timestamp":1713385259000},"page":"392-410","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":7,"title":["Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine Translation"],"prefix":"10.1162","volume":"12","author":[{"given":"Lukas","family":"Edman","sequence":"first","affiliation":[{"name":"Center for Language and Cognition, University of Groningen, the Netherlands. j.l.edman@rug.nl"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gabriele","family":"Sarti","sequence":"additional","affiliation":[{"name":"Center for Language and Cognition, University of Groningen, the Netherlands. g.sarti@rug.nl"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Antonio","family":"Toral","sequence":"additional","affiliation":[{"name":"Center for Language and Cognition, University of Groningen, the Netherlands. a.toral.ruiz@rug.nl"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gertjan van","family":"Noord","sequence":"additional","affiliation":[{"name":"Center for Language and Cognition, University of Groningen, the Netherlands. g.j.m.van.noord@rug.nl"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arianna","family":"Bisazza","sequence":"additional","affiliation":[{"name":"Center for Language and Cognition, University of Groningen, the Netherlands. a.bisazza@rug.nl"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2024,4,16]]},"reference":[{"key":"2024041720204444600_bib1","doi-asserted-by":"publisher","first-page":"976","DOI":"10.18653\/v1\/2022.emnlp-main.64","article-title":"\u201cWill you find these shortcuts?\u201d A protocol for evaluating the faithfulness of input salience methods for text classification","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Bastings","year":"2022"},{"key":"2024041720204444600_bib2","article-title":"Synthetic and natural noise both break neural machine translation","author":"Belinkov","year":"2017","journal-title":"arXiv preprint arXiv: 1711.02173"},{"key":"2024041720204444600_bib3","first-page":"131","article-title":"On the effectiveness of quasi character-level models for machine translation","volume-title":"Proceedings of the 15th Biennial Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track)","author":"Carri\u00f3n-Ponz","year":"2022"},{"key":"2024041720204444600_bib4","doi-asserted-by":"publisher","first-page":"4295","DOI":"10.18653\/v1\/D18-1461","article-title":"Revisiting character-based neural machine translation with capacity and compression","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Cherry","year":"2018"},{"key":"2024041720204444600_bib5","doi-asserted-by":"publisher","first-page":"1693","DOI":"10.18653\/v1\/P16-1160","article-title":"A character-level decoder without explicit segmentation for neural machine translation","volume-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chung","year":"2016"},{"key":"2024041720204444600_bib6","doi-asserted-by":"publisher","first-page":"73","DOI":"10.1162\/tacl_a_00448","article-title":"Canine: Pre-training an efficient tokenization-free encoder for language representation","volume":"10","author":"Clark","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024041720204444600_bib7","doi-asserted-by":"publisher","first-page":"154","DOI":"10.18653\/v1\/W17-4123","article-title":"Byte-based neural machine translation","volume-title":"Proceedings of the First Workshop on Subword and Character Level Models in NLP","author":"Costa-juss\u00e0","year":"2017"},{"key":"2024041720204444600_bib8","doi-asserted-by":"publisher","first-page":"357","DOI":"10.18653\/v1\/P16-2058","article-title":"Character-based neural machine translation","volume-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Costa-juss\u00e0","year":"2016"},{"key":"2024041720204444600_bib9","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18653\/v1\/W19-5201","article-title":"Saliency-driven word alignment interpretation for neural machine translation","volume-title":"Proceedings of the Fourth Conference on Machine Translation (Volume 1: Research Papers)","author":"Ding","year":"2019"},{"key":"2024041720204444600_bib10","doi-asserted-by":"publisher","first-page":"2112","DOI":"10.18653\/v1\/2021.eacl-main.181","article-title":"Word alignment by fine-tuning embeddings on parallel corpora","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Dou","year":"2021"},{"key":"2024041720204444600_bib11","doi-asserted-by":"publisher","first-page":"1504","DOI":"10.18653\/v1\/N19-1154","article-title":"One size does not fit all: Comparing NMT representations of different granularities","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Durrani","year":"2019"},{"key":"2024041720204444600_bib12","first-page":"465","article-title":"Hindi-to-Urdu machine translation through transliteration","volume-title":"Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics","author":"Durrani","year":"2010"},{"key":"2024041720204444600_bib13","doi-asserted-by":"publisher","first-page":"981","DOI":"10.18653\/v1\/2022.findings-emnlp.69","article-title":"Subword-delimited downsampling for better character-level translation","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Edman","year":"2022"},{"key":"2024041720204444600_bib14","doi-asserted-by":"publisher","first-page":"6903","DOI":"10.18653\/v1\/2020.coling-main.609","article-title":"CharacterBERT: Reconciling ELMo and BERT for word-level open-vocabulary representations from characters","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"El Boukkouri","year":"2020"},{"key":"2024041720204444600_bib15","first-page":"46","article-title":"Results of WMT22 metrics shared task: Stop using BLEU \u2013 neural metrics are better and more robust","volume-title":"Proceedings of the Seventh Conference on Machine Translation (WMT)","author":"Freitag","year":"2022"},{"key":"2024041720204444600_bib16","doi-asserted-by":"publisher","first-page":"1754","DOI":"10.18653\/v1\/2021.emnlp-main.132","article-title":"Cross-attention is all you need: Adapting pretrained Transformers for machine translation","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Gheini","year":"2021"},{"key":"2024041720204444600_bib17","doi-asserted-by":"publisher","DOI":"10.4324\/9780203501887","volume-title":"Translation: An Advanced Resource Book","author":"Hatim","year":"2004"},{"key":"2024041720204444600_bib18","doi-asserted-by":"publisher","first-page":"953","DOI":"10.18653\/v1\/D19-1088","article-title":"Towards understanding neural machine translation with word importance","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"He","year":"2019"},{"key":"2024041720204444600_bib19","doi-asserted-by":"publisher","first-page":"238","DOI":"10.18653\/v1\/2022.blackboxnlp-1.19","article-title":"Investigating the characteristics of a transformer in a few-shot setup: Does freezing layers in RoBERTa help?","volume-title":"Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"Ingle","year":"2022"},{"key":"2024041720204444600_bib20","article-title":"Exploring the limits of language modeling","author":"Jozefowicz","year":"2016","journal-title":"arXiv preprint arXiv:1602.02410"},{"key":"2024041720204444600_bib21","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v30i1.10362","article-title":"Character-aware neural language models","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Kim","year":"2016"},{"key":"2024041720204444600_bib22","article-title":"Traducci\u00f3n autom\u00e1tica basada en caracteres y redes neuronales","author":"Larriba Flor","year":"2017"},{"key":"2024041720204444600_bib23","doi-asserted-by":"publisher","first-page":"365","DOI":"10.1162\/tacl_a_00067","article-title":"Fully character-level neural machine translation without explicit segmentation","volume":"5","author":"Lee","year":"2017","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024041720204444600_bib24","doi-asserted-by":"publisher","first-page":"543","DOI":"10.18653\/v1\/2021.acl-short.69","article-title":"When is char better than subword: A systematic study of segmentation algorithms for neural machine translation","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)","author":"Li","year":"2021"},{"key":"2024041720204444600_bib25","doi-asserted-by":"publisher","first-page":"2470","DOI":"10.18653\/v1\/2022.findings-acl.194","article-title":"Why don\u2019t people use character-level machine translation?","volume-title":"Findings of the Association for Computational Linguistics: ACL 2022","author":"Libovick\u00fd","year":"2022"},{"key":"2024041720204444600_bib26","doi-asserted-by":"publisher","first-page":"726","DOI":"10.1162\/tacl_a_00343","article-title":"Multilingual denoising pre-training for neural machine translation","volume":"8","author":"Liu","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"issue":"8","key":"2024041720204444600_bib27","doi-asserted-by":"publisher","DOI":"10.1145\/3546577","article-title":"Post-hoc interpretability for neural nlp: A survey","volume":"55","author":"Madsen","year":"2022","journal-title":"ACM Computing Surveys"},{"key":"2024041720204444600_bib28","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2207.04672","article-title":"No language left behind: Scaling human-centered machine translation","volume":"abs\/2207.04672","author":"Team","year":"2022","journal-title":"ArXiv"},{"key":"2024041720204444600_bib29","doi-asserted-by":"publisher","first-page":"612","DOI":"10.18653\/v1\/W17-4770","article-title":"chrF++: words helping character n-grams","volume-title":"Proceedings of the Second Conference on Machine Translation","author":"Popovi\u0107","year":"2017"},{"issue":"1","key":"2024041720204444600_bib30","first-page":"5485","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"The Journal of Machine Learning Research"},{"key":"2024041720204444600_bib31","first-page":"578","article-title":"COMET-22: Unbabel-IST 2022 submission for the metrics shared task","volume-title":"Proceedings of the Seventh Conference on Machine Translation (WMT)","author":"Rei","year":"2022"},{"key":"2024041720204444600_bib32","doi-asserted-by":"publisher","first-page":"2685","DOI":"10.18653\/v1\/2020.emnlp-main.213","article-title":"COMET: A neural framework for MT evaluation","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Rei","year":"2020"},{"key":"2024041720204444600_bib33","doi-asserted-by":"publisher","first-page":"12336","DOI":"10.18653\/v1\/2023.acl-long.689","article-title":"Accelerating transformer inference for translation via parallel decoding","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Santilli","year":"2023"},{"key":"2024041720204444600_bib34","doi-asserted-by":"publisher","first-page":"421","DOI":"10.18653\/v1\/2023.acl-demo.40","article-title":"Inseq: An interpretability toolkit for sequence generation models","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations)","author":"Sarti","year":"2023"},{"key":"2024041720204444600_bib35","doi-asserted-by":"publisher","first-page":"376","DOI":"10.18653\/v1\/E17-2060","article-title":"How grammatical is character-level neural machine translation? Assessing MT quality with contrastive translation pairs","volume-title":"Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers","author":"Sennrich","year":"2017"},{"key":"2024041720204444600_bib36","doi-asserted-by":"publisher","first-page":"1715","DOI":"10.18653\/v1\/P16-1162","article-title":"Neural machine translation of rare words with subword units","volume-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Sennrich","year":"2016"},{"key":"2024041720204444600_bib37","first-page":"4596","article-title":"Adafactor: Adaptive learning rates with sublinear memory cost","volume-title":"Proceedings of the 35th International Conference on Machine Learning","author":"Shazeer","year":"2018"},{"key":"2024041720204444600_bib38","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1312.6034","article-title":"Deep inside convolutional networks: Visualising image classification models and saliency maps","volume-title":"2nd International Conference on Learning Representations (ICLR 2014): Workshop Track Proceedings","author":"Simonyan","year":"2014"},{"key":"2024041720204444600_bib39","unstructured":"Mittul\n              Singh\n            \n          . 2017. Handling Long-term Dependencies and Rare Words in Low-resource Language Modelling. Ph.D. thesis, Saarland University."},{"key":"2024041720204444600_bib40","article-title":"Blockwise parallel decoding for deep autoregressive models","volume-title":"Advances in Neural Information Processing Systems","author":"Stern","year":"2018"},{"key":"2024041720204444600_bib41","article-title":"Charformer: Fast character transformers via gradient-based subword tokenization","volume-title":"Proceedings of the Tenth International Conference on Learning Representations (ICLR)","author":"Yi","year":"2022"},{"key":"2024041720204444600_bib42","first-page":"676","article-title":"Analyzing the use of character-level translation with sparse and noisy datasets","volume-title":"Proceedings of the International Conference Recent Advances in Natural Language Processing RANLP 2013","author":"Tiedemann","year":"2013"},{"key":"2024041720204444600_bib43","doi-asserted-by":"publisher","first-page":"1126","DOI":"10.18653\/v1\/2021.acl-long.91","article-title":"Analyzing the source and target contributions to predictions in neural machine translation","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Voita","year":"2021"},{"key":"2024041720204444600_bib44","doi-asserted-by":"publisher","first-page":"8478","DOI":"10.18653\/v1\/2021.emnlp-main.667","article-title":"Language modeling, lexical translation, reordering: The training process of NMT through the lens of classical SMT","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Voita","year":"2021"},{"key":"2024041720204444600_bib45","doi-asserted-by":"publisher","first-page":"291","DOI":"10.1162\/tacl_a_00461","article-title":"ByT5: Towards a token-free future with pre-trained byte-to-byte models","volume":"10","author":"Xue","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024041720204444600_bib46","doi-asserted-by":"publisher","first-page":"483","DOI":"10.18653\/v1\/2021.naacl-main.41","article-title":"mT5: A massively multilingual pre-trained text-to-text transformer","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Xue","year":"2021"},{"key":"2024041720204444600_bib47","article-title":"Megabyte: Predicting million-byte sequences with multiscale transformers","author":"Lili","year":"2023","journal-title":"arXiv preprint arXiv:2305.07185"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00651\/2364129\/tacl_a_00651.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00651\/2364129\/tacl_a_00651.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,4,17]],"date-time":"2024-04-17T20:21:06Z","timestamp":1713385266000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00651\/120650\/Are-Character-level-Translations-Worth-the-Wait"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":47,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00651","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}