{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,10,7]],"date-time":"2026-10-07T15:53:21Z","timestamp":1791388401674,"version":"4.3.3"},"reference-count":103,"publisher":"Springer Science and Business Media LLC","license":[{"start":{"date-parts":[[2026,10,7]],"date-time":"2026-10-07T00:00:00Z","timestamp":1791331200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,10,7]],"date-time":"2026-10-07T00:00:00Z","timestamp":1791331200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Nature"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Recent advances in artificial intelligence (AI) have largely been driven by large language models, deep neural networks that operate over discrete units called tokens. To represent text, most large language models use words or word fragments as the tokens, known as subword tokenization\n                    <jats:sup>1<\/jats:sup>\n                    . Subword tokenization obscures fine-grained information, which is problematic, especially for scientific data\u2014such as computer code or biological sequences\u2014where meaning depends on the individual characters or bytes\n                    <jats:sup>2<\/jats:sup>\n                    . Models that instead operate directly on the byte encoding of text avoid these limitations, but until now they have lagged behind subword-based models in performance. Here we introduce a general method for creating byte-level large language models through byteification that approach the capabilities of subword-based systems. We use a two-stage conversion procedure to retrofit existing subword-based models into byte-level models with minimal extra training. The resulting models outperform earlier byte-level approaches and excel on character-level reasoning tasks, achieving practical inference speeds by efficiently processing byte-level information and adaptability by reusing the existing ecosystem around the source large language model. Our results remove a long-standing performance barrier to end-to-end byte-level language modelling, demonstrating that models operating on raw text encodings can scale competitively while offering advantages in domains requiring fine-grained textual understanding.\n                  <\/jats:p>","DOI":"10.1038\/s41586-026-11111-4","type":"journal-article","created":{"date-parts":[[2026,10,7]],"date-time":"2026-10-07T15:03:25Z","timestamp":1791385405000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Retrofitting language models to operate over bytes"],"prefix":"10.1038","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6520-4245","authenticated-orcid":false,"given":"Benjamin","family":"Minixhofer","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tyler","family":"Murray","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tomasz","family":"Limisiewicz","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anna","family":"Korhonen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Luke","family":"Zettlemoyer","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Noah A.","family":"Smith","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Edoardo M.","family":"Ponti","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Luca","family":"Soldaini","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6603-3428","authenticated-orcid":false,"given":"Valentin","family":"Hofmann","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,10,7]]},"reference":[{"key":"11111_CR1","doi-asserted-by":"crossref","unstructured":"Sennrich, R., Haddow, B. & Birch, A. Neural machine translation of rare words with subword units. In Proc. 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Erk, K. & Smith, N. A.) 1715\u20131725 (Association for Computational Linguistics, 2016).","DOI":"10.18653\/v1\/P16-1162"},{"key":"11111_CR2","unstructured":"Hwang, S., Wang, B. & Gu, A. Dynamic chunking for end-to-end hierarchical sequence modeling. In Proc. Fourteenth International Conference on Learning Representations (eds Vondrick, C. et al.) 149273\u2013149313 (International Conference on Learning Representations, 2026)."},{"key":"11111_CR3","unstructured":"Brown, T. et al. Language models are few-shot learners. In Proc. Advances in Neural Information Processing Systems, Vol. 33 (eds Larochelle, H. et al.) 1877\u20131901 (Curran Associates, 2020)."},{"key":"11111_CR4","doi-asserted-by":"publisher","first-page":"633","DOI":"10.1038\/s41586-025-09422-z","volume":"645","author":"D Guo","year":"2025","unstructured":"Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 633\u2013638 (2025).","journal-title":"Nature"},{"key":"11111_CR5","doi-asserted-by":"crossref","unstructured":"Hofmann, V., Pierrehumbert, J. & Sch\u00fctze, H. Superbizarre is not superb: derivational morphology improves BERT's interpretation of complex words. In Proc. 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (eds Zong, C. et al.) 3594\u20133608 (Association for Computational Linguistics, 2021).","DOI":"10.18653\/v1\/2021.acl-long.279"},{"key":"11111_CR6","doi-asserted-by":"crossref","unstructured":"Ahia, O. et al. Do all languages cost the same? Tokenization in the era of commercial language models. In Proc. 2023 Conference on Empirical Methods in Natural Language Processing (eds Bouamor, H. et al.) 9904\u20139923 (Association for Computational Linguistics, 2023).","DOI":"10.18653\/v1\/2023.emnlp-main.614"},{"key":"11111_CR7","doi-asserted-by":"crossref","unstructured":"Land, S. & Bartolo, M. Fishing for magikarp: automatically detecting under-trained tokens in large language models. In Proc. 2024 Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y. et al.) 11631\u201311646 (Association for Computational Linguistics, 2024).","DOI":"10.18653\/v1\/2024.emnlp-main.649"},{"key":"11111_CR8","doi-asserted-by":"crossref","unstructured":"Peng, Q., Chai, Y. & S\u00f8gaard, A. Understanding subword compositionality of large language models. In Proc. 2025 Conference on Empirical Methods in Natural Language Processing (eds Christodoulopoulos, C. et al.) 22524\u201322535 (Association for Computational Linguistics, 2025).","DOI":"10.18653\/v1\/2025.emnlp-main.1146"},{"key":"11111_CR9","doi-asserted-by":"crossref","unstructured":"Zheng, B. S. et al. Broken tokens? Your language model can secretly handle non-canonical tokenizations. In Advances in Neural Information Processing Systems, Vol. 38 (eds Belgrave, D. et al.) 30322\u201330349 (Curran Associates, 2025).","DOI":"10.52202\/085713-1017"},{"key":"11111_CR10","doi-asserted-by":"crossref","unstructured":"Kudo, T. Subword regularization: improving neural network translation models with multiple subword candidates. In Proc. 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Gurevych, I. & Miyao, Y.) 66\u201375 (Association for Computational Linguistics, 2018).","DOI":"10.18653\/v1\/P18-1007"},{"key":"11111_CR11","doi-asserted-by":"crossref","unstructured":"Edman, L., Schmid, H. & Fraser, A. CUTE: measuring LLMs' understanding of their tokens. In Proc. 2024 Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y. et al.) 3017\u20133026 (Association for Computational Linguistics, 2024).","DOI":"10.18653\/v1\/2024.emnlp-main.177"},{"key":"11111_CR12","doi-asserted-by":"crossref","unstructured":"Cosma, A., Ruseti, S., Radoi, E. & Dascalu, M. The strawberry problem: emergence of character-level understanding in tokenized language models. In Proc. 2025 Conference on Empirical Methods in Natural Language Processing (eds Christodoulopoulos, C. et al.) 28252\u201328263 (Association for Computational Linguistics, 2025).","DOI":"10.18653\/v1\/2025.emnlp-main.1434"},{"key":"11111_CR13","doi-asserted-by":"crossref","unstructured":"Uzan, O. & Pinter, Y. CharBench: evaluating the role of tokenization in character-level tasks. In Proc. AAAI Conference on Artificial Intelligence, Vol. 40 (eds Koenig, S. et al.) 33296\u201333304 (AAAI Press, 2026).","DOI":"10.1609\/aaai.v40i39.40615"},{"key":"11111_CR14","unstructured":"Chirkova, N. & Troshin, S. CodeBPE: investigating subtokenization options for large language model pretraining on source code. In Proc. Eleventh International Conference on Learning Representations (eds Liu, Y. et al.) (International Conference on Learning Representations, 2023)."},{"key":"11111_CR15","doi-asserted-by":"crossref","unstructured":"Zilio, L., Qian, S., Kanojia, D. & Orasan, C. Using character-level models for efficient abbreviation and long-form detection. In Proc. 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (eds Calzolari, N. et al.) 3028\u20133037 (ELRA and ICCL, 2024).","DOI":"10.63317\/2tuwvz3mosgi"},{"key":"11111_CR16","unstructured":"Dagan, G., Synnaeve, G. & Roziere, B. Getting the most out of your tokenizer for pre-training and domain adaptation. In Proc. 41st International Conference on Machine Learning, Vol. 235 (eds Salakhutdinov, R. et al.) 9784\u20139805 (PMLR, 2024)."},{"key":"11111_CR17","doi-asserted-by":"publisher","first-page":"btaf456","DOI":"10.1093\/bioinformatics\/btaf456","volume":"41","author":"LM Lindsey","year":"2025","unstructured":"Lindsey, L. M. et al. The impact of tokenizer selection in genomic language models. Bioinformatics 41, btaf456 (2025).","journal-title":"Bioinformatics"},{"key":"11111_CR18","unstructured":"Phan, B. et al. Exact byte-level probabilities from tokenized language models for fim-tasks and model ensembles. In Proc. Thirteenth International Conference on Learning Representations (eds Yue, Y. et al.) 38145\u201338166 (International Conference on Learning Representations, 2025)."},{"key":"11111_CR19","unstructured":"Hayase, J., Liu, A., Smith, N. A. & Oh, S. Sampling from your language model one byte at a time. In Proc. Forty-third International Conference on Machine Learning (2026)."},{"key":"11111_CR20","unstructured":"Vieira, T. et al. From language models over tokens to language models over characters. In Proc. Forty-second International Conference on Machine Learning, Vol. 267 (eds Singh, A. et al.) 61391\u201361412 (PMLR, 2025)."},{"key":"11111_CR21","doi-asserted-by":"crossref","unstructured":"Liang, D. et al. XLM-V: overcoming the vocabulary bottleneck in multilingual masked language models. In Proc. 2023 Conference on Empirical Methods in Natural Language Processing (eds Bouamor, H. et al.) 13142\u201313152 (Association for Computational Linguistics, 2023).","DOI":"10.18653\/v1\/2023.emnlp-main.813"},{"key":"11111_CR22","doi-asserted-by":"crossref","unstructured":"Pagnoni, A. et al. Byte latent transformer: patches scale better than tokens. In Proc. 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Che, W. et al.) 9238\u20139258 (Association for Computational Linguistics, 2025).","DOI":"10.18653\/v1\/2025.acl-long.453"},{"key":"11111_CR23","doi-asserted-by":"crossref","unstructured":"Yergeau, F. UTF-8, A Transformation Format of ISO 10646. Technical Report RFC 3629 (2003).","DOI":"10.17487\/rfc3629"},{"key":"11111_CR24","doi-asserted-by":"crossref","unstructured":"Nawrot, P., Chorowski, J., Lancucki, A. & Ponti, E. M. Efficient transformers with dynamic token pooling. In Proc. 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Rogers, A. et al.) 6403\u20136417 (Association for Computational Linguistics, 2023).","DOI":"10.18653\/v1\/2023.acl-long.353"},{"key":"11111_CR25","doi-asserted-by":"crossref","unstructured":"Slagle, K. SpaceByte: towards deleting tokenization from large language modeling. In Advances in Neural Information Processing Systems, Vol. 37 (eds Globerson, A. et al.) 124925\u2013124950 (Curran Associates, 2024).","DOI":"10.52202\/079017-3967"},{"key":"11111_CR26","unstructured":"Wang, J., Gangavarapu, T., Yan, J. N. & Rush, A. M. MambaByte: token-free selective state space model. In Proc. First Conference on Language Modeling (2024)."},{"key":"11111_CR27","unstructured":"Zheng, L. et al. EvaByte: Efficient Byte-level Language Models at Scale. HKU NLP Group https:\/\/hkunlp.github.io\/blog\/2025\/evabyte (2025)."},{"key":"11111_CR28","doi-asserted-by":"publisher","unstructured":"Olmo Team. Olmo 3. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2512.13961 (2025).","DOI":"10.48550\/arXiv.2512.13961"},{"key":"11111_CR29","unstructured":"Walsh, E. P. et al. 2 OLMo 2 furious (COLM's version). In Proc. Second Conference on Language Modeling (2025)."},{"key":"11111_CR30","doi-asserted-by":"publisher","unstructured":"Yang, A. et al. Qwen3 technical report. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2505.09388 (2025).","DOI":"10.48550\/arXiv.2505.09388"},{"key":"11111_CR31","doi-asserted-by":"publisher","unstructured":"Grattafiori, A. et al. The Llama 3 herd of models. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2407.21783 (2024).","DOI":"10.48550\/arXiv.2407.21783"},{"key":"11111_CR32","unstructured":"Vaswani, A. et al. Attention is all you need. In Advances in Neural Information Processing Systems, Vol. 30 (eds Guyon, I. et al.) (Curran Associates, 2017)."},{"key":"11111_CR33","unstructured":"Beck, M. et al. xLSTM 7B: a recurrent LLM for fast and efficient inference. In Proc. Forty-second International Conference on Machine Learning, Vol. 267 (eds Singh, A. et al.) 3335\u20133357 (PMLR, 2025)."},{"key":"11111_CR34","doi-asserted-by":"crossref","unstructured":"Minixhofer, B., Pfeiffer, J. & Vuli\u0107, I. CompoundPiece: evaluating and improving decompounding performance of language models. In Proc. 2023 Conference on Empirical Methods in Natural Language Processing (eds Bouamor, H. et al.) 343\u2013359 (Association for Computational Linguistics, 2023).","DOI":"10.18653\/v1\/2023.emnlp-main.24"},{"key":"11111_CR35","doi-asserted-by":"publisher","unstructured":"Fleshman, W. & Van Durme, B. Toucan: token-aware character level language modeling. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2311.08620 (2023).","DOI":"10.48550\/arXiv.2311.08620"},{"key":"11111_CR36","unstructured":"Holtzman, A., Buys, J., Du, L., Forbes, M. & Choi, Y. The curious case of neural text degeneration. In Proc. International Conference on Learning Representations (eds Rush, A. et al.) (International Conference on Learning Representations, 2020)."},{"key":"11111_CR37","unstructured":"Neitemeier, P., Deiseroth, B., Eichenberg, C. & Balles, L. Hierarchical autoregressive transformers: combining byte- and word-level processing for robust, adaptable language models. In Proc. Thirteenth International Conference on Learning Representations (eds Yue, Y. et al.) 51088\u201351111 (International Conference on Learning Representations, 2025)."},{"key":"11111_CR38","unstructured":"Liu, A. et al. SuperBPE: space travel for language models. In Proc. Second Conference on Language Modeling (2025)."},{"key":"11111_CR39","doi-asserted-by":"crossref","unstructured":"Dobler, K. & de Melo, G. FOCUS: effective embedding initialization for monolingual specialization of multilingual models. In Proc. 2023 Conference on Empirical Methods in Natural Language Processing (eds Bouamor, H. et al.) 13440\u201313454 (Association for Computational Linguistics, 2023).","DOI":"10.18653\/v1\/2023.emnlp-main.829"},{"key":"11111_CR40","unstructured":"Morin, F. & Bengio, Y. Hierarchical probabilistic neural network language model. In Proc. Tenth International Workshop on Artificial Intelligence and Statistics, Vol. R5 (eds Cowell, R. G. & Ghahramani, Z.) 246\u2013252 (PMLR, 2005)."},{"key":"11111_CR41","unstructured":"Grave, \u00c9., Joulin, A., Ciss\u00e9, M., Grangier, D. & J\u00e9gou, H. Efficient softmax approximation for GPUs. In Proc. 34th International Conference on Machine Learning, Vol. 70 (eds Precup, D. & Teh, Y. W.) 1302\u20131310 (PMLR, 2017)."},{"key":"11111_CR42","unstructured":"Ilharco, G. et al. Editing models with task arithmetic. In Proc. Eleventh International Conference on Learning Representations (eds Liu, Y. et al.) (International Conference on Learning Representations, 2023)."},{"key":"11111_CR43","unstructured":"Lambert, N. et al. Tulu 3: pushing frontiers in open language model post-training. In Proc. Second Conference on Language Modeling (2025)."},{"key":"11111_CR44","unstructured":"Pich\u00e9, A., Kamalloo, E., Pardinas, R., Chen, X. & Bahdanau, D. PipelineRL: faster on-policy reinforcement learning for long sequence generation. Trans. Mach. Learn. Res. https:\/\/openreview.net\/forum?id=A35ak14Cyp (2026)."},{"key":"11111_CR45","doi-asserted-by":"publisher","unstructured":"Zhou, J. et al. Instruction-following evaluation for large language models. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2311.07911 (2023).","DOI":"10.48550\/arXiv.2311.07911"},{"key":"11111_CR46","unstructured":"Huang, H. et al. Over-tokenized transformer: vocabulary is generally worth scaling. In Proc. 42nd International Conference on Machine Learning, Vol. 267 (eds Singh, A. et al.) 26261\u201326282 (PMLR, 2025)."},{"key":"11111_CR47","unstructured":"Tito Svenstrup, D., Hansen, J. & Winther, O. Hash embeddings for efficient word representations. In Advances in Neural Information Processing Systems, Vol. 30 (eds Guyon, I. et al.) (Curran Associates, 2017)."},{"key":"11111_CR48","unstructured":"Athanasiadis, I., Karmush, A. & Felsberg, M. Grounding functional similarity by invariance-aware model stitching. In Proc. Forty-third International Conference on Machine Learning (2026)."},{"key":"11111_CR49","doi-asserted-by":"crossref","unstructured":"Minixhofer, B., Vuli\u0107, I. & Ponti, E. Universal cross-tokenizer distillation via approximate likelihood matching. In Advances in Neural Information Processing Systems, Vol. 38 (eds Belgrave, D. et al.) 79297\u201379326 (Curran Associates, 2025).","DOI":"10.52202\/085713-2653"},{"key":"11111_CR50","doi-asserted-by":"crossref","unstructured":"Feher, D., Vuli\u0107, I. & Minixhofer, B. Retrofitting large language models with dynamic tokenization. In Proc. 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Che, W. et al.) 29866\u201329883 (Association for Computational Linguistics, 2025).","DOI":"10.18653\/v1\/2025.acl-long.1444"},{"key":"11111_CR51","doi-asserted-by":"publisher","first-page":"5586","DOI":"10.1109\/TKDE.2021.3070203","volume":"34","author":"Y Zhang","year":"2022","unstructured":"Zhang, Y. & Yang, Q. A survey on multi-task learning. IEEE Trans. Knowl. Data Eng. 34, 5586\u20135609 (2022).","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"11111_CR52","doi-asserted-by":"crossref","unstructured":"Edman, L., Schmid, H. & Fraser, A. EXECUTE: a multilingual benchmark for LLM token understanding. In Proc. Findings of the Association for Computational Linguistics: ACL 2025 (eds Che, W. et al.) 1878\u20131887 (Association for Computational Linguistics, 2025).","DOI":"10.18653\/v1\/2025.findings-acl.95"},{"key":"11111_CR53","doi-asserted-by":"publisher","unstructured":"Clark, P. et al. Think you have solved question answering? Try ARC, the AI2 reasoning challenge. Preprint at https:\/\/doi.org\/10.48550\/arXiv.1803.05457 (2018).","DOI":"10.48550\/arXiv.1803.05457"},{"key":"11111_CR54","unstructured":"Hendrycks, D. et al. Measuring massive multitask language understanding. In Proc. International Conference on Learning Representations (ICLR) (eds Mohamed, S. et al.) (International Conference on Learning Representations, 2021)."},{"key":"11111_CR55","doi-asserted-by":"crossref","unstructured":"Talmor, A., Herzig, J., Lourie, N. & Berant, J. CommonsenseQA: a question answering challenge targeting commonsense knowledge. In Proc. 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (eds Burstein, J. et al.) 4149\u20134158 (Association for Computational Linguistics, 2019).","DOI":"10.18653\/v1\/N19-1421"},{"key":"11111_CR56","doi-asserted-by":"crossref","unstructured":"Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A. & Choi, Y. HellaSwag: Can a machine really finish your sentence? In Proc. 57th Annual Meeting of the Association for Computational Linguistics (eds Korhonen, A. et al.) 4791\u20134800 (Association for Computational Linguistics, 2019).","DOI":"10.18653\/v1\/P19-1472"},{"key":"11111_CR57","doi-asserted-by":"crossref","unstructured":"Sakaguchi, K., Le Bras, R., Bhagavatula, C. & Choi, Y. WinoGrande: an adversarial winograd schema challenge at scale. In Proc. AAAI Conference on Artificial Intelligence, Vol. 34 (eds Conitzer, V. & Sha, F.) 8732\u20138740 (AAAI Press, 2020).","DOI":"10.1609\/aaai.v34i05.6399"},{"key":"11111_CR58","doi-asserted-by":"crossref","unstructured":"Sap, M., Rashkin, H., Chen, D., Le Bras, R. & Choi, Y. Social IQa: commonsense reasoning about social interactions. In Proc. 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (eds Inui, K. et al.) 4463\u20134473 (Association for Computational Linguistics, 2019).","DOI":"10.18653\/v1\/D19-1454"},{"key":"11111_CR59","doi-asserted-by":"crossref","unstructured":"Bisk, Y., Zellers, R., Le Bras, R., Gao, J. & Choi, Y. PIQA: reasoning about physical commonsense in natural language. In Proc. AAAI Conference on Artificial Intelligence, Vol. 34 (eds Conitzer, V. & Sha, F.) 7432\u20137439 (AAAI Press, 2020).","DOI":"10.1609\/aaai.v34i05.6239"},{"key":"11111_CR60","doi-asserted-by":"publisher","unstructured":"Chen, M. et al. Evaluating large language models trained on code. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2107.03374 (2021).","DOI":"10.48550\/arXiv.2107.03374"},{"key":"11111_CR61","doi-asserted-by":"publisher","unstructured":"Austin, J. et al. Program synthesis with large language models. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2108.07732 (2021).","DOI":"10.48550\/arXiv.2108.07732"},{"key":"11111_CR62","unstructured":"Lai, Y. et al. DS-1000: a natural and reliable benchmark for data science code generation. In Proc. Fortieth International Conference on Machine Learning, Vol. 202 (eds Krause, A. et al.) 18319\u201318345 (PMLR, 2023)."},{"key":"11111_CR63","doi-asserted-by":"publisher","unstructured":"Guo, D. et al. DeepSeek-Coder: when the large language model meets programming\u2014the rise of code intelligence. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2401.14196 (2024).","DOI":"10.48550\/arXiv.2401.14196"},{"key":"11111_CR64","doi-asserted-by":"crossref","unstructured":"Cassano, F. et al. MultiPL-E: a scalable and polyglot approach to benchmarking neural code generation. IEEE Trans. Softw. Eng. 49, 3675\u20133691 (2023).","DOI":"10.1109\/TSE.2023.3267446"},{"key":"11111_CR65","doi-asserted-by":"publisher","unstructured":"Cobbe, K. et al. Training verifiers to solve math word problems. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2110.14168 (2021).","DOI":"10.48550\/arXiv.2110.14168"},{"key":"11111_CR66","doi-asserted-by":"crossref","unstructured":"Lewkowycz, A. et al. Solving quantitative reasoning problems with language models. In Advances in Neural Information Processing Systems, Vol. 35 (eds Koyejo, S. et al.) 3843\u20133857 (Curran Associates, 2022).","DOI":"10.52202\/068431-0278"},{"key":"11111_CR67","unstructured":"Pal, A., Umapathi, L. K. & Sankarasubbu, M. MedMCQA: a large-scale multi-subject multi-choice dataset for medical domain question answering. In Proc. Conference on Health, Inference, and Learning, Vol. 174 (eds Flores, G. et al.) 248\u2013260 (PMLR, 2022)."},{"key":"11111_CR68","doi-asserted-by":"publisher","first-page":"6421","DOI":"10.3390\/app11146421","volume":"11","author":"D Jin","year":"2021","unstructured":"Jin, D. et al. What disease does this patient have? A large-scale open domain question answering dataset from medical exams. Appl. Sci. 11, 6421 (2021).","journal-title":"Appl. Sci."},{"key":"11111_CR69","doi-asserted-by":"crossref","unstructured":"Welbl, J., Liu, N. F. & Gardner, M. Crowdsourcing multiple choice science questions. In Proc. 3rd Workshop on Noisy User-generated Text (eds Derczynski, L. et al.) 94\u2013106 (Association for Computational Linguistics, 2017).","DOI":"10.18653\/v1\/W17-4413"},{"key":"11111_CR70","doi-asserted-by":"crossref","unstructured":"Dua, D. et al. DROP: a reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proc. 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (eds Burstein, J. et al.) 2368\u20132378 (Association for Computational Linguistics, 2019).","DOI":"10.18653\/v1\/N19-1246"},{"key":"11111_CR71","first-page":"452","volume":"7","author":"T Kwiatkowski","year":"2019","unstructured":"Kwiatkowski, T. et al. Natural questions: a benchmark for question answering research. Trans. Assoc. Comput. Linguist. 7, 452\u2013466 (2019).","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"11111_CR72","doi-asserted-by":"crossref","unstructured":"Rajpurkar, P., Zhang, J., Lopyrev, K. & Liang, P. SQuAD: 100,000+ questions for machine comprehension of text. In Proc. 2016 Conference on Empirical Methods in Natural Language Processing, 2383\u20132392 (Association for Computational Linguistics, 2016).","DOI":"10.18653\/v1\/D16-1264"},{"key":"11111_CR73","doi-asserted-by":"publisher","first-page":"249","DOI":"10.1162\/tacl_a_00266","volume":"7","author":"S Reddy","year":"2019","unstructured":"Reddy, S., Chen, D. & Manning, C. D. CoQA: a conversational question answering challenge. Trans. Assoc. Comput. Linguist. 7, 249\u2013266 (2019).","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"11111_CR74","doi-asserted-by":"crossref","unstructured":"Paperno, D. et al. The LAMBADA dataset: word prediction requiring a broad discourse context. In Proc. 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Erk, K. & Smith, N. A.) 1525\u20131534 (Association for Computational Linguistics, 2016).","DOI":"10.18653\/v1\/P16-1144"},{"key":"11111_CR75","doi-asserted-by":"crossref","unstructured":"Ansel, J. et al. PyTorch 2: faster machine learning through dynamic Python bytecode transformation and graph compilation. In Proc. 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Vol. 2 (eds Abu-Ghazaleh, N. et al.) 929\u2013947 (Association for Computing Machinery, 2024).","DOI":"10.1145\/3620665.3640366"},{"key":"11111_CR76","doi-asserted-by":"crossref","unstructured":"Kwon, W. et al. Efficient memory management for large language model serving with PagedAttention. In Proc. 29th Symposium on Operating Systems Principles (eds Flinn, J. et al.) 611\u2013626 (Association for Computing Machinery, 2023).","DOI":"10.1145\/3600006.3613165"},{"key":"11111_CR77","doi-asserted-by":"publisher","first-page":"2523","DOI":"10.1109\/TASLP.2023.3288409","volume":"31","author":"Z Borsos","year":"2023","unstructured":"Borsos, Z. et al. AudioLM: a language modeling approach to audio generation. IEEE\/ACM Trans. Audio Speech Lang. Process. 31, 2523\u20132533 (2023).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"11111_CR78","unstructured":"Dosovitskiy, A. et al. An image is worth 16 \u00d7 16 words: transformers for image recognition at scale. In Proc. International Conference on Learning Representations (eds Mohamed, S. et al.) (International Conference on Learning Representations, 2021)."},{"key":"11111_CR79","doi-asserted-by":"crossref","unstructured":"Kaushal, A. & Mahowald, K. What do tokens know about their characters and how do they know it? In Proc. 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (eds Carpuat, M. et al.) 2487\u20132507 (Association for Computational Linguistics, 2022).","DOI":"10.18653\/v1\/2022.naacl-main.179"},{"key":"11111_CR80","doi-asserted-by":"crossref","unstructured":"Xu, Z. et al. Enhancing character-level understanding in LLMs through token internal structure learning. In Proc. 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Che, W. et al.) 3839\u20133853 (Association for Computational Linguistics, 2025).","DOI":"10.18653\/v1\/2025.acl-long.194"},{"key":"11111_CR81","doi-asserted-by":"crossref","unstructured":"Hiraoka, T. & Inui, K. Spelling-out is not straightforward: LLMs' capability of tokenization from token to characters. In Proc. Findings of the Association for Computational Linguistics: EMNLP 2025 (eds Christodoulopoulos, C. et al.) 13340\u201313353 (Association for Computational Linguistics, 2025).","DOI":"10.18653\/v1\/2025.findings-emnlp.719"},{"key":"11111_CR82","doi-asserted-by":"crossref","unstructured":"\u0141a\u0144cucki, A., Staniszewski, K., Nawrot, P. & Ponti, E. Inference-time hyper-scaling with KV cache compression. In Advances in Neural Information Processing Systems, Vol. 38 (eds Belgrave, D. et al.) 9365\u20139397 (Curran Associates, 2025).","DOI":"10.52202\/085713-0317"},{"key":"11111_CR83","unstructured":"Gloeckle, F., Youbi Idrissi, B., Roziere, B., Lopez-Paz, D. & Synnaeve, G. Better & faster large language models via multi-token prediction. In Proc. Forty-first International Conference on Machine Learning, Vol. 235 (eds Salakhutdinov, R. et al.) 15706\u201315734 (PMLR, 2024)."},{"key":"11111_CR84","doi-asserted-by":"crossref","unstructured":"Lotz, J., Salesky, E., Rust, P. & Elliott, D. Text rendering strategies for pixel language models. In Proc. 2023 Conference on Empirical Methods in Natural Language Processing (eds Bouamor, H. et al.) 10155\u201310172 (Association for Computational Linguistics, 2023).","DOI":"10.18653\/v1\/2023.emnlp-main.628"},{"key":"11111_CR85","unstructured":"Rust, P. et al. Language modelling with pixels. In Proc. Eleventh International Conference on Learning Representations (eds Liu, Y. et al.) (International Conference on Learning Representations, 2023)."},{"key":"11111_CR86","doi-asserted-by":"publisher","unstructured":"Wei, H., Sun, Y. & Li, Y. DeepSeek-OCR: contexts optical compression. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2510.18234 (2025).","DOI":"10.48550\/arXiv.2510.18234"},{"key":"11111_CR87","doi-asserted-by":"publisher","first-page":"291","DOI":"10.1162\/tacl_a_00461","volume":"10","author":"L Xue","year":"2022","unstructured":"Xue, L. et al. ByT5: towards a token-free future with pre-trained byte-to-byte models. Trans. Assoc. Comput. Linguist. 10, 291\u2013306 (2022).","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"11111_CR88","doi-asserted-by":"crossref","unstructured":"Limisiewicz, T., Blevins, T., Gonen, H., Ahia, O. & Zettlemoyer, L. MYTE: morphology-driven byte encoding for better and fairer multilingual language modeling. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Ku, L.-W. et al.) 15059\u201315076 (Association for Computational Linguistics, 2024).","DOI":"10.18653\/v1\/2024.acl-long.804"},{"key":"11111_CR89","doi-asserted-by":"publisher","unstructured":"Land, S. & Arnett, C. BPE stays on SCRIPT: structured encoding for robust multilingual pretokenization. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2505.24689 (2025).","DOI":"10.48550\/arXiv.2505.24689"},{"key":"11111_CR90","doi-asserted-by":"crossref","unstructured":"Nawrot, P. et al. Hierarchical transformers are more efficient language models. In Proc. Findings of the Association for Computational Linguistics: NAACL 2022 (eds Carpuat, M. et al.) 1559\u20131571 (Association for Computational Linguistics, 2022).","DOI":"10.18653\/v1\/2022.findings-naacl.117"},{"key":"11111_CR91","doi-asserted-by":"crossref","unstructured":"Yu, L. et al. MEGABYTE: predicting million-byte sequences with multiscale transformers. In Advances in Neural Information Processing Systems, Vol. 36 (eds Oh, A. et al.) 78808\u201378823 (Curran Associates, 2023).","DOI":"10.52202\/075280-3447"},{"key":"11111_CR92","doi-asserted-by":"crossref","unstructured":"Ho, N. et al. Block transformer: global-to-local language modeling for fast inference. In Advances in Neural Information Processing Systems, Vol. 37 (eds Globerson, A. et al.) 48740\u201348783 (Curran Associates, 2024).","DOI":"10.52202\/079017-1545"},{"key":"11111_CR93","doi-asserted-by":"crossref","unstructured":"Dolga, R., Maystre, L., Berariu, T. & Barber, D. From characters to tokens: dynamic grouping with hierarchical BPE. In Proc. Findings of the Association for Computational Linguistics: EMNLP 2025 (eds Christodoulopoulos, C. et al.) 11154\u201311162 (Association for Computational Linguistics, 2025).","DOI":"10.18653\/v1\/2025.findings-emnlp.595"},{"key":"11111_CR94","unstructured":"Kallini, J., Murty, S., Manning, C. D., Potts, C. & Csord\u00e1s, R. MrT5: dynamic token merging for efficient byte-level language models. In Proc. Thirteenth International Conference on Learning Representations (eds Yue, Y. et al.) 56646\u201356669 (International Conference on Learning Representations, 2025)."},{"key":"11111_CR95","doi-asserted-by":"crossref","unstructured":"Geng, S. et al. zip2zip: inference-time adaptive tokenization via online compression. In Advances in Neural Information Processing Systems, Vol. 38 (eds Belgrave, D. et al.) 151818\u2013151847 (Curran Associates, 2025).","DOI":"10.52202\/085713-5077"},{"key":"11111_CR96","doi-asserted-by":"crossref","unstructured":"Bick, A., Li, K. Y., Xing, E. P., Kolter, J. Z. & Gu, A. Transformers to SSMs: distilling quadratic knowledge to subquadratic models. In Advances in Neural Information Processing Systems, Vol. 37 (eds Globerson, A. et al.) 31788\u201331812 (Curran Associates, 2024).","DOI":"10.52202\/079017-0999"},{"key":"11111_CR97","doi-asserted-by":"publisher","unstructured":"Tran, K. From English to foreign languages: transferring pre-trained language models. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2002.07306 (2020).","DOI":"10.48550\/arXiv.2002.07306"},{"key":"11111_CR98","doi-asserted-by":"crossref","unstructured":"Minixhofer, B., Paischer, F. & Rekabsaz, N. WECHSEL: effective initialization of subword embeddings for cross-lingual transfer of monolingual language models. In Proc. 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (eds Carpuat, M. et al.) 3992\u20134006 (Association for Computational Linguistics, 2022).","DOI":"10.18653\/v1\/2022.naacl-main.293"},{"key":"11111_CR99","doi-asserted-by":"crossref","unstructured":"Minixhofer, B., Ponti, E. M. & Vuli\u0107, I. Zero-shot tokenizer transfer. In Advances in Neural Information Processing Systems, Vol. 37 (eds Globerson, A. et al.) 46791\u201346818 (Curran Associates, 2024).","DOI":"10.52202\/079017-1484"},{"key":"11111_CR100","unstructured":"Dobler, K., Elliott, D. & de Melo, G. Token distillation: attention-aware input embeddings for new tokens. In Proc. Fourteenth International Conference on Learning Representations (eds Vondrick, C. et al.) 140590\u2013140615 (International Conference on Learning Representations, 2026)."},{"key":"11111_CR101","doi-asserted-by":"publisher","unstructured":"Haltiuk, M. & Smywi\u0144ski-Pohl, A. Model-aware tokenizer transfer. Preprint at https:\/\/doi.org\/10.48550\/arXiv.2510.21954 (2025).","DOI":"10.48550\/arXiv.2510.21954"},{"key":"11111_CR102","unstructured":"Shenfeld, I., Pari, J. & Agrawal, P. RL's razor: why online reinforcement learning forgets less. In Proc. Fourteenth International Conference on Learning Representations (eds Vondrick, C. et al.) 59839\u201359864 (International Conference on Learning Representations, 2026)."},{"key":"11111_CR103","doi-asserted-by":"crossref","unstructured":"Gu, Y. et al. OLMES: a standard for language model evaluations. In Proc. Findings of the Association for Computational Linguistics: NAACL 2025 (eds Chiruzzo, L. et al.) 5020\u20135048 (Association for Computational Linguistics, 2025).","DOI":"10.18653\/v1\/2025.findings-naacl.282"}],"container-title":["Nature"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s41586-026-11111-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41586-026-11111-4","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s41586-026-11111-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,10,7]],"date-time":"2026-10-07T15:03:40Z","timestamp":1791385420000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s41586-026-11111-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,10,7]]},"references-count":103,"alternative-id":["11111"],"URL":"https:\/\/doi.org\/10.1038\/s41586-026-11111-4","relation":{},"ISSN":["0028-0836","1476-4687"],"issn-type":[{"value":"0028-0836","type":"print"},{"value":"1476-4687","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,10,7]]},"assertion":[{"value":"13 February 2026","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 September 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 October 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","label":"Competing interests","group":{"name":"EthicsHeading","label":"Ethics"}}]}}