{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T13:09:10Z","timestamp":1779887350111,"version":"3.53.1"},"reference-count":63,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T00:00:00Z","timestamp":1779840000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>The success of large neural networks trained for sequential prediction via log-loss minimization over massive and diverse datasets has sparked debate regarding the fundamental limits of this paradigm. While these models are not explicitly programmed to perform planning and search, their behavior increasingly resembles complex reasoning and adaptive problem-solving. This paper reviews a series of theoretical and empirical works, aiming to bridge the gap between the practical success of LLMs and formal theories of computation and intelligence\u2014that is, algorithmic information theory and Universal Artificial Intelligence. Grounded in the framework of memory-based meta-learning, the main argument is that training sequence models to predict the next token across diverse tasks implicitly meta-trains them to perform algorithmic compression, thereby performing (amortized) Bayesian inference over the task in-context. Consequently, when pretrained on a sufficiently rich data distribution, the resulting neural networks behave as if compressing by inferring the generative algorithm producing the observed data. We discuss recent theoretical and empirical evidence demonstrating that this approach can approximate Solomonoff induction in the theoretical limit, match exact Bayesian inference on complex sources in practice, achieve strong compression on out-of-distribution data, and synthesize complex in-context algorithms like chessboard evaluations. As models become more capable and general, the theoretical understanding through the lens of algorithmic information theory, including hard theoretical limits and how far practical models are from them, becomes increasingly relevant. We thus conclude our paper by outlining a number of open research questions to further bridge the gap from well-understood theory to modern machine learning practice.<\/jats:p>","DOI":"10.3390\/e28060596","type":"journal-article","created":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T12:31:56Z","timestamp":1779885116000},"page":"596","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Algorithmic Compression via Pretrained Neural Networks"],"prefix":"10.3390","volume":"28","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8039-4027","authenticated-orcid":false,"given":"Tim","family":"Genewein","sequence":"first","affiliation":[{"name":"Google DeepMind, London N1C 4DJ, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8868-8270","authenticated-orcid":false,"given":"Jordi","family":"Grau-Moya","sequence":"additional","affiliation":[{"name":"Google DeepMind, London N1C 4DJ, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7090-3078","authenticated-orcid":false,"given":"Li Kevin","family":"Wenliang","sequence":"additional","affiliation":[{"name":"Google DeepMind, London N1C 4DJ, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0857-4039","authenticated-orcid":false,"given":"Laurent","family":"Orseau","sequence":"additional","affiliation":[{"name":"Google DeepMind, London N1C 4DJ, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3263-4097","authenticated-orcid":false,"given":"Marcus","family":"Hutter","sequence":"additional","affiliation":[{"name":"Google DeepMind, London N1C 4DJ, UK"},{"name":"ANU School of Computing, Australian National University, Canberra, ACT 2601, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,5,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/S0019-9958(64)90223-2","article-title":"A formal theory of inductive inference. Part I","volume":"7","author":"Solomonoff","year":"1964","journal-title":"Inf. Control"},{"key":"ref_2","unstructured":"Hutter, M. (2004). Universal Artificial Intelligence: Sequential Decisions Based on Algorithmic Probability, Springer Science & Business Media."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Hutter, M., Quarel, D., and Catt, E. (2024). An Introduction to Universal Artificial Intelligence, Chapman & Hall.","DOI":"10.1201\/9781003460299"},{"key":"ref_4","unstructured":"Ortega, P.A., Wang, J.X., Rowland, M., Genewein, T., Kurth-Nelson, Z., Pascanu, R., Heess, N., Veness, J., Pritzel, A., and Sprechmann, P. (2019). Meta-learning of sequential strategies. arXiv."},{"key":"ref_5","first-page":"166910","article-title":"Understanding Prompt Tuning and In-Context Learning via Meta-Learning","volume":"38","author":"Genewein","year":"2025","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_6","unstructured":"Hutter, M. (2026, May 23). The Hutter Prize. Prize for Compressing Human Knowledge. Available online: http:\/\/prize.hutter1.net\/."},{"key":"ref_7","unstructured":"Grau-Moya, J., Genewein, T., Hutter, M., Orseau, L., Deletang, G., Catt, E., Ruoss, A., Wenliang, L.K., Mattern, C., and Aitchison, M. (2024). Learning Universal Predictors. Proceedings of the International Conference on Machine Learning, Vienna, Austria, 21\u201327 July 2024, PMLR."},{"key":"ref_8","unstructured":"Genewein, T., Del\u00e9tang, G., Ruoss, A., Wenliang, L.K., Catt, E., Dutordoir, V., Grau-Moya, J., Orseau, L., Hutter, M., and Veness, J. (2023). Memory-based meta-learning on non-stationary distributions. Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA, 23\u201329 July 2023, PMLR."},{"key":"ref_9","unstructured":"Del\u00e9tang, G., Ruoss, A., Duquenne, P., Catt, E., Genewein, T., Mattern, C., Grau-Moya, J., Wenliang, L.K., Aitchison, M., and Orseau, L. (2024, January 7\u201311). Language Modeling Is Compression. Proceedings of the Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria."},{"key":"ref_10","unstructured":"Heurtel-Depeiges, D., Ruoss, A., Veness, J., and Genewein, T. (2025, January 13\u201319). Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data. Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada."},{"key":"ref_11","first-page":"65765","article-title":"Amortized planning with large-scale transformers: A case study on chess","volume":"37","author":"Ruoss","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_12","first-page":"18691","article-title":"Meta-trained agents implement bayes-optimal agents","volume":"33","author":"Mikulik","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_13","first-page":"27181","article-title":"Self-predictive universal AI","volume":"36","author":"Catt","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_14","unstructured":"Ortega, P.A., Kunesch, M., Del\u00e9tang, G., Genewein, T., Grau-Moya, J., Veness, J., Buchli, J., Degrave, J., Piot, B., and Perolat, J. (2021). Shaking the foundations: Delusions in sequence models for interaction and control. arXiv."},{"key":"ref_15","unstructured":"Ruoss, A., Pardo, F., Chan, H., Li, B., Mnih, V., and Genewein, T. (2025, January 13\u201319). LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations. Proceedings of the Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Li, M., and Vit\u00e1nyi, P. (2008). An Introduction to Kolmogorov Complexity and Its Applications, Springer.","DOI":"10.1007\/978-0-387-49820-1"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"198","DOI":"10.1147\/rd.203.0198","article-title":"Generalized Kraft inequality and arithmetic coding","volume":"20","author":"Rissanen","year":"1976","journal-title":"IBM J. Res. Dev."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"465","DOI":"10.1016\/0005-1098(78)90005-5","article-title":"Modeling by shortest data description","volume":"14","author":"Rissanen","year":"1978","journal-title":"Automatica"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Gr\u00fcnwald, P.D. (2007). The Minimum Description Length Principle, MIT Press.","DOI":"10.7551\/mitpress\/4643.001.0001"},{"key":"ref_20","unstructured":"Blier, L., and Ollivier, Y. (2018). The Description Length of Deep Learning Models. Proceedings of the Advances in Neural Information Processing Systems, Montr\u00e9al, QC, Canada, 3\u20138 December 2018, Curran Associates Inc."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Jiang, Z., Yang, M.Y., Tsirlin, M., Tang, R., Dai, Y., and Lin, J. (2023). \u201cLow-Resource\u201d Text Classification: A Parameter-Free Classification Method with Compressors. Proceedings of the Findings of the Association for Computational Linguistics: ACL 2023, Toronto, ON, Canada, 9\u201314 July 2023, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2023.findings-acl.426"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1002\/j.1538-7305.1948.tb01338.x","article-title":"A mathematical theory of communication","volume":"27","author":"Shannon","year":"1948","journal-title":"Bell Syst. Tech. J."},{"key":"ref_23","unstructured":"von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M. (2023). Transformers learn in-context by gradient descent. Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA, 23\u201329 July 2023, PMLR."},{"key":"ref_24","unstructured":"Aky\u00fcrek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D. (2023, January 1\u20135). What learning algorithm is in-context learning? Investigations with linear models. Proceedings of the International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_25","unstructured":"von Oswald, J., Schlegel, M., Meulemans, A., Kobayashi, S., Niklasson, E., Zucchet, N., Scherrer, N., Miller, N., Sandler, M., and y Arcas, B.A. (2024). Uncovering mesa-optimization algorithms in Transformers. arXiv."},{"key":"ref_26","unstructured":"Laskin, M., Wang, L., Oh, J., Parisotto, E., Spencer, S., Steiber, R., Strouse, D., Hansen, S.S., Filos, A., and Brooks, E. (2023, January 1\u20135). In-Context Reinforcement Learning with Algorithm Distillation. Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_27","unstructured":"Lampinen, A.K., Chan, S.C.Y., Singh, A.K., and Shanahan, M. (2024). The broader spectrum of in-context learning. arXiv."},{"key":"ref_28","unstructured":"Falck, F., Wang, Z., and Holmes, C.C. (2024). Is In-Context Learning in Large Language Models Bayesian? A Martingale Perspective. Proceedings of the International Conference on Machine Learning, Vienna, Austria, 21\u201327 July 2024, PMLR."},{"key":"ref_29","unstructured":"Lampinen, A.K., Chaudhry, A., Chan, S.C.Y., Wild, C., Wan, D., Ku, A., Bornschein, J., Pascanu, R., Shanahan, M., and McClelland, J.L. (2025). On the Generalization of Language Models from In-Context Learning and Finetuning: A Controlled Study. arXiv."},{"key":"ref_30","unstructured":"Wenliang, L.K., Ruoss, A., Grau-Moya, J., Hutter, M., and Genewein, T. (2026, January 2\u20134). Why is prompting hard? Understanding prompts on binary sequence predictors. Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS), Valencia, Spain."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Veness, J., White, M., Bowling, M., and Gy\u00f6rgy, A. (2012). Partition Tree Weighting. arXiv.","DOI":"10.1109\/DCC.2013.40"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"653","DOI":"10.1109\/18.382012","article-title":"The context-tree weighting method: Basic properties","volume":"41","author":"Willems","year":"1995","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_33","unstructured":"Del\u00e9tang, G., Ruoss, A., Grau-Moya, J., Genewein, T., Wenliang, L.K., Catt, E., Cundy, C., Hutter, M., Legg, S., and Veness, J. (2023, January 1\u20135). Neural Networks and the Chomsky Hierarchy. Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_34","unstructured":"Shaw, P., Cohan, J., Eisenstein, J., and Toutanova, K. (2026, January 23\u201327). Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers. Proceedings of the International Conference on Learning Representations (ICLR), Rio de Janeiro, Brazil."},{"key":"ref_35","unstructured":"Bornschein, J., Li, Y., and Hutter, M. (2023, January 1\u20135). Sequential Learning of Neural Networks for Prequential MDL. Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda."},{"key":"ref_36","unstructured":"Shinnick, Z., Jiang, L., Saratchandran, H., van den Hengel, A., and Teney, D. (2025, January 19). Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning. Proceedings of the ICML 2025 Workshop on Methods and Opportunities at Small Scale, Vancouver, BC, Canada."},{"key":"ref_37","unstructured":"Li, K., Hopkins, A.K., Bau, D., Vi\u00e9gas, F., Pfister, H., and Wattenberg, M. (2023, January 1\u20135). Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task. Proceedings of the International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_38","unstructured":"Jenner, E., Kapur, S., Georgiev, V., Allen, C., Emmons, S., and Russell, S. (2024). Evidence of Learned Look-Ahead in a Chess-Playing Neural Network. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 10\u201315 December 2024, Curran Associates, Inc."},{"key":"ref_39","unstructured":"Cruz, D. (2025). Understanding the learned look-ahead behavior of chess neural networks. arXiv."},{"key":"ref_40","unstructured":"Monroe, D., and Chalmers, P.A. (2024). Mastering Chess with a Transformer Model. arXiv."},{"key":"ref_41","unstructured":"Deora, P., Vasudeva, B., Behnia, T., and Thrampoulidis, C. (2025, January 7\u201310). In-Context Occam\u2019s Razor: How Transformers Prefer Simpler Hypotheses on the Fly. Proceedings of the Conference on Language Modeling, Montreal, QC, Canada."},{"key":"ref_42","unstructured":"Elmoznino, E., Marty, T., Kasetty, T., Gagnon, L., Mittal, S., Fathi, M., Sridhar, D., and Lajoie, G. (2025, January 13\u201319). In-context learning and Occam\u2019s razor. Proceedings of the International Conference on Machine Learning, Vancouver, BC, Canada."},{"key":"ref_43","unstructured":"Adaptive Agent Team, Bauer, J., Baumli, K., Baveja, S., Behbahani, F., Bhatt, A., Bhoopchand, A., Chang, M., Clay, N., and Collister, A. (2023). Human-Timescale Adaptation in an Open-Ended Task Space. Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA, 23\u201329 July 2023, PMLR."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Chan, S.C.Y., Santoro, A., Lampinen, A.K., Wang, J.X., Singh, A.K., Richemond, P.H., McClelland, J.L., and Hill, F. (2022). Data Distributional Properties Drive Emergent In-Context Learning in Transformers. Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA, 28 November\u20139 December 2022, Curran Associates, Inc.","DOI":"10.52202\/068431-1371"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Pearl, J. (2009). Causality: Models, Reasoning, and Inference, Cambridge University Press.","DOI":"10.1017\/CBO9780511803161"},{"key":"ref_46","unstructured":"de Haan, P., Jayaraman, D., and Levine, S. (2019). Causal Confusion in Imitation Learning. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 8\u201314 November 2019, Curran Associates, Inc."},{"key":"ref_47","unstructured":"Reed, S., Zolna, K., Parisotto, E., Gomez Colmenarejo, S., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., and Springenberg, J.T. (2022). A Generalist Agent. arXiv."},{"key":"ref_48","unstructured":"Paglieri, D., Cupia\u0142, B., Coward, S., Piterbarg, U., Wolczyk, M., Khan, A., Pignatelli, E., Kuci\u0144ski, \u0141., Pinto, L., and Fergus, R. (2025, January 24\u201328). BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games. Proceedings of the International Conference on Learning Representations, Singapore."},{"key":"ref_49","unstructured":"Ortega, P.A. (2026). Universal Artificial Intelligence as Imitation, Daios Technologies. Available online: https:\/\/www.adaptiveagents.org\/uiai."},{"key":"ref_50","unstructured":"Shao, D., Kleine Buening, T., and Kwiatkowska, M. (2025, January 28). A Unifying Framework for Causal Imitation Learning with Hidden Confounders. Proceedings of the ICLR 2025 Workshop on Spurious Correlation and Shortcut Learning, Singapore."},{"key":"ref_51","unstructured":"Xie, S.M., Raghunathan, A., Liang, P., and Ma, T. (2022, January 25\u201329). An Explanation of In-context Learning as Implicit Bayesian Inference. Proceedings of the International Conference on Learning Representations, Virtual."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Garg, S., Tsipras, D., Liang, P., and Valiant, G. (2022). What Can Transformers Learn In-Context? A Case Study of Simple Function Classes. Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA, 28 November\u20139 December 2022, Curran Associates, Inc.","DOI":"10.52202\/068431-2217"},{"key":"ref_53","unstructured":"Kirsch, L., Harrison, J., Sohl-Dickstein, J., and Metz, L. (2022). General-Purpose In-Context Learning by Meta-Learning Transformers. arXiv."},{"key":"ref_54","unstructured":"Yadlowsky, S., Doshi, L., and Tripuraneni, N. (2023). Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models. arXiv."},{"key":"ref_55","first-page":"65189","article-title":"Meta-in-context learning in large language models","volume":"36","author":"Binz","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_56","unstructured":"Mirchandani, S., Xia, F., Florence, P., Ichter, B., Driess, D., Arenas, M.G., Rao, K., Sadigh, D., and Zeng, A. (2023). Large Language Models as General Pattern Machines. Proceedings of the Conference on Robot Learning, Atlanta, GA, USA, 6\u20139 November 2023, PMLR."},{"key":"ref_57","unstructured":"Ravi, S., and Beatson, A. (2019, January 6\u20139). Amortized Bayesian Meta-Learning. Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA."},{"key":"ref_58","unstructured":"Grant, E., Finn, C., Levine, S., Darrell, T., and Griffiths, T.L. (May, January 30). Recasting Gradient-Based Meta-Learning as Hierarchical Bayes. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_59","unstructured":"M\u00fcller, S., Hollmann, N., Pineda Arango, S., Grabocka, J., and Hutter, F. (2022, January 25\u201329). Transformers can do Bayesian inference. Proceedings of the International Conference on Learning Representations, Virtual."},{"key":"ref_60","unstructured":"Hollmann, N., M\u00fcller, S., Eggensperger, K., and Hutter, F. (2023, January 1\u20135). TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second. Proceedings of the International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"1452","DOI":"10.1109\/TNNLS.2020.3042395","article-title":"BayesFlow: Learning complex stochastic models with invertible neural networks","volume":"33","author":"Radev","year":"2020","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_62","unstructured":"Reuter, A., Rudner, T.G.J., Fortuin, V., and R\u00fcgamer, D. (2025). Can Transformers Learn Full Bayesian Inference in Context?. Proceedings of the International Conference on Machine Learning, Vancouver, BC, Canada, 13\u201319 July 2025, PMLR."},{"key":"ref_63","unstructured":"Wan, J., and Mei, L. (2025). Large Language Models as Computable Approximations to Solomonoff Induction. arXiv."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/28\/6\/596\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T12:44:21Z","timestamp":1779885861000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/28\/6\/596"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,27]]},"references-count":63,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2026,6]]}},"alternative-id":["e28060596"],"URL":"https:\/\/doi.org\/10.3390\/e28060596","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,27]]}}}