{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T07:01:21Z","timestamp":1780038081873,"version":"3.53.1"},"reference-count":57,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T00:00:00Z","timestamp":1780012800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Comput. Neurosci."],"abstract":"<jats:p>Understanding and controlling the complexity of neural networks is a central challenge in machine learning, with implications for generalization, optimization, and model capacity. While most approaches rely on entropy-based loss functions and statistical metrics, these measures often fail to capture deeper, causally relevant algorithmic regularities embedded in network structure. We propose a shift toward algorithmic information theory, using binarized neural networks (BNNs) as a first proxy. Grounded in algorithmic probability (AP) and the universal distribution it defines, our approach characterizes learning dynamics through a formal, causally grounded lens. We apply the Block Decomposition Method (BDM), a scalable approximation of algorithmic complexity based on AP, and demonstrate that it more closely tracks structural changes during training than entropy, generally exhibiting stronger correlations with training loss across a wide range of architectures, datasets, and randomized training runs. These results support the view of training in BNNs as a process of algorithmic compression, where learning corresponds to the progressive internalization of structured regularities. In doing so, our work offers a principled estimate of learning progression and suggests a framework for complexity-aware learning and regularization, grounded in first principles from information theory, complexity, and computability.<\/jats:p>","DOI":"10.3389\/fncom.2026.1791546","type":"journal-article","created":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T06:01:22Z","timestamp":1780034482000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Binarized neural networks converge toward algorithmic simplicity: empirical support for the learning-as-compression hypothesis"],"prefix":"10.3389","volume":"20","author":[{"given":"Eduardo Y.","family":"Sakabe","sequence":"first","affiliation":[{"name":"Faculdade de Engenharia El\u00e9trica e de Computa\u00e7\u00e3o (FEEC), Universidade Estadual de Campinas (UNICAMP)","place":["Campinas, Brazil"]},{"name":"Algorithmic Dynamics Lab, King's College London","place":["London, United Kingdom"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Felipe S.","family":"Abrah\u00e3o","sequence":"additional","affiliation":[{"name":"Algorithmic Dynamics Lab, King's College London","place":["London, United Kingdom"]},{"name":"Centro de L\u00f3gica, Epistemologia e Hist\u00f3ria da Ci\u00eancia (CLE), Universidade Estadual de Campinas (UNICAMP)","place":["Campinas, Brazil"]},{"name":"Oxford Immune Algorithmics, Oxford University Innovation & London Institute for Healthcare Engineering","place":["London, United Kingdom"]},{"name":"Data Extreme Lab (DEXL), National Laboratory for Scientific Computing (LNCC)","place":["Petr\u00f3polis, Brazil"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alexandre","family":"Sim\u00f5es","sequence":"additional","affiliation":[{"name":"S\u00e3o Paulo State University (UNESP), Institute of Science and Technology","place":["Sorocaba, Brazil"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Esther","family":"Colombini","sequence":"additional","affiliation":[{"name":"Instituto de Computa\u00e7\u00e3o (IC), Universidade Estadual de Campinas (UNICAMP)","place":["Campinas, Brazil"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Paula","family":"Costa","sequence":"additional","affiliation":[{"name":"Faculdade de Engenharia El\u00e9trica e de Computa\u00e7\u00e3o (FEEC), Universidade Estadual de Campinas (UNICAMP)","place":["Campinas, Brazil"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ricardo","family":"Gudwin","sequence":"additional","affiliation":[{"name":"Faculdade de Engenharia El\u00e9trica e de Computa\u00e7\u00e3o (FEEC), Universidade Estadual de Campinas (UNICAMP)","place":["Campinas, Brazil"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hector","family":"Zenil","sequence":"additional","affiliation":[{"name":"Algorithmic Dynamics Lab, King's College London","place":["London, United Kingdom"]},{"name":"Oxford Immune Algorithmics, Oxford University Innovation & London Institute for Healthcare Engineering","place":["London, United Kingdom"]},{"name":"Research Department of Biomedical Computing, School of Biomedical Engineering and Imaging Sciences, King's College London","place":["London, United Kingdom"]},{"name":"Research Department Digital Twins, School of Biomedical Engineering and Imaging Sciences, King's College London","place":["London, United Kingdom"]},{"name":"King's Institute for Artificial Intelligence, King's College London","place":["London, United Kingdom"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2026,5,29]]},"reference":[{"key":"B1","first-page":"3","article-title":"\u201cA public domain dataset for human activity recognition using smartphones,\u201d","volume-title":"Esann 2013 proceedings, european symposium on artificial neural networks, computational intelligence and machine learning","author":"Anguita","year":"2013"},{"key":"B2","article-title":"Estimating or propagating gradients through stochastic neurons","author":"Bengio","year":"2013","journal-title":"arXiv [preprint"},{"key":"B3","article-title":"Information bottleneck analysis of deep neural networks via lossy compression","author":"Butakov","year":"2024","journal-title":"arXiv [preprint"},{"key":"B4","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-662-04978-5","volume-title":"Information and Randomness: An algorithmic Perspective","author":"Calude","year":"2002"},{"key":"B5","doi-asserted-by":"publisher","DOI":"10.1002\/0471667196.ess0029","author":"Chaitin","year":"2004","journal-title":"Algorithmic Information Theory"},{"key":"B6","article-title":"Towards the limit of network quantization","author":"Choi","year":"2017","journal-title":"arXiv [preprint"},{"key":"B7","doi-asserted-by":"publisher","first-page":"63","DOI":"10.1016\/j.amc.2011.10.006","article-title":"Numerical evaluation of algorithmic complexity for short strings: A glance into the innermost structure of randomness","volume":"219","author":"Delahaye","year":"2012","journal-title":"Appl. Math. Comput"},{"key":"B8","volume-title":"Algorithmic Randomness and Complexity. Theory and Applications of Computability","author":"Downey","year":"2010"},{"key":"B9","article-title":"The lottery ticket hypothesis: finding sparse, trainable neural networks","author":"Frankle","year":"2019","journal-title":"arXiv [preprint"},{"key":"B10","article-title":"Deep compression: compressing deep neural networks with pruning, trained quantization and huffman coding","author":"Han","year":"2016","journal-title":"arXiv [preprint"},{"key":"B11","article-title":"SuperARC: an agnostic test for narrow, general, and super intelligence based on the principles of recursive compression and algorithmic probability","author":"Hern\u00e1ndez-Espinosa","year":"2025","journal-title":"arXiv [preprint"},{"key":"B12","doi-asserted-by":"publisher","first-page":"180399","DOI":"10.1098\/rsos.180399","article-title":"Algorithmically probable mutations reproduce aspects of evolution, such as convergence rate, genetic memory and modularity","volume":"5","author":"Hern\u00e1ndez-Orozco","year":"2018","journal-title":"R. Soc. Open Sci"},{"key":"B13","doi-asserted-by":"publisher","first-page":"567656","DOI":"10.3389\/frai.2020.567356","article-title":"Algorithmic probability-guided machine learning on non-differentiable spaces","volume":"3","author":"Hern\u00e1ndez-Orozco","year":"2021","journal-title":"Front. Artif. Intell"},{"key":"B14","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1145\/168304.168306","article-title":"\u201cKeeping the neural networks simple by minimizing the description length of the weights,\u201d","author":"Hinton","year":"1993","journal-title":"Proceedings of the sixth annual conference on Computational learning theory, COLT '93"},{"key":"B15","first-page":"4107","article-title":"\u201cBinarized neural networks,\u201d","volume-title":"Advances in Neural Information Processing Systems, Vol. 29","author":"Hubara","year":"2016"},{"key":"B16","doi-asserted-by":"crossref","DOI":"10.1201\/9781003460299","volume-title":"An Introduction to Universal Artificial Intelligence","author":"Hutter","year":"2024"},{"key":"B17","doi-asserted-by":"publisher","first-page":"7","DOI":"10.1007\/BF03024407","article-title":"The miraculous universal distribution","volume":"19","author":"Kirchherr","year":"1997","journal-title":"Math. Intelligencer"},{"key":"B18","article-title":"Simulation intelligence: towards a new generation of scientific methods","author":"Lavin","year":"2022","journal-title":"arXiv [preprint"},{"key":"B19","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"Lecun","year":"1998","journal-title":"Proc. IEEE"},{"key":"B20","article-title":"An additively optimal interpreter for approximating Kolmogorov prefix complexity","author":"Leyva-Acosta","year":"2024","journal-title":"arXiv [preprint"},{"key":"B21","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-11298-1","author":"Li","year":"2019","journal-title":"An Introduction to Kolmogorov Complexity and Its Applications. Texts in Computer Science"},{"key":"B22","unstructured":"KAN: Kolmogorov-Arnold networks\n          \n          \n            \n              Liu\n              Z.\n            \n            \n              Wang\n              Y.\n            \n            \n              Vaidya\n              S.\n            \n            \n              Ruehle\n              F.\n            \n            \n              Halverson\n              J.\n            \n            \n              Solja\u010di\u0107\n              M.\n            \n          \n          arXiv [preprint\n          \n          2025"},{"key":"B23","unstructured":"The Limits of Understanding. Marvin Minsky Discusses the Importance of Algorithmic Probability and Universal Induction in This Panel Discussion\n          \n          2014"},{"key":"B24","first-page":"2498","article-title":"\u201cVariational dropout sparsifies deep neural networks,\u201d","volume-title":"Proceedings of the 34th international conference on machine learning","author":"Molchanov","year":"2017"},{"key":"B25","doi-asserted-by":"publisher","first-page":"473","DOI":"10.1162\/neco.1992.4.4.473","article-title":"Simplifying Neural networks by soft weight-sharing","volume":"4","author":"Nowlan","year":"1992","journal-title":"Neural Comput"},{"key":"B26","article-title":"Scalable model compression by entropy penalized reparameterization","author":"Oktay","year":"2020","journal-title":"arXiv [preprint"},{"key":"B27","article-title":"Assembly theory reduced to shannon entropy and rendered redundant by naive statistical algorithms","author":"Ozelim","year":"2025","journal-title":"arXiv [preprint"},{"key":"B28","doi-asserted-by":"publisher","first-page":"1080","DOI":"10.1214\/aos\/1176350051","article-title":"Stochastic complexity and modeling","volume":"14","author":"Rissanen","year":"1986","journal-title":"Ann. Stat"},{"key":"B29","doi-asserted-by":"publisher","first-page":"857","DOI":"10.1016\/S0893-6080(96)00127-X","article-title":"Discovering neural nets with low kolmogorov complexity and high generalization capability","volume":"10","author":"Schmidhuber","year":"1997","journal-title":"Neural Netw"},{"key":"B30","doi-asserted-by":"publisher","first-page":"379","DOI":"10.1002\/j.1538-7305.1948.tb01338.x","article-title":"A mathematical theory of communication","volume":"27","author":"Shannon","year":"1948","journal-title":"Bell Syst. Tech. J"},{"key":"B31","article-title":"Outrageously large neural networks: the sparsely-gated mixture-of-experts layer","author":"Shazeer","year":"2017","journal-title":"arXiv [preprint"},{"key":"B32","article-title":"opening the black box of deep neural networks via information","author":"Shwartz-Ziv","year":"2017","journal-title":"arXiv [preprint"},{"key":"B33","doi-asserted-by":"publisher","first-page":"455","DOI":"10.1016\/S0304-3975(01)00102-5","article-title":"The Kolmogorov complexity of real numbers","volume":"284","author":"Staiger","year":"2002","journal-title":"Theor. Comput. Sci"},{"key":"B34","doi-asserted-by":"publisher","first-page":"491","DOI":"10.1109\/CSNT.2014.104","article-title":"\u201cDynamic growth of hidden-layer neurons using the non-extensive entropy,\u201d","author":"Susan","year":"2014","journal-title":"2014 fourth international conference on communication systems and network technologies"},{"key":"B35","doi-asserted-by":"publisher","first-page":"201","DOI":"10.1007\/978-981-13-1135-2_16","article-title":"\u201cNeural net optimization by weight-entropy monitoring,\u201d","author":"Susan","year":"2019","journal-title":"Computational Intelligence: Theories, Applications and Future Directions"},{"key":"B36","unstructured":"07\n          \n            \n              Sutskever\n              I.\n            \n          \n          Berkeley, CA\n          Simons Institute for the Theory of Computing\n          Talk at the Simons Institute: Ilya Sutskever (OpenAI).\n          24\n          2023"},{"key":"B37","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.10652065","article-title":"sztal\/pybdm: v0.1.0 (v0.1.0)","author":"Talaga","year":"2024","journal-title":"Zenodo"},{"key":"B38","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/ITW.2015.7133169","article-title":"\u201cDeep learning and the information bottleneck principle,\u201d","author":"Tishby","year":"2015","journal-title":"2015 IEEE information theory workshop (ITW)"},{"key":"B39","article-title":"Attention is all you need","author":"Vaswani","year":"2017","journal-title":"arXiv [preprint"},{"key":"B40","doi-asserted-by":"publisher","first-page":"185","DOI":"10.1093\/comjnl\/11.2.185","article-title":"An information measure for classification","volume":"11","author":"Wallace","year":"1968","journal-title":"Comput. J"},{"key":"B41","doi-asserted-by":"publisher","first-page":"8","DOI":"10.1109\/MC.1984.1659158","article-title":"A technique for high-performance data compression","volume":"17","year":"1984","journal-title":"Computer"},{"key":"B42","doi-asserted-by":"publisher","first-page":"700","DOI":"10.1109\/JSTSP.2020.2969554","article-title":"DeepCABAC: a universal compression algorithm for deep neural networks","volume":"14","author":"Wiedemann","year":"","journal-title":"IEEE J. Sel. Top. Signal Process"},{"key":"B43","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/IJCNN.2019.8852119","article-title":"\u201cEntropy-constrained training of deep neural networks,\u201d","author":"Wiedemann","year":"2019","journal-title":"2019 international joint conference on neural networks (IJCNN)"},{"key":"B44","doi-asserted-by":"publisher","first-page":"772","DOI":"10.1109\/TNNLS.2019.2910073","article-title":"Compact and computationally efficient representation of deep neural networks","volume":"31","author":"Wiedemann","year":"","journal-title":"IEEE Trans. Neural Netw. Learn. Syst"},{"key":"B45","unstructured":"Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms\n          \n          \n            \n              Xiao\n              H.\n            \n            \n              Rasul\n              K.\n            \n            \n              Vollgraf\n              R.\n            \n          \n          arXiv [preprint\n          \n          2017"},{"key":"B46","doi-asserted-by":"publisher","first-page":"014320","DOI":"10.1103\/PhysRevE.111.014320","article-title":"Evolution beats random chance: performance-dependent network evolution for enhanced computational capacity","volume":"111","author":"Yadav","year":"2025","journal-title":"Phys. Rev. E"},{"key":"B47","doi-asserted-by":"publisher","first-page":"083123","DOI":"10.1063\/5.0273535","article-title":"Task-specific node pruning enhances computational efficiency of reservoir computing networks","volume":"35","author":"Yadav","year":"2025","journal-title":"Chaos"},{"key":"B48","doi-asserted-by":"publisher","first-page":"612","DOI":"10.3390\/e22060612","article-title":"A review of methods for estimating algorithmic complexity: options, challenges, and new directions","volume":"22","author":"Zenil","year":"2020","journal-title":"Entropy"},{"key":"B49","doi-asserted-by":"publisher","first-page":"605","DOI":"10.3390\/e20080605","article-title":"A decomposition method for global evaluation of shannon entropy and local estimations of algorithmic complexity","volume":"20","author":"Zenil","year":"","journal-title":"Entropy"},{"key":"B50","doi-asserted-by":"publisher","first-page":"53143","DOI":"10.4249\/scholarpedia.53143","article-title":"Algorithmic information dynamics","volume":"15","author":"Zenil","year":"2020","journal-title":"Scholarpedia J"},{"key":"B51","doi-asserted-by":"publisher","first-page":"551","DOI":"10.3390\/e20080551","article-title":"A review of graph and network complexity from an algorithmic information perspective","volume":"20","author":"Zenil","year":"","journal-title":"Entropy"},{"key":"B52","doi-asserted-by":"publisher","first-page":"1160","DOI":"10.1016\/j.isci.2019.07.043","article-title":"An algorithmic information calculus for causal discovery and reprogramming systems","volume":"19","author":"Zenil","year":"","journal-title":"iScience"},{"key":"B53","article-title":"Numerical investigation of graph spectra and information interpretability of Eigenvalues","author":"Zenil","year":"","journal-title":"arXiv [preprint"},{"key":"B54","doi-asserted-by":"publisher","first-page":"012308","DOI":"10.1103\/PhysRevE.96.012308","article-title":"Low-algorithmic-complexity entropy-deceiving graphs","volume":"96","author":"Zenil","year":"2017","journal-title":"Phys. Rev. E"},{"key":"B55","doi-asserted-by":"publisher","first-page":"58","DOI":"10.1038\/s42256-018-0005-0","article-title":"Causal deconvolution by algorithmic generative models","volume":"1","author":"Zenil","year":"","journal-title":"Nat. Mach. Intell"},{"key":"B56","doi-asserted-by":"publisher","first-page":"e23","DOI":"10.7717\/peerj-cs.23","article-title":"Two-dimensional Kolmogorov complexity and an empirical validation of the Coding theorem method by compressibility","volume":"1","author":"Zenil","year":"","journal-title":"PeerJ Comput. Sci"},{"key":"B57","doi-asserted-by":"publisher","first-page":"337","DOI":"10.1109\/TIT.1977.1055714","article-title":"A universal algorithm for sequential data compression","volume":"23","author":"Ziv","year":"1977","journal-title":"IEEE Trans. Inf. Theory"}],"container-title":["Frontiers in Computational Neuroscience"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fncom.2026.1791546\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T06:01:27Z","timestamp":1780034487000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fncom.2026.1791546\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,29]]},"references-count":57,"alternative-id":["10.3389\/fncom.2026.1791546"],"URL":"https:\/\/doi.org\/10.3389\/fncom.2026.1791546","relation":{},"ISSN":["1662-5188"],"issn-type":[{"value":"1662-5188","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,29]]},"article-number":"1791546"}}