{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T14:13:50Z","timestamp":1784211230530,"version":"3.55.0"},"reference-count":30,"publisher":"Association for Computing Machinery (ACM)","issue":"POPL","license":[{"start":{"date-parts":[[2026,1,8]],"date-time":"2026-01-08T00:00:00Z","timestamp":1767830400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Program. Lang."],"published-print":{"date-parts":[[2026,1,8]]},"abstract":"<jats:p>\n                    We don\u2019t program neural networks directly. Instead, we rely on an indirect style where learning algorithms, like gradient descent, determine a neural network\u2019s function by learning from data. This indirect style is often a virtue; it empowers us to solve problems that were previously impossible. But it lacks discrete structure. We can\u2019t compile most algorithms into a neural network\u2014even if these algorithms could help the network learn. This limitation occurs because discrete algorithms are not obviously differentiable, making them incompatible with the gradient-based learning algorithms that determine a neural network\u2019s function. To address this, we introduce\n                    <jats:styled-content style=\"color:#11559a\">Cajal<\/jats:styled-content>\n                    (\n                    <jats:styled-content style=\"color:#11559a\">\u22b8<\/jats:styled-content>\n                    ,\n                    <jats:styled-content style=\"color:#11559a\">\ud835\udfda<\/jats:styled-content>\n                    ): a typed, higher-order and linear programming language intended to be a minimal vehicle for exploring a direct style of programming neural networks. We prove\n                    <jats:styled-content style=\"color:#11559a\">Cajal<\/jats:styled-content>\n                    (\n                    <jats:styled-content style=\"color:#11559a\">\u22b8<\/jats:styled-content>\n                    ,\n                    <jats:styled-content style=\"color:#11559a\">\ud835\udfda<\/jats:styled-content>\n                    ) programs compile to linear neurons, allowing discrete algorithms to be expressed in a differentiable form compatible with gradient-based learning. With our implementation of\n                    <jats:styled-content style=\"color:#11559a\">Cajal<\/jats:styled-content>\n                    (\n                    <jats:styled-content style=\"color:#11559a\">\u22b8<\/jats:styled-content>\n                    ,\n                    <jats:styled-content style=\"color:#11559a\">\ud835\udfda<\/jats:styled-content>\n                    ), we conduct several experiments where we link these linear neurons against other neural networks to determine part of their function prior to learning. Linking with these neurons allows networks to learn faster, with greater data-efficiency, and in a way that\u2019s easier to debug. A key lesson is that linear programming languages provide a path towards directly programming neural networks, enabling a rich interplay between learning and the discrete structures of ordinary programming.\n                  <\/jats:p>","DOI":"10.1145\/3776677","type":"journal-article","created":{"date-parts":[[2026,1,8]],"date-time":"2026-01-08T18:59:43Z","timestamp":1767898783000},"page":"1010-1035","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Compiling to Linear Neurons"],"prefix":"10.1145","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-6451-5107","authenticated-orcid":false,"given":"Joey","family":"Velez-Ginorio","sequence":"first","affiliation":[{"name":"University of Pennsylvania, Philadelphia, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0830-7248","authenticated-orcid":false,"given":"Nada","family":"Amin","sequence":"additional","affiliation":[{"name":"Harvard University, Cambridge, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8408-4499","authenticated-orcid":false,"given":"Konrad","family":"Kording","sequence":"additional","affiliation":[{"name":"University of Pennsylvania, Philadelphia, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3516-1512","authenticated-orcid":false,"given":"Steve","family":"Zdancewic","sequence":"additional","affiliation":[{"name":"University of Pennsylvania, Philadelphia, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,1,8]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.quant-ph\/0402130"},{"key":"e_1_3_2_3_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et al. 2023. GPT-4 Technical Report. (2023)."},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","unstructured":"Sami Alabed Daniel Belov Bart Chrzaszcz Juliana Franco Dominik Grewe Dougal Maclaurin James Molloy Tom Natan Tamara Norman Xiaoyue Pan Adam Paszke Norman A. Rink Michael Schaarschmidt Timur Sitdikov Agnieszka Swietlik Dimitrios Vytiniotis and Joel Wee. 2025. PartIR: Composing SPMD Partitioning Strategies for Machine Learning. International Conference on Architectural Support for Programming Languages and Operating Systems. doi:10.1145\/3669940.3707284","DOI":"10.1145\/3669940.3707284"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-41026-0"},{"key":"e_1_3_2_6_2","volume-title":"The lambda calculus","author":"Barendregt Hendrik P","year":"1984","unstructured":"Hendrik P Barendregt et al. 1984. The lambda calculus. Vol. 3. North-Holland Amsterdam."},{"key":"e_1_3_2_7_2","volume-title":"Deep Learning","author":"Bengio Yoshua","year":"2017","unstructured":"Yoshua Bengio, Ian Goodfellow, and Aaron Courville. 2017. Deep Learning. Vol. 1. MIT press Cambridge, MA, USA."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","unstructured":"Leonard F. Bereska and Efstratios Gavves. 2024. Mechanistic Interpretability for AI Safety \u2014 A Review. TMLR (April 2024). doi:10.48550\/arXiv.2404.14082","DOI":"10.48550\/arXiv.2404.14082"},{"key":"e_1_3_2_9_2","first-page":"547","volume-title":"International conference on machine learning","author":"Bosnjak Matko","year":"2017","unstructured":"Matko Bosnjak, Tim Rocktaschel, Jason Naradowsky, and Sebastian Riedel. 2017. Programming with a differentiable forth interpreter. In International conference on machine learning. PMLR, 547\u2013556."},{"key":"e_1_3_2_10_2","unstructured":"James Bradbury Roy Frostig Peter Hawkins Matthew James Johnson Chris Leary Dougal Maclaurin George Necula Adam Paszke Jake VanderPlas Skye Wanderman-Milne and Qiao Zhang. 2018. JAX: composable transformations of Python+NumPy programs. http:\/\/github.com\/jax-ml\/jax"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","unstructured":"Kai Fong Ernest Chong. 2020. A closer look at the approximation capabilities of neural networks. International Conference on Learning Representations (2020). doi:10.48550\/arXiv.2002.06505","DOI":"10.48550\/arXiv.2002.06505"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF02551274"},{"key":"e_1_3_2_13_2","first-page":"1213","volume-title":"In International Conference on Machine Learning","author":"Gaunt Alexander L","year":"2017","unstructured":"Alexander L Gaunt, Marc Brockschmidt, Nate Kushman, and Daniel Tarlow. 2017. Differentiable programs with neural libraries. In International Conference on Machine Learning. PMLR, 1213-1222."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","unstructured":"David Ha Andrew Dai and Quoc V Le. 2016. Hypernetworks. arXiv preprint arXiv:1609.09106 (2016). doi:10.48550\/arXiv.1609.09106","DOI":"10.48550\/arXiv.1609.09106"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2015. Delving deep into rectifiers: Surpassing humanlevel performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision. 1026-1034. doi:10.1109\/ICCV.2015.123","DOI":"10.1109\/ICCV.2015.123"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_3_2_18_2","unstructured":"Yann LeCun Corinna Cortes and C. J. Burges. 2010. MNIST handwritten digit database. ATT Labs [Online]. Available: http:\/\/yann.lecun.com\/exdb\/mnist 2 (2010)."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3591280"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","unstructured":"Nicholas D Matsakis and Felix S Klock. 2014. The rust language. In Proceedings of the 2014 ACM SIGAda annual conference on High integrity language technology. 103-104. doi:10.1145\/2663171.2663188","DOI":"10.1145\/2663171.2663188"},{"key":"e_1_3_2_21_2","first-page":"15","article-title":"Categorical semantics of linear logic","volume":"27","author":"Mellies Paul-Andr\u00e9","year":"2009","unstructured":"Paul-Andr\u00e9 Mellies. 2009. Categorical semantics of linear logic. Panoramas et syntheses 27 (2009), 15-215.","journal-title":"Panoramas et syntheses"},{"key":"e_1_3_2_22_2","unstructured":"Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in PyTorch. (2017)."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3341689"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.5555\/509043"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1017\/9781009023405"},{"key":"e_1_3_2_26_2","volume-title":"Deep linear neural networks: A theory of learning in the brain and mind","author":"Saxe Andrew Michael","year":"2015","unstructured":"Andrew Michael Saxe. 2015. Deep linear neural networks: A theory of learning in the brain and mind. Stanford University."},{"key":"e_1_3_2_27_2","article-title":"Hidden technical debt in machine learning systems","volume":"28","author":"Sculley David","year":"2015","unstructured":"David Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-Francois Crespo, and Dan Dennison. 2015. Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems 28 (2015).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10882-7_26"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10882-7_26"},{"key":"e_1_3_2_30_2","first-page":"66","article-title":"Dimensionality reduction: A comparative review","volume":"10","author":"Van Der Maaten Laurens","year":"2009","unstructured":"Laurens Van Der Maaten, Eric O Postma, H Jaap Van Den Herik, et al. 2009. Dimensionality reduction: A comparative review. Journal of machine learning research 10, 66-71 (2009), 13.","journal-title":"Journal of machine learning research"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","unstructured":"David Walker. 2005. Substructural type systems. Advanced Topics in Types and Programming Languages (2005) 3-44.","DOI":"10.7551\/mitpress\/1104.003.0003"}],"container-title":["Proceedings of the ACM on Programming Languages"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3776677","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T13:43:38Z","timestamp":1784209418000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3776677"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,8]]},"references-count":30,"journal-issue":{"issue":"POPL","published-print":{"date-parts":[[2026,1,8]]}},"alternative-id":["10.1145\/3776677"],"URL":"https:\/\/doi.org\/10.1145\/3776677","relation":{},"ISSN":["2475-1421"],"issn-type":[{"value":"2475-1421","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,8]]},"assertion":[{"value":"2025-07-10","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-06","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}