{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T02:42:53Z","timestamp":1784342573279,"version":"3.55.0"},"reference-count":73,"publisher":"Association for Computing Machinery (ACM)","issue":"PLDI","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Program. Lang."],"published-print":{"date-parts":[[2025,6,10]]},"abstract":"<jats:p>Large language models (LLMs) have achieved notable success in code generation. However, they still frequently produce uncompilable output because their next-token inference procedure does not model formal aspects of code. Although constrained decoding is a promising approach to alleviate this issue, it has only been applied to handle either domain-specific languages or syntactic features of general-purpose programming languages. However, LLMs frequently generate code with typing errors, which are beyond the domain of syntax and generally hard to adequately constrain. To address this challenge, we introduce a type-constrained decoding approach that leverages type systems to guide code generation. For this purpose, we develop novel prefix automata and a search over inhabitable types, forming a sound approach to enforce well-typedness on LLM-generated code. We formalize our approach on a foundational simply-typed language and extend it to TypeScript to demonstrate practicality. Our evaluation on the HumanEval and MBPP datasets shows that our approach reduces compilation errors by more than half and significantly increases functional correctness in code synthesis, translation, and repair tasks across LLMs of various sizes and model families, including state-of-the-art open-weight models with more than 30B parameters. The results demonstrate the generality and effectiveness of our approach in constraining LLM code generation with formal rules of type systems.<\/jats:p>","DOI":"10.1145\/3729274","type":"journal-article","created":{"date-parts":[[2025,6,13]],"date-time":"2025-06-13T16:02:27Z","timestamp":1749830547000},"page":"601-626","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Type-Constrained Code Generation with Language Models"],"prefix":"10.1145","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3851-2557","authenticated-orcid":false,"given":"Niels","family":"M\u00fcndler","sequence":"first","affiliation":[{"name":"ETH Zurich, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-9858-5640","authenticated-orcid":false,"given":"Jingxuan","family":"He","sequence":"additional","affiliation":[{"name":"University of California at Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8480-8167","authenticated-orcid":false,"given":"Hao","family":"Wang","sequence":"additional","affiliation":[{"name":"University of California at Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4539-9188","authenticated-orcid":false,"given":"Koushik","family":"Sen","sequence":"additional","affiliation":[{"name":"University of California at Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9745-6802","authenticated-orcid":false,"given":"Dawn","family":"Song","sequence":"additional","affiliation":[{"name":"University of California at Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0054-9568","authenticated-orcid":false,"given":"Martin","family":"Vechev","sequence":"additional","affiliation":[{"name":"ETH Zurich, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,13]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","unstructured":"LakshyaAgrawal AdityaKanade NavinGoyal Shuvendu KLahiri and SriramRajamani. 2023.Monitor-Guided Decoding of Code LMs with Static Analysis of Repository Context. In NeurIPS.https:\/\/openreview.net\/forum?id=qPUbKxKvXq","DOI":"10.52202\/075280-1401"},{"key":"e_1_3_2_3_2","unstructured":"Anthropic.[n. d.].Claude 3 Model Card.https:\/\/assets.anthropic.com\/m\/61e7d27f8c8f5919\/original\/Claude-3-Model-Card.pdf Accessed: March 10 2025."},{"key":"e_1_3_2_4_2","unstructured":"Anthropic. 2025. JSON Mode https:\/\/docs.anthropic.eom\/en\/docs\/build-with-claude\/tool-use#json-mode Accessed: March 10 2025."},{"key":"e_1_3_2_5_2","unstructured":"Ken Arnold and James Gosling. 1996. The Java Programming Language.."},{"key":"e_1_3_2_6_2","unstructured":"JacobAustin AugustusOdena Maxwell I.Nye MaartenBosma HenrykMichalewski DavidDohan EllenJiang Carrie J.Cai MichaelTerry Quoc V.Le et al. 2021. Program Synthesis with Large Language Models. arXiv Preprint https:\/\/arxiv.org\/abs\/2108.07732"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","unstructured":"NastaranBassamzadeh and ChhayaMethani. 2024. A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation. arXiv Preprint https:\/\/doi.org\/10.48550\/arXiv.2407.02742 10.48550\/arXiv.2407.02742","DOI":"10.48550\/arXiv.2407.02742"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","unstructured":"LucaBeurer-Kellner MarcFischer and MartinVechev. 2023. Prompting Is Programming: A Query Language for Large Language Models. In PLDI (2023). https:\/\/doi.org\/10.1145\/3591300 10.1145\/3591300","DOI":"10.1145\/3591300"},{"key":"e_1_3_2_9_2","unstructured":"LucaBeurer-Kellner MarcFischer and MartinVechev. 2024. Guiding LLMs The Right Way: Fast Non-Invasive Constrained Generation. In ICML. https:\/\/openreview.net\/forum?id=pXaEYzrFae"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","unstructured":"SatwikBhattamishra KabirAhuja and NavinGoyal. 2020. On the Ability and Limitations of Transformers to Recognize Formal Languages. In EMNLP. https:\/\/doi.org\/10.18653\/vl\/2020.emnlp-main.576 10.18653\/vl\/2020.emnlp-main.576","DOI":"10.18653\/vl\/2020.emnlp-main.576"},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Gavin M.Bierman MartinAbadi and MadsTorgersen. 2014. Understanding TypeScript. In ECOOP.","DOI":"10.1007\/978-3-662-44202-9_11"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","unstructured":"AndrewBlinn XiangLi June HyungKim and CyrusOmar. 2024. Statically Contextualizing Large Language Models with Typed Holes. OOPSLA (2024). https:\/\/doi.org\/10.1145\/3689728 10.1145\/3689728","DOI":"10.1145\/3689728"},{"key":"e_1_3_2_13_2","unstructured":"Tom B.Brown BenjaminMann NickRyder MelanieSubbiah JaredKaplan PrafullaDhariwal ArvindNeelakantan PranavShyam GirishSastry AmandaAskell et al.. 2020. Language Models are Few-Shot Learners. In NeurIPS. https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html"},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","unstructured":"FedericoCassano JohnGouwar DanielNguyen SydneyNguyen LunaPhipps-Costin DonaldPinckney Ming-HoYee YangtianZi Carolyn JaneAnderson Molly Q.Feldman. et al.2023. MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation. IEEE Trans. Software Eng. (2023).","DOI":"10.1109\/TSE.2023.3267446"},{"key":"e_1_3_2_15_2","unstructured":"MarkChen JerryTworek HeewooJun QimingYuan Henrique Pond\u00e9de Oliveira Pinto JaredKaplan HarriEdwards YuriBurda NicholasJoseph GregBrockman et al.2021. Evaluating Large Language Models Trained on Code. arXiv Preprint (2021). https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","unstructured":"DeepSeek-Al DayaGuo DejianYang HaoweiZhang JunxiaoSong RuoyuZhang RunxinXu QihaoZhu ShirongMa PeiyiWang et al. 2025. DeepSeek-Rl: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv Preprint (2025). https:\/\/doi.org\/10.48550\/arXiv.2501.12948 10.48550\/arXiv.2501.12948","DOI":"10.48550\/arXiv.2501.12948"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"PantazisDeligiannis AkashLal NikitaMehrotra RishiPoddar and AseemRastogi. 2025. RustAssistant: Using LLMs to Fix Compilation Errors in Rust Code. In ICSE. https:\/\/www.microsoft.com\/en-us\/research\/publication\/rustassistant-using-llms-to-fix-compilation-errors-in-rust-code\/","DOI":"10.1109\/ICSE55347.2025.00022"},{"key":"e_1_3_2_18_2","unstructured":"TypeScript Developers. [n. d.]. TypeScript: Documentation - More on Functions https:\/\/www.typescriptlang.org\/docs\/handbook\/2\/functions.html#function-type-expressions Accessed: March 10 2025."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","unstructured":"YixinDong Charlie F.Ruan YaxingCai RuihangLai ZiyiXu YilongZhao and TianqiChen. 2024. XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models. arXiv Preprint (2024). https:\/\/doi.org\/10.48550\/arXiv.2411.15100 10.48550\/arXiv.2411.15100","DOI":"10.48550\/arXiv.2411.15100"},{"key":"e_1_3_2_20_2","unstructured":"Alan AADonovan and Brian WKernighan. 2015. The Go programming language."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","unstructured":"ShihanDou HaoxiangJia ShenxiWu HuiyuanZheng WeikangZhou MulingWu MingxuChai JessicaFan CaishuangHuang YunboTao et al. 2024. What\u2019s Wrong with Your Code Generated by Large Language Models? An Extensive Study. arXiv Preprint (2024). https:\/\/doi.org\/10.48550\/arXiv.2407.06153 10.48550\/arXiv.2407.06153","DOI":"10.48550\/arXiv.2407.06153"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"JavidEbrahimi DhruvGelda and WeiZhang. 2020. How Can Self-Attention Networks Recognize Dyck-n Languages?. In EMNLP. https:\/\/aclanthology.org\/2020.findings-emnlp.384\/","DOI":"10.18653\/v1\/2020.findings-emnlp.384"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","unstructured":"Jon\u00e2sFiala ShacharItzhaky PeterM\u00fcller NadiaPolikarpova and IlyaSergey. 2023. Leveraging Rust Types for Program Synthesis. PLDI (2023). https:\/\/doi.org\/10.1145\/3591278 10.1145\/3591278","DOI":"10.1145\/3591278"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","unstructured":"ZhengGao ChristianBird and Earl T.Barr. 2017. To type or not to type: quantifying detectable bugs in JavaScript. In ICSE. https:\/\/doi.org\/10.1109\/ICSE.2017.75 10.1109\/ICSE.2017.75","DOI":"10.1109\/ICSE.2017.75"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","unstructured":"AlessandroGiagnorio AlbertoMartin-Lopez and GabrieleBavota. 2025. Enhancing Code Generation for Low-Resource Languages: No Silver Bullet. arXiv Preprint (2025). https:\/\/doi.org\/10.48550\/arXiv.2501.19085 10.48550\/arXiv.2501.19085","DOI":"10.48550\/arXiv.2501.19085"},{"key":"e_1_3_2_26_2","unstructured":"GitHub. [n. d.]. https:\/\/github.com\/features\/copilot"},{"key":"e_1_3_2_27_2","unstructured":"GitHub. 2022. The top programming languages https:\/\/octoverse.github.com\/2022\/top-programming-languages"},{"key":"e_1_3_2_28_2","unstructured":"AaronGrattafiori AbhimanyuDubey AbhinavJauhri AbhinavPandey AbhishekRadian AhmadAl-Dahle AieshaLetman AkhilMathur AlanSchelten AlexVaughan et al. 2024. The Llama 3 Herd of Models. ArXiv Preprint (2024). https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","unstructured":"DayaGuo QihaoZhu DejianYang ZhendaXie KaiDong WentaoZhang GuantingChen XiaoBi Y.Wu Y. K.Liet al. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence. arXiv Preprint (2024). https:\/\/doi.org\/10.48550\/arXiv.2401.14196 10.48550\/arXiv.2401.14196","DOI":"10.48550\/arXiv.2401.14196"},{"key":"e_1_3_2_30_2","unstructured":"Gusanidas.[n. d.]. Compilation Benchmark. https:\/\/github.com\/Gusanidas\/compilation-benchmark Accessed: March 10 2025."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","unstructured":"TihomirGvero ViktorKuncak IvanKuraj and Ru\u017eicaPiskac. 2013. Complete completion using types and weights. In PLDI. https:\/\/doi.org\/10.1145\/2491956.2462192 10.1145\/2491956.2462192","DOI":"10.1145\/2491956.2462192"},{"key":"e_1_3_2_32_2","unstructured":"John E.Hopcroft and Jeffrey D.Ullman. 1979. Introduction to Automata Theory Languages and Computation."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","unstructured":"LeiHuang WeijiangYu WeitaoMa WeihongZhong ZhangyinFeng HaotianWang QianglongChen WeihuaPeng XiaochengFeng BingQinet al. 2023. A Survey on Hallucination in Large Language Models: Principles Taxonomy Challenges and Open Questions. arXiv Preprint (2023). https:\/\/doi.org\/10.48550\/arXiv.2311.05232 10.48550\/arXiv.2311.05232","DOI":"10.48550\/arXiv.2311.05232"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","unstructured":"AaronJaech AdamKalai AdamLerer AdamRichardson AhmedEl-Kishky AidenLow AlecHelyar AleksanderMadry AlexBeutel AlexCarney et al. 2024. OpenAI o1 System Card. arXiv Preprint (2024). https:\/\/doi.org\/10.48550\/arXiv.2412.16720 10.48550\/arXiv.2412.16720","DOI":"10.48550\/arXiv.2412.16720"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","unstructured":"PrithwishJana PiyushJha HaoyangJu GauthamKishore AryanMahajan and VijayGanesh. 2024. CoTran: An LLM-Based Code Translator Using Reinforcement Learning with Feedback from Compiler and Symbolic Execution. In ECAI (Frontiers in Artificial Intelligence and Applications). https:\/\/doi.org\/10.3233\/FAIA240968 10.3233\/FAIA240968","DOI":"10.3233\/FAIA240968"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","unstructured":"JuyongJiang FanWang JiasiShen SungjuKim and SunghunKim. 2024. A Survey on Large Language Models for Code Generation. arXiv Preprint (2024). https:\/\/doi.org\/10.48550\/arXiv.2406.00515 10.48550\/arXiv.2406.00515","DOI":"10.48550\/arXiv.2406.00515"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","unstructured":"SathvikJoel Jie JWWu and Fatemeh H.Fard. 2024. Survey on Code Generation for Low resource and Domain Specific Programming Languages. arXiv Preprint (2024). https:\/\/doi.org\/10.48550\/arXiv.2410.03981 10.48550\/arXiv.2410.03981","DOI":"10.48550\/arXiv.2410.03981"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","unstructured":"AntonLozhkov RaymondLi LoubnaBen Allai FedericoCassano JoelLamy-Poirier NouamaneTazi AoTang DmytroPykhtar JiaweiLiu YuxiangWei et al. 2024. StarCoder 2 and The Stack v2: The Next Generation. arXiv Preprint (2024). https:\/\/doi.org\/10.48550\/arXiv.2402.19173 10.48550\/arXiv.2402.19173","DOI":"10.48550\/arXiv.2402.19173"},{"key":"e_1_3_2_39_2","unstructured":"Madnight. 2024. GitHut 2.0. https:\/\/madnight.github.io\/githut\/#\/pull_requests\/2024\/l"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","unstructured":"Harry G.Mairson. 2004. Linear lambda calculus and PTIME-completeness. J. Fund. Program. (2004). https:\/\/doi.org\/10.1017\/S0956796804005131 10.1017\/S0956796804005131","DOI":"10.1017\/S0956796804005131"},{"key":"e_1_3_2_41_2","doi-asserted-by":"crossref","unstructured":"Nicholas DMatsakis and Felix SKlock. 2014. The rust language. IACM SIGAda Ada Letters.(2014)","DOI":"10.1145\/2692956.2663188"},{"key":"e_1_3_2_42_2","unstructured":"DanielMelcer NathanFulton SanjayKrishna Gouda and HaifengQian. 2024. Constrained Decoding for Fill-in-the-Middle Code Language Models via Efficient Left and Right Quotienting of Context-Sensitive Grammars. arXiv Preprint (2024). https:\/\/arxiv.org\/abs\/2402.17988"},{"key":"e_1_3_2_43_2","unstructured":"Microsoft. 2024. TypeScript. https:\/\/github.com\/microsoft\/TypeScript Accessed on November 9 2024 commit #ef802bl."},{"key":"e_1_3_2_44_2","doi-asserted-by":"crossref","unstructured":"John C.MITCHELL. 1990. Type Systems for Programming Languages. In Formal Models and Semantics https:\/\/www.sciencedirect.com\/science\/article\/pii\/B9780444880741500135","DOI":"10.1016\/B978-0-444-88074-1.50013-5"},{"key":"e_1_3_2_45_2","unstructured":"NiklasMuennighoff QianLiu Armel RandyZebaze QinkaiZheng BinyuanHui Terry YueZhuo SwayamSingh XiangruTang Leandro vonWerra and ShayneLongpre. 2024. OctoPack: Instruction Tuning Code Large Language Models. In ICLR. https:\/\/openreview.net\/forum?id=mwlPWNSWZP"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","unstructured":"NielsM\u00fcnder JingxuanHe HaoWang KoushikSen DawnSong and MartinVechev. 2025. Reproduction Package for \u2033Type-Constrained Code Generation with Language Models\u2033. https:\/\/doi.org\/10.5281\/zenodo.l5355889 10.5281\/zenodo.l5355889","DOI":"10.5281\/zenodo.l5355889"},{"key":"e_1_3_2_47_2","unstructured":"NielsM\u00fcnder Mark NiklasM\u00fcller JingxuanHe and MartinVechev. 2024. SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents. In NeurIPS. http:\/\/papers.nips.ee\/paper_files\/paper\/2024\/hash\/94f093b41fc2666376fblf667fe282f3-Abstract-Conference.html"},{"key":"e_1_3_2_48_2","unstructured":"nielstron. 2024. Incorrect type deducted for accumulator in reduce. https:\/\/github.com\/microsoft\/TypeScript\/issues\/59999"},{"key":"e_1_3_2_49_2","unstructured":"nop33. 2024. Wrong inferred initial value in reduce. https:\/\/github.com\/microsoft\/TypeScript\/issues\/59863"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","unstructured":"OpenAI. 2023. GPT-4 Technical Report. arXiv Preprint (2023). https:\/\/doi.org\/10.48550\/arXiv.2303.08774 10.48550\/arXiv.2303.08774","DOI":"10.48550\/arXiv.2303.08774"},{"key":"e_1_3_2_51_2","unstructured":"OpenAI. 2025. Structured Outputs https:\/\/platform.openai.com\/docs\/guides\/structured-outputs Accessed: March 10 2025."},{"key":"e_1_3_2_52_2","unstructured":"GabrielOrlanski KefanXiao XavierGarcia JeffreyHui JoshuaHowland JonathanMalmaud JacobAustin RishabhSingh and MicheleCatasta. 2023. Measuring the Impact of Programming Language Distribution. In ICML. https:\/\/proceedings.mlr.press\/v202\/orlanski23a.html"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","unstructured":"RangeetPan Ali RezaIbrahimzada RahulKrishna DivyaSankar Lambert PouguemWassi MicheleMerler BorisSobolev RajuPavuluri SaurabhSinha and ReyhanehJabbarvand. 2024. Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code. In ICSE. https:\/\/doi.org\/10.1145\/3597503.3639226 10.1145\/3597503.3639226","DOI":"10.1145\/3597503.3639226"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","unstructured":"RangeetPan Ali RezaIbrahimzada RahulKrishna DivyaSankar Lambert PouguemWassi MicheleMerler BorisSobolev RajuPavuluri SaurabhSinha and ReyhanehJabbarvand. 2024. Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code. In ICSE. https:\/\/doi.org\/10.1145\/3597503.3639226 10.1145\/3597503.3639226","DOI":"10.1145\/3597503.3639226"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","unstructured":"HammondPearce BaleeghAhmad BenjaminTan BrendanDolan-Gavitt and RameshKarri. 2022. Asleep at the Keyboard? Assessing the Security of GitHub Copilot\u2019s Code Contributions. In S&P. https:\/\/doi.org\/10.1109\/SP46214.2022.9833571 10.1109\/SP46214.2022.9833571","DOI":"10.1109\/SP46214.2022.9833571"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","unstructured":"DanielPerelman SumitGulwani ThomasBall and DanGrossman. 2012. Type-directed completion of partial expressions. In PLDI. https:\/\/doi.org\/10.1145\/2254064.2254098 10.1145\/2254064.2254098","DOI":"10.1145\/2254064.2254098"},{"key":"e_1_3_2_57_2","unstructured":"GabrielPoesia AlexPolozov VuLe AshishTiwari GustavoSoares ChristopherMeek and SumitGulwani. 2022. Synchromesh: Reliable Code Generation from Pre-trained Language Models. In ICLR. https:\/\/openreview.net\/forum?id=KmtVD97J43e"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","unstructured":"VipulaRawte Amit P.Sheth and AmitavaDas. 2023. A Survey of Hallucination in Large Foundation Models. arXiv Preprint (2023). https:\/\/doi.org\/10.48550\/arXiv.2309.05922 10.48550\/arXiv.2309.05922","DOI":"10.48550\/arXiv.2309.05922"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","unstructured":"BaptisteRozi\u00e8re JonasGehring FabianGloeckle StenSootla ItaiGat Xiaoqing EllenTan YossiAdi JingyuLiu TalRemez J\u00e9r\u00e9myRapin et al. 2023. Code Llama: Open Foundation Models for Code. arXiv Preprint (2023). https:\/\/doi.org\/10.48550\/arXiv.2308.12950 10.48550\/arXiv.2308.12950","DOI":"10.48550\/arXiv.2308.12950"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","unstructured":"RicoSennrich BarryHaddow and AlexandraBirch. 2016. Neural Machine Translation of Rare Words with Subword Units. In ACL. https:\/\/doi.org\/10.18653\/vl\/pl6-1162 10.18653\/vl\/pl6-1162","DOI":"10.18653\/vl\/pl6-1162"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","unstructured":"ManishShetty NamanJain AdwaitGodbole Sanjit A.Seshia and KoushikSen. 2024. Syzygy: Dual Code-Test C to (safe) Rust Translation using LLMs and Dynamic Analysis. arXiv Preprint (2024). https:\/\/doi.org\/10.48550\/arXiv.2412.14234 10.48550\/arXiv.2412.14234","DOI":"10.48550\/arXiv.2412.14234"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","unstructured":"VinceSzabo DominikWinterer and ZhendongSu. 2024. Compilation Quotient (CQ): A Metric for the Compilation Hardness of Programming Languages. arXiv Preprint (2024). https:\/\/doi.org\/10.48550\/arXiv.2406.04778 10.48550\/arXiv.2406.04778","DOI":"10.48550\/arXiv.2406.04778"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1007\/sl0664-025-10614-4"},{"key":"e_1_3_2_64_2","unstructured":"Gemma Team MorganeRiviere ShreyaPathak Pier GiuseppeSessa CassidyHardin SuryaBhupatiraju L\u00e9onardHussenot ThomasMesnard BobakShahriari AlexandreRam\u00e9et al. 2024. Gemma 2: Improving Open Language Models at a Practical Size. arXiv Preprint (2024). https:\/\/arxiv.org\/abs\/2408.00118"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","unstructured":"Yun-DaTsai MingjieLiu and HaoxingRen. 2024. Code Less Align More: Efficient LLM Fine-tuning for Code Generation with Data Pruning. (2024). https:\/\/doi.org\/10.48550\/arXiv.2407.05040 10.48550\/arXiv.2407.05040","DOI":"10.48550\/arXiv.2407.05040"},{"key":"e_1_3_2_66_2","unstructured":"ShubhamUgare TarunSuresh HangooKang SasaMisailovic and GagandeepSingh. 2024. SynCode: LLM Generation with Grammar Augmentation. ArXiv Preprint (2024). https:\/\/arxiv.org\/abs\/2403.01632"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","unstructured":"PawelUrzyczyn. 1997. Inhabitation in Typed Lambda-Calculi (A Syntactic Approach). In TLCA (Lecture Notes in Computer Science) https:\/\/doi.org\/10.1007\/3-540-62688-3_47 10.1007\/3-540-62688-3_47","DOI":"10.1007\/3-540-62688-3_47"},{"key":"e_1_3_2_68_2","unstructured":"HeidiVella. 2024. Google turns to AI to write new code; Workforce reduced https:\/\/aibusiness.com\/data\/google-turns-to-ai-to-write-new-code-workforce-reduced"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","unstructured":"YuxiangWei Chunqiu StevenXia and LingmingZhang. 2023. Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program Repair. In ESEC\/FSE. https:\/\/doi.org\/10.1145\/3611643.3616271 10.1145\/3611643.3616271","DOI":"10.1145\/3611643.3616271"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","unstructured":"MartinWeyssow XinZhou KisubKim DavidLo and Houari A.Sahraoui. 2023. Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models. arXiv Preprint (2023). https:\/\/doi.org\/10.48550\/arXiv.2308.10462 10.48550\/arXiv.2308.10462","DOI":"10.48550\/arXiv.2308.10462"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","unstructured":"Brandon T.Willard and R\u00e9miLouf. 2023. Efficient Guided Generation for Large Language Models. arXiv Preprint (2023). https:\/\/doi.org\/10.48550\/arXiv.2307.09702 10.48550\/arXiv.2307.09702","DOI":"10.48550\/arXiv.2307.09702"},{"key":"e_1_3_2_72_2","unstructured":"AndyYang DavidChiang and DanaAngluin. 2024. Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages. In NeurIPS. http:\/\/papers.nips.cc\/paper_files\/paper\/2024\/hash\/13d7f172259b11b230cc5da8768abc5f-Abstract-Conference.html"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","unstructured":"AnYang BaosongYang BeichenZhang BinyuanHui BoZheng BowenYu ChengyuanLi DayihengLiu FeiHuang HaoranWeiet al. 2024. Qwen2.5 Technical Report. arXiv Preprint (2024). https:\/\/doi.org\/10.48550\/arXiv.2412.15115 10.48550\/arXiv.2412.15115","DOI":"10.48550\/arXiv.2412.15115"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","unstructured":"QuanjunZhang ChunrongFang YangXie YuxiangMa WeisongSun YunYang and ZhenyuChen. 2024. A Systematic Literature Review on Large Language Models for Automated Program Repair. arXiv Preprint (2024). https:\/\/doi.org\/10.48550\/arXiv.2405.01466 10.48550\/arXiv.2405.01466","DOI":"10.48550\/arXiv.2405.01466"}],"container-title":["Proceedings of the ACM on Programming Languages"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729274","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T10:08:32Z","timestamp":1784196512000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3729274"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,10]]},"references-count":73,"journal-issue":{"issue":"PLDI","published-print":{"date-parts":[[2025,6,10]]}},"alternative-id":["10.1145\/3729274"],"URL":"https:\/\/doi.org\/10.1145\/3729274","relation":{},"ISSN":["2475-1421"],"issn-type":[{"value":"2475-1421","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,10]]},"assertion":[{"value":"2024-11-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-06","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}