{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T17:33:43Z","timestamp":1778693623131,"version":"3.51.4"},"reference-count":31,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,10,30]],"date-time":"2024-10-30T00:00:00Z","timestamp":1730246400000},"content-version":"vor","delay-in-days":303,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,10,23]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Current Text-to-Code models demonstrate impressive capabilities in generating executable code from natural language snippets. However, current studies focus on technical instructions and programmer-oriented language, and it is an open question whether these models can effectively translate natural language descriptions given by non-technical users and express complex goals, to an executable program that contains an intricate flow\u2014composed of API access and control structures as loops, conditions, and sequences. To unlock the challenge of generating a complete program from a plain non-technical description we present NoviCode, a novel NL Programming task, which takes as input an API and a natural language description by a novice non-programmer, and provides an executable program as output. To assess the efficacy of models on this task, we provide a novel benchmark accompanied by test suites wherein the generated program code is assessed not according to their form, but according to their functional execution. Our experiments show that, first, NoviCode is indeed a challenging task in the code synthesis domain, and that generating complex code from non-technical instructions goes beyond the current Text-to-Code paradigm. Second, we show that a novel approach wherein we align the NL utterances with the compositional hierarchical structure of the code, greatly enhances the performance of LLMs on this task, compared with the end-to-end Text-to-Code counterparts.<\/jats:p>","DOI":"10.1162\/tacl_a_00694","type":"journal-article","created":{"date-parts":[[2024,10,30]],"date-time":"2024-10-30T14:50:10Z","timestamp":1730299810000},"page":"1330-1345","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":2,"title":["NoviCode: Generating Programs from Natural Language Utterances by Novices"],"prefix":"10.1162","volume":"12","author":[{"given":"Asaf Achi","family":"Mordechai","sequence":"first","affiliation":[{"name":"Computer Science Department Bar-Ilan University Ramat-Gan, Israel. asaf.achimordechai@gmail.com"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yoav","family":"Goldberg","sequence":"additional","affiliation":[{"name":"Computer Science Department Bar-Ilan University Ramat-Gan, Israel. yoav.goldberg@gmail.com"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Reut","family":"Tsarfaty","sequence":"additional","affiliation":[{"name":"Computer Science Department Bar-Ilan University Ramat-Gan, Israel. reut.tsarfaty@gmail.com"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2024,10,23]]},"reference":[{"key":"2024103014160154600_bib1","article-title":"Program synthesis with large language models","author":"Austin","year":"2021"},{"key":"2024103014160154600_bib2","article-title":"Studenteval: A benchmark of student-written prompts for large language models of code","author":"Babe","year":"2023"},{"key":"2024103014160154600_bib3","doi-asserted-by":"publisher","first-page":"1415","DOI":"10.3115\/v1\/P14-1133","article-title":"Semantic parsing via paraphrasing","volume-title":"Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Berant","year":"2014"},{"key":"2024103014160154600_bib4","article-title":"Language models are few-shot learners","author":"Brown","year":"2020"},{"key":"2024103014160154600_bib5","article-title":"Evaluating large language models trained on code","author":"Chen","year":"2021","journal-title":"CoRR"},{"key":"2024103014160154600_bib6","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.413","article-title":"Low-resource domain adaptation for compositional task-oriented semantic parsing","author":"Chen","year":"2020","journal-title":"CoRR"},{"key":"2024103014160154600_bib7","article-title":"Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation","author":"Xueying","year":"2023"},{"key":"2024103014160154600_bib8","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1300","article-title":"Semantic parsing for task oriented dialog using hierarchical representations","author":"Gupta","year":"2018","journal-title":"CoRR"},{"issue":"1","key":"2024103014160154600_bib9","doi-asserted-by":"publisher","first-page":"28","DOI":"10.1109\/MC.2008.10","article-title":"Can programming be liberated, period?","volume":"41","author":"Harel","year":"2008","journal-title":"Computer"},{"key":"2024103014160154600_bib10","article-title":"Measuring coding challenge competence with apps","author":"Hendrycks","year":"2021"},{"key":"2024103014160154600_bib11","article-title":"Span-based semantic parsing for compositional generalization","author":"Herzig","year":"2020","journal-title":"CoRR"},{"key":"2024103014160154600_bib12","article-title":"Codesearchnet challenge: Evaluating the state of semantic code search","author":"Husain","year":"2019","journal-title":"CoRR"},{"key":"2024103014160154600_bib13","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1192","article-title":"Mapping language to code in programmatic context","author":"Iyer","year":"2018","journal-title":"CoRR"},{"key":"2024103014160154600_bib14","article-title":"Spoc: Search-based pseudocode to code","volume":"32","author":"Kulal","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024103014160154600_bib15","article-title":"Ds-1000: A natural and reliable benchmark for data science code generation","author":"Lai","year":"2022"},{"issue":"6624","key":"2024103014160154600_bib16","doi-asserted-by":"publisher","first-page":"1092","DOI":"10.1126\/science.abq1158","article-title":"Competition-level code generation with alphacode","volume":"378","author":"Li","year":"2022","journal-title":"Science"},{"key":"2024103014160154600_bib17","article-title":"Is your code generated by chatgpt really correct? Rigorous evaluation of large language models for code generation","author":"Liu","year":"2023"},{"key":"2024103014160154600_bib18","article-title":"Decoupled weight decay regularization","author":"Loshchilov","year":"2019"},{"key":"2024103014160154600_bib19","article-title":"Codexglue: A machine learning benchmark dataset for code understanding and generation","author":"Shuai","year":"2021"},{"key":"2024103014160154600_bib20","article-title":"Codegen2: Lessons for training LLMs on programming and natural languages","author":"Nijkamp","year":"2023","journal-title":"ICLR"},{"key":"2024103014160154600_bib21","doi-asserted-by":"publisher","first-page":"311","DOI":"10.3115\/1073083.1073135","article-title":"Bleu: A method for automatic evaluation of machine translation","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni","year":"2002"},{"key":"2024103014160154600_bib22","doi-asserted-by":"publisher","first-page":"878","DOI":"10.3115\/v1\/P15-1085","article-title":"Language to code: Learning semantic parsers for if-this-then-that recipes","volume-title":"Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Quirk","year":"2015"},{"key":"2024103014160154600_bib23","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","author":"Raffel","year":"2019","journal-title":"CoRR"},{"key":"2024103014160154600_bib24","article-title":"Codebleu: A method for automatic evaluation of code synthesis","author":"Ren","year":"2020","journal-title":"CoRR"},{"key":"2024103014160154600_bib25","doi-asserted-by":"publisher","first-page":"165","DOI":"10.1109\/ICPC.2019.00034","article-title":"Does bleu score work for code migration?","volume-title":"2019 IEEE\/ ACM 27th International Conference on Program Comprehension (ICPC)","author":"Tran","year":"2019"},{"key":"2024103014160154600_bib26","article-title":"Codet5+: Open code large language models for code understanding and generation","author":"Wang","year":"2023","journal-title":"arXiv preprint arXiv:2305.07922"},{"key":"2024103014160154600_bib27","article-title":"Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation","author":"Wang","year":"2021","journal-title":"arXiv preprint arXiv:2109.00859"},{"key":"2024103014160154600_bib28","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.89","article-title":"Execution-based evaluation for open-domain code generation","author":"Wang","year":"2023"},{"key":"2024103014160154600_bib29","doi-asserted-by":"publisher","DOI":"10.1145\/3196398.3196408","article-title":"Learning to mine aligned code and natural language pairs from stack overflow","author":"Yin","year":"2018","journal-title":"CoRR"},{"key":"2024103014160154600_bib30","doi-asserted-by":"crossref","first-page":"440","DOI":"10.18653\/v1\/P17-1041","article-title":"A syntactic neural model for general-purpose code generation","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Yin","year":"2017"},{"key":"2024103014160154600_bib31","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/D18-2002","article-title":"Tranx: A transition-based neural abstract syntax parser for semantic parsing and code generation","author":"Yin","year":"2018"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00694\/2477387\/tacl_a_00694.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00694\/2477387\/tacl_a_00694.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,30]],"date-time":"2024-10-30T14:50:15Z","timestamp":1730299815000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00694\/125031\/NoviCode-Generating-Programs-from-Natural-Language"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":31,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00694","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}