{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T06:24:35Z","timestamp":1783405475966,"version":"3.54.6"},"reference-count":83,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,10,10]],"date-time":"2024-10-10T00:00:00Z","timestamp":1728518400000},"content-version":"vor","delay-in-days":283,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,10,2]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Recently, large language models (LLMs), especially those that are pretrained on code, have demonstrated strong capabilities in generating programs from natural language inputs. Despite promising results, there is a notable lack of a comprehensive evaluation of these models\u2019 language-to-code generation capabilities. Existing studies often focus on specific tasks, model architectures, or learning paradigms, leading to a fragmented understanding of the overall landscape. In this work, we present L2CEval, a systematic evaluation of the language-to-code generation capabilities of LLMs on 7 tasks across the domain spectrum of semantic parsing, math reasoning, and Python programming, analyzing the factors that potentially affect their performance, such as model size, pretraining data, instruction tuning, and different prompting methods. In addition, we assess confidence calibration, and conduct human evaluations to identify typical failures across different tasks and models. L2CEval offers a comprehensive understanding of the capabilities and limitations of LLMs in language-to-code generation. We release the evaluation framework1 and all model outputs, hoping to lay the groundwork for further future research.<\/jats:p><jats:p>All future evaluations (e.g., LLaMA-3, StarCoder2, etc) will be updated on the project website: https:\/\/l2c-eval.github.io\/.<\/jats:p>","DOI":"10.1162\/tacl_a_00705","type":"journal-article","created":{"date-parts":[[2024,10,10]],"date-time":"2024-10-10T16:24:58Z","timestamp":1728577498000},"page":"1311-1329","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":16,"title":["<tt>L2CEval<\/tt>: Evaluating Language-to-Code Generation Capabilities of Large Language Models"],"prefix":"10.1162","volume":"12","author":[{"given":"Ansong","family":"Ni","sequence":"first","affiliation":[{"name":"Yale University, USA. ansong.ni@yale.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pengcheng","family":"Yin","sequence":"additional","affiliation":[{"name":"Google DeepMind, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yilun","family":"Zhao","sequence":"additional","affiliation":[{"name":"Yale University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Martin","family":"Riddell","sequence":"additional","affiliation":[{"name":"Yale University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Troy","family":"Feng","sequence":"additional","affiliation":[{"name":"Yale University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rui","family":"Shen","sequence":"additional","affiliation":[{"name":"Yale University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stephen","family":"Yin","sequence":"additional","affiliation":[{"name":"Yale University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ye","family":"Liu","sequence":"additional","affiliation":[{"name":"Salesforce Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Semih","family":"Yavuz","sequence":"additional","affiliation":[{"name":"Salesforce Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Caiming","family":"Xiong","sequence":"additional","affiliation":[{"name":"Salesforce Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shafiq","family":"Joty","sequence":"additional","affiliation":[{"name":"Salesforce Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yingbo","family":"Zhou","sequence":"additional","affiliation":[{"name":"Salesforce Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dragomir","family":"Radev","sequence":"additional","affiliation":[{"name":"Yale University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Arman","family":"Cohan","sequence":"additional","affiliation":[{"name":"Yale University, USA. arman.cohan@yale.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Arman","family":"Cohan","sequence":"additional","affiliation":[{"name":"Allen Institute for AI, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,10,2]]},"reference":[{"key":"2024101016244652000_bib1","unstructured":"Replit model. https:\/\/huggingface.co\/replit\/replit-code-v1-3b. Accessed: 2023-09-30."},{"key":"2024101016244652000_bib2","doi-asserted-by":"publisher","first-page":"793","DOI":"10.1007\/s00778-019-00567-8","article-title":"A comparative survey of recent natural language interfaces for databases","volume":"28","author":"Affolter","year":"2019","journal-title":"The VLDB Journal"},{"key":"2024101016244652000_bib3","doi-asserted-by":"publisher","first-page":"5436","DOI":"10.18653\/v1\/D19-154","article-title":"JuICe: A large scale distantly supervised dataset for open domain context-based code generation","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Agashe","year":"2019"},{"key":"2024101016244652000_bib4","article-title":"Santacoder: Don\u2019t reach for the stars!","author":"Allal","year":"2023","journal-title":"arXiv preprint arXiv:2301.03988"},{"key":"2024101016244652000_bib5","article-title":"The falcon series of language models: Towards open frontier models","author":"Almazrouei","year":"2023"},{"key":"2024101016244652000_bib6","doi-asserted-by":"publisher","first-page":"556","DOI":"10.1162\/tacl_a_00333","article-title":"Task-oriented dialogue as dataflow synthesis","volume":"8","author":"Andreas","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"issue":"1","key":"2024101016244652000_bib7","doi-asserted-by":"publisher","first-page":"29","DOI":"10.1017\/S135132490000005X","article-title":"Natural language interfaces to databases\u2013an introduction","volume":"1","author":"Androutsopoulos","year":"1995","journal-title":"Natural Language Engineering"},{"key":"2024101016244652000_bib8","article-title":"Multi-lingual evaluation of code generation models","volume-title":"The Eleventh International Conference on Learning Representations","author":"Athiwaratkun","year":"2023"},{"key":"2024101016244652000_bib9","article-title":"Program synthesis with large language models","author":"Austin","year":"2021","journal-title":"arXiv preprint arXiv:2108.07732"},{"key":"2024101016244652000_bib10","unstructured":"Loubna Ben Allal , NiklasMuennighoff, Logesh KumarUmapathi, BenLipkin, and Leandrovon Werra. 2022. A framework for the evaluation of code generation models. https:\/\/github.com\/bigcode-project\/bigcode-evaluation-harness."},{"key":"2024101016244652000_bib11","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/D13-1160","article-title":"Semantic parsing on Freebase from question-answer pairs","volume-title":"Empirical Methods in Natural Language Processing (EMNLP)","author":"Berant","year":"2013"},{"key":"2024101016244652000_bib12","first-page":"2397","article-title":"Pythia: A suite for analyzing large language models across training and scaling","volume-title":"International Conference on Machine Learning","author":"Biderman","year":"2023"},{"key":"2024101016244652000_bib13","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.bigscience-1.9","article-title":"Gpt-neox-20b: An open- source autoregressive language model","author":"Black","year":"2022","journal-title":"arXiv preprint arXiv:2204.06745"},{"key":"2024101016244652000_bib14","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TSE.2023.3267446","article-title":"Multipl-e: A scalable and polyglot approach to benchmarking neural code generation","author":"Cassano","year":"2023","journal-title":"IEEE Transactions on Software Engineering"},{"key":"2024101016244652000_bib15","article-title":"Evaluating large language models trained on code","author":"Chen","year":"2021","journal-title":"arXiv preprint arXiv:2107.03374"},{"key":"2024101016244652000_bib16","article-title":"Binding language models in symbolic languages","author":"Cheng","journal-title":"ArXiv preprint arXiv:arXiv:2210.02875"},{"issue":"240","key":"2024101016244652000_bib17","first-page":"1","article-title":"Palm: Scaling language modeling with pathways","volume":"24","author":"Chowdhery","year":"2023","journal-title":"Journal of Machine Learning Research"},{"key":"2024101016244652000_bib18","article-title":"Training verifiers to solve math word problems","author":"Cobbe","year":"2021","journal-title":"arXiv preprint arXiv:2110.14168"},{"key":"2024101016244652000_bib19","article-title":"Free dolly: Introducing the world\u2019s first truly open instruction-tuned llm","author":"Conover","year":"2023"},{"key":"2024101016244652000_bib20","article-title":"Incoder: A generative model for code infilling and synthesis","author":"Fried","year":"2023"},{"key":"2024101016244652000_bib21","article-title":"How does gpt obtain its ability? Tracing emergent abilities of language models to their sources","author":"Yao","year":"2022","journal-title":"Yao Fu\u2019s Notion"},{"key":"2024101016244652000_bib22","first-page":"10764","article-title":"Pal: Program-aided language models","volume-title":"International Conference on Machine Learning","author":"Gao","year":"2023"},{"key":"2024101016244652000_bib23","first-page":"1321","article-title":"On calibration of modern neural networks","volume-title":"International Conference on Machine Learning","author":"Guo","year":"2017"},{"key":"2024101016244652000_bib24","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.10742","article-title":"DeepFix: Fixing common C language errors by deep learning","volume-title":"AAAI Conference on Artificial Intelligence","author":"Gupta","year":"2017"},{"key":"2024101016244652000_bib25","article-title":"Folio: Natural language reasoning with first-order logic","author":"Han","year":"2022","journal-title":"arXiv preprint arXiv:2209.00840"},{"key":"2024101016244652000_bib26","article-title":"Measuring coding challenge competence with apps","author":"Hendrycks","year":"2021","journal-title":"arXiv preprint arXiv:2105.09938"},{"key":"2024101016244652000_bib27","article-title":"Training compute-optimal large language models","author":"Hoffmann","year":"2022"},{"key":"2024101016244652000_bib28","first-page":"1769","article-title":"Inner monologue: Embodied reasoning through planning with language models","volume-title":"Conference on Robot Learning","author":"Huang","year":"2023"},{"key":"2024101016244652000_bib29","article-title":"Codesearchnet challenge: Evaluating the state of semantic code search","author":"Husain","year":"2019","journal-title":"ArXiv"},{"key":"2024101016244652000_bib30","doi-asserted-by":"publisher","first-page":"1643","DOI":"10.18653\/v1\/D18-1192","article-title":"Mapping language to code in programmatic context","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Iyer","year":"2018"},{"issue":"12","key":"2024101016244652000_bib31","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3571730","article-title":"Survey of hallucination in natural language generation","volume":"55","author":"Ji","year":"2023","journal-title":"ACM Computing Surveys"},{"key":"2024101016244652000_bib32","unstructured":"Albert Q. Jiang , AlexandreSablayrolles, ArthurMensch, ChrisBamford, Devendra SinghChaplot, Diegode las Casas, FlorianBressand, GiannaLengyel, GuillaumeLample, LucileSaulnier, L\u00e9lio RenardLavaud, Marie-AnneLachaux, PierreStock, Teven LeScao, ThibautLavril, ThomasWang, Timoth\u00e9eLacroix, and William ElSayed. 2023. Mistral 7b."},{"key":"2024101016244652000_bib33","article-title":"Scaling laws for neural language models","author":"Kaplan","year":"2020","journal-title":"arXiv preprint arXiv:2001.08361"},{"key":"2024101016244652000_bib34","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/IWSC55060.2022.00008","article-title":"Bigclonebench considered harmful for machine learning","author":"Krinke","year":"2022","journal-title":"2022 IEEE 16th International Workshop on Software Clones (IWSC)"},{"key":"2024101016244652000_bib35","first-page":"18319","article-title":"Ds-1000: A natural and reliable benchmark for data science code generation","volume-title":"International Conference on Machine Learning","author":"Lai","year":"2023"},{"key":"2024101016244652000_bib36","article-title":"Can llm already serve as a database interface? A big bench for large-scale database grounded text-to-sqls","volume":"36","author":"Li","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024101016244652000_bib37","article-title":"Starcoder: May the source be with you!","author":"Li","year":"2023","journal-title":"arXiv preprint arXiv:2305.06161"},{"key":"2024101016244652000_bib38","article-title":"On the advance of making language models better reasoners","author":"Li","year":"2022","journal-title":"arXiv preprint arXiv:2206.02336"},{"key":"2024101016244652000_bib39","article-title":"Competition-level code generation with alphacode","author":"Li","year":"2022","journal-title":"arXiv preprint arXiv:2203.07814"},{"key":"2024101016244652000_bib40","article-title":"Holistic evaluation of language models","author":"Liang","year":"2022","journal-title":"arXiv preprint arXiv:2211.09110"},{"key":"2024101016244652000_bib41","article-title":"Codexglue: A machine learning benchmark dataset for code understanding and generation","volume-title":"Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1)","author":"Shuai","year":"2023"},{"key":"2024101016244652000_bib42","article-title":"Wizardcoder: Empowering code large language models with evol-instruct","author":"Luo","year":"2023"},{"key":"2024101016244652000_bib43","doi-asserted-by":"publisher","first-page":"5807","DOI":"10.18653\/v1\/2022.emnlp-main.392","article-title":"Lila: A unified benchmark for mathematical reasoning","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Mishra","year":"2022"},{"key":"2024101016244652000_bib44","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v29i1.9602","article-title":"Obtaining well calibrated probabilities using Bayesian binning","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Naeini","year":"2015"},{"key":"2024101016244652000_bib45","article-title":"Learning math reasoning from self-sampled correct and partially-correct solutions","volume-title":"The Eleventh International Conference on Learning Representations","author":"Ni","year":"2023"},{"key":"2024101016244652000_bib46","first-page":"26106","article-title":"Lever: Learning to verify language-to-code generation with execution","volume-title":"International Conference on Machine Learning","author":"Ni","year":"2023"},{"key":"2024101016244652000_bib47","doi-asserted-by":"publisher","first-page":"8536","DOI":"10.1609\/aaai.v34i05.6375","article-title":"Merging weak and active supervision for semantic parsing","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Ni","year":"2020"},{"key":"2024101016244652000_bib48","article-title":"Codegen2: Lessons for training llms on programming and natural languages","author":"Nijkamp","year":"2023"},{"key":"2024101016244652000_bib49","article-title":"A conversational paradigm for program synthesis","author":"Nijkamp","year":"2022","journal-title":"arXiv preprint arXiv:2203.13474"},{"key":"2024101016244652000_bib50","unstructured":"OpenAI. 2022. Chatgpt: Optimizing language models for dialogue."},{"key":"2024101016244652000_bib51","unstructured":"OpenAI. 2023. Gpt-4 technical report."},{"key":"2024101016244652000_bib52","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume":"35","author":"Ouyang","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024101016244652000_bib53","article-title":"Art: Automatic multi-step reasoning and tool-use for large language models","author":"Paranjape","year":"2023","journal-title":"arXiv preprint arXiv:2303.09014"},{"key":"2024101016244652000_bib54","doi-asserted-by":"publisher","first-page":"1470","DOI":"10.3115\/v1\/P15-1142","article-title":"Compositional semantic parsing on semi-structured tables","volume-title":"Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Pasupat","year":"2015"},{"key":"2024101016244652000_bib55","doi-asserted-by":"publisher","first-page":"2080","DOI":"10.18653\/v1\/2021.naacl-main.168","article-title":"Are nlp models really able to solve simple math word problems?","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Patel","year":"2021"},{"key":"2024101016244652000_bib56","article-title":"Evaluating the text-to-sql capabilities of large language models","author":"Rajkumar","year":"2022","journal-title":"arXiv preprint arXiv:2204.00498"},{"key":"2024101016244652000_bib57","article-title":"Code llama: Open foundation models for code","author":"Rozi\u00e8re","year":"2023","journal-title":"arXiv preprint arXiv:2308.12950"},{"key":"2024101016244652000_bib58","first-page":"20601","article-title":"Unsupervised translation of programming languages","volume":"33","author":"Roziere","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024101016244652000_bib59","article-title":"Toolformer: Language models can teach themselves to use tools","volume":"36","author":"Schick","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024101016244652000_bib60","doi-asserted-by":"publisher","first-page":"3533","DOI":"10.18653\/v1\/2022.emnlp-main.231","article-title":"Natural language to code translation with execution","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Shi","year":"2022"},{"key":"2024101016244652000_bib61","doi-asserted-by":"publisher","first-page":"10737","DOI":"10.1109\/CVPR42600.2020.01075","article-title":"Alfred: A benchmark for interpreting grounded instructions for everyday tasks","volume-title":"2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Shridhar","year":"2020"},{"key":"2024101016244652000_bib62","article-title":"Stanford alpaca: An instruction-following llama model","author":"Taori","year":"2023"},{"key":"2024101016244652000_bib63","article-title":"Introducing mpt-30b: Raising the bar for open-source foundation models","author":"MosaicML NLP Team","year":"2023"},{"key":"2024101016244652000_bib64","article-title":"Introducing mpt-7b: A new standard for open-source, commercially usable llms","author":"MosaicML NLP Team","year":"2023"},{"key":"2024101016244652000_bib65","article-title":"Llama: Open and efficient foundation language models","author":"Touvron","year":"2023","journal-title":"arXiv preprint arXiv:2302.13971"},{"key":"2024101016244652000_bib66","article-title":"Llama 2: Open foundation and fine-tuned chat models","author":"Touvron","year":"2023","journal-title":"arXiv preprint arXiv:2307.09288"},{"key":"2024101016244652000_bib67","unstructured":"Ben Wang and AranKomatsuzaki. 2021. GPT-J-6B: A 6 billion parameter autoregressive language model. https:\/\/github.com\/kingoflolz\/mesh-transformer-jax"},{"key":"2024101016244652000_bib68","article-title":"Self-consistency improves chain of thought reasoning in language models","volume-title":"The Eleventh International Conference on Learning Representations","author":"Wang","year":"2023"},{"key":"2024101016244652000_bib69","doi-asserted-by":"publisher","first-page":"1271","DOI":"10.18653\/v1\/2023.findings-emnlp.89","article-title":"Execution-based evaluation for open-domain code generation","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Wang","year":"2023"},{"key":"2024101016244652000_bib70","article-title":"Emergent abilities of large language models","author":"Wei","year":"2022","journal-title":"Transactions on Machine Learning Research"},{"key":"2024101016244652000_bib71","article-title":"Chain of thought prompting elicits reasoning in large language models","author":"Wei","year":"2022","journal-title":"arXiv preprint arXiv:2201.11903"},{"key":"2024101016244652000_bib72","article-title":"Generating sequences by learning to self-correct","volume-title":"The Eleventh International Conference on Learning Representations","author":"Welleck","year":"2023"},{"key":"2024101016244652000_bib73","doi-asserted-by":"publisher","first-page":"602","DOI":"10.18653\/v1\/2022.emnlp-main.39","article-title":"Unifiedskg: Unifying and multi-tasking structured knowledge grounding with text-to-text language models","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Xie","year":"2022"},{"key":"2024101016244652000_bib74","article-title":"Lemur: Harmonizing natural language and code for language agents","volume-title":"The Twelfth International Conference on Learning Representations","author":"Yiheng","year":"2023"},{"key":"2024101016244652000_bib75","doi-asserted-by":"publisher","first-page":"476","DOI":"10.1145\/3196398.3196408","article-title":"Learning to mine aligned code and natural language pairs from stack overflow","author":"Yin","year":"2018","journal-title":"2018 IEEE\/ACM 15th International Conference on Mining Software Repositories (MSR)"},{"key":"2024101016244652000_bib76","doi-asserted-by":"crossref","first-page":"440","DOI":"10.18653\/v1\/P17-1041","article-title":"A syntactic neural model for general-purpose code generation","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Yin","year":"2017"},{"key":"2024101016244652000_bib77","doi-asserted-by":"publisher","first-page":"8413","DOI":"10.18653\/v1\/2020.acl-main.745","article-title":"Tabert: Pretraining for joint understanding of textual and tabular data","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Yin","year":"2020"},{"key":"2024101016244652000_bib78","doi-asserted-by":"publisher","first-page":"3911","DOI":"10.18653\/v1\/D18-1425","article-title":"Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Tao","year":"2018"},{"key":"2024101016244652000_bib79","first-page":"658","article-title":"Learning to map sentences to logical form: Structured classification with probabilistic categorial grammars","volume-title":"UAI \u201905, Proceedings of the 21st Conference in Uncertainty in Artificial Intelligence, Edinburgh, Scotland, July 26\u201329, 2005","author":"Zettlemoyer","year":"2005"},{"key":"2024101016244652000_bib80","first-page":"41832","article-title":"Coder reviewer reranking for code generation","volume-title":"International Conference on Machine Learning","author":"Zhang","year":"2023"},{"key":"2024101016244652000_bib81","article-title":"Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x","volume":"abs\/2303.17568","author":"Zheng","year":"2023","journal-title":"ArXiv"},{"key":"2024101016244652000_bib82","doi-asserted-by":"publisher","first-page":"396","DOI":"10.18653\/v1\/2020.emnlp-main.29","article-title":"Semantic evaluation for text-to-sql with distilled test suites","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Zhong","year":"2020"},{"key":"2024101016244652000_bib83","doi-asserted-by":"publisher","first-page":"67","DOI":"10.18653\/v1\/2022.suki-1.8","article-title":"Hierarchical control of situated agents through natural language","volume-title":"Proceedings of the Workshop on Structured and Unstructured Knowledge Integration (SUKI)","author":"Zhou","year":"2022"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00705\/2474777\/tacl_a_00705.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00705\/2474777\/tacl_a_00705.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,29]],"date-time":"2024-11-29T12:19:43Z","timestamp":1732882783000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00705\/124835\/L2CEval-Evaluating-Language-to-Code-Generation"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":83,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00705","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}