{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T03:16:02Z","timestamp":1785381362931,"version":"3.55.0"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"abstract":"<jats:p>\n            Recent code large language models (LLMs) have shown promising performance in generating standalone functions. However, they face limitations in repository-level code generation due to their lack of awareness of\n            <jats:italic>repository-level dependencies<\/jats:italic>\n            (\n            <jats:italic>e.g.,<\/jats:italic>\n            user-defined attributes), resulting in\n            <jats:italic>dependency errors<\/jats:italic>\n            such as undefined-variable and no-member errors. In this work, we introduce\n            <jats:sc>ToolGen<\/jats:sc>\n            , an approach that integrates autocompletion tools into the code LLM generation process to address these dependencies.\n            <jats:sc>ToolGen<\/jats:sc>\n            comprises two main phases: Trigger Insertion and Model Fine-tuning (Offline), and Tool-integrated Code Generation (Online). During the offline phase,\n            <jats:sc>ToolGen<\/jats:sc>\n            augments functions within a given code corpus with a special mark token, indicating positions to trigger autocompletion tools. These augmented functions, along with their corresponding descriptions, are then used to fine-tune a selected code LLM. In the online phase,\n            <jats:sc>ToolGen<\/jats:sc>\n            iteratively generates functions by predicting tokens step-by-step using the fine-tuned LLM. Whenever a mark token is encountered,\n            <jats:sc>ToolGen<\/jats:sc>\n            invokes the autocompletion tool to suggest code completions and selects the most appropriate one through constrained greedy search.\n          <\/jats:p>\n          <jats:p>\n            We conduct comprehensive experiments to evaluate\n            <jats:sc>ToolGen<\/jats:sc>\n            \u2019s effectiveness in repository-level code generation across three distinct code LLMs: CodeGPT, CodeT5, and CodeLlama. To facilitate this evaluation, we create a benchmark comprising 671 real-world code repositories and introduce two new dependency-based metrics:\n            <jats:italic>Dependency Coverage<\/jats:italic>\n            and\n            <jats:italic>Static Validity Rate<\/jats:italic>\n            . The results demonstrate that\n            <jats:sc>ToolGen<\/jats:sc>\n            significantly improves\n            <jats:italic>Dependency Coverage<\/jats:italic>\n            by 31.4% to 39.1% and\n            <jats:italic>Static Validity Rate<\/jats:italic>\n            by 44.9% to 57.7% across the three LLMs, while maintaining competitive or improved performance in widely recognized similarity metrics such as BLEU-4, CodeBLEU, Edit Similarity, and Exact Match. On the CoderEval dataset,\n            <jats:sc>ToolGen<\/jats:sc>\n            achieves improvements of 40.0% and 25.0% in test pass rate (Pass@1) for CodeT5 and CodeLlama, respectively, while maintaining the same pass rate for CodeGPT.\n            <jats:sc>ToolGen<\/jats:sc>\n            also demonstrates high efficiency in repository-level code generation, with latency ranging from 0.63 to 2.34 seconds for generating each function. Furthermore, our generalizability evaluation confirms\n            <jats:sc>ToolGen<\/jats:sc>\n            \u2019s consistent performance when applied to diverse code LLMs, encompassing various model architectures and scales.\n          <\/jats:p>","DOI":"10.1145\/3714462","type":"journal-article","created":{"date-parts":[[2025,1,27]],"date-time":"2025-01-27T15:53:31Z","timestamp":1737993211000},"update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":34,"title":["Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1424-6290","authenticated-orcid":false,"given":"Chong","family":"Wang","sequence":"first","affiliation":[{"name":"Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8316-1894","authenticated-orcid":false,"given":"Jian","family":"Zhang","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7235-2377","authenticated-orcid":false,"given":"Yebo","family":"Feng","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2207-1622","authenticated-orcid":false,"given":"Tianlin","family":"Li","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9236-8264","authenticated-orcid":false,"given":"Weisong","family":"Sun","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7300-9215","authenticated-orcid":false,"given":"Yang","family":"Liu","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3376-2581","authenticated-orcid":false,"given":"Xin","family":"Peng","sequence":"additional","affiliation":[{"name":"Fudan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,1,27]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2024. Argmax Function. https:\/\/en.wikipedia.org\/wiki\/Arg_max"},{"key":"e_1_2_1_2_1","unstructured":"2024. Jedi - an awesome autocompletion static analysis and refactoring library for Python. https:\/\/jedi.readthedocs.io\/"},{"key":"e_1_2_1_3_1","unstructured":"2024. Pylint. https:\/\/github.com\/pylint-dev\/pylint"},{"key":"e_1_2_1_4_1","unstructured":"2024. Replication Package. https:\/\/github.com\/cs-wangchong\/ToolGen-Replication\/"},{"key":"e_1_2_1_5_1","unstructured":"2024. Trie Structure. https:\/\/en.wikipedia.org\/wiki\/Trie"},{"key":"e_1_2_1_6_1","volume-title":"Guiding language models of code with global context using monitors. arXiv preprint arXiv:2306.10763","author":"Agrawal A","year":"2023","unstructured":"Lakshya\u00a0A Agrawal, Aditya Kanade, Navin Goyal, Shuvendu\u00a0K Lahiri, and Sriram\u00a0K Rajamani. 2023. Guiding language models of code with global context using monitors. arXiv preprint arXiv:2306.10763 (2023)."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2301.03988"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.18653\/V1"},{"key":"e_1_2_1_9_1","volume-title":"Program Synthesis with Large Language Models. CoRR abs\/2108.07732","author":"Austin Jacob","year":"2021","unstructured":"Jacob Austin, Augustus Odena, Maxwell\u00a0I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie\u00a0J. Cai, Michael Terry, Quoc\u00a0V. Le, and Charles Sutton. 2021. Program Synthesis with Large Language Models. CoRR abs\/2108.07732 (2021). arXiv:2108.07732 https:\/\/arxiv.org\/abs\/2108.07732"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2309.12499"},{"key":"e_1_2_1_11_1","volume-title":"Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015","author":"Bengio Samy","year":"2015","unstructured":"Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, Corinna Cortes, Neil\u00a0D. Lawrence, Daniel\u00a0D. Lee, Masashi Sugiyama, and Roman Garnett (Eds.). 1171\u20131179. https:\/\/proceedings.neurips.cc\/paper\/2015\/hash\/e995f98d56967d946471af29d7bf99f1-Abstract.html"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2204.06745"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","unstructured":"Sid Black Leo Gao Phil Wang Connor Leahy and Stella Biderman. 2021. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow. https:\/\/doi.org\/10.5281\/zenodo.5297715 If you use this software please cite it using these metadata..","DOI":"10.5281\/zenodo.5297715"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.JSS.2019.03.010"},{"key":"e_1_2_1_15_1","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique\u00a0Pond\u00e9 de Oliveira\u00a0Pinto Jared Kaplan Harrison Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mohammad Bavarian Clemens Winter Philippe Tillet Felipe\u00a0Petroski Such Dave Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William\u00a0Hebgen Guss Alex Nichol Alex Paino Nikolas Tezak Jie Tang Igor Babuschkin Suchir Balaji Shantanu Jain William Saunders Christopher Hesse Andrew\u00a0N. Carr Jan Leike Joshua Achiam Vedant Misra Evan Morikawa Alec Radford Matthew Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating Large Language Models Trained on Code. CoRR abs\/2107.03374 (2021). arXiv:2107.03374 https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2211.12588"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2207.11280"},{"key":"e_1_2_1_18_1","volume-title":"Training Verifiers to Solve Math Word Problems. CoRR abs\/2110.14168","author":"Cobbe Karl","year":"2021","unstructured":"Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training Verifiers to Solve Math Word Problems. CoRR abs\/2110.14168 (2021). arXiv:2110.14168 https:\/\/arxiv.org\/abs\/2110.14168"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639219"},{"key":"e_1_2_1_20_1","volume-title":"Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG. arXiv preprint arXiv:2406.11147","author":"Du Xueying","year":"2024","unstructured":"Xueying Du, Geng Zheng, Kaixin Wang, Jiayi Feng, Wentai Deng, Mingwei Liu, Bihuan Chen, Xin Peng, Tao Ma, and Yiling Lou. 2024. Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG. arXiv preprint arXiv:2406.11147 (2024)."},{"key":"e_1_2_1_21_1","volume-title":"InCoder: A Generative Model for Code Infilling and Synthesis. In The Eleventh International Conference on Learning Representations, ICLR 2023","author":"Fried Daniel","year":"2023","unstructured":"Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, and Mike Lewis. 2023. InCoder: A Generative Model for Code Infilling and Synthesis. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. https:\/\/openreview.net\/pdf?id=hQwb-lbM6EL"},{"key":"e_1_2_1_22_1","volume-title":"PAL: Program-aided Language Models. In International Conference on Machine Learning, ICML 2023","author":"Gao Luyu","year":"2023","unstructured":"Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. PAL: Program-aided Language Models. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Machine Learning Research, Vol.\u00a0202), Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 10764\u201310799. https:\/\/proceedings.mlr.press\/v202\/gao23f.html"},{"key":"e_1_2_1_23_1","volume-title":"LoRA: Low-Rank Adaptation of Large Language Models. In The Tenth International Conference on Learning Representations, ICLR 2022","author":"Hu J.","year":"2022","unstructured":"Edward\u00a0J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net. https:\/\/openreview.net\/forum?id=nZeVKeeFYf9"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3582570"},{"key":"e_1_2_1_25_1","volume-title":"CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. CoRR abs\/1909.09436","author":"Husain Hamel","year":"2019","unstructured":"Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019. CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. CoRR abs\/1909.09436 (2019). arXiv:1909.09436 http:\/\/arxiv.org\/abs\/1909.09436"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.579"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2203.05115"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2305.06161"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2203.07814"},{"key":"e_1_2_1_30_1","volume-title":"Boosting LLM-based Repository-level Code Completion with Static Analysis. arXiv preprint arXiv:2406.10018","author":"Liu Junwei","year":"2024","unstructured":"Junwei Liu, Yixuan Chen, Mingwei Liu, Xin Peng, and Yiling Lou. 2024. STALL+: Boosting LLM-based Repository-level Code Completion with Static Analysis. arXiv preprint arXiv:2406.10018 (2024)."},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021","author":"Lu Shuai","year":"2021","unstructured":"Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin\u00a0B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao\u00a0Kun Deng, Shengyu Fu, and Shujie Liu. 2021. CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual, Joaquin Vanschoren and Sai-Kit Yeung (Eds.). https:\/\/datasets-benchmarks-proceedings.neurips.cc\/paper\/2021\/hash\/c16a5320fa475530d9583c34fd356ef5-Abstract-round1.html"},{"key":"e_1_2_1_32_1","volume-title":"WebGPT: Browser-assisted question-answering with human feedback. CoRR abs\/2112.09332","author":"Nakano Reiichiro","year":"2021","unstructured":"Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021. WebGPT: Browser-assisted question-answering with human feedback. CoRR abs\/2112.09332 (2021). arXiv:2112.09332 https:\/\/arxiv.org\/abs\/2112.09332"},{"key":"e_1_2_1_33_1","volume-title":"CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. In The Eleventh International Conference on Learning Representations, ICLR 2023","author":"Nijkamp Erik","year":"2023","unstructured":"Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023. CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. https:\/\/openreview.net\/pdf?id=iaYcJKpY2B_"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2303.09014"},{"key":"e_1_2_1_36_1","volume-title":"6th International Conference on Learning Representations, ICLR","author":"Paulus Romain","year":"2018","unstructured":"Romain Paulus, Caiming Xiong, and Richard Socher. 2018. A Deep Reinforced Model for Abstractive Summarization. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net. https:\/\/openreview.net\/forum?id=HkAClQgA-"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2308.12950"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2302.04761"},{"key":"e_1_2_1_39_1","volume-title":"Repository-Level Prompt Generation for Large Language Models of Code. In International Conference on Machine Learning, ICML 2023","author":"Shrivastava Disha","year":"2023","unstructured":"Disha Shrivastava, Hugo Larochelle, and Daniel Tarlow. 2023. Repository-Level Prompt Generation for Large Language Models of Code. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Machine Learning Research, Vol.\u00a0202), Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 31693\u201331715. https:\/\/proceedings.mlr.press\/v202\/shrivastava23a.html"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-emnlp.27"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3417058"},{"key":"e_1_2_1_42_1","volume-title":"Claire Cui, Marian Croak, Ed\u00a0H. Chi, and Quoc Le.","author":"Thoppilan Romal","year":"2022","unstructured":"Romal Thoppilan, Daniel\u00a0De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu\u00a0Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Kathleen\u00a0S. Meier-Hellstern, Meredith\u00a0Ringel Morris, Tulsee Doshi, Renelito\u00a0Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise\u00a0Ag\u00fcera y Arcas, Claire Cui, Marian Croak, Ed\u00a0H. Chi, and Quoc Le. 2022. LaMDA: Language Models for Dialog Applications. CoRR abs\/2201.08239 (2022). arXiv:2201.08239 https:\/\/arxiv.org\/abs\/2201.08239"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2307.09288"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3491101.3519665"},{"key":"e_1_2_1_45_1","volume-title":"Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan\u00a0N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna\u00a0M. Wallach, Rob Fergus, S.\u00a0V.\u00a0N. Vishwanathan, and Roman Garnett (Eds.). 5998\u20136008. https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html"},{"key":"e_1_2_1_46_1","unstructured":"Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https:\/\/github.com\/kingoflolz\/mesh-transformer-jax."},{"key":"e_1_2_1_47_1","volume-title":"How and Why LLMs Use Deprecated APIs in Code Completion? An Empirical Study. arXiv preprint arXiv:2406.09834","author":"Wang Chong","year":"2024","unstructured":"Chong Wang, Kaifeng Huang, Jian Zhang, Yebo Feng, Lyuye Zhang, Yang Liu, and Xin Peng. 2024. How and Why LLMs Use Deprecated APIs in Code Completion? An Empirical Study. arXiv preprint arXiv:2406.09834 (2024)."},{"key":"e_1_2_1_48_1","volume-title":"Boosting Static Resource Leak Detection via LLM-based Resource-Oriented Intention Inference. arXiv preprint arXiv:2311.04448","author":"Wang Chong","year":"2023","unstructured":"Chong Wang, Jianan Liu, Xin Peng, Yang Liu, and Yiling Lou. 2023. Boosting Static Resource Leak Detection via LLM-based Resource-Oriented Intention Inference. arXiv preprint arXiv:2311.04448 (2023)."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE56229.2023.00130"},{"key":"e_1_2_1_50_1","volume-title":"TIGER: A Generating-Then-Ranking Framework for Practical Python Type Inference. arXiv preprint arXiv:2407.02095","author":"Wang Chong","year":"2024","unstructured":"Chong Wang, Jian Zhang, Yiling Lou, Mingwei Liu, Weisong Sun, Yang Liu, and Xin Peng. 2024. TIGER: A Generating-Then-Ranking Framework for Practical Python Type Inference. arXiv preprint arXiv:2407.02095 (2024)."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.18653\/V1"},{"key":"e_1_2_1_52_1","volume-title":"Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models. arXiv preprint arXiv:2407.00456","author":"Wang Yanlin","year":"2024","unstructured":"Yanlin Wang, Tianyue Jiang, Mingwei Liu, Jiachi Chen, and Zibin Zheng. 2024. Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models. arXiv preprint arXiv:2407.00456 (2024)."},{"key":"e_1_2_1_53_1","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023","author":"Wang Yue","year":"2023","unstructured":"Yue Wang, Hung Le, Akhilesh Gotmare, Nghi D.\u00a0Q. Bui, Junnan Li, and Steven C.\u00a0H. Hoi. 2023. CodeT5+: Open Code Large Language Models for Code Understanding and Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, 1069\u20131088. https:\/\/aclanthology.org\/2023.emnlp-main.68"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.18653\/V1"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3611643.3616271"},{"key":"e_1_2_1_56_1","volume-title":"CoderEval: A Benchmark of Pragmatic Code Generation with Generative Pre-trained Models. arXiv preprint arXiv:2302.00288","author":"Yu Hao","year":"2023","unstructured":"Hao Yu, Bo Shen, Dezhi Ran, Jiaxin Zhang, Qi Zhang, Yuchi Ma, Guangtai Liang, Ying Li, Tao Xie, and Qianxiang Wang. 2023. CoderEval: A Benchmark of Pragmatic Code Generation with Generative Pre-trained Models. arXiv preprint arXiv:2302.00288 (2023)."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2305.04207"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.151"},{"key":"e_1_2_1_59_1","volume-title":"Malicious Package Detection in NPM and PyPI using a Single Model of Malicious Behavior Sequence. arXiv preprint arXiv:2309.02637","author":"Zhang Junan","year":"2023","unstructured":"Junan Zhang, Kaifeng Huang, Bihuan Chen, Chong Wang, Zhenhao Tian, and Xin Peng. 2023. Malicious Package Detection in NPM and PyPI using a Single Model of Malicious Behavior Sequence. arXiv preprint arXiv:2309.02637 (2023)."},{"key":"e_1_2_1_60_1","volume-title":"An Empirical Study of Automated Vulnerability Localization with Large Language Models. arXiv preprint arXiv:2404.00287","author":"Zhang Jian","year":"2024","unstructured":"Jian Zhang, Chong Wang, Anran Li, Weisong Sun, Cen Zhang, Wei Ma, and Yang Liu. 2024. An Empirical Study of Automated Vulnerability Localization with Large Language Models. arXiv preprint arXiv:2404.00287 (2024)."},{"key":"e_1_2_1_61_1","volume-title":"ToolCoder: Teach Code Generation Models to use APIs with search tools. arXiv preprint arXiv:2305.04032","author":"Zhang Kechi","year":"2023","unstructured":"Kechi Zhang, Ge Li, Jia Li, Zhuo Li, and Zhi Jin. 2023. ToolCoder: Teach Code Generation Models to use APIs with search tools. arXiv preprint arXiv:2305.04032 (2023)."}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3714462","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,27]],"date-time":"2025-01-27T15:58:40Z","timestamp":1737993520000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3714462"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,27]]},"references-count":61,"alternative-id":["10.1145\/3714462"],"URL":"https:\/\/doi.org\/10.1145\/3714462","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,27]]},"assertion":[{"value":"2024-01-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-02","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"3714462"}}