{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T22:42:11Z","timestamp":1785537731537,"version":"3.56.0"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:00:00Z","timestamp":1720742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2024,7,12]]},"abstract":"<jats:p>As software projects progress, quality of code assumes paramount importance as it affects reliability, maintainability and security of software. For this reason, static analysis tools are used in developer workflows to flag code quality issues. However, developers need to spend extra efforts to revise their code to improve code quality based on the tool findings. In this work, we investigate the use of (instruction-following) large language models (LLMs) to assist developers in revising code to resolve code quality issues.<\/jats:p>\n                  <jats:p>We present a tool, CORE (short for COde REvisions), architected using a pair of LLMs organized as a duo comprised of a proposer and a ranker. Providers of static analysis tools recommend ways to mitigate the tool warnings and developers follow them to revise their code. The proposer LLM of CORE takes the same set of recommendations and applies them to generate candidate code revisions. The candidates which pass the static quality checks are retained. However, the LLM may introduce subtle, unintended functionality changes which may go un-detected by the static analysis. The ranker LLM evaluates the changes made by the proposer using a rubric that closely follows the acceptance criteria that a developer would enforce. CORE uses the scores assigned by the ranker LLM to rank the candidate revisions before presenting them to the developer.<\/jats:p>\n                  <jats:p>\n                    We conduct a variety of experiments on two public benchmarks to show the ability of CORE: (1) to generate code revisions acceptable to both static analysis tools and human reviewers (the latter evaluated with user study on a subset of the Python benchmark), (2) to reduce human review efforts by detecting and eliminating revisions with unintended changes, (3) to readily work across multiple languages (Python and Java), static analysis tools (CodeQL and SonarQube) and quality checks (52 and 10 checks, respectively), and (4) to achieve fix rate comparable to a rule-based automated program repair tool but with much smaller engineering efforts (on the Java benchmark). CORE could revise 59.2% Python files (across 52 quality checks) so that they pass scrutiny by both a tool and a human reviewer. The ranker LLM reduced false positives by 25.8% in these cases. CORE produced revisions that passed the static analysis tool in 76.8% Java files (across 10 quality checks) comparable to 78.3% of a specialized program repair tool, with significantly much less engineering efforts. We release code, data, and supplementary material publicly at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"http:\/\/aka.ms\/COREMSRI\">http:\/\/aka.ms\/COREMSRI<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3643762","type":"journal-article","created":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T10:22:09Z","timestamp":1720779729000},"page":"789-811","source":"Crossref","is-referenced-by-count":49,"title":["CORE: Resolving Code Quality Issues using LLMs"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-1206-9851","authenticated-orcid":false,"given":"Nalin","family":"Wadhwa","sequence":"first","affiliation":[{"name":"Microsoft Research, Bangalore, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-7753-6210","authenticated-orcid":false,"given":"Jui","family":"Pradhan","sequence":"additional","affiliation":[{"name":"Microsoft Research, Bangalore, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-2132-0458","authenticated-orcid":false,"given":"Atharv","family":"Sonwane","sequence":"additional","affiliation":[{"name":"Microsoft Research, Bangalore, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-9943-5222","authenticated-orcid":false,"given":"Surya Prakash","family":"Sahu","sequence":"additional","affiliation":[{"name":"Microsoft Research, Bangalore, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6435-245X","authenticated-orcid":false,"given":"Nagarajan","family":"Natarajan","sequence":"additional","affiliation":[{"name":"Microsoft Research, Bangalore, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8701-6977","authenticated-orcid":false,"given":"Aditya","family":"Kanade","sequence":"additional","affiliation":[{"name":"Microsoft Research, Bangalore, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6675-3219","authenticated-orcid":false,"given":"Suresh","family":"Parthasarathy","sequence":"additional","affiliation":[{"name":"Microsoft Research, Bangalore, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1400-7065","authenticated-orcid":false,"given":"Sriram","family":"Rajamani","sequence":"additional","affiliation":[{"name":"Microsoft Research, Bangalore, India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,7,12]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"[n.d.]. CodeQL website. https:\/\/codeql.github.com\/. Accessed: September 15 2023."},{"key":"e_1_3_1_3_2","unstructured":"[n.d.]. Coverity Static Analysis. https:\/\/www.synopsys.com\/software-integrity\/security-testing\/static-analysis-sast.html. Accessed: September 15 2023."},{"key":"e_1_3_1_4_2","unstructured":"[n.d.]. __eq__ not Overridden when adding attributes. https:\/\/codeql.github.com\/codeql-query-help\/python\/py-missing-equals\/. Accessed: September 15 2023."},{"key":"e_1_3_1_5_2","unstructured":"[n.d.]. FindBugs Project. https:\/\/spotbugs.github.io\/. Accessed: September 15 2023."},{"key":"e_1_3_1_6_2","unstructured":"[n.d.]. Infer static analyzer. https:\/\/fbinfer.com\/. Accessed: September 15 2023."},{"key":"e_1_3_1_7_2","unstructured":"[n.d.]. InferSharp static analyzer. https:\/\/github.com\/microsoft\/infersharp. Accessed: September 15 2023."},{"key":"e_1_3_1_8_2","unstructured":"[n.d.]. PMD: An extensible cross-language static code analyzer. https:\/\/pmd.github.io\/. Accessed: September 15 2023."},{"key":"e_1_3_1_9_2","unstructured":"[n.d.]. SonarQube. https:\/\/docs.SonarQube.org\/latest\/. Accessed: September 15 2023."},{"key":"e_1_3_1_10_2","unstructured":"[n.d.]. Sorald Tool Source. https:\/\/github.com\/ASSERT-KTH\/sorald\/releases\/tag\/sorald-0.8.5."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/SCAM.2012.28"},{"key":"e_1_3_1_12_2","unstructured":"Lakshya A Agrawal Aditya Kanade Navin Goyal Shuvendu K Lahiri and Sriram K Rajamani. 2023. Guiding Language Models of Code with Global Context using Monitors. arXiv preprint arXiv:2306.10763 2023."},{"key":"e_1_3_1_13_2","volume-title":"30th European Conference on Object-Oriented Programming","author":"Avgustinov Pavel","year":"2016","unstructured":"Pavel Avgustinov, Oege de Moor, Michael Peyton Jones, and Max Sch\u00e4r. 2016. QL: Object-oriented Queries on Relational Data. In 30th European Conference on Object-Oriented Programming. Schloss Dagstuhl - Leibniz-Zentrum f\u00fcr Informatik."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","unstructured":"Johannes Bader Andrew Scott Michael Pradel and Satish Chandra. 2019. Getafix: Learning to Fix Bugs Automatically. Proc. ACM Program. Lang. 3 OOPSLA (oct 2019) 27 pages. https:\/\/doi.org\/10.1145\/3360585 10.1145\/3360585","DOI":"10.1145\/3360585"},{"key":"e_1_3_1_15_2","unstructured":"Yuntao Bai Saurav Kadavath Sandipan Kundu Amanda Askell Jackson Kernion Andy Jones Anna Chen Anna Goldie Azalia Mirhoseini Cameron McKinnon et al.. 2022. Constitutional AI: Harmlessness from AI Feedback. arXiv preprint arXiv:2212.08073 (2022)."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3338906.3338952"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al.. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020) 1877\u20131901. https:\/\/doi.org\/10.1145\/3373574.3373576 10.1145\/3373574.3373576","DOI":"10.1145\/3373574.3373576"},{"key":"e_1_3_1_18_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman et al.. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)."},{"key":"e_1_3_1_19_2","unstructured":"Aakanksha Chowdhery Sharan Narang Jacob Devlin Maarten Bosma Gaurav Mishra Adam Roberts Paul Barham Hyung Won Chung Charles Sutton Sebastian Gehrmann et al.. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311 (2022)."},{"key":"e_1_3_1_20_2","unstructured":"Paul F Christiano Jan Leike Tom Brown Miljan Martic Shane Legg and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Andreea Costea Abhishek Tiwari Sigmund Chianasta Abhik Roychoudhury and Ilya Sergey. 2023. Hippodrome: Data race repair using static analysis summaries. ACM Transactions on Software Engineering and Methodology 32 2 (2023) 1\u201333.","DOI":"10.1145\/3546942"},{"key":"e_1_3_1_22_2","unstructured":"Paul M Duvall Steve Matyas and Andrew Glover. 2007. Continuous integration: improving software quality and reducing risk. Pearson Education."},{"key":"e_1_3_1_23_2","unstructured":"Zhiyu Fan Xiang Gao Abhik Roychoudhury and Shin Hwei Tan. 2022. Automated Repair of Programs from Large Language Models. arXiv preprint arXiv:2205.10583 (2022)."},{"key":"e_1_3_1_24_2","unstructured":"Martin Fowler. 2018. Refactoring: Improving the Design of Existing Code. Addison-Wesley Professional."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3318162"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","unstructured":"Yang Hong Chakkrit Tantithamthavorn Patanamon Thongtanunam and Aldeida Aleti. 2022. CommentFinder: A simpler faster more accurate code review comments recommendation (ESEC\/FSE 2022). Association for Computing Machinery New York NY USA 507\u2013519. https:\/\/doi.org\/10.1145\/3540250.3549119 10.1145\/3540250.3549119","DOI":"10.1145\/3540250.3549119"},{"key":"e_1_3_1_27_2","unstructured":"Kai Huang Zhengzi Xu Su Yang Hongyu Sun Xuejun Li Zheng Yan and Yuqing Zhang. 2023. A Survey on Automated Program Repair Techniques. arXiv preprint arXiv:2303.18184 (2023)."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","unstructured":"Zhen Huang David Lie Gang Tan and Trent Jaeger. 2019. Using Safety Properties to Generate Vulnerability Patches. In 2019 IEEE Symposium on Security and Privacy (SP) 539\u2013554. https:\/\/doi.org\/10.1109\/SP.2019.00071 10.1109\/SP.2019.00071","DOI":"10.1109\/SP.2019.00071"},{"key":"e_1_3_1_29_2","unstructured":"Naman Jain Shubham Gandhi Atharv Sonwane Aditya Kanade Nagarajan Natarajan Suresh Parthasarathy Sriram Rajamani and Rahul Sharma. 2023. StaticFixer: From Static Analysis to Static Repair. arXiv:2307.12465 [cs.SE]"},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","unstructured":"Naman Jain Skanda Vaidyanath Arun Iyer Nagarajan Natarajan Suresh Parthasarathy Sriram Rajamani and Rahul Sharma. 2022. Jigsaw: Large language models meet program synthesis. In Proceedings of the 44th International Conference on Software Engineering 1219\u20131231.","DOI":"10.1145\/3510003.3510203"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00107"},{"key":"e_1_3_1_32_2","unstructured":"Matthew Jin Syed Shahriar Michele Tufano Xin Shi Shuai Lu Neel Sundaresan and Alexey Svyatkovskiy. 2023. InferFix: End-to-End Program Repair with LLMs. arXiv preprint arXiv:2303.07263 (2023)."},{"key":"e_1_3_1_33_2","unstructured":"Harshit Joshi Jos\u00e9 Cambronero Sumit Gulwani Vu Le Ivan Radicek and Gust Verbruggen. 2022. Repair is nearly generation: Multilingual program repair with llms. arXiv preprint arXiv:2208.11640 (2022)."},{"key":"e_1_3_1_34_2","unstructured":"Tom Kocmi and Christian Federmann. 2023. Large language models are state-of-the-art evaluators of translation quality. arXiv preprint arXiv:2302.14520 (2023)."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597207"},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","unstructured":"Lingwei Li Li Yang Huaxi Jiang Jun Yan Tiejian Luo Zihan Hua Geng Liang and Chun Zuo. 2022. AUGER: Automatically Generating Review Comments with Pre-training Models. arXiv:2208.08014 [cs.SE]","DOI":"10.1145\/3540250.3549099"},{"key":"e_1_3_1_37_2","unstructured":"Zhiyu Li Shuai Lu Daya Guo Nan Duan Shailesh Jannu Grant Jenks Deep Majumder Jared Green Alexey Svyatkovskiy Shengyu Fu and Neel Sundaresan. 2022. Automating Code Review Activities by Large-Scale Pre-training. arXiv:2203.09095 [cs.SE]"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2018.2884955"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","unstructured":"Kui Liu Anil Koyuncu Dongsun Kim and Tegawende F. Bissyand. 2019. AVATAR: Fixing Semantic Bugs with Fix Patterns of Static Analysis Violations. In 2019 IEEE 26th International Conference on Software Analysis Evolution and Reengineering (SANER). 1\u201312. https:\/\/doi.org\/10.1109\/SANER.2019.8667970 10.1109\/SANER.2019.8667970","DOI":"10.1109\/SANER.2019.8667970"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","unstructured":"Vadim Liventsev Anastasiia Grishina Aki H\u00e4rm\u00e4 and Leon Moonen. 2023. Fully Autonomous Programming with Large Language Models. arXiv preprint arXiv:2304.10423 (2023).","DOI":"10.1145\/3583131.3590481"},{"key":"e_1_3_1_42_2","doi-asserted-by":"crossref","unstructured":"Fan Long and Martin Rinard. 2015. Staged program repair with condition synthesis. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering 166\u2013178.","DOI":"10.1145\/2786805.2786811"},{"key":"e_1_3_1_43_2","unstructured":"Ziyang Luo Can Xu Pu Zhao Qingfeng Sun Xiubo Geng Wenxiang Hu Chongyang Tao Jing Ma Qingwei Lin and Daxin Jiang. 2023. WizardCoder: Empowering Code Large Language Models with Evol-Instruct. arXiv:2306.08568 [cs.CL]"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/SCAM.2019.00013"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3105906"},{"key":"e_1_3_1_46_2","unstructured":"Theo X. Olausson Jeevana Priya Inala Chenglong Wang Jianfeng Gao and Armando Solar-Lezama. 2023. Demystifying GPT Self-Repair for Code Generation. arXiv:2306.09896 [cs.CL]"},{"key":"e_1_3_1_47_2","unstructured":"OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL]"},{"key":"e_1_3_1_48_2","unstructured":"Long Ouyang Jeffrey Wu Xu Jiang Diogo Almeida Carroll Wainwright Pamela Mishkin Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray et al.. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35 (2022) 27730\u201327744."},{"key":"e_1_3_1_49_2","first-page":"1","volume-title":"2023 IEEE Symposium on Security and Privacy (SP)","author":"Pearce Hammond","year":"2022","unstructured":"Hammond Pearce, Benjamin Tan, Baleegh Ahmad, Ramesh Karri, and Brendan Dolan-Gavitt. 2022. Examining Zero-Shot Vulnerability Repair with Large Language Models. In 2023 IEEE Symposium on Security and Privacy (SP), IEEE Computer Society, 1\u201318."},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","unstructured":"Jeff H Perkins Sunghun Kim Sam Larsen Saman Amarasinghe Jonathan Bachrach Michael Carbin Carlos Pacheco Frank Sherwood Stelios Sidiroglou Greg Sullivan et al.. 2009. Automatically patching errors in deployed software. In Proceedings of the ACM SIGOPS 22nd symposium on Operating systems principles 87\u2013102.","DOI":"10.1145\/1629575.1629585"},{"key":"e_1_3_1_51_2","unstructured":"Reudismam Rolim Gustavo Soares Rohit Gheyi Titus Barik and Loris D\u2019Antoni. 2018. Learning Quick Fixes from Code Repositories. arXiv:1803.03806 [cs.SE]"},{"key":"e_1_3_1_52_2","unstructured":"Surya Prakash Sahu Madhurima Mandal Shikhar Bharadwaj Aditya Kanade Petros Maniatis and Shirish Shevade. 2024. CodeQueries: A Dataset of Semantic Queries over Code. In 17th Innovations in Software Engineering Conference ACM."},{"key":"e_1_3_1_53_2","unstructured":"Noah Shinn Federico Cassano Beck Labash Ashwin Gopinath Karthik Narasimhan and Shunyu Yao. 2023. Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366 [cs.AI]"},{"key":"e_1_3_1_54_2","article-title":"Sorald: Automatic Patch Suggestions for Sonar\n                  Qube Static Analysis Violations","author":"Someoliayi Khashayar Etemadi","year":"2022","unstructured":"Khashayar Etemadi Someoliayi, Nicolas Yves Maurice Harrand, Simon Larsen, Haris Adzemovic, Henry Luong Phu, Ashutosh Verma, Fernanda Madeiral, Douglas Wikstrom, and Martin Monperrus. 2022. Sorald: Automatic Patch Suggestions for SonarQube Static Analysis Violations. IEEE Transactions on Dependable and Secure Computing (2022).","journal-title":"IEEE Transactions on Dependable and Secure Computing"},{"key":"e_1_3_1_55_2","unstructured":"Nisan Stiennon Long Ouyang Jeffrey Wu Daniel Ziegler Ryan Lowe Chelsea Voss Alec Radford Dario Amodei and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems 33 (2020) 3008\u20133021."},{"key":"e_1_3_1_56_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e9re Naman Goyal Eric Hambro Faisal Azhar et al.. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)."},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340544"},{"key":"e_1_3_1_58_2","doi-asserted-by":"crossref","unstructured":"Rosalia Tufano Simone Masiero Antonio Mastropaolo Luca Pascarella Denys Poshyvanyk and Gabriele Bavota. 2022. Using Pre-Trained Models to Boost Code Review Automation. arXiv:2201.06850 [cs.SE]","DOI":"10.1145\/3510003.3510621"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3180155.3180250"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-019-09750-5"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00129"},{"key":"e_1_3_1_62_2","unstructured":"Chunqiu Steven Xia and Lingming Zhang. 2023. Conversational Automated Program Repair. arXiv preprint arXiv:2301.13246 (2023)."},{"key":"e_1_3_1_63_2","doi-asserted-by":"crossref","unstructured":"He Ye Matias Martinez and Martin Monperrus. 2022. Neural Program Repair with Execution-Based Backpropagation. In Proceedings of the 44th International Conference on Software Engineering. 1506\u20131518.","DOI":"10.1145\/3510003.3510222"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","unstructured":"Fiorella Zampetti Simone Scalabrino Rocco Oliveto Gerardo Canfora and Massimiliano Di Penta. 2017. How Open Source Projects Use Static Code Analysis Tools in Continuous Integration Pipelines. In 2017 IEEE\/ACM 14th International Conference on Mining Software Repositories (MSR). 334\u2013344. https:\/\/doi.org\/10.1109\/MSR.2017.2 10.1109\/MSR.2017.2","DOI":"10.1109\/MSR.2017.2"},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","unstructured":"Qihao Zhu Zeyu Sun Yuan-an Xiao Wenjie Zhang Kang Yuan Yingfei Xiong and Lu Zhang. 2021. A Syntax-Guided Edit Decoder for Neural Program Repair. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 341\u2013353.","DOI":"10.1145\/3468264.3468544"},{"key":"e_1_3_1_66_2","unstructured":"Terry Yue Zhuo. 2023. Large Language Models Are State-of-the-Art Evaluators of Code Generation. arXiv:2304.14317 [cs.AI]."}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643762","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3643762","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T07:59:18Z","timestamp":1770191958000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643762"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,12]]},"references-count":65,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2024,7,12]]}},"alternative-id":["10.1145\/3643762"],"URL":"https:\/\/doi.org\/10.1145\/3643762","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,12]]}}}