{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T03:20:30Z","timestamp":1785381630582,"version":"3.55.0"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"abstract":"<jats:p>\n            AI programming has become a popular topic in recent years. Code suggestion, with code suggestion being a key capability of AI programming. Copilot, an \u201cAI programmer\u201d that provides code suggestions from natural language descriptions, has been launched by GitHub and OpenAI. By far, Copilot has been widely used by millions of developers. However, little work has systematically evaluated the correctness of Copilot's suggestions. We conducted an empirical study on all 2,033\n            <jats:italic>LeetCode<\/jats:italic>\n            problems to assess Copilot's code generation across four mainstream languages: C, Java, JavaScript, and Python. We have found that: 1) 70.0% of problems received at least one correct suggestion, with language-specific rates of 29.7% (C), 57.7% (Java), 54.1% (JavaScript), and 41.0% (Python); 2) Correctness decreases as problem difficulty increases, with acceptance rates of 89.3% (Easy), 72.1% (Medium), and 43.4% (Hard); 3) Acceptance rates vary across problem domains from 49.5% to 90.1%, while\n            <jats:italic>Graph<\/jats:italic>\n            problems challenge C and Python most, and\n            <jats:italic>Prefix Sum<\/jats:italic>\n            and\n            <jats:italic>Heap<\/jats:italic>\n            challenge Java and JavaScript most; 4) For the incorrect suggestions, we further summarize 17 types of error reasons accounting for their incorrectness and analyzed possible causes for why these errors occur. We believe our study can provide valuable insights into Copilot's capabilities and limitations.\n          <\/jats:p>","DOI":"10.1145\/3715108","type":"journal-article","created":{"date-parts":[[2025,1,27]],"date-time":"2025-01-27T15:53:31Z","timestamp":1737993211000},"update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Assessing and Analyzing the Correctness of GitHub Copilot\u2019s Code Suggestions"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7556-153X","authenticated-orcid":false,"given":"Ran","family":"Mo","sequence":"first","affiliation":[{"name":"Central China Normal University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6896-1687","authenticated-orcid":false,"given":"Dongyu","family":"Wang","sequence":"additional","affiliation":[{"name":"Central China Normal University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-1966-5517","authenticated-orcid":false,"given":"Wenjing","family":"Zhan","sequence":"additional","affiliation":[{"name":"Central China Normal University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-3385-7403","authenticated-orcid":false,"given":"Yingjie","family":"Jiang","sequence":"additional","affiliation":[{"name":"Central China Normal University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-8055-2425","authenticated-orcid":false,"given":"Yepeng","family":"Wang","sequence":"additional","affiliation":[{"name":"University of South Carolina, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2642-5109","authenticated-orcid":false,"given":"Yuqi","family":"Zhao","sequence":"additional","affiliation":[{"name":"Central China Normal University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7258-993X","authenticated-orcid":false,"given":"Zengyang","family":"Li","sequence":"additional","affiliation":[{"name":"Central China Normal University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4239-2009","authenticated-orcid":false,"given":"Yutao","family":"Ma","sequence":"additional","affiliation":[{"name":"Central China Normal University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,1,27]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3212695"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-023-10380-1"},{"key":"e_1_2_1_3_1","first-page":"4818","article-title":"An empirical study on the usage of transformer models for code completion","volume":"48","author":"Ciniselli Matteo","year":"2021","unstructured":"Matteo Ciniselli, Nathan Cooper, Luca Pascarella, Antonio Mastropaolo, Emad Aghajani, Denys Poshyvanyk, Massimiliano Di\u00a0Penta, and Gabriele Bavota. 2021. An empirical study on the usage of transformer models for code completion. IEEE Transactions on Software Engineering 48, 12 (2021), 4818\u20134837.","journal-title":"IEEE Transactions on Software Engineering"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSR52588.2021.00024"},{"key":"e_1_2_1_5_1","volume-title":"Retrieved","author":"Copilot Github","year":"2022","unstructured":"Github Copilot. 2022. About GitHub Copilot Individual. Retrieved March 25, 2023 from https:\/\/docs.github.com\/en\/copilot\/overview-of-github-copilot\/about-github-copilot-for-individuals"},{"key":"e_1_2_1_6_1","volume-title":"Retrieved","author":"Copilot Github","year":"2022","unstructured":"Github Copilot. 2022. your AI pair programmer. Retrieved December 25, 2023 from https:\/\/github.com\/features\/copilot"},{"key":"e_1_2_1_7_1","volume-title":"Retrieved","author":"Coporation MITRE","year":"2021","unstructured":"The\u00a0MITRE Coporation. 2021. 2021 CWE top 25 Most Dangerous Software Weakness. Retrieved March 25, 2023 from https:\/\/cwe.mitre.org\/top25\/archive\/2021\/2021cwetop25.html"},{"key":"e_1_2_1_8_1","volume-title":"Introduction to algorithms","author":"Cormen H","unstructured":"Thomas\u00a0H Cormen, Charles\u00a0E Leiserson, Ronald\u00a0L Rivest, and Clifford Stein. 2022. Introduction to algorithms. MIT press."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2023.111734"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2442754.2442763"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3510454.3528648"},{"key":"e_1_2_1_12_1","volume-title":"Retrieved","year":"2024","unstructured":"FrozenFish. 2024. author's github repository. Retrieved Feb, 2024 from https:\/\/github.com\/frozenfishDY\/CopilotEvaluation"},{"key":"e_1_2_1_13_1","volume-title":"Retrieved","year":"2022","unstructured":"Github. 2022. The top programming languages. Retrieved March 25, 2023 from https:\/\/octoverse.github.com\/2022\/top-programming-languages"},{"key":"e_1_2_1_14_1","volume-title":"Retrieved","author":"Hanley Mike","year":"2022","unstructured":"Mike Hanley. 2022. How we use GitHub to be more productive, collaborative, and secure. Retrieved March 25, 2023 from https:\/\/github.blog\/2022-12-20-how-we-use-github-to-be-more-productive-collaborative-and-secure\/"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3106237.3106290"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3449639.3459285"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2739480.2754769"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3510454.3522684"},{"key":"e_1_2_1_19_1","volume-title":"Retrieved","author":"Kalliamvakou Eirini","year":"2022","unstructured":"Eirini Kalliamvakou. 2022. Research: quantifying GitHub Copilot's impact on developer productivity and happiness. Retrieved March 25, 2023 from https:\/\/github.blog\/2022-09-07-research-quantifying-github-copilots-impact-on-developer-productivity-and-happiness\/"},{"key":"e_1_2_1_20_1","volume-title":"Maybe deep neural networks are the best choice for modeling source code. arXiv preprint arXiv:1903.05734","author":"Karampatsis Rafael-Michael","year":"2019","unstructured":"Rafael-Michael Karampatsis and Charles Sutton. 2019. Maybe deep neural networks are the best choice for modeling source code. arXiv preprint arXiv:1903.05734 (2019)."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00026"},{"key":"e_1_2_1_22_1","volume-title":"Applying large language models and chain-of-thought for automatic scoring. Computers and Education: Artificial Intelligence","author":"Lee Gyeong-Geon","year":"2024","unstructured":"Gyeong-Geon Lee, Ehsan Latif, Xuansheng Wu, Ninghao Liu, and Xiaoming Zhai. 2024. Applying large language models and chain-of-thought for automatic scoring. Computers and Education: Artificial Intelligence (2024), 100213."},{"key":"e_1_2_1_23_1","volume-title":"Retrieved","year":"2011","unstructured":"LeetCode. 2011. The world's leading online programming learning platform. Retrieved March 25, 2023 from https:\/\/leetcode.com\/"},{"key":"e_1_2_1_24_1","volume-title":"Retrieved","year":"2016","unstructured":"LeetCode. 2016. Remove Duplicates from Sorted List. Retrieved March 25, 2023 from https:\/\/leetcode.com\/problems\/remove-duplicates-from-sorted-list\/"},{"key":"e_1_2_1_25_1","volume-title":"Code completion with neural attention and pointer networks. arXiv preprint arXiv:1711.09573","author":"Li Jian","year":"2017","unstructured":"Jian Li, Yue Wang, Michael\u00a0R Lyu, and Irwin King. 2017. Code completion with neural attention and pointer networks. arXiv preprint arXiv:1711.09573 (2017)."},{"key":"e_1_2_1_26_1","volume-title":"Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74\u201381.","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74\u201381."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-022-10140-7"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416591"},{"key":"e_1_2_1_29_1","volume-title":"Exploring and evaluating hallucinations in llm-powered code generation. arXiv preprint arXiv:2404.00971","author":"Liu Fang","year":"2024","unstructured":"Fang Liu, Yang Liu, Lin Shi, Houkun Huang, Ruifeng Wang, Zhen Yang, and Li Zhang. 2024. Exploring and evaluating hallucinations in llm-powered code generation. arXiv preprint arXiv:2404.00971 (2024)."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3360578"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00181"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3524842.3528470"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409690"},{"key":"e_1_2_1_34_1","volume-title":"Retrieved","author":"AI.","year":"2023","unstructured":"OpenAI. 200. OpenAI CodeX. Retrieved December 25, 2023 from https:\/\/openai.com\/blog\/openai-codex\/"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.17706\/jsw.11.11.1083-1088"},{"key":"e_1_2_1_36_1","volume-title":"LLM is Like a Box of Chocolates: the Non-determinism of ChatGPT in Code Generation. arXiv preprint arXiv:2308.02828","author":"Ouyang Shuyin","year":"2023","unstructured":"Shuyin Ouyang, Jie\u00a0M Zhang, Mark Harman, and Meng Wang. 2023. LLM is Like a Box of Chocolates: the Non-determinism of ChatGPT in Code Generation. arXiv preprint arXiv:2308.02828 (2023)."},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311\u2013318","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP46214.2022.9833571"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10515-010-0064-x"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3512290.3528700"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2635868.2635875"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSR59073.2023.00035"},{"key":"e_1_2_1_43_1","volume-title":"Denny Zhou, et\u00a0al.","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc\u00a0V Le, Denny Zhou, et\u00a0al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824\u201324837."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3545945.3569830"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300864"},{"key":"e_1_2_1_46_1","volume-title":"Towards better chain-of-thought prompting strategies: A survey. arXiv preprint arXiv:2310.04959","author":"Yu Zihan","year":"2023","unstructured":"Zihan Yu, Liang He, Zhen Wu, Xinyu Dai, and Jiajun Chen. 2023. Towards better chain-of-thought prompting strategies: A survey. arXiv preprint arXiv:2310.04959 (2023)."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2007.1078"},{"key":"e_1_2_1_48_1","volume-title":"International conference on machine learning. PMLR, 11328\u201311339","author":"Zhang Jingqing","year":"2020","unstructured":"Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In International conference on machine learning. PMLR, 11328\u201311339."}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715108","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,27]],"date-time":"2025-01-27T15:56:32Z","timestamp":1737993392000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715108"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,27]]},"references-count":48,"alternative-id":["10.1145\/3715108"],"URL":"https:\/\/doi.org\/10.1145\/3715108","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,27]]},"assertion":[{"value":"2024-05-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-18","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"3715108"}}