{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,25]],"date-time":"2026-02-25T17:13:06Z","timestamp":1772039586417,"version":"3.50.1"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"3","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2026,3,31]]},"abstract":"<jats:p>\n                    Recent years have witnessed significant progress in developing deep learning-based models for automated code completion. Examples of such models include CodeGPT and StarCoder. These models are typically trained from a large amount of source code collected from open source communities such as GitHub. Although using source code in GitHub has been a common practice for training deep-learning-based models for code completion, it may induce some legal and ethical issues such as copyright infringement. In this article, we investigate the legal and ethical issues of current neural code completion models by answering the following question:\n                    <jats:italic toggle=\"yes\">Is my code used to train your neural code completion model<\/jats:italic>\n                    ?\n                  <\/jats:p>\n                  <jats:p>\n                    To this end, we tailor a membership inference approach (termed\n                    <jats:sc>CodeMI<\/jats:sc>\n                    ) that was originally crafted for classification tasks to a more challenging task of code completion. In particular, since the target code completion models perform as opaque black boxes, preventing access to their training data and parameters, we opt to train multiple shadow models to mimic their behavior. The acquired posteriors from these shadow models are subsequently employed to train a membership classifier. After that, the membership classifier can be effectively employed to deduce the membership status of a given code sample based on the output of a target code completion model. We comprehensively evaluate the effectiveness of this adapted approach across a diverse array of neural code completion models (i.e., LSTM-based, CodeGPT, CodeGen, and StarCoder). Experimental results demonstrate that our approach effectively detects data membership, achieving accuracies of 0.842 and 0.730 for LSTM-based and CodeGPT models, respectively. Interestingly, our experiments also show that the data membership of current large language models of code, e.g., CodeGen and StarCoder, is difficult to detect, leaving ample space for further improvement. Finally, we also try to explain the findings from the perspective of model memorization.\n                  <\/jats:p>","DOI":"10.1145\/3742785","type":"journal-article","created":{"date-parts":[[2025,6,25]],"date-time":"2025-06-25T11:15:02Z","timestamp":1750850102000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Does Your Neural Code Completion Model Use My Code? A Membership Inference Approach"],"prefix":"10.1145","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6937-4180","authenticated-orcid":false,"given":"Yao","family":"Wan","sequence":"first","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4515-4398","authenticated-orcid":false,"given":"Guanghua","family":"Wan","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-4661-5900","authenticated-orcid":false,"given":"Shijie","family":"Zhang","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3063-9425","authenticated-orcid":false,"given":"Hongyu","family":"Zhang","sequence":"additional","affiliation":[{"name":"Chongqing University, Chongqing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9510-6574","authenticated-orcid":false,"given":"Yulei","family":"Sui","sequence":"additional","affiliation":[{"name":"University of New South Wales, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8747-9912","authenticated-orcid":false,"given":"Pan","family":"Zhou","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3934-7605","authenticated-orcid":false,"given":"Hai","family":"Jin","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1539-7939","authenticated-orcid":false,"given":"Lichao","family":"Sun","sequence":"additional","affiliation":[{"name":"Lehigh University, Bethlehem, Pennsylvania, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,2,13]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"About Github Copilot telemetry. 2022. Retrieved September 1 2023 from https:\/\/ tinyurl.com\/37w8nfnz"},{"key":"e_1_3_2_3_2","unstructured":"ChatGPT. 2022. Retrieved December 1 2022 from https:\/\/openai.com\/blog\/chatgpt\/"},{"key":"e_1_3_2_4_2","unstructured":"Codex. 2022. Retrieved December 1 2022 from https:\/\/openai.com\/blog\/openai-codex\/"},{"key":"e_1_3_2_5_2","unstructured":"Copilot. 2022. Retrieved December 1 2022 from https:\/\/github.com\/features\/copilot\/"},{"key":"e_1_3_2_6_2","unstructured":"IntelliCode. 2022. Retrieved December 1 2022 from https:\/\/visualstudio.microsoft.com\/services\/intellicode\/"},{"key":"e_1_3_2_7_2","unstructured":"Tabnine. 2022. Retrieved December 1 2022 from https:\/\/www.tabnine.com"},{"key":"e_1_3_2_8_2","first-page":"245","volume-title":"Proceedings of the 37th International Conference on Machine Learning","volume":"119","author":"Alon Uri","year":"2020","unstructured":"Uri Alon, Roy Sadaka, Omer Levy, and Eran Yahav. 2020. Structural language models of code. In Proceedings of the 37th International Conference on Machine Learning, Vol. 119 (2020), 245\u2013256."},{"key":"e_1_3_2_9_2","unstructured":"Gareth Ari Aye and Gail E. Kaiser. 2020. Sequence model design for code completion in the modern IDE. arXiv:2004.05249. Retrieved from https:\/\/arxiv.org\/abs\/2004.05249"},{"key":"e_1_3_2_10_2","first-page":"1877","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"33","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 33, 1877\u20131901."},{"key":"e_1_3_2_11_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Carlini Nicholas","year":"2023","unstructured":"Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tram\u00e8r, and Chiyuan Zhang. 2023. Quantifying memorization across neural language models. In Proceedings of the 11th International Conference on Learning Representations."},{"key":"e_1_3_2_12_2","first-page":"2633","volume-title":"Proceedings of the 30th USENIX Security Symposium","author":"Carlini Nicholas","year":"2021","unstructured":"Nicholas Carlini, Florian Tram\u00e8r, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Song, \u00dalfar Erlingsson, et al. 2021. Extracting training data from large language models. In Proceedings of the 30th USENIX Security Symposium, 2633\u20132650."},{"key":"e_1_3_2_13_2","unstructured":"Junyoung Chung \u00c7aglar G\u00fcl\u00e7ehre KyungHyun Cho and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv:1412.3555. Retrieved from https:\/\/arxiv.org\/abs\/1412.3555"},{"key":"e_1_3_2_14_2","first-page":"9243","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP)","author":"Guan Batu","year":"2024","unstructured":"Batu Guan, Yao Wan, Zhangqian Bi, Zheng Wang, Hongyu Zhang, Pan Zhou, and Lichao Sun. 2024. CodeIP: A grammar-guided multi-bit watermark for large language models of code. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP), 9243\u20139258."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2022.3141725"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_17_2","volume-title":"Proceedings of the 31st International Joint Conference on Artificial Intelligence","author":"Hu Hongsheng","year":"2022","unstructured":"Hongsheng Hu, Zoran Salcic, Gillian Dobbie, Chen Jinjun, Lichao Sun, and Xuyun Zhang. 2022. Membership inference via backdooring. In Proceedings of the 31st International Joint Conference on Artificial Intelligence."},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3523273"},{"issue":"9","key":"e_1_3_2_19_2","article-title":"Deep learning for code generation: A survey","volume":"67","author":"Huangzhao Zhang","year":"2024","unstructured":"Zhang Huangzhao, Zhang Kechi, Li Zhuo, Li Jia, Li Yongmin, Zhao Yunfei, Zhu Yuqi, Liu Fang, Li Ge, and Jin Zhi. 2024. Deep learning for code generation: A survey. Science China Information Sciences 67, 9 (2024).","journal-title":"Science China Information Sciences"},{"key":"e_1_3_2_20_2","unstructured":"Hamel Husain Ho-Hsiang Wu Tiferet Gazit Miltiadis Allamanis and Marc Brockschmidt. 2019. CodeSearchNet challenge: Evaluating the state of semantic code search. arXiv:1909.09436. Retrieved from https:\/\/arxiv.org\/abs\/1909.09436"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00026"},{"key":"e_1_3_2_22_2","article-title":"The stack: 3 TB of permissively licensed source code","author":"Kocetkov Denis","year":"2023","unstructured":"Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Mu\u00f1oz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, et al. 2023. The stack: 3 TB of permissively licensed source code. IEEE Transactions on Machine Learning Research 2023 (2023).","journal-title":"IEEE Transactions on Machine Learning Research"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3643735"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.5555\/3304222.3304348"},{"key":"e_1_3_2_25_2","unstructured":"Raymond Li Loubna Ben Allal Yangtian Zi Niklas Muennighoff Denis Kocetkov Chenghao Mou Marc Marone Christopher Akiki Jia Li Jenny Chim et al. 2023. StarCoder: May the source be with you! IEEE Transactions on Machine Learning Research 2023 (2023)."},{"key":"e_1_3_2_26_2","volume-title":"Proceedings of the 4th International Conference on Learning Representations","author":"Li Yujia","year":"2016","unstructured":"Yujia Li, Daniel Tarlow, Marc Brockschmidt, and Richard S. Zemel. 2016. Gated graph sequence neural networks. In Proceedings of the 4th International Conference on Learning Representations."},{"issue":"2021","key":"e_1_3_2_27_2","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1016\/j.neucom.2021.07.051","article-title":"A survey of deep neural network watermarking techniques","volume":"461","author":"Li Yue","year":"2021","unstructured":"Yue Li, Hongxia Wang, and Mauro Barni. 2021. A survey of deep neural network watermarking techniques. Neurocomputing 461, 2021 (2021), 171\u2013193.","journal-title":"Neurocomputing"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3387904.3389261"},{"key":"e_1_3_2_29_2","first-page":"473","volume-title":"Proceedings of the International Conference on Automated Software Engineering","author":"Liu Fang","year":"2020","unstructured":"Fang Liu, Ge Li, Yunfei Zhao, and Zhi Jin. 2020. Multi-task learning based pre-trained language model for code completion. In Proceedings of the International Conference on Automated Software Engineering, 473\u2013485."},{"key":"e_1_3_2_30_2","first-page":"6227","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics","author":"Lu Shuai","year":"2022","unstructured":"Shuai Lu, Nan Duan, Hojae Han, Daya Guo, Seung-Won Hwang, and Alexey Svyatkovskiy. 2022. ReACC: A retrieval-augmented code completion framework. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 6227\u20136240."},{"key":"e_1_3_2_31_2","volume-title":"Proceedings of the 26th Annual Network and Distributed System Security Symposium","author":"Meli Michael","year":"2019","unstructured":"Michael Meli, Matthew R. McNiece, and Bradley Reaves. 2019. How bad can it git? Characterizing secret leakage in public GitHub repositories. In Proceedings of the 26th Annual Network and Distributed System Security Symposium."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.measurement.2018.12.030"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13735-018-0147-1"},{"key":"e_1_3_2_34_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Nijkamp Erik","year":"2023","unstructured":"Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023. CodeGen: An open large language model for code with multi-turn program synthesis. In Proceedings of the 11th International Conference on Learning Representations."},{"key":"e_1_3_2_35_2","first-page":"2133","volume-title":"Proceedings of the 32nd USENIX Security Symposium","author":"Niu Liang","year":"2023","unstructured":"Liang Niu, Muhammad Shujaat Mirza, Zayd Maradni, and Christina P\u00f6pper. 2023. CodexLeaks: Privacy leaks from code generation language models in GitHub copilot. In Proceedings of the 32nd USENIX Security Symposium, 2133\u20132150."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cose.2023.103102"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/2594291.2594321"},{"key":"e_1_3_2_38_2","unstructured":"Baptiste Rozi\u00e8re Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing Ellen Tan Yossi Adi Jingyu Liu Tal Remez J\u00e9r\u00e9my Rapin et al. 2023. Code Llama: Open foundation models for code. arXiv:2308.12950. Retrieved from https:\/\/arxiv.org\/abs\/2308.12950"},{"key":"e_1_3_2_39_2","volume-title":"Proceedings of the 26th Annual Network and Distributed System Security Symposium","author":"Salem Ahmed","year":"2019","unstructured":"Ahmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang, Mario Fritz, and Michael Backes. 2019. ML-Leaks: Model and data independent membership inference attacks and defenses on machine learning models. In Proceedings of the 26th Annual Network and Distributed System Security Symposium."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2017.41"},{"issue":"9","key":"e_1_3_2_41_2","first-page":"165","article-title":"A survey of digital watermarking techniques, applications and attacks","volume":"2","author":"Singh Prabhishek","year":"2013","unstructured":"Prabhishek Singh and Ramneet Singh Chadha. 2013. A survey of digital watermarking techniques, applications and attacks. International Journal of Engineering and Innovative Technology (IJEIT) 2, 9 (2013), 165\u2013175.","journal-title":"International Journal of Engineering and Innovative Technology (IJEIT)"},{"key":"e_1_3_2_42_2","first-page":"2615","volume-title":"Proceedings of the 30th USENIX Security Symposium","author":"Song Liwei","year":"2021","unstructured":"Liwei Song and Prateek Mittal. 2021. Systematic evaluation of privacy risks of machine learning models. In Proceedings of the 30th USENIX Security Symposium, 2615\u20132632."},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3611643.3616297"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3485447.3512225"},{"key":"e_1_3_2_45_2","first-page":"1433","volume-title":"Proceedings of the 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering","author":"Svyatkovskiy Alexey","year":"2020","unstructured":"Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020. IntelliCode compose: Code generation using transformer. In Proceedings of the 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 1433\u20131443."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330699"},{"key":"e_1_3_2_47_2","first-page":"5998","volume-title":"Proceedings of Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of Advances in Neural Information Processing Systems, 5998\u20136008."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664597"},{"key":"e_1_3_2_49_2","volume-title":"Proceedings of 44th International Conference on Software Engineering, Companion Volume","author":"Wan Yao","year":"2022","unstructured":"Yao Wan, Yang He, Zhangqian Bi, Jianguo Zhang, Yulei Sui, Hongyu Zhang, Kazuma Hashimoto, Hai Jin, Guandong Xu, Caiming Xiong, et al. 2022. NaturalCC: An open-source toolkit for code intelligence. In Proceedings of 44th International Conference on Software Engineering, Companion Volume."},{"key":"e_1_3_2_50_2","doi-asserted-by":"crossref","first-page":"9","DOI":"10.18653\/v1\/2022.findings-acl.2","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (ACL \u201922)","author":"Wang Xin","year":"2022","unstructured":"Xin Wang, Yasheng Wang, Yao Wan, Fei Mi, Yitong Li, Pingyi Zhou, Jin Liu, Hao Wu, Xin Jiang, and Qun Liu. 2022. Compilable neural code generation with compiler feedback. In Proceedings of the Findings of the Association for Computational Linguistics (ACL \u201922), 9\u201319."},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.68"},{"key":"e_1_3_2_52_2","first-page":"14015","volume-title":"Proceedings of the 35th AAAI Conference on Artificial Intelligence, 33rd Conference on Innovative Applications of Artificial Intelligence, and the 11th Symposium on Educational Advances in Artificial Intelligence","author":"Wang Yanlin","year":"2021","unstructured":"Yanlin Wang and Hui Li. 2021. Code completion by modeling flattened abstract syntax trees as graphs. In Proceedings of the 35th AAAI Conference on Artificial Intelligence, 33rd Conference on Innovative Applications of Artificial Intelligence, and the 11th Symposium on Educational Advances in Artificial Intelligence, 14015\u201314023."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3520312.3534862"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2024.3482719"},{"key":"e_1_3_2_55_2","first-page":"856","volume-title":"Proceedings of the 46th International Conference on Software Engineering","author":"Yang Zhou","year":"2024","unstructured":"Zhou Yang, Zhipeng Zhao, Chenyu Wang, Jieke Shi, Dongsun Kim, Donggyun Han, and David Lo. 2024. Unveiling memorization in code models. In Proceedings of the 46th International Conference on Software Engineering, 856\u2013856."},{"key":"e_1_3_2_56_2","first-page":"268","volume-title":"Proceedings of the 31st Computer Security Foundations Symposium","author":"Yeom Samuel","year":"2018","unstructured":"Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha. 2018. Privacy risk in machine learning: Analyzing the connection to overfitting. In Proceedings of the 31st Computer Security Foundations Symposium, 268\u2013282."},{"key":"e_1_3_2_57_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Zhang Chiyuan","year":"2023","unstructured":"Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tram\u00e8r, and Nicholas Carlini. 2023. Counterfactual memorization in neural language models. In Proceedings of the Advances in Neural Information Processing Systems."}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3742785","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,13]],"date-time":"2026-02-13T14:36:04Z","timestamp":1770993364000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3742785"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,13]]},"references-count":56,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,3,31]]}},"alternative-id":["10.1145\/3742785"],"URL":"https:\/\/doi.org\/10.1145\/3742785","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,13]]},"assertion":[{"value":"2024-04-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}