{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,19]],"date-time":"2026-07-19T09:44:51Z","timestamp":1784454291436,"version":"3.55.0"},"reference-count":78,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2025,1,17]],"date-time":"2025-01-17T00:00:00Z","timestamp":1737072000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["No. 92267201, No. 62206042, and No. U23B2019"],"award-info":[{"award-number":["No. 92267201, No. 62206042, and No. U23B2019"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Joint Funds of Natural Science Foundation of Liaoning Province","award":["2023-MSBA-081"],"award-info":[{"award-number":["2023-MSBA-081"]}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["No. N2416012"],"award-info":[{"award-number":["No. N2416012"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>Pretrained language models have shown strong effectiveness in code-related tasks, such as code retrieval, code generation, code summarization, and code completion tasks. In this article, we propose COde assistaNt viA retrieval-augmeNted language model (CONAN), which aims to build a code assistant by mimicking the knowledge-seeking behaviors of humans during coding. Specifically, it consists of a code structure-aware retriever (CONAN-R) and a dual-view code representation-based retrieval-augmented generation model (CONAN-G). CONAN-R pretrains CodeT5 using Code-Documentation Alignment and Masked Entity Prediction tasks to make language models code structure-aware and learn effective representations for code snippets and documentation. Then CONAN-G designs a dual-view code representation mechanism for implementing a retrieval-augmented code generation model. CONAN-G regards the code documentation descriptions as prompts, which help language models better understand the code semantics. Our experiments show that CONAN achieves convincing performance on different code generation tasks and significantly outperforms previous retrieval augmented code generation models. Our further analyses show that CONAN learns tailored representations for both code snippets and documentation by aligning code-documentation data pairs and capturing structural semantics by masking and predicting entities in the code data. Additionally, the retrieved code snippets and documentation provide necessary information from both program language and natural language to assist the code generation process. CONAN can also be used as an assistant for Large Language Models (LLMs), providing LLMs with external knowledge in shorter code document lengths to improve their effectiveness on various code tasks. It shows the ability of CONAN to extract necessary information and help filter out the noise from retrieved code documents.<\/jats:p>","DOI":"10.1145\/3695868","type":"journal-article","created":{"date-parts":[[2024,9,16]],"date-time":"2024-09-16T13:01:16Z","timestamp":1726491676000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["Building a Coding Assistant via the Retrieval-Augmented Language Model"],"prefix":"10.1145","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3014-0875","authenticated-orcid":false,"given":"Xinze","family":"Li","sequence":"first","affiliation":[{"name":"Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-0720-1610","authenticated-orcid":false,"given":"Hanbin","family":"Wang","sequence":"additional","affiliation":[{"name":"Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0083-3224","authenticated-orcid":false,"given":"Zhenghao","family":"Liu","sequence":"additional","affiliation":[{"name":"Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6335-1076","authenticated-orcid":false,"given":"Shi","family":"Yu","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5408-3145","authenticated-orcid":false,"given":"Shuo","family":"Wang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2421-7098","authenticated-orcid":false,"given":"Yukun","family":"Yan","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1094-9335","authenticated-orcid":false,"given":"Yukai","family":"Fu","sequence":"additional","affiliation":[{"name":"Chinese Academy of Sciences, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7422-6254","authenticated-orcid":false,"given":"Yu","family":"Gu","sequence":"additional","affiliation":[{"name":"Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3171-8889","authenticated-orcid":false,"given":"Ge","family":"Yu","sequence":"additional","affiliation":[{"name":"Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,1,17]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"2655","volume-title":"Proceedings of NAACL-HLT","author":"Ahmad Wasi","year":"2021","unstructured":"Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Unified pre-training for program understanding and generation. In Proceedings of NAACL-HLT, 2655\u20132668."},{"key":"e_1_3_2_3_2","first-page":"207","volume-title":"Proceedings of MSR","author":"Allamanis Miltiadis","year":"2013","unstructured":"Miltiadis Allamanis and Charles Sutton. 2013. Mining source code repositories at massive scale using language modeling. In Proceedings of MSR, 207\u2013216."},{"key":"e_1_3_2_4_2","unstructured":"Jacob Austin Augustus Odena Maxwell Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie Cai Michael Terry Quoc Le and Charles Sutton. 2021. Program synthesis with large language models."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2007.70720"},{"key":"e_1_3_2_6_2","first-page":"513","volume-title":"Proceedings of CHI","author":"Brandt Joel","year":"2010","unstructured":"Joel Brandt, Mira Dontcheva, Marcos Weskamp, and Scott R. Klemmer. 2010. Example-centric programming: Integrating web search into the development environment. In Proceedings of CHI, 513\u2013522."},{"key":"e_1_3_2_7_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mohammad Bavarian Clemens Winter Philippe Tillet Felipe Petroski Such Dave Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William Hebgen Guss Alex Nichol Alex Paino Nikolas Tezak Jie Tang Igor Babuschkin Suchir Balaji Shantanu Jain William Saunders Christopher Hesse Andrew N. Carr Jan Leike Josh Achiam Vedant Misra Evan Morikawa Alec Radford Matthew Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating large language models trained on code."},{"key":"e_1_3_2_8_2","first-page":"15750","volume-title":"Proceedings of CVPR","author":"Chen Xinlei","year":"2021","unstructured":"Xinlei Chen and Kaiming He. 2021. Exploring simple Siamese representation learning. In Proceedings of CVPR, 15750\u201315758."},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of ICLR","author":"Clark Kevin","year":"2020","unstructured":"Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: Pre-training text encoders as discriminators rather than generators. In Proceedings of ICLR."},{"key":"e_1_3_2_10_2","first-page":"4171","volume-title":"Proceedings of NAACL-HLT","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, 4171\u20134186."},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Hongchao Fang Sicheng Wang Meng Zhou Jiayuan Ding and Pengtao Xie. 2020. Cert: Contrastive self-supervised learning for language understanding.","DOI":"10.36227\/techrxiv.12308378.v1"},{"key":"e_1_3_2_12_2","first-page":"1536","volume-title":"Proceedings of EMNLP Findings","author":"Feng Zhangyin","year":"2020","unstructured":"Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. CodeBERT: A pre-trained model for programming and natural languages. In Proceedings of EMNLP Findings, 1536\u20131547."},{"key":"e_1_3_2_13_2","volume-title":"Proceedings of ICLR","author":"Gao Jun","year":"2019","unstructured":"Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2019. Representation degeneration problem in training natural language generation models. In Proceedings of ICLR."},{"key":"e_1_3_2_14_2","first-page":"981","volume-title":"Proceedings of EMNLP","author":"Gao Luyu","year":"2021","unstructured":"Luyu Gao and Jamie Callan. 2021. Condenser: A pre-training architecture for dense retrieval. In Proceedings of EMNLP, 981\u2013993."},{"key":"e_1_3_2_15_2","first-page":"6894","volume-title":"Proceedings of EMNLP","author":"Gao Tianyu","year":"2021","unstructured":"Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of EMNLP, 6894\u20136910."},{"key":"e_1_3_2_16_2","first-page":"7212","volume-title":"Proceedings of ACL","author":"Guo Daya","year":"2022","unstructured":"Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. 2022. UniXcoder: Unified cross-modal pre-training for code representation. In Proceedings of ACL, 7212\u20137225."},{"key":"e_1_3_2_17_2","volume-title":"Proceedings of ICLR","author":"Guo Daya","year":"2021","unstructured":"Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou. 2021. GraphCodeBERT: Pre-training code representations with data flow. In Proceedings of ICLR."},{"key":"e_1_3_2_18_2","unstructured":"Daya Guo Qihao Zhu Dejian Yang Zhenda Xie Kai Dong Wentao Zhang Guanting Chen Xiao Bi Y. Wu Y. K. Li Fuli Luo Yingfei Xiong and Wenfeng Liang. 2024. DeepSeek-Coder: When the large language model meets programming \u2013 The rise of code intelligence."},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","unstructured":"Yucan Guo Zixuan Li Xiaolong Jin Yantao Liu Yutao Zeng Wenxuan Liu Xiang Li Pan Yang Long Bai Jiafeng Guo and Xueqi Cheng. 2023. Retrieval-augmented code generation for universal information extraction.","DOI":"10.1007\/978-981-97-9434-8_3"},{"key":"e_1_3_2_20_2","first-page":"3929","volume-title":"Proceedings of ICML","author":"Guu Kelvin","year":"2020","unstructured":"Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Retrieval augmented language model pre-training. In Proceedings of ICML, 3929\u20133938."},{"key":"e_1_3_2_21_2","unstructured":"Hamel Husain Ho-Hsiang Wu Tiferet Gazit Miltiadis Allamanis and Marc Brockschmidt. 2020. CodeSearchNet challenge: Evaluating the state of semantic code search."},{"key":"e_1_3_2_22_2","first-page":"1643","volume-title":"Proceedings of EMNLP","author":"Iyer Srinivasan","year":"2018","unstructured":"Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2018. Mapping language to code in programmatic context. In Proceedings of EMNLP, 1643\u20131652."},{"key":"e_1_3_2_23_2","first-page":"874","volume-title":"Proceedings of EACL","author":"Izacard Gautier","year":"2021","unstructured":"Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of EACL, 874\u2013880."},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"Zhengbao Jiang Frank F. Xu Luyu Gao Zhiqing Sun Qian Liu Jane Dwivedi-Yu Yiming Yang Jamie Callan and Graham Neubig. 2023. Active retrieval augmented generation.","DOI":"10.18653\/v1\/2023.emnlp-main.495"},{"key":"e_1_3_2_25_2","first-page":"535","article-title":"Billion-scale similarity search with GPUs","volume":"21","author":"Johnson Jeff","year":"2019","unstructured":"Jeff Johnson, Matthijs Douze, and Herv\u00e9 J\u00e9gou. 2019. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data 21 (2019), 535\u2013547.","journal-title":"IEEE Transactions on Big Data"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.550"},{"key":"e_1_3_2_27_2","first-page":"14967","volume-title":"Proceedings of NeurIPS","author":"Lachaux Marie-Anne","year":"2021","unstructured":"Marie-Anne Lachaux, Baptiste Rozi\u00e8re, Marc Szafraniec, and Guillaume Lample. 2021. DOBF: A deobfuscation pre-training objective for programming languages. In Proceedings of NeurIPS, 14967\u201314979."},{"key":"e_1_3_2_28_2","volume-title":"Proceedings of NeurIPS","author":"Lewis Patrick S. H.","year":"2020","unstructured":"Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K\u00fcttler, Mike Lewis, Wen-tau Yih, Tim Rockt\u00e4schel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Proceedings of NeurIPS."},{"key":"e_1_3_2_29_2","first-page":"9119","volume-title":"Proceedings of EMNLP","author":"Li Bohan","year":"2020","unstructured":"Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020. On the sentence embeddings from pre-trained language models. In Proceedings of EMNLP, 9119\u20139130."},{"key":"e_1_3_2_30_2","first-page":"142","volume-title":"Proceedings of WCRE","author":"Li Hongwei","year":"2013","unstructured":"Hongwei Li, Zhenchang Xing, Xin Peng, and Wenyun Zhao. 2013. What help do developers seek, when and how? In Proceedings of WCRE, 142\u2013151."},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","unstructured":"Jia Li Yongmin Li Ge Li Zhi Jin Yiyang Hao and Xing Hu. 2023. SkCoder: A sketch-based approach for automatic code generation.","DOI":"10.1109\/ICSE48619.2023.00179"},{"key":"e_1_3_2_32_2","first-page":"2898","volume-title":"Proceedings of EMNLP","author":"Li Xiaonan","year":"2022","unstructured":"Xiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu, Hang Zhang, Bolun Yao, Weizhen Qi, Daxin Jiang, Weizhu Chen, and Nan Duan. 2022. CodeRetriever: A large scale contrastive pre-training method for code search. In Proceedings of EMNLP, 2898\u20132910."},{"key":"e_1_3_2_33_2","first-page":"11560","volume-title":"Proceedings of ACL","author":"Li Xinze","year":"2023","unstructured":"Xinze Li, Zhenghao Liu, Chenyan Xiong, Shi Yu, Yu Gu, Zhiyuan Liu, and Ge Yu. 2023. Structure-aware language model pretraining improves dense retrieval on structured data. In Proceedings of ACL, 11560\u201311574."},{"key":"e_1_3_2_34_2","first-page":"287","volume-title":"Proceedings of SIGIR","author":"Li Yizhi","year":"2021","unstructured":"Yizhi Li, Zhenghao Liu, Chenyan Xiong, and Zhiyuan Liu. 2021. More robust dense retrieval with contrastive dual learning. In Proceedings of SIGIR, 287\u2013296."},{"key":"e_1_3_2_35_2","unstructured":"Dianshu Liao Shidong Pan Qing Huang Xiaoxue Ren Zhenchang Xing Huan Jin and Qinying Li. 2023. Context-aware code generation framework for code repositories: Local global and third-party library awareness."},{"key":"e_1_3_2_36_2","first-page":"501","volume-title":"Proceedings of COLING","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin and Franz Josef Och. 2004. ORANGE: A method for evaluating automatic evaluation metrics for machine translation. In Proceedings of COLING, 501\u2013507."},{"key":"e_1_3_2_37_2","volume-title":"Proceedings of ICLR","author":"Liu Shangqing","year":"2021","unstructured":"Shangqing Liu, Yu Chen, Xiaofei Xie, Jing Kai Siow, and Yang Liu. 2021. Retrieval-augmented generation for code summarization via hybrid GNN. In Proceedings of ICLR."},{"key":"e_1_3_2_38_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach."},{"key":"e_1_3_2_39_2","volume-title":"Proceedings of ICLR","author":"Liu Zhenghao","year":"2023","unstructured":"Zhenghao Liu, Chenyan Xiong, Yuanhuiyi Lv, Zhiyuan Liu, and Ge Yu. 2023. Universal vision-language dense retrieval: Learning a unified representation space for multi-modal retrieval. In Proceedings of ICLR."},{"key":"e_1_3_2_40_2","first-page":"6227","volume-title":"Proceedings of ACL","author":"Lu Shuai","year":"2022","unstructured":"Shuai Lu, Nan Duan, Hojae Han, Daya Guo, Seung-won Hwang, and Alexey Svyatkovskiy. 2022. ReACC: A retrieval-augmented code completion framework. In Proceedings of ACL, 6227\u20136240."},{"key":"e_1_3_2_41_2","volume-title":"Proceedings of NeurIPS","author":"Lu Shuai","year":"2021","unstructured":"Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021. CodeXGLUE: A machine learning benchmark dataset for code understanding and generation. In Proceedings of NeurIPS."},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","first-page":"329","DOI":"10.1162\/tacl_a_00369","article-title":"Sparse, dense, and attentional representations for text retrieval","author":"Luan Yi","year":"2021","unstructured":"Yi Luan, Jacob Eisenstein, Kristina Toutanova, and Michael Collins. 2021. Sparse, dense, and attentional representations for text retrieval. Transactions of the Association for Computational Linguistics (2021), 329\u2013345.","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"e_1_3_2_43_2","unstructured":"Hongyin Luo Yung-Sung Chuang Yuan Gong Tianhua Zhang Yoon Kim Xixin Wu Danny Fox Helen Meng and James Glass. 2023. SAIL: Search-augmented instruction learning."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1166"},{"key":"e_1_3_2_45_2","first-page":"1735","article-title":"Long short-term memory","volume":"8","author":"Hochreiter Sepp","year":"2010","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber. 2010. Long short-term memory. Neural Computation 8 (2010), 1735\u20131780.","journal-title":"Neural Computation"},{"key":"e_1_3_2_46_2","first-page":"23102","volume-title":"Proceedings of NeurIPS","author":"Meng Yu","year":"2021","unstructured":"Yu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary, Paul Bennett, Jiawei Han, and Xia Song. 2021. COCO-LM: Correcting and contrasting text sequences for language model pretraining. In Proceedings of NeurIPS, 23102\u201323114."},{"key":"e_1_3_2_47_2","unstructured":"OpenAI. 2022. ChatGPT: Optimizing Language Models for Dialogue. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2021\/hash\/c2c2a04512b35d13102459f8784f1a2d-Abstract.html"},{"key":"e_1_3_2_48_2","first-page":"311","volume-title":"Proceedings of ACL","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of ACL, 311\u2013318."},{"key":"e_1_3_2_49_2","first-page":"2719","volume-title":"Proceedings of EMNLP Findings","author":"Parvez Md Rizwan","year":"2021","unstructured":"Md Rizwan Parvez, Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Retrieval augmented code generation and summarization. In Proceedings of EMNLP Findings, 2719\u20132734."},{"key":"e_1_3_2_50_2","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 8 (2019) 9."},{"key":"e_1_3_2_51_2","first-page":"140:1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21 (2020), 140:1\u2013140:67.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3022671.2984041"},{"key":"e_1_3_2_53_2","first-page":"3982","volume-title":"Proceedings of EMNLP","author":"Reimers Nils","year":"2019","unstructured":"Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of EMNLP, 3982\u20133992."},{"key":"e_1_3_2_54_2","unstructured":"Shuo Ren Daya Guo Shuai Lu Long Zhou Shujie Liu Duyu Tang Neel Sundaresan Ming Zhou Ambrosio Blanco and Shuai Ma. 2020. CodeBLEU: A method for automatic evaluation of code synthesis."},{"key":"e_1_3_2_55_2","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1561\/1500000019","article-title":"The probabilistic relevance framework: BM25 and beyond","volume":"4","author":"Robertson Stephen","year":"2009","unstructured":"Stephen Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: BM25 and beyond. Foundations and Trends\u00ae in Information Retrieval 4 (2009), 333\u2013389.","journal-title":"Foundations and Trends\u00ae in Information Retrieval"},{"key":"e_1_3_2_56_2","first-page":"81","volume-title":"Proceedings of WCRE","author":"Roy Chanchal K.","year":"2008","unstructured":"Chanchal K. Roy and James R. Cordy. 2008. An empirical study of function clones in open source software. In Proceedings of WCRE, 81\u201390."},{"key":"e_1_3_2_57_2","unstructured":"Baptiste Roziere Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing Ellen Tan Yossi Adi Jingyu Liu Romain Sauvestre Tal Remez J\u00e9r\u00e9my Rapin Artyom Kozhevnikov Ivan Evtimov Joanna Bitton Manish Bhatt Cristian Canton Ferrer Aaron Grattafiori Wenhan Xiong Alexandre D\u00e9fossez Jade Copet Faisal Azhar Hugo Touvron Louis Martin Nicolas Usunier Thomas Scialom and Gabriel Synnaeve. 2023. Code Llama: Open foundation models for code."},{"key":"e_1_3_2_58_2","first-page":"191","volume-title":"Proceedings of FSE","author":"Sadowski Caitlin","year":"2015","unstructured":"Caitlin Sadowski, Kathryn T. Stolee, and Sebastian Elbaum. 2015. How developers search for code: A case study. In Proceedings of FSE, 191\u2013201."},{"key":"e_1_3_2_59_2","first-page":"6138","volume-title":"Proceedings of EMNLP","author":"Sciavolino Christopher","year":"2021","unstructured":"Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen. 2021. Simple entity-centric questions challenge dense retrievers. In Proceedings of EMNLP, 6138\u20136148."},{"key":"e_1_3_2_60_2","unstructured":"Anton Shapkin Denis Litvinov and Timofey Bryksin. 2023. Entity-augmented code generation."},{"key":"e_1_3_2_61_2","unstructured":"Disha Shrivastava Denis Kocetkov Harm de Vries Dzmitry Bahdanau and Torsten Scholak. 2023. RepoFusion: Training code models to understand your repository."},{"key":"e_1_3_2_62_2","first-page":"131","volume-title":"Proceedings of ICSME","author":"Svajlenko Jeffrey","year":"2015","unstructured":"Jeffrey Svajlenko and Chanchal K. Roy. 2015. Evaluating clone detection tools with BigCloneBench. In Proceedings of ICSME, 131\u2013140."},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","unstructured":"Qwen Team. 2024. Code with CodeQwen1.5. DOI: 10.1109\/ICSM.2015.7332459","DOI":"10.1109\/ICSM.2015.7332459"},{"key":"e_1_3_2_64_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale Dan Bikel Lukas Blecher Cristian Canton Ferrer Moya Chen Guillem Cucurull David Esiobu Jude Fernandes Jeremy Fu Wenyin Fu Brian Fuller Cynthia Gao Vedanuj Goswami Naman Goyal Anthony Hartshorn Saghar Hosseini Rui Hou Hakan Inan Marcin Kardas Viktor Kerkez Madian Khabsa Isabel Kloumann Artem Korenev Punit Singh Koura Marie-Anne Lachaux Thibaut Lavril Jenya Lee Diana Liskovich Yinghai Lu Yuning Mao Xavier Martinet Todor Mihaylov Pushkar Mishra Igor Molybog Yixin Nie Andrew Poulton Jeremy Reizenstein Rashi Rungta Kalyan Saladi Alan Schelten Ruan Silva Eric Michael Smith Ranjan Subramanian Xiaoqing Ellen Tan Binh Tang Ross Taylor Adina Williams Jian Xiang Kuan Puxin Xu Zheng Yan Iliyan Zarov Yuchen Zhang Angela Fan Melanie Kambadur Sharan Narang Aurelien Rodriguez Robert Stojnic Sergey Edunov and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models."},{"key":"e_1_3_2_65_2","first-page":"5998","volume-title":"Proceedings of NeurIPS","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of NeurIPS, 5998\u20136008."},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","unstructured":"Yue Wang Hung Le Akhilesh Deepak Gotmare Nghi D. Q. Bui Junnan Li and Steven C. H. Hoi. 2023. CodeT5+: Open code large language models for code understanding and generation.","DOI":"10.18653\/v1\/2023.emnlp-main.68"},{"key":"e_1_3_2_67_2","first-page":"8696","volume-title":"Proceedings of EMNLP","author":"Wang Yue","year":"2021","unstructured":"Yue Wang, Weishi Wang, Shafiq Joty, and Steven C.H. Hoi. 2021. CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. In Proceedings of EMNLP, 8696\u20138708."},{"key":"e_1_3_2_68_2","doi-asserted-by":"crossref","unstructured":"Thomas Wolf Lysandre Debut Victor Sanh Julien Chaumond Clement Delangue Anthony Moi Pierric Cistac Tim Rault R\u00e9mi Louf Morgan Funtowicz Joe Davison Sam Shleifer Patrick von Platen Clara Ma Yacine Jernite Julien Plu Canwen Xu Teven Le Scao Sylvain Gugger Mariama Drame Quentin Lhoest and Alexander M. Rush. 2020. HuggingFace\u2019s transformers: State-of-the-art natural language processing.","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"e_1_3_2_69_2","unstructured":"Zhuofeng Wu Sinong Wang Jiatao Gu Madian Khabsa Fei Sun and Hao Ma. 2020. Clear: Contrastive learning for sentence representation."},{"key":"e_1_3_2_70_2","volume-title":"Proceedings of ICLR","author":"Xiong Lee","year":"2021","unstructured":"Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate nearest neighbor negative contrastive learning for dense text retrieval. In Proceedings of ICLR."},{"key":"e_1_3_2_71_2","volume-title":"Proceedings of ICLR","author":"Xiong Wenhan","year":"2021","unstructured":"Wenhan Xiong, Xiang Lorraine Li, Srini Iyer, Jingfei Du, Patrick S. H. Lewis, William Yang Wang, Yashar Mehdad, Scott Yih, Sebastian Riedel, Douwe Kiela, and Barlas Oguz. 2021. Answering complex open-domain questions with multi-hop dense retrieval. In Proceedings of ICLR."},{"key":"e_1_3_2_72_2","first-page":"5065","volume-title":"Proceedings of ACL","author":"Yan Yuanmeng","year":"2021","unstructured":"Yuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang, Wei Wu, and Weiran Xu. 2021. ConSERT: A contrastive framework for self-supervised sentence representation transfer. In Proceedings of ACL, 5065\u20135075."},{"key":"e_1_3_2_73_2","first-page":"7170","volume-title":"Proceedings of EMNLP","author":"Ye Deming","year":"2020","unstructured":"Deming Ye, Yankai Lin, Jiaju Du, Zhenghao Liu, Peng Li, Maosong Sun, and Zhiyuan Liu. 2020. Coreferential reasoning learning for language representation. In Proceedings of EMNLP, 7170\u20137186."},{"key":"e_1_3_2_74_2","volume-title":"Proceedings of SIGIR","author":"Yu Shi","year":"2021","unstructured":"Shi Yu, Zhenghao Liu, Chenyan Xiong, Tao Feng, and Zhiyuan Liu. 2021. Few-shot conversational dense retrieval. In Proceedings of SIGIR."},{"key":"e_1_3_2_75_2","first-page":"3160","volume-title":"Proceedings of SIGIR","author":"Yu Shi","year":"2023","unstructured":"Shi Yu, Zhenghao Liu, Chenyan Xiong, and Zhiyuan Liu. 2023. OpenMatch-v2: An all-in-one multi-modality PLM-based information retrieval toolkit. In Proceedings of SIGIR, 3160\u20133164."},{"key":"e_1_3_2_76_2","first-page":"3160","volume-title":"Proceedings of IJCAI","author":"Zan Daoguang","year":"2022","unstructured":"Daoguang Zan, Bei Chen, Dejian Yang, Zeqi Lin, Minsu Kim, Bei Guan, Yongji Wang, Weizhu Chen, and Jian-Guang Lou. 2022. CERT: Continual pre-training on sketches for library-oriented code generation. In Proceedings of IJCAI, 3160\u20133164."},{"key":"e_1_3_2_77_2","doi-asserted-by":"crossref","unstructured":"Fengji Zhang Bei Chen Yue Zhang Jacky Keung Jin Liu Daoguang Zan Yi Mao Jian-Guang Lou and Weizhu Chen. 2023. RepoCoder: Repository-level code completion through iterative retrieval and generation.","DOI":"10.18653\/v1\/2023.emnlp-main.151"},{"key":"e_1_3_2_78_2","first-page":"1441","volume-title":"Proceedings of ACL","author":"Zhang Zhengyan","year":"2019","unstructured":"Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019. ERNIE: Enhanced language representation with informative entities. In Proceedings of ACL, 1441\u20131451."},{"key":"e_1_3_2_79_2","first-page":"1441","volume-title":"Proceedings of ICLR","author":"Zhou Shuyan","year":"2022","unstructured":"Shuyan Zhou, Uri Alon, Frank F. Xu, Zhengbao Jiang, and Graham Neubig. 2022. DocPrompting: Generating code by retrieving the docs. In Proceedings of ICLR, 1441\u20131451."}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3695868","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3695868","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:04:29Z","timestamp":1750291469000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3695868"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,17]]},"references-count":78,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3695868"],"URL":"https:\/\/doi.org\/10.1145\/3695868","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,17]]},"assertion":[{"value":"2024-01-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-08-31","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-17","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}