{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T18:36:29Z","timestamp":1784226989442,"version":"3.55.0"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","funder":[{"name":"Zhejiang Provincial Natural Science Foundation of China","award":["LZ25F020003"],"award-info":[{"award-number":["LZ25F020003"]}]},{"name":"National Natural Science Foundation of China","award":["62202420, 62202074"],"award-info":[{"award-number":["62202420, 62202074"]}]},{"name":"National Research Foundation Singapore","award":["NRF-NRFI08-2022-0002"],"award-info":[{"award-number":["NRF-NRFI08-2022-0002"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,19]]},"abstract":"<jats:p>Code search is a crucial task in software engineering, aiming to retrieve code snippets that are semantically relevant to a natural language query. Recently, Pre-trained Language Models (PLMs) have shown remarkable success and are widely adopted for code search tasks. However, PLM-based methods often struggle in cross-domain scenarios. When applied to a new domain, they typically require extensive fine-tuning with substantial data. Even worse, the data scarcity problem in new domains often forces these methods to operate in a zero-shot setting, resulting in a significant decline in performance. RAPID, which generates synthetic data for model fine-tuning, is currently the only effective method for zero-shot cross-domain code search. Despite its effectiveness, RAPID demands substantial computational resources for fine-tuning and needs to maintain specialized models for each domain, underscoring the need for a zero-shot, fine-tuning-free approach for cross-domain code search.<\/jats:p>\n          <jats:p>The key to tackling zero-shot cross-domain code search lies in bridging the gaps among domains. In this work, we propose to break the query-code matching process of code search into two simpler tasks: query-comment matching and code-code matching. We first conduct an empirical study to investigate the effectiveness of these two matching schemas in zero-shot cross-domain code search. Our findings highlight the strong complementarity among the three matching schemas, i.e., query-code, query-comment, and code-code matching. Based on the findings, we propose CodeBridge, a zero-shot, fine-tuning-free approach for cross-domain code search. Specifically, CodeBridge first employs zero-shot prompting to guide Large Language Models (LLMs) to generate a comment for each code snippet in the codebase and produce a code for each query. Subsequently, it encodes queries, code snippets, comments, and the generated code using PLMs and assesses similarities through three matching schemas: query-code, query-comment, and generated code-code. Lastly, CodeBridge leverages a sampling-based fusion approach that combines these three similarity scores to rank the final search outcomes. Experimental results show that our approach outperforms the state-of-the-art PLM-based code search approaches, i.e., CoCoSoDa and UniXcoder, by an average of 21.4% and 24.9% in MRR, respectively, across three datasets. Our approach also yields results that are better than or comparable to those of the zero-shot cross-domain code search approach RAPID, which requires costly fine-tuning.<\/jats:p>","DOI":"10.1145\/3729357","type":"journal-article","created":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T15:15:34Z","timestamp":1750346134000},"page":"1937-1959","source":"Crossref","is-referenced-by-count":3,"title":["Zero-Shot Cross-Domain Code Search without Fine-Tuning"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-4613-247X","authenticated-orcid":false,"given":"Keyu","family":"Liang","sequence":"first","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1981-1626","authenticated-orcid":false,"given":"Zhongxin","family":"Liu","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8283-9146","authenticated-orcid":false,"given":"Chao","family":"Liu","sequence":"additional","affiliation":[{"name":"Chongqing University, Chongqing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7657-6653","authenticated-orcid":false,"given":"Zhiyuan","family":"Wan","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4367-7201","authenticated-orcid":false,"given":"David","family":"Lo","sequence":"additional","affiliation":[{"name":"Singapore Management University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4111-4189","authenticated-orcid":false,"given":"Xiaohu","family":"Yang","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,19]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Diogo Almeida, Janko Altenschmidt, Sam Altman, and Shyamal Anadkat.","author":"Achiam Josh","year":"2023","unstructured":"Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, and Shyamal Anadkat. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774."},{"key":"e_1_2_1_2_1","volume-title":"Proceedings of the ACM\/IEEE 42nd International Conference on Software Engineering. 590\u2013601","author":"Aghajani Emad","year":"2020","unstructured":"Emad Aghajani, Csaba Nagy, Mario Linares-V\u00e1squez, Laura Moreno, Gabriele Bavota, Michele Lanza, and David C Shepherd. 2020. Software documentation: the practitioners\u2019 perspective. In Proceedings of the ACM\/IEEE 42nd International Conference on Software Engineering. 590\u2013601. https:\/\/doi.org\/10.1145\/3377811.3380405 10.1145\/3377811.3380405"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.scico.2012.04.008"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the 22nd Australasian Document Computing Symposium. 1\u20138. https:\/\/doi.org\/10","author":"Benham Rodger","year":"2017","unstructured":"Rodger Benham and J Shane Culpepper. 2017. Risk-reward trade-offs in rank fusion. In Proceedings of the 22nd Australasian Document Computing Symposium. 1\u20138. https:\/\/doi.org\/10.1145\/3166072.3166084 10.1145\/3166072.3166084"},{"key":"e_1_2_1_5_1","unstructured":"Bing. 2024. Bing Search Engine. https:\/\/bing.com\/ [Accessed 26-08-2024]"},{"key":"e_1_2_1_6_1","volume-title":"Histoire de l\u2019Acad\u00e9mie royale des sciences pour 1781.","author":"Borda JC","year":"1953","unstructured":"JC Borda. 1784. M\u00e9moire sur les \u00e9lections au scrutin, Histoire de l\u2019Acad\u00e9mie royale des sciences pour 1781. Paris (English Transl. by Grazia, A. 1953. Isis 44)."},{"key":"e_1_2_1_7_1","volume-title":"Seventh European Conference onSoftware Maintenance and Reengineering, 2003. Proceedings.. 13\u201315","author":"Briand Lionel C","year":"2003","unstructured":"Lionel C Briand. 2003. Software documentation: how much is enough? In Seventh European Conference onSoftware Maintenance and Reengineering, 2003. Proceedings.. 13\u201315. https:\/\/doi.org\/10.1109\/csmr.2003.1192406 10.1109\/csmr.2003.1192406"},{"key":"e_1_2_1_8_1","volume-title":"Language models are few-shot learners. Advances in neural information processing systems, 33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, and Amanda Askell. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33 (2020), 1877\u20131901."},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the 44th International Conference on Software Engineering. 487\u2013498","author":"Chai Yitian","year":"2022","unstructured":"Yitian Chai, Hongyu Zhang, Beijun Shen, and Xiaodong Gu. 2022. Cross-domain deep code search with meta learning. In Proceedings of the 44th International Conference on Software Engineering. 487\u2013498. https:\/\/doi.org\/10.1145\/3510003.3510125 10.1145\/3510003.3510125"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the 30th IEEE\/ACM International Conference on Program Comprehension. 533\u2013542","author":"Cheng Yi","year":"2022","unstructured":"Yi Cheng and Li Kuang. 2022. CSRS: code search with relevance matching and semantic matching. In Proceedings of the 30th IEEE\/ACM International Conference on Program Comprehension. 533\u2013542. https:\/\/doi.org\/10.1145\/3524610.3527889 10.1145\/3524610.3527889"},{"key":"e_1_2_1_11_1","unstructured":"CodeBridge. 2024. Replication Package. https:\/\/github.com\/ZJU-CTAG\/CodeBridge"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval. 758\u2013759","author":"Cormack Gordon V","year":"2009","unstructured":"Gordon V Cormack, Charles LA Clarke, and Stefan Buettcher. 2009. Reciprocal rank fusion outperforms condorcet and individual rank learning methods. In Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval. 758\u2013759. https:\/\/doi.org\/10.1145\/1571941.1572114 10.1145\/1571941.1572114"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3641542"},{"key":"e_1_2_1_14_1","first-page":"139","volume-title":"CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Feng Zhangyin","year":"2020","unstructured":"Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, and Daxin Jiang. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP 2020. 1536\u20131547. https:\/\/doi.org\/10.18653\/v1\/2020.findings-emnlp.139 10.18653\/v1\/2020.findings-emnlp.139"},{"key":"e_1_2_1_15_1","volume-title":"International conference on machine learning. 1126\u20131135","author":"Finn Chelsea","year":"2017","unstructured":"Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning. 1126\u20131135."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","unstructured":"Edward Fox and Joseph Shaw. 1994. Combination of multiple searches. NIST special publication SP 243\u2013243. https:\/\/doi.org\/10.6028\/nist.sp.500-225.vpi 10.6028\/nist.sp.500-225.vpi","DOI":"10.6028\/nist.sp.500-225.vpi"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the 46th IEEE\/ACM International Conference on Software Engineering. 1\u201313","author":"Geng Mingyang","year":"2024","unstructured":"Mingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang, Ge Li, Zhi Jin, Xiaoguang Mao, and Xiangke Liao. 2024. Large language models are few-shot summarizers: Multi-intent comment generation via in-context learning. In Proceedings of the 46th IEEE\/ACM International Conference on Software Engineering. 1\u201313. https:\/\/doi.org\/10.1145\/3597503.3608134 10.1145\/3597503.3608134"},{"key":"e_1_2_1_18_1","unstructured":"Github. 2024. Github website. https:\/\/github.com\/ [Accessed 26-08-2024]"},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the 40th International Conference on Software Engineering. 933\u2013944","author":"Gu Xiaodong","year":"2018","unstructured":"Xiaodong Gu, Hongyu Zhang, and Sunghun Kim. 2018. Deep code search. In Proceedings of the 40th International Conference on Software Engineering. 933\u2013944. https:\/\/doi.org\/10.1145\/3180155.3180167 10.1145\/3180155.3180167"},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 7212\u20137225","author":"Guo Daya","year":"2022","unstructured":"Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. 2022. UniXcoder: Unified Cross-Modal Pre-training for Code Representation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 7212\u20137225. https:\/\/doi.org\/10.18653\/v1\/2022.acl-long.499 10.18653\/v1\/2022.acl-long.499"},{"key":"e_1_2_1_21_1","volume-title":"Graphcodebert: Pre-training code representations with data flow. arXiv preprint arXiv:2009.08366.","author":"Guo Daya","year":"2020","unstructured":"Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, and Shengyu Fu. 2020. Graphcodebert: Pre-training code representations with data flow. arXiv preprint arXiv:2009.08366."},{"key":"e_1_2_1_22_1","unstructured":"Daya Guo Qihao Zhu Dejian Yang Zhenda Xie Kai Dong Wentao Zhang Guanting Chen Xiao Bi Yu Wu and YK Li. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming\u2013The Rise of Code Intelligence. arXiv preprint arXiv:2401.14196."},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 5690\u20135700","author":"Huang Junjie","year":"2021","unstructured":"Junjie Huang, Duyu Tang, Linjun Shou, Ming Gong, Ke Xu, Daxin Jiang, Ming Zhou, and Nan Duan. 2021. CoSQA: 20,000+ Web Queries for Code Search and Question Answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 5690\u20135700. https:\/\/doi.org\/10.18653\/v1\/2021.acl-long.442 10.18653\/v1\/2021.acl-long.442"},{"key":"e_1_2_1_24_1","unstructured":"Hamel Husain Ho-Hsiang Wu Tiferet Gazit Miltiadis Allamanis and Marc Brockschmidt. 2019. Codesearchnet challenge: Evaluating the state of semantic code search. arXiv preprint arXiv:1909.09436."},{"key":"e_1_2_1_25_1","volume-title":"TIOBE Index for","author":"Index TIOBE","year":"2025","unstructured":"TIOBE Index. 2025. TIOBE Index for January 2025. https:\/\/www.tiobe.com\/tiobe-index\/ [Accessed 08-02-2025]"},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the 29th Symposium on Operating Systems Principles. 611\u2013626","author":"Kwon Woosuk","year":"2023","unstructured":"Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles. 611\u2013626. https:\/\/doi.org\/10.1145\/3600006.3613165 10.1145\/3600006.3613165"},{"key":"e_1_2_1_27_1","volume-title":"Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33","author":"Lewis Patrick","year":"2020","unstructured":"Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K\u00fcttler, Mike Lewis, Wen-tau Yih, and Tim Rockt\u00e4schel. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33 (2020), 9459\u20139474."},{"key":"e_1_2_1_28_1","volume-title":"Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, and Jenny Chim.","author":"Li Raymond","year":"2023","unstructured":"Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, and Jenny Chim. 2023. Starcoder: may the source be with you!. arXiv preprint arXiv:2305.06161."},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2898\u20132910","author":"Li Xiaonan","year":"2022","unstructured":"Xiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu, Hang Zhang, Bolun Yao, Weizhen Qi, Daxin Jiang, Weizhu Chen, and Nan Duan. 2022. Coderetriever: A large scale contrastive pre-training method for code search. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2898\u20132910. https:\/\/doi.org\/10.18653\/v1\/2022.emnlp-main.187 10.18653\/v1\/2022.emnlp-main.187"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3480027"},{"key":"e_1_2_1_31_1","first-page":"1","article-title":"Codematcher: Searching code based on sequential semantics of important query words","volume":"31","author":"Liu Chao","year":"2021","unstructured":"Chao Liu, Xin Xia, David Lo, Zhiwe Liu, Ahmed E Hassan, and Shanping Li. 2021. Codematcher: Searching code based on sequential semantics of important query words. ACM Transactions on Software Engineering and Methodology (TOSEM), 31, 1 (2021), 1\u201337.","journal-title":"ACM Transactions on Software Engineering and Methodology (TOSEM)"},{"key":"e_1_2_1_32_1","unstructured":"Steve Lohr. 2012. For impatient web users an eye blink is just too long to wait. The New York Times A1\u2013L."},{"key":"e_1_2_1_33_1","volume-title":"Wizardcoder: Empowering code large language models with evol-instruct. arXiv preprint arXiv:2306.08568.","author":"Luo Ziyang","year":"2023","unstructured":"Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023. Wizardcoder: Empowering code large language models with evol-instruct. arXiv preprint arXiv:2306.08568."},{"key":"e_1_2_1_34_1","volume-title":"2015 30th IEEE\/ACM International Conference on Automated Software Engineering (ASE). 260\u2013270","author":"Lv Fei","year":"2015","unstructured":"Fei Lv, Hongyu Zhang, Jian-guang Lou, Shaowei Wang, Dongmei Zhang, and Jianjun Zhao. 2015. Codehow: Effective code search based on api understanding and extended boolean model (e). In 2015 30th IEEE\/ACM International Conference on Automated Software Engineering (ASE). 260\u2013270. https:\/\/doi.org\/10.1109\/ase.2015.42 10.1109\/ase.2015.42"},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 1075\u20131088","author":"Ma Ji","year":"2021","unstructured":"Ji Ma, Ivan Korotkov, Yinfei Yang, Keith Hall, and Ryan McDonald. 2021. Zero-shot Neural Passage Retrieval via Domain-targeted Synthetic Question Generation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 1075\u20131088. https:\/\/doi.org\/10.18653\/v1\/2021.eacl-main.92 10.18653\/v1\/2021.eacl-main.92"},{"key":"e_1_2_1_36_1","volume-title":"Sgpt: Gpt sentence embeddings for semantic search. arXiv preprint arXiv:2202.08904.","author":"Muennighoff Niklas","year":"2022","unstructured":"Niklas Muennighoff. 2022. Sgpt: Gpt sentence embeddings for semantic search. arXiv preprint arXiv:2202.08904."},{"key":"e_1_2_1_37_1","volume-title":"Codegen: An open large language model for code with multi-turn program synthesis. arXiv preprint arXiv:2203.13474.","author":"Nijkamp Erik","year":"2022","unstructured":"Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2022. Codegen: An open large language model for code with multi-turn program synthesis. arXiv preprint arXiv:2203.13474."},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2908\u20132926","author":"Patel Arkil","year":"2024","unstructured":"Arkil Patel, Siva Reddy, Dzmitry Bahdanau, and Pradeep Dasigi. 2024. Evaluating In-Context Learning of Libraries for Code Generation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2908\u20132926."},{"key":"e_1_2_1_39_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence. 37","author":"Pian Weiguo","year":"2023","unstructured":"Weiguo Pian, Hanyu Peng, Xunzhu Tang, Tiezhu Sun, Haoye Tian, Andrew Habib, Jacques Klein, and Tegawend\u00e9 F Bissyand\u00e9. 2023. MetaTPTrans: A meta learning approach for multilingual code representation learning. In Proceedings of the AAAI Conference on Artificial Intelligence. 37, 5239\u20135247. https:\/\/doi.org\/10.1609\/aaai.v37i4.25654 10.1609\/aaai.v37i4.25654"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","unstructured":"N Reimers. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv preprint arXiv:1908.10084 https:\/\/doi.org\/10.18653\/v1\/d19-1410 10.18653\/v1\/d19-1410","DOI":"10.18653\/v1"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1561\/1500000019"},{"key":"e_1_2_1_42_1","volume-title":"Yossi Adi, Jingyu Liu, Tal Remez, and J\u00e9r\u00e9my Rapin.","author":"Roziere Baptiste","year":"2023","unstructured":"Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, and J\u00e9r\u00e9my Rapin. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950."},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the 2nd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages. 31\u201341","author":"Sachdev Saksham","year":"2018","unstructured":"Saksham Sachdev, Hongyu Li, Sifei Luan, Seohyun Kim, Koushik Sen, and Satish Chandra. 2018. Retrieval on source code: a neural code search. In Proceedings of the 2nd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages. 31\u201341. https:\/\/doi.org\/10.1145\/3211346.3211353 10.1145\/3211346.3211353"},{"key":"e_1_2_1_44_1","unstructured":"Sebastian Schelter Felix Biessmann Tim Januschowski David Salinas Stephan Seufert and Gyuri Szarvas. 2015. On challenges in machine learning model management."},{"key":"e_1_2_1_45_1","volume-title":"Proceedings of the 30th ACM International Conference on Information & Knowledge Management. 4104\u20134113","author":"Sheng Xiang-Rong","year":"2021","unstructured":"Xiang-Rong Sheng, Liqin Zhao, Guorui Zhou, Xinyao Ding, Binding Dai, Qiang Luo, Siran Yang, Jingshan Lv, Chi Zhang, and Hongbo Deng. 2021. One model to serve all: Star topology adaptive recommender for multi-domain ctr prediction. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management. 4104\u20134113. https:\/\/doi.org\/10.1145\/3459637.3481941 10.1145\/3459637.3481941"},{"key":"e_1_2_1_46_1","volume-title":"2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). 2198\u20132210","author":"Shi Ensheng","year":"2023","unstructured":"Ensheng Shi, Yanlin Wang, Wenchao Gu, Lun Du, Hongyu Zhang, Shi Han, Dongmei Zhang, and Hongbin Sun. 2023. Cocosoda: Effective contrastive learning for code search. In 2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). 2198\u20132210. https:\/\/doi.org\/10.1109\/icse48619.2023.00185 10.1109\/icse48619.2023.00185"},{"key":"e_1_2_1_47_1","volume-title":"Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis. 39\u201351","author":"Shi Ensheng","year":"2023","unstructured":"Ensheng Shi, Yanlin Wang, Hongyu Zhang, Lun Du, Shi Han, Dongmei Zhang, and Hongbin Sun. 2023. Towards efficient fine-tuning of pre-trained code models: An experimental study and beyond. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis. 39\u201351. https:\/\/doi.org\/10.1145\/3597926.3598036 10.1145\/3597926.3598036"},{"key":"e_1_2_1_48_1","volume-title":"Proceedings of the 28th International Conference on Program Comprehension. 196\u2013207","author":"Shuai Jianhang","year":"2020","unstructured":"Jianhang Shuai, Ling Xu, Chao Liu, Meng Yan, Xin Xia, and Yan Lei. 2020. Improving code search with co-attentive representation learning. In Proceedings of the 28th International Conference on Program Comprehension. 196\u2013207."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/ms.2010.95"},{"key":"e_1_2_1_50_1","unstructured":"Jacob Mitchell Springer Suhas Kotha Daniel Fried Graham Neubig and Aditi Raghunathan. 2024. Repetition improves language model embeddings. arXiv preprint arXiv:2402.15449."},{"key":"e_1_2_1_51_1","volume-title":"2013 21st international conference on program comprehension (icpc). 83\u201392","author":"Steidl Daniela","year":"2013","unstructured":"Daniela Steidl, Benjamin Hummel, and Elmar Juergens. 2013. Quality analysis of source code comments. In 2013 21st international conference on program comprehension (icpc). 83\u201392. https:\/\/doi.org\/10.1109\/icpc.2013.6613836 10.1109\/icpc.2013.6613836"},{"key":"e_1_2_1_52_1","unstructured":"Weisong Sun Chunrong Fang Yudu You Yun Miao Yi Liu Yuekang Li Gelei Deng Shenghan Huang Yuchen Chen and Quanjun Zhang. 2023. Automatic code summarization via chatgpt: How far are we? arXiv preprint arXiv:2305.12865."},{"key":"e_1_2_1_53_1","volume-title":"2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). 5\u201316","author":"Wang Deze","year":"2023","unstructured":"Deze Wang, Boxing Chen, Shanshan Li, Wei Luo, Shaoliang Peng, Wei Dong, and Xiangke Liao. 2023. One adapter for all programming languages? adapter tuning for code search and summarization. In 2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). 5\u201316. https:\/\/doi.org\/10.1109\/icse48619.2023.00013 10.1109\/icse48619.2023.00013"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","unstructured":"Kexin Wang Nandan Thakur Nils Reimers and Iryna Gurevych. 2021. GPL: Generative pseudo labeling for unsupervised domain adaptation of dense retrieval. arXiv preprint arXiv:2112.07577 https:\/\/doi.org\/10.18653\/v1\/2022.naacl-main.168 10.18653\/v1\/2022.naacl-main.168","DOI":"10.18653\/v1"},{"key":"e_1_2_1_55_1","volume-title":"You Augment Me: Exploring ChatGPT-based Data Augmentation for Semantic Code Search. In 2023 IEEE International Conference on Software Maintenance and Evolution (ICSME). 14\u201325","author":"Wang Yanlin","year":"2023","unstructured":"Yanlin Wang, Lianghong Guo, Ensheng Shi, Wenqing Chen, Jiachi Chen, Wanjun Zhong, Menghan Wang, Hui Li, Hongyu Zhang, and Ziyu Lyu. 2023. You Augment Me: Exploring ChatGPT-based Data Augmentation for Semantic Code Search. In 2023 IEEE International Conference on Software Maintenance and Evolution (ICSME). 14\u201325. https:\/\/doi.org\/10.1109\/icsme58846.2023.00014 10.1109\/icsme58846.2023.00014"},{"key":"e_1_2_1_56_1","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 1069\u20131088","author":"Wang Yue","year":"2023","unstructured":"Yue Wang, Hung Le, Akhilesh Gotmare, Nghi Bui, Junnan Li, and Steven Hoi. 2023. CodeT5+: Open Code Large Language Models for Code Understanding and Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 1069\u20131088. https:\/\/doi.org\/10.18653\/v1\/2023.emnlp-main.68 10.18653\/v1\/2023.emnlp-main.68"},{"key":"e_1_2_1_57_1","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 8696\u20138708","author":"Wang Yue","year":"2021","unstructured":"Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 8696\u20138708. https:\/\/doi.org\/10.18653\/v1\/2021.emnlp-main.685 10.18653\/v1\/2021.emnlp-main.685"},{"key":"e_1_2_1_58_1","volume-title":"2018 International Workshop on Blockchain Oriented Software Engineering (IWBOSE). 2\u20138.","author":"Wohrer Maximilian","year":"2018","unstructured":"Maximilian Wohrer and Uwe Zdun. 2018. Smart contracts: security patterns in the ethereum ecosystem and solidity. In 2018 International Workshop on Blockchain Oriented Software Engineering (IWBOSE). 2\u20138."},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-017-9514-4"},{"key":"e_1_2_1_60_1","volume-title":"Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 641\u2013649","author":"Xiao Shitao","year":"2024","unstructured":"Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. 2024. C-Pack: Packed Resources For General Chinese Embeddings. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 641\u2013649."},{"key":"e_1_2_1_61_1","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2471\u20132484","author":"Zhang Fengji","year":"2023","unstructured":"Fengji Zhang, Bei Chen, Yue Zhang, Jacky Keung, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen. 2023. RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2471\u20132484. https:\/\/doi.org\/10.18653\/v1\/2023.emnlp-main.151 10.18653\/v1\/2023.emnlp-main.151"},{"key":"e_1_2_1_62_1","volume-title":"Proceedings of the 1st Workshop on Natural Language Processing for Programming (NLP4Prog 2021","author":"Zhang Xinyu","year":"2021","unstructured":"Xinyu Zhang, Ji Xin, Andrew Yates, and Jimmy Lin. 2021. Bag-of-Words Baselines for Semantic Code Search. In Proceedings of the 1st Workshop on Natural Language Processing for Programming (NLP4Prog 2021). 88\u201394. https:\/\/doi.org\/10.18653\/v1\/2021.nlp4prog-1.10 10.18653\/v1\/2021.nlp4prog-1.10"},{"key":"e_1_2_1_63_1","volume-title":"Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 6344\u20136355","author":"Zhao Yao","year":"2024","unstructured":"Yao Zhao, Zhitian Xie, Chen Liang, Chenyi Zhuang, and Jinjie Gu. 2024. Lookahead: An inference acceleration framework for large language model with lossless generation accuracy. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 6344\u20136355. https:\/\/doi.org\/10.1145\/3637528.3671614 10.1145\/3637528.3671614"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729357","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T15:28:34Z","timestamp":1750346914000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3729357"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,19]]},"references-count":63,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2025,6,19]]}},"alternative-id":["10.1145\/3729357"],"URL":"https:\/\/doi.org\/10.1145\/3729357","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,19]]}}}