{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T03:48:18Z","timestamp":1781927298392,"version":"3.54.5"},"reference-count":86,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2024,6,3]],"date-time":"2024-06-03T00:00:00Z","timestamp":1717372800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Key Foundation of China","doi-asserted-by":"crossref","award":["62032016"],"award-info":[{"award-number":["62032016"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2024,6,30]]},"abstract":"<jats:p>Code search, which refers to the process of identifying the most relevant code snippets for a given natural language query, plays a crucial role in software maintenance. However, current approaches heavily rely on labeled data for training, which results in performance decreases when confronted with cross-domain scenarios including domain- or project-specific situations. This decline can be attributed to their limited ability to effectively capture the semantics associated with such scenarios. To tackle the aforementioned problem, we propose a ze<jats:bold>R<\/jats:bold>o-shot dom<jats:bold>A<\/jats:bold>in ada<jats:bold>P<\/jats:bold>tion with pre-tra<jats:bold>I<\/jats:bold>ned mo<jats:bold>D<\/jats:bold>els framework for code search named RAPID. The framework first generates synthetic data by pseudo labeling, then trains the CodeBERT with sampled synthetic data. To avoid the influence of noisy synthetic data and enhance the model performance, we propose a mixture sampling strategy to obtain hard negative samples during training. Specifically, the mixture sampling strategy considers both relevancy and diversity to select the data that are hard to be distinguished by the models. To validate the effectiveness of our approach in zero-shot settings, we conduct extensive experiments and find that RAPID outperforms the CoCoSoDa and UniXcoder model by an average of 15.7% and 10%, respectively, as measured by the MRR metric. When trained on full data, our approach results in an average improvement of 7.5% under the MRR metric using CodeBERT. We observe that as the model\u2019s performance in zero-shot tasks improves, the impact of hard negatives diminishes. Our observation also indicates that fine-tuning CodeT5 for generating pseudo labels can enhance the performance of the code search model, and using only 100-shot samples can yield comparable results to the supervised baseline. Furthermore, we evaluate the effectiveness of RAPID in real-world code search tasks in three GitHub projects through both human and automated assessments. Our findings reveal RAPID exhibits superior performance, e.g., an average improvement of 18% under the MRR metric over the top-performing model.<\/jats:p><jats:p\/>","DOI":"10.1145\/3641542","type":"journal-article","created":{"date-parts":[[2024,1,18]],"date-time":"2024-01-18T12:16:59Z","timestamp":1705580219000},"page":"1-35","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":16,"title":["RAPID: Zero-Shot Domain Adaptation for Code Search with Pre-Trained Models"],"prefix":"10.1145","volume":"33","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2031-0615","authenticated-orcid":false,"given":"Guodong","family":"Fan","sequence":"first","affiliation":[{"name":"College of Intelligence and Computing, Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4430-4765","authenticated-orcid":false,"given":"Shizhan","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Intelligence and Computing, Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8513-6836","authenticated-orcid":false,"given":"Cuiyun","family":"Gao","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3741-8104","authenticated-orcid":false,"given":"Jianmao","family":"Xiao","sequence":"additional","affiliation":[{"name":"School of Software, Jiangxi Normal University, Nanchang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6272-4069","authenticated-orcid":false,"given":"Tao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Macau University of Science and Technology, Macau, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8158-7453","authenticated-orcid":false,"given":"Zhiyong","family":"Feng","sequence":"additional","affiliation":[{"name":"College of Intelligence and Computing, Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,6,3]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","first-page":"590","DOI":"10.1145\/3377811.3380405","volume-title":"Proceedings of the ACM\/IEEE 42nd International Conference on Software Engineering","author":"Aghajani Emad","year":"2020","unstructured":"Emad Aghajani, Csaba Nagy, Mario Linares-V\u00e1squez, Laura Moreno, Gabriele Bavota, Michele Lanza, and David C. Shepherd. 2020. Software documentation: The practitioners\u2019 perspective. In Proceedings of the ACM\/IEEE 42nd International Conference on Software Engineering. 590\u2013601."},{"key":"e_1_3_2_3_2","unstructured":"D. Bahdanau K. Cho and Y. Bengio. 2014. Neural machine translation by jointly learning to align and translate. ICLR 2015 (2014) 1\u201315."},{"key":"e_1_3_2_4_2","first-page":"1877","article-title":"Language models are few-shot learners","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems . 1877\u20131901.","journal-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510125"},{"key":"e_1_3_2_6_2","first-page":"167","volume-title":"Proceedings of the 2021 36th IEEE\/ACM International Conference on Automated Software Engineering","author":"Chen Binger","year":"2021","unstructured":"Binger Chen and Ziawasch Abedjan. 2021. Interactive cross-language code retrieval with auto-encoders. In Proceedings of the 2021 36th IEEE\/ACM International Conference on Automated Software Engineering. IEEE, 167\u2013178."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/2207676.2208589"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3565971"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939719"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482127"},{"key":"e_1_3_2_11_2","first-page":"29","volume-title":"Proceedings of the 2019 IEEE\/ACM 16th International Conference on Mining Software Repositories","author":"Efstathiou Vasiliki","year":"2019","unstructured":"Vasiliki Efstathiou and Diomidis Spinellis. 2019. Semantic source code models using identifier embeddings. In Proceedings of the 2019 IEEE\/ACM 16th International Conference on Mining Software Repositories. IEEE, 29\u201333."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2021.106542"},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","unstructured":"Zhangyin Feng Daya Guo Duyu Tang Nan Duan Xiaocheng Feng Ming Gong Linjun Shou Bing Qin Ting Liu Daxin Jiang and Ming Zhou. 2020. Codebert: A pre-trained model for programming and natural languages. CodeBERT: A Pre-Trained Model for Programming and Natural Languages[C]\/\/Findings of the Association for Computational Linguistics: EMNLP 2020. 1536\u20131547.","DOI":"10.18653\/v1\/2020.findings-emnlp.139"},{"key":"e_1_3_2_14_2","first-page":"1126","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Finn Chelsea","year":"2017","unstructured":"Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the International Conference on Machine Learning. PMLR, 1126\u20131135."},{"key":"e_1_3_2_15_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Gao Jun","year":"2018","unstructured":"Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tieyan Liu. 2018. Representation degeneration problem in training natural language generation models. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3522674"},{"key":"e_1_3_2_17_2","article-title":"Code search: A survey of techniques for finding code","author":"Grazia Luca Di","year":"2022","unstructured":"Luca Di Grazia and Michael Pradel. 2022. Code search: A survey of techniques for finding code. ACM Computing Surveys (CSUR) 55, 11, Article No. 220 (2022), 1\u201331.","journal-title":"ACM Computing Surveys (CSUR)"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2021.04.019"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.181"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3180155.3180167"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.499"},{"key":"e_1_3_2_22_2","unstructured":"Daya Guo Shuo Ren Shuai Lu Zhangyin Feng Duyu Tang Shujie Liu Long Zhou Nan Duan Alexey Svyatkovskiy Shengyu Fu Michele Tufano Shao Kun Deng Colin Clement Dawn Drain Neel Sundaresan Jian Yin Daxin Jiang and Ming Zhou. 2020. GraphCodeBERT: Pre-training Code Representations with Data Flow[C]\/\/International Conference on Learning Representations. 2020. 1\u201318."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.740"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.100"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462891"},{"key":"e_1_3_2_26_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Holtzman Ari","year":"2019","unstructured":"Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_27_2","first-page":"994","volume-title":"Proceedings of the 16th ACM International Conference on Web Search and Data Mining","author":"Hu Fan","year":"2023","unstructured":"Fan Hu, Yanlin Wang, Lun Du, Xirong Li, Hongyu Zhang, Shi Han, and Dongmei Zhang. 2023. Revisiting code search in a two-stage paradigm. In Proceedings of the 16th ACM International Conference on Web Search and Data Mining. 994\u20131002."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549141"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3196321.3196334"},{"key":"e_1_3_2_30_2","first-page":"5690","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Huang Junjie","year":"2021","unstructured":"Junjie Huang, Duyu Tang, Linjun Shou, Ming Gong, Ke Xu, Daxin Jiang, Ming Zhou, and Nan Duan. 2021. CoSQA: 20,000+ web queries for code search and question answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 5690\u20135700."},{"key":"e_1_3_2_31_2","first-page":"138","volume-title":"Proceedings of the 44th International Conference on Software Engineering","author":"Huo Yintong","year":"2022","unstructured":"Yintong Huo, Yuxin Su, Hongming Zhang, and Michael R. Lyu. 2022. ARCLIN: Automated API mention resolution for unformatted texts. In Proceedings of the 44th International Conference on Software Engineering. 138\u2013149."},{"key":"e_1_3_2_32_2","unstructured":"Hamel Husain Ho-Hsiang Wu Tiferet Gazit Miltiadis Allamanis and Marc Brockschmidt. 2019. Codesearchnet challenge: Evaluating the state of semantic code search. arXiv:1909.09436. Retrieved from https:\/\/arxiv.org\/abs\/1909.09436"},{"key":"e_1_3_2_33_2","first-page":"21798","article-title":"Hard negative mixing for contrastive learning","author":"Kalantidis Yannis","year":"2020","unstructured":"Yannis Kalantidis, Mert Bulent Sariyildiz, Noe Pion, Philippe Weinzaepfel, and Diane Larlus. 2020. Hard negative mixing for contrastive learning. In Proceedings of the 34th International Conference on Neural Information Processing Systems . 21798\u201321809.","journal-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems"},{"key":"e_1_3_2_34_2","first-page":"4171","volume-title":"Proceedings of the NAACL-HLT","author":"Kenton Jacob Devlin Ming-Wei Chang","year":"2019","unstructured":"Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the NAACL-HLT. 4171\u20134186."},{"issue":"3","key":"e_1_3_2_35_2","first-page":"1","article-title":"A brief overview of universal sentence representation methods: A linguistic view","volume":"55","author":"Li Ruiqi","year":"2022","unstructured":"Ruiqi Li, Xiang Zhao, and Marie-Francine Moens. 2022. A brief overview of universal sentence representation methods: A linguistic view. ACM Computing Surveys 55, 3 (2022), 1\u201342.","journal-title":"ACM Computing Surveys"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/2950290.2950341"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.353"},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","unstructured":"Zongjie Li Chaozheng Wang Zhibo Liu Haoxuan Wang Shuai Wang and Cuiyun Gao. 2022. CCTEST: Testing and repairing code completion systems[C]\/\/2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). IEEE 1238\u20131250.","DOI":"10.1109\/ICSE48619.2023.00110"},{"key":"e_1_3_2_39_2","first-page":"74","volume-title":"Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out. 74\u201381."},{"key":"e_1_3_2_40_2","article-title":"Efficient training of retrieval models using negative cache","author":"Lindgren Erik","year":"2021","unstructured":"Erik Lindgren, Sashank Reddi, Ruiqi Guo, and Sanjiv Kumar. 2021. Efficient training of retrieval models using negative cache. In Advances in Neural Information Processing Systems, Marc\u2019Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (Eds.). Vol. 34, Curran Associates, Inc., 4134\u20134146.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3480027"},{"issue":"9","key":"e_1_3_2_42_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3560815","article-title":"Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing","volume":"55","author":"Liu Pengfei","year":"2023","unstructured":"Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. Computing Surveys 55, 9 (2023), 1\u201335.","journal-title":"Computing Surveys"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-short.8"},{"key":"e_1_3_2_44_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.492"},{"key":"e_1_3_2_46_2","unstructured":"Shuai Lu Daya Guo Shuo Ren Junjie Huang Alexey Svyatkovskiy Ambrosio Blanco Colin Clement Dawn Drain Daxin Jiang Duyu Tang Ge Li Lidong Zhou Linjun Shou Long Zhou Michele Tufano Ming Gong Ming Zhou Nan Duan Neel Sundaresan Shao Kun Deng Shengyu Fu and Shujie Liu. 2021. Codexglue: A machine learning benchmark dataset for code understanding and generation. NeurIPS Datasets and Benchmarks 2021."},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3412747"},{"key":"e_1_3_2_48_2","first-page":"260","volume-title":"Proceedings of the 2015 30th IEEE\/ACM International Conference on Automated Software Engineering (ASE)","author":"Lv Fei","year":"2015","unstructured":"Fei Lv, Hongyu Zhang, Jian-guang Lou, Shaowei Wang, Dongmei Zhang, and Jianjun Zhao. 2015. CodeHow: Effective code search based on API understanding and extended boolean model (E). In Proceedings of the 2015 30th IEEE\/ACM International Conference on Automated Software Engineering (ASE). IEEE, 260\u2013270."},{"key":"e_1_3_2_49_2","first-page":"1075","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Ma Ji","year":"2021","unstructured":"Ji Ma, Ivan Korotkov, Yinfei Yang, Keith Hall, and Ryan McDonald. 2021. Zero-shot neural passage retrieval via domain-targeted synthetic question generation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 1075\u20131088."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.eacl-main.92"},{"key":"e_1_3_2_51_2","doi-asserted-by":"crossref","first-page":"336","DOI":"10.1109\/ICSE43902.2021.00041","volume-title":"Proceedings of the 2021 IEEE\/ACM 43rd International Conference on Software Engineering (ICSE)","author":"Mastropaolo Antonio","year":"2021","unstructured":"Antonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader Palacio, Denys Poshyvanyk, Rocco Oliveto, and Gabriele Bavota. 2021. Studying the usage of text-to-text transfer transformer to support code-related tasks. In Proceedings of the 2021 IEEE\/ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 336\u2013347."},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3468264.3468538"},{"key":"e_1_3_2_53_2","doi-asserted-by":"crossref","unstructured":"Fangwen Mu Xiao Chen Lin Shi Song Wang and Qing Wang. 2022. Automatic comment generation via multi-pass deliberation. Proceedings of the 37th IEEE\/ACM International Conference on Automated Software Engineering. 1\u201312.","DOI":"10.1145\/3551349.3556917"},{"key":"e_1_3_2_54_2","unstructured":"Erik Nijkamp Bo Pang Hiroaki Hayashi Lifu Tu Huan Wang Yingbo Zhou Silvio Savarese and Caiming Xiong. 2022. CodeGen: An open large language model for code with multi-turn program synthesis. The 11th International Conference on Learning Representations. 2022."},{"key":"e_1_3_2_55_2","unstructured":"Long Ouyang Jeffrey Wu Xu Jiang Diogo Almeida Carroll Wainwright Pamela Mishkin Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray John Schulman Jacob Hilton Fraser Kelton Luke Miller Maddie Simens Amanda Askell Peter Welinder Paul F. Christiano Jan Leike and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35 (2022) 27730\u201327744."},{"key":"e_1_3_2_56_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_3_2_57_2","first-page":"103","volume-title":"Proceedings of the Transfer Learning for Natural Language Processing Workshop","author":"Patil Rajaswa","year":"2023","unstructured":"Rajaswa Patil, Manasi Patwardhan, Shirish Karande, Lovekesh Vig, and Gautam Shroff. 2023. Exploring dimensions of generalizability and few-shot transfer for text-to-SQL semantic parsing. In Proceedings of the Transfer Learning for Natural Language Processing Workshop. PMLR, 103\u2013114."},{"key":"e_1_3_2_58_2","first-page":"701","volume-title":"Proceedings of the 2021 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER)","author":"Pierro Giuseppe Antonio","year":"2021","unstructured":"Giuseppe Antonio Pierro and Roberto Tonelli. 2021. Analysis of source code duplication in Ethreum smart contracts. In Proceedings of the 2021 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE, 701\u2013707."},{"issue":"8","key":"e_1_3_2_59_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_2_60_2","first-page":"140:1\u2013140:67","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21, 1 (2020), 140:1\u2013140:67. Retrieved from http:\/\/jmlr.org\/papers\/v21\/20-074.html","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_61_2","first-page":"207","volume-title":"Advances in Information Retrieval: 44th European Conference on IR Research, ECIR 2022","author":"Rau David","year":"2022","unstructured":"David Rau and Jaap Kamps. 2022. How different are pre-trained transformers for text ranking?. In Advances in Information Retrieval: 44th European Conference on IR Research, ECIR 2022, Matthias Hagen, Suzan Verberne, Craig Macdonald, Christin Seifert, Krisztian Balog, Kjetil N\u00f8rv\u00e5g, and Vinay Setty (Eds.). Springer, 207\u2013214."},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1561\/1500000019"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3211346.3211353"},{"key":"e_1_3_2_66_2","article-title":"On the effectiveness of transfer learning for code search","author":"Salza Pasquale","year":"2023","unstructured":"Pasquale Salza, Christoph Schwizer, Jian Gu, and Harald C. Gall. 2023. On the effectiveness of transfer learning for code search. IEEE Transactions on Software Engineering 49, 4 (2023), 1804\u20131822.","journal-title":"IEEE Transactions on Software Engineering"},{"key":"e_1_3_2_67_2","unstructured":"Victor Sanh Lysandre Debut Julien Chaumond and Thomas Wolf. 2019. DistilBERT a distilled version of BERT: Smaller faster cheaper and lighter. arXiv:1910.01108. Retrieved from https:\/\/arxiv.org\/abs\/1910.01108"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1145\/3468264.3468591"},{"key":"e_1_3_2_69_2","doi-asserted-by":"crossref","unstructured":"Ensheng Shi Wenchao Gub Yanlin Wang Lun Du Hongyu Zhang Shi Han Dongmei Zhang and Hongbin Sun. 2023. CoCoSoDa: Effective contrastive learning for code search. IEEE\/ACM 45th International Conference on Software Engineering (ICSE\u201923). IEEE 2198\u20132210.","DOI":"10.1109\/ICSE48619.2023.00185"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549087"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/3387904.3389269"},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/MS.2010.95"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1145\/203241.203256"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510160"},{"key":"e_1_3_2_75_2","unstructured":"Nandan Thakur Nils Reimers Andreas R\u00fcckl\u00e9 Abhishek Srivastava and Iryna Gurevych. 2021. BEIR: A heterogenous benchmark for zero-shot evaluation of information retrieval models. 35th Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2). 1\u201316."},{"issue":"11","key":"e_1_3_2_76_2","article-title":"Visualizing data using t-SNE.","volume":"9","author":"Maaten Laurens Van der","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, 11 (2008), 2579\u20132605.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_77_2","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017. Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (Eds.), Curran Associates, Inc., 5998\u20136008."},{"key":"e_1_3_2_78_2","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1109\/ASE.2019.00012","volume-title":"Proceedings of the 2019 34th IEEE\/ACM International Conference on Automated Software Engineering (ASE)","author":"Wan Yao","year":"2019","unstructured":"Yao Wan, Jingdong Shu, Yulei Sui, Guandong Xu, Zhou Zhao, Jian Wu, and Philip Yu. 2019. Multi-modal attention network learning for semantic source code retrieval. In Proceedings of the 2019 34th IEEE\/ACM International Conference on Automated Software Engineering (ASE). IEEE, 13\u201325."},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549113"},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-emnlp.59"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.168"},{"key":"e_1_3_2_82_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.685"},{"key":"e_1_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2007.1078"},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462880"},{"key":"e_1_3_2_85_2","article-title":"An accurate identifier renaming prediction and suggestion approach","author":"Zhang Jingxuan","year":"2023","unstructured":"Jingxuan Zhang, Junpeng Luo, Jiahui Liang, Lina Gong, and Zhiqiu Huang. 2023. An accurate identifier renaming prediction and suggestion approach. ACM Transactions on Software Engineering and Methodology (2023).","journal-title":"ACM Transactions on Software Engineering and Methodology"},{"key":"e_1_3_2_86_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Zhang Tianyi","year":"2019","unstructured":"Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2019. BERTScore: Evaluating text generation with BERT. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_87_2","first-page":"11730","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"36","author":"Zhang Yanzhao","year":"2022","unstructured":"Yanzhao Zhang, Richong Zhang, Samuel Mensah, Xudong Liu, and Yongyi Mao. 2022. Unsupervised sentence representation via contrastive learning with mixing negatives. In Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 36, 11730\u201311738."}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3641542","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3641542","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:50:16Z","timestamp":1750287016000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3641542"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,3]]},"references-count":86,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,6,30]]}},"alternative-id":["10.1145\/3641542"],"URL":"https:\/\/doi.org\/10.1145\/3641542","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,3]]},"assertion":[{"value":"2023-05-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-10","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-06-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}