{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,17]],"date-time":"2026-03-17T00:54:31Z","timestamp":1773708871372,"version":"3.50.1"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"abstract":"<jats:p>Code generation models based on the pre-training and fine-tuning paradigm have been increasingly attempted by both academia and industry, resulting in well-known industrial models such as Codex, CodeGen, and PanGu-Coder. After being pre-trained on a large-scale corpus of code, a model is further fine-tuned with datasets specifically for the target downstream task, e.g., generating code from natural language description. The target code being generated can be classified into two types: a standalone function, i.e., a function that invokes or accesses only built-in functions and standard libraries, and a non-standalone function, i.e., a function that invokes or accesses user-defined functions or third-party libraries.<\/jats:p>\n          <jats:p>To effectively generate code especially non-standalone functions (largely ignored by existing work), in this article, we present Wenwang, an approach to improving the capability of a pre-trained model on generating code beyond standalone functions. Wenwang consists of two components: a fine-tuning dataset named WenwangData and a fine-tuned model named WenwangCoder. Compared with existing fine-tuning datasets, WenwangData additionally covers non-standalone functions. Besides the docstring and code snippet for a function, WenwangData also includes its contextual information collected via program analysis. Based on PanGu-Coder, we produce WenwangCoder by fine-tuning PanGu-Coder on WenwangData with our context-aware fine-tuning technique so that the contextual information can be fully leveraged during code generation. On CoderEval and HumanEval, WenwangCoder outperforms three state-of-the-art models with similar parameter sizes (at the scale of around 300M), namely CodeGen, PanGu-Coder, and PanGu-FT. Although WenwangCoder does not outperform ChatGPT on HumanEval, WenwangCoder with smaller model parameter sizes can achieve similar effects to ChatGPT on CoderEval. Our experimental results also shed light on a number of promising optimization directions based on existing pre-trained models.<\/jats:p>","DOI":"10.1145\/3725213","type":"journal-article","created":{"date-parts":[[2025,3,20]],"date-time":"2025-03-20T14:21:05Z","timestamp":1742480465000},"update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Wenwang: Toward Effectively Generating Code Beyond Standalone Functions via Generative Pre-trained Models"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3828-7612","authenticated-orcid":false,"given":"Hao","family":"Yu","sequence":"first","affiliation":[{"name":"The Hong Kong University of Science and Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0825-8001","authenticated-orcid":false,"given":"Bo","family":"Shen","sequence":"additional","affiliation":[{"name":"Huawei Cloud Computing Technologies Co., Ltd., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-6438-360X","authenticated-orcid":false,"given":"Jiaxin","family":"Zhang","sequence":"additional","affiliation":[{"name":"Huawei Noah\u2019s Ark Lab, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-0310-855X","authenticated-orcid":false,"given":"Shaoxin","family":"Lin","sequence":"additional","affiliation":[{"name":"Huawei Cloud Computing Technologies Co., Ltd., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-1572-6640","authenticated-orcid":false,"given":"Lin","family":"Li","sequence":"additional","affiliation":[{"name":"Huawei Cloud Computing Technologies Co., Ltd., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-2454-1706","authenticated-orcid":false,"given":"Guangtai","family":"Liang","sequence":"additional","affiliation":[{"name":"Huawei Cloud Computing Technologies Co., Ltd., China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6278-2357","authenticated-orcid":false,"given":"Ying","family":"Li","sequence":"additional","affiliation":[{"name":"National Research Center of Software Engineering, Peking University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6598-0041","authenticated-orcid":false,"given":"Qianxiang","family":"Wang","sequence":"additional","affiliation":[{"name":"National Research Center of Software Engineering, Peking University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6731-216X","authenticated-orcid":false,"given":"Tao","family":"Xie","sequence":"additional","affiliation":[{"name":"Key Lab of HCST (PKU), MOE; SCS; Peking University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,3,20]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2021. https:\/\/github.com\/openai\/human-eval\/."},{"key":"e_1_2_1_2_1","unstructured":"2022. https:\/\/cloud.google.com\/blog\/topics\/public-datasets\/github-on-bigquery-analyze-all-the-open-source-code."},{"key":"e_1_2_1_3_1","unstructured":"2022. https:\/\/github.com\/."},{"key":"e_1_2_1_4_1","unstructured":"2022. https:\/\/www.travis-ci.com."},{"key":"e_1_2_1_5_1","unstructured":"2022. https:\/\/tox.wiki\/en\/latest."},{"key":"e_1_2_1_6_1","unstructured":"2024. https:\/\/zenodo.org\/records\/11238957."},{"key":"e_1_2_1_7_1","unstructured":"2024. https:\/\/www.graphpad.com\/quickcalcs\/McNemar1.cfm."},{"key":"e_1_2_1_8_1","volume-title":"Guiding language models of code with global context using monitors. arXiv preprint arXiv:2306.10763","author":"Agrawal Lakshya A","year":"2023","unstructured":"Lakshya A Agrawal, Aditya Kanade, Navin Goyal, Shuvendu K Lahiri, and Sriram K Rajamani. 2023. Guiding language models of code with global context using monitors. arXiv preprint arXiv:2306.10763 (2023)."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445922"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","unstructured":"Sid Black Leo Gao Phil Wang Connor Leahy and Stella Biderman. 2021. GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow. https:\/\/doi.org\/10.5281\/zenodo.5297715","DOI":"10.5281\/zenodo.5297715"},{"key":"e_1_2_1_11_1","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems 33 (2020), 1877\u20131901.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_12_1","volume-title":"Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, Arjun Guha, Michael Greenberg, and Abhinav Jangda.","author":"Cassano Federico","year":"2022","unstructured":"Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, Arjun Guha, Michael Greenberg, and Abhinav Jangda. 2022. A scalable and extensible approach to benchmarking NL2Code for 18 programming languages. arXiv:2208.08227"},{"key":"e_1_2_1_13_1","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harrison Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mohammad Bavarian Clemens Winter Philippe Tillet Felipe Petroski Such Dave Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William Hebgen Guss Alex Nichol Alex Paino Nikolas Tezak Jie Tang Igor Babuschkin Suchir Balaji Shantanu Jain William Saunders Christopher Hesse Andrew N. Carr Jan Leike Joshua Achiam Vedant Misra Evan Morikawa Alec Radford Matthew Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating large language models trained on code. (2021). arXiv:2107.03374"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2022.106987"},{"key":"e_1_2_1_15_1","volume-title":"PanGu-Coder: Program synthesis with function-level language modeling. arXiv preprint arXiv:2207.11280","author":"Christopoulou Fenia","year":"2022","unstructured":"Fenia Christopoulou, Gerasimos Lampouras, Milan Gritta, Guchun Zhang, Yinpeng Guo, Zhongqi Li, Qi Zhang, Meng Xiao, Bo Shen, Lin Li, Hao Yu, Li Yan, Pingyi Zhou, Xin Wang, Yuchi Ma, Ignacio Iacobacci, Yasheng Wang, Guangtai Liang, Jiansheng Wei, Xin Jiang, Qianxiang Wang, and Qun Liu. 2022. PanGu-Coder: Program synthesis with function-level language modeling. arXiv preprint arXiv:2207.11280 (2022)."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3545945.3569823"},{"key":"e_1_2_1_17_1","volume-title":"Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, and Bing Xiang.","author":"Ding Yangruibo","year":"2022","unstructured":"Yangruibo Ding, Zijian Wang, Wasi Uddin Ahmad, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, and Bing Xiang. 2022. Cocomic: Code completion by jointly modeling in-file and cross-file context. arXiv preprint arXiv:2212.10007 (2022)."},{"key":"e_1_2_1_18_1","unstructured":"Daniel Fried Armen Aghajanyan Jessy Lin Sida Wang Eric Wallace Freda Shi Ruiqi Zhong Wen-tau Yih Luke Zettlemoyer and Mike Lewis. 2022. InCoder: A generative model for code infilling and synthesis. arXiv:2204.05999"},{"key":"e_1_2_1_19_1","volume-title":"The Pile: An 800GB dataset of diverse text for language modeling.","author":"Gao Leo","year":"2021","unstructured":"Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2021. The Pile: An 800GB dataset of diverse text for language modeling. (2021). arXiv:2101.00027"},{"key":"e_1_2_1_20_1","doi-asserted-by":"crossref","unstructured":"Daya Guo Shuai Lu Nan Duan Yanlin Wang Ming Zhou and Jian Yin. 2022. UniXcoder: Unified cross-modal pre-training for code representation. (2022). arXiv:2203.03850","DOI":"10.18653\/v1\/2022.acl-long.499"},{"key":"e_1_2_1_21_1","unstructured":"Yiyang Hao Ge Li Yongqiang Liu Xiaowei Miao He Zong Siyuan Jiang Yang Liu and He Wei. 2022. AixBench: a code generation benchmark dataset. arXiv:2206.13179"},{"key":"e_1_2_1_22_1","volume-title":"Preceedings of the 8th International Conference on Learning Representations.","author":"Holtzman Ari","year":"2020","unstructured":"Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In Preceedings of the 8th International Conference on Learning Representations."},{"key":"e_1_2_1_23_1","doi-asserted-by":"crossref","unstructured":"Max Hort Anastasiia Grishina and Leon Moonen. 2023. An exploratory literature study on sharing and energy use of language models for source code. (2023). arXiv:2307.02443","DOI":"10.1109\/ESEM56168.2023.10304803"},{"key":"e_1_2_1_24_1","unstructured":"Srinivasan Iyer Ioannis Konstas Alvin Cheung and Luke Zettlemoyer. 2018. Mapping language to code in programmatic context. (2018). arXiv:1808.09588"},{"key":"e_1_2_1_25_1","unstructured":"Woojeong Jin Yu Cheng Yelong Shen Weizhu Chen and Xiang Ren. 2021. A good prompt is worth millions of parameters: Low-resource prompt-based learning for vision-language models. (2021). arXiv:2110.08484"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.5555\/3524938.3525412"},{"key":"e_1_2_1_27_1","volume-title":"Preceeding of the 3rd International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.).","author":"Diederik","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Preceeding of the 3rd International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.)."},{"key":"e_1_2_1_28_1","unstructured":"Tomasz Korbak Hady Elsahar Germ\u00e1n Kruszewski and Marc Dymetman. [n.d.]. On reward maximization and distribution matching for fine-tuning language models."},{"key":"e_1_2_1_29_1","volume-title":"Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems.","author":"Kulal Sumith","year":"2019","unstructured":"Sumith Kulal, Panupong Pasupat, Kartik Chandra, Mina Lee, Oded Padon, Alex Aiken, and Percy Liang. 2019. SPoC: search-based pseudocode to code. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems."},{"key":"e_1_2_1_30_1","volume-title":"Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al.","author":"Li Raymond","year":"2023","unstructured":"Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. 2023. StarCoder: may the source be with you! (2023). arXiv:2305.06161"},{"key":"e_1_2_1_31_1","unstructured":"Suichan Li Dongdong Chen Yinpeng Chen Lu Yuan Lei Zhang Qi Chu Bin Liu and Nenghai Yu. 2021. Unsupervised Finetuning. arXiv:2110.09510"},{"key":"e_1_2_1_32_1","unstructured":"Yujia Li David Choi Junyoung Chung Nate Kushman Julian Schrittwieser R\u00e9mi Leblond Tom Eccles James Keeling Felix Gimeno Agustin Dal Lago Thomas Hubert Peter Choy Cyprien de Masson d\u2019Autume Igor Babuschkin Xinyun Chen Po-Sen Huang Johannes Welbl Sven Gowal Alexey Cherepanov James Molloy Daniel J. Mankowitz Esme Sutherland Robson Pushmeet Kohli Nando de Freitas Koray Kavukcuoglu and Oriol Vinyals. 2022. Competition-level code generation with AlphaCode. arXiv:2203.07814"},{"key":"e_1_2_1_33_1","doi-asserted-by":"crossref","unstructured":"Zehao Lin Guodun Li Jingfeng Zhang Yue Deng Xiangji Zeng Yin Zhang and Yao Wan. 2022. XCode: Towards cross-language code representation with large-scale pre-training. ACM Trans. Softw. Eng. Methodol. (2022).","DOI":"10.1145\/3506696"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416591"},{"key":"e_1_2_1_35_1","volume-title":"RepoBench: Benchmarking repository-level code auto-completion systems. arXiv preprint arXiv:2306.03091","author":"Liu Tianyang","year":"2023","unstructured":"Tianyang Liu, Canwen Xu, and Julian McAuley. 2023. RepoBench: Benchmarking repository-level code auto-completion systems. arXiv preprint arXiv:2306.03091 (2023)."},{"key":"e_1_2_1_36_1","unstructured":"Yang Liu. 2021. Unsupervised Finetuning. arXiv:1903.10318"},{"key":"e_1_2_1_37_1","volume-title":"Shengyu Fu, and Shujie Liu.","author":"Lu Shuai","year":"2021","unstructured":"Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021. CodeXGLUE: A machine learning benchmark dataset for code understanding and generation. arXiv:2102.04664"},{"key":"e_1_2_1_38_1","unstructured":"Ziyang Luo Can Xu Pu Zhao Qingfeng Sun Xiubo Geng Wenxiang Hu Chongyang Tao Jing Ma Qingwei Lin and Daxin Jiang. 2023. WizardCoder: Empowering code large language models with evol-instruct. (2023). arXiv:2306.08568"},{"key":"e_1_2_1_39_1","volume-title":"On the robustness of code generation techniques: An empirical study on GitHub Copilot. arXiv preprint arXiv:2302.00438","author":"Mastropaolo Antonio","year":"2023","unstructured":"Antonio Mastropaolo, Luca Pascarella, Emanuela Guglielmi, Matteo Ciniselli, Simone Scalabrino, Rocco Oliveto, and Gabriele Bavota. 2023. On the robustness of code generation techniques: An empirical study on GitHub Copilot. arXiv preprint arXiv:2302.00438 (2023)."},{"key":"e_1_2_1_40_1","volume-title":"BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv:1910.13461","author":"Mike Lewis Naman Goyal","year":"2019","unstructured":"Naman Goyal Mike Lewis, Yinhan Liu. 2019. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv:1910.13461"},{"key":"e_1_2_1_41_1","volume-title":"A conversational paradigm for program synthesis. arXiv preprint arXiv:2203.13474","author":"Nijkamp Erik","year":"2022","unstructured":"Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2022. A conversational paradigm for program synthesis. arXiv preprint arXiv:2203.13474 (2022)."},{"key":"e_1_2_1_42_1","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. 311\u2013318","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_2_1_43_1","volume-title":"Byeongwook Kim, Youngjoo Lee, and Dongsoo Lee.","author":"Park Gunho","year":"2022","unstructured":"Gunho Park, Baeseong Park, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, and Dongsoo Lee. 2022. nuQmm: Quantized matmul for efficient inference of large-scale generative language models. arXiv preprint arXiv:2206.09557 (2022)."},{"key":"e_1_2_1_44_1","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever. 2018. Improving language understanding by generative pre-training. (2018)."},{"key":"e_1_2_1_45_1","unstructured":"Shuo Ren Daya Guo Shuai Lu Long Zhou Shujie Liu Duyu Tang Neel Sundaresan Ming Zhou Ambrosio Blanco and Shuai Ma. 2020. CodeBLEU: A method for automatic evaluation of code synthesis. (2020). arXiv:2009.10297"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1021\/acs.est.3c01106"},{"key":"e_1_2_1_47_1","unstructured":"Disha Shrivastava Denis Kocetkov Harm de Vries Dzmitry Bahdanau and Torsten Scholak. 2023. RepoFusion: Training code models to understand your repository. (2023). arXiv:2306.10998"},{"key":"e_1_2_1_48_1","volume-title":"Preceeding of International Conference on Machine Learning. PMLR, 31693\u201331715","author":"Shrivastava Disha","year":"2023","unstructured":"Disha Shrivastava, Hugo Larochelle, and Daniel Tarlow. 2023. Repository-level prompt generation for large language models of code. In Preceeding of International Conference on Machine Learning. PMLR, 31693\u201331715."},{"key":"e_1_2_1_49_1","doi-asserted-by":"crossref","unstructured":"Chi Sun Xipeng Qiu Yige Xu and Xuanjing Huang. 2019. How to fine-tune BERT for text classification?. In Chinese Computational Linguistics. 194\u2013206.","DOI":"10.1007\/978-3-030-32381-3_16"},{"key":"e_1_2_1_50_1","volume-title":"Attention is all you need. Advances in Neural Information Processing Systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017)."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3552326.3587438"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","unstructured":"Yanlin Wang Ensheng Shi Lun Du Xiaodi Yang Yuxuan Hu Shi Han Hongyu Zhang and Dongmei Zhang. 2021. CoCoSum: Contextual Code Summarization with Multi-Relational Graph Neural Network. https:\/\/doi.org\/10.48550\/arXiv.1804.07461","DOI":"10.48550\/arXiv.1804.07461"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3623316"},{"key":"e_1_2_1_54_1","unstructured":"Wei Zeng Xiaozhe Ren Teng Su Hui Wang Yi Liao Zhiwei Wang Xin Jiang ZhenZhang Yang Kaisheng Wang Xiaoda Zhang Chen Li Ziyan Gong Yifan Yao Xinjing Huang Jun Wang Jianfeng Yu Qi Guo Yue Yu Yan Zhang Jin Wang Hengtao Tao Dasen Yan Zexuan Yi Fang Peng Fangqing Jiang Han Zhang Lingfeng Deng Yehong Zhang Zhe Lin Chao Zhang Shaojie Zhang Mingyue Guo Shanzhi Gu Gaojun Fan Yaowei Wang Xuefeng Jin Qun Liu and Yonghong Tian. [n.d.]. PanGu-\u03b1: Large-scale autoregressive pretrained chinese language models with auto-parallel computation. ([n. d.]). arXiv:2104.12369"},{"key":"e_1_2_1_55_1","volume-title":"Repocoder: Repository-level code completion through iterative retrieval and generation. arXiv preprint arXiv:2303.12570","author":"Zhang Fengji","year":"2023","unstructured":"Fengji Zhang, Bei Chen, Yue Zhang, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen. 2023. Repocoder: Repository-level code completion through iterative retrieval and generation. arXiv preprint arXiv:2303.12570 (2023)."},{"key":"e_1_2_1_56_1","volume-title":"interact, fine-tune: a novel interaction representation for text classification. Information Processing & Management 57, 6","author":"Zheng Jianming","year":"2020","unstructured":"Jianming Zheng, Fei Cai, Honghui Chen, and Maarten de Rijke. 2020. Pre-train, interact, fine-tune: a novel interaction representation for text classification. Information Processing & Management 57, 6 (2020)."},{"key":"e_1_2_1_57_1","doi-asserted-by":"crossref","unstructured":"Qinkai Zheng Xiao Xia Xu Zou Yuxiao Dong Shan Wang Yufei Xue Zihan Wang Lei Shen Andi Wang Yang Li et al. 2023. CodeGeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x. arXiv preprint arXiv:2303.17568 (2023).","DOI":"10.1145\/3580305.3599790"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3725213","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,3,20]],"date-time":"2025-03-20T15:33:27Z","timestamp":1742484807000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3725213"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,20]]},"references-count":57,"alternative-id":["10.1145\/3725213"],"URL":"https:\/\/doi.org\/10.1145\/3725213","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,20]]},"assertion":[{"value":"2023-11-13","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-08-30","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"3725213"}}