{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,24]],"date-time":"2026-08-24T17:24:56Z","timestamp":1787592296675,"version":"build-2736575974"},"reference-count":90,"publisher":"Association for Computing Machinery (ACM)","issue":"OOPSLA1","license":[{"start":{"date-parts":[[2025,4,9]],"date-time":"2025-04-09T00:00:00Z","timestamp":1744156800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Program. Lang."],"published-print":{"date-parts":[[2025,4,9]]},"abstract":"<jats:p>Large code models (LCMs), pre-trained on vast code corpora, have demonstrated remarkable performance across a wide array of code-related tasks. Supervised fine-tuning (SFT) plays a vital role in aligning these models with specific requirements and enhancing their performance in particular domains. However, synthesizing high-quality SFT datasets poses a significant challenge due to the uneven quality of datasets and the scarcity of domain-specific datasets.<\/jats:p>\n                  <jats:p>\n                    Inspired by APIs as high-level abstractions of code that encapsulate rich semantic information in a concise structure, we propose\n                    <jats:sc>DataScope<\/jats:sc>\n                    , an API-guided dataset synthesis framework designed to enhance the SFT process for LCMs in both general and domain-specific scenarios.\n                    <jats:sc>DataScope<\/jats:sc>\n                    comprises two main components:\n                    <jats:sc>Dslt<\/jats:sc>\n                    and\n                    <jats:sc>Dgen<\/jats:sc>\n                    . On the one hand,\n                    <jats:sc>Dslt<\/jats:sc>\n                    employs API coverage as a core metric, enabling efficient dataset synthesis in general scenarios by selecting subsets of existing (uneven-quality) datasets with higher API coverage. On the other hand,\n                    <jats:sc>Dgen<\/jats:sc>\n                    recasts domain dataset synthesis as a process of using API-specified high-level functionality and deliberately constituted code skeletons to synthesize concrete code.\n                  <\/jats:p>\n                  <jats:p>\n                    Extensive experiments demonstrate\n                    <jats:sc>DataScope\u2019s<\/jats:sc>\n                    effectiveness, with models fine-tuned on its synthesized datasets outperforming those tuned on unoptimized datasets five times larger. Furthermore, a series of analyses on model internals, relevant hyperparameters, and case studies provide additional evidence for the efficacy of our proposed methods. These findings underscore the significance of dataset quality in SFT and advance the field of LCMs by providing an efficient, cost-effective framework for constructing high-quality datasets, which in turn lead to more powerful and tailored LCMs for both general and domain-specific scenarios.\n                  <\/jats:p>","DOI":"10.1145\/3720449","type":"journal-article","created":{"date-parts":[[2025,4,9]],"date-time":"2025-04-09T13:48:26Z","timestamp":1744206506000},"page":"786-815","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["API-Guided Dataset Synthesis to Finetune Large Code Models"],"prefix":"10.1145","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9897-4086","authenticated-orcid":false,"given":"Zongjie","family":"Li","sequence":"first","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3752-0718","authenticated-orcid":false,"given":"Daoyuan","family":"Wu","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0866-0308","authenticated-orcid":false,"given":"Shuai","family":"Wang","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2970-1391","authenticated-orcid":false,"given":"Zhendong","family":"Su","sequence":"additional","affiliation":[{"name":"ETH Zurich, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,4,9]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1002\/wics.101"},{"key":"e_1_3_1_3_2","unstructured":"Bo Adler Niket Agarwal Ashwath Aithal Dong H Anh Pallab Bhattacharya Annika Brundyn Jared Casper Bryan Catanzaro Sharon Clay Jonathan Cohen et al. 2024. Nemotron-4 340B Technical Report. arXiv preprint arXiv:2406.11704 (2024)."},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","unstructured":"Ali Al-Kaswan Maliheh Izadi and Arie Van Deursen. 2024. Traces of memorisation in large language models for code. In Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. 1\u201312.","DOI":"10.1145\/3597503.3639133"},{"key":"e_1_3_1_5_2","unstructured":"Jacob Austin Augustus Odena Maxwell Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie Cai Michael Terry Quoc Le et al. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732 (2021)."},{"key":"e_1_3_1_6_2","unstructured":"Jinze Bai Shuai Bai Yunfei Chu Zeyu Cui Kai Dang Xiaodong Deng Yang Fan Wenbin Ge Yu Han Fei Huang et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609 (2023)."},{"key":"e_1_3_1_7_2","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020) 1877\u20131901."},{"key":"e_1_3_1_8_2","unstructured":"Collin Burns Pavel Izmailov Jan Hendrik Kirchner Bowen Baker Leo Gao Leopold Aschenbrenner Yining Chen Adrien Ecoffet Manas Joglekar Jan Leike et al. 2023. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. arXiv preprint arXiv:2312.09390 (2023)."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1080\/03610927408827101"},{"key":"e_1_3_1_10_2","article-title":"Instruction mining: When data mining meets large language model finetuning","volume":"2307","author":"Cao Yihan","year":"2023","unstructured":"Yihan Cao, Yanbin Kang, Chi Wang, and Lichao Sun. 2023. Instruction mining: When data mining meets large language model finetuning. arXiv preprint arXiv 2307 (2023).","journal-title":"arXiv preprint arXiv"},{"key":"e_1_3_1_11_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)."},{"key":"e_1_3_1_12_2","unstructured":"Daixuan Cheng Shaohan Huang and Furu Wei. 2024. Adapting Large Language Models via Reading Comprehension. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=y886UXPEZ0"},{"issue":"70","key":"e_1_3_1_13_2","first-page":"1","article-title":"Scaling instruction-finetuned language models","volume":"25","author":"Chung Hyung Won","year":"2024","unstructured":"Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research 25, 70 (2024), 1\u201353.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_14_2","unstructured":"codefuse ai. 2023. CodeExercise-Python-27k. https:\/\/huggingface.co\/datasets\/codefuse-ai\/CodeExercise-Python-27k\/"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3148148"},{"key":"e_1_3_1_16_2","unstructured":"Defog AI. 2024. Open-sourcing SQLCoder2-15b and SQLCoder-7b. https:\/\/defog.ai\/blog\/open-sourcing-sqlcoder2-7b\/"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3649825"},{"key":"e_1_3_1_18_2","unstructured":"Guanting Dong Hongyi Yuan Keming Lu Chengpeng Li Mingfeng Xue Dayiheng Liu Wei Wang Zheng Yuan Chang Zhou and Jingren Zhou. 2023. How abilities in large language models are affected by supervised fine-tuning data composition. arXiv preprint arXiv:2310.05492 (2023)."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3296979.3192382"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3009837.3009851"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1037\/h0031619"},{"key":"e_1_3_1_22_2","doi-asserted-by":"crossref","unstructured":"Gordon Fraser and Andrea Arcuri. 2011. Evosuite: automatic test suite generation for object-oriented software. In Proceedings of the 19th ACM SIGSOFT symposium and the 13th European conference on Foundations of software engineering. 416\u2013419.","DOI":"10.1145\/2025113.2025179"},{"key":"e_1_3_1_23_2","unstructured":"Yuan Ge Yilun Liu Chi Hu Weibin Meng Shimin Tao Xiaofeng Zhao Hongxia Ma Li Zhang Hao Yang and Tong Xiao. 2024. Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation. arXiv preprint arXiv:2402.18191 (2024)."},{"key":"e_1_3_1_24_2","unstructured":"Sreyan Ghosh Chandra Kiran Reddy Evuru Sonal Kumar Deepali Aneja Zeyu Jin Ramani Duraiswami Dinesh Manocha et al. 2024. A Closer Look at the Limitations of Instruction Tuning. arXiv preprint arXiv:2402.05119 (2024)."},{"key":"e_1_3_1_25_2","unstructured":"Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics. JMLR Workshop and Conference Proceedings 249\u2013256."},{"key":"e_1_3_1_26_2","unstructured":"Google Cloud. 2023. Supercharging security with generative AI. https:\/\/cloud.google.com\/blog\/products\/identitysecurity\/rsa-google-cloud-security-ai-workbench-generative-ai"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3519939.3523450"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3371080"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3591285"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/2491956.2462192"},{"key":"e_1_3_1_31_2","unstructured":"Zeyu Han Chao Gao Jinyang Liu Sai Qian Zhang et al. 2024. Parameter-efficient fine-tuning for large models: A comprehensive survey. arXiv preprint arXiv:2403.14608 (2024)."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.2307\/2346830"},{"key":"e_1_3_1_33_2","unstructured":"Zhenlan Ji Pingchuan Ma Zongjie Li and Shuai Wang. 2023. Benchmarking and Explaining Large Language Modelbased Code Generation: A Causality-Centric Approach. arXiv e-prints (2023) arXiv\u20132310.06680."},{"key":"e_1_3_1_34_2","unstructured":"Wenyu Jiang Zhenlong Liu Zejian Xie Songxin Zhang Bingyi Jing and Hongxin Wei. 2024. Exploring Learning Complexity for Downstream Data Pruning. arXiv preprint arXiv:2402.05356 (2024)."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1613"},{"key":"e_1_3_1_36_2","article-title":"Openassistant conversations-democratizing large language model alignment","volume":"36","author":"K\u00f6pf Andreas","year":"2024","unstructured":"Andreas K\u00f6pf, Yannic Kilcher, Dimitri von R\u00fctte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Rich\u00e1rd Nagyfi, et al. 2024. Openassistant conversations-democratizing large language model alignment. Advances in Neural Information Processing Systems 36 (2024).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_37_2","unstructured":"Yuhang Lai Chengxi Li Yiming Wang Tianyi Zhang Ruiqi Zhong Luke Zettlemoyer Wen-tau Yih Daniel Fried Sida Wang and Tao Yu. 2023. DS-1000: a natural and reliable benchmark for data science code generation. In Proceedings of the 40th International Conference on Machine Learning (Honolulu Hawaii USA) (ICML\u201923). JMLR.org Article 756 27 pages."},{"key":"e_1_3_1_38_2","unstructured":"Bin Lei Yuchen Li and Qiuwu Chen. 2024. AutoCoder: Enhancing Code Large Language Model with AIEV-INSTRUCT. arXiv preprint arXiv:2405.14906 (2024)."},{"key":"e_1_3_1_39_2","unstructured":"Raymond Li Loubna Ben Allal Yangtian Zi Niklas Muennighoff Denis Kocetkov Chenghao Mou Marc Marone Christopher Akiki Jia Li Jenny Chim et al. 2023. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161 (2023)."},{"key":"e_1_3_1_40_2","unstructured":"Raymond Li Loubna Ben Allal Yangtian Zi Niklas Muennighoff Denis Kocetkov Chenghao Mou Marc Marone Christopher Akiki Jia Li Jenny Chim et al. 2023. StarCoder: may the source be with you! arXiv preprint arXiv:2305.06161 (2023)."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00110"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.621"},{"key":"e_1_3_1_43_2","unstructured":"Zongjie Li Daoyuan Wu Shuai Wang and Zhendong Su. 2024. API-Guided Dataset Synthesis to Finetune Large Code Models. arXiv preprint arXiv:2408.08343 (2024)."},{"key":"e_1_3_1_44_2","unstructured":"Yilun Liu Shimin Tao Xiaofeng Zhao Ming Zhu Wenbing Ma Junhao Zhu Chang Su Yutai Hou Miao Zhang Min Zhang et al. 2023. Automatic instruction optimization for open-source llm instruction tuning. arXiv preprint arXiv:2311.13246 (2023)."},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","unstructured":"Yilun Liu Shimin Tao Xiaofeng Zhao Ming Zhu Wenbing Ma Junhao Zhu Chang Su Yutai Hou Miao Zhang Min Zhang Hongxia Ma Li Zhang Hao Yang and Yanfei Jiang. 2024. CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning. arXiv:2311.13246","DOI":"10.1109\/ICDE60146.2024.00390"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE.2015.7381798"},{"key":"e_1_3_1_47_2","unstructured":"Ziyang Luo Can Xu Pu Zhao Qingfeng Sun Xiubo Geng Wenxiang Hu Chongyang Tao Jing Ma Qingwei Lin and Daxin Jiang. 2023. WizardCoder: Empowering Code Large Language Models with Evol-Instruct. CoRR abs\/2306.08568 (2023)."},{"key":"e_1_3_1_48_2","unstructured":"Ziyang Luo Can Xu Pu Zhao Qingfeng Sun Xiubo Geng Wenxiang Hu Chongyang Tao Jing Ma Qingwei Lin and Daxin Jiang. 2023. Wizardcoder: Empowering code large language models with evol-instruct. arXiv preprint arXiv:2306.08568 (2023)."},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","unstructured":"Thomas McCabe. 1977. A Complexity Measure. Software Engineering IEEE Transactions on SE-2 (01 1977) 308\u2013320. doi:10.1109\/TSE.1976.233837","DOI":"10.1109\/TSE.1976.233837"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3632858"},{"key":"e_1_3_1_51_2","doi-asserted-by":"crossref","unstructured":"Tomoki Nakamaru Tomomasa Matsunaga Tetsuro Yamazaki Soramichi Akiyama and Shigeru Chiba. 2020. An empirical study of method chaining in java. In Proceedings of the 17th International Conference on Mining Software Repositories. 93\u2013102.","DOI":"10.1145\/3379597.3387441"},{"key":"e_1_3_1_52_2","doi-asserted-by":"crossref","unstructured":"Sydney Nguyen Hannah McLean Babe Yangtian Zi Arjun Guha Carolyn Jane Anderson and Molly Q Feldman. 2024. How Beginning Programmers and Code LLMs (Mis) read Each Other. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1\u201326.","DOI":"10.1145\/3613904.3642706"},{"key":"e_1_3_1_53_2","unstructured":"OpenAI. 2023. Codex. https:\/\/openai.com\/blog\/openai-codex\/"},{"key":"e_1_3_1_54_2","unstructured":"OpenAI. 2023. gpt4. https:\/\/cdn.openai.com\/papers\/gpt-4-system-card.pdf"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-2011"},{"key":"e_1_3_1_56_2","doi-asserted-by":"crossref","unstructured":"Carlos Pacheco and Michael D Ernst. 2007. Randoop: feedback-directed random testing for Java. In Companion to the 22nd ACM SIGPLAN conference on Object-oriented programming systems and applications companion. 815\u2013816.","DOI":"10.1145\/1297846.1297902"},{"key":"e_1_3_1_57_2","unstructured":"Pecan. 2024. Pecan GenAI Business. https:\/\/www.pecan.ai\/"},{"key":"e_1_3_1_58_2","unstructured":"Baolin Peng Chunyuan Li Pengcheng He Michel Galley and Jianfeng Gao. 2023. Instruction Tuning with GPT-4. arXiv preprint arXiv:2304.03277 (2023)."},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1037\/1082-989X.2.4.329"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/2254064.2254098"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/2980983.2908093"},{"key":"e_1_3_1_62_2","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans Ilya Sutskever et al. 2018. Improving language understanding by generative pre-training. (2018)."},{"key":"e_1_3_1_63_2","article-title":"Direct preference optimization: Your language model is secretly a reward model","volume":"36","author":"Rafailov Rafael","year":"2024","unstructured":"Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems 36 (2024).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_64_2","doi-asserted-by":"crossref","unstructured":"Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084 (2019).","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_1_65_2","unstructured":"Shuo Ren Daya Guo Shuai Lu Long Zhou Shujie Liu Duyu Tang Neel Sundaresan Ming Zhou Ambrosio Blanco and Shuai Ma. 2020. Codebleu: a method for automatic evaluation of code synthesis. arXiv preprint arXiv:2009.10297 (2020)."},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1016\/0377-0427(87)90125-7"},{"key":"e_1_3_1_67_2","unstructured":"Baptiste Roziere Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing Ellen Tan Yossi Adi Jingyu Liu Tal Remez J\u00e9r\u00e9my Rapin et al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950 (2023)."},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2017.41"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-023-06291-2"},{"key":"e_1_3_1_70_2","doi-asserted-by":"crossref","unstructured":"Armando Solar-Lezama Liviu Tancau Rastislav Bodik Sanjit Seshia and Vijay Saraswat. 2006. Combinatorial sketching for finite programs. In Proceedings of the 12th international conference on Architectural support for programming languages and operating systems. 404\u2013415.","DOI":"10.1145\/1168857.1168907"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/3632904"},{"key":"e_1_3_1_72_2","unstructured":"Stack Exchange Team. 2023. stackexchange QA communities. https:\/\/huggingface.co\/datasets\/codefuse-ai\/CodeExercise-Python-27k\/"},{"key":"e_1_3_1_73_2","doi-asserted-by":"crossref","unstructured":"Zhensu Sun Xiaoning Du Fu Song Shangwen Wang and Li Li. 2024. When Neural Code Completion Models Size up the Situation: Attaining Cheaper and Faster Completion through Dynamic Model Inference. In Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. 1\u201312.","DOI":"10.1145\/3597503.3639120"},{"key":"e_1_3_1_74_2","unstructured":"Surendrabikram Thapa Usman Naseem and Mehwish Nasim. 2023. From humans to machines: can ChatGPT-like LLMs effectively replace human annotators in NLP tasks. In Workshop Proceedings of the 17th International AAAI Conference on Web and Social Media."},{"key":"e_1_3_1_75_2","doi-asserted-by":"crossref","unstructured":"Saad Ullah Mingji Han Saurabh Pujar Hammond Pearce Ayse Coskun and Gianluca Stringhini. 2024. LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation Framework and Benchmarks. In IEEE Symposium on Security and Privacy.","DOI":"10.1109\/SP54263.2024.00210"},{"issue":"11","key":"e_1_3_1_76_2","article-title":"Visualizing data using t-SNE","volume":"9","author":"Van der Maaten Laurens","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of machine learning research 9, 11 (2008).","journal-title":"Journal of machine learning research"},{"key":"e_1_3_1_77_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017).","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_1_78_2","unstructured":"Chaozheng Wang Zongjie Li Cuiyun Gao Wenxuan Wang Ting Peng Hailiang Huang Yuetang Deng Shuai Wang and Michael R Lyu. 2024. Exploring Multi-Lingual Bias of Large Code Models in Code Generation. arXiv preprint arXiv:2404.19368 (2024)."},{"key":"e_1_3_1_79_2","unstructured":"Peiqi Wang Yikang Shen Zhen Guo Matthew Stallone Yoon Kim Polina Golland and Rameswar Panda. 2024. Diversity Measurement and Subset Selection for Instruction Tuning Datasets. arXiv preprint arXiv:2402.02318 (2024)."},{"key":"e_1_3_1_80_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3158151","article-title":"Program synthesis using abstraction refinement","volume":"2","author":"Wang Xinyu","year":"2017","unstructured":"Xinyu Wang, Isil Dillig, and Rishabh Singh. 2017. Program synthesis using abstraction refinement. Proceedings of the ACM on Programming Languages 2, POPL (2017), 1\u201330.","journal-title":"Proceedings of the ACM on Programming Languages"},{"key":"e_1_3_1_81_2","doi-asserted-by":"crossref","unstructured":"Yizhong Wang Yeganeh Kordi Swaroop Mishra Alisa Liu Noah A Smith Daniel Khashabi and Hannaneh Hajishirzi. 2023. Self-Instruct: Aligning Language Models with Self-Generated Instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 13484\u201313508.","DOI":"10.18653\/v1\/2023.acl-long.754"},{"key":"e_1_3_1_82_2","unstructured":"Jason Wei Maarten Bosma Vincent Y Zhao Kelvin Guu Adams Wei Yu Brian Lester Nan Du Andrew M Dai and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652 (2021)."},{"key":"e_1_3_1_83_2","unstructured":"Yuxiang Wei Zhe Wang Jiawei Liu Yifeng Ding and Lingming Zhang. [n. d.]. Magicoder: Empowering Code Generation with OSS-Instruct. In Forty-first International Conference on Machine Learning."},{"key":"e_1_3_1_84_2","unstructured":"Yuxiang Wei Zhe Wang Jiawei Liu Yifeng Ding and Lingming Zhang. 2023. Magicoder: Source code is all you need. arXiv preprint arXiv:2312.02120 (2023)."},{"key":"e_1_3_1_85_2","unstructured":"Can Xu Qingfeng Sun Kai Zheng Xiubo Geng Pu Zhao Jiazhan Feng Chongyang Tao and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244 (2023)."},{"key":"e_1_3_1_86_2","doi-asserted-by":"crossref","unstructured":"Shin Yoo and Mark Harman. 2007. Pareto efficient multi-objective test case selection. In Proceedings of the 2007 international symposium on Software testing and analysis. 140\u2013150.","DOI":"10.1145\/1273463.1273483"},{"key":"e_1_3_1_87_2","unstructured":"Biao Zhang Zhongtao Liu Colin Cherry and Orhan Firat. 2024. When Scaling Meets LLM Finetuning: The Effect of Data Model and Finetuning Method. arXiv preprint arXiv:2402.17193 (2024)."},{"key":"e_1_3_1_88_2","doi-asserted-by":"publisher","DOI":"10.1109\/SP54263.2024.00208"},{"key":"e_1_3_1_89_2","unstructured":"Lianmin Zheng Wei-Lin Chiang Ying Sheng Tianle Li Siyuan Zhuang Zhanghao Wu Yonghao Zhuang Zhuohan Li Zi Lin Eric Xing et al. 2023. Lmsys-chat-1m: A large-scale real-world llm conversation dataset. arXiv preprint arXiv:2309.11998 (2023)."},{"key":"e_1_3_1_90_2","article-title":"Judging llm-as-a-judge with mt-bench and chatbot arena","volume":"36","author":"Zheng Lianmin","year":"2024","unstructured":"Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems 36 (2024).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_91_2","article-title":"Lima: Less is more for alignment","volume":"36","author":"Zhou Chunting","year":"2024","unstructured":"Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2024. Lima: Less is more for alignment. Advances in Neural Information Processing Systems 36 (2024).","journal-title":"Advances in Neural Information Processing Systems"}],"container-title":["Proceedings of the ACM on Programming Languages"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3720449","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3720449","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,24]],"date-time":"2026-08-24T16:30:10Z","timestamp":1787589010000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3720449"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,9]]},"references-count":90,"journal-issue":{"issue":"OOPSLA1","published-print":{"date-parts":[[2025,4,9]]}},"alternative-id":["10.1145\/3720449"],"URL":"https:\/\/doi.org\/10.1145\/3720449","relation":{},"ISSN":["2475-1421"],"issn-type":[{"value":"2475-1421","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,9]]},"assertion":[{"value":"2024-10-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-18","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}