{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T06:10:26Z","timestamp":1768284626187,"version":"3.49.0"},"reference-count":91,"publisher":"Association for Computing Machinery (ACM)","issue":"ISSTA","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,22]]},"abstract":"<jats:p>While code generation has been widely used in various software development scenarios, the quality of the generated code is not guaranteed. This has been a particular concern in the era of large language models (LLM)-based code generation, where LLMs, deemed a complex and powerful black-box model, are instructed by a high-level natural language specification, namely a prompt, to generate code. Nevertheless, effectively evaluating and explaining the code generation capability of LLMs is inherently challenging, given the complexity of LLMs and the lack of transparency.<\/jats:p>\n          <jats:p>Inspired by recent progress in causality analysis and its software engineering applications, this paper proposes a causality-driven approach to systematically analyze prompt-code causal relationships. However, this endeavor faces three key technical challenges: (1) representing textual prompts and code in a canonical form, (2) establishing causal relations between high-level concepts and code features, and (3) systematically analyzing diverse prompt variations. To address these challenges, we first propose a novel causal graph-based representation of the prompt and the generated code, which is established over the fine-grained, human-understandable concepts in the input prompts. The formed causal graph is then used to identify the causal relations between the prompt and the derived code. We illustrate the insights that our framework can provide by studying over four popular LLMs with over 12 prompt adjustment strategies. The results of these studies illustrate the potential of our technique to provide insights into LLM effectiveness and aid end-users in understanding predictions. Additionally, we demonstrate that our approach provides actionable insights to improve the quality of the LLM-generated code by properly calibrating the prompt.<\/jats:p>","DOI":"10.1145\/3728938","type":"journal-article","created":{"date-parts":[[2025,6,22]],"date-time":"2025-06-22T10:52:56Z","timestamp":1750589576000},"page":"1374-1397","source":"Crossref","is-referenced-by-count":1,"title":["Causality-Aided Evaluation and Explanation of Large Language Model-Based Code Generation"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3167-0480","authenticated-orcid":false,"given":"Zhenlan","family":"Ji","sequence":"first","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7680-2817","authenticated-orcid":false,"given":"Pingchuan","family":"Ma","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9897-4086","authenticated-orcid":false,"given":"Zongjie","family":"Li","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-6892-1264","authenticated-orcid":false,"given":"Zhaoyu","family":"Wang","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0866-0308","authenticated-orcid":false,"given":"Shuai","family":"Wang","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,6,22]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"[n. d.]. Black - The uncompromising code formatter. https:\/\/black.readthedocs.io\/en\/stable\/"},{"key":"e_1_2_1_2_1","unstructured":"[n. d.]. Semgrep \u2014 Find bugs and enforce code standards. https:\/\/semgrep.dev\/"},{"key":"e_1_2_1_3_1","unstructured":"[n. d.]. Tree-sitter. https:\/\/tree-sitter.github.io\/tree-sitter\/"},{"key":"e_1_2_1_4_1","unstructured":"2022. EconML: A Python Package for ML-Based Heterogeneous Treatment Effects Estimation. https:\/\/github.com\/microsoft\/EconML"},{"key":"e_1_2_1_5_1","unstructured":"2023. ChatGPT. https:\/\/chat.openai.com\/chat"},{"key":"e_1_2_1_6_1","unstructured":"2023. Rephrasing Tool - QuillBot AI. https:\/\/quillbot.com\/"},{"key":"e_1_2_1_7_1","unstructured":"2023. Research Artifact. https:\/\/anonymous.4open.science\/r\/CALL-E7F2\/"},{"key":"e_1_2_1_8_1","first-page":"6022","article-title":"Towards Robust NLG Bias Evaluation with Syntactically-diverse Prompts","volume":"2022","author":"Aggarwal Arshiya","year":"2022","unstructured":"Arshiya Aggarwal, Jiao Sun, and Nanyun Peng. 2022. Towards Robust NLG Bias Evaluation with Syntactically-diverse Prompts. In Findings of the Association for Computational Linguistics: EMNLP 2022. 6022\u20136032.","journal-title":"Findings of the Association for Computational Linguistics: EMNLP"},{"key":"e_1_2_1_9_1","unstructured":"Jacob Austin Augustus Odena Maxwell Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie Cai Michael Terry and Quoc Le. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732."},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security. 249\u2013262","author":"Baluta Teodora","year":"2022","unstructured":"Teodora Baluta, Shiqi Shen, S Hitarth, Shruti Tople, and Prateek Saxena. 2022. Membership inference attacks and generalization: A causal perspective. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security. 249\u2013262."},{"key":"e_1_2_1_11_1","volume-title":"Advances in NeurIPS","author":"Brown Tom","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Advances in NeurIPS, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.). 33, Curran Associates, Inc.."},{"key":"e_1_2_1_12_1","volume-title":"Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, and Greg Brockman.","author":"Chen Mark","year":"2021","unstructured":"Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, and Greg Brockman. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374."},{"key":"e_1_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Victor Chernozhukov Denis Chetverikov Mert Demirer Esther Duflo Christian Hansen Whitney Newey and James Robins. 2016. Double\/debiased machine learning for treatment and causal parameters. arXiv preprint arXiv:1608.00060.","DOI":"10.3386\/w23564"},{"key":"e_1_2_1_14_1","unstructured":"Cheng-Han Chiang and Hung-yi Lee. 2023. Can large language models be an alternative to human evaluations? arXiv preprint arXiv:2305.01937."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/3176764.3176769"},{"key":"e_1_2_1_16_1","unstructured":"Yilun Du Shuang Li Antonio Torralba Joshua B Tenenbaum and Igor Mordatch. 2023. Improving Factuality and Reasoning in Language Models through Multiagent Debate. arXiv preprint arXiv:2305.14325."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510200"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00128"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389694"},{"key":"e_1_2_1_20_1","volume-title":"Measuring nominal scale agreement among many raters.. Psychological bulletin, 76, 5","author":"Fleiss Joseph L","year":"1971","unstructured":"Joseph L Fleiss. 1971. Measuring nominal scale agreement among many raters.. Psychological bulletin, 76, 5 (1971), 378."},{"key":"e_1_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Carlo A Furia Richard Torkar and Robert Feldt. 2023. Towards Causal Analysis of Empirical Software Engineering Data: The Impact of Programming Languages on Coding Competitions. arXiv preprint arXiv:2301.07524.","DOI":"10.1145\/3611667"},{"key":"e_1_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Luca Giamattei Antonio Guerriero Roberto Pietrantuono and Stefano Russo. 2024. Causal reasoning in Software Quality Assurance: A systematic review. Information and Software Technology 107599.","DOI":"10.1016\/j.infsof.2024.107599"},{"key":"e_1_2_1_23_1","volume-title":"Readings in Artificial Intelligence","author":"Green Cordell","unstructured":"Cordell Green. 1981. Application of theorem proving to problem solving. In Readings in Artificial Intelligence. Elsevier, 202\u2013222."},{"key":"e_1_2_1_24_1","volume-title":"Program synthesis. Foundations and Trends\u00ae in Programming Languages, 4, 1-2","author":"Gulwani Sumit","year":"2017","unstructured":"Sumit Gulwani, Oleksandr Polozov, and Rishabh Singh. 2017. Program synthesis. Foundations and Trends\u00ae in Programming Languages, 4, 1-2 (2017), 1\u2013119."},{"key":"e_1_2_1_25_1","unstructured":"Dan Hendrycks Steven Basart Saurav Kadavath Mantas Mazeika Akul Arora Ethan Guo Collin Burns Samir Puranik Horace He Dawn Song and Jacob Steinhardt. 2021. Measuring Coding Challenge Competence With APPS. NeurIPS."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3093336.3037712"},{"key":"e_1_2_1_27_1","unstructured":"Binyuan Hui Jian Yang Zeyu Cui Jiaxi Yang Dayiheng Liu Lei Zhang Tianyu Liu Jiajun Zhang Bowen Yu and Kai Dang. 2024. Qwen2. 5-Coder Technical Report. arXiv preprint arXiv:2409.12186."},{"key":"e_1_2_1_28_1","unstructured":"Hamel Husain Ho-Hsiang Wu Tiferet Gazit Miltiadis Allamanis and Marc Brockschmidt. 2019. Codesearchnet challenge: Evaluating the state of semantic code search. arXiv preprint arXiv:1909.09436."},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2365\u20132376","author":"Ishibashi Yoichi","year":"2023","unstructured":"Yoichi Ishibashi, Danushka Bollegala, Katsuhito Sudoh, and Satoshi Nakamura. 2023. Evaluating the Robustness of Discrete Prompts. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2365\u20132376."},{"key":"e_1_2_1_30_1","volume-title":"2023 38th IEEE\/ACM International Conference on Automated Software Engineering (ASE). 1454\u20131466","author":"Ji Zhenlan","year":"2023","unstructured":"Zhenlan Ji, Pingchuan Ma, and Shuai Wang. 2023. Perfce: Performance debugging on databases with chaos engineering-enhanced causality analysis. In 2023 38th IEEE\/ACM International Conference on Automated Software Engineering (ASE). 1454\u20131466."},{"key":"e_1_2_1_31_1","volume-title":"2023 38th IEEE\/ACM International Conference on Automated Software Engineering (ASE). 371\u2013383","author":"Ji Zhenlan","year":"2023","unstructured":"Zhenlan Ji, Pingchuan Ma, Shuai Wang, and Yanhui Li. 2023. Causality-aided trade-off analysis for machine learning fairness. In 2023 38th IEEE\/ACM International Conference on Automated Software Engineering (ASE). 371\u2013383."},{"key":"e_1_2_1_32_1","volume-title":"CC: Causality-Aware Coverage Criterion for Deep Neural Networks. In 2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). 1788\u20131800","author":"Ji Zhenlan","year":"2023","unstructured":"Zhenlan Ji, Pingchuan Ma, Yuanyuan Yuan, and Shuai Wang. 2023. CC: Causality-Aware Coverage Criterion for Deep Neural Networks. In 2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). 1788\u20131800."},{"key":"e_1_2_1_33_1","unstructured":"Juyong Jiang Fan Wang Jiasi Shen Sungju Kim and Sunghun Kim. 2024. A survey on large language models for code generation. arXiv preprint arXiv:2406.00515."},{"key":"e_1_2_1_34_1","unstructured":"Wenxiang Jiao Wenxuan Wang Jen-tse Huang Xing Wang and Zhaopeng Tu. 2023. Is ChatGPT a good translator? A preliminary study. arXiv preprint arXiv:2301.08745."},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the ACM\/IEEE 42nd International Conference on Software Engineering. 87\u201399","author":"Johnson Brittany","year":"2020","unstructured":"Brittany Johnson, Yuriy Brun, and Alexandra Meliou. 2020. Causal testing: understanding defects\u2019 root causes. In Proceedings of the ACM\/IEEE 42nd International Conference on Software Engineering. 87\u201399."},{"key":"e_1_2_1_36_1","first-page":"21314","article-title":"Coderl: Mastering code generation through pretrained models and deep reinforcement learning","volume":"35","author":"Le Hung","year":"2022","unstructured":"Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven Chu Hong Hoi. 2022. Coderl: Mastering code generation through pretrained models and deep reinforcement learning. Advances in Neural Information Processing Systems, 35 (2022), 21314\u201321328.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 10669\u201310686","author":"Lee Bruce W","year":"2021","unstructured":"Bruce W Lee, Yoo Sung Jang, and Jason Lee. 2021. Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic Features. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 10669\u201310686."},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA","author":"Bruce","year":"2023","unstructured":"Bruce W. Lee and Jason Lee. 2023. LFTK: Handcrafted Features in Computational Linguistics. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023). 1\u201319."},{"key":"e_1_2_1_39_1","unstructured":"Yujia Li David Choi Junyoung Chung Nate Kushman Julian Schrittwieser R\u00e9mi Leblond Tom Eccles James Keeling Felix Gimeno Agustin Dal Lago Thomas Hubert Peter Choy Cyprien de Masson d\u2019Autume Igor Babuschkin Xinyun Chen Po-Sen Huang Johannes Welbl Sven Gowal Alexey Cherepanov James Molloy Daniel Mankowitz Esme Sutherland Robson Pushmeet Kohli Nando de Freitas Koray Kavukcuoglu and Oriol Vinyals. 2022. Competition-Level Code Generation with AlphaCode. arXiv preprint arXiv:2203.07814."},{"key":"e_1_2_1_40_1","volume-title":"2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). 1238\u20131250","author":"Li Zongjie","year":"2023","unstructured":"Zongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang, Dong Chen, Shuai Wang, and Cuiyun Gao. 2023. Cctest: Testing and repairing code completion systems. In 2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). 1238\u20131250."},{"key":"e_1_2_1_41_1","volume-title":"CCTEST: Testing and Repairing Code Completion Systems. In 45th IEEE\/ACM International Conference on Software Engineering, ICSE 2023","author":"Li Zongjie","year":"2023","unstructured":"Zongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang, Dong Chen, Shuai Wang, and Cuiyun Gao. 2023. CCTEST: Testing and Repairing Code Completion Systems. In 45th IEEE\/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023. IEEE, 1238\u20131250."},{"key":"e_1_2_1_42_1","volume-title":"Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering (ICSE \u201924)","author":"Li Zongjie","year":"2024","unstructured":"Zongjie Li, Chaozheng Wang, Pingchuan Ma, Chaowei Liu, Shuai Wang, Daoyuan Wu, Cuiyun Gao, and Yang Liu. 2024. On Extracting Specialized Code Abilities from Large Language Models: A Feasibility Study. In Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering (ICSE \u201924). Association for Computing Machinery, New York, NY, USA. Article 74, 13 pages."},{"key":"e_1_2_1_43_1","unstructured":"Pengfei Liu Weizhe Yuan Jinlan Fu Zhengbao Jiang Hiroaki Hayashi and Graham Neubig. 2023. Pre-train Prompt and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Comput. Surv.."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3678167"},{"key":"e_1_2_1_45_1","first-page":"24111","article-title":"Dibs: Differentiable bayesian structure learning","volume":"34","author":"Lorch Lars","year":"2021","unstructured":"Lars Lorch, Jonas Rothfuss, Bernhard Sch\u00f6lkopf, and Andreas Krause. 2021. Dibs: Differentiable bayesian structure learning. Advances in Neural Information Processing Systems, 34 (2021), 24111\u201324123.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_46_1","volume-title":"Codexglue: A machine learning benchmark dataset for code understanding and generation. arXiv preprint arXiv:2102.04664.","author":"Lu Shuai","year":"2021","unstructured":"Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, and Duyu Tang. 2021. Codexglue: A machine learning benchmark dataset for code understanding and generation. arXiv preprint arXiv:2102.04664."},{"key":"e_1_2_1_47_1","volume-title":"International journal of corpus linguistics, 15, 4","author":"Xiaofei Lu.","year":"2010","unstructured":"Xiaofei Lu. 2010. Automatic analysis of syntactic complexity in second language writing. International journal of corpus linguistics, 15, 4 (2010), 474\u2013496."},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/362566.362568"},{"key":"e_1_2_1_49_1","unstructured":"Brady Neal. 2015. Introduction to Causal Inference."},{"key":"e_1_2_1_50_1","unstructured":"R OpenAI. 2023. GPT-4 technical report. arXiv 2303\u201308774."},{"key":"e_1_2_1_51_1","unstructured":"Shuyin Ouyang Jie M Zhang Mark Harman and Meng Wang. 2023. LLM is Like a Box of Chocolates: the Non-determinism of ChatGPT in Code Generation. arXiv preprint arXiv:2308.02828."},{"key":"e_1_2_1_52_1","volume-title":"Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311\u2013318","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP46214.2022.9833571"},{"key":"e_1_2_1_54_1","unstructured":"Judea Pearl. 2009. Causality. Cambridge university press."},{"key":"e_1_2_1_55_1","volume-title":"Elements of causal inference: foundations and learning algorithms","author":"Peters Jonas","unstructured":"Jonas Peters, Dominik Janzing, and Bernhard Sch\u00f6lkopf. 2017. Elements of causal inference: foundations and learning algorithms. The MIT Press."},{"key":"e_1_2_1_56_1","first-page":"117","article-title":"Some new test criteria in multivariate analysis","author":"Sreedharan Pillai KC","year":"1955","unstructured":"KC Sreedharan Pillai. 1955. Some new test criteria in multivariate analysis. The Annals of Mathematical Statistics, 117\u2013121.","journal-title":"The Annals of Mathematical Statistics"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-6319"},{"key":"e_1_2_1_58_1","volume-title":"Chenguang Zhu, and Michael Zeng.","author":"Pryzant Reid","year":"2023","unstructured":"Reid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee, Chenguang Zhu, and Michael Zeng. 2023. Automatic prompt optimization with\" gradient descent\" and beam search. arXiv preprint arXiv:2305.03495."},{"key":"e_1_2_1_59_1","unstructured":"Shuo Ren Daya Guo Shuai Lu Long Zhou Shujie Liu Duyu Tang Neel Sundaresan Ming Zhou Ambrosio Blanco and Shuai Ma. 2020. Codebleu: a method for automatic evaluation of code synthesis. arXiv preprint arXiv:2009.10297."},{"key":"e_1_2_1_60_1","volume-title":"Multitask Prompted Training Enables Zero-Shot Task Generalization. In ICLR 2022","author":"Sanh Victor","year":"2022","unstructured":"Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault F\u00e9vry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush. 2022. Multitask Prompted Training Enables Zero-Shot Task Generalization. In ICLR 2022, Virtual Event, April 25-29, 2022."},{"key":"e_1_2_1_61_1","unstructured":"Mauro Scanagatta Cassio P de Campos Giorgio Corani and Marco Zaffalon. 2015. Learning Bayesian Networks with Thousands of Variables.. In NIPS. 1864\u20131872."},{"key":"e_1_2_1_62_1","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 4222\u20134235","author":"Shin Taylor","year":"2020","unstructured":"Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 4222\u20134235."},{"key":"e_1_2_1_63_1","doi-asserted-by":"crossref","unstructured":"Julien Siebert. 2023. Applications of statistical causal inference in software engineering. Information and Software Technology 107198.","DOI":"10.1016\/j.infsof.2023.107198"},{"key":"e_1_2_1_64_1","volume-title":"Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, and Stephen Pfohl.","author":"Singhal Karan","year":"2023","unstructured":"Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, and Stephen Pfohl. 2023. Large language models encode clinical knowledge. Nature, 1\u20139."},{"key":"e_1_2_1_65_1","volume-title":"Program synthesis by sketching","author":"Solar-Lezama Armando","unstructured":"Armando Solar-Lezama. 2008. Program synthesis by sketching. University of California, Berkeley."},{"key":"e_1_2_1_66_1","doi-asserted-by":"crossref","unstructured":"Peter Spirtes Clark N Glymour Richard Scheines and David Heckerman. 2000. Causation prediction and search.","DOI":"10.7551\/mitpress\/1754.001.0001"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510080"},{"key":"e_1_2_1_68_1","unstructured":"Wannita Takerngsaksiri Jirat Pasuksmit Patanamon Thongtanunam Chakkrit Tantithamthavorn Ruixiong Zhang Fan Jiang Jing Li Evan Cook Kun Chen and Ming Wu. 2024. Human-In-the-Loop Software Development Agents. arXiv preprint arXiv:2411.12924."},{"key":"e_1_2_1_69_1","unstructured":"Yashar Talebirad and Amirhossein Nadiri. 2023. Multi-agent collaboration: Harnessing the power of intelligent llm agents. arXiv preprint arXiv:2306.03314."},{"key":"e_1_2_1_70_1","volume-title":"The max-min hill-climbing Bayesian network structure learning algorithm. Machine learning, 65, 1","author":"Tsamardinos Ioannis","year":"2006","unstructured":"Ioannis Tsamardinos, Laura E Brown, and Constantin F Aliferis. 2006. The max-min hill-climbing Bayesian network structure learning algorithm. Machine learning, 65, 1 (2006), 31\u201378."},{"key":"e_1_2_1_71_1","volume-title":"2017 ACM\/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM). 436\u2013441","author":"Tsunoda Masateru","year":"2017","unstructured":"Masateru Tsunoda and Sousuke Amasaki. 2017. On software productivity analysis with propensity score matching. In 2017 ACM\/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM). 436\u2013441."},{"key":"e_1_2_1_72_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2492248.2492277","article-title":"Search based software test data generation for structural testing: a perspective","volume":"38","author":"Varshney Sapna","year":"2013","unstructured":"Sapna Varshney and Monica Mehrotra. 2013. Search based software test data generation for structural testing: a perspective. ACM SIGSOFT Software Engineering Notes, 38, 4 (2013), 1\u20136.","journal-title":"ACM SIGSOFT Software Engineering Notes"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549113"},{"key":"e_1_2_1_74_1","unstructured":"Ruoke Wang Zongjie Li Chaozheng Wang Yang Xiao and Cuiyun Gao. 2024. NAVRepair: Node-type Aware C\/C++ Code Vulnerability Repair. arXiv preprint arXiv:2405.04994."},{"key":"e_1_2_1_75_1","volume-title":"Chi, and Denny Zhou","author":"Wang Xuezhi","year":"2022","unstructured":"Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, and Denny Zhou. 2022. Self-Consistency Improves Chain of Thought Reasoning in Language Models. CoRR, abs\/2203.11171 (2022)."},{"key":"e_1_2_1_76_1","doi-asserted-by":"crossref","unstructured":"Yue Wang Weishi Wang Shafiq Joty and Steven CH Hoi. 2021. Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. arXiv preprint arXiv:2109.00859.","DOI":"10.18653\/v1\/2021.emnlp-main.685"},{"key":"e_1_2_1_77_1","volume-title":"Code generation as a dual task of code summarization. Advances in neural information processing systems, 32","author":"Wei Bolin","year":"2019","unstructured":"Bolin Wei, Ge Li, Xin Xia, Zhiyi Fu, and Zhi Jin. 2019. Code generation as a dual task of code summarization. Advances in neural information processing systems, 32 (2019)."},{"key":"e_1_2_1_78_1","unstructured":"Jason Wei Yi Tay Rishi Bommasani Colin Raffel Barret Zoph Sebastian Borgeaud Dani Yogatama Maarten Bosma Denny Zhou and Donald Metzler. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682."},{"key":"e_1_2_1_79_1","volume-title":"Chi, Quoc Le, and Denny Zhou","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903."},{"key":"e_1_2_1_80_1","unstructured":"Jules White Quchen Fu Sam Hays Michael Sandborn Carlos Olea Henry Gilbert Ashraf Elnashar Jesse Spencer-Smith and Douglas C Schmidt. 2023. A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382."},{"key":"e_1_2_1_81_1","unstructured":"Wai Kin Wong Huaijin Wang Zongjie Li Zhibo Liu Shuai Wang Qiyi Tang Sen Nie and Shi Wu. 2023. Refining decompiled c code with large language models. arXiv preprint arXiv:2310.06530."},{"key":"e_1_2_1_82_1","volume-title":"Umang Bhatt, Willie Neiswanger, Ruslan Salakhutdinov, and Louis-Philippe Morency.","author":"Xiao Yuxin","year":"2022","unstructured":"Yuxin Xiao, Paul Pu Liang, Umang Bhatt, Willie Neiswanger, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2022. Uncertainty quantification with pre-trained language models: A large-scale empirical analysis. arXiv preprint arXiv:2210.04714."},{"key":"e_1_2_1_83_1","unstructured":"An Yang Baosong Yang Binyuan Hui Bo Zheng Bowen Yu Chang Zhou Chengpeng Li Chengyuan Li Dayiheng Liu Fei Huang Guanting Dong Haoran Wei Huan Lin Jialong Tang Jialin Wang Jian Yang Jianhong Tu Jianwei Zhang Jianxin Ma Jin Xu Jingren Zhou Jinze Bai Jinzheng He Junyang Lin Kai Dang Keming Lu Keqin Chen Kexin Yang Mei Li Mingfeng Xue Na Ni Pei Zhang Peng Wang Ru Peng Rui Men Ruize Gao Runji Lin Shijie Wang Shuai Bai Sinan Tan Tianhang Zhu Tianhao Li Tianyu Liu Wenbin Ge Xiaodong Deng Xiaohuan Zhou Xingzhang Ren Xinyu Zhang Xipin Wei Xuancheng Ren Yang Fan Yang Yao Yichang Zhang Yu Wan Yunfei Chu Yuqiong Liu Zeyu Cui Zhenru Zhang and Zhihao Fan. 2024. Qwen2 Technical Report. arXiv preprint arXiv:2407.10671."},{"key":"e_1_2_1_84_1","unstructured":"Chengrun Yang Xuezhi Wang Yifeng Lu Hanxiao Liu Quoc V Le Denny Zhou and Xinyun Chen. 2023. Large language models as optimizers. arXiv preprint arXiv:2309.03409."},{"key":"e_1_2_1_85_1","volume-title":"International Conference on Machine Learning. 7154\u20137163","author":"Yu Yue","year":"2019","unstructured":"Yue Yu, Jie Chen, Tian Gao, and Mo Yu. 2019. DAG-GNN: DAG structure learning with graph neural networks. In International Conference on Machine Learning. 7154\u20137163."},{"key":"e_1_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549103"},{"key":"e_1_2_1_87_1","unstructured":"Yifan Zhang Jingqin Yang Yang Yuan and Andrew Chi-Chih Yao. 2023. Cumulative Reasoning With Large Language Models. arXiv preprint arXiv:2308.04371."},{"key":"e_1_2_1_88_1","volume-title":"Dags with no tears: Continuous optimization for structure learning. Advances in Neural Information Processing Systems, 31","author":"Zheng Xun","year":"2018","unstructured":"Xun Zheng, Bryon Aragam, Pradeep K Ravikumar, and Eric P Xing. 2018. Dags with no tears: Continuous optimization for structure learning. Advances in Neural Information Processing Systems, 31 (2018)."},{"key":"e_1_2_1_89_1","volume-title":"Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba.","author":"Zhou Yongchao","year":"2022","unstructured":"Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022. Large language models are human-level prompt engineers. arXiv preprint arXiv:2211.01910."},{"key":"e_1_2_1_90_1","volume-title":"Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, and Indraneil Paul.","author":"Zhuo Terry Yue","year":"2024","unstructured":"Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, and Indraneil Paul. 2024. BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions. arXiv preprint arXiv:2406.15877."},{"key":"e_1_2_1_91_1","unstructured":"Daniel M Ziegler Nisan Stiennon Jeffrey Wu Tom B Brown Alec Radford Dario Amodei Paul Christiano and Geoffrey Irving. 2019. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593."}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728938","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,16]],"date-time":"2025-07-16T16:51:28Z","timestamp":1752684688000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728938"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,22]]},"references-count":91,"journal-issue":{"issue":"ISSTA","published-print":{"date-parts":[[2025,6,22]]}},"alternative-id":["10.1145\/3728938"],"URL":"https:\/\/doi.org\/10.1145\/3728938","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,22]]}}}