{"status":"ok","message-type":"work-list","message-version":"1.0.0","message":{"facets":{},"total-results":575,"items":[{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:58:00Z","timestamp":1782845880224,"version":"3.54.5"},"reference-count":70,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No. 62402423"],"award-info":[{"award-number":["No. 62402423"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["No. 226202400143"],"award-info":[{"award-number":["No. 226202400143"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Binary code similarity detection (BCSD) serves as a fundamental technique for various software engineering tasks, e.g., vulnerability detection and classification. Attacks against such BCSD models have therefore drawn extensive attention, aiming at misleading the models to generate erroneous predictions. Prior works have explored various approaches to generating semantic-preserving variants, i.e., adversarial samples, to evaluate the robustness of the models against adversarial attacks. However, they have mainly relied on heuristic criteria or iterative greedy algorithms to locate salient code influencing the model output, which often leads to inefficient search and high computational cost. Moreover, when processing programs with high complexities, such attacks tend to be time-consuming.<\/jats:p>\n                  <jats:p>In this work, we unveil the fragility of BCSD models through a novel attack framework guided by model explanations. In particular, we focus on targeted attacks where the attack goal is to mislead the model\u2019s predictions to a specific target. Our attack leverages explainers to pinpoint critical code snippet for perturbations, reducing the exploration overhead. The evaluation results demonstrate that the proposed attacks effectively improve the attack efficiency, while maintaining comparable or higher success rates. Importantly, the speedup for perturbation target selection achieves up to 63.66\u00d7, demonstrating the practical value of explanation-guided localization. Our real-world case studies on vulnerability detection and classification further demonstrate the security implications of our attacks, highlighting fundamental robustness limitations in current BCSD models, and the urgent need for more robust designs.<\/jats:p>","DOI":"10.1145\/3808188","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"4116-4139","source":"Crossref","is-referenced-by-count":0,"title":["Unveiling the Fragility of Binary Code Similarity Detection via Targeted Attacks with Model Explanations"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-6103-5681","authenticated-orcid":false,"given":"Mingjie","family":"Chen","sequence":"first","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-5323-1562","authenticated-orcid":false,"given":"Tiancheng","family":"Zhu","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8863-8751","authenticated-orcid":false,"given":"Mingxue","family":"Zhang","sequence":"additional","affiliation":[{"name":"Zhejiang University, The State Key Laboratory of Blockchain and Data Security, Hangzhou, China"},{"name":"Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5977-1489","authenticated-orcid":false,"given":"Yiling","family":"He","sequence":"additional","affiliation":[{"name":"University College London, London, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-5776-4789","authenticated-orcid":false,"given":"Minghao","family":"Lin","sequence":"additional","affiliation":[{"name":"Independent Researcher, Los Angeles, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3077-5697","authenticated-orcid":false,"given":"Penghui","family":"Li","sequence":"additional","affiliation":[{"name":"Columbia University, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3441-6277","authenticated-orcid":false,"given":"Kui","family":"Ren","sequence":"additional","affiliation":[{"name":"Zhejiang University, The State Key Laboratory of Blockchain and Data Security, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2025. Angr. https:\/\/github.com\/angr\/angr."},{"key":"e_1_2_1_2_1","unstructured":"2025. National Vulnerability Database (NVD). https:\/\/nvd.nist.gov\/."},{"key":"e_1_2_1_3_1","unstructured":"2025. Radare2. https:\/\/github.com\/radareorg\/radare2."},{"key":"e_1_2_1_4_1","unstructured":"2025. Software Assurance Reference Dataset (SARD). https:\/\/samate.nist.gov\/SARD\/."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.14722\/ndss.2014"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3243734.3264418"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/access.2024.3488204"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/eurosp63326.2025.00060"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","unstructured":"Mingjie Chen Tiancheng Zhu Mingxue Zhang Yiling He Minghao Lin Penghui Li and Kui Ren. 2026. Artifact for \"Unveiling the Fragility of Binary Code Similarity Detection via Targeted Attacks with Model Explanations\". doi:10.5281\/zenodo.19709683 10.5281\/zenodo.19709683","DOI":"10.5281\/zenodo.19709683"},{"key":"e_1_2_1_10_1","unstructured":"Mingjie Chen Tiancheng Zhu Mingxue Zhang Yiling He Minghao Lin Penghui Li and Kui Ren. 2026. Explainer- Guided-Adv-Attack-BCSD (Code Repository). https:\/\/github.com\/zju-ws-seclab\/Explainer-Guided-Adv-Attack-BCSD. GitHub repository."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598051"},{"key":"e_1_2_1_12_1","unstructured":"CVE. 2025. CWE Top 25 Most Dangerous Software Weaknesses. https:\/\/cwe.mitre.org\/top25\/."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3672608.3707944"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3428206"},{"key":"e_1_2_1_15_1","unstructured":"FFmpeg. https:\/\/ffmpeg.org\/. Accessed: 2025-05-30."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2024.3392391"},{"key":"e_1_2_1_17_1","unstructured":"Gsl. https:\/\/www.gnu.org\/software\/gsl\/. Accessed: 2025-05-30."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3243734.3243792"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2513228.2513294"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/tdsc.2022.3168285"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3576915.3616599"},{"key":"e_1_2_1_22_1","unstructured":"Hex-Rays. [n. d.]."},{"key":"e_1_2_1_23_1","unstructured":"IDA Pro. https:\/\/www.hex-rays.com\/products\/ida\/. Accessed: 2025-05-30."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2208.14191"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-emnlp.673"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639100"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6311"},{"key":"e_1_2_1_28_1","unstructured":"junk code. 2021. Foudation of CTF reverse engineering. https:\/\/blog.csdn.net\/u011642058\/article\/details\/114757503."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.23919\/EUSIPCO.2018.8553214"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1802.04528"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3460120.3484587"},{"key":"e_1_2_1_33_1","volume-title":"International conference on machine learning. PMLR, 3835-3845","author":"Li Yujia","year":"2019","unstructured":"Yujia Li, Chenjie Gu, Thomas Dullien, Oriol Vinyals, and Pushmeet Kohli. 2019. Graph matching networks for learning the similarity of graph structured objects. In International conference on machine learning. PMLR, 3835-3845."},{"key":"e_1_2_1_34_1","unstructured":"Libconfig Project. [n. d.]. Libconfig. https:\/\/github.com\/hyperrealm\/libconfig. Accessed: 2025-05-30."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3433210.3453086"},{"key":"e_1_2_1_36_1","first-page":"I","article-title":"A Unified Approach to Interpreting Model Predictions","volume":"30","author":"Lundberg Scott M","year":"2017","unstructured":"Scott M Lundberg and Su-In Lee. 2017. A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems 30, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.). Curran Associates, Inc., 4765-4774. https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/ 8a20a8621978632d76c43dfd28b67767-Abstract.html","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_37_1","first-page":"400","article-title":"Parameterized explainer for graph neural network","volume":"33","author":"Luo Dongsheng","year":"2020","unstructured":"Dongsheng Luo, Wei Cheng, Dongkuan Xu, Wenchao Yu, Bo Zong, Haifeng Chen, and Xiang Zhang. 2020. Parameterized explainer for graph neural network. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 33. 400-411. https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/e37b08dd3015330dcbb5d6663667b8b8-Abstract.html","journal-title":"Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/tse.2017.2655046"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/tdsc.2021.3051852"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/1273442.1250746"},{"key":"e_1_2_1_41_1","unstructured":"OpenSSL Project. [n. d.]. OpenSSL: The Open Source Toolkit for SSL\/TLS. https:\/\/www.openssl.org\/. Accessed: 2025-05-30."},{"key":"e_1_2_1_42_1","volume-title":"Proceedings of the 36th International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research","volume":"4979","author":"Pang Tianyu","year":"2019","unstructured":"Tianyu Pang, Kun Xu, Chao Du, Ning Chen, and Jun Zhu. 2019. Improving Adversarial Robustness via Promoting Ensemble Diversity. In Proceedings of the 36th International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 97). 4970-4979."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/eurosp.2016.36"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/tse.2022.3231621"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/sp40000.2020.00073"},{"key":"e_1_2_1_46_1","unstructured":"Postgresql Project. [n. d.]. Postgresql. https:\/\/www.postgresql.org\/. Accessed: 2025-05-30."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3579856.3582818"},{"key":"e_1_2_1_48_1","volume-title":"28th USENIX Security Symposium (USENIX Security 19)","author":"Quiring Erwin","year":"2019","unstructured":"Erwin Quiring, Alwin Maier, and Konrad Rieck. 2019. Misleading authorship attribution of source code using adversarial learning. In 28th USENIX Security Symposium (USENIX Security 19). 479-496. https:\/\/www.usenix.org\/conference\/ usenixsecurity19\/presentation\/quiring"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/p19-1103"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939778"},{"key":"e_1_2_1_51_1","volume-title":"30th USENIX security symposium (USENIX security 21). 1487-1504.","author":"Severi Giorgio","unstructured":"Giorgio Severi, Jim Meyer, Scott Coull, and Alina Oprea. 2021. {Explanation-Guided} backdoor poisoning attacks against malware classifiers. In 30th USENIX security symposium (USENIX security 21). 1487-1504."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3264820.3264821"},{"key":"e_1_2_1_53_1","volume-title":"33rd USENIX Security Symposium (USENIX Security 24)","author":"Shimmi Samiha","year":"2024","unstructured":"Samiha Shimmi, Ashiqur Rahman, Mohan Gadde, Hamed Okhravi, and Mona Rahimi. 2024. {VulSim}: Leveraging Similarity of {Multi-Dimensional} Neighbor Embeddings for Vulnerability Detection. In 33rd USENIX Security Symposium (USENIX Security 24). 1777-1794. https:\/\/www.usenix.org\/conference\/usenixsecurity24\/presentation\/shimmi"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3488932.3497768"},{"key":"e_1_2_1_55_1","unstructured":"Sqlite Project. [n. d.]."},{"key":"e_1_2_1_56_1","unstructured":"Sqlite. https:\/\/www.sqlite.org\/. Accessed: 2025-05-30."},{"key":"e_1_2_1_57_1","volume-title":"Proceedings of the 44th International Conference on Software Engineering (ICSE). ACM, 1-12","author":"Sun Zeyu","year":"2022","unstructured":"Zeyu Sun, Changjian Li, Junda Yao, Yin Wang, Qingshan Zheng, and Yang Liu. 2022. Understanding and Improving Graph Neural Networks for Vulnerability Detection. In Proceedings of the 44th International Conference on Software Engineering (ICSE). ACM, 1-12."},{"key":"e_1_2_1_58_1","unstructured":"Vector 35 Inc. [n. d.]. Binary Ninja. https:\/\/binary.ninja\/. Accessed: 2025-05-30."},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3721481"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11416-007-0074-9"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3650212.3652145"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3533767.3534367"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/FG52635.2021.9667076"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/icsme55016.2022.00019"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133956.3134018"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3428230"},{"key":"e_1_2_1_67_1","volume-title":"Gnnexplainer: Generating explanations for graph neural networks. Advances in neural information processing systems 32","author":"Ying Zhitao","year":"2019","unstructured":"Zhitao Ying, Dylan Bourgeois, Jiaxuan You, Marinka Zitnik, and Jure Leskovec. 2019. Gnnexplainer: Generating explanations for graph neural networks. Advances in neural information processing systems 32 (2019). https:\/\/ proceedings.neurips.cc\/paper_files\/paper\/2019\/hash\/d80b7040b773199015de6d3b4293c8ff-Abstract.html"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/tpami.2022.3204236"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1109\/tse.2023.3240118"},{"key":"e_1_2_1_70_1","volume-title":"32nd USENIX Security Symposium (USENIX Security 23)","author":"Zhang Zhuo","year":"2023","unstructured":"Zhuo Zhang, Guanhong Tao, Guangyu Shen, Shengwei An, Qiuling Xu, Yingqi Liu, Yapeng Ye, Yaoxuan Wu, and Xiangyu Zhang. 2023. {PELICAN}: Exploiting Backdoors of Naturally Trained Deep Learning Models In Binary Code Analysis. In 32nd USENIX Security Symposium (USENIX Security 23). 2365-2382. https:\/\/www.usenix.org\/conference\/ usenixsecurity23\/presentation\/zhang-zhuo-pelican"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3808188","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:59:43Z","timestamp":1782842383000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3808188"}},"issued":{"date-parts":[[2026,6,30]]},"references-count":70,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3808188"],"URL":"https:\/\/doi.org\/10.1145\/3808188","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2026,6,30]]}},{"indexed":{"date-parts":[[2026,2,24]],"date-time":"2026-02-24T17:45:36Z","timestamp":1771955136795,"version":"3.50.1"},"reference-count":35,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:00:00Z","timestamp":1720742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62372218"],"award-info":[{"award-number":["62372218"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Shenzhen Science and Technology Program","award":["SGDX20201103095408029"],"award-info":[{"award-number":["SGDX20201103095408029"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2024,7,12]]},"abstract":"<jats:p>\n                    As a powerful tool for developers, interactive debuggers help locate and fix errors in software. By using debugging information included in binaries, debuggers can retrieve necessary program states about the program. Unlike\n                    <jats:monospace>printf<\/jats:monospace>\n                    -style debugging, debuggers allow for more flexible inspection and modification of program execution states. However, debuggers may incorrectly retrieve and interpret program execution, causing confusion and hindering the debugging process.\n                  <\/jats:p>\n                  <jats:p>\n                    Despite the wide usage of interactive debuggers,\n                    <jats:italic toggle=\"yes\">a scalable and comprehensive measurement of their functionality correctness<\/jats:italic>\n                    does not exist yet. Existing works either fall short in scalability or focus more on the \u201ccompiler\u2014side\u201d defects instead of debugger bugs. To facilitate a better assessment of debugger correctness, we first propose and advocate a set of debugger testing criteria, covering both comprehensiveness (in terms of debug information covered) and scalability (in terms of testing overhead). Moreover, we design comparative experiments to show that fulfilling these criteria is not only theoretically appealing, but also brings major improvement to debugger testing. Furthermore, based on these criteria, we present DTD, a differential testing (DT) framework for detecting bugs in interactive debuggers. DTD compares the behaviors of two mainstream debuggers when processing an identical C executable \u2014 discrepancies indicate bugs in one of the two debuggers.\n                  <\/jats:p>\n                  <jats:p>DTD leverages a novel heuristic method to avoid the repetitive structures (e.g., loops) that exist in C programs, which facilitates DTD to achieve full debug information coverage efficiently. Moreover, we have also designed a Temporal Differential Filtering method to practically filter out the false positives caused by the uninitialized variables in common C programs. With these carefully designed techniques, DTD fulfills our proposed testing requirements and, therefore, achieves high scalability and testing comprehensiveness. For the first time, it offers large-scale testing for C debuggers to detect debugger behavior discrepancies when inspecting millions of program states. An empirical comparison shows that DTD finds 17\u00d7 more errortriggering cases and detects 5\u00d7 more bugs than the state-of-the-art debugger testing technique. We have used DTD to detect 13 bugs in the LLVM toolchain (Clang\/LLDB) and 5 bugs in the GNU toolchain (GCC\/GDB). One of our fixes has already landed in the latest LLDB development branch.<\/jats:p>","DOI":"10.1145\/3643779","type":"journal-article","created":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T10:22:09Z","timestamp":1720779729000},"page":"1172-1193","source":"Crossref","is-referenced-by-count":2,"title":["DTD: Comprehensive and Scalable Testing for Debuggers"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6864-5409","authenticated-orcid":false,"given":"Hongyi","family":"Lu","sequence":"first","affiliation":[{"name":"Southern University of Science and Technology, Shenzhen, China"},{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7872-1129","authenticated-orcid":false,"given":"Zhibo","family":"Liu","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0866-0308","authenticated-orcid":false,"given":"Shuai","family":"Wang","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3365-2526","authenticated-orcid":false,"given":"Fengwei","family":"Zhang","sequence":"additional","affiliation":[{"name":"Southern University of Science and Technology, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,7,12]]},"reference":[{"key":"e_1_3_1_2_1","doi-asserted-by":"publisher","unstructured":"Cristian Assaiante Daniele Cono D'Elia Giuseppe Antonio Di Luna and Leonardo Querzoni. 2023. Where Did My Variable Go? Poking Holes in Incomplete Debug Information. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems Volume 2 (Vancouver BC Canada) (ASPLOS 2023). Association for Computing Machinery New York NY USA 935-947. https:\/\/doi.org\/10.1145\/3575693.3575720 10.1145\/3575693.3575720","DOI":"10.1145\/3575693.3575720"},{"key":"e_1_3_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178372.3179521"},{"key":"e_1_3_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3363562"},{"key":"e_1_3_1_5_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2002.12543"},{"key":"e_1_3_1_6_1","doi-asserted-by":"publisher","unstructured":"Yuting Chen Ting Su and Zhendong Su. 2019. Deep differential testing of JVM implementations. In 2019 IEEE\/ACM 41st International Conference on Software Engineering (ICSE). IEEE 1257-1268. https:\/\/doi.org\/10.1109\/ICSE.2019.00127 10.1109\/ICSE.2019.00127","DOI":"10.1109\/ICSE.2019.00127"},{"key":"e_1_3_1_7_1","doi-asserted-by":"publisher","unstructured":"Yuting Chen Ting Su Chengnian Sun Zhendong Su and Jianjun Zhao. 2016. Coverage-directed differential testing of JVM implementations. In proceedings of the 37th ACM SIGPLAN Conference on Programming Language Design and Implementation. 85-99. https:\/\/doi.org\/10.1145\/2908080.2908095 10.1145\/2908080.2908095","DOI":"10.1145\/2908080.2908095"},{"key":"e_1_3_1_8_1","unstructured":"Weidong Cui Xinyang Ge Baris Kasikci Ben Niu Upamanyu Sharma Ruoyu Wang and Insu Yun. 2018. REPT: Reverse Debugging of Failures in Deployed Software. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association Carlsbad CA 17-32. https:\/\/www.usenix.org\/conference\/osdi18\/presentation\/weidong"},{"key":"e_1_3_1_9_1","unstructured":"Albert Danial 2021. cloc. https:\/\/github.com\/AlDanial\/cloc"},{"key":"e_1_3_1_10_1","doi-asserted-by":"publisher","unstructured":"Giuseppe Antonio Di Luna Davide Italiano Luca Massarelli Sebastian Osterlund Cristiano Giuffrida and Leonardo Querzoni. 2021. Who's Debugging the Debuggers? Exposing Debug Information Bugs in Optimized Binaries. In Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (Virtual USA) (ASPLOS 21). Association for Computing Machinery New York NY USA 1034-1045. https:\/\/doi.org\/10.1145\/3445814.3446695 10.1145\/3445814.3446695","DOI":"10.1145\/3445814.3446695"},{"key":"e_1_3_1_11_1","unstructured":"DTD. 2023. DTD: Supplementary Website. https:\/\/sites.google.com\/view\/dtd-supplementary\/"},{"key":"e_1_3_1_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/3357377.3357382"},{"key":"e_1_3_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2666356.2594334"},{"key":"e_1_3_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2858965.2814319"},{"key":"e_1_3_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3236024.3236037"},{"key":"e_1_3_1_16_1","doi-asserted-by":"publisher","unstructured":"Xavier Leroy. 2009. Formal Verification of a Realistic Compiler. Commun. ACM 52 7 (jul 2009) 107-115. https:\/\/doi.org\/10.1145\/1538788.1538814 10.1145\/1538788.1538814","DOI":"10.1145\/1538788.1538814"},{"key":"e_1_3_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3582016.3582053"},{"key":"e_1_3_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3385412.3386020"},{"key":"e_1_3_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598068"},{"key":"e_1_3_1_20_1","unstructured":"LLVM. 2023. The LLVM Project is a collection of modular and reusable compiler and toolchain technologies. https:\/\/github.com\/llvm\/llvm-project"},{"key":"e_1_3_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3489517.3530583"},{"issue":"1","key":"e_1_3_1_22_1","first-page":"100","article-title":"Differential testing for software. Digital Technical","volume":"10","author":"William M. McKeeman","year":"1998","unstructured":"William M. McKeeman 1998 Differential testing for software. Digital Technical Journal 10, 1 (1998), 100-107.","journal-title":"Journal"},{"key":"e_1_3_1_23_1","article-title":"Ninja: Towards Transparent Tracing and Debugging on ARM","author":"Ning Zhenyu","year":"2017","unstructured":"Zhenyu Ning and Fengwei Zhang. 2017. Ninja: Towards Transparent Tracing and Debugging on ARM. In Proceedings of The 26th USENIX Security Symposium (USENIX-Security 17). https:\/\/www.usenix.org\/conference\/usenixsecurity17\/technical-sessions\/presentation\/ning","journal-title":"Proceedings of The 26th USENIX Security Symposium (USENIX-Security 17)"},{"key":"e_1_3_1_24_1","unstructured":"Radare. 2023. Radare2: UNIX-like reverse engineering framework and command-line toolset. https:\/\/github.com\/radareorg\/radare2"},{"key":"e_1_3_1_25_1","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1145\/2254064.2254104","article-title":"Test-case reduction for C compiler bugs","author":"Regehr John","year":"2012","unstructured":"John Regehr, Yang Chen, Pascal Cuoq, Eric Eide, Chucky Ellison, and Xuejun Yang. 2012. Test-case reduction for C compiler bugs. In Proceedings of the 33rd ACM SIGPLAN conference on Programming Language Design and Implementation. 335-346.","journal-title":"Proceedings of the 33rd ACM SIGPLAN conference on Programming Language Design and Implementation"},{"key":"e_1_3_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409710"},{"key":"e_1_3_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2983990.2984038"},{"key":"e_1_3_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2931037.2931074"},{"key":"e_1_3_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503222.3507764"},{"key":"e_1_3_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3293882.3330567"},{"key":"e_1_3_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3575693.3575740"},{"key":"e_1_3_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00201"},{"key":"e_1_3_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3508035"},{"key":"e_1_3_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/1993316.1993532"},{"key":"e_1_3_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2015.11"},{"key":"e_1_3_1_36_1","doi-asserted-by":"crossref","unstructured":"Yiming Zhang Yuxin Hu Haonan Li Wenxuan Shi Zhenyu Ning Xiapu Luo and Fengwei Zhang. 2023. Alligator in Vest: A Practical Failure-Diagnosis Framework via Arm Hardware Features. 917-928. https:\/\/doi.org\/10.1145\/3597926.3598106","DOI":"10.1145\/3597926.3598106"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643779","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3643779","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T08:01:40Z","timestamp":1770192100000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643779"}},"issued":{"date-parts":[[2024,7,12]]},"references-count":35,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2024,7,12]]}},"alternative-id":["10.1145\/3643779"],"URL":"https:\/\/doi.org\/10.1145\/3643779","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2024,7,12]]}},{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T19:03:43Z","timestamp":1782846223485,"version":"3.54.5"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Hong Kong General Research Fund","award":["15224121"],"award-info":[{"award-number":["15224121"]}]},{"name":"Hong Kong General Research Fund","award":["15231223"],"award-info":[{"award-number":["15231223"]}]},{"name":"Hong Kong General Research Fund","award":["15204225"],"award-info":[{"award-number":["15204225"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62372490"],"award-info":[{"award-number":["62372490"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["W2412110"],"award-info":[{"award-number":["W2412110"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Smart contracts, predominantly written in Solidity and executed on blockchains like Ethereum, are immutable, making functional correctness paramount: once deployed, bugs and vulnerabilities become permanent. Despite rapid progress in transformer-based code LLMs, existing evaluations of Solidity code completion rely heavily on surface-form metrics (e.g., BLEU, CrystalBLEU) or hand-grading, which poorly correlate with functional correctness. Unlike Python, Solidity lacks large-scale and execution-based benchmarks, hindering systematic assessment and optimization of LLMs for smart contract development.<\/jats:p>\n                  <jats:p>To bridge this research gap, we introduce SolBench, a comprehensive benchmark and automated testing pipeline for Solidity, designed to emphasize functional correctness via differential fuzzing. SolBench contains 28,825 functions from 7,604 contracts collected from Etherscan (genesis to 2024), spanning 10 popular domains. We benchmark 14 diverse LLMs (open\/closed, 1.3B to 671B parameters, general\/code-specific, with\/without reasoning). The dominant failure mode is missing crucial details (e.g., type definitions, state variables) in intra-contract context. Providing full-contract context mitigates this and improves code completion accuracy.<\/jats:p>\n                  <jats:p>However, full-context inference can be prohibitively expensive in practice. Generating outputs with large context windows using state-of-the-art models often incurs significant costs, rendering naive context scaling economically impractical. Crucially, most of a contract is irrelevant to implementing a given function; only a small subset of details is needed. To exploit this, we propose Retrieval-Augmented Repair (RAR), which integrates retrieval into code repair: it uses the executor's error messages to extract only the most relevant snippets from the full contract. RAR sharply reduces input length for function completion, improving accuracy while significantly cutting computational cost. We further analyze retrieval and code repair strategies within RAR, showing substantial improvements in accuracy and efficiency. SolBench and our RAR framework enable principled evaluation and cost-effective improvement of Solidity code generation. Dataset and code are available at https:\/\/github.com\/ZaoyuChen\/SolBench.<\/jats:p>","DOI":"10.1145\/3797068","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"116-137","source":"Crossref","is-referenced-by-count":0,"title":["Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-0397-8952","authenticated-orcid":false,"given":"Zaoyu","family":"Chen","sequence":"first","affiliation":[{"name":"The Hong Kong Polytechnic University, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2350-7698","authenticated-orcid":false,"given":"Haoran","family":"Qin","sequence":"additional","affiliation":[{"name":"The Hong Kong Polytechnic University, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8600-8203","authenticated-orcid":false,"given":"Nuo","family":"Chen","sequence":"additional","affiliation":[{"name":"The Hong Kong Polytechnic University, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4468-5529","authenticated-orcid":false,"given":"Xiangyu","family":"Zhao","sequence":"additional","affiliation":[{"name":"The Hong Kong Polytechnic University, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5321-5740","authenticated-orcid":false,"given":"Lei","family":"Xue","sequence":"additional","affiliation":[{"name":"Sun Yat-sen University, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9082-3208","authenticated-orcid":false,"given":"Xiapu","family":"Luo","sequence":"additional","affiliation":[{"name":"The Hong Kong Polytechnic University, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3130-0554","authenticated-orcid":false,"given":"Xiao-Ming","family":"Wu","sequence":"additional","affiliation":[{"name":"The Hong Kong Polytechnic University, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Jacob Austin Augustus Odena Maxwell Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie Cai Michael Terry Quoc Le et al. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732 (2021)."},{"key":"e_1_2_1_2_1","volume-title":"A parallel corpus of python functions and documentation strings for automated code documentation and code generation. arXiv preprint arXiv:1707.02275","author":"Miceli Barone Antonio Valerio","year":"2017","unstructured":"Antonio Valerio Miceli Barone and Rico Sennrich. 2017. A parallel corpus of python functions and documentation strings for automated code documentation and code generation. arXiv preprint arXiv:1707.02275 (2017)."},{"key":"e_1_2_1_3_1","volume-title":"Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al.","author":"Chen Mark","year":"2021","unstructured":"Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)."},{"key":"e_1_2_1_4_1","unstructured":"Xinyun Chen Maxwell Lin Nathanael Sch\u00e4rli and Denny Zhou. 2023. Teaching Large Language Models to Self-Debug. arXiv:2304.05128 [cs.CL] https:\/\/arxiv.org\/abs\/2304.05128"},{"key":"e_1_2_1_5_1","unstructured":"Crytic. 2024. Crytic\/Diffusc. GitHub. Accessed: 2024-06-13 https:\/\/github.com\/crytic\/diffusc."},{"key":"e_1_2_1_6_1","volume-title":"Emmanuel Teye-Kofi Odonkor, and Paul Ammah","author":"Osae Dade Nii Osae","year":"2023","unstructured":"Nii Osae Osae Dade, Margaret Lartey-Quaye, Emmanuel Teye-Kofi Odonkor, and Paul Ammah. 2023. Optimizing Large Language Models to Expedite the Development of Smart Contracts. arXiv preprint arXiv:2310.05178 (2023)."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/BRAINS63024.2024.10732686"},{"key":"e_1_2_1_8_1","volume-title":"Ramesh Nallapati, Parminder Bhatia, Dan Roth, et al.","author":"Ding Yangruibo","year":"2024","unstructured":"Yangruibo Ding, Zijian Wang, Wasi Ahmad, Hantian Ding, Ming Tan, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, et al. 2024. Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion. Advances in Neural Information Processing Systems 36 (2024)."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3551349.3556903"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/wetseb.2019.00008"},{"key":"e_1_2_1_11_1","volume-title":"Codebert: A pre-trained model for programming and natural languages. arXiv preprint arXiv:2002.08155","author":"Feng Zhangyin","year":"2020","unstructured":"Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al. 2020. Codebert: A pre-trained model for programming and natural languages. arXiv preprint arXiv:2002.08155 (2020)."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3415298"},{"key":"e_1_2_1_13_1","volume-title":"Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 2","author":"Gao Yunfan","year":"2023","unstructured":"Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 2 (2023)."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3395363.3404366"},{"key":"e_1_2_1_15_1","volume-title":"Unixcoder: Unified cross-modal pre-training for code representation. arXiv preprint arXiv:2203.03850","author":"Guo Daya","year":"2022","unstructured":"Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. 2022. Unixcoder: Unified cross-modal pre-training for code representation. arXiv preprint arXiv:2203.03850 (2022)."},{"key":"e_1_2_1_16_1","unstructured":"Daya Guo Qihao Zhu Dejian Yang Zhenda Xie Kai Dong Wentao Zhang Guanting Chen Xiao Bi Y Wu YK Li et al. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming-The Rise of Code Intelligence. arXiv preprint arXiv:2401.14196 (2024)."},{"key":"e_1_2_1_17_1","unstructured":"Dan Hendrycks Steven Basart Saurav Kadavath Mantas Mazeika Akul Arora Ethan Guo Collin Burns Samir Puranik Horace He Dawn Song et al. 2021. Measuring coding challenge competence with apps. arXiv preprint arXiv:2105.09938 (2021)."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2024.3491578"},{"key":"e_1_2_1_19_1","volume-title":"Jia Li, Chenghao Mou, Carlos Mu\u00f1oz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries.","author":"Kocetkov Denis","year":"2022","unstructured":"Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Mu\u00f1oz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries. 2022. The Stack: 3 TB of permissively licensed source code. Preprint (2022)."},{"key":"e_1_2_1_20_1","volume-title":"International Conference on Machine Learning. PMLR","author":"Lai Yuhang","year":"2023","unstructured":"Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Wen-tau Yih, Daniel Fried, Sida Wang, and Tao Yu. 2023. DS-1000: A natural and reliable benchmark for data science code generation. In International Conference on Machine Learning. PMLR, 18319-18345."},{"key":"e_1_2_1_21_1","unstructured":"Patrick Lewis Ethan Perez Aleksandra Piktus Fabio Petroni Vladimir Karpukhin Naman Goyal Heinrich K\u00fcttler Mike Lewis Wen-tau Yih Tim Rockt\u00e4schel et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems 33 (2020) 9459-9474."},{"key":"e_1_2_1_22_1","unstructured":"Haoyang Li Jing Zhang Hanbing Liu Ju Fan Xiaokang Zhang Jun Zhu Renjie Wei Hongyan Pan Cuiping Li and Hong Chen. 2024. CodeS: Towards Building Open-source Language Models for Text-to-SQL. arXiv:2402.16347 [cs.CL] https:\/\/arxiv.org\/abs\/2402.16347"},{"key":"e_1_2_1_23_1","volume-title":"Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al.","author":"Li Raymond","year":"2023","unstructured":"Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. 2023. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161 (2023)."},{"key":"e_1_2_1_24_1","volume-title":"Repobench: Benchmarking repository-level code auto-completion systems. arXiv preprint arXiv:2306.03091","author":"Liu Tianyang","year":"2023","unstructured":"Tianyang Liu, Canwen Xu, and Julian McAuley. 2023. Repobench: Benchmarking repository-level code auto-completion systems. arXiv preprint arXiv:2306.03091 (2023)."},{"key":"e_1_2_1_25_1","volume-title":"Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark.","author":"Madaan Aman","year":"2023","unstructured":"Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023. Self-Refine: Iterative Refinement with Self-Feedback. arXiv:2303.17651 [cs.CL] https:\/\/arxiv.org\/abs\/2303.17651"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3674805.3686686"},{"key":"e_1_2_1_27_1","volume-title":"DISL: Fueling Research with A Large Dataset of Solidity Smart Contracts. arXiv preprint arXiv:2403.16861","author":"Morello Gabriele","year":"2024","unstructured":"Gabriele Morello, Mojtaba Eshghie, Sofia Bobadilla, and Martin Monperrus. 2024. DISL: Fueling Research with A Large Dataset of Solidity Smart Contracts. arXiv preprint arXiv:2403.16861 (2024)."},{"key":"e_1_2_1_28_1","volume-title":"DISL: Fueling Research with A Large Dataset of Solidity Smart Contracts. arXiv:2403.16861 [cs.SE] https:\/\/arxiv.org\/abs\/2403.16861","author":"Morello Gabriele","year":"2024","unstructured":"Gabriele Morello, Mojtaba Eshghie, Sofia Bobadilla, and Martin Monperrus. 2024. DISL: Fueling Research with A Large Dataset of Solidity Smart Contracts. arXiv:2403.16861 [cs.SE] https:\/\/arxiv.org\/abs\/2403.16861"},{"key":"e_1_2_1_29_1","volume-title":"Codegen: An open large language model for code with multi-turn program synthesis. arXiv preprint arXiv:2203.13474","author":"Nijkamp Erik","year":"2022","unstructured":"Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2022. Codegen: An open large language model for code with multi-turn program synthesis. arXiv preprint arXiv:2203.13474 (2022)."},{"key":"e_1_2_1_30_1","volume-title":"Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama.","author":"Olausson Theo X.","year":"2024","unstructured":"Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama. 2024. Is Self-Repair a Silver Bullet for Code Generation? arXiv:2306.09896 [cs.CL] https:\/\/arxiv.org\/abs\/2306.09896"},{"key":"e_1_2_1_31_1","doi-asserted-by":"crossref","unstructured":"Zhiyuan Peng Xin Yin Rui Qian Peiqin Lin Yongkang Liu Chenhao Ying and Yuan Luo. 2025. SolEval: Benchmarking Large Language Models for Repository-level Solidity Code Generation. (2025). arXiv:2502.18793 [cs.SE] https: \/\/arxiv.org\/abs\/2502.18793 unpublished.","DOI":"10.18653\/v1\/2025.emnlp-main.218"},{"key":"e_1_2_1_32_1","volume-title":"Yossi Adi, Jingyu Liu, Tal Remez, J\u00e9r\u00e9my Rapin, et al.","author":"Roziere Baptiste","year":"2023","unstructured":"Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, J\u00e9r\u00e9my Rapin, et al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950 (2023)."},{"key":"e_1_2_1_33_1","volume-title":"LLM4Fuzz: Guided Fuzzing of Smart Contracts with Large Language Models. arXiv preprint arXiv:2401.11108","author":"Shou Chaofan","year":"2024","unstructured":"Chaofan Shou, Jing Liu, Doudou Lu, and Koushik Sen. 2024. LLM4Fuzz: Guided Fuzzing of Smart Contracts with Large Language Models. arXiv preprint arXiv:2401.11108 (2024)."},{"key":"e_1_2_1_34_1","volume-title":"International Conference on Machine Learning. PMLR, 31693-31715","author":"Shrivastava Disha","year":"2023","unstructured":"Disha Shrivastava, Hugo Larochelle, and Daniel Tarlow. 2023. Repository-level prompt generation for large language models of code. In International Conference on Machine Learning. PMLR, 31693-31715."},{"key":"e_1_2_1_35_1","volume-title":"2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 683-693","author":"Storhaug Andr\u00e9","year":"2023","unstructured":"Andr\u00e9 Storhaug, Jingyue Li, and Tianyuan Hu. 2023. Efficient avoidance of vulnerabilities in auto-completed smart contract code using vulnerability-constrained decoding. In 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 683-693."},{"key":"e_1_2_1_36_1","unstructured":"Sean Welleck Ximing Lu Peter West Faeze Brahman Tianxiao Shen Daniel Khashabi and Yejin Choi. 2022. Generating Sequences by Learning to Self-Correct. arXiv:2211.00053 [cs.CL] https:\/\/arxiv.org\/abs\/2211.00053"},{"key":"e_1_2_1_37_1","volume-title":"Can We Verify Step by Step for Incorrect Answer Detection? arXiv preprint arXiv:2402.10528","author":"Xu Xin","year":"2024","unstructured":"Xin Xu, Shizhe Diao, Can Yang, and Yang Wang. 2024. Can We Verify Step by Step for Incorrect Answer Detection? arXiv preprint arXiv:2402.10528 (2024)."},{"key":"e_1_2_1_38_1","unstructured":"Lei Yu Shiqi Chen Hang Yuan Peng Wang Zhirong Huang Jingyuan Zhang Chenjie Shen Fengjun Zhang Li Yang and Jiajia Ma. 2024. Smart-LLaMA: Two-Stage Post-Training of Large Language Models for Smart Contract Vulnerability Detection and Explanation. arXiv:2411.06221 [cs.CR] https:\/\/arxiv.org\/abs\/2411.06221"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.151"},{"key":"e_1_2_1_40_1","doi-asserted-by":"crossref","unstructured":"Hainan Zhang Qinnan Zhang Ziwei Wang Hongwei Zheng Jin Dong Zhiming Zheng et al. 2025. CodeBC: A More Secure Large Language Model for Smart Contract Code Generation in Blockchain. arXiv preprint arXiv:2504.21043 (2025).","DOI":"10.1016\/j.neucom.2026.133741"},{"key":"e_1_2_1_41_1","volume-title":"Self-Edit: Fault-Aware Code Editor","author":"Zhang Kechi","unstructured":"Kechi Zhang, Zhuo Li, Jia Li, Ge Li, and Zhi Jin. 2023. Self-Edit: Fault-Aware Code Editor for Code Generation. arXiv:2305.04087 [cs.SE] https:\/\/arxiv.org\/abs\/2305.04087"},{"key":"e_1_2_1_42_1","volume-title":"Acfix: Guiding llms with mined common rbac practices for context-aware repair of access control vulnerabilities in smart contracts. arXiv preprint arXiv:2403.06838","author":"Zhang Lyuye","year":"2024","unstructured":"Lyuye Zhang, Kaixuan Li, Kairan Sun, Daoyuan Wu, Ye Liu, Haoye Tian, and Yang Liu. 2024. Acfix: Guiding llms with mined common rbac practices for context-aware repair of access control vulnerabilities in smart contracts. arXiv preprint arXiv:2403.06838 (2024)."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2024.107405"},{"key":"e_1_2_1_44_1","volume-title":"Verify-and-edit: A knowledge-enhanced chain-of-thought framework. arXiv preprint arXiv:2305.03268","author":"Zhao Ruochen","year":"2023","unstructured":"Ruochen Zhao, Xingxuan Li, Shafiq Joty, Chengwei Qin, and Lidong Bing. 2023. Verify-and-edit: A knowledge-enhanced chain-of-thought framework. arXiv preprint arXiv:2305.03268 (2023)."}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3797068","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:08:18Z","timestamp":1782842898000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3797068"}},"issued":{"date-parts":[[2026,6,30]]},"references-count":44,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3797068"],"URL":"https:\/\/doi.org\/10.1145\/3797068","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2026,6,30]]}},{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T15:00:12Z","timestamp":1784300412017,"version":"3.55.0"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:00:00Z","timestamp":1720742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2024,7,12]]},"abstract":"<jats:p>Continuous Integration (CI) is a common practice adopted by modern software organizations. It plays an especially important role for large corporations like Ubisoft, where thousands of build jobs are submitted daily. Indeed, the cadence of development progress is constrained by the pace at which CI services process build jobs. To provide faster CI feedback, recent work explores how build outcomes can be anticipated. Although early results show plenty of promise, the distinct characteristics of Project X\u2014a AAA video game project at Ubisoft\u2014present new challenges for build outcome prediction. In the Project X setting, changes that do not modify source code also incur build failures. We also observe that the code changes that have an impact that crosses the source-data boundary are more prone to build failures than code changes that do not impact data files. Since such changes are not fully characterized by the existing set of features for build outcome prediction, state-of-the-art models tend to underperform.<\/jats:p>\n                  <jats:p>To incorporate the data context, we propose RavenBuild\u2014a novel approach to build outcome prediction that leverages context-, relevance-, and dependency-aware features. In the Project X context, we observe that RavenBuild improves the F1-score of the failing class by 50%, the recall of the failing class by 105%, and the AUC by 11% with respect to the state-of-the-art BuildFast approach. To ease adoption in settings with heterogeneous project sets, we also provide a simplified alternative RavenBuild-CR, which excludes dependency-aware features. We observe across-the-board improvements when RavenBuild-CR is applied to 22 open-source projects and Project X. On the other hand, we find that a na\u00efve Parrot approach, which simply echoes the previous build outcome as its prediction, is surprisingly competitive with BuildFast and RavenBuild. Though Parrot fails to predict when the build outcome differs from their immediate predecessor, Parrot serves well as a tendency indicator of the sequences in build outcome datasets. Thus, we recommend that future studies also compare to the Parrot approach as a baseline when evaluating build outcome prediction models.<\/jats:p>","DOI":"10.1145\/3643771","type":"journal-article","created":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T10:22:09Z","timestamp":1720779729000},"page":"996-1018","source":"Crossref","is-referenced-by-count":9,"title":["RavenBuild: Context, Relevance, and Dependency Aware Build Outcome Prediction"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-6143-9862","authenticated-orcid":false,"given":"Gengyi","family":"Sun","sequence":"first","affiliation":[{"name":"University of Waterloo, Waterloo, Canada"},{"name":"Ubisoft, Toronto, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5989-1413","authenticated-orcid":false,"given":"Sarra","family":"Habchi","sequence":"additional","affiliation":[{"name":"Ubisoft, Montr\u00e9al, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0193-3975","authenticated-orcid":false,"given":"Shane","family":"McIntosh","sequence":"additional","affiliation":[{"name":"University of Waterloo, Waterloo, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,7,12]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2020.2967380"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2019.2897300"},{"key":"e_1_3_1_4_2","first-page":"42","article-title":"BUILDFAST: History-Aware Build Outcome Prediction for Fast Feedback and Reduced Cost in Continuous Integration","author":"Chen Bihuan","year":"2020","unstructured":"Bihuan Chen, Linlin Chen, Chen Zhang, and Xin Peng. 2020. BUILDFAST: History-Aware Build Outcome Prediction for Fast Feedback and Reduced Cost in Continuous Integration. In 2020 35th IEEE\/ACM International Conference on Automated Software Engineering (ASE). 42-53.","journal-title":"2020 35th IEEE\/ACM International Conference on Automated Software Engineering (ASE)"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_3_1_6_2","unstructured":"CircleCI. 2023. CircleCI. https:\/\/circleci.com\/ Accessed on Date September 26 2023."},{"key":"e_1_3_1_7_2","unstructured":"Paul M Duvall Steve Matyas and Andrew Glover. 2007. Continuous integration: improving software quality and reducing risk. Pearson Education."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/2635868.2635910"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510457.3513078"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2020.3048335"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","first-page":"1330","DOI":"10.1145\/3510003.3510211","article-title":"Lessons from Eight Years of Operational Data from a Continuous Integration Service: An Exploratory Case Study of CircleCI","author":"Gallaba Keheliya","year":"2022","unstructured":"Keheliya Gallaba, Maxime Lamothe, and Shane McIntosh. 2022. Lessons from Eight Years of Operational Data from a Continuous Integration Service: An Exploratory Case Study of CircleCI. In Proc. of the International Conference on Software Engineering (ICSE). 1330-1342.","journal-title":"Proc. of the International Conference on Software Engineering (ICSE)"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2018.2838131"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/2771783.2771784"},{"key":"e_1_3_1_14_2","unstructured":"Priscilla E Greenwood and Michael S Nikulin. 1996. A guide to chi-squared testing. Vol. 280. John Wiley & Sons."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2006.72"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ESEM.2017.23"},{"key":"e_1_3_1_17_2","doi-asserted-by":"crossref","unstructured":"Kim Herzig Michaela Greiler Jacek Czerwonka and Brendan Murphy. 2015. The Art of Testing Less without Sacrificing Quality. In Proceedings of the 37th International Conference on Software Engineering - Volume 1 (Florence Italy) (ICSE \u201815). IEEE Press 483-483.","DOI":"10.1109\/ICSE.2015.66"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3106237.3106270"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","unstructured":"Michael Hilton Timothy Tunnell Kai Huang Darko Marinov and Danny Dig. 2016. Usage Costs and Benefits of Continuous Integration in Open-Source Projects. (2016) 426-437. https:\/\/doi.org\/10.1145\/2970276.2970358 10.1145\/2970276.2970358","DOI":"10.1145\/2970276.2970358"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/2970276.2970358"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/1370750.1370773"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3468264.3473103"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377811.3380437"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3576038"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2022.111292"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2004.08.006"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/MS.2015.50"},{"issue":"12","key":"e_1_3_1_28_2","first-page":"2346","article-title":"Learning under concept drift: A review","volume":"31","author":"Lu Jie","year":"2018","unstructured":"Jie Lu, Anjin Liu, Fan Dong, Feng Gu, Joao Gama, and Guangquan Zhang. 2018. Learning under concept drift: A review. IEEE Transactions on Knowledge and Data Engineering 31, 12 (2018), 2346-2363.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/1985793.1985813"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-SEIP.2017.16"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/2568225.2568226"},{"key":"e_1_3_1_32_2","first-page":"153","article-title":"CLEVER: Combining Code Metrics with Clone Detection for Just-in-Time Fault Prevention and Resolution in Large Industrial Projects","author":"Nayrolles Mathieu","year":"2018","unstructured":"Mathieu Nayrolles and Abdelwahab Hamou-Lhadj. 2018. CLEVER: Combining Code Metrics with Clone Detection for Just-in-Time Fault Prevention and Resolution in Large Industrial Projects. In 2018 IEEE\/ACM 15th International Conference on Mining Software Repositories (MSR). 153-164.","journal-title":"2018 IEEE\/ACM 15th International Conference on Mining Software Repositories (MSR)"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSR.2017.26"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510122"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460319.3464840"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3468264.3473115"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10515-021-00319-5"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3639477.3639726"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/2786805.2786850"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/2393596.2393642"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE.2007.19"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/2702123.2702593"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643771","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3643771","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T08:03:36Z","timestamp":1770192216000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643771"}},"issued":{"date-parts":[[2024,7,12]]},"references-count":41,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2024,7,12]]}},"alternative-id":["10.1145\/3643771"],"URL":"https:\/\/doi.org\/10.1145\/3643771","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2024,7,12]]}},{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:57:34Z","timestamp":1782845854121,"version":"3.54.5"},"reference-count":85,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Software deployment is a critical software engineering practice, particularly for high performance computing (HPC) software.The deployment determines the software execution performance because the deployment maps a number of software components to multiple CPUs in a server. Inappropriate mapping decreases software parallelism and increases resource contention due to software co-location on the same CPUs. However, calculating the mapping to maximize the software performance is challenging, primarily due to the lack of a joint performance model that accounts for both software parallelism and co-location. Consequently, existing industry practice has to rely on experienced engineers to manually tune the mapping during deployment, resulting in substantial human resource waste of man-months and suboptimal software performance.<\/jats:p>\n                  <jats:p>This paper proposes a holistic approach to mapping multiple CPUs among multiple software components  \nto achieve better applicability and performance. We develop a performance model for predicting performance impact of different CPU mapping configurations, along with a search algorithm to identify the best mapping scheme. Our performance model jointly considers software parallelism and co-location, breaks the performance estimation into regularized execution and interference coefficient to improve accuracy, and integrates expert knowledge to reduce the model complexity. Our search algorithm employs nested iterative packing algorithm to explore all possible mapping schemes, thereby uncovering the optimal solution. Evaluation on a multi module HPC application shows 17% better performance than its default CPU mapping Our solution has been deployed in a commercial HPC cluster with more than 50K CPU cores, delivering 26.5% performance improvement and saving many man-months effort spent on performance tuning.<\/jats:p>","DOI":"10.1145\/3808183","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"4001-4024","source":"Crossref","is-referenced-by-count":0,"title":["Unleashing HPC Application Performance through Software Deployment: A Joint Model of Software Parallelism and Co-location"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2678-9225","authenticated-orcid":false,"given":"Yuxin","family":"Ren","sequence":"first","affiliation":[{"name":"Huawei Technologies, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7707-3468","authenticated-orcid":false,"given":"Li","family":"Zhou","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0139-9055","authenticated-orcid":false,"given":"Chumin","family":"Sun","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-7059-5234","authenticated-orcid":false,"given":"Rui","family":"Fan","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2553-1804","authenticated-orcid":false,"given":"Jie","family":"Sun","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-8246-4713","authenticated-orcid":false,"given":"Ning","family":"Jia","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-6909-9511","authenticated-orcid":false,"given":"Xinwei","family":"Hu","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","unstructured":"S. Balsamo A. Di Marco P. Inverardi and M. Simeoni. 2004. Model-based performance prediction in software development: a survey. IEEE Transactions on Software Engineering (2004). doi:10.1109\/TSE.2004.9 10.1109\/TSE.2004.9","DOI":"10.1109\/TSE.2004.9"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2741948.2741962"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ECRTS.2009.14"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1137\/080738970"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3392717.3392764"},{"key":"e_1_2_1_6_1","unstructured":"CESM. [n. d.]. Community Earth System Model. https:\/\/www.cesm.ucar.edu\/"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00036"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00068"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3243176.3243199"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2938172"},{"key":"e_1_2_1_11_1","unstructured":"Patent citation network. [n. d.]. https:\/\/snap.stanford.edu\/data\/cit-Patents.html."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1137\/0209062"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2540708.2540737"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1029\/2019ms001916"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2451116.2451157"},{"key":"e_1_2_1_17_1","unstructured":"CESM Input Data. [n. d.]. https:\/\/ftp.cgd.ucar.edu\/cesm\/inputdata\/."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517349.2522714"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2694344.2694359"},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the 6th Conference on Symposium on Operating Systems Design and Implementation (OSDI'04)","author":"Dean Jeffrey","year":"2004","unstructured":"Jeffrey Dean and Sanjay Ghemawat. 2004. MapReduce: Simplified Data Processing on Large Clusters. In Proceedings of the 6th Conference on Symposium on Operating Systems Design and Implementation (OSDI'04)."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS62706.2024.00034"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3337821.3337893"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2018.00050"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416620"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2012.11"},{"key":"e_1_2_1_26_1","unstructured":"Twitter follower network. [n. d.]. https:\/\/snap.stanford.edu\/data\/twitter-2010.html."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2013.6606597"},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the 2018 USENIX Conference on Usenix Annual Technical Conference (USENIX ATC'18).","author":"Funston Justin","year":"2018","unstructured":"Justin Funston, Maxime Lorrillere, Alexandra Fedorova, Baptiste Lepers, David Vengerov, Jean-Pierre Lozi, and Vivien Qu\u00e9ma. 2018. Placement of Virtual Containers on NUMA Systems: A Practical and Comprehensive Model. In Proceedings of the 2018 USENIX Conference on Usenix Annual Technical Conference (USENIX ATC'18)."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS.2019.00015"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3064176"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2013.6693089"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","unstructured":"Qianyu Guo Sen Chen Xiaofei Xie Lei Ma Qiang Hu Hongtao Liu Yang Liu Jianjun Zhao and Xiaohong Li. 2019. An Empirical Study Towards Characterizing Deep Learning Development and Deployment Across Different Frameworks and Platforms. In 2019 34th IEEE\/ACM International Conference on Automated Software Engineering (ASE'19). doi:10.1109\/ASE.2019.00080 10.1109\/ASE.2019.00080","DOI":"10.1109\/ASE.2019.00080"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2592798.2592807"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3731569.3764800"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1175\/BAMS-D-12-00121.1"},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the 2018 USENIX Conference on Usenix Annual Technical Conference (USENIX ATC'18).","author":"Iorgulescu C\u0103lin","year":"2018","unstructured":"C\u0103lin Iorgulescu, Reza Azimi, Youngjin Kwon, Sameh Elnikety, Manoj Syamala, Vivek Narasayya, Herodotos Herodotou, Paulo Tomita, Alex Chen, Jack Zhang, and Junhua Wang. 2018. PerfIso: Performance Isolation for Commercial Latency- Sensitive Services. In Proceedings of the 2018 USENIX Conference on Usenix Annual Technical Conference (USENIX ATC'18)."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-Companion58688"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-Companion.2019.00028"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2400682.2400704"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2807591.2807632"},{"key":"e_1_2_1_42_1","volume-title":"Thread and Memory Placement on NUMA Systems: Asymmetry Matters. In 2015 USENIX Annual Technical Conference (USENIX ATC'15)","author":"Lepers Baptiste","year":"2015","unstructured":"Baptiste Lepers, Vivien Quema, and Alexandra Fedorova. 2015. Thread and Memory Placement on NUMA Systems: Asymmetry Matters. In 2015 USENIX Annual Technical Conference (USENIX ATC'15)."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3307681.3325409"},{"key":"e_1_2_1_44_1","unstructured":"Lightweight Python library for in-memory matrix completion. [n. d.]. https:\/\/github.com\/tonyduan\/matrix-completion."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.75"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.84"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2815400.2815406"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/1294904.1294911"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/103727.103729"},{"key":"e_1_2_1_50_1","unstructured":"The Community Earth System Model. [n. d.]. https:\/\/github.com\/ESCOMP\/CESM."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2019.00065"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00176"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-4380-9_2"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/1048935.1050204"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824043"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.14778\/3015274.3015275"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2254064.2254082"},{"key":"e_1_2_1_58_1","volume-title":"Multi-tenant Edge Clouds. In 2020 USENIX Annual Technical Conference (USENIX ATC'20)","author":"Ren Yuxin","year":"2020","unstructured":"Yuxin Ren, Guyue Liu, Vlad Nitu, Wenyuan Shao, Riley Kennedy, Gabriel Parmer, Timothy Wood, and Alain Tchana. 2020. Fine-Grained Isolation for Scalable, Dynamic, Multi-tenant Edge Clouds. In 2020 USENIX Annual Technical Conference (USENIX ATC'20)."},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS.2018.00025"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3361525.3361537"},{"key":"e_1_2_1_61_1","unstructured":"California road network. [n. d.]. https:\/\/snap.stanford.edu\/data\/roadNet-CA.html."},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132747.3132771"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/3392717.3392765"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/2370816.2370833"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/2889160.2889223"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/2786805.2786845"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2012.6227196"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/3431379.3460635"},{"key":"e_1_2_1_69_1","unstructured":"Phoronix Test Suite. [n. d.]. https:\/\/www.phoronix-test-suite.com\/."},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/1508244.1508274"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356152"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS59052.2023.00030"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/1088149.1088190"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00100"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00072"},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2016.7446083"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/1504176.1504189"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00099"},{"key":"e_1_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/1882362.1882445"},{"key":"e_1_2_1_80_1","doi-asserted-by":"crossref","first-page":"977","DOI":"10.5194\/gmd-13-977-2020","article-title":"Beijing Climate Center Earth System Model version 1 (BCC-ESM1): model description and evaluation of aerosol simulations","volume":"13","author":"Wu T.","year":"2020","unstructured":"T. Wu, F. Zhang, J. Zhang, W. Jie, Y. Zhang, F. Wu, L. Li, J. Yan, X. Liu, X. Lu, H. Tan, L. Zhang, J. Wang, and A. Hu. 2020. Beijing Climate Center Earth System Model version 1 (BCC-ESM1): model description and evaluation of aerosol simulations. Geoscientific Model Development 13, 3 (2020), 977-1005. https:\/\/gmd.copernicus.org\/articles\/13\/977\/2020\/","journal-title":"Geoscientific Model Development"},{"key":"e_1_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-SEIP.2019.00010"},{"key":"e_1_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517349.2522737"},{"key":"e_1_2_1_83_1","unstructured":"Lei Zhang Takayuki Okamoto Shuuichirou Ishii Kouichi Hirai Shinji Sumimoto Balazs Gerofi Masamichi Takagi and Yutaka Ishikawa. [n. d.]. OS Enhancement in Supercomputer Fugaku https:\/\/www.fujitsu.com\/global\/about\/ resources\/publications\/technicalreview\/2020-03\/article06.html."},{"key":"e_1_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2015.15"},{"key":"e_1_2_1_85_1","volume-title":"Gemini: A Computation-Centric Distributed Graph Processing System. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI'16)","author":"Zhu Xiaowei","year":"2016","unstructured":"Xiaowei Zhu, Wenguang Chen, Weimin Zheng, and Xiaosong Ma. 2016. Gemini: A Computation-Centric Distributed Graph Processing System. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI'16)."}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3808183","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:57:03Z","timestamp":1782842223000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3808183"}},"issued":{"date-parts":[[2026,6,30]]},"references-count":85,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3808183"],"URL":"https:\/\/doi.org\/10.1145\/3808183","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2026,6,30]]}},{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T19:11:35Z","timestamp":1782846695158,"version":"3.54.5"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["Grant No. 62372114"],"award-info":[{"award-number":["Grant No. 62372114"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["Grant No. 62332005"],"award-info":[{"award-number":["Grant No. 62332005"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["Grant No. 62402342"],"award-info":[{"award-number":["Grant No. 62402342"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Static analysis tools can identify potential vulnerabilities, but they often fall short in providing concrete proofs-of-concept (PoCs) to validate their findings. Directed greybox fuzzing (DGF) has emerged as a promising solution by systematically guiding execution toward suspicious code locations and generating reproducible PoCs that can trigger the target vulnerabilities. However, DGF tools often overlook the influence of configurable options on reaching target locations. Besides, option-aware greybox fuzzing (GF) tools suffer from ineffective option extraction to target locations, and inefficient coordination between options and file fuzzing.<\/jats:p>\n                  <jats:p>To address these limitations, we present CoupleFuzz, a novel option-aware DGF tool that redefines PoC inputs as the combination of option input (OI) and file input (FI). CoupleFuzz adopts a two-phase workflow. The static analysis phase extracts option knowledge for guiding the fuzzing. The option-aware fuzzing phase employs taint analysis to dynamically prioritize effective option combinations and file bytes to target locations, and introduces a novel cross-guided fuzzing strategy that coordinates OI and FI fuzzing modules and enables each module to adapt to and benefit from its counterpart's advances, iteratively driving execution toward the target locations efficiently. Our evaluation has demonstrated that CoupleFuzz significantly outperforms the state-of-the-art DGF tools in generating PoCs for 22 real-world vulnerabilities, generating 15 (a 3.1\u00d7 improvement) more PoCs than the best traditional DGF baseline and achieves an average speedup of 5.6\u00d7 to reach target locations, with 6 0-day vulnerabilities confirmed by developers and 1 CVE identifier assigned.<\/jats:p>","DOI":"10.1145\/3808105","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"2189-2210","source":"Crossref","is-referenced-by-count":0,"title":["It Takes Two: Option-Aware Directed Greybox Fuzzing for Vulnerability PoC Generation"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-1112-0700","authenticated-orcid":false,"given":"Susheng","family":"Wu","sequence":"first","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-2794-9623","authenticated-orcid":false,"given":"Xin","family":"Hu","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-6101-8270","authenticated-orcid":false,"given":"Yiheng","family":"Cao","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-7819-9656","authenticated-orcid":false,"given":"Zhuotong","family":"Zhou","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-4722-3658","authenticated-orcid":false,"given":"Yiheng","family":"Huang","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9290-2068","authenticated-orcid":false,"given":"Yijian","family":"Wu","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6362-2582","authenticated-orcid":false,"given":"Bihuan","family":"Chen","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-8422-7677","authenticated-orcid":false,"given":"Zhijia","family":"Zhao","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3376-2581","authenticated-orcid":false,"given":"Xin","family":"Peng","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","first-page":"6977","volume-title":"34th USENIX Security Symposium (USENIX Security 25)","author":"Bao Andrew","year":"2025","unstructured":"Andrew Bao, Wenjia Zhao, Yanhao Wang, Yueqiang Cheng, Stephen McCamant, and Pen-Chung Yew. 2025. From Alarms to Real Bugs: Multi-target Multi-step Directed Greybox Fuzzing for Static Analysis Result Verification. In 34th USENIX Security Symposium (USENIX Security 25). 6977-6997."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133956.3134020"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3243734.3243849"},{"key":"e_1_2_1_4_1","unstructured":"CodeQL. 2025. CodeQL. Retrieved Sep 6 2025 from https:\/\/codeql.github.com\/"},{"key":"e_1_2_1_5_1","unstructured":"CoupleFuzz. 2025. CoupleFuzz. Retrieved Sep 6 2025 from https:\/\/github.com\/optionGo\/CoupleFuzz"},{"key":"e_1_2_1_6_1","volume-title":"A option-free PoC of libde265. Retrieved","year":"2025","unstructured":"cve. 2025. A option-free PoC of libde265. Retrieved Sep 6, 2025 from https:\/\/www.cve.org\/CVERecord?id=CVE-2022- 43249"},{"key":"e_1_2_1_7_1","unstructured":"Dataflowsanitizer. 2024. Dataflowsanitizer. Retrieved Sep 6 2025 from https:\/\/clang.llvm.org\/docs\/ DataFlowSanitizerDesign.html"},{"key":"e_1_2_1_8_1","first-page":"6199","volume-title":"34th USENIX Security Symposium (USENIX Security 25)","author":"Deng Peng","year":"2025","unstructured":"Peng Deng, Lei Zhang, Yuchuan Meng, Zhemin Yang, and Yuan Zhang. 2025. {ChainFuzz}: Exploiting Upstream Vulnerabilities in {Open-Source} Supply Chains. In 34th USENIX Security Symposium (USENIX Security 25). 6199-6218."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510197"},{"key":"e_1_2_1_10_1","volume-title":"14th USENIX workshop on offensive technologies (WOOT 20)","author":"Fioraldi Andrea","year":"2020","unstructured":"Andrea Fioraldi, Dominik Maier, Heiko Ei\u00dffeldt, and Marc Heuse. 2020. {AFL++}: Combining incremental steps of fuzzing research. In 14th USENIX workshop on offensive technologies (WOOT 20)."},{"key":"e_1_2_1_11_1","volume-title":"29th USENIX security symposium (USENIX Security 20). 2577-2594.","author":"Gan Shuitao","unstructured":"Shuitao Gan, Chao Zhang, Peng Chen, Bodong Zhao, Xiaojun Qin, Dong Wu, and Zuoning Chen. 2020. {GREYONE}: Data flow sensitive fuzzing. In 29th USENIX security symposium (USENIX Security 20). 2577-2594."},{"key":"e_1_2_1_12_1","volume-title":"ruleset of CodeQL. Retrieved","year":"2025","unstructured":"github. 2025. ruleset of CodeQL. Retrieved Sep 6, 2025 from https:\/\/codeql.github.com\/docs\/writing-codeql-queries\/ codeql-queries\/"},{"key":"e_1_2_1_13_1","volume-title":"2022 IEEE Symposium on Security and Privacy (SP). IEEE, 36-50","author":"Huang Heqing","year":"2022","unstructured":"Heqing Huang, Yiyuan Guo, Qingkai Shi, Peisen Yao, Rongxin Wu, and Charles Zhang. 2022. Beacon: Directed grey-box fuzzing with provable path pruning. In 2022 IEEE Symposium on Security and Privacy (SP). IEEE, 36-50."},{"key":"e_1_2_1_14_1","volume-title":"2024 IEEE Symposium on Security and Privacy (SP). IEEE","author":"Huang Heqing","year":"2024","unstructured":"Heqing Huang, Peisen Yao, Hung-Chun Chiu, Yiyuan Guo, and Charles Zhang. 2024. Titan: Efficient multi-target directed greybox fuzzing. In 2024 IEEE Symposium on Security and Privacy (SP). IEEE, 1849-1864."},{"key":"e_1_2_1_15_1","volume-title":"2024 IEEE Symposium on Security and Privacy (SP). IEEE","author":"Huang Heqing","year":"2024","unstructured":"Heqing Huang, Anshunkang Zhou, Mathias Payer, and Charles Zhang. 2024. Everything is good for something: Counterexample-guided directed fuzzing via likely invariant inference. In 2024 IEEE Symposium on Security and Privacy (SP). IEEE, 1956-1973."},{"key":"e_1_2_1_16_1","first-page":"4931","volume-title":"32nd USENIX Security Symposium (USENIX Security 23)","author":"Kim Tae Eun","year":"2023","unstructured":"Tae Eun Kim, Jaeseung Choi, Kihong Heo, and Sang Kil Cha. 2023. {DAFL}: Directed Grey-box Fuzzing guided by Data Dependency. In 32nd USENIX Security Symposium (USENIX Security 23). 4931-4948."},{"key":"e_1_2_1_17_1","first-page":"220","article-title":"POWER: Program Option-Aware Fuzzer for High Bug Detection Ability","author":"Lee Ahcheong","year":"2022","unstructured":"Ahcheong Lee, Irfan Ariq, Yunho Kim, and Moonzoo Kim. 2022. POWER: Program Option-Aware Fuzzer for High Bug Detection Ability.. In ICST. 220-231.","journal-title":"ICST."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3697014"},{"key":"e_1_2_1_19_1","first-page":"3559","volume-title":"30th USENIX Security Symposium (USENIX Security 21)","author":"Lee Gwangmu","year":"2021","unstructured":"Gwangmu Lee, Woochul Shim, and Byoungyoung Lee. 2021. Constraint-guided directed greybox fuzzing. In 30th USENIX Security Symposium (USENIX Security 21). 3559-3576."},{"key":"e_1_2_1_20_1","first-page":"489","volume-title":"34th USENIX Security Symposium (USENIX Security 25)","author":"Lekssays Ahmed","year":"2025","unstructured":"Ahmed Lekssays, Hamza Mouhcine, Khang Tran, Ting Yu, and Issa Khalil. 2025. {LLMxCPG}:{Context-Aware} Vulnerability Detection Through Code Property {Graph-Guided} Large Language Models. In 34th USENIX Security Symposium (USENIX Security 25). 489-507."},{"key":"e_1_2_1_21_1","first-page":"2441","volume-title":"33rd USENIX Security Symposium (USENIX Security 24)","author":"Li Penghui","year":"2024","unstructured":"Penghui Li, Wei Meng, and Chao Zhang. 2024. {SDFuzz}: Target States Driven Directed Fuzzing. In 33rd USENIX Security Symposium (USENIX Security 24). 2441-2457."},{"key":"e_1_2_1_22_1","volume-title":"SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis. In 2025 IEEE Symposium on Security and Privacy (SP). IEEE, 3014-3032","author":"Li Yansong","year":"2025","unstructured":"Yansong Li, Paula Branco, Alexander M Hoole, Manish Marwah, Hari Manassery Koduvely, Guy-Vincent Jourdan, and Stephan Jou. 2025. SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis. In 2025 IEEE Symposium on Security and Privacy (SP). IEEE, 3014-3032."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/TDSC.2024.3354789"},{"key":"e_1_2_1_24_1","volume-title":"A default option PoC of libde265. Retrieved","year":"2025","unstructured":"libde265. 2025. A default option PoC of libde265. Retrieved Sep 6, 2025 from https:\/\/github.com\/strukturag\/libde265\/ issues\/484"},{"key":"e_1_2_1_25_1","volume-title":"A buffer overflow Vulnerability in 'combineSeparateSamplesBytes' of libtiff. Retrieved","year":"2025","unstructured":"libtiff. 2025. A buffer overflow Vulnerability in 'combineSeparateSamplesBytes' of libtiff. Retrieved Sep 6, 2025 from https:\/\/gitlab.com\/libtiff\/libtiff\/-\/issues\/740"},{"key":"e_1_2_1_26_1","volume-title":"A file dominated PoC of libtiff. Retrieved","year":"2025","unstructured":"libtiff. 2025. A file dominated PoC of libtiff. Retrieved Sep 6, 2025 from https:\/\/gitlab.com\/libtiff\/libtiff\/-\/issues\/741"},{"key":"e_1_2_1_27_1","volume-title":"A hybrid PoC Generated by CoupleFuzz. Retrieved","year":"2025","unstructured":"libtiff. 2025. A hybrid PoC Generated by CoupleFuzz. Retrieved Sep 6, 2025 from https:\/\/gitlab.com\/libtiff\/libtiff\/- \/issues\/595"},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the 31st Network and Distributed System Security Symposium (NDSS).","author":"Lin Peihong","year":"2024","unstructured":"Peihong Lin, Pengfei Wang, Xu Zhou, Wei Xie, Gen Zhang, and Kai Lu. 2024. DeepGo: Predictive Directed Greybox Fuzzing. In Proceedings of the 31st Network and Distributed System Security Symposium (NDSS)."},{"key":"e_1_2_1_29_1","volume-title":"American Fuzzy Lop. Retrieved","author":"Lop American Fuzzy","year":"2025","unstructured":"American Fuzzy Lop. 2024. American Fuzzy Lop. Retrieved Sep 6, 2025 from https:\/\/github.com\/google\/AFL"},{"key":"e_1_2_1_30_1","volume-title":"2023 IEEE Symposium on Security and Privacy (SP). IEEE, 2693-2707","author":"Luo Changhua","year":"2023","unstructured":"Changhua Luo, Wei Meng, and Penghui Li. 2023. Selectfuzz: Efficient directed fuzzing with selective path exploration. In 2023 IEEE Symposium on Security and Privacy (SP). IEEE, 2693-2707."},{"key":"e_1_2_1_31_1","first-page":"2289","volume-title":"29th USENIX Security Symposium (USENIX Security 20)","author":"\u00d6sterlund Sebastian","year":"2020","unstructured":"Sebastian \u00d6sterlund, Kaveh Razavi, Herbert Bos, and Cristiano Giuffrida. 2020. {ParmeSan}: Sanitizer-guided greybox fuzzing. In 29th USENIX Security Symposium (USENIX Security 20). 2289-2306."},{"key":"e_1_2_1_32_1","first-page":"2475","volume-title":"33rd USENIX Security Symposium (USENIX Security 24)","author":"Rong Huanyao","year":"2024","unstructured":"Huanyao Rong, Wei You, Xiaofeng Wang, and Tianhao Mao. 2024. Toward Unbiased {Multiple-Target} Fuzzing with Path Diversity. In 33rd USENIX Security Symposium (USENIX Security 24). 2475-2492."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2892208.2892235"},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 1716-1726","author":"Szab\u00f3 Tam\u00e1s","year":"2023","unstructured":"Tam\u00e1s Szab\u00f3. 2023. Incrementalizing production codeql analyses. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 1716-1726."},{"key":"e_1_2_1_35_1","volume-title":"32nd USENIX Security Symposium (USENIX Security 23)","author":"Wang Dawei","year":"2023","unstructured":"Dawei Wang, Ying Li, Zhiyu Zhang, and Kai Chen. 2023. {CarpetFuzz}: Automatic program option constraint extraction from documentation for fuzzing. In 32nd USENIX Security Symposium (USENIX Security 23). 1919-1936."},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 735-749","author":"Wang Dawei","year":"2024","unstructured":"Dawei Wang, Geng Zhou, Li Chen, Dan Li, and Yukai Miao. 2024. Prophetfuzz: Fully automated prediction and fuzzing of high-risk option combinations with only documentation via large language model. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 735-749."},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 705-719","author":"Wang Kelin","year":"2024","unstructured":"Kelin Wang, Mengda Chen, Liang He, Purui Su, Yan Cai, Jiongyi Chen, Bin Zhang, Chao Feng, and Chaojing Tang. 2024. OSmart: Whitebox Program Option Fuzzing. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 705-719."},{"key":"e_1_2_1_38_1","volume-title":"Tofu: Target-oriented fuzzer. arXiv preprint arXiv:2004.14375","author":"Wang Zi","year":"2020","unstructured":"Zi Wang, Ben Liblit, and Thomas Reps. 2020. Tofu: Target-oriented fuzzer. arXiv preprint arXiv:2004.14375 (2020)."},{"key":"e_1_2_1_39_1","first-page":"2459","volume-title":"33rd USENIX Security Symposium (USENIX Security 24)","author":"Xiang Yi","year":"2024","unstructured":"Yi Xiang, Xuhong Zhang, Peiyu Liu, Shouling Ji, Hong Liang, Jiacheng Xu, and Wenhai Wang. 2024. Critical code guided directed greybox fuzzing for commits. In 33rd USENIX Security Symposium (USENIX Security 24). 2459-2474."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598102"},{"key":"e_1_2_1_41_1","volume-title":"2024 IEEE Symposium on Security and Privacy (SP). IEEE","author":"Zhang Yujian","year":"2024","unstructured":"Yujian Zhang, Yaokun Liu, Jinyu Xu, and Yanhao Wang. 2024. Predecessor-aware directed greybox fuzzing. In 2024 IEEE Symposium on Security and Privacy (SP). IEEE, 1884-1900."},{"key":"e_1_2_1_42_1","volume-title":"USA: International Fuzzing Workshop (FUZZING).","author":"Zhang Zenong","year":"2022","unstructured":"Zenong Zhang, George Klees, Eric Wang, Michael Hicks, and Shiyi Wei. 2022. Registered report: Fuzzing configurations of program options. In San Diego, CA, USA: International Fuzzing Workshop (FUZZING)."},{"key":"e_1_2_1_43_1","first-page":"1343","volume-title":"32nd USENIX Security Symposium (USENIX Security 23)","author":"Zheng Han","year":"2023","unstructured":"Han Zheng, Jiayuan Zhang, Yuhang Huang, Zezhong Ren, He Wang, Chunjie Cao, Yuqing Zhang, Flavio Toffalini, and Mathias Payer. 2023. {FISHFUZZ}: Catch deeper bugs by throwing larger nets. In 32nd USENIX Security Symposium (USENIX Security 23). 1343-1360."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3708522"},{"key":"e_1_2_1_45_1","volume-title":"29th USENIX security symposium (USENIX security 20). 2255-2269.","author":"Zong Peiyuan","unstructured":"Peiyuan Zong, Tao Lv, Dawei Wang, Zizhuang Deng, Ruigang Liang, and Kai Chen. 2020. {FuzzGuard}: Filtering out unreachable inputs in directed grey-box fuzzing through deep learning. In 29th USENIX security symposium (USENIX security 20). 2255-2269."}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3808105","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:35:24Z","timestamp":1782844524000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3808105"}},"issued":{"date-parts":[[2026,6,30]]},"references-count":45,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3808105"],"URL":"https:\/\/doi.org\/10.1145\/3808105","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2026,6,30]]}},{"indexed":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T15:06:24Z","timestamp":1786633584474,"version":"build-2736575974"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:00:00Z","timestamp":1720742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2024,7,12]]},"abstract":"<jats:p>\n                    Developers heavily rely on Application Programming Interfaces (APIs) from libraries to build their software. As software evolves, developers may need to replace the used libraries with alternate libraries, a process known as\n                    <jats:italic toggle=\"yes\">library migration<\/jats:italic>\n                    . Doing this manually can be tedious, time-consuming, and prone to errors. Automated migration techniques can help alleviate some of this burden. However, designing effective automated migration techniques requires understanding the types of code changes required to transform client code that used the old library to the new library. This paper contributes an empirical study that provides a holistic view of Python library migrations, both in terms of the code changes required in a migration and the typical development effort involved. We manually label 3,096 migration-related code changes in 335 Python library migrations from 311 client repositories spanning 141 library pairs from 35 domains. Based on our labeled data, we derive a taxonomy for describing migration-related code changes,\n                    <jats:sc>PyMigTax<\/jats:sc>\n                    . Leveraging\n                    <jats:sc>PyMigTax<\/jats:sc>\n                    and our labeled data, we investigate various characteristics of Python library migrations, such as the types of program elements and properties of API mappings, the combinations of types of migration-related code changes in a migration, and the typical development effort required for a migration. Our findings highlight various potential shortcomings of current library migration tools. For example, we find that 40% of library pairs have API mappings that involve non-function program elements, while most library migration techniques typically assume that function calls from the source library will map into (one or more) function calls from the target library. As an approximation for the development effort involved, we find that, on average, a developer needs to learn about 4 APIs and 2 API mappings to perform a migration, and change 8 lines of code. However, we also found cases of migrations that involve up to 43 unique APIs, 22 API mappings, and 758 lines of code, making them harder to manually implement. Overall, our contributions provide the necessary knowledge and foundations for developing automated Python library migration techniques. We make all data and scripts related to this study publicly available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/doi.org\/10.6084\/m9.figshare.24216858.v2\">https:\/\/doi.org\/10.6084\/m9.figshare.24216858.v2<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3643731","type":"journal-article","created":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T10:22:09Z","timestamp":1720779729000},"page":"92-114","source":"Crossref","is-referenced-by-count":11,"title":["Characterizing Python Library Migrations"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6822-2270","authenticated-orcid":false,"given":"Mohayeminul","family":"Islam","sequence":"first","affiliation":[{"name":"University of Alberta, Edmonton, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2771-353X","authenticated-orcid":false,"given":"Ajay Kumar","family":"Jha","sequence":"additional","affiliation":[{"name":"North Dakota State University, Fargo, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6660-8890","authenticated-orcid":false,"given":"Ildar","family":"Akhmetov","sequence":"additional","affiliation":[{"name":"Northeastern University, Vancouver, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0091-6030","authenticated-orcid":false,"given":"Sarah","family":"Nadi","sequence":"additional","affiliation":[{"name":"University of Alberta, Edmonton, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,7,12]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3106237.3106267"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1002\/9780470606834"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377811.3380405"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.1983.235271"},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","unstructured":"Hussein Alrubaye and Mohamed Wiem Mkaouer. 2018. Automating the detection of third-party Java library migration at the function level.. In CASCON. 60\u201371.","DOI":"10.1109\/ICSME.2019.00072"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","unstructured":"Hussein Alrubaye Mohamed Wiem Mkaouer Igor Khokhlov Leon Reznik Ali Ouni and Jason Mcgoff. 2020. Learning to recommend third-party library migration opportunities at the API level. Applied Soft Computing 90 (2020) 106140. https:\/\/doi.org\/10.1016\/j.asoc.2020.106140 10.1016\/j.asoc.2020.106140","DOI":"10.1016\/j.asoc.2020.106140"},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","first-page":"414","DOI":"10.1109\/ICSME.2019.00072","volume-title":"2019 IEEE international conference on software maintenance and evolution (ICSME)","author":"Alrubaye Hussein","year":"2019","unstructured":"Hussein Alrubaye, Mohamed Wiem Mkaouer, and Ali Ouni. 2019. Migrationminer: An automated detection tool of third-party java library migration at the method level. In 2019 IEEE international conference on software maintenance and evolution (ICSME). IEEE, 414\u2013417."},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","unstructured":"Hussein Alrubaye Mohamed Wiem Mkaouer and Ali Ouni. 2019. On the use of information retrieval to automate the detection of third-party java library migration at the method level. In 2019 IEEE\/ACM 27th International Conference on Program Comprehension (ICPC). IEEE 347\u2013357.","DOI":"10.1109\/ICPC.2019.00053"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/1103845.1094832"},{"issue":"4","key":"e_1_3_1_11_2","doi-asserted-by":"crossref","first-page":"94","DOI":"10.1007\/s10664-021-10072-8","article-title":"Sampling in software engineering research: A critical review and guidelines.","volume":"27","author":"Baltes Sebastian","year":"2022","unstructured":"Sebastian Baltes and Paul Ralph. 2022. Sampling in software engineering research: A critical review and guidelines. Empirical Software Engineering 27, 4 (2022), 94.","journal-title":"Empirical Software Engineering"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","first-page":"435","DOI":"10.1109\/CSMR.2012.55","volume-title":"2012 16th European Conference on Software Maintenance and Reengineering","author":"Bauer Veronika","year":"2012","unstructured":"Veronika Bauer and Lars Heinemann. 2012. Understanding API usage to support informed decision making in software maintenance. In 2012 16th European Conference on Software Maintenance and Reengineering. IEEE, 435\u2013440."},{"issue":"6","key":"e_1_3_1_13_2","doi-asserted-by":"crossref","first-page":"349","DOI":"10.1002\/smr.412","article-title":"Understanding software maintenance and evolution by analyzing individual changes: a literature review.","volume":"21","author":"Benestad Hans Christian","year":"2009","unstructured":"Hans Christian Benestad, Bente Anda, and Erik Arisholm. 2009. Understanding software maintenance and evolution by analyzing individual changes: a literature review. Journal of Software Maintenance and Evolution: Research and Practice 21, 6 (2009), 349\u2013378.","journal-title":"Journal of Software Maintenance and Evolution: Research and Practice"},{"key":"e_1_3_1_14_2","unstructured":"B Boehm and D Reifer. 2000. Software Cost Estimation with COCOMO II. Prentice Hall. Upper Saddle River NJ (2000)."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","unstructured":"Aline Brito Laerte Xavier Andre Hora and Marco Tulio Valente. 2018. APIDiff: Detecting API breaking changes. In 2018 IEEE 25th International Conference on Software Analysis Evolution and Reengineering (SANER). 507\u2013511. https:\/\/doi.org\/10.1109\/SANER.2018.8330249 10.1109\/SANER.2018.8330249","DOI":"10.1109\/SANER.2018.8330249"},{"issue":"3","key":"e_1_3_1_16_2","doi-asserted-by":"crossref","first-page":"432","DOI":"10.1109\/TSE.2019.2896123","article-title":"Mining likely analogical apis across third-party libraries via large-scale unsupervised api semantics embedding.","volume":"47","author":"Chen Chunyang","year":"2019","unstructured":"Chunyang Chen, Zhenchang Xing, Yang Liu, and Kent Ong Long Xiong. 2019. Mining likely analogical apis across third-party libraries via large-scale unsupervised api semantics embedding. IEEE Transactions on Software Engineering 47, 3 (2019), 432\u2013447.","journal-title":"IEEE Transactions on Software Engineering"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2019.2896123"},{"issue":"1","key":"e_1_3_1_18_2","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1177\/001316446002000104","article-title":"A coefficient of agreement for nominal scales.","volume":"20","author":"Cohen Jacob","year":"1960","unstructured":"Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and psychological measurement 20, 1 (1960), 37\u201346.","journal-title":"Educational and psychological measurement"},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1109\/TSE.1979.234165","article-title":"Measuring the psychological complexity of software maintenance tasks with the Halstead and McCabe metrics.","volume":"2","author":"Curtis Bill","year":"1979","unstructured":"Bill Curtis, Sylvia B. Sheppard, Phil Milliman, MA Borst, and Tom Love. 1979. Measuring the psychological complexity of software maintenance tasks with the Halstead and McCabe metrics. IEEE Transactions on software engineering 2 (1979), 96\u2013104.","journal-title":"IEEE Transactions on software engineering"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSR52588.2021.00074"},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Erik Derr Sven Bugiel Sascha Fahl Yasemin Acar and Michael Backes. 2017. Keep me updated: An empirical study of third-party library updatability on android. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security. 2187\u20132200.","DOI":"10.1145\/3133956.3134059"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.5555\/2700539"},{"key":"e_1_3_1_23_2","unstructured":"Python Software Foundation. [n. d.]. Python Package Index - PyPI. Retrieved March 31 2022 from https:\/\/pypi.org"},{"issue":"12","key":"e_1_3_1_24_2","article-title":"The state-of-the-art in software development effort estimation.","volume":"30","author":"Gautam Swarnima Singh","year":"2018","unstructured":"Swarnima Singh Gautam and Vrijendra Singh. 2018. The state-of-the-art in software development effort estimation. Journal of Software: Evolution and Process 30, 12 (2018), e1983.","journal-title":"Journal of Software: Evolution and Process"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","unstructured":"Georgios Gousios and Diomidis Spinellis. 2012. GHTorrent: Github\u2019s data from a firehose. In 2012 9th IEEE Working Conference on Mining Software Repositories (MSR). 12\u201321. https:\/\/doi.org\/10.1109\/MSR.2012.6224294 10.1109\/MSR.2012.6224294","DOI":"10.1109\/MSR.2012.6224294"},{"key":"e_1_3_1_26_2","first-page":"627","volume-title":"2023 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER)","author":"Gu Haiqiao","year":"2023","unstructured":"Haiqiao Gu, Hao He, and Minghui Zhou. 2023. Self-Admitted Library Migrations in Java, JavaScript, and Python Packaging Ecosystems: A Comparative Study. In 2023 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE, 627\u2013638."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.5555\/540137"},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","first-page":"1109","DOI":"10.1109\/ICOEI.2017.8300883","volume-title":"2017 international conference on trends in electronics and informatics (ICEI)","author":"Hariprasad T","year":"2017","unstructured":"T Hariprasad, G Vidhyagaran, K Seenu, and Chandrasegar Thirumalai. 2017. Software complexity analysis using halstead metrics. In 2017 international conference on trends in electronics and informatics (ICEI). IEEE, 1109\u20131113."},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","unstructured":"Hao He Runzhi He Haiqiao Gu and Minghui Zhou. 2021. A large-scale empirical study on Java library migrations: prevalence trends and rationales. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 478\u2013490.","DOI":"10.1145\/3468264.3468571"},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1109\/SANER50967.2021.00016","volume-title":"2021 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER)","author":"He Hao","year":"2021","unstructured":"Hao He, Yulin Xu, Yixiao Ma, Yifei Xu, Guangtai Liang, and Minghui Zhou. 2021. A multi-metric ranking approach for library migration recommendations. In 2021 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE, 72\u201383."},{"key":"e_1_3_1_31_2","volume-title":"Practical software project estimation: a toolkit for estimating software development effort & duration","author":"Hill Peter R","year":"2011","unstructured":"Peter R Hill. 2011. Practical software project estimation: a toolkit for estimating software development effort & duration. McGraw-Hill Education."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSR59073.2023.00075"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","unstructured":"Mohayeminul Islam Sarah Nadi Ildar Akhmetov and Ajay Kumar Jha. 2024. Characterizing Python Library Migrations - artifacts. (1 2024). https:\/\/doi.org\/10.6084\/m9.figshare.24216858.v2 10.6084\/m9.figshare.24216858.v2","DOI":"10.6084\/m9.figshare.24216858.v2"},{"issue":"2","key":"e_1_3_1_34_2","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1109\/MS.2014.49","article-title":"What we do and don\u2019t know about software development effort estimation.","volume":"31","author":"J\u00f8rgensen Magne","year":"2014","unstructured":"Magne J\u00f8rgensen. 2014. What we do and don\u2019t know about software development effort estimation. IEEE software 31, 2 (2014), 37\u201340.","journal-title":"IEEE software"},{"key":"e_1_3_1_35_2","first-page":"154","volume-title":"2016 IEEE\/ACM 13th Working Conference on Mining Software Repositories (MSR)","author":"Kabinna Suhas","year":"2016","unstructured":"Suhas Kabinna, Cor-Paul Bezemer, Weiyi Shang, and Ahmed E Hassan. 2016. Logging library migrations: A case study for the apache software foundation projects. In 2016 IEEE\/ACM 13th Working Conference on Mining Software Repositories (MSR). IEEE, 154\u2013164."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","unstructured":"Jeremy Katz. 2020. Libraries.io Open Source Repository and Dependency Metadata. https:\/\/doi.org\/10.5281\/ZENODO.808272 10.5281\/ZENODO.808272","DOI":"10.5281\/ZENODO.808272"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","unstructured":"Ameya Ketkar Oleg Smirnov Nikolaos Tsantalis Danny Dig and Timofey Bryksin. 2022. Inferring and applying type changes. In Proceedings of the 44th International Conference on Software Engineering. 1206\u20131218. https:\/\/doi.org\/10.1145\/3510003.3510115 10.1145\/3510003.3510115","DOI":"10.1145\/3510003.3510115"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/571681.571686"},{"key":"e_1_3_1_39_2","first-page":"221","volume-title":"Content analysis: An introduction to its methodology","author":"Krippendorff Klaus","year":"2013","unstructured":"Klaus Krippendorff. 2013. Content analysis: An introduction to its methodology (3rd ed.). Sage publications, Thousand Oaks, California. 221\u2013250 pages."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-017-9521-5"},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","unstructured":"J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data. biometrics (1977) 159\u2013174.","DOI":"10.2307\/2529310"},{"key":"e_1_3_1_42_2","doi-asserted-by":"crossref","unstructured":"Enrique Larios Vargas Maur\u00edcio Aniche Christoph Treude Magiel Bruntink and Georgios Gousios. 2020. Selecting third-party libraries: The practitioners\u2019 perspective. In Proceedings of the 28th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering. 245\u2013256.","DOI":"10.1145\/3368089.3409711"},{"issue":"7","key":"e_1_3_1_43_2","article-title":"A systematic review of studies on use case points and expert-based estimation of software development effort.","volume":"32","author":"Mahmood Yasir","year":"2020","unstructured":"Yasir Mahmood, Nazri Kama, and Azri Azmi. 2020. A systematic review of studies on use case points and expert-based estimation of software development effort. Journal of Software: Evolution and Process 32, 7 (2020), e2245.","journal-title":"Journal of Software: Evolution and Process"},{"key":"e_1_3_1_44_2","doi-asserted-by":"crossref","first-page":"308","DOI":"10.1109\/TSE.1976.233837","article-title":"A complexity measure.","volume":"4","author":"McCabe Thomas J","year":"1976","unstructured":"Thomas J McCabe. 1976. A complexity measure. IEEE Transactions on software Engineering 4 (1976), 308\u2013320.","journal-title":"IEEE Transactions on software Engineering"},{"issue":"2","key":"e_1_3_1_45_2","doi-asserted-by":"crossref","first-page":"78","DOI":"10.1109\/MS.2013.142","article-title":"A large-scale empirical study on software reuse in mobile apps.","volume":"31","author":"Mojica Israel J","year":"2013","unstructured":"Israel J Mojica, Bram Adams, Meiyappan Nagappan, Steffen Dienst, Thorsten Berger, and Ahmed E Hassan. 2013. A large-scale empirical study on software reuse in mobile apps. IEEE software 31, 2 (2013), 78\u201386.","journal-title":"IEEE software"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/2889160.2892661"},{"key":"e_1_3_1_47_2","first-page":"112","volume-title":"2021 IEEE\/ACM 43rd International Conference on Software Engineering (ICSE)","author":"Ni Ansong","year":"2021","unstructured":"Ansong Ni, Daniel Ramos, Aidan ZH Yang, In\u00eas Lynce, Vasco Manquinho, Ruben Martins, and Claire Le Goues. 2021. Soar: a synthesis approach for data science api refactoring. In 2021 IEEE\/ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 112\u2013124."},{"key":"e_1_3_1_48_2","doi-asserted-by":"crossref","first-page":"74","DOI":"10.1109\/ICSE43902.2021.00020","volume-title":"2021 IEEE\/ACM 43rd International Conference on Software Engineering (ICSE)","author":"Nielsen Benjamin Barslev","year":"2021","unstructured":"Benjamin Barslev Nielsen Martin Toldam Torp and Anders M\u00f8ller. 2021. Semantic patches for adaptation of javascript programs to evolving libraries. In 2021 IEEE\/ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 74\u201385."},{"key":"e_1_3_1_49_2","unstructured":"R OpenAI. 2023. GPT-4 technical report. arXiv (2023) 2303\u201308774."},{"key":"e_1_3_1_50_2","unstructured":"Qualtrics. [n. d.]. How to use stratified random sampling in 2023. https:\/\/www.qualtrics.com\/experience-management\/research\/stratified-random-sampling\/."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-019-09713-w"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/32.799955"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.protcy.2012.05.116"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/WCRE.2012.38"},{"key":"e_1_3_1_55_2","doi-asserted-by":"crossref","first-page":"192","DOI":"10.1109\/WCRE.2013.6671294","volume-title":"2013 20th Working Conference on Reverse Engineering (WCRE)","author":"Teyton C\u00e9dric","year":"2013","unstructured":"C\u00e9dric Teyton, Jean-R\u00e9my Falleri, and Xavier Blanc. 2013. Automatic discovery of function mappings between similar libraries. In 2013 20th Working Conference on Reverse Engineering (WCRE). IEEE, 192\u2013201. Proc. ACM Softw. Eng., Vol. 1, No. FSE, Article 5. Publication date: July 2024."},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1002\/smr.1660"},{"key":"e_1_3_1_57_2","doi-asserted-by":"crossref","unstructured":"Jiawei Wang Li Li Kui Liu and Haipeng Cai. 2020. Exploring how deprecated python library apis are (not) handled. In Proceedings of the 28th acm joint meeting on european software engineering conference and symposium on the foundations of software engineering. 233\u2013244.","DOI":"10.1145\/3368089.3409735"},{"key":"e_1_3_1_58_2","doi-asserted-by":"crossref","first-page":"35","DOI":"10.1109\/ICSME46990.2020.00014","volume-title":"2020 IEEE International Conference on Software Maintenance and Evolution (ICSME)","author":"Wang Ying","year":"2020","unstructured":"Ying Wang, Bihuan Chen, Kaifeng Huang, Bowen Shi, Congying Xu, Xin Peng, Yijian Wu, and Yang Liu. 2020. An empirical study of usages, updates and risks of third-party libraries in java projects. In 2020 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 35\u201345."},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2007.70747"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-019-09771-0"},{"key":"e_1_3_1_61_2","doi-asserted-by":"crossref","unstructured":"Zejun Zhang Minxue Pan Tian Zhang Xinyu Zhou and Xuandong Li. 2020. Deep-diving into documentation to develop improved java-to-swift api mapping. In Proceedings of the 28th International Conference on Program Comprehension. 106\u2013116.","DOI":"10.1145\/3387904.3389282"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643731","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3643731","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T07:52:46Z","timestamp":1770191566000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643731"}},"issued":{"date-parts":[[2024,7,12]]},"references-count":60,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2024,7,12]]}},"alternative-id":["10.1145\/3643731"],"URL":"https:\/\/doi.org\/10.1145\/3643731","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2024,7,12]]}},{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T01:47:23Z","timestamp":1787017643232,"version":"build-2736575974"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:00:00Z","timestamp":1720742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2024,7,12]]},"abstract":"<jats:p>\n                    Deep learning (DL) is a critical tool for real-world applications, and comprehensive testing of DL models is vital to ensure their quality before deployment. However, recent studies have shown that even subtle deviations in DL operators can result in catastrophic consequences, underscoring the importance of rigorous testing of these components. Unlike testing other DL system components, operator analysis poses unique challenges due to complex inputs and uncertain outputs. The existing DL operator testing approach has limitations in terms of testing efficiency and error localization. In this paper, we propose\n                    <jats:italic toggle=\"yes\">Meta<\/jats:italic>\n                    , a novel operator testing framework based on metamorphic testing that automatically tests and assists bug location based on metamorphic relations (MRs). Meta distinguishes itself in three key ways: (1) it considers both parameters and input tensors to detect operator errors, enabling it to identify both implementation and precision errors; (2) it uses MRs to guide the generation of more effective inputs (i.e., tensors and parameters) in less time; (3) it assists the precision error localization by tracing the error to the input level of the operator based on MR violations. We designed 18 MRs for testing 10 widely used DL operators. To assess the effectiveness of Meta, we conducted experiments on 13 released versions of 5 popular DL libraries. Our results revealed that Meta successfully detected 41 errors, including 14 new ones that were reported to the respective platforms and 8 of them are confirmed\/fixed. Additionally, Meta demonstrated high efficiency, outperforming the baseline by detecting\n                    <jats:inline-formula>\n                      <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"inline\">\n                        <mml:mrow>\n                          <mml:mo>\u223c<\/mml:mo>\n                          <mml:mn>2<\/mml:mn>\n                        <\/mml:mrow>\n                      <\/mml:math>\n                    <\/jats:inline-formula>\n                    times more errors of the baseline. Meta is open-sourced and available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/TDY-raedae\/Medi-Test\">https:\/\/github.com\/TDY-raedae\/Medi-Test<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3660796","type":"journal-article","created":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T10:22:09Z","timestamp":1720779729000},"page":"2005-2027","source":"Crossref","is-referenced-by-count":10,"title":["A Miss Is as Good as A Mile: Metamorphic Testing for Deep Learning Operators"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7153-2755","authenticated-orcid":false,"given":"Jinyin","family":"Chen","sequence":"first","affiliation":[{"name":"Zhejiang University of Technology, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-6960-0047","authenticated-orcid":false,"given":"Chengyu","family":"Jia","sequence":"additional","affiliation":[{"name":"Zhejiang University of Technology, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-8058-6383","authenticated-orcid":false,"given":"Yunjie","family":"Yan","sequence":"additional","affiliation":[{"name":"Zhejiang University of Technology, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-5163-8765","authenticated-orcid":false,"given":"Jie","family":"Ge","sequence":"additional","affiliation":[{"name":"Zhejiang University of Technology, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8997-5343","authenticated-orcid":false,"given":"Haibin","family":"Zheng","sequence":"additional","affiliation":[{"name":"Zhejiang University of Technology, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5781-5185","authenticated-orcid":false,"given":"Yao","family":"Cheng","sequence":"additional","affiliation":[{"name":"T\u00dcV S\u00dcD Asia Pacific, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,7,12]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/IEEESTD.2020.9091348"},{"key":"e_1_3_1_3_2","first-page":"265","volume-title":"OSDI","author":"Abadi Martin","year":"2016","unstructured":"Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: A System for Large-Scale Machine Learning. In OSDI. USENIX Association, 265\u2013283."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2019.00042"},{"key":"e_1_3_1_5_2","unstructured":"Tianqi Chen Mu Li Yutian Li Min Lin Naiyan Wang Minjie Wang Tianjun Xiao Bing Xu Chiyuan Zhang and Zheng Zhang. 2015. MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. In NeurIPS. 1\u20136."},{"key":"e_1_3_1_6_2","first-page":"430","volume-title":"ASE","author":"Chen Zhuangbin","year":"2021","unstructured":"Zhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang, Xuemin Wen, Xiao Ling, Yongqiang Yang, and Michael R. Lyu. 2021. Graph-based Incident Aggregation for Large-Scale Online Service Systems. In ASE. IEEE, 430\u2013442."},{"key":"e_1_3_1_7_2","unstructured":"Francois Chollet. 2015. Keras: Deep learning library for theano and tensorflow. https:\/\/github.com\/keras-team\/keras."},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","unstructured":"Yinlin Deng Chunqiu Steven Xia Haoran Peng Chenyuan Yang and Lingming Zhang. 2023. Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models. In Proceedings of the 32nd ACM SIGSOFT international symposium on software testing and analysis. 423\u2013435.","DOI":"10.1145\/3597926.3598067"},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","unstructured":"Yinlin Deng Chunqiu Steven Xia Chenyuan Yang Shizhuo Dylan Zhang Shujing Yang and Lingming Zhang. 2024. Large language models are edge-case generators: Crafting unusual programs for fuzzing deep learning libraries. In Proceedings of the 46th IEEE\/ACM International Conference on Software Engineering. 1\u201313.","DOI":"10.1145\/3597503.3623343"},{"key":"e_1_3_1_10_2","first-page":"44","volume-title":"ESEC\/SIGSOFT FSE","author":"Deng Yinlin","year":"2022","unstructured":"Yinlin Deng, Chenyuan Yang, Anjiang Wei, and Lingming Zhang. 2022. Fuzzing deep-learning libraries via automated relational API inference. In ESEC\/SIGSOFT FSE. ACM, 44\u201356."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/MET.2019.00008"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3213846.3213858"},{"key":"e_1_3_1_13_2","first-page":"5539","volume-title":"ACL (1)","author":"Gao Fei","year":"2019","unstructured":"Fei Gao, Jinhua Zhu, Lijun Wu, Yingce Xia, Tao Qin, Xueqi Cheng, Wengang Zhou, and Tie-Yan Liu. 2019. Soft Contextual Data Augmentation for Neural Machine Translation. In ACL (1). Association for Computational Linguistics, 5539\u20135544."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510092"},{"key":"e_1_3_1_15_2","first-page":"71","volume-title":"ICSE (SEIP)","author":"Gulzar Muhammad Ali","year":"2019","unstructured":"Muhammad Ali Gulzar, Yongkang Zhu, and Xiaofeng Han. 2019. Perception and practices of differential testing. In ICSE (SEIP). IEEE \/ ACM, 71\u201380."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416571"},{"key":"e_1_3_1_17_2","unstructured":"Huawei. 2020. MindSpore. https:\/\/gitee.com\/mindspore\/mindspore."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3548606.3560578"},{"key":"e_1_3_1_19_2","first-page":"510","volume-title":"ESEC\/SIGSOFT FSE","author":"Islam Johirul","year":"2019","unstructured":"Md Johirul Islam, Giang Nguyen, Rangeet Pan, and Hridesh Rajan. 2019. A comprehensive study on deep learning bug characteristics. In ESEC\/SIGSOFT FSE. ACM, 510\u2013520."},{"key":"e_1_3_1_20_2","first-page":"1135","volume-title":"ICSE","author":"Islam Johirul","year":"2020","unstructured":"Md Johirul Islam, Rangeet Pan, Giang Nguyen, and Hridesh Rajan. 2020. Repairing deep neural networks: fix patterns and challenges. In ICSE. ACM, 1135\u20131146."},{"key":"e_1_3_1_21_2","first-page":"604","volume-title":"DASFAA (1) (Lecture Notes in Computer Science","author":"Jia Li","year":"2020","unstructured":"Li Jia, Hao Zhong, Xiaoyin Wang, Linpeng Huang, and Xuansheng Lu. 2020. An Empirical Study on Bugs Inside TensorFlow. In DASFAA (1) (Lecture Notes in Computer Science, Vol. 12112). Springer, 604\u2013620."},{"key":"e_1_3_1_22_2","first-page":"1","volume-title":"Proceedings of Machine Learning and Systems 2020, MLSys 2020, Austin, TX, USA, March 2-4, 2020","author":"Jiang Xiaotang","year":"2020","unstructured":"Xiaotang Jiang, Huan Wang, Yiliu Chen, Ziqi Wu, Lichuan Wang, Bin Zou, Yafeng Yang, Zongyang Cui, Yu Cai, Tianhang Yu, Chengfei Lyu, and Zhihua Wu. 2020. MNN: A Universal and Efficient Inference Engine. In Proceedings of Machine Learning and Systems 2020, MLSys 2020, Austin, TX, USA, March 2-4, 2020, Inderjit S. Dhillon, Dimitris S. Papailiopoulos, and Vivienne Sze (Eds.), Vol. 2. mlsys.org, 1\u201313."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2208.01508"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3524846.3527341"},{"key":"e_1_3_1_25_2","unstructured":"Christian Murphy and Gail E Kaiser. 2010. Empirical evaluation of approaches to testing applications without test oracles. (2010)."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2896880"},{"key":"e_1_3_1_27_2","unstructured":"Adam Paszke Sam Gross Francisco Massa Adam Lerer James Bradbury Gregory Chanan Trevor Killeen Zeming Lin Natalia Gimelshein Luca Antiga Alban Desmaison Andreas Kopf Edward Z. Yang Zachary DeVito Martin Raison Alykhan Tejani Sasank Chilamkurthy Benoit Steiner Lu Fang Junjie Bai and Soumith Chintala. 2019. PyTorch: An Imperative Style High-Performance Deep Learning Library. In NeurIPS. 8024\u20138035."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3361566"},{"key":"e_1_3_1_29_2","first-page":"1027","volume-title":"ICSE","author":"Pham Hung Viet","year":"2019","unstructured":"Hung Viet Pham, Thibaud Lutellier, Weizhen Qi, and Lin Tan. 2019. CRADLE: cross-backend validation to detect and localize bugs in deep learning libraries. In ICSE. IEEE \/ ACM, 1027\u20131038."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510454.3516835"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2016.2532875"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2301.08653"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1504\/IJWGS.2020.110945"},{"key":"e_1_3_1_34_2","doi-asserted-by":"crossref","unstructured":"Jiannan Wang Thibaud Lutellier Shangshu Qian Hung Viet Pham and Lin Tan. 2022. EAGLE: creating equivalent graphs to test deep learning libraries. In Proceedings of the 44th International Conference on Software Engineering. 798\u2013810.","DOI":"10.1145\/3510003.3510165"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2020.07.042"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409761"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510041"},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"761","DOI":"10.1145\/3591251","article-title":"Optimal Reads-From Consistency Checking for C11-Style Memory Models","volume":"7","author":"Windsor Matt","year":"2023","unstructured":"Matt Windsor, Alastair F. Donaldson, and John Wickerson. 2023. Optimal Reads-From Consistency Checking for C11-Style Memory Models. Proceedings of the ACM on Programming Languages 7, PLDI (2023), 761\u2013785.","journal-title":"Proceedings of the ACM on Programming Languages"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3533767.3534220"},{"key":"e_1_3_1_40_2","first-page":"135","volume-title":"QSIC","author":"Xie Xiaoyuan","year":"2009","unstructured":"Xiaoyuan Xie, Joshua Wing Kei Ho, Christian Murphy, Gail E. Kaiser, Baowen Xu, and Tsong Yueh Chen. 2009. Application of Metamorphic Testing to Supervised Classifiers. In QSIC. IEEE Computer Society, 135\u2013144."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","unstructured":"Xiaoyuan Xie Joshua Wing Kei Ho Christian Murphy Gail E. Kaiser Baowen Xu and Tsong Yueh Chen. 2011. Testing and validating machine learning classifiers by metamorphic testing. J. Syst. Softw. 84 4 (2011) 544\u2013558. https:\/\/doi.org\/10.1016\/j.jss.2010.11.920 10.1016\/j.jss.2010.11.920","DOI":"10.1016\/j.jss.2010.11.920"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3468264.3468612"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2302.04351"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TR.2021.3107165"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460319.3464843"},{"key":"e_1_3_1_46_2","first-page":"129","volume-title":"ISSTA","author":"Zhang Yuhao","year":"2018","unstructured":"Yuhao Zhang, Yifan Chen, Shing-Chi Cheung, Yingfei Xiong, and Lu Zhang. 2018. An empirical study on TensorFlow program bugs. In ISSTA. ACM, 129\u2013140."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409720"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/MET52542.2021.00010"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3533767.3534409"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3660796","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3660796","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T07:55:02Z","timestamp":1770191702000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3660796"}},"issued":{"date-parts":[[2024,7,12]]},"references-count":48,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2024,7,12]]}},"alternative-id":["10.1145\/3660796"],"URL":"https:\/\/doi.org\/10.1145\/3660796","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2024,7,12]]}},{"indexed":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T14:30:56Z","timestamp":1787495456240,"version":"build-2736575974"},"reference-count":47,"publisher":"Association for Computing Machinery (ACM)","issue":"ISSTA","license":[{"start":{"date-parts":[[2025,6,22]],"date-time":"2025-06-22T00:00:00Z","timestamp":1750550400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,22]]},"abstract":"<jats:p>\n                    Ransomware encrypts files on infected systems and demands a hefty ransom for decryption, posing a significant threat to both enterprises and individuals. However, existing methods fail to capture the encryption preferences of diverse ransomware families, lacking an efficient and systematic proactive defense method. In this paper, we propose\n                    <jats:bold>Pepper,<\/jats:bold>\n                    a preference-aware active ransomware trapping method, covering decoy file generation, deployment, and monitoring. Through examination of numerous ransomware families, we have identified two prevalent encryption preferences: encryption file preferences and encryption path preferences. Deploying decoy files aligned with ransomware\u2019s encryption preferences within its preferred pathways provides an opportunity for efficient and early trapping of ransomware. Pepper combines a GNN-based recommendation model with expert insights to unveil the encryption file and path preferences across various ransomware families, guiding the generation and deployment of decoy files. Moreover, a decoy file monitor is designed to continuously monitor decoy file changes and promptly respond to anomalies. Extensive experiments show that Pepper achieves a 98.68% detection rate for ransomware, with an average file loss of 2.27. Moreover, it exhibits robustness in detecting unknown ransomware variants and does not interfere with regular users.\n                  <\/jats:p>","DOI":"10.1145\/3728932","type":"journal-article","created":{"date-parts":[[2025,6,22]],"date-time":"2025-06-22T10:52:56Z","timestamp":1750589576000},"page":"1280-1302","source":"Crossref","is-referenced-by-count":1,"title":["Pepper: Preference-Aware Active Trapping for Ransomware"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-0950-7514","authenticated-orcid":false,"given":"Huan","family":"Zhang","sequence":"first","affiliation":[{"name":"Institute of Information Engineering at Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, School of Cyber Security, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4008-4263","authenticated-orcid":false,"given":"Zhengkai","family":"Qin","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering at Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, School of Cyber Security, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-0566-791X","authenticated-orcid":false,"given":"Lixin","family":"Zhao","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering at Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5521-4757","authenticated-orcid":false,"given":"Aimin","family":"Yu","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering at Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, School of Cyber Security, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1710-8881","authenticated-orcid":false,"given":"Lijun","family":"Cai","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering at Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, School of Cyber Security, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-9868-5353","authenticated-orcid":false,"given":"Dan","family":"Meng","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering at Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, School of Cyber Security, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,22]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/b978-0-12-815739-8.00012-2"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-88418-5_12"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2022.3181278"},{"key":"e_1_3_1_5_2","first-page":"3005","volume-title":"30th USENIX Security Symposium, USENIX Security 2021, August 11-13, 2021","author":"Alsaheel Abdulellah","year":"2021","unstructured":"Abdulellah Alsaheel, Yuhong Nan, Shiqing Ma, Le Yu, Gregory Walkup, Z. Berkay Celik, Xiangyu Zhang, and Dongyan Xu. 2021. ATLAS: A Sequence-based Learning Approach for Attack Investigation. In 30th USENIX Security Symposium, USENIX Security 2021, August 11-13, 2021, Michael D. Bailey and Rachel Greenstadt (Eds.). USENIX Association, 3005\u20133022. https:\/\/www.usenix.org\/conference\/usenixsecurity21\/presentation\/alsaheel"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1016\/J.COSE.2024.104203"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/2991079.2991110"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.1136800"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2023.3240025"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.diin.2009.06.016"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-22038-9_11"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/J.COSE.2017.11.019"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","unstructured":"Wajih Ul Hassan Shengjian Guo Ding Li Zhengzhang Chen Kangkook Jee Zhichun Li and Adam Bates. 2019. NoDoze: Combatting Threat Alert Fatigue with Automated Provenance Triage. In 26th Annual Network and Distributed System Security Symposium NDSS 2019 San Diego California USA February 24-27 2019. The Internet Society. doi:10.14722\/ndss.2019.23349","DOI":"10.14722\/ndss.2019.23349"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-25133-2"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639090"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eij.2015.06.005"},{"key":"e_1_3_1_17_2","first-page":"757","volume-title":"25th USENIX Security Symposium, USENIX Security 16, Austin, TX, USA, August 10-12, 2016","author":"Kharraz Amin","year":"2016","unstructured":"Amin Kharraz, Sajjad Arshad, Collin Mulliner, William K. Robertson, and Engin Kirda. 2016. UNVEIL: A Large-Scale, Automated Approach to Detecting Ransomware. In 25th USENIX Security Symposium, USENIX Security 16, Austin, TX, USA, August 10-12, 2016, Thorsten Holz and Stefan Savage (Eds.). USENIX Association, 757\u2013772. https:\/\/www.usenix.org\/conference\/usenixsecurity16\/technical-sessions\/presentation\/kharaz"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-66332-6_5"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3129676.3129713"},{"key":"e_1_3_1_20_2","first-page":"212","article-title":"Ransomware Detection and Prevention through Strategically Hidden Decoy File","volume":"25","author":"Lin Yung-She","year":"2023","unstructured":"Yung-She Lin and Chin-Feng Lee. 2023. Ransomware Detection and Prevention through Strategically Hidden Decoy File. Int. J. Netw. Secur 25 (2023), 212\u2013220. http:\/\/ijns.jalaxy.com.tw\/contents\/ijns-v25-n2\/ijns-2023-v25-n2-p212-220.pdf","journal-title":"Int. J. Netw. Secur"},{"key":"e_1_3_1_21_2","unstructured":"MalwareBazaar. 2023. https:\/\/bazaar.abuse.ch\/."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/DASC\/PICOM\/DATACOM\/CYBERSCITEC.2018.00124"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-00470-5_6"},{"key":"e_1_3_1_24_2","unstructured":"Microsoft. 2021. Event Tracing for Windows (ETW). https:\/\/learn.microsoft.com\/en-us\/windows-hardware\/drivers\/devtest\/event-tracing-for-windows--etw-."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3319535.3363217"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACSAC.2007.21"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3514229"},{"key":"e_1_3_1_28_2","first-page":"729","volume-title":"28th USENIX Security Symposium, USENIX Security 2019, Santa Clara, CA, USA, August 14-16, 2019","author":"Pendlebury Feargus","year":"2019","unstructured":"Feargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder, and Lorenzo Cavallaro. 2019. TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time. In 28th USENIX Security Symposium, USENIX Security 2019, Santa Clara, CA, USA, August 14-16, 2019, Nadia Heninger and Patrick Traynor (Eds.). USENIX Association, 729\u2013746. https:\/\/www.usenix.org\/conference\/usenixsecurity19\/presentation\/pendlebury"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/SSCI.2018.8628743"},{"key":"e_1_3_1_30_2","unstructured":"PyTorch. 2024. PyG Documentation. https:\/\/pytorch-geometric.readthedocs.io\/en\/latest\/."},{"key":"e_1_3_1_31_2","unstructured":"PyTorch. 2024. PyTorch. https:\/\/pytorch.org\/."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-41187-3"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS.2016.46"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1009804230409"},{"key":"e_1_3_1_35_2","unstructured":"Daniele Sgandurra Luis Mu\u00f1oz-Gonz\u00e1lez Rabih Mohsen and Emil C. Lupu. 2016. Automated Dynamic Analysis of Ransomware: Benefits Limitations and use for Detection. CoRR abs\/1609.03020 (2016). arXiv:1609.03020 http:\/\/arxiv.org\/abs\/1609.03020"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMSNETS.2018.8328219"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/J.COMPELECENG.2022.108346"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICACCI.2018.8554938"},{"key":"e_1_3_1_39_2","unstructured":"Sophos. 2024. The State of Ransomware 2023. https:\/\/assets.sophos.com\/X24WTUEQ\/at\/c949g7693gsnjh9rb9gr8\/sophos-state-of-ransomware-2023-wp.pdf."},{"key":"e_1_3_1_40_2","unstructured":"Statista. 2024. Annual share of organizations affected by ransomware attacks worldwide from 2018 to 2023. https:\/\/statista.com\/statistics\/204457\/businesses-ransomware-attack-rate\/."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1016\/J.COSE.2020.101997"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/NOMS56928.2023.10154378"},{"key":"e_1_3_1_43_2","unstructured":"VirusShare. 2023. https:\/\/virusshare.com\/."},{"key":"e_1_3_1_44_2","unstructured":"VirusTotal. 2023. https:\/\/www.virustotal.com\/."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3658644.3690269"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3331184.3331267"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2024.3410511"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/SP46215.2023.10179372"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728932","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T13:49:08Z","timestamp":1787492948000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728932"}},"issued":{"date-parts":[[2025,6,22]]},"references-count":47,"journal-issue":{"issue":"ISSTA","published-print":{"date-parts":[[2025,6,22]]}},"alternative-id":["10.1145\/3728932"],"URL":"https:\/\/doi.org\/10.1145\/3728932","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2025,6,22]]}},{"indexed":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T11:21:03Z","timestamp":1787484063945,"version":"build-2736575974"},"reference-count":80,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:00:00Z","timestamp":1750291200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["2211454"],"award-info":[{"award-number":["2211454"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,19]]},"abstract":"<jats:p>Bug report reproduction is a crucial but time-consuming task to be carried out during mobile app maintenance. To accelerate this process, researchers have developed automated techniques for reproducing mobile app bug reports. However, due to the lack of an effective mechanism to recognize different buggy behaviors described in the report, existing work is limited to reproducing crash bug reports, or requires developers to manually analyze execution traces to determine if a bug was successfully reproduced. To address this limitation, we introduce a novel technique to automatically identify and extract the buggy behavior from the bug report and detect it during the automated reproduction process. To accommodate various buggy behaviors of mobile app bugs, we conducted an empirical study and created a standardized representation for expressing the bug behavior identified from our study. Given a report, our approach first transforms the documented buggy behavior into this standardized representation, then matches it against real-time device and UI information during the reproduction to recognize the bug. Our empirical evaluation demonstrated that our approach achieved over 90% precision and recall in generating the standardized representation of buggy behaviors. It correctly identified bugs in 83% of the bug reports and enhanced existing reproduction techniques, allowing them to reproduce four times more bug reports.<\/jats:p>","DOI":"10.1145\/3729370","type":"journal-article","created":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T11:15:34Z","timestamp":1750331734000},"page":"2240-2263","source":"Crossref","is-referenced-by-count":1,"title":["Automated Recognition of Buggy Behaviors from Mobile Bug Reports"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1333-1637","authenticated-orcid":false,"given":"Zhaoxu","family":"Zhang","sequence":"first","affiliation":[{"name":"University of Southern California, Los Angeles, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0199-3424","authenticated-orcid":false,"given":"Komei","family":"Ryu","sequence":"additional","affiliation":[{"name":"University of Southern California, Los Angeles, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9461-4251","authenticated-orcid":false,"given":"Tingting","family":"Yu","sequence":"additional","affiliation":[{"name":"University of Connecticut, Storrs, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4951-9367","authenticated-orcid":false,"given":"William G.J.","family":"Halfond","sequence":"additional","affiliation":[{"name":"University of Southern California, Los Angeles, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,19]]},"reference":[{"key":"e_1_3_1_1_2","unstructured":"2014. Replication Package. https:\/\/github.com\/USC-SQL\/BugSpot-Artifact."},{"key":"e_1_3_1_2_2","unstructured":"2019. GPT-4 website. https:\/\/openai.com\/index\/gpt-4\/."},{"key":"e_1_3_1_3_2","unstructured":"2019. Issue 121 for ATimeTracker Github Repository. https:\/\/github.com\/netmackan\/ATimeTracker\/issues\/121."},{"key":"e_1_3_1_4_2","unstructured":"2019. Issue 184 for MaterialFiles Github Repository. https:\/\/github.com\/zhanghai\/MaterialFiles\/issues\/184."},{"key":"e_1_3_1_5_2","unstructured":"2019. Issue 2191 for Collect Github Repository. https:\/\/github.com\/getodk\/collect\/issues\/2191."},{"key":"e_1_3_1_6_2","unstructured":"2019. Issue 634 for Omni-Notes Github Repository. https:\/\/github.com\/federicoiosue\/Omni-Notes\/issues\/634."},{"key":"e_1_3_1_7_2","unstructured":"2019. Notifications are not cleared on display of chat. https:\/\/github.com\/deltachat\/deltachat-android\/issues\/725."},{"key":"e_1_3_1_8_2","unstructured":"2019. Podcast Cover Disappears on Device Rotation #2992. https:\/\/github.com\/AntennaPod\/AntennaPod\/issues\/2992."},{"key":"e_1_3_1_9_2","unstructured":"2019. ReCDroid Github repository. https:\/\/github.com\/AndroidTestBugReport\/ReCDroid."},{"key":"e_1_3_1_10_2","unstructured":"2019. ReproBot\u2019s Github repository. https:\/\/github.com\/USC-SQL\/ReproBot-Artifact."},{"key":"e_1_3_1_11_2","unstructured":"2019. Roam\u2019s Replication Package. https:\/\/zenodo.org\/records\/11068809."},{"key":"e_1_3_1_12_2","unstructured":"2020. Chromecast controls disappear immediately. https:\/\/github.com\/jellyfin\/jellyfin-android\/issues\/459."},{"key":"e_1_3_1_13_2","unstructured":"2020. I can\u2019t gesture type words with the first letter upper-case and others lower-case. https:\/\/github.com\/AnySoftKeyboard\/AnySoftKeyboard\/issues\/2825."},{"key":"e_1_3_1_14_2","unstructured":"2020. Impossible to reproduce any video in Android App. https:\/\/github.com\/nextcloud\/android\/issues\/7602."},{"key":"e_1_3_1_15_2","unstructured":"2021. Audio file playing can\u2019t be stopped unless app is closed. No media controls to be found. https:\/\/github.com\/nextcloud\/android\/issues\/8905."},{"key":"e_1_3_1_16_2","unstructured":"2023. Github Issue Tracker. https:\/\/github.com\/issues."},{"key":"e_1_3_1_17_2","unstructured":"2023. Google Code Issue Tracker. https:\/\/code.google.com\/archive\/."},{"key":"e_1_3_1_18_2","unstructured":"2023. Sample size determination. https:\/\/en.wikipedia.org\/wiki\/Sample_size_determination."},{"key":"e_1_3_1_19_2","unstructured":"2024. Android ADB Shell. https:\/\/developer.android.com\/tools\/adb."},{"key":"e_1_3_1_20_2","unstructured":"2024. Android Checkbox. https:\/\/developer.android.com\/reference\/android\/widget\/CheckBox."},{"key":"e_1_3_1_21_2","unstructured":"2024. Android Crash Handler. https:\/\/developer.android.com\/games\/optimize\/crash."},{"key":"e_1_3_1_22_2","unstructured":"2024. FDroid. https:\/\/f-droid.org\/en\/."},{"key":"e_1_3_1_23_2","unstructured":"2024. GitHub. https:\/\/github.com."},{"key":"e_1_3_1_24_2","unstructured":"2024. Hugging Face: Llama-3.1-70B. https:\/\/huggingface.co\/meta-llama\/Llama-3.1-70B."},{"key":"e_1_3_1_25_2","unstructured":"2024. Keyboard keeps showing after opening settings menu. https:\/\/github.com\/flex3r\/DankChat\/issues\/66."},{"key":"e_1_3_1_26_2","unstructured":"2024. Openhab Android Issue #2523. https:\/\/github.com\/openhab\/openhab-android\/issues\/2583."},{"key":"e_1_3_1_27_2","unstructured":"2024. SpotiFlyer Issue #2523. https:\/\/github.com\/Shabinder\/SpotiFlyer\/issues\/764."},{"key":"e_1_3_1_28_2","unstructured":"2024. UI Automator. https:\/\/developer.android.com\/training\/testing\/other-components\/ui-automator."},{"key":"e_1_3_1_29_2","unstructured":"2024. Unstoppable Wallet Issue #3763. https:\/\/github.com\/horizontalsystems\/unstoppable-wallet-android\/issues\/3763."},{"key":"e_1_3_1_30_2","volume-title":"Proceedings of the 21st International Conference on Mining Software Repositories","author":"Baral Kesina","year":"2024","unstructured":"Kesina Baral, Jack Johnson, Mattia Fazzini, Julia Rubin, Junayed Mahmud, Sabiha Salma, Jeff Offutt, and Kevin Moran. 2024. Automating GUI-based Test Oracles for Mobile Apps. In Proceedings of the 21st International Conference on Mining Software Repositories. Lisbon Portugal."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377811.3380328"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","unstructured":"Pamela Bhattacharya Liudmila Ulanova Iulian Neamtiu and Sai Charan Koduru. 2013. An Empirical Analysis of Bug Reports and Bug Fixing in Open Source Android Apps. In 2013 17th European Conference on Software Maintenance and Reengineering. 133\u2013143. https:\/\/doi.org\/10.1109\/CSMR.2013.23 10.1109\/CSMR.2013.23 ISSN: 1534-5351.","DOI":"10.1109\/CSMR.2013.23"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","unstructured":"Charles E. Brown. 1998. Coefficient of Variation. Springer Berlin Heidelberg Berlin Heidelberg 155\u2013157. https:\/\/doi.org\/10.1007\/978-3-642-80328-4_13 10.1007\/978-3-642-80328-4_13","DOI":"10.1007\/978-3-642-80328-4_13"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3338906.3338947"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3106237.3106285"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2014.2363469"},{"key":"e_1_3_1_37_2","volume-title":"Basics of qualitative research: Techniques and procedures for developing grounded theory","author":"Corbin Juliet M.","year":"2015","unstructured":"Juliet M. Corbin and Anselm L. Strauss. 2015. Basics of qualitative research: Techniques and procedures for developing grounded theory. SAGE."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","unstructured":"Mattia Fazzini and Alessandro Orso. 2017. Automated cross-platform inconsistency detection for mobile apps. In 2017 32nd IEEE\/ACM International Conference on Automated Software Engineering (ASE). 308\u2013318. https:\/\/doi.org\/10.1109\/ASE.2017.8115644 10.1109\/ASE.2017.8115644","DOI":"10.1109\/ASE.2017.8115644"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3213846.3213869"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510048"},{"key":"e_1_3_1_41_2","volume-title":"Proceedings of the 46th International Conference on Software Engineering","author":"Feng Sidong","year":"2024","unstructured":"Sidong Feng and Chunyang Chen. 2024. Prompting Is All Your Need: Automated Android Bug Replay with Large Language Models. In Proceedings of the 46th International Conference on Software Engineering. ACM, Portugal."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3533767.3534402"},{"key":"e_1_3_1_43_2","volume-title":"Proceedings of the 46th IEEE\/ACM International Conference on Software Engineering","author":"Huang Yuchao","year":"2024","unstructured":"Yuchao Huang, Junjie Wang, Zhe Liu, Yawen Wang, Song Wang, Chunyang Chen, Yuanzhe Hu, and Qing Wang. 2024. CrashTranslator: Automatically Reproducing Mobile Application Crashes Directly from Stack Trace. In Proceedings of the 46th IEEE\/ACM International Conference on Software Engineering. ACM. https:\/\/dl.acm.org\/doi\/abs\/10.1145\/3597503.3623298"},{"key":"e_1_3_1_44_2","volume-title":"2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE)","author":"Huang Yuchao","year":"2023","unstructured":"Yuchao Huang, Junjie Wang, Liu Zhe, Song Wang, Chunyang Chen, Mingyang Li, and Qing Wang. 2023. Context-aware Bug Reproduction for Mobile Apps. In 2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). IEEE, Melbourne, Australia."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","unstructured":"Ajay Kumar Jha Sunghee Lee and Woo Jin Lee. 2019. Characterizing Android-Specific Crash Bugs. In 2019 IEEE\/ACM 6th International Conference on Mobile Software Engineering and Systems (MOBILESoft). 111\u2013122. https:\/\/doi.org\/10.1109\/MOBILESoft.2019.00024 10.1109\/MOBILESoft.2019.00024","DOI":"10.1109\/MOBILESoft.2019.00024"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/SANER53432.2022.00048"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2407.19053"},{"key":"e_1_3_1_48_2","volume-title":"Content Analysis: An Introduction to Its Methodology (second edition)","author":"Krippendorff Klaus","year":"2004","unstructured":"Klaus Krippendorff. 2004. Content Analysis: An Introduction to Its Methodology (second edition). Sage Publications."},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639167"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3691620.3695476"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","unstructured":"Gaoyi Lin Zhihua Zhang and Zhanqi Cui. 2023. Widget Hierarchy Graph Guided Crash Reproduction Method for Android Applications (S). In The 35th International Conference on Software Engineering and Knowledge Engineering {SEKE} 2023. 584\u2013587. https:\/\/doi.org\/10.18293\/SEKE2023-066 10.18293\/SEKE2023-066","DOI":"10.18293\/SEKE2023-066"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","unstructured":"Mario Linares-V\u00e1squez Gabriele Bavota and Camilo Escobar-Vel\u00e1squez. 2017. An Empirical Study on Android-Related Vulnerabilities. In 2017 IEEE\/ACM 14th International Conference on Mining Software Repositories (MSR). 2\u201313. https:\/\/doi.org\/10.1109\/MSR.2017.60 10.1109\/MSR.2017.60","DOI":"10.1109\/MSR.2017.60"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/2568225.2568229"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00119"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639180"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416547"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2024.3428543"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3551349.3560424"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSME46990.2020.00063"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.5555\/3294996.3295163"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3551349.3556935"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3611643.3616286"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460319.3464806"},{"issue":"5","key":"e_1_3_1_64_2","first-page":"360","article-title":"Understanding interobserver agreement: the kappa statistic","volume":"37","author":"Viera Anthony J.","year":"2005","unstructured":"Anthony J. Viera and Joanne Mills Garrett. 2005. Understanding interobserver agreement: the kappa statistic. Family medicine 37 5 (2005), 360\u2013363. https:\/\/api.semanticscholar.org\/CorpusID:38150955","journal-title":"Family medicine"},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","unstructured":"Dingbang Wang Zhaoxu Zhang Sidong Feng William G.J. Halfond and Tingting Yu. 2025. An Empirical Study on Leveraging Images in Automated Bug Report Reproduction. In Proceedings of the 22rd International Conference on Mining Software Repositories.","DOI":"10.1109\/MSR66628.2025.00019"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1145\/3650212.3680341"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549170"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.5555\/3600270.3602070"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSR52588.2021.00082"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3691620.3694986"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598138"},{"key":"e_1_3_1_72_2","doi-asserted-by":"publisher","unstructured":"Ali Asghar Yarifard Saeed Araban Samad Paydar Vahid Garousi Maurizio Morisio and Riccardo Coppola. 2024. Extraction and empirical evaluation of GUI-level invariants as GUI Oracles in mobile app testing. Information and Software Technology (July 2024) 107531. https:\/\/doi.org\/10.1016\/j.infsof.2024.107531 10.1016\/j.infsof.2024.107531","DOI":"10.1016\/j.infsof.2024.107531"},{"key":"e_1_3_1_73_2","unstructured":"Juyeon Yoon Robert Feldt and Shin Yoo. 2023. Autonomous Large Language Model Agents Enabling Intent-Driven Mobile GUI Testing. http:\/\/arxiv.org\/abs\/2311.08649 arXiv:2311.08649 [cs]."},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICST.2014.31"},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/3660824"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598066"},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.1145\/3488244"},{"key":"e_1_3_1_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019.00030"},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00049"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2010.63"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729370","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729370","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T10:57:26Z","timestamp":1787482646000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3729370"}},"issued":{"date-parts":[[2025,6,19]]},"references-count":80,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2025,6,19]]}},"alternative-id":["10.1145\/3729370"],"URL":"https:\/\/doi.org\/10.1145\/3729370","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2025,6,19]]}},{"indexed":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T11:21:13Z","timestamp":1787484073466,"version":"build-2736575974"},"reference-count":91,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:00:00Z","timestamp":1750291200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,19]]},"abstract":"<jats:p>Programming is an essential activity in data science (DS). Unlike regular software developers, DS programmers often use Jupyter notebooks instead of conventional IDEs. Moreover, DS programmers focus on statistics, data analytics, and modeling rather than writing production-ready code following best practices in software engineering. Thus, in order to provide effective tool support to improve their productivity, it is important to understand what kinds of errors they make and how they fix them. Previous studies have analyzed DS code from public code-sharing platforms such as GitHub and Kaggle. However, they only accounted for code changes committed to the version history, omitting many programming mistakes that are resolved before code commits. To bridge the gap, we present an in-depth analysis of the fine-grained logs of a DS competition, which includes 390 Jupyter Notebooks written by 67 participants over six weeks. In addition, we conducted semi-structured interviews with 10 DS programmers from different domains to understand the reasons behind their programming mistakes. We identified several unique programming mistakes and fix patterns that had not been reported before, highlighting opportunities for designing new tool support for DS programming.<\/jats:p>","DOI":"10.1145\/3729352","type":"journal-article","created":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T11:15:34Z","timestamp":1750331734000},"page":"1824-1846","source":"Crossref","is-referenced-by-count":2,"title":["Towards Understanding Fine-Grained Programming Mistakes and Fixing Patterns in Data Science"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-9108-8473","authenticated-orcid":false,"given":"Wei-Hao","family":"Chen","sequence":"first","affiliation":[{"name":"Purdue University, Computer Science, West Lafayette, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4626-5509","authenticated-orcid":false,"given":"Jia Lin","family":"Cheoh","sequence":"additional","affiliation":[{"name":"Purdue University, West Lafayette, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3582-5091","authenticated-orcid":false,"given":"Manthan","family":"Keim","sequence":"additional","affiliation":[{"name":"Purdue University, West Lafayette, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8631-0955","authenticated-orcid":false,"given":"Sabine","family":"Brunswicker","sequence":"additional","affiliation":[{"name":"Purdue University, West Lafayette, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5468-9347","authenticated-orcid":false,"given":"Tianyi","family":"Zhang","sequence":"additional","affiliation":[{"name":"Purdue University, West Lafayette, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,19]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2019. The Data Scientist Profile 2019 - Skills Experience Education Of 1 001 Data Scientists. https:\/\/365datascience.com\/career-advice\/career-guides\/data-scientist-profile\/"},{"key":"e_1_3_1_3_2","unstructured":"2024. IPython. https:\/\/ipython.readthedocs.io\/"},{"key":"e_1_3_1_4_2","unstructured":"2024. Jupyter Notebook. https:\/\/jupyter.org\/"},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","unstructured":"Iago Abai Claus Brabrand and Andrzej Wasowski. 2014. 42 variability bugs in the linux kernel: a qualitative analysis. In Proceedings of the 29th ACM\/IEEE international conference on Automated software engineering. 421\u2013432.","DOI":"10.1145\/2642937.2642990"},{"key":"e_1_3_1_6_2","unstructured":"Shibbir Ahmed Mohammad Wardat Hamid Bagheri Breno Dantas Cruz and Hridesh Rajan. 2023. Characterizing Bugs in Python and R Data Analytics Programs. arXiv preprint arXiv:2306.08632 (2023)."},{"key":"e_1_3_1_7_2","doi-asserted-by":"crossref","first-page":"293","DOI":"10.1109\/ASE.2019.00036","volume-title":"2019 34th IEEE\/ACM International Conference on Automated Software Engineering (ASE)","author":"Rahat Tamjid Al","year":"2019","unstructured":"Tamjid Al Rahat, Yu Feng, and Yuan Tian. 2019. Oauthlint: An empirical study on oauth bugs in android applications. In 2019 34th IEEE\/ACM International Conference on Automated Software Engineering (ASE). IEEE, 293\u2013304."},{"issue":"5","key":"e_1_3_1_8_2","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1007\/s10664-023-10352-5","article-title":"What constitutes debugging? An exploratory study of debugging episodes","volume":"28","author":"Alaboudi Abdulaziz","year":"2023","unstructured":"Abdulaziz Alaboudi and Thomas D LaToza. 2023. What constitutes debugging? An exploratory study of debugging episodes. Empirical Software Engineering 28, 5 (2023), 117.","journal-title":"Empirical Software Engineering"},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","first-page":"129","DOI":"10.1109\/ASE51524.2021.9678696","volume-title":"2021 36th IEEE\/ACMInternational Conference on Automated Software Engineering (ASE)","author":"Bavishi Rohan","year":"2021","unstructured":"Rohan Bavishi, Shadaj Laddad, Hiroaki Yoshida, Mukul R Prasad, and Koushik Sen. 2021. Vizsmith: Automated visualization synthesis by mining data-science notebooks. In 2021 36th IEEE\/ACMInternational Conference on Automated Software Engineering (ASE). IEEE, 129\u2013141."},{"key":"e_1_3_1_10_2","doi-asserted-by":"crossref","unstructured":"Moritz Beller Niels Spruit Diomidis Spinellis and Andy Zaidman. 2018. On the dichotomy of debugging behavior among programmers. In Proceedings of the 40th International Conference on Software Engineering. 572\u2013583.","DOI":"10.1145\/3180155.3180175"},{"key":"e_1_3_1_11_2","volume-title":"Qualitative research methods for the social sciences","author":"Berg Bruce Lawrence","year":"2001","unstructured":"Bruce Lawrence Berg. 2001. Qualitative research methods for the social sciences. Allyn & Bacon."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CSMR.2013.23"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Sumon Biswas Mohammad Wardat and Hridesh Rajan. 2022. The art and practice of data science pipelines: A comprehensive study of data science pipelines in theory in-the-small and in-the-large. In Proceedings of the 44th International Conference on Software Engineering. 2091\u20132103.","DOI":"10.1145\/3510003.3510057"},{"key":"e_1_3_1_14_2","unstructured":"Kelly Nicole Bodwin Ian Flores Siaca Amelia McNamara Philipp Burckhardt Allison Theobold Amal Abdel-Ghani and Greg Wilson. 2022. \"Looks okay to me\": A study of best practice in data analysis code review. In ICOTS."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1191\/1478088706qp063oa"},{"key":"e_1_3_1_16_2","first-page":"25","volume-title":"2019 IEEE symposium on visual languages and human-centric computing (VL\/HCC)","author":"Cai Carrie J","year":"2019","unstructured":"Carrie J Cai and Philip J Guo. 2019. Software developers learning machine learning: Motivations, hurdles, and desires. In 2019 IEEE symposium on visual languages and human-centric computing (VL\/HCC). IEEE, 25\u201334."},{"key":"e_1_3_1_17_2","doi-asserted-by":"crossref","unstructured":"Steven P Callahan Juliana Freire Emanuele Santos Carlos E Scheidegger Cl\u00e1udio T Silva and Huy T Vo. 2006. VisTrails: visualization meets data management. In Proceedings of the 2006 ACM SIGMOD international conference on Management of data. 745\u2013747.","DOI":"10.1145\/1142473.1142574"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","unstructured":"Souti Chattopadhyay Ishita Prasad Austin Z Henley Anita Sarma and Titus Barik. 2020. What\u2019s wrong with computational notebooks? Pain points needs and design opportunities. In Proceedings of the 2020 CHI conference on human factors in computing systems. 1\u201312.","DOI":"10.1145\/3313831.3376729"},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","unstructured":"Bhavya Chopra Anna Fariha Sumit Gulwani Austin Z Henley Daniel Perelman Mohammad Raza Sherry Shi Danny Simmons and Ashish Tiwari. 2023. CoWrangler: Recommender System for Data-Wrangling Scripts. In Companion of the 2023 International Conference on Management of Data. 147\u2013150.","DOI":"10.1145\/3555041.3589722"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.5555\/861869"},{"key":"e_1_3_1_21_2","unstructured":"Taijara Loiola de Santana Paulo Anselmo da Mota Silveira Neto Eduardo Santana de Almeida and Iftekhar Ahmed. 2022. Bug Analysis in Jupyter Notebook Projects: An Empirical Study. ACM Transactions on Software Engineering and Methodology (2022)."},{"key":"e_1_3_1_22_2","doi-asserted-by":"crossref","unstructured":"Will Epperson April Yi Wang Robert DeLine and Steven M Drucker. 2022. Strategies for reuse and sharing among data scientists in software teams. In Proceedings of the 44th International Conference on Software Engineering: Software Engineering in Practice. 243\u2013252.","DOI":"10.1145\/3510457.3513042"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"Robert B Evans and Alberto Savoia. 2007. Differential testing: a new approach to change detection. In The 6th Joint Meeting on European software engineering conference and the ACM SIGSOFT Symposium on the Foundations of Software Engineering: Companion Papers. 549\u2013552.","DOI":"10.1145\/1295014.1295038"},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","unstructured":"Jean-R\u00e9my Falleri Flor\u00e9al Morandat Xavier Blanc Matias Martinez and Martin Monperrus. 2014. Fine-grained and accurate source code differencing. In Proceedings of the 29th ACM\/IEEE international conference on Automated software engineering. 313\u2013324.","DOI":"10.1145\/2642937.2642982"},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","unstructured":"Konstantin Grotov Sergey Titov Vladimir Sotnikov Yaroslav Golubev and Timofey Bryksin. 2022. A large-scale comparison of Python code in Jupyter notebooks and scripts. In Proceedings of the 19th International Conference on Mining Software Repositories. 353\u2013364.","DOI":"10.1145\/3524842.3528447"},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","unstructured":"Ken Gu Eunice Jun and Tim Althoff. 2023. Understanding and supporting debugging workflows in multiverse analysis. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1\u201319.","DOI":"10.1145\/3544548.3581099"},{"key":"e_1_3_1_27_2","first-page":"137","volume-title":"Dependable Software Systems Engineering","author":"Gulwani Sumit","year":"2016","unstructured":"Sumit Gulwani. 2016. Programming by examples-and its applications in data wrangling. In Dependable Software Systems Engineering. IOS Press, 137\u2013158."},{"key":"e_1_3_1_28_2","first-page":"71","volume-title":"2019 IEEE\/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP)","author":"Gulzar Muhammad Ali","year":"2019","unstructured":"Muhammad Ali Gulzar, Yongkang Zhu, and Xiaofeng Han. 2019. Perception and practices of differential testing. In 2019 IEEE\/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 71\u201380."},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","unstructured":"Junjie Huang Daya Guo Chenglong Wang Jiazhen Gu Shuai Lu Jeevana Priya Inala Cong Yan Jianfeng Gao Nan Duan and Michael R Lyu. 2024. Contextualized Data-Wrangling Code Generation in Computational Notebooks. In Proceedings of the 39th IEEE\/ACM International Conference on Automated Software Engineering. 1282\u20131294.","DOI":"10.1145\/3691620.3695503"},{"key":"e_1_3_1_30_2","first-page":"254","volume-title":"2024 IEEE\/ACM 21st International Conference on Mining Software Repositories (MSR)","author":"Islam Md Anaytul","year":"2024","unstructured":"Md Anaytul Islam, Muhammad Asaduzzman, and Shaowei Wang. 2024. On the Executability of R Markdown Files. In 2024 IEEE\/ACM 21st International Conference on Mining Software Repositories (MSR). IEEE, 254\u2013264."},{"key":"e_1_3_1_31_2","doi-asserted-by":"crossref","unstructured":"Md Johirul Islam Giang Nguyen Rangeet Pan and Hridesh Rajan. 2019. A comprehensive study on deep learning bug characteristics. In Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 510\u2013520.","DOI":"10.1145\/3338906.3338955"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/APSEC.2016.025"},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","unstructured":"Ren\u00e9 Just Darioush Jalali and Michael D Ernst. 2014. Defects4J: A database of existing faults to enable controlled testing studies for Java programs. In Proceedings of the 2014 international symposium on software testing and analysis. 437\u2013440.","DOI":"10.1145\/2610384.2628055"},{"issue":"3","key":"e_1_3_1_34_2","article-title":"Jupyter notebooks on github: characteristics and code clones","volume":"5","author":"K\u00e4ll\u00e9n Malin","year":"2021","unstructured":"Malin K\u00e4ll\u00e9n, Ulf Sigvardsson, and Tobias Wrigstad. 2021. Jupyter notebooks on github: characteristics and code clones. The Art, Science, and Engineering of Programming 5, 3 (2021).","journal-title":"The Art, Science, and Engineering of Programming"},{"key":"e_1_3_1_35_2","doi-asserted-by":"crossref","unstructured":"Sean Kandel Andreas Paepcke Joseph Hellerstein and Jeffrey Heer. 2011. Wrangler: Interactive visual specification of data transformation scripts. In Proceedings of the sigchi conference on human factors in computing systems. 3363\u20133372.","DOI":"10.1145\/1978942.1979444"},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","unstructured":"Rafael-Michael Karampatsis and Charles Sutton. 2020. How often do single-statement bugs occur? the manysstubs4j dataset. In Proceedings of the 17th International Conference on Mining Software Repositories. 573\u2013577.","DOI":"10.1145\/3379597.3387491"},{"key":"e_1_3_1_37_2","unstructured":"Staffs Keele et al.. 2007. Guidelines for performing systematic literature reviews in software engineering. Technical Report. Technical report ver. 2.3 ebse technical report ebse."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/VLHCC.2017.8103446"},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","unstructured":"Miryung Kim Thomas Zimmermann Robert DeLine and Andrew Begel. 2016. The emerging role of data scientists on software development teams. In Proceedings of the 38th International Conference on Software Engineering. 96\u2013107.","DOI":"10.1145\/2884781.2884783"},{"key":"e_1_3_1_40_2","unstructured":"Thomas Kluyver Benjamin Ragan-Kelley Fernando P\u00e9rez Brian E Granger Matthias Bussonnier Jonathan Frederic Kyle Kelley Jessica B Hamrick Jason Grout Sylvain Corlay et al.. 2016. Jupyter Notebooks-a publishing format for reproducible computational workflows. In Elpub. 87\u201390."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1093\/comjnl\/27.2.97"},{"issue":"4","key":"e_1_3_1_42_2","doi-asserted-by":"crossref","first-page":"2020","DOI":"10.1109\/TSE.2022.3208210","article-title":"Impact of software engineering research in practice: A patent and author survey analysis","volume":"49","author":"Kotti Zoe","year":"2022","unstructured":"Zoe Kotti, Georgios Gousios, and Diomidis Spinellis. 2022. Impact of software engineering research in practice: A patent and author survey analysis. IEEE Transactions on Software Engineering 49, 4 (2022), 2020\u20132038.","journal-title":"IEEE Transactions on Software Engineering"},{"issue":"1","key":"e_1_3_1_43_2","first-page":"427","article-title":"Duet: Helping data analysis novices conduct pairwise comparisons by minimal specification","volume":"25","author":"Law Po-Ming","year":"2018","unstructured":"Po-Ming Law, Rahul C Basole, and Yanhong Wu. 2018. Duet: Helping data analysis novices conduct pairwise comparisons by minimal specification. IEEE transactions on visualization and computer graphics 25, 1 (2018), 427\u2013437.","journal-title":"IEEE transactions on visualization and computer graphics"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2015.2454513"},{"key":"e_1_3_1_45_2","first-page":"2078","volume-title":"2022 IEEE Symposium on Security and Privacy (SP)","author":"Lin Zhenpeng","year":"2022","unstructured":"Zhenpeng Lin, Yueqi Chen, Yuhang Wu, Dongliang Mu, Chensheng Yu, Xinyu Xing, and Kang Li. 2022. GREBE: Unveiling exploitation potential for Linux kernel bugs. In 2022 IEEE Symposium on Security and Privacy (SP). IEEE, 2078\u20132095."},{"key":"e_1_3_1_46_2","first-page":"381","volume-title":"2017 IEEE\/ACM 39th International Conference on Software Engineering (ICSE)","author":"Ma Wanwangying","year":"2017","unstructured":"Wanwangying Ma, Lin Chen, Xiangyu Zhang, Yuming Zhou, and Baowen Xu. 2017. How do developers fix cross-project correlated bugs? a case study on the github scientific python ecosystem. In 2017 IEEE\/ACM 39th International Conference on Software Engineering (ICSE). IEEE, 381\u2013392."},{"issue":"1","key":"e_1_3_1_47_2","first-page":"100","article-title":"Differential testing for software","volume":"10","author":"McKeeman William M","year":"1998","unstructured":"William M McKeeman. 1998. Differential testing for software. Digital Technical Journal 10, 1 (1998), 100\u2013107.","journal-title":"Digital Technical Journal"},{"issue":"9","key":"e_1_3_1_48_2","first-page":"1","article-title":"pandas: a foundational Python library for data analysis and statistics","volume":"14","author":"McKinney Wes","year":"2011","unstructured":"Wes McKinney et al.. 2011. pandas: a foundational Python library for data analysis and statistics. Python for high performance and scientific computing 14, 9 (2011), 1\u20139.","journal-title":"Python for high performance and scientific computing"},{"key":"e_1_3_1_49_2","doi-asserted-by":"crossref","unstructured":"Hailie Mitchell. 2022. Automatically Fixing Breaking Changes of Data Science Libraries. In Proceedings of the 37th IEEE\/ACM International Conference on Automated Software Engineering. 1\u20133.","DOI":"10.1145\/3551349.3559507"},{"key":"e_1_3_1_50_2","unstructured":"Paul Timothy Mooney. 2022. Kaggle Survey 2022: All Results. https:\/\/www.kaggle.com\/code\/paultimothymooney\/kaggle-survey-2022-all-results. Accessed: 2025."},{"key":"e_1_3_1_51_2","first-page":"112","volume-title":"2021 IEEE\/ACM 43rd International Conference on Software Engineering (ICSE)","author":"Ni Ansong","year":"2021","unstructured":"Ansong Ni, Daniel Ramos, Aidan ZH Yang, In\u00e9s Lynce, Vasco Manquinho, Ruben Martins, and Claire Le Goues. 2021. Soar: a synthesis approach for data science api refactoring. In 2021 IEEE\/ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 112\u2013124."},{"key":"e_1_3_1_52_2","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1109\/ESEM.2013.18","volume-title":"2013 ACM\/IEEE International Symposium on Empirical Software Engineering and Measurement","author":"Ocariza Frolin","year":"2013","unstructured":"Frolin Ocariza, Kartik Bajaj, Karthik Pattabiraman, and Ali Mesbah. 2013. An empirical study of client-side JavaScript bugs. In 2013 ACM\/IEEE International Symposium on Empirical Software Engineering and Measurement. IEEE, 55\u201364."},{"key":"e_1_3_1_53_2","doi-asserted-by":"crossref","first-page":"100","DOI":"10.1109\/ISSRE.2011.28","volume-title":"2011 IEEE 22nd International Symposium on Software Reliability Engineering","author":"Jr Frolin S Ocariza","year":"2011","unstructured":"Frolin S Ocariza Jr, Karthik Pattabiraman, and Benjamin Zorn. 2011. JavaScript errors in the wild: An empirical study. In 2011 IEEE 22nd International Symposium on Software Reliability Engineering. IEEE, 100\u2013109."},{"key":"e_1_3_1_54_2","doi-asserted-by":"crossref","unstructured":"Wonseok Oh and Hakjoo Oh. 2022. PyTER: effective program repair for Python type errors. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 922\u2013934.","DOI":"10.1145\/3540250.3549130"},{"key":"e_1_3_1_55_2","doi-asserted-by":"crossref","unstructured":"Yun Peng Shuzheng Gao Cuiyun Gao Yintong Huo and Michael Lyu. 2024. Domain knowledge matters: Improving prompts with fix templates for repairing python type errors. In Proceedings of the 46th IEEE\/ACM International Conference on Software Engineering. 1\u201313.","DOI":"10.1145\/3597503.3608132"},{"key":"e_1_3_1_56_2","doi-asserted-by":"crossref","unstructured":"Deepthi Raghunandan Aayushi Roy Shenzhi Shi Niklas Elmqvist and Leilani Battle. 2022. Code code evolution: Understanding how people change data science notebooks over time. arXiv preprint arXiv:2209.02851 (2022).","DOI":"10.1145\/3544548.3580997"},{"issue":"1","key":"e_1_3_1_57_2","first-page":"1","article-title":"Workflow analysis of data science code in public GitHub repositories","volume":"28","author":"Ramasamy Dhivyabharathi","year":"2023","unstructured":"Dhivyabharathi Ramasamy, Cristina Sarasua, Alberto Bacchelli, and Abraham Bernstein. 2023. Workflow analysis of data science code in public GitHub repositories. Empirical Software Engineering 28, 1 (2023), 1\u201347.","journal-title":"Empirical Software Engineering"},{"key":"e_1_3_1_58_2","first-page":"72","volume-title":"2023 IEEE\/ACM 45th International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER)","author":"Reimann Lars","year":"2023","unstructured":"Lars Reimann and G\u00fcnter Kniesei-W\u00fcnsche. 2023. Safe-DS: A Domain Specific Language to Make Data Science Safe. In 2023 IEEE\/ACM 45th International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER). IEEE, 72\u201377."},{"key":"e_1_3_1_59_2","doi-asserted-by":"crossref","unstructured":"Derek Robinson Neil A Ernst Enrique Larios Vargas and Margaret-Anne D Storey. 2022. Error identification strategies for Python Jupyter notebooks. In Proceedings of the 30th IEEE\/ACM International Conference on Program Comprehension. 253\u2013263.","DOI":"10.1145\/3524610.3529156"},{"key":"e_1_3_1_60_2","doi-asserted-by":"crossref","unstructured":"Adam Rule Aur\u00e9lien Tabard and James D Hollan. 2018. Exploration and explanation in computational notebooks. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems. 1\u201312.","DOI":"10.1145\/3173574.3173606"},{"issue":"2009","key":"e_1_3_1_61_2","doi-asserted-by":"crossref","first-page":"131","DOI":"10.1007\/s10664-008-9102-8","article-title":"Guidelines for conducting and reporting case study research in software engineering","volume":"14","author":"Runeson Per","year":"2009","unstructured":"Per Runeson and Martin H\u00f6st. 2009. Guidelines for conducting and reporting case study research in software engineering. Empirical software engineering 14 (2009), 131\u2013164.","journal-title":"Empirical software engineering"},{"key":"e_1_3_1_62_2","doi-asserted-by":"crossref","unstructured":"Ripon K Saha Yingjun Lyu Wing Lam Hiroaki Yoshida and Mukul R Prasad. 2018. Bugs jar: A large-scale diverse dataset of re al-world java bugs. In Proceedings of the 15th international conference on mining software repositories. 10\u201313.","DOI":"10.1145\/3196398.3196473"},{"key":"e_1_3_1_63_2","doi-asserted-by":"crossref","unstructured":"Marija Selakovic and Michael Pradel. 2016. Performance issues and optimizations in javascript: an empirical study. In Proceedings of the 38th International Conference on Software Engineering.61\u201372.","DOI":"10.1145\/2884781.2884829"},{"key":"e_1_3_1_64_2","doi-asserted-by":"crossref","unstructured":"Forrest Shull Janice Singer and Dag IK Sjoberg. 2008. Guide to advanced empirical software engineering. Vol. 93.Springer.","DOI":"10.1007\/978-1-84800-044-5"},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","first-page":"1033","DOI":"10.1109\/ASE51524.2021.9678873","volume-title":"2021 36th IEEE\/ACM International Conference on Automated Software Engineering (ASE)","author":"Sivasothy Shangeetha","year":"2021","unstructured":"Shangeetha Sivasothy. 2021. DSInfoSearch: supporting experimentation process of data scientists. In 2021 36th IEEE\/ACM International Conference on Automated Software Engineering (ASE). IEEE, 1033\u20131037."},{"key":"e_1_3_1_66_2","first-page":"1","article-title":"What\u2019s the difference? evaluating variations of multi-series bar charts for visual comparison tasks","author":"Srinivasan Arjun","year":"2018","unstructured":"Arjun Srinivasan, Matthew Brehmer, Bongshin Lee, and Steven M Drucker. 2018. What\u2019s the difference? evaluating variations of multi-series bar charts for visual comparison tasks. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems. 1\u201312.","journal-title":"Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems"},{"issue":"1","key":"e_1_3_1_67_2","doi-asserted-by":"crossref","first-page":"120","DOI":"10.1109\/TVCG.2018.2865024","article-title":"Knowledgepearls: Provenancebased visualization retrieval","volume":"25","author":"Stitz Holger","year":"2018","unstructured":"Holger Stitz, Samuel Gratzl, Harald Piringer, Thomas Zichner, and Marc Streit. 2018. Knowledgepearls: Provenancebased visualization retrieval. IEEE Transactions on Visualization and Computer Graphics 25, 1 (2018), 120\u2013130.","journal-title":"IEEE Transactions on Visualization and Computer Graphics"},{"key":"e_1_3_1_68_2","doi-asserted-by":"crossref","unstructured":"Pavle Suboti\u0107 Lazar Miliki\u0107 and Milan Stoji\u0107. 2022. A static analysis framework for data science notebooks. In Proceedings of the 44th International Conference on Software Engineering: Software Engineering in Practice. 13\u201322.","DOI":"10.1145\/3510457.3513032"},{"key":"e_1_3_1_69_2","doi-asserted-by":"crossref","first-page":"1964","DOI":"10.1109\/ICDE.2019.00215","volume-title":"2019 IEEE 35th International Conference on Data Engineering (ICDE)","author":"Tang Mingjie","year":"2019","unstructured":"Mingjie Tang, Saisai Shao, Weiqing Yang, Yanbo Liang, Yongyang Yu, Bikas Saha, and Dongjoon Hyun. 2019. Sac: A system for big data lineage tracking. In 2019 IEEE 35th International Conference on Data Engineering (ICDE). IEEE, 1964\u20131967."},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE.2012.22"},{"key":"e_1_3_1_71_2","unstructured":"John Wilder Tukey et al.. 1977. Exploratory data analysis. Vol. 2. Springer."},{"key":"e_1_3_1_72_2","volume-title":"Data science in action","author":"Aalst Wil Van Der","year":"2016","unstructured":"Wil Van Der Aalst and Wil van der Aalst. 2016. Data science in action. Springer."},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCSE.2011.37"},{"key":"e_1_3_1_74_2","first-page":"179","volume-title":"2021 IEEE\/ACM 18th International Conference on Mining Software Repositories (MSR)","author":"Vidoni Melina","year":"2021","unstructured":"Melina Vidoni. 2021. Self-admitted technical debt in r packages: An exploratory study. In 2021 IEEE\/ACM 18th International Conference on Mining Software Repositories (MSR). IEEE, 179\u2013189."},{"key":"e_1_3_1_75_2","doi-asserted-by":"crossref","unstructured":"Anh Duc Vu Timo Kehrer and Christos Tsigkanos. 2022. Outcome-preserving input reduction for scientific data analysis workflows. In Proceedings of the 37th IEEE\/ACM International Conference on Automated Software Engineering. 1\u20135.","DOI":"10.1145\/3551349.3559558"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSR.2017.59"},{"key":"e_1_3_1_77_2","doi-asserted-by":"crossref","unstructured":"April Yi Wang Will Epperson Robert A DeLine and Steven M Drucker. 2022. Diff in the loop: Supporting data comparison in exploratory data analysis. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. 1\u201310.","DOI":"10.1145\/3491102.3502123"},{"key":"e_1_3_1_78_2","doi-asserted-by":"crossref","unstructured":"Jiawei Wang Tzu-yang Kuo Li Li and Andreas Zeller. 2020. Assessing and restoring reproducibility of Jupyter notebooks. In Proceedings of the 35th IEEE\/ACM International Conference on Automated Software Engineering. 138-\u2013149.","DOI":"10.1145\/3324884.3416585"},{"key":"e_1_3_1_79_2","doi-asserted-by":"crossref","unstructured":"Jiawei Wang Li Li and Andreas Zeller. 2020. Better code better sharing: on the need of analyzing jupyter notebooks. In Proceedings of the ACM\/IEEE 42nd international conference on software engineering: new ideas and emerging results. 53\u201356.","DOI":"10.1145\/3377816.3381724"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445527"},{"key":"e_1_3_1_81_2","doi-asserted-by":"crossref","unstructured":"Ratnadira Widyasari Sheng Qin Sim Camellia Lok Haodi Qi Jack Phan Qijin Tay Constance Tan Fiona Wee Jodie Ethelda Tan Yuheng Yieh et al.. 2020. BugsInPy: A database of existing bugs in Python programs to enable controlled testing and debugging studies. In Proceedings of the 28th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering. 1556\u20131560.","DOI":"10.1145\/3368089.3417943"},{"issue":"1","key":"e_1_3_1_82_2","first-page":"45","article-title":"The art of coding and thematic exploration in qualitative research","volume":"15","author":"Williams Michael","year":"2019","unstructured":"Michael Williams and Tami Moser. 2019. The art of coding and thematic exploration in qualitative research. International management review 15, 1 (2019), 45\u201355.","journal-title":"International management review"},{"key":"e_1_3_1_83_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00129"},{"key":"e_1_3_1_84_2","doi-asserted-by":"crossref","unstructured":"Chunqiu Steven Xia and Lingming Zhang. 2022. Less training more repairing please: revisiting automated program repair via zero-shot learning. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 959\u2013971.","DOI":"10.1145\/3540250.3549101"},{"key":"e_1_3_1_85_2","doi-asserted-by":"crossref","DOI":"10.1201\/9781138359444","volume-title":"R markdown: The definitive guide","author":"Xie Yihui","year":"2018","unstructured":"Yihui Xie, Joseph J Allaire, and Garrett Grolemund. 2018. R markdown: The definitive guide. CRC Press."},{"key":"e_1_3_1_86_2","doi-asserted-by":"crossref","unstructured":"Chenyang Yang Rachel A Brower-Sinning Grace Lewis and Christian K\u00e4stner. 2022. Data leakage in notebooks: Static detection and better processes. In Proceedings of the 37th IEEE\/ACM International Conference on Automated Software Engineering. 1\u201312.","DOI":"10.1145\/3551349.3556918"},{"key":"e_1_3_1_87_2","doi-asserted-by":"crossref","first-page":"304","DOI":"10.1109\/ASE51524.2021.9678520","volume-title":"2021 36th IEEE\/ACM International Conference on Automated Software Engineering (ASE)","author":"Yang Chenyang","year":"2021","unstructured":"Chenyang Yang, Shurui Zhou, Jin LC Guo, and Christian K\u00e4stner. 2021. Subtle bugs everywhere: Generating documentation for data wrangling code. In 2021 36th IEEE\/ACM International Conference on Automated Software Engineering (ASE). IEEE, 304\u2013316."},{"issue":"2","key":"e_1_3_1_88_2","article-title":"Mining Python fix patterns via analyzing fine-grained source code changes","volume":"27","author":"Yang Yilin","year":"2022","unstructured":"Yilin Yang, Tianxing He, Yang Feng, Shaoying Liu, and Baowen Xu. 2022. Mining Python fix patterns via analyzing fine-grained source code changes. Empirical Software Engineering 27, 2 (2022), 48.","journal-title":"Empirical Software Engineering"},{"key":"e_1_3_1_89_2","doi-asserted-by":"crossref","unstructured":"Carmen Zannier Grigori Melnik and Frank Maurer. 2006. On the success of empirical studies in the international conference on software engineering. In Proceedings of the 28th international conference on Software engineering. 341\u2013350.","DOI":"10.1145\/1134285.1134333"},{"key":"e_1_3_1_90_2","doi-asserted-by":"publisher","DOI":"10.1145\/3392826"},{"key":"e_1_3_1_91_2","doi-asserted-by":"crossref","unstructured":"Yuhao Zhang Yifan Chen Shing-Chi Cheung Yingfei Xiong and Lu Zhang. 2018. An empirical study on TensorFlow program bugs. In Proceedings of the 27th ACM SIGSOFT International Symposium on Software Testing and Analysis. 129\u2013140.","DOI":"10.1145\/3213846.3213866"},{"key":"e_1_3_1_92_2","doi-asserted-by":"crossref","unstructured":"Bo Zhou Iulian Neamtiu and Rajiv Gupta. 2015. A cross-platform analysis of bugs and bug-fixing in open source projects: Desktop vs. android vs. ios. In Proceedings of the 19th International Conference on Evaluation and Assessment in Software Engineering. 1\u201310.","DOI":"10.1145\/2745802.2745808"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729352","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T11:01:59Z","timestamp":1787482919000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3729352"}},"issued":{"date-parts":[[2025,6,19]]},"references-count":91,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2025,6,19]]}},"alternative-id":["10.1145\/3729352"],"URL":"https:\/\/doi.org\/10.1145\/3729352","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2025,6,19]]}},{"indexed":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T11:21:14Z","timestamp":1787484074415,"version":"build-2736575974"},"reference-count":68,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:00:00Z","timestamp":1750291200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2022YFB4501903"],"award-info":[{"award-number":["2022YFB4501903"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["62172429 & 62032024"],"award-info":[{"award-number":["62172429 & 62032024"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,19]]},"abstract":"<jats:p>Floating-point constraint solving is challenging due to the complex representation and non-linear computations. Search-based constraint solving provides an effective method for solving floating-point constraints. In this paper, we propose QSF to improve the efficiency of search-based solving for floating-point constraints. The key idea of QSF is to model the floating-point constraint solving problem as a multi-objective optimization problem. Specifically, QSF considers both the number of unsatisfied constraints and the sum of the violation degrees of unsatisfied constraints as the objectives for search-based optimization. Besides, we propose a new evolutionary algorithm in which the mutation operators are specially designed for floating-point numbers, aiming to solve the multi-objective problem more efficiently. We have implemented QSF and conducted extensive experiments on both the SMT-COMP benchmark and the benchmark from real-world floating-point programs. The results demonstrate that compared to SOTA floating-point solvers, QSF achieved an average speedup of 15.72X under a 60-second timeout and an impressive 87.48X under a 600-second timeout on the first benchmark. Similarly, on the second benchmark, QSF delivered an average speedup of 22.44X and 29.23X, respectively, under the two timeout configurations. Furthermore, QSF has also enhanced the performance of symbolic execution for floating-point programs.<\/jats:p>","DOI":"10.1145\/3715739","type":"journal-article","created":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T11:15:34Z","timestamp":1750331734000},"page":"511-532","source":"Crossref","is-referenced-by-count":1,"title":["QSF: Multi-objective Optimization Based Efficient Solving for Floating-Point Constraints"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6177-9164","authenticated-orcid":false,"given":"Xu","family":"Yang","sequence":"first","affiliation":[{"name":"National University of Defense Technology, College of Computer Science and Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4066-7892","authenticated-orcid":false,"given":"Zhenbang","family":"Chen","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, College of Computer Science and Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8033-7943","authenticated-orcid":false,"given":"Wei","family":"Dong","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, College of Computer Science and Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0637-8744","authenticated-orcid":false,"given":"Ji","family":"Wang","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, College of Computer Science and Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,19]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"1985. IEEE standard for binary floating-point arithmetic - IEEE standard 754-1985. Beuth."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1020281327116"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1007\/S10703-023-00423-0"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-99524-9_24"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","unstructured":"Earl T. Barr Thanh Vo Vu Le and Zhendong Su. 2013. Automatic detection of floating-point exceptions. (2013) 549\u2013560. doi:10.1145\/2429069.2429133","DOI":"10.1145\/2429069.2429133"},{"key":"e_1_3_1_7_2","article-title":"The smt-lib standard: Version 2.0","volume":"13","author":"Barrett Clark","year":"2010","unstructured":"Clark Barrett, Aaron Stump, Cesare Tinelli, et al. 2010. The smt-lib standard: Version 2.0. In Proceedings of the 8th international workshop on satisfiability modulo theories (Edinburgh, UK), Vol. 13. 14.","journal-title":"Proceedings of the 8th international workshop on satisfiability modulo theories (Edinburgh, UK)"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","unstructured":"Clark W. Barrett David L. Dill and Jeremy R. Levitt. 1998. A Decision Procedure for Bit-Vector Arithmetic. In Proceedings of the 35th Conference on Design Automation Moscone center San Francico California USA June 15-19 1998 Basant R. Chawla Randal E. Bryant and Jan M. Rabaey (Eds.). ACM Press 522\u2013527. doi:10.1145\/277044.277186","DOI":"10.1145\/277044.277186"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","unstructured":"Aaron Bembenek Michael Greenberg and Stephen Chong. 2023. From SMT to ASP:Solver-Based Approaches to Solving Datalog Synthesis-as-Rule-Selection Problems. Proc. ACM Program. Lang. 7 POPL (2023) 185\u2013217. doi:10.1145\/3571200","DOI":"10.1145\/3571200"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.23919\/FMCAD.2017.8102235"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1214\/ss\/1177011077"},{"key":"e_1_3_1_12_2","volume-title":"Handbook of Satisfiability. Frontiers in Artificial Intelligence and Applications","author":"Biere Armin","year":"2009","unstructured":"Armin Biere, Marijn Heule, Hans van Maaren, and Toby Walsh (Eds.). 2009. Handbook of Satisfiability. Frontiers in Artificial Intelligence and Applications, Vol. 185. IOS Press."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/FMCAD.2008.ECP.18"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-74113-8"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-17462-0_5"},{"key":"e_1_3_1_16_2","unstructured":"Cristian Cadar Daniel Dunbar and Dawson R. Engler. 2008. KLEE: Unassisted and Automatic Generation of High- Coverage Tests for Complex Systems Programs. In 8th USENIX Symposium on Operating Systems Design and Implementation OSDI 2008 December 8-10 2008 San Diego California USA Proceedings Richard Draves and Robbert van Renesse (Eds.). USENIX Association 209\u2013224. http:\/\/www.usenix.org\/events\/osdi08\/tech\/full_papers\/cadar\/cadar.pdf"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF01442131"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","unstructured":"Tao Chen and Miqing Li. 2024. Adapting Multi-objectivized Software Configuration Tuning. Proc. ACM Softw. Eng. 1 FSE (2024) 539\u2013561. doi:10.1145\/3643751","DOI":"10.1145\/3643751"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-36742-7_7"},{"key":"e_1_3_1_20_2","first-page":"143","volume-title":"The Complexity of Theorem-Proving Procedures","author":"Cook Stephen A.","year":"2023","unstructured":"Stephen A. Cook. 2023. The Complexity of Theorem-Proving Procedures (1 ed.). Association for Computing Machinery, New York, NY, USA, 143\u2013152.","edition":"1"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.4324\/9781315693361"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-78800-3_24"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/4235.996017"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/1276958.1277190"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-00234-2_1"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-662-44874-8_3"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/S1568-4946(02)00021-2"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503221.3508424"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1162\/EVCO.2008.16.3.355"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-41540-6_11"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-63618-0_11"},{"key":"e_1_3_1_32_2","unstructured":"Mark Galassi Jim Davies James Theiler Brian Gough Gerard Jungman Patrick Alken Michael Booth Fabrice Rossi and Rhys Ulerich. 2002. GNU scientific library. Network Theory Limited Godalming."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-38574-2_14"},{"key":"e_1_3_1_34_2","unstructured":"Steven G. Johnson. 2007. The NLopt nonlinear-optimization package. https:\/\/github.com\/stevengj\/nlopt."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10957-006-9101-0"},{"issue":"94720","key":"e_1_3_1_36_2","first-page":"11","article-title":"IEEE standard 754 for binary floating-point arithmetic","volume":"754","author":"Kahan William","year":"1996","unstructured":"William Kahan. 1996. IEEE standard 754 for binary floating-point arithmetic. Lecture Notes on the Status of IEEE 754, 94720-1776 (1996), 11.","journal-title":"Lecture Notes on the Status of IEEE"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICNN.1995.488968"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","unstructured":"Mifa Kim Tomoyuki Hiroyasu Mitsunori Miki and Shinya Watanabe. 2004. SPEA2+: Improving the Performance of the Strength Pareto Evolutionary Algorithm 2. 3242 (2004) 742\u2013751. doi:10.1007\/978-3-540-30217-9_75","DOI":"10.1007\/978-3-540-30217-9_75"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.5120\/ijca2017913370"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-662-50497-0"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-16573-3_11"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2004.1281665"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.proeng.2012.01.172"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3338906.3338921"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1007\/S10479-021-04033-Z"},{"key":"e_1_3_1_46_2","volume-title":"The SMT Workshop","author":"Marre Bruno","year":"2017","unstructured":"Bruno Marre, Fran\u00e7ois Bobot, and Zakaria Chihani. 2017. Real Behavior of Floating Point Numbers. In The SMT Workshop. SMT 2017, 15th International Workshop on Satisfiability Modulo Theories, Heidelberg, Germany. https:\/\/cea.hal.science\/cea-01795760"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","unstructured":"Claude Michel Michel Rueher and Yahia Lebbah. 2001. Solving Constraints over Floating-Point Numbers. In Principles and Practice of Constraint Programming - CP 2001 7th International Conference CP 2001 Paphos Cyprus November 26 - December 1 2001 Proceedings (Lecture Notes in Computer Science Vol. 2239) Toby Walsh (Ed.). Springer 524\u2013538. doi:10.1007\/3-540-45578-7_36","DOI":"10.1007\/3-540-45578-7_36"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3611643.3616357"},{"key":"e_1_3_1_49_2","doi-asserted-by":"crossref","unstructured":"Aina Niemetz and Mathias Preiner. 2023. Bitwuzla. In Computer Aided Verification - 35th International Conference CAV 2023 Paris France July 17-22 2023 Proceedings Part II (Lecture Notes in Computer Science Vol. 13965) Constantin Enea and Akash Lal (Eds.). Springer 3\u201317.","DOI":"10.1007\/978-3-031-37703-7_1"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-20398-5_22"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1162\/evco.1998.6.3.231"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1007\/S10515-014-0154-2"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1007\/S10703-017-0270-2"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4757-4145-2_5"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-54013-4_19"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","unstructured":"Ziqi Shuai Zhenbang Chen Kelin Ma Kunlin Liu Yufeng Zhang Jun Sun and Ji Wang. 2024. Partial Solution Based Constraint Solving Cache in Symbolic Execution. Proc. ACM Softw. Eng. 1 FSE (2024) 2493\u20132514. doi:10.1145\/3660817","DOI":"10.1145\/3660817"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-73190-0_2"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-20398-5_26"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3457784.3457823"},{"key":"e_1_3_1_60_2","unstructured":"Gilbert Syswerda. 1989. Uniform Crossover in Genetic Algorithms. In Proceedings of the 3rd International Conference on Genetic Algorithms George Mason University Fairfax Virginia USA June 1989 J. David Schaffer (Ed.). Morgan Kaufmann 2\u20139."},{"key":"e_1_3_1_61_2","doi-asserted-by":"crossref","first-page":"134","DOI":"10.1007\/978-3-540-79124-9_10","volume-title":"Tests and Proofs,","author":"Tillmann Nikolai","year":"2008","unstructured":"Nikolai Tillmann and Jonathan de Halleux. 2008. Pex\u2013White Box Test Generation for .NET. In Tests and Proofs, Bernhard Beckert and Reiner H\u00e4hnle (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 134\u2013153."},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1002\/MALQ.200610007"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/TEVC.2023.3252612"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","unstructured":"Xu Yang Guofeng Zhang Ziqi Shuai Zhenbang Chen and Ji Wang. 2025. Symbolic execution of floating-point programs: How far are we? J. Syst. Softw. 220 112242. doi:10.1016\/J.JSS.2024.112242","DOI":"10.1016\/J.JSS.2024.112242"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/TEVC.2021.3118593"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/TEVC.2007.892759"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416645"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","unstructured":"Shasha Zhou Mingyu Huang Yanan Sun and Ke Li. 2024. Evolutionary Multi-objective Optimization for Contextual Adversarial Example Generation. Proc. ACM Softw. Eng. 1 FSE (2024) 2285\u20132308. doi:10.1145\/3660808","DOI":"10.1145\/3660808"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-66158-2_45"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715739","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T11:05:02Z","timestamp":1787483102000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715739"}},"issued":{"date-parts":[[2025,6,19]]},"references-count":68,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2025,6,19]]}},"alternative-id":["10.1145\/3715739"],"URL":"https:\/\/doi.org\/10.1145\/3715739","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2025,6,19]]}},{"indexed":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T16:00:16Z","timestamp":1787500816645,"version":"build-2736575974"},"reference-count":64,"publisher":"Association for Computing Machinery (ACM)","issue":"ISSTA","license":[{"start":{"date-parts":[[2025,6,22]],"date-time":"2025-06-22T00:00:00Z","timestamp":1750550400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,22]]},"abstract":"<jats:p>\n                    The Notification Listener Service (NLS) in Android allows third-party apps to monitor and process device notifications, enabling powerful features but also introducing security and privacy risks. Despite the special permission required to access NLS, it has been recurrently exploited by malicious actors. However, there is a lack of systematic investigation into NLS usage patterns and their security implications. In this paper, we propose\n                    <jats:sc>NLRadar<\/jats:sc>\n                    , a hybrid approach combining static analysis and LLM to examine NLS usage in Android apps. We apply\n                    <jats:sc>NLRadar<\/jats:sc>\n                    to a large scale of apps, including both malware and regular apps, to demystify NLS usage and to uncover abuses. Our analysis reveals that NLS is heavily abused, with interesting discoveries such as apps insecurely storing social media messages, exploiting NLS for destructive competition or SMS credential stealing, and leveraging NLS to spread promotional messages or even malicious links. We also find undisclosed changes in NLS usage through app updates and inadequate disclosure in privacy policies. Our findings emphasize the need for more rigorous vetting of NLS usage and better developer education on responsible NLS practices.\n                  <\/jats:p>","DOI":"10.1145\/3728898","type":"journal-article","created":{"date-parts":[[2025,6,22]],"date-time":"2025-06-22T10:52:56Z","timestamp":1750589576000},"page":"434-456","source":"Crossref","is-referenced-by-count":2,"title":["Walls Have Ears: Demystifying Notification Listener Usage in Android Apps"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-1876-9285","authenticated-orcid":false,"given":"Jiapeng","family":"Deng","sequence":"first","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5216-933X","authenticated-orcid":false,"given":"Tianming","family":"Liu","sequence":"additional","affiliation":[{"name":"Monash University, Melbourne, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8793-5367","authenticated-orcid":false,"given":"Yanjie","family":"Zhao","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8117-0352","authenticated-orcid":false,"given":"Chao","family":"Wang","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-6642-2238","authenticated-orcid":false,"given":"Lin","family":"Zhang","sequence":"additional","affiliation":[{"name":"The National Computer Emergency Response Team\/Coordination Center of China (CNCERT\/CC), Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1100-8633","authenticated-orcid":false,"given":"Haoyu","family":"Wang","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,22]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2019. Expanding target API level requirements in 2019. https:\/\/android-developers.googleblog.com\/2019\/02\/expanding-target-api-level-requirements.html."},{"key":"e_1_3_1_3_2","unstructured":"2019. Malware sidesteps Google policy with new 2FA bypass technique. https:\/\/www.welivesecurity.com\/2019\/06\/17\/malware-google-permissions-2fa-bypass\/."},{"key":"e_1_3_1_4_2","unstructured":"2021. Beware \u2014 A New Wormable Android Malware Spreading Through WhatsApp. https:\/\/thehackernews.com\/2021\/01\/beware-new-wormable-android-malware.html."},{"key":"e_1_3_1_5_2","unstructured":"2022. Android target API level requirements. https:\/\/seller.samsungapps.com\/notice\/getNoticeDetail.as?csNoticeID=0000007234."},{"key":"e_1_3_1_6_2","unstructured":"2023. Alien Android Banking Trojan Sidesteps 2FA | Threatpost. https:\/\/threatpost.com\/alien-android-2fa\/159517\/."},{"key":"e_1_3_1_7_2","unstructured":"2023. Promiscuous Permissions: Catching Your Android Apps in the Act | Keysight Blogs. https:\/\/www.keysight.com\/blogs\/en\/tech\/nwvs\/2023\/03\/24\/promiscuous-permissions-catching-your-android-apps-in-the-act."},{"key":"e_1_3_1_8_2","unstructured":"2023. Scoped Storage. https:\/\/source.android.com\/docs\/core\/storage\/scoped."},{"key":"e_1_3_1_9_2","unstructured":"2023. Sneaky DogeRAT Trojan Poses as Popular Apps Targets Indian Android Users. https:\/\/thehackernews.com\/2023\/05\/sneaky-dogerat-trojan-poses-as-popular.html."},{"key":"e_1_3_1_10_2","unstructured":"2023. Target API level requirements for Google Play apps. https:\/\/support.google.com\/googleplay\/android-developer\/answer\/11926878."},{"key":"e_1_3_1_11_2","unstructured":"2023. Technical analysis of SOVA android malware. https:\/\/muha2xmad.github.io\/malware-analysis\/sova\/."},{"key":"e_1_3_1_12_2","unstructured":"2024. BIND_NOTIFICATION_LISTENER_SERVICE | Manifest.permission | Android Developers. https:\/\/developer.android.com\/reference\/android\/Manifest.permission#BIND_NOTIFICATION_LISTENER_SERVICE."},{"key":"e_1_3_1_13_2","unstructured":"2024. Gift Offer Results APK (Android App) - Free Download. https:\/\/apkcombo.com\/gift-offer-results\/com.magis.app\/."},{"key":"e_1_3_1_14_2","unstructured":"2024. GPT-4o | OpenAI. https:\/\/openai.com\/index\/hello-gpt-4o\/."},{"key":"e_1_3_1_15_2","unstructured":"2024. Jelly Bean | Android Developers. https:\/\/developer.android.com\/about\/versions\/jelly-bean."},{"key":"e_1_3_1_16_2","unstructured":"2024. Market Distribution of the Regular App Dataset | NLRadar. https:\/\/github.com\/security-pride\/NLRadar\/tree\/master\/ApkInfo."},{"key":"e_1_3_1_17_2","unstructured":"2024. NotificationListenerService | Android Developers. https:\/\/developer.android.com\/reference\/android\/service\/notification\/NotificationListenerService."},{"key":"e_1_3_1_18_2","unstructured":"2024. Notifications overview | Android Developers. https:\/\/developer.android.com\/develop\/ui\/views\/notifications."},{"key":"e_1_3_1_19_2","unstructured":"2024. Prompting Questions for Assessing NLS Usage Security and Chain-of-thought Reasoning Example | NLRadar. https:\/\/github.com\/security-pride\/NLRadar\/tree\/master\/NLRadar\/LLM_Evaluation."},{"key":"e_1_3_1_20_2","unstructured":"2024. Save data in a local database using Room - Android Developers. https:\/\/developer.android.com\/training\/data-storage\/room."},{"key":"e_1_3_1_21_2","unstructured":"2024. SharedPreferences | Android Developers. https:\/\/developer.android.com\/training\/data-storage\/shared-preferences."},{"key":"e_1_3_1_22_2","unstructured":"2024. StatusBarNotification | Android Developers. https:\/\/developer.android.com\/reference\/android\/service\/notification\/StatusBarNotification."},{"key":"e_1_3_1_23_2","unstructured":"2024. Using Binder IPC | Android Open Source Project. https:\/\/source.android.com\/docs\/core\/architecture\/hidl\/binderipc."},{"key":"e_1_3_1_24_2","unstructured":"2025. Code Snippets of Identified NLS Abuse Examples | NLRadar. https:\/\/github.com\/security-pride\/NLRadar\/tree\/master\/NLRadar\/LLM_Evaluation\/Abuse_Example."},{"key":"e_1_3_1_25_2","unstructured":"2025. deepseek-ai\/DeepSeek-R1. https:\/\/github.com\/deepseek-ai\/DeepSeek-R1."},{"key":"e_1_3_1_26_2","unstructured":"2025. Fiddler B. https:\/\/www.telerik.com\/fiddler-b."},{"key":"e_1_3_1_27_2","unstructured":"2025. GitHub - skylot\/jadx: Dex to Java decompiler. https:\/\/github.com\/skylot\/jadx."},{"key":"e_1_3_1_28_2","unstructured":"2025. Intent | API reference | Android Developers. https:\/\/developer.android.com\/reference\/android\/content\/Intent."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","unstructured":"Huda Abualola Hessa Alhawai Maha Kadadha Hadi Otrok and Azzam Mourad. 2016. An Android-based Trojan Spyware to study the notificationlistener service vulnerability. Procedia Computer Science 83 (2016) 465\u2013471. doi:10.1016\/J.PROCS.2016.04.210","DOI":"10.1016\/J.PROCS.2016.04.210"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","unstructured":"Mansour Ahmadi Battista Biggio Steven Arzt Davide Ariu and Giorgio Giacinto. 2016. Detecting misuse of google cloud messaging in android badware. In Proceedings of the 6th Workshop on Security and Privacy in Smartphones and Mobile Devices. 103\u2013112. doi:10.1145\/2994459.2994469","DOI":"10.1145\/2994459.2994469"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","unstructured":"Kevin Allix Tegawend\u00e9 F Bissyand\u00e9 Jacques Klein and Yves Le Traon. 2016. Androzoo: Collecting millions of android apps for the research community. In Proceedings of the 13th international conference on mining software repositories. 468\u2013471. doi:10.1145\/2901739.2903508","DOI":"10.1145\/2901739.2903508"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","unstructured":"Steven Arzt Siegfried Rasthofer Christian Fritz Eric Bodden Alexandre Bartel Jacques Klein Yves Le Traon Damien Octeau and Patrick McDaniel. 2014. Flowdroid: Precise context flow field object-sensitive and lifecycle-aware taint analysis for android apps. ACM sigplan notices 49 6 (2014) 259\u2013269. doi:10.1145\/2666356.2594299","DOI":"10.1145\/2666356.2594299"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","unstructured":"Yangyi Chen Tongxin Li XiaoFeng Wang Kai Chen and Xinhui Han. 2015. Perplexed messengers from the cloud: Automated security analysis of push-messaging integrations. In Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security . 1260\u20131272. doi:10.1145\/2810103.2813652","DOI":"10.1145\/2810103.2813652"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","unstructured":"Yongliang Chen Ruoqin Tang Chaoshun Zuo Xiaokuan Zhang Lei Xue Xiapu Luo and Qingchuan Zhao. 2024. Attention! Your Copied Data is Under Monitoring: A Systematic Study of Clipboard Usage in Android Apps. In Proceedings of the 46th IEEE\/ACM International Conference on Software Engineering. 1\u201313. doi:10.1145\/3597503.3623317","DOI":"10.1145\/3597503.3623317"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","unstructured":"Kyle Denney A Selcuk Uluagac Kemal Akkaya and Shekhar Bhansali. 2016. A novel storage covert channel on wearable devices using status bar notifications. In 2016 13th IEEE Annual Consumer Communications & Networking Conference (CCNC). IEEE 845\u2013848. doi:10.1109\/CCNC.2016.7444898","DOI":"10.1109\/CCNC.2016.7444898"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","unstructured":"Kyle Denney A Selcuk Uluagac Hidayet Aksu and Kemal Akkaya. 2018. An Android-Based Covert Channel Framework on Wearables Using Status Bar Notifications. Versatile Cybersecurity (2018) 1\u201317. doi:10.1007\/978-3-319-97643-3_1","DOI":"10.1007\/978-3-319-97643-3_1"},{"key":"e_1_3_1_37_2","unstructured":"Wenrui Diao Yue Zhang Li Zhang Zhou Li Fenghao Xu Xiaorui Pan Xiangyu Liu Jian Weng Kehuan Zhang and XiaoFeng Wang. 2019. Kindness is a Risky Business: On the Usage of the Accessibility {APIs} in Android. In 22nd International Symposium on Research in Attacks Intrusions and Defenses (RAID 2019). 261\u2013275."},{"key":"e_1_3_1_38_2","unstructured":"Zikan Dong Tianming Liu Jiapeng Deng Li Li Minghui Yang Meng Wang Guosheng Xu and Guoai Xu. 2024. Exploring Covert Third-party Identifiers through External Storage in the Android New Era. In 33rd USENIX Security Symposium (USENIX Security 24). 4535\u20134552."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","unstructured":"Xinyi Hou Yanjie Zhao Yue Liu Zhou Yang Kailong Wang Li Li Xiapu Luo David Lo John Grundy and Haoyu Wang. 2024. Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology 33 8 (2024) 1\u201379. doi:10.1145\/3695988","DOI":"10.1145\/3695988"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","unstructured":"Sangwon Hyun Junsung Cho Geumhwan Cho and Hyoungshick Kim. 2018. Design and Analysis of Push Notification-Based Malware on Android. Security and Communication Networks 2018 1 (2018) 8510256. doi:10.1155\/2018\/8510256","DOI":"10.1155\/2018\/8510256"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","unstructured":"Yeongjin Jang Chengyu Song Simon P Chung Tielei Wang and Wenke Lee. 2014. A11y attacks: Exploiting accessibility in operating systems. In Proceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security. 103\u2013115. doi:10.1145\/2660267.2660295","DOI":"10.1145\/2660267.2660295"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","unstructured":"Hayoung Lee Taeho Kang Sangho Lee Jong Kim and Yoonho Kim. 2014. Punobot: Mobile botnet using push notification service in android. In Information Security Applications: 14th International Workshop WISA 2013 Jeju Island Korea August 19-21 2013 Revised Selected Papers 14. Springer 124\u2013137. doi:10.1007\/978-3-319-05149-9_8","DOI":"10.1007\/978-3-319-05149-9_8"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","unstructured":"Li Li Alexandre Bartel Tegawend\u00e9 F Bissyand\u00e9 Jacques Klein Yves Le Traon Steven Arzt Siegfried Rasthofer Eric Bodden Damien Octeau and Patrick McDaniel. 2015. Iccta: Detecting inter-component privacy leaks in android apps. In 2015 IEEE\/ACM 37th IEEE International Conference on Software Engineering Vol. 1. IEEE 280\u2013291. doi:10.1109\/ICSE.2015.48","DOI":"10.1109\/ICSE.2015.48"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","unstructured":"Li Li Tegawend\u00e9 F Bissyand\u00e9 Damien Octeau and Jacques Klein. 2016. Reflection-Aware Static Analysis of Android Apps. In The 31st IEEE\/ACM International Conference on Automated Software Engineering Demo Track (ASE 2016). doi:10.1145\/2970276.2970277","DOI":"10.1145\/2970276.2970277"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","unstructured":"Tongxin Li Xiaoyong Zhou Luyi Xing Yeonjoon Lee Muhammad Naveed XiaoFeng Wang and Xinhui Han. 2014. Mayhem in the push clouds: Understanding and mitigating security hazards in mobile push-messaging services. In Proceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security. 978\u2013989. doi:10.1145\/2660267.2660302","DOI":"10.1145\/2660267.2660302"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","unstructured":"Keke Lian Lei Zhang Guangliang Yang Shuo Mao Xinjie Wang Yuan Zhang and Min Yang. 2024. Component Security Ten Years Later: An Empirical Study of Cross-Layer Threats in Real-World Mobile Applications. Proceedings of the ACM on Software Engineering 1 FSE (2024) 70\u201391. doi:10.1145\/3643730","DOI":"10.1145\/3643730"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","unstructured":"Tianming Liu Haoyu Wang Li Li Guangdong Bai Yao Guo and Guoai Xu. 2019. Dapanda: Detecting aggressive push notifications in android apps. In 2019 34th IEEE\/ACM International Conference on Automated Software Engineering (ASE). IEEE 66\u201378. doi:10.1109\/ASE.2019.00017","DOI":"10.1109\/ASE.2019.00017"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","unstructured":"Jiadong Lou Xiaohan Zhang Yihe Zhang Xinghua Li Xu Yuan and Ning Zhang. 2023. Devils in Your Apps: Vulnerabilities and User Privacy Exposure in Mobile Notification Systems. In 2023 53rd Annual IEEE\/IFIP International Conference on Dependable Systems and Networks (DSN). IEEE 28\u201341. doi:10.1109\/DSN58367.2023.00017","DOI":"10.1109\/DSN58367.2023.00017"},{"key":"e_1_3_1_49_2","unstructured":"Allan Lyons Julien Gamba Austin Shawaga Joel Reardon Juan Tapiador Serge Egelman and Narseo Vallina-Rodr\u00edguez. 2023. Log:{It\u2019s} Big {It\u2019s} Heavy {It\u2019s} Filled with Personal Data! Measuring the Logging of Sensitive Information in the Android Ecosystem. In 32nd USENIX Security Symposium (USENIX Security 23). 2115\u20132132."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","unstructured":"Daye Nam Andrew Macvean Vincent Hellendoorn Bogdan Vasilescu and Brad Myers. 2024. Using an llm to help with code understanding. In Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. 1\u201313. doi:10.1145\/3597503.3639187","DOI":"10.1145\/3597503.3639187"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","unstructured":"Mohammad Naseri Nataniel P Borges Jr Andreas Zeller and Romain Rouvoy. 2019. Accessileaks: Investigating privacy leaks exposed by the android accessibility service. In PETS 2019-The 19th Privacy Enhancing Technologies Symposium. doi:10.2478\/popets-2019-0031","DOI":"10.2478\/popets-2019-0031"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","unstructured":"Thomas Neteler Sascha Fahl and Luigi Lo Iacono. 2024. \u201cYou received $100 000 from Johnny\u201d: A Mixed-Methods Study on Push Notification Security and Privacy in Android Apps. IEEE Access (2024). doi:10.1109\/ACCESS.2024.3439095","DOI":"10.1109\/ACCESS.2024.3439095"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","unstructured":"Siegfried Rasthofer Steven Arzt and Eric Bodden. 2014. A machine-learning approach for classifying and categorizing android sources and sinks. In NDSS Vol. 14. 1125. doi:10.14722\/ndss.2014.23039","DOI":"10.14722\/ndss.2014.23039"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","unstructured":"Nikita Samarin Alex Sanchez Trinity Chung Akshay Dan Bhavish Juleemun Conor Gilsenan Nick Merrill Joel Reardon and Serge Egelman. [n. d.]. The Medium is the Message: How Secure Messaging Apps Leak Sensitive Data to Push Notification Services. Proceedings on Privacy Enhancing Technologies 2024 4 ([n. d.]). doi:10.56553\/popets-2024-0151","DOI":"10.56553\/popets-2024-0151"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","unstructured":"Xiaoyu Sun Li Li Tegawend\u00e9 F Bissyand\u00e9 Jacques Klein Damien Octeau and John Grundy. 2020. Taming Reflection: An Essential Step Towards Whole-Program Analysis of Android Apps. ACM Transactions on Software Engineering and Methodology (TOSEM) (2020). doi:10.1145\/3440033","DOI":"10.1145\/3440033"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","unstructured":"Xiaoyu Tan Yongxin Deng Xihe Qiu Weidi Xu Chao Qu Wei Chu Yinghui Xu and Yuan Qi. 2024. Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Driven Prolog-based Chain-of-Though. arXiv preprint arXiv:2407.14562 (2024). doi:10.48550\/arXiv.2407.14562","DOI":"10.48550\/arXiv.2407.14562"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","unstructured":"Raja Vall\u00e9e-Rai Phong Co Etienne Gagnon Laurie Hendren Patrick Lam and Vijay Sundaresan. 2010. Soot: A Java bytecode optimization framework. In CASCON First Decade High Impact Papers. 214\u2013224. doi:10.1145\/1925805.1925818","DOI":"10.1145\/1925805.1925818"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","unstructured":"Liu Wang Haoyu Wang Ren He Ran Tao Guozhu Meng Xiapu Luo and Xuanzhe Liu. 2022. MalRadar: Demystifying android malware in the new era. Proceedings of the ACM on Measurement and Analysis of Computing Systems 6 2 (2022) 1\u201327. doi:10.1145\/3530906","DOI":"10.1145\/3530906"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","unstructured":"Yaqing Wang Quanming Yao James T Kwok and Lionel M Ni. 2020. Generalizing from a few examples: A survey on few-shot learning. ACM computing surveys (csur) 53 3 (2020) 1\u201334. doi:10.1145\/3386252","DOI":"10.1145\/3386252"},{"key":"e_1_3_1_60_2","doi-asserted-by":"crossref","unstructured":"Jason Wei Xuezhi Wang Dale Schuurmans Maarten Bosma Fei Xia Ed Chi Quoc V Le Denny Zhou et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022) 24824\u201324837.","DOI":"10.52202\/068431-1800"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","unstructured":"Lili Wei Yepang Liu and Shing-Chi Cheung. 2016. Taming android fragmentation: Characterizing and detecting compatibility issues for android apps. In Proceedings of the 31st IEEE\/ACM international conference on automated software engineering. 226\u2013237. doi:10.1145\/2970276.2970312","DOI":"10.1145\/2970276.2970312"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","unstructured":"HanXiang Xu ShenAo Wang Ningke Li Yanjie Zhao Kai Chen Kailong Wang Yang Liu Ting Yu and HaoYu Wang. 2024. Large language models for cyber security: A systematic literature review. arXiv preprint arXiv:2405.04760 (2024). doi:10.48550\/arXiv.2405.04760","DOI":"10.48550\/arXiv.2405.04760"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","unstructured":"Jiwei Yan Shixin Zhang Yepang Liu Jun Yan and Jian Zhang. 2022. Iccbot: fragment-aware and context-sensitive icc resolution for android applications. In Proceedings of the ACM\/IEEE 44th International Conference on Software Engineering: Companion Proceedings. 105\u2013109. doi:10.1145\/3510454.3516864","DOI":"10.1145\/3510454.3516864"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","unstructured":"Qingchuan Zhao Chaoshun Zuo Brendan Dolan-Gavitt Giancarlo Pellegrino and Zhiqiang Lin. 2020. Automatic uncovering of hidden behaviors from input validation in mobile apps. In 2020 IEEE Symposium on Security and Privacy (SP). IEEE 1106\u20131120. doi:10.1109\/SP40000.2020.00072","DOI":"10.1109\/SP40000.2020.00072"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","unstructured":"Hao Zhou Xiapu Luo Haoyu Wang and Haipeng Cai. 2022. Uncovering Intent based Leak of Sensitive Data in Android Framework. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security. 3239\u20133252. doi:10.1145\/3548606.3560601","DOI":"10.1145\/3548606.3560601"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728898","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T13:49:30Z","timestamp":1787492970000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728898"}},"issued":{"date-parts":[[2025,6,22]]},"references-count":64,"journal-issue":{"issue":"ISSTA","published-print":{"date-parts":[[2025,6,22]]}},"alternative-id":["10.1145\/3728898"],"URL":"https:\/\/doi.org\/10.1145\/3728898","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2025,6,22]]}},{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T19:51:31Z","timestamp":1782849091909,"version":"3.54.5"},"reference-count":72,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2023YFB4503802"],"award-info":[{"award-number":["2023YFB4503802"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62302515"],"award-info":[{"award-number":["62302515"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62172426"],"award-info":[{"award-number":["62172426"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62332005"],"award-info":[{"award-number":["62332005"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Code adaptation is a fundamental but challenging task in software development, requiring developers to modify existing code for new contexts. A key challenge is to resolve\n                    <jats:bold>Context Adaptation Bugs (CtxBugs)<\/jats:bold>\n                    , which occurs when code correct in its original context violates constraints in the target environment. Unlike isolated bugs,\n                    <jats:italic toggle=\"yes\">CtxBugs<\/jats:italic>\n                    cannot be resolved through local fixes and require cross-context reasoning to identify semantic mismatches. Overlooking them may lead to critical failures in adaptation. Although Large Language Models (LLMs) show great potential in automating code-related tasks, their ability to resolve\n                    <jats:italic toggle=\"yes\">CtxBugs<\/jats:italic>\n                    remains a significant and unexplored obstacle to their practical use in code adaptation.\n                  <\/jats:p>\n                  <jats:p>\n                    To bridge this gap, we propose\n                    <jats:italic toggle=\"yes\">CtxBugGen<\/jats:italic>\n                    , a novel framework for generating\n                    <jats:italic toggle=\"yes\">CtxBugs<\/jats:italic>\n                    to evaluate LLMs. Its core idea is to leverage LLMs\u2019 tendency to generate plausible but context-free code when contextual constraints are absent. The framework generates\n                    <jats:italic toggle=\"yes\">CtxBugs<\/jats:italic>\n                    through a four-step process to ensure their relevance and validity: (1) Selection of four established context-aware adaptation tasks from the literature, (2) Perturbation via task-specific rules to induce\n                    <jats:italic toggle=\"yes\">CtxBugs<\/jats:italic>\n                    from LLMs while ensuring their plausibility, (3) Generation of candidate variants by prompting LLMs without any context constraints and (4) Identification of valid\n                    <jats:italic toggle=\"yes\">CtxBugs<\/jats:italic>\n                    through syntactic differencing and test execution in the target context. Based on the benchmark constructed by\n                    <jats:italic toggle=\"yes\">CtxBugGen<\/jats:italic>\n                    , we conduct an empirical study with four state-of-the-art LLMs. Our results reveal their unsatisfactory performance in\n                    <jats:italic toggle=\"yes\">CtxBug<\/jats:italic>\n                    resolution. The best performing LLM, Kimi-K2, achieves 55.93% on Pass@1 and resolves just 52.47% of\n                    <jats:italic toggle=\"yes\">CtxBugs<\/jats:italic>\n                    . The presence of\n                    <jats:italic toggle=\"yes\">CtxBugs<\/jats:italic>\n                    degrades LLMs\u2019 adaptation performance by up to 30%. Failure analysis indicates that LLMs often overlook\n                    <jats:italic toggle=\"yes\">CtxBugs<\/jats:italic>\n                    and replicate them in their outputs. This suggests that LLMs overly focus on the local code correctness of the reused code while ignoring its compatibility in the target context. Our study highlights a critical weakness in LLMs\u2019 cross-context reasoning and emphasize the need for new methods to enhance their context awareness for reliable code adaptation. The replication package for this paper is at https:\/\/github.com\/ztwater\/CtxBugGen.\n                  <\/jats:p>","DOI":"10.1145\/3797148","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"229-252","source":"Crossref","is-referenced-by-count":0,"title":["Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs during Code Adaptation"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7241-9730","authenticated-orcid":false,"given":"Tanghaoran","family":"Zhang","sequence":"first","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6003-5748","authenticated-orcid":false,"given":"Xinjun","family":"Mao","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1469-2063","authenticated-orcid":false,"given":"Shangwen","family":"Wang","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0061-9457","authenticated-orcid":false,"given":"Yuxin","family":"Zhao","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3520-5829","authenticated-orcid":false,"given":"Yao","family":"Lu","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6902-3270","authenticated-orcid":false,"given":"Zezhou","family":"Tang","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2088-4976","authenticated-orcid":false,"given":"Wenyu","family":"Xu","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8686-818X","authenticated-orcid":false,"given":"Longfei","family":"Sun","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-6739-7100","authenticated-orcid":false,"given":"Changrong","family":"Xie","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2313-7141","authenticated-orcid":false,"given":"Kang","family":"Yang","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9865-2212","authenticated-orcid":false,"given":"Yue","family":"Yu","sequence":"additional","affiliation":[{"name":"Peng Cheng Laboratory, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Miltiadis Allamanis and Marc Brockschmidt. 2017. SmartPaste: Learning to Adapt Source Code. http:\/\/arxiv.org\/abs\/ 1705.07867 arXiv:1705.07867 [cs]."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/SANER.2017.7884629"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2591062.2591130"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3586030"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/MS.2009.147"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1191\/1478088706qp063oa"},{"key":"e_1_2_1_7_1","first-page":"1877","volume-title":"Lin (Eds.)","volume":"33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 1877-1901."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE55347.2025.00012"},{"key":"e_1_2_1_9_1","volume-title":"Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al.","author":"Chen Mark","year":"2021","unstructured":"Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating Large Language Models Trained on Code. (2021). arXiv:2107.03374 [cs.LG]"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1177\/001316446002000104"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPC66645.2025.00060"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1370175.1370194"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1453101.1453130"},{"key":"e_1_2_1_14_1","unstructured":"DeepSeek-AI. 2024. DeepSeek-V3 Technical Report. arXiv:2412.19437 [cs.CL] https:\/\/arxiv.org\/abs\/2412.19437"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2688204.2688208"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639219"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2642937.2642982"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSR.2019.00039"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/sp.2017.31"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","unstructured":"Shuzheng Gao Cuiyun Gao Wenchao Gu and Michael Lyu. 2024. Search-Based LLMs for Code Optimization. doi:10.48550\/arXiv.2408.12159 arXiv:2408.12159 [cs]. 10.48550\/arXiv.2408.12159","DOI":"10.48550\/arXiv.2408.12159"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSR.2017.15"},{"key":"e_1_2_1_22_1","volume-title":"CodeEditor","author":"Guo Jiawei","unstructured":"Jiawei Guo, Ziming Li, Xueling Liu, Kaijing Ma, Tianyu Zheng, Zhouliang Yu, Ding Pan, Yizhi LI, Ruibo Liu, Yue Wang, Shuyue Guo, Xingwei Qu, Xiang Yue, Ge Zhang, Wenhu Chen, and Jie Fu. 2024. CodeEditorBench: Evaluating Code Editing Capability of Large Language Models. arXiv:2404.03543 [cs.SE] https:\/\/arxiv.org\/abs\/2404.03543"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1002\/9781119196037"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2377656.2377657"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-Companion58688.2023.00013"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3524610.3527923"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/1328279.1328283"},{"key":"e_1_2_1_28_1","unstructured":"Barbara Kitchenham and Stuart Charters. 2007. Guidelines for performing Systematic Literature Reviews in Software Engineering. Technical Report. Keele University."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3729390"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549081"},{"key":"e_1_2_1_31_1","volume-title":"Ahmed Musa Awon, Daniela Damian, and Bowen Xu.","author":"Li Ze Shi","year":"2024","unstructured":"Ze Shi Li, Nowshin Nawar Arony, Ahmed Musa Awon, Daniela Damian, and Bowen Xu. 2024. AI Tool Use and Adoption in Software Development by Individuals and Organizations: A Grounded Theory Study. arXiv:2406.17325 [cs.SE] https:\/\/arxiv.org\/abs\/2406.17325"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3611643.3616339"},{"key":"e_1_2_1_33_1","unstructured":"Pengfei Liu Weizhe Yuan Jinlan Fu Zhengbao Jiang Hiroaki Hayashi and Graham Neubig. 2021. Pre-train Prompt and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. http:\/\/arxiv.org\/abs\/2107.13586 arXiv:2107.13586 [cs]."},{"key":"e_1_2_1_34_1","unstructured":"Xiaoyu Liu Jinu Jang Neel Sundaresan Miltiadis Allamanis and Alexey Svyatkovskiy. 2022. AdaptivePaste: Code Adaptation through Learning Semantics-aware Variable Usage Representations. http:\/\/arxiv.org\/abs\/2205.11023 arXiv:2205.11023 [cs]."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASWEC.2018.00027"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3660811"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSME.2019.00026"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/1932682.1869486"},{"key":"e_1_2_1_39_1","unstructured":"OpenAI. 2024. gpt-4o. https:\/\/platform.openai.com\/docs\/models\/gpt-4o"},{"key":"e_1_2_1_40_1","unstructured":"Shuyin Ouyang Jie M. Zhang Mark Harman and Meng Wang. 2023. LLM is Like a Box of Chocolates: the Nondeterminism of ChatGPT in Code Generation. arXiv:2308.02828 [cs.SE] https:\/\/arxiv.org\/abs\/2308.02828"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2019"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2023.3248113"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377929.3398087"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510216"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPC52881.2021.00017"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2043174.2043193"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE55347.2025.00040"},{"key":"e_1_2_1_49_1","unstructured":"Kimi Team. 2025. Kimi K2: Open Agentic Intelligence. arXiv:2507.20534 [cs.LG] https:\/\/arxiv.org\/abs\/2507.20534"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2931037.2931058"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE51524.2021.9678576"},{"key":"e_1_2_1_52_1","doi-asserted-by":"crossref","unstructured":"Runchu Tian Yining Ye Yujia Qin Xin Cong Yankai Lin Yinxu Pan Yesai Wu Haotian Hui Weichuan Liu Zhiyuan Liu and Maosong Sun. 2024. DebugBench: Evaluating Debugging Capability of Large Language Models. arXiv:2401.04621 [cs.SE] https:\/\/arxiv.org\/abs\/2401.04621","DOI":"10.18653\/v1\/2024.findings-acl.247"},{"key":"e_1_2_1_53_1","unstructured":"Marko Vasic Aditya Kanade Petros Maniatis David Bieber and Rishabh Singh. 2019. Neural Program Repair by Jointly Learning to Localize and Repair. arXiv:1904.01720 [cs.LG] https:\/\/arxiv.org\/abs\/1904.01720"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","unstructured":"Taiming Wang Yanjie Jiang Chunhao Dong Yuxia Zhang and Hui Liu. 2025. Context-Aware Code Wiring Recommendation with LLM-based Agent. doi:10.48550\/arXiv.2507.01315 arXiv:2507.01315 [cs]. 10.48550\/arXiv.2507.01315","DOI":"10.48550\/arXiv.2507.01315"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/2950290.2983934"},{"key":"e_1_2_1_56_1","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems","author":"Wei Jason","year":"2024","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2024. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the 36th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS '22). Curran Associates Inc., Red Hook, NY, USA, Article 1800, 14 pages."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2380116.2380145"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-018-9634-5"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3540250"},{"key":"e_1_2_1_60_1","unstructured":"Chunqiu Steven Xia and Lingming Zhang. 2023. Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT. http:\/\/arxiv.org\/abs\/2304.00385 arXiv:2304.00385 [cs]."},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2404.08877"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/2901739.2901767"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSR.2017.13"},{"key":"e_1_2_1_64_1","doi-asserted-by":"crossref","unstructured":"Mingke Yang Yuming Zhou Bixin Li and Yutian Tang. 2023. On Code Reuse from StackOverflow: An Exploratory Study on Jupyter Notebook. http:\/\/arxiv.org\/abs\/2302.11732 arXiv:2302.11732 [cs].","DOI":"10.22541\/au.167748671.10484077\/v1"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3505243"},{"key":"e_1_2_1_66_1","unstructured":"Zezhou Yang Cuiyun Gao Zhaoqiang Guo Zhenhao Li Kui Liu Xin Xia and Yuming Zhou. 2024. A Survey on Modern Code Review: Progresses Challenges and Opportunities. http:\/\/arxiv.org\/abs\/2405.18216 arXiv:2405.18216 [cs]."},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2024.3395519"},{"key":"e_1_2_1_68_1","unstructured":"Tanghaoran Zhang Xinjun Mao Shangwen Wang Yuxin Zhao Yao Lu Jin Zhang Zhang Zhang Kang Yang and Yue Yu. 2026. AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation. arXiv:2601.04540 [cs.SE] https:\/\/arxiv.org\/abs\/2601.04540"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3180155.3180260"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019.00046"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE55347.2025"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE55347.2025.00082"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3797148","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T19:01:06Z","timestamp":1782846066000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3797148"}},"issued":{"date-parts":[[2026,6,30]]},"references-count":72,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3797148"],"URL":"https:\/\/doi.org\/10.1145\/3797148","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2026,6,30]]}},{"indexed":{"date-parts":[[2026,8,24]],"date-time":"2026-08-24T00:23:32Z","timestamp":1787531012555,"version":"build-2736575974"},"reference-count":71,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:00:00Z","timestamp":1720742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100000781","name":"European Research Council","doi-asserted-by":"publisher","award":["851895"],"award-info":[{"award-number":["851895"]}],"id":[{"id":"10.13039\/501100000781","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2024,7,12]]},"abstract":"<jats:p>Python has emerged as one of the most popular programming languages, extensively utilized in domains such as machine learning, data analysis, and web applications. Python\u2019s dynamic nature and extensive usage make it an attractive candidate for dynamic program analysis. However, unlike for other popular languages, there currently is no comprehensive benchmark suite of executable Python projects, which hinders the development of dynamic analyses. This work addresses this gap by presenting DyPyBench, the first benchmark of Python projects that is large-scale, diverse, ready-to-run (i.e., with fully configured and prepared test suites), and ready-to-analyze (by integrating with the DynaPyt dynamic analysis framework). The benchmark encompasses 50 popular open-source projects from various application domains, with a total of 681k lines of Python code, and 30k test cases. DyPyBench enables various applications in testing and dynamic analysis, of which we explore three in this work: (i) Gathering dynamic call graphs and empirically comparing them to statically computed call graphs, which exposes and quantifies limitations of existing call graph construction techniques for Python. (ii) Using DyPyBench to build a training data set for LExecutor, a neural model that learns to predict values that otherwise would be missing at runtime. (iii) Using dynamically gathered execution traces to mine API usage specifications, which establishes a baseline for future work on specification mining for Python. We envision DyPyBench to provide a basis for other dynamic analyses and for studying the runtime behavior of Python code.<\/jats:p>","DOI":"10.1145\/3643742","type":"journal-article","created":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T10:22:09Z","timestamp":1720779729000},"page":"338-358","source":"Crossref","is-referenced-by-count":12,"title":["DyPyBench: A Benchmark of Executable Python Software"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3920-3839","authenticated-orcid":false,"given":"Islem","family":"Bouzenia","sequence":"first","affiliation":[{"name":"University of Stuttgart, Stuttgart, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-1227-2265","authenticated-orcid":false,"given":"Bajaj Piyush","family":"Krishan","sequence":"additional","affiliation":[{"name":"University of Stuttgart, Stuttgart, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1623-498X","authenticated-orcid":false,"given":"Michael","family":"Pradel","sequence":"additional","affiliation":[{"name":"University of Stuttgart, Stuttgart, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,7,12]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CSF51468.2021.00043"},{"key":"e_1_3_1_3_2","first-page":"54","volume-title":"European Conference on Object-Oriented Programming","author":"Ali Karim","year":"2014","unstructured":"Karim Ali, Marianna Rapoport, Ond\u0159ej Lhot\u00e1k, Julian Dolby, and Frank Tip. 2014. Constructing call graphs of Scala programs. In European Conference on Object-Oriented Programming. Springer, 54\u201379."},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","unstructured":"Miltiadis Allamanis Earl T. Barr Soline Ducousso and Zheng Gao. 2020. Typilus: Neural Type Hints. In PLDI.","DOI":"10.1145\/3385412.3385997"},{"key":"e_1_3_1_5_2","first-page":"4","volume-title":"Symposium on Principles of Programming Languages (POPL)","author":"Ammons Glenn","year":"2002","unstructured":"Glenn Ammons, Rastislav Bod\u00edk, and James R. Larus. 2002. Mining specifications. In Symposium on Principles of Programming Languages (POPL). ACM, 4\u201316."},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","unstructured":"Jong-hoon (David) An Avik Chaudhuri Jeffrey S. Foster and Michael Hicks. 2011. Dynamic inference of static types for Ruby.. In POPL. 459\u2013472.","DOI":"10.1145\/1926385.1926437"},{"key":"e_1_3_1_7_2","doi-asserted-by":"crossref","unstructured":"David F Bacon and Peter F Sweeney. 1996. Fast static analysis of C++ virtual function calls. In Proceedings of the 11th ACM SIGPLAN conference on Object-oriented programming systems languages and applications. 324\u2013341.","DOI":"10.1145\/236337.236371"},{"key":"e_1_3_1_8_2","unstructured":"Emery D Berger Sam Stern and Juan Altmayer Pizzorno. 2023. Triangulating Python Performance Issues with {SCALENE}. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). 51\u201364."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00096"},{"key":"e_1_3_1_10_2","volume-title":"Benchmarking Modern Multiprocessors","author":"Bienia Christian","year":"2011","unstructured":"Christian Bienia. 2011. Benchmarking Modern Multiprocessors. Ph.D. Dissertation. Princeton University."},{"key":"e_1_3_1_11_2","first-page":"169","volume-title":"Conference on Object-Oriented Programming, Systems, Languages, and Applications (OOPSLA)","author":"Blackburn Stephen M.","year":"2006","unstructured":"Stephen M. Blackburn, Robin Garner, Chris Hoffmann, Asjad M. Khan, Kathryn S. McKinley, Rotem Bentzur, Amer Diwan, Daniel Feinberg, Daniel Frampton, Samuel Z. Guyer, Martin Hirzel, Antony L. Hosking, Maria Jump, Han Bok Lee, J. Eliot B. Moss, Aashish Phansalkar, Darko Stefanovic, Thomas VanDrunen, Daniel von Dincklage, and Ben Wiedermann. 2006. The DaCapo Benchmarks: Java Benchmarking Development and Analysis. In Conference on Object-Oriented Programming, Systems, Languages, and Applications (OOPSLA). ACM, 169\u2013190."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","unstructured":"Islem Bouzenia Bajaj Piyush Krishan and Michael Pradel. 2024. DyPyBench Docker Image. https:\/\/doi.org\/10.5281\/zenodo.11097202 10.5281\/zenodo.11097202","DOI":"10.5281\/zenodo.11097202"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2003.1191551"},{"key":"e_1_3_1_14_2","volume-title":"36th European Conference on Object-Oriented Programming (ECOOP 2022)","author":"Chakraborty Madhurima","year":"2022","unstructured":"Madhurima Chakraborty, Renzo Olivares, Manu Sridharan, and Behnaz Hassanshahi. 2022. Automatic root cause quantification for missing edges in javascript call graphs. In 36th European Conference on Object-Oriented Programming (ECOOP 2022). Schloss Dagstuhl-Leibniz-Zentrum f\u00fcr Informatik."},{"key":"e_1_3_1_15_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMPSAC.2014.30"},{"key":"e_1_3_1_17_2","doi-asserted-by":"crossref","first-page":"219","DOI":"10.1109\/COMPSAC.2014.30","volume-title":"2014 IEEE 38th Annual Computer Software and Applications Conference.","author":"Chen Zhifei","year":"2014","unstructured":"Zhifei Chen, Lin Chen, Yuming Zhou, Zhaogui Xu, William C Chu, and Baowen Xu. 2014. Dynamic slicing of Python programs. In 2014 IEEE 38th Annual Computer Software and Applications Conference. IEEE, 219\u2013228."},{"key":"e_1_3_1_18_2","first-page":"196","volume-title":"International Symposium on Software Testing and Analysis (ISSTA)","author":"Clause James A.","year":"2007","unstructured":"James A. Clause, Wanchun Li, and Alessandro Orso. 2007. Dytan: a generic dynamic taint analysis framework. In International Symposium on Software Testing and Analysis (ISSTA). ACM, 196\u2013206."},{"key":"e_1_3_1_19_2","unstructured":"Naji Dmeiri David A. Tomassi Yichen Wang Antara Bhowmick Yen-Chuan Liu Premkumar T. Devanbu Bogdan Vasilescu and Cindy Rubio-Gonz\u00e1lez. 2019. BugSwarm: Mining and Continuously Growing a Dataset of Reproducible Failures and Fixes. CoRR abs\/1903.06725 (2019). arXiv:1903.06725 http:\/\/arxiv.org\/abs\/1903.06725"},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","unstructured":"Brendan Dolan-Gavitt Patrick Hulin Engin Kirda Tim Leek Andrea Mambretti William K. Robertson Frederick Ulrich and Ryan Whelan. 2016. LAVA: Large-Scale Automated Vulnerability Addition. In IEEE Symposium on Security and Privacy SP 2016 San Jose CA USA May 22-26 2016. 110\u2013121.","DOI":"10.1109\/SP.2016.15"},{"key":"e_1_3_1_21_2","unstructured":"Manuel Egele Maverick Woo Peter Chapman and David Brumley. 2014. Blanket Execution: Dynamic Similarity Testing for Program Binaries and Components. In Proceedings of the 23rd USENIX Security Symposium San Diego CA USA August 20-22 2014. 303\u2013317."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549126"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"Asger Feldthaus Max Sch\u00e4fer Manu Sridharan Julian Dolby and Frank Tip. 2013. Efficient construction of approximate call graphs for JavaScript IDE services. In 35th International Conference on Software Engineering ICSE \u201913 San Francisco CA USA May 18-26 2013. 752\u2013761.","DOI":"10.1109\/ICSE.2013.6606621"},{"key":"e_1_3_1_24_2","first-page":"256","volume-title":"Symposium on Principles of Programming Languages (POPL)","author":"Flanagan Cormac","year":"2004","unstructured":"Cormac Flanagan and Stephen N. Freund. 2004. Atomizer: a dynamic atomicity checker for multithreaded programs. In Symposium on Principles of Programming Languages (POPL). ACM, 256\u2013267."},{"key":"e_1_3_1_25_2","first-page":"1","volume-title":"Workshop on Program Analysis for Software Tools and Engineering (PASTE)","author":"Flanagan Cormac","year":"2010","unstructured":"Cormac Flanagan and Stephen N. Freund. 2010. The RoadRunner dynamic analysis framework for concurrent programs. In Workshop on Program Analysis for Software Tools and Engineering (PASTE). ACM, 1\u20138."},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","unstructured":"Liang Gong Michael Pradel Manu Sridharan and Koushik Sen. 2015. DLint: Dynamically Checking Bad Coding Practices in JavaScript. In International Symposium on Software Testing and Analysis (ISSTA). 94\u2013105.","DOI":"10.1145\/2771783.2771809"},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","unstructured":"Luca Di Grazia and Michael Pradel. 2022. The Evolution of Type Annotations in Python: An Empirical Study.. In ESEC\/FSE.","DOI":"10.1145\/3540250.3549114"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICST49551.2021.00026"},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","first-page":"90","DOI":"10.1109\/ICST.2019.00019","volume-title":"2019 12th IEEE Conference on Software Testing, Validation and Verification (ICST)","author":"Gyimesi P\u00e9ter","year":"2019","unstructured":"P\u00e9ter Gyimesi, B\u00e9la Vancsics, Andrea Stocco, Davood Mazinanian, Arp\u00e1d Besz\u00e9des, Rudolf Ferenc, and Ali Mesbah. 2019. Bugsjs: a benchmark of javascript bugs. In 2019 12th IEEE Conference on Software Testing, Validation and Verification (ICST). IEEE, 90\u2013101."},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","first-page":"215","DOI":"10.1109\/ICDE.2001.914830","volume-title":"proceedings of the 17th international conference on data engineering","author":"Han Jiawei","year":"2001","unstructured":"Jiawei Han, Jian Pei, Behzad Mortazavi-Asl, Helen Pinto, Qiming Chen, Umeshwar Dayal, and Meichun Hsu. 2001. Prefixspan: Mining sequential patterns efficiently by prefix-projected pattern growth. In proceedings of the 17th international conference on data engineering. IEEE, 215\u2013224."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3428334"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/1186736.1186737"},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","unstructured":"Matthias H\u00f6schele and Andreas Zeller. 2016. Mining input grammars from dynamic taints. In Proceedings of the 31st IEEE\/ACM International Conference on Automated Software Engineering ASE 2016 Singapore September 3-7 2016. 720\u2013725.","DOI":"10.1145\/2970276.2970321"},{"key":"e_1_3_1_34_2","doi-asserted-by":"crossref","unstructured":"Ren\u00e9 Just Darioush Jalali and Michael D. Ernst. 2014. Defects4J: a database of existing faults to enable controlled testing studies for Java programs. In International Symposium on Software Testing and Analysis ISSTA \u201914 San Jose CA USA - July 21 - 26 2014. 437\u2013440.","DOI":"10.1145\/2610384.2628055"},{"key":"e_1_3_1_35_2","doi-asserted-by":"crossref","unstructured":"Daniel Lehmann and Michael Pradel. 2019. Wasabi: A Framework for Dynamically Analyzing WebAssembly. In ASPLOS.","DOI":"10.1145\/3297858.3304068"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","unstructured":"Daniel Lehmann Michelle Thalakottur Frank Tip and Michael Pradel. 2023. That\u2019s a Tough Call: Studying the Challenges of Call Graph Construction for WebAssembly. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis ISSTA 2023. 892\u2013903. https:\/\/doi.org\/10.1145\/3597926.3598104 10.1145\/3597926.3598104","DOI":"10.1145\/3597926.3598104"},{"key":"e_1_3_1_37_2","unstructured":"Shuai Lu Daya Guo Shuo Ren Junjie Huang Alexey Svyatkovskiy Ambrosio Blanco Colin Clement Dawn Drain Daxin Jiang Duyu Tang et al. 2021. Codexglue: A machine learning benchmark dataset for code understanding and generation. arXiv preprint arXiv:2102.04664 (2021)."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/1064978.1065034"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2019.00146"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/2162049.2162077"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/279310.279314"},{"key":"e_1_3_1_42_2","first-page":"89","volume-title":"Conference on Programming Language Design and Implementation (PLDI)","author":"Nethercote Nicholas","year":"2007","unstructured":"Nicholas Nethercote and Julian Seward. 2007. Valgrind: a framework for heavyweight dynamic binary instrumentation. In Conference on Programming Language Design and Implementation (PLDI). ACM, 89\u2013100."},{"key":"e_1_3_1_43_2","first-page":"167","volume-title":"Symposium on Principles and Practice of Parallel Programming (PPOPP)","author":"O\u2019Callahan Robert","year":"2003","unstructured":"Robert O\u2019Callahan and Jong-Deok Choi. 2003. Hybrid dynamic data race detection. In Symposium on Principles and Practice of Parallel Programming (PPOPP). ACM, 167\u2013178."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3468264.3468623"},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","unstructured":"Kexin Pei Jonas Guan Matthew Broughton Zhongtian Chen Songchen Yao David Williams-King Vikas Ummadisetty Junfeng Yang Baishakhi Ray and Suman Jana. 2021. StateFormer: Fine-grained type recovery from binaries using generative state modeling. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 690\u2013702.","DOI":"10.1145\/3468264.3468607"},{"key":"e_1_3_1_46_2","unstructured":"Kexin Pei Zhou Xuan Junfeng Yang Suman Jana and Baishakhi Ray. 2020. Trex: Learning execution semantics from micro-traces for binary similarity. arXiv preprint arXiv:2012.08680 (2020)."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460348"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","unstructured":"Michael Pradel Georgios Gousios Jason Liu and Satish Chandra. 2020. TypeWriter: Neural Type Prediction with Search-based Validation. In ESEC\/FSE \u201920: 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering Virtual Event USA November 8-13 2020. 209\u2013220. https:\/\/doi.org\/10.1145\/3368089.3409715 10.1145\/3368089.3409715","DOI":"10.1145\/3368089.3409715"},{"key":"e_1_3_1_49_2","doi-asserted-by":"crossref","unstructured":"Michael Pradel and Thomas R. Gross. 2009. Automatic Generation of Object Usage Specifications from Large Method Traces. In International Conference on Automated Software Engineering (ASE). 371\u2013382.","DOI":"10.1109\/ASE.2009.60"},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","unstructured":"Michael Pradel Parker Schuh and Koushik Sen. 2015. TypeDevil: Dynamic Type Inconsistency Analysis for JavaScript. In International Conference on Software Engineering (ICSE).","DOI":"10.1109\/ICSE.2015.51"},{"key":"e_1_3_1_51_2","doi-asserted-by":"crossref","unstructured":"Ingkarat Rak-amnouykit Daniel McCrevan Ana Milanova Martin Hirzel and Julian Dolby. 2020. Python 3 Types in the Wild: A Tale of Two Type Systems. In DLS.","DOI":"10.1145\/3426422.3426981"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3293882.3330555"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2012.63"},{"key":"e_1_3_1_54_2","doi-asserted-by":"crossref","unstructured":"Ripon K Saha Yingjun Lyu Wing Lam Hiroaki Yoshida and Mukul R Prasad. 2018. Bugs. jar: A large-scale diverse dataset of real-world java bugs. In Proceedings of the 15th international conference on mining software repositories. 10\u201313.","DOI":"10.1145\/3196398.3196473"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00146"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/265924.265927"},{"key":"e_1_3_1_57_2","doi-asserted-by":"crossref","unstructured":"Marija Selakovic and Michael Pradel. 2016. Performance Issues and Optimizations in JavaScript: An Empirical Study. In International Conference on Software Engineering (ICSE). 61\u201372.","DOI":"10.1145\/2884781.2884829"},{"key":"e_1_3_1_58_2","doi-asserted-by":"crossref","unstructured":"Koushik Sen Swaroop Kalasapur Tasneem Brutch and Simon Gibbs. 2013. Jalangi: A Selective Record-Replay and Dynamic Analysis Framework for JavaScript. In European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC\/FSE). 488\u2013498.","DOI":"10.1145\/2491411.2491447"},{"key":"e_1_3_1_59_2","doi-asserted-by":"crossref","unstructured":"Andreas Sewe Mira Mezini Aibek Sarimbekov and Walter Binder. 2011. Da capo con scala: Design and analysis of a scala benchmark suite for the java virtual machine. In Proceedings of the 2011 ACM international conference on Object oriented programming systems languages and applications. 657\u2013676.","DOI":"10.1145\/2048066.2048118"},{"key":"e_1_3_1_60_2","doi-asserted-by":"crossref","unstructured":"Beatriz Souza and Michael Pradel. 2023. LExecutor: Learning-Guided Execution. In FSE.","DOI":"10.1145\/3611643.3616254"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377811.3380441"},{"key":"e_1_3_1_62_2","doi-asserted-by":"crossref","unstructured":"Frank Tip and Jens Palsberg. 2000. Scalable propagation-based call graph construction algorithms. In Proceedings of the 15th ACM SIGPLAN conference on Object-oriented programming systems languages and applications. 281\u2013293.","DOI":"10.1145\/353171.353190"},{"key":"e_1_3_1_63_2","doi-asserted-by":"crossref","unstructured":"Luca Della Toffola Michael Pradel and Thomas R. Gross. 2015. Performance Problems You Can Fix: A Dynamic Analysis of Memoization Opportunities. In Conference on Object-Oriented Programming Systems Languages and Applications (OOPSLA). 607\u2013622.","DOI":"10.1145\/2814270.2814290"},{"key":"e_1_3_1_64_2","doi-asserted-by":"crossref","unstructured":"Akshay Utture Shuyang Liu Christian Gram Kalhauge and Jens Palsberg. 2022. Striking a Balance: Pruning False- Positives from Static Call Graphs. In ICSE.","DOI":"10.1145\/3510003.3510166"},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","unstructured":"Ratnadira Widyasari Sheng Qin Sim Camellia Lok Haodi Qi Jack Phan Qijin Tay Constance Tan Fiona Wee Jodie Ethelda Tan Yuheng Yieh et al. 2020. Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies. In Proceedings of the 28th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering. 1556\u20131560.","DOI":"10.1145\/3368089.3417943"},{"key":"e_1_3_1_66_2","first-page":"419","volume-title":"Conference on Programming Language Design and Implementation (PLDI)","author":"Xu Guoqing (Harry)","year":"2009","unstructured":"Guoqing (Harry) Xu, Matthew Arnold, Nick Mitchell, Atanas Rountev, and Gary Sevitsky. 2009. Go with the flow: profiling copies to find runtime bloat. In Conference on Programming Language Design and Implementation (PLDI). ACM, 419\u2013430."},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/2950290.2950357"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","unstructured":"Zhaogui Xu Xiangyu Zhang Lin Chen Kexin Pei and Baowen Xu. 2016. Python probabilistic type inference with natural language support. In Proceedings of the 24th ACM SIGSOFT International Symposium on Foundations of Software Engineering FSE 2016 Seattle WA USA November 13-18 2016. 607\u2013618. https:\/\/doi.org\/10.1145\/2950290.2950343 10.1145\/2950290.2950343","DOI":"10.1145\/2950290.2950343"},{"key":"e_1_3_1_69_2","doi-asserted-by":"crossref","unstructured":"Yanyan Yan Yang Feng Hongcheng Fan and Baowen Xu. 2023. DLInfer: Deep Learning with Static Slicing for Python Type Inference. In ICSE.","DOI":"10.1109\/ICSE48619.2023.00170"},{"key":"e_1_3_1_70_2","first-page":"282","volume-title":"International Conference on Software Engineering (ICSE)","author":"Yang Jinlin","year":"2006","unstructured":"Jinlin Yang, David Evans, Deepali Bhardwaj, Thirumalesh Bhat, and Manuvir Das. 2006. Perracotta: Mining temporal API rules from imperfect traces. In International Conference on Software Engineering (ICSE). ACM, 282\u2013291."},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/CGO51591.2021.9370317"},{"key":"e_1_3_1_72_2","volume-title":"arXiv preprint arXiv:2301.12633","author":"Zhang Zejun","year":"2023","unstructured":"Zejun Zhang, Zhenchang Xing, Xin Xia, Xiwei Xu, Liming Zhu, and Qinghua Lu. 2023. Faster or Slower? Performance Mystery of Python Idioms Unveiled with Empirical Evidence, In ICSE. arXiv preprint arXiv:2301.12633."}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643742","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3643742","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T08:04:56Z","timestamp":1770192296000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643742"}},"issued":{"date-parts":[[2024,7,12]]},"references-count":71,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2024,7,12]]}},"alternative-id":["10.1145\/3643742"],"URL":"https:\/\/doi.org\/10.1145\/3643742","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2024,7,12]]}},{"indexed":{"date-parts":[[2026,9,10]],"date-time":"2026-09-10T10:08:09Z","timestamp":1789034889220,"version":"build-2803163510"},"reference-count":80,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:00:00Z","timestamp":1750291200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/publication-rights-and-licensing-policy"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China Grant","doi-asserted-by":"crossref","award":["No.62402483"],"award-info":[{"award-number":["No.62402483"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,19]]},"abstract":"<jats:p>\n                    In software development, similar apps often encounter similar bugs due to shared functionalities and implementation methods. However, current automated GUI testing methods mainly focus on generating test scripts to cover more pages by analyzing the internal structure of the app, without targeted exploration of paths that may trigger bugs, resulting in low efficiency in bug discovery. Considering that a large number of bug reports on open source platforms can provide external knowledge for testing, this paper proposes\n                    <jats:monospace>BugHunter<\/jats:monospace>\n                    , a novel bug-aware automated GUI testing approach that generates exploration paths guided by bug reports from similar apps, utilizing a combination of Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG). Instead of focusing solely on coverage,\n                    <jats:monospace>BugHunter<\/jats:monospace>\n                    dynamically adapts the testing process to target bug paths, thereby increasing bug detection efficiency.\n                    <jats:monospace>BugHunter<\/jats:monospace>\n                    first builds a high-quality bug knowledge base from historical bug reports. Then it retrieves relevant reports from this large bug knowledge base using a two-stage retrieval process, and generates test paths based on similar apps\u2019 bug reports.\n                    <jats:monospace>BugHunter<\/jats:monospace>\n                    also introduces a local and global path-planning mechanism to handle differences in functionality and UI design across apps, and the ambiguous behavior or missing steps in the online bug reports. We evaluate\n                    <jats:monospace>BugHunter<\/jats:monospace>\n                    on 121 bugs across 71 apps and compare its performance against 16 state-of-the-art baselines.\n                    <jats:monospace>BugHunter<\/jats:monospace>\n                    achieves 60% improvement in bug detection over the best baseline, with comparable or higher coverage against the baselines. Furthermore,\n                    <jats:monospace>BugHunter<\/jats:monospace>\n                    successfully detects 49 new crash bugs in real-world apps from Google Play, with 33 bugs fixed, 9 confirmed, and 7 pending feedback.\n                  <\/jats:p>","DOI":"10.1145\/3715755","type":"journal-article","created":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T11:15:34Z","timestamp":1750331734000},"page":"825-846","source":"Crossref","is-referenced-by-count":9,"title":["Standing on the Shoulders of Giants: Bug-Aware Automated GUI Testing via Retrieval Augmentation"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-4397-750X","authenticated-orcid":false,"given":"Mengzhuo","family":"Chen","sequence":"first","affiliation":[{"name":"Institute of Software at Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9709-8275","authenticated-orcid":false,"given":"Zhe","family":"Liu","sequence":"additional","affiliation":[{"name":"Institute of Software at Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2011-9618","authenticated-orcid":false,"given":"Chunyang","family":"Chen","sequence":"additional","affiliation":[{"name":"Technical University of Munich, Munich, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9941-6713","authenticated-orcid":false,"given":"Junjie","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Software at Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-9285-3419","authenticated-orcid":false,"given":"Boyu","family":"Wu","sequence":"additional","affiliation":[{"name":"Institute of Software at Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-1530-7499","authenticated-orcid":false,"given":"Jun","family":"Hu","sequence":"additional","affiliation":[{"name":"Institute of Software at Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2618-5694","authenticated-orcid":false,"given":"Qing","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Software at Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,19]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2020. Android Debug Bridge (adb). https:\/\/developer.android.com\/studio\/command-line\/adb."},{"key":"e_1_3_1_3_2","unstructured":"2022. Android. https:\/\/developer.android.google\/topic\/."},{"key":"e_1_3_1_4_2","unstructured":"2023. pyvbox. https:\/\/pypi.org\/project\/pyvbox\/."},{"key":"e_1_3_1_5_2","unstructured":"2024. facebook 0 followers glitch. https:\/\/www.reddit.com\/r\/facebook\/comments\/1b8szmc\/follower_count_not_seen_on_facebook_profile_any\/."},{"key":"e_1_3_1_6_2","unstructured":"2024. Instagram 0 followers glitch. https:\/\/www.reddit.com\/r\/Instagram\/comments\/12sjp31\/why_does_my_acc_say_0_followers\/."},{"key":"e_1_3_1_7_2","unstructured":"2024. Tiktok 0 followers glitch. https:\/\/www.dexerto.com\/entertainment\/tiktok-0-followers-glitch-fix-for-account-bug-as-profiles-show-no-followers-1566312\/."},{"key":"e_1_3_1_8_2","unstructured":"2024. Twitters 0 followers glitch. https:\/\/medium.com\/@max-fowler\/how-to-fix-twitters-0-following-bug-a-step-by-step-guide-07e90ca65dd9."},{"key":"e_1_3_1_9_2","article-title":"A3Test: Assertion-Augmented Automated Test Case Generation","author":"Alagarsamy Saranya","year":"2023","unstructured":"Saranya Alagarsamy, Chakkrit Tantithamthavorn, and Aldeida Aleti. 2023. A3Test: Assertion-Augmented Automated Test Case Generation. arXiv preprint arXiv:2302.10352 (2023).","journal-title":"arXiv preprint arXiv:2302.10352"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1002\/spe.2564"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2019.00016"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","unstructured":"Tianqin Cai Zhao Zhang and Ping Yang. 2020. Fastbot: A Multi-Agent Model-Based Test Generation System Beijing Bytedance Network Technology Co. Ltd.. In Proceedings of the IEEE\/ACM 1st International Conference on Automation of Software Test. 93\u201396.","DOI":"10.1145\/3387903.3389308"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Yinghao Chen Zehao Hu Chen Zhi Junxiao Han Shuiguang Deng and Jianwei Yin. 2024. Chatunitest: A framework for llm-based test generation. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering. 572\u2013576.","DOI":"10.1145\/3663529.3663801"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2024.107468"},{"key":"e_1_3_1_15_2","unstructured":"Yinlin Deng Chunqiu Steven Xia Haoran Peng Chenyuan Yang and Lingming Zhang. 2022. Fuzzing Deep-Learning Libraries via Large Language Models. arXiv preprint arXiv:2212.14834 (2022)."},{"key":"e_1_3_1_16_2","unstructured":"Android Developers. 2012. Ui\/application exerciser monkey."},{"key":"e_1_3_1_17_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. NAACL-HLT 2019 (2018)."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377811.3380402"},{"key":"e_1_3_1_19_2","first-page":"408","volume-title":"2018 IEEE\/ACM 40th International Conference on Software Engineering (ICSE)","author":"Fan Lingling","year":"2018","unstructured":"Lingling Fan, Ting Su, Sen Chen, Guozhu Meng, Yang Liu, Lihua Xu, Geguang Pu, and Zhendong Su. 2018. Large-scale analysis of framework-specific exceptions in android apps. In 2018 IEEE\/ACM 40th International Conference on Software Engineering (ICSE). IEEE, 408\u2013419."},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","unstructured":"Mattia Fazzini Martin Prammer Marcelo d\u2019Amorim and Alessandro Orso. 2018. Automatically translating bug reports into test cases for mobile apps. In Proceedings of the 27th ACM SIGSOFT International Symposium on Software Testing and Analysis. 141\u2013152.","DOI":"10.1145\/3213846.3213869"},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Sidong Feng and Chunyang Chen. 2024. Prompting is all you need: Automated android bug replay with large language models. In Proceedings of the 46th IEEE\/ACM International Conference on Software Engineering. 1\u201313.","DOI":"10.1145\/3597503.3608137"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019.00042"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","first-page":"1071","DOI":"10.1109\/SP40000.2020.00071","volume-title":"2020 IEEE Symposium on Security and Privacy (SP)","author":"He Yuyu","year":"2020","unstructured":"Yuyu He, Lei Zhang, Zhemin Yang, Yinzhi Cao, Keke Lian, Shuai Li, Wei Yang, Zhibo Zhang, Min Yang, Yuan Zhang,et al. 2020. TextExerciser: feedback-driven text input exercising for android applications. In 2020 IEEE Symposium on Security and Privacy (SP). IEEE, 1071\u20131087."},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","unstructured":"Gang Hu Linjie Zhu and Junfeng Yang. 2018. AppFlow: using machine learning to synthesize robust reusable UI tests. In Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 269\u2013282.","DOI":"10.1145\/3236024.3236055"},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","first-page":"2336","DOI":"10.1109\/ICSE48619.2023.00196","volume-title":"2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE)","author":"Huang Yuchao","year":"2023","unstructured":"Yuchao Huang, Junjie Wang, Zhe Liu, Song Wang, Chunyang Chen, Mingyang Li, and Qing Wang. 2023. Context-aware bug reproduction for mobile apps. In 2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 2336\u20132348."},{"key":"e_1_3_1_26_2","volume-title":"Complete Essays: Aldous Huxley, 1938-1956","author":"Huxley Aldous","year":"2023","unstructured":"Aldous Huxley. 2023. Complete Essays: Aldous Huxley, 1938-1956. Rowman & Littlefield."},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","unstructured":"Sungmin Kang Bei Chen Shin Yoo and Jian-Guang Lou. 2023. Explainable Automated Debugging via Large Language Model-driven Scientific Debugging. arXiv preprint arXiv:2304.02195 (2023).","DOI":"10.1007\/s10664-024-10594-x"},{"key":"e_1_3_1_28_2","unstructured":"Sungmin Kang Juyeon Yoon and Shin Yoo. 2022. Large Language Models are Few-shot Testers: Exploring LLM-based General Bug Reproduction. CoRR abs\/2209.11515 (2022)."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TR.2018.2865733"},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","unstructured":"Caroline Lemieux Jeevana Priya Inala Shuvendu K Lahiri and Siddhartha Sen. 2023. CODAMOSA: Escaping coverage plateaus in test generation with pre-trained large language models. (2023).","DOI":"10.1109\/ICSE48619.2023.00085"},{"key":"e_1_3_1_31_2","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1109\/ICSE-C.2017.8","volume-title":"2017 IEEE\/ACM 39th International Conference on Software Engineering Companion (ICSE-C)","author":"Li Yuanchun","year":"2017","unstructured":"Yuanchun Li, Ziyue Yang, Yao Guo, and Xiangqun Chen. 2017. Droidbot: a lightweight ui-guided test input generator for android. In 2017 IEEE\/ACM 39th International Conference on Software Engineering Companion (ICSE-C). IEEE, 23\u201326."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2019.00104"},{"key":"e_1_3_1_33_2","first-page":"42","volume-title":"2019 34th IEEE\/ACM International Conference on Automated Software Engineering (ASE)","author":"Lin Jun-Wei","year":"2019","unstructured":"Jun-Wei Lin, Reyhaneh Jabbarvand, and Sam Malek. 2019. Test transfer across mobile apps through semantic mapping. In 2019 34th IEEE\/ACM International Conference on Automated Software Engineering (ASE). IEEE, 42\u201353."},{"key":"e_1_3_1_34_2","volume-title":"International Conference on Software Engineering","author":"Liu Changlin","year":"2022","unstructured":"Changlin Liu and Xusheng Xiao. 2022. ProMal: precise window transition graphs for Android via synergy of program analysis and machine learning. In International Conference on Software Engineering. IEEE."},{"key":"e_1_3_1_35_2","first-page":"643","volume-title":"2017 IEEE\/ACM 39th International Conference on Software Engineering (ICSE)","author":"Liu Peng","year":"2017","unstructured":"Peng Liu, Xiangyu Zhang, Marco Pistoia, Yunhui Zheng, Manoel Marques, and Lingfei Zeng. 2017. Automatic text input generation for mobile testing. In 2017 IEEE\/ACM 39th International Conference on Software Engineering (ICSE). IEEE, 643\u2013653."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00119"},{"key":"e_1_3_1_37_2","unstructured":"Zhe Liu Chunyang Chen Junjie Wang Mengzhuo Chen Boyu Wu Xing Che Dandan Wang and Qing Wang. 2023. Chatting with gpt-3 for zero-shot human-like mobile automated gui testing. arXiv preprint arXiv:2305.09434 (2023)."},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","unstructured":"Zhe Liu Chunyang Chen Junjie Wang Mengzhuo Chen Boyu Wu Xing Che Dandan Wang and Qing Wang. 2024. Make llm a testing expert: Bringing human-like interaction to mobile gui testing via functionality-aware decisions. In Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. 1\u201313.","DOI":"10.1145\/3597503.3639180"},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","unstructured":"Zhe Liu Chunyang Chen Junjie Wang Mengzhuo Chen Boyu Wu Yuekai Huang Jun Hu and Qing Wang. 2024. Unblind Text Inputs: Predicting Hint-text of Text Input in Mobile Apps via LLM. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1\u201320.","DOI":"10.1145\/3613904.3642939"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416547"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","unstructured":"Zhe Liu Chunyang Chen Junjie Wang Yuekai Huang Jun Hu and Qing Wang. 2022. Nighthawk: Fully Automated Localizing UI Display Issues via Visual Understanding. IEEE Transactions on Software Engineering 1\u201316. https:\/\/doi.org\/10.1109\/TSE.2022.3150876 10.1109\/TSE.2022.3150876","DOI":"10.1109\/TSE.2022.3150876"},{"key":"e_1_3_1_42_2","first-page":"1983","volume-title":"2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE)","author":"Liu Zhe","year":"2023","unstructured":"Zhe Liu, Chunyang Chen, Junjie Wang, Yuhui Su, Yuekai Huang, Jun Hu, and Qing Wang. 2023. Ex pede Herculem: Augmenting Activity Transition Graph for Apps via Graph Convolution Network. In 2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 1983\u20131995."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","unstructured":"Zhe Liu Chunyang Chen Junjie Wang and Qing Wang. 2022. Guided Bug Crush: Assist Manual GUI Testing of Android Apps via Hint Moves. In CHI 2022. https:\/\/doi.org\/10.1145\/3491102.3501903 10.1145\/3491102.3501903","DOI":"10.1145\/3491102.3501903"},{"key":"e_1_3_1_44_2","unstructured":"Zhe Liu Cheng Li Chunyang Chen Junjie Wang Boyu Wu Yawen Wang Jun Hu and Qing Wang. 2024. Vision-driven automated mobile gui testing via multimodal large language model. arXiv preprint arXiv:2407.03037 (2024)."},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","unstructured":"Aravind Machiry Rohan Tahiliani and Mayur Naik. 2013. Dynodroid: An input generation system for android apps. In Proceedings of the 2013 9th Joint Meeting on Foundations of Software Engineering. 224\u2013234.","DOI":"10.1145\/2491411.2491450"},{"key":"e_1_3_1_46_2","doi-asserted-by":"crossref","unstructured":"Ke Mao Mark Harman and Yue Jia. 2016. Sapienz: Multi-objective automated testing for Android applications. In Proceedings of the 25th International Symposium on Software Testing and Analysis. 94\u2013105.","DOI":"10.1145\/2931037.2931054"},{"key":"e_1_3_1_47_2","doi-asserted-by":"crossref","unstructured":"Nariman Mirzaei Joshua Garcia Hamid Bagheri Alireza Sadeghi and Sam Malek. 2016. Reducing combinatorics in GUI testing of android applications. In 2016 IEEE\/ACM 38th International Conference on Software Engineering (ICSE). IEEE 559\u2013570.","DOI":"10.1145\/2884781.2884853"},{"key":"e_1_3_1_48_2","unstructured":"Inc. NetEase Youdao. 2023. BCEmbedding: Bilingual and Crosslingual Embedding for RAG. https:\/\/github.com\/netease-youdao\/BCEmbedding."},{"key":"e_1_3_1_49_2","doi-asserted-by":"crossref","unstructured":"Minxue Pan An Huang Guoxin Wang Tian Zhang and Xuandong Li. 2020. Reinforcement learning based curiosity-driven testing of Android applications. In Proceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis. 153\u2013164.","DOI":"10.1145\/3395363.3397354"},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","unstructured":"Andrea Romdhana Alessio Merlo Mariano Ceccato and Paolo Tonella. 2022. Deep reinforcement learning for black-box testing of android apps. TOSEM (2022).","DOI":"10.1145\/3502868"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2018.2141024"},{"key":"e_1_3_1_52_2","unstructured":"Max Sch\u00e4fer Sarah Nadi Aryaz Eghbali and Frank Tip. 2023. Adaptive test generation using a large language model. arXiv preprint arXiv:2302.06527 (2023)."},{"key":"e_1_3_1_53_2","unstructured":"Max Sch\u00e4fer Sarah Nadi Aryaz Eghbali and Frank Tip. 2023. An empirical evaluation of using large language models for automated unit test generation. IEEE Transactions on Software Engineering (2023)."},{"key":"e_1_3_1_54_2","unstructured":"Mohammed Latif Siddiq Joanna Santos Ridwanul Hasan Tanvir Noshin Ulfat Fahmid Al Rifat and Vinicius Carvalho Lopes. 2023. Exploring the Effectiveness of Large Language Models in Generating Unit Tests. arXiv preprint arXiv:2305.00418 (2023)."},{"key":"e_1_3_1_55_2","doi-asserted-by":"crossref","unstructured":"Ting Su Guozhu Meng Yuting Chen Ke Wu Weiming Yang Yao Yao Geguang Pu Yang Liu and Zhendong Su. 2017. Guided stochastic model-based GUI testing of Android apps. In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering. 245\u2013256.","DOI":"10.1145\/3106237.3106298"},{"key":"e_1_3_1_56_2","doi-asserted-by":"crossref","unstructured":"Ting Su Jue Wang and Zhendong Su. 2021. Benchmarking automated GUI testing for Android against real-world bugs. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 119\u2013130.","DOI":"10.1145\/3468264.3468620"},{"key":"e_1_3_1_57_2","doi-asserted-by":"crossref","unstructured":"Yanqi Su Zheming Han Zhenchang Xing Xin Xia Xiwei Xu Liming Zhu and Qinghua Lu. 2022. Constructing a system knowledge graph of user tasks and failures from bug reports to support soap opera testing. In Proceedings of the 37th IEEE\/ACM International Conference on Automated Software Engineering. 1\u201313.","DOI":"10.1145\/3551349.3556967"},{"key":"e_1_3_1_58_2","doi-asserted-by":"crossref","unstructured":"Yanqi Su Dianshu Liao Zhenchang Xing Qing Huang Mulong Xie Qinghua Lu and Xiwei Xu. 2024. Enhancing exploratory testing by large language model and knowledge graph. In Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. 1\u201312.","DOI":"10.1145\/3597503.3639157"},{"key":"e_1_3_1_59_2","doi-asserted-by":"crossref","unstructured":"Michele Tufano Dawn Drain Alexey Svyatkovskiy and Neel Sundaresan. 2022. Generating accurate assert statements for unit test cases using pretrained transformers. In Proceedings of the 3rd ACM\/IEEE International Conference on Automation of Software Test. 54\u201364.","DOI":"10.1145\/3524481.3527220"},{"key":"e_1_3_1_60_2","unstructured":"UIAutomator. 2021. Python wrapper of Android uiautomator test tool. https:\/\/github.com\/xiaocong\/uiautomator."},{"key":"e_1_3_1_61_2","doi-asserted-by":"crossref","unstructured":"Junjie Wang Yuchao Huang Chunyang Chen Zhe Liu Song Wang and Qing Wang. 2024. Software testing with large language models: Survey landscape and vision. IEEE Transactions on Software Engineering (2024).","DOI":"10.1109\/TSE.2024.3368208"},{"key":"e_1_3_1_62_2","doi-asserted-by":"crossref","unstructured":"Jue Wang Yanyan Jiang Chang Xu Chun Cao Xiaoxing Ma and Jian Lu. 2020. Combodroid: generating high-quality test inputs for android apps via use case combinations. In ICSE. 469\u2013480.","DOI":"10.1145\/3377811.3380382"},{"key":"e_1_3_1_63_2","doi-asserted-by":"crossref","unstructured":"Jun Wang Yanhui Li Zhifei Chen Lin Chen Xiaofang Zhang and Yuming Zhou. 2024. Knowledge Graph Driven Inference Testing for Question Answering Software. In Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. 1\u201313.","DOI":"10.1145\/3597503.3639109"},{"key":"e_1_3_1_64_2","unstructured":"Junyang Wang Haiyang Xu Jiabo Ye Ming Yan Weizhou Shen Ji Zhang Fei Huang and Jitao Sang. 2024. Mobile-agent: Autonomous multi-modal mobile device agent with visual perception. arXiv preprint arXiv:2401.16158 (2024)."},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","unstructured":"Wenyu Wang Wing Lam and Tao Xie. 2021. An infrastructure approach to improving effectiveness of Android UI testing tools. In Proceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis. 165\u2013176.","DOI":"10.1145\/3460319.3464828"},{"key":"e_1_3_1_66_2","doi-asserted-by":"crossref","unstructured":"Wenyu Wang Wei Yang Tianyin Xu and Tao Xie. 2021. Vet: identifying and avoiding UI exploration tarpits. In FSE. 83\u201394.","DOI":"10.1145\/3468264.3468554"},{"key":"e_1_3_1_67_2","article-title":"Empowering llm to use smartphone for intelligent task automation","author":"Wen Hao","year":"2023","unstructured":"Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2023. Empowering llm to use smartphone for intelligent task automation. arXiv e-prints (2023), arXiv\u20132308.","journal-title":"arXiv e-prints"},{"issue":"1","key":"e_1_3_1_68_2","article-title":"Designing and comparing automated test oracles for GUI-based software applications","volume":"16","author":"Xie Qing","year":"2007","unstructured":"Qing Xie and AtifMMemon. 2007. Designing and comparing automated test oracles for GUI-based software applications. ACM Transactions on Software Engineering and Methodology (TOSEM) 16, 1 (2007), 4\u2013es.","journal-title":"ACM Transactions on Software Engineering and Methodology (TOSEM)"},{"key":"e_1_3_1_69_2","unstructured":"Zhuokui Xie Yinghao Chen Chen Zhi Shuiguang Deng and Jianwei Yin. 2023. ChatUniTest: a ChatGPT-based automated unit test generation tool. arXiv preprint arXiv:2305.04764 (2023)."},{"key":"e_1_3_1_70_2","doi-asserted-by":"crossref","unstructured":"Jiwei Yan Hao Liu Linjie Pan Jun Yan Jian Zhang and Bin Liang. 2020. Multiple-entry testing of android applications by constructing activity launching contexts. In 2020 IEEE\/ACM 42nd International Conference on Software Engineering (ICSE). IEEE 457\u2013468.","DOI":"10.1145\/3377811.3380347"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.5555\/3288647.3288710"},{"key":"e_1_3_1_72_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-37057-1_19"},{"key":"e_1_3_1_73_2","unstructured":"Zhao Yang Jiaxuan Liu Yucheng Han Xin Chen Zebiao Huang Bin Fu and Gang Yu. 2023. Appagent: Multimodal agents as smartphone users. arXiv preprint arXiv:2312.13771 (2023)."},{"key":"e_1_3_1_74_2","doi-asserted-by":"crossref","unstructured":"Husam N Yasin Siti Hafizah Ab Hamid and Raja Jamilah Raja Yusof. 2021. Droidbotx: Test case generation tool for android applications using Q-learning. Symmetry (2021).","DOI":"10.3390\/sym13020310"},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/QRS60937.2023.00029"},{"key":"e_1_3_1_76_2","unstructured":"Zhiqiang Yuan Yiling Lou Mingwei Liu Shiji Ding Kaixin Wang Yixuan Chen and Xin Peng. 2023. No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation. arXiv preprint arXiv:2305.04207 (2023)."},{"key":"e_1_3_1_77_2","doi-asserted-by":"crossref","unstructured":"Xia Zeng Dengfeng Li Wujie Zheng Fan Xia Yuetang Deng Wing Lam Wei Yang and Tao Xie. 2016. Automated test input generation for android: Are we really there yet in an industrial case?. In Proceedings of the 2016 24th ACM SIGSOFT International Symposium on Foundations of Software Engineering. 987\u2013992.","DOI":"10.1145\/2950290.2983958"},{"key":"e_1_3_1_78_2","doi-asserted-by":"crossref","unstructured":"Yakun Zhang Wenjie Zhang Dezhi Ran Qihao Zhu Chengfeng Dou Dan Hao Tao Xie and Lu Zhang. 2024. Learningbased Widget Matching for Migrating GUI Test Cases. In Proceedings of the 46th IEEE\/ACM International Conference on Software Engineering. 1\u201313.","DOI":"10.1145\/3597503.3623322"},{"key":"e_1_3_1_79_2","doi-asserted-by":"crossref","unstructured":"Yu Zhao Ting Su Yang Liu Wei Zheng Xiaoxue Wu Ramakanth Kavuluru William GJ Halfond and Tingting Yu. 2022. ReCDroid+: Automated End-to-End Crash Reproduction from Bug Reports for Android Apps. TOSEM (2022).","DOI":"10.1145\/3488244"},{"key":"e_1_3_1_80_2","doi-asserted-by":"crossref","first-page":"128","DOI":"10.1109\/ICSE.2019.00030","volume-title":"2019 IEEE\/ACM 41st International Conference on Software Engineering (ICSE)","author":"Zhao Yu","year":"2019","unstructured":"Yu Zhao, Tingting Yu, Ting Su, Yang Liu, Wei Zheng, Jingzhi Zhang, and William GJ Halfond. 2019. Recdroid: automatically reproducing android application crashes from bug reports. In 2019 IEEE\/ACM 41st International Conference on Software Engineering (ICSE). IEEE, 128\u2013139."},{"key":"e_1_3_1_81_2","first-page":"253","volume-title":"2017 IEEE\/ACM 39th International Conference on Software Engineering: Software Engineering in Practice Track (ICSE-SEIP)","author":"Zheng Haibing","year":"2017","unstructured":"Haibing Zheng, Dengfeng Li, Beihai Liang, Xia Zeng, Wujie Zheng, Yuetang Deng, Wing Lam, Wei Yang, and Tao Xie. 2017. Automated test input generation for android: Towards getting there in an industrial case. In 2017 IEEE\/ACM 39th International Conference on Software Engineering: Software Engineering in Practice Track (ICSE-SEIP). IEEE, 253\u2013262."}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715755","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T10:59:46Z","timestamp":1787482786000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715755"}},"issued":{"date-parts":[[2025,6,19]]},"references-count":80,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2025,6,19]]}},"alternative-id":["10.1145\/3715755"],"URL":"https:\/\/doi.org\/10.1145\/3715755","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2025,6,19]]}},{"indexed":{"date-parts":[[2026,9,10]],"date-time":"2026-09-10T23:12:43Z","timestamp":1789081963414,"version":"build-2803163510"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"ISSTA","license":[{"start":{"date-parts":[[2025,6,22]],"date-time":"2025-06-22T00:00:00Z","timestamp":1750550400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/publication-rights-and-licensing-policy"}],"funder":[{"DOI":"10.13039\/501100001711","name":"Swiss National Science Foundation","doi-asserted-by":"crossref","award":["200021_215487"],"award-info":[{"award-number":["200021_215487"]}],"id":[{"id":"10.13039\/501100001711","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000001","name":"NSF","doi-asserted-by":"publisher","award":["CNS-2120070"],"award-info":[{"award-number":["CNS-2120070"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000185","name":"DARPA","doi-asserted-by":"crossref","award":["FA8750-20-C-0226"],"award-info":[{"award-number":["FA8750-20-C-0226"]}],"id":[{"id":"10.13039\/100000185","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,22]]},"abstract":"<jats:p>\n                    This paper presents\n                    <jats:sc>Tratto,<\/jats:sc>\n                    a neuro-symbolic approach that generates assertions (boolean expressions) that can serve as axiomatic oracles, from source code and documentation. The symbolic module of\n                    <jats:sc>Tratto<\/jats:sc>\n                    takes advantage of the grammar of the programming language, the unit under test, and the context of the unit (its class and available APIs) to restrict the search space of the tokens that can be successfully used to generate valid oracles. The neural module of\n                    <jats:sc>Tratto<\/jats:sc>\n                    uses transformers fine-tuned for both deciding whether to output an oracle or not and selecting the next lexical token to incrementally build the oracle from the set of tokens returned by the symbolic module. Our experiments show that\n                    <jats:sc>Tratto<\/jats:sc>\n                    outperforms the state-of-the-art axiomatic oracle generation approaches, with 73% accuracy, 72% precision, and 61% F1-score, largely higher than the best results of the symbolic and neural approaches considered in our study (61%, 62%, and 37%, respectively).\n                    <jats:sc>Tratto<\/jats:sc>\n                    can generate three times more axiomatic oracles than current symbolic approaches, while generating 10 times less false positives than GPT4 complemented with few-shot learning and Chain-of-Thought prompting.\n                  <\/jats:p>","DOI":"10.1145\/3728960","type":"journal-article","created":{"date-parts":[[2025,6,22]],"date-time":"2025-06-22T10:52:56Z","timestamp":1750589576000},"page":"1887-1909","source":"Crossref","is-referenced-by-count":4,"title":["Tratto: A Neuro-Symbolic Approach to Deriving Axiomatic Test Oracles"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3995-8529","authenticated-orcid":false,"given":"Davide","family":"Molinelli","sequence":"first","affiliation":[{"name":"Constructor Institute, Schaffhausen, Switzerland"},{"name":"Universit\u00e0 della Svizzera italiana, Lugano, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5501-9225","authenticated-orcid":false,"given":"Alberto","family":"Martin-Lopez","sequence":"additional","affiliation":[{"name":"Universit\u00e0 della Svizzera italiana, Lugano, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3071-0480","authenticated-orcid":false,"given":"Elliott","family":"Zackrone","sequence":"additional","affiliation":[{"name":"University of Washington, Seattle, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6824-2765","authenticated-orcid":false,"given":"Beyza","family":"Eken","sequence":"additional","affiliation":[{"name":"Universit\u00e0 della Svizzera italiana, Lugano, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9379-277X","authenticated-orcid":false,"given":"Michael D.","family":"Ernst","sequence":"additional","affiliation":[{"name":"University of Washington, Seattle, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5193-7379","authenticated-orcid":false,"given":"Mauro","family":"Pezz\u00e8","sequence":"additional","affiliation":[{"name":"Constructor Institute, Schaffhausen, Switzerland"},{"name":"Universit\u00e0 della Svizzera italiana, Lugano, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,22]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"[n.d.]. Dataset of procedure specifications by Blasi et al. https:\/\/github.com\/albertogoffi\/toradocu\/tree\/master\/src\/test\/resources\/goal-output."},{"key":"e_1_3_1_3_2","unstructured":"Anonymous. 2025. [Replication Package] Tratto: A Neuro-Symbolic Approach to Deriving Axiomatic Test Oracles. https:\/\/anonymous.4open.science\/r\/tratto-replication-package-31F6."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/32.825766"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2014.2372785"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3558968"},{"key":"e_1_3_1_7_2","volume-title":"Proceedings of the International Symposium on Software Testing and Analysis (ISSTA \u201918)","author":"Blasi Arianna","year":"2018","unstructured":"Arianna Blasi, Alberto Goffi, Konstantin Kuznetsov, Alessandra Gorla, Michael D. Ernst, Mauro Pezz\u00e8, and Sergio Delgado Castellanos. 2018. Translating Code Comments to Procedure Specifications. In Proceedings of the International Symposium on Software Testing and Analysis (ISSTA \u201918). ACM."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3551349.3556961"},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","unstructured":"Arianna Blasi Alessandra Gorla Michael D Ernst Mauro Pezze and Antonio Carzaniga. 2021. MeMo: Automatically identifying metamorphic relations in Javadoc comments for test automation. Journal of Systems and Software 181 (2021) 111041.","DOI":"10.1016\/j.jss.2021.111041"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3571226"},{"key":"e_1_3_1_11_2","volume-title":"Technical Report HKUST-CS98-01","author":"Chen T. Y.","year":"1998","unstructured":"T. Y. Chen, S. C. Cheung, and S. M. Yiu. 1998. Metamorphic testing: A new approach for generating next test cases. Technical Report HKUST-CS98-01. HKUST Department of Computer Science, Hong Kong."},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","unstructured":"Betty HC Cheng and Joanne M Atlee. 2007. Research directions in requirements engineering. Future of software engineering (FOSE\u201907) (2007) 285\u2013303.","DOI":"10.1109\/FOSE.2007.17"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Yoonsik Cheon and Gary T. Leavens. 2002. A Simple and Practical Approach to Unit Testing: The JML and JUnit Way. In Proceedings of the European Conference on Object-Oriented Programming (ECOOP \u201902). 231\u2013255.","DOI":"10.1007\/3-540-47993-7_10"},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1007\/3-540-47993-7_10","volume-title":"ECOOP 2002 \u2013 Object-Oriented Programming, 16th European Conference","author":"Cheon Yoonsik","year":"2002","unstructured":"Yoonsik Cheon and Gary T. Leavens. 2002. A simple and practical approach to unit testing: The JML and JUnit way. In ECOOP 2002 \u2013 Object-Oriented Programming, 16th European Conference. M\u00e1laga, Spain, 231\u2013255."},{"key":"e_1_3_1_15_2","first-page":"268","volume-title":"ICFP 2000: Proceedings of the fifth ACM SIGPLAN International Conference on Functional Programming","author":"Claessen Koen","year":"2000","unstructured":"Koen Claessen and John Hughes. 2000. QuickCheck: A lightweight tool for random testing of Haskell programs. In ICFP 2000: Proceedings of the fifth ACM SIGPLAN International Conference on Functional Programming. Montreal, Canada, 268\u2013279."},{"key":"e_1_3_1_16_2","doi-asserted-by":"crossref","unstructured":"Henry Coles Thomas Laurent Christopher Henard Mike Papadakis and Anthony Ventresque. 2016. PIT: a practical mutation testing tool for Java. In Proceedings of the 25th international symposium on software testing and analysis. 449\u2013452.","DOI":"10.1145\/2931037.2948707"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2009.28"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","first-page":"281","DOI":"10.1145\/1368088.1368127","volume-title":"Proceedings of the International Conference on Software Engineering (ICSE \u201908)","author":"Csallner Christoph","year":"2008","unstructured":"Christoph Csallner, Nikolai Tillmann, and Yannis Smaragdakis. 2008. DySy: Dynamic Symbolic Execution for Invariant Inference. In Proceedings of the International Conference on Software Engineering (ICSE \u201908). ACM, 281\u2013290."},{"key":"e_1_3_1_19_2","unstructured":"J. D. Day and J. D. Gannon. 1985. A Test Oracle Based on Formal Specifications. In Proceedings of the Conference on Software Development Tools Techniques and Alternatives (SOFTAIR \u201985). 126\u2013130."},{"key":"e_1_3_1_20_2","volume-title":"ICSE 2022","author":"Dinella Elizabeth","year":"2022","unstructured":"Elizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, and Shuvendu Lahiri. 2022. TOGA: A Neural Method for Test Oracle Generation. In ICSE 2022. ACM. https:\/\/www.microsoft.com\/en-us\/research\/publication\/toga-a-neural-method-fortest-oracle-generation\/"},{"issue":"1","key":"e_1_3_1_21_2","doi-asserted-by":"crossref","first-page":"35","DOI":"10.1016\/j.scico.2007.01.015","article-title":"The Daikon system for dynamic detection of likely invariants.","volume":"69","author":"Ernst Michael D.","year":"2007","unstructured":"Michael D. Ernst, Jeff H. Perkins, Philip J. Guo, Stephen McCamant, Carlos Pacheco, Matthew S. Tschantz, and Chen Xiao. 2007. The Daikon system for dynamic detection of likely invariants. Science of Computer Programming 69, 1\u20193 (2007), 35\u201345.","journal-title":"Science of Computer Programming"},{"key":"e_1_3_1_22_2","unstructured":"Roy Thomas Fielding. 2000. Architectural Styles and the Design of Network-based Software Architectures. Ph.D. Dissertation."},{"issue":"4","key":"e_1_3_1_23_2","doi-asserted-by":"crossref","first-page":"74","DOI":"10.1145\/263244.263267","article-title":"Property-based testing: A new approach to testing for assurance.","volume":"22","author":"Fink George","year":"1997","unstructured":"George Fink and Matt Bishop. 1997. Property-based testing: A new approach to testing for assurance. ACM SIGSOFT Software Engineering Notes 22, 4 (July 1997), 74\u201380.","journal-title":"ACM SIGSOFT Software Engineering Notes"},{"key":"e_1_3_1_24_2","unstructured":"Apache Software Foundation. 2023. Apache Commons Collections: The Apache Commons Collections Library. https:\/\/commons.apache.org\/proper\/commons-collections\/. Accessed: 2024-06-07."},{"key":"e_1_3_1_25_2","unstructured":"Apache Software Foundation. 2023. Apache Commons Math: The Apache Commons Mathematics Library. https:\/\/commons.apache.org\/proper\/commons-math\/. Accessed: 2024-06-07."},{"key":"e_1_3_1_26_2","first-page":"416","volume-title":"Proceedings of the European Software Engineering Conference held jointly with the ACM SIGSOFT International Symposium on Foundations of Software Engineering (ESEC\/FSE \u201911)","author":"Fraser Gordon","year":"2011","unstructured":"Gordon Fraser and Andrea Arcuri. 2011. EvoSuite: Automatic Test Suite Generation for Object-Oriented Software. In Proceedings of the European Software Engineering Conference held jointly with the ACM SIGSOFT International Symposium on Foundations of Software Engineering (ESEC\/FSE \u201911). ACM, 416\u2013419."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2012.14"},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","first-page":"213","DOI":"10.1145\/2931037.2931061","volume-title":"Proceedings of the International Symposium on Software Testing and Analysis (ISSTA \u201916)","author":"Goffi Alberto","year":"2016","unstructured":"Alberto Goffi, Alessandra Gorla, Michael D. Ernst, and Mauro Pezz\u00e8. 2016. Automatic Generation of Oracles for Exceptional Behaviors. In Proceedings of the International Symposium on Software Testing and Analysis (ISSTA \u201916). ACM, 213\u2013224."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3611643.3616265"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2010.62"},{"key":"e_1_3_1_31_2","doi-asserted-by":"crossref","first-page":"437","DOI":"10.1145\/2610384.2628055","volume-title":"Proceedings of the International Symposium on Software Testing and Analysis (ISSTA \u201914)","author":"Just Ren\u00e9","year":"2014","unstructured":"Ren\u00e9 Just, Darioush Jalali, and Michael D. Ernst. 2014. Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs. In Proceedings of the International Symposium on Software Testing and Analysis (ISSTA \u201914). ACM, 437\u2013440."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSEA.2009.26"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00026"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jlap.2008.08.004"},{"key":"e_1_3_1_35_2","unstructured":"Anton Lozhkov Raymond Li Loubna Ben Allal Federico Cassano Joel Lamy-Poirier Nouamane Tazi Ao Tang Dmytro Pykhtar Jiawei Liu Yuxiang Wei Tianyang Liu Max Tian Denis Kocetkov Arthur Zucker Younes Belkada Zijian Wang Qian Liu Dmitry Abulkhanov Indraneil Paul Zhuang Li Wen-Ding Li Megan Risdal Jia Li Jian Zhu Terry Yue Zhuo Evgenii Zheltonozhskii Nii Osae Osae Dade Wenhao Yu Lucas Krau\u00df Naman Jain Yixuan Su Xuanli He Manan Dey Edoardo Abati Yekun Chai Niklas Muennighoff Xiangru Tang Muhtasham Oblokulov Christopher Akiki Marc Marone Chenghao Mou Mayank Mishra Alex Gu Binyuan Hui Tri Dao Armel Zebaze Olivier Dehaene Nicolas Patry Canwen Xu Julian McAuley Han Hu Torsten Scholak Sebastien Paquet Jennifer Robinson Carolyn Jane Anderson Nicolas Chapados Mostofa Patwary Nima Tajbakhsh Yacine Jernite Carlos Mu\u00f1oz Ferrandis Lingming Zhang Sean Hughes Thomas Wolf Arjun Guha Leandro von Werra and Harm de Vries. 2024. StarCoder 2 and The Stack v2: The Next Generation. arXiv:2402.19173 [cs.SE] https:\/\/arxiv.org\/abs\/2402.19173"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2022.3183297"},{"issue":"2","key":"e_1_3_1_37_2","doi-asserted-by":"crossref","first-page":"153","DOI":"10.1007\/BF02295996","article-title":"Note on the sampling error of the difference between correlated proportions or percentages.","volume":"12","author":"McNemar Quinn","year":"1947","unstructured":"Quinn McNemar. 1947. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12, 2 (1947), 153\u2013157.","journal-title":"Psychometrika"},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"1223","DOI":"10.1109\/ICSE43902.2021.00112","volume-title":"2021 IEEE\/ACM 43rd International Conference on Software Engineering (ICSE)","author":"Molina Facundo","year":"2021","unstructured":"Facundo Molina, Pablo Ponzio, Nazareno Aguirre, and Marcelo Frias. 2021. EvoSpex: An evolutionary algorithm for learning postconditions. In 2021 IEEE\/ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 1223\u20131235."},{"key":"e_1_3_1_39_2","unstructured":"OpenAI. 2023. ChatGPT: An AI Language Model. https:\/\/www.openai.com\/chatgpt. Accessed: 2024-06-07."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","unstructured":"OpenAI. 2023. GPT-4 Technical Report. CoRR abs\/2303.08774 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2303.08774 10.48550\/ARXIV.2303.08774 arXiv:2303.08774","DOI":"10.48550\/ARXIV.2303.08774"},{"key":"e_1_3_1_41_2","unstructured":"OpenAI. 2024. GPT-4 | OpenAI. https:\/\/openai.com\/index\/gpt-4\/. Accessed: 2024-10-07."},{"key":"e_1_3_1_42_2","first-page":"75","volume-title":"Proceedings of the International Conference on Software Engineering (ICSE \u201907)","author":"Pacheco Carlos","year":"2007","unstructured":"Carlos Pacheco, Shuvendu K. Lahiri, Michael D. Ernst, and Thomas Ball. 2007. Feedback-Directed Random Test Generation. In Proceedings of the International Conference on Software Engineering (ICSE \u201907). ACM, 75\u201384."},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","first-page":"815","DOI":"10.1109\/ICSE.2012.6227137","volume-title":"2012 34th international conference on software engineering (ICSE)","author":"Pandita Rahul","year":"2012","unstructured":"Rahul Pandita, Xusheng Xiao, Hao Zhong, Tao Xie, Stephen Oney, and Amit Paradkar. 2012. Inferring method specifications from natural language API descriptions. In 2012 34th international conference on software engineering (ICSE). IEEE, 815\u2013825."},{"key":"e_1_3_1_44_2","unstructured":"Emilio Parisotto Abdel-rahman Mohamed Rishabh Singh Lihong Li Dengyong Zhou and Pushmeet Kohli. 2016. Neuro-symbolic program synthesis. In International Conference on Learning Representations."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/32.667877"},{"key":"e_1_3_1_46_2","first-page":"1","article-title":"Automated Test Oracles: A Survey.","volume":"95","author":"Pezz\u00e8 Mauro","year":"2015","unstructured":"Mauro Pezz\u00e8 and Cheng Zhang. 2015. Automated Test Oracles: A Survey. In Advances in Computers. Vol. 95. Elsevier, 1\u201348.","journal-title":"Advances in Computers"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.5555\/2566876"},{"key":"e_1_3_1_48_2","unstructured":"Baptiste Rozi\u00e8re Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing Ellen Tan Yossi Adi Jingyu Liu Romain Sauvestre Tal Remez J\u00e9r\u00e9my Rapin Artyom Kozhevnikov Ivan Evtimov Joanna Bitton Manish Bhatt Cristian Canton Ferrer Aaron Grattafiori Wenhan Xiong Alexandre D\u00e9fossez Jade Copet Faisal Azhar Hugo Touvron Louis Martin Nicolas Usunier Thomas Scialom and Gabriel Synnaeve. 2024. Code Llama: Open Foundation Models for Code. arXiv:2308.12950 [cs.CL] https:\/\/arxiv.org\/abs\/2308.12950"},{"key":"e_1_3_1_49_2","first-page":"846","volume-title":"OOPSLA Companion: Object-Oriented Programming Systems, Languages, and Applications","author":"Saff David","year":"2007","unstructured":"David Saff. 2007. Theory-infected: Or how I learned to stop worrying and love universal quantification. In OOPSLA Companion: Object-Oriented Programming Systems, Languages, and Applications. Montreal, Canada, 846\u2013847."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSR52588.2021.00045"},{"key":"e_1_3_1_51_2","first-page":"260","volume-title":"Proceedings of the International Conference on Software Testing, Verification and Validation (ICST \u201912)","author":"Tan Shin Hwei","year":"2012","unstructured":"Shin Hwei Tan, Darko Marinov, Lin Tan, and Gary T. Leavens. 2012. @tComment: Testing Javadoc Comments to Detect Comment-Code Inconsistencies. In Proceedings of the International Conference on Software Testing, Verification and Validation (ICST \u201912). IEEE Computer Society, 260\u2013269."},{"key":"e_1_3_1_52_2","unstructured":"CodeGemma Team Heri Zhao Jeffrey Hui Joshua Howland Nam Nguyen Siqi Zuo Andrea Hu Christopher A. Choquette-Choo Jingyue Shen Joe Kelley Kshitij Bansal Luke Vilnis Mateo Wirth Paul Michel Peter Choy Pratik Joshi Ravin Kumar Sarmad Hashmi Shubham Agrawal Zhitao Gong Jane Fine Tris Warkentin Ale Jakse Hartman Bin Ni Kathy Korevec Kelly Schaefer and Scott Huffman. 2024. CodeGemma: Open Code Models Based on Gemma. arXiv:2406.11409 [cs.CL] https:\/\/arxiv.org\/abs\/2406.11409"},{"key":"e_1_3_1_53_2","first-page":"1178","volume-title":"Proceedings of the Joint Meeting on Foundations of Software Engineering (ESEC\/FSE \u201920)","author":"Terragni Valerio","year":"2020","unstructured":"Valerio Terragni, Gunel Jahangirova, Paolo Tonella, and Mauro Pezz\u00e8. 2020. Evolutionary Improvement of Assertion Oracles. In Proceedings of the Joint Meeting on Foundations of Software Engineering (ESEC\/FSE \u201920). ACM, 1178\u20131189."},{"key":"e_1_3_1_54_2","first-page":"253","volume-title":"ESEC\/FSE 2005: Proceedings of the 10th European Software Engineering Conference and the 13th ACM SIGSOFT Symposium on the Foundations of Software Engineering","author":"Tillmann Nikolai","year":"2005","unstructured":"Nikolai Tillmann and Wolfram Schulte. 2005. Parameterized unit tests. In ESEC\/FSE 2005: Proceedings of the 10th European Software Engineering Conference and the 13th ACM SIGSOFT Symposium on the Foundations of Software Engineering. Lisbon, Portugal, 253\u2013262."},{"key":"e_1_3_1_55_2","unstructured":"Michele Tufano Dawn Drain Alexey Svyatkovskiy Shao Kun Deng and Neel Sundaresan. 2021. Unit Test Case Generation with Transformers and Focal Context. arXiv:2009.05617 [cs.SE]"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3524481.3527220"},{"issue":"1","key":"e_1_3_1_57_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3280987","article-title":"Oracles for testing software timeliness with uncertainty.","volume":"28","author":"Wang Chunhui","year":"2018","unstructured":"Chunhui Wang, Fabrizio Pastore, and Lionel Briand. 2018. Oracles for testing software timeliness with uncertainty. ACM Transactions on Software Engineering and Methodology (TOSEM) 28, 1 (2018), 1\u201330.","journal-title":"ACM Transactions on Software Engineering and Methodology (TOSEM)"},{"key":"e_1_3_1_58_2","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1007\/978-3-540-30232-2_13","volume-title":"Formal Techniques for Networked and Distributed Systems \u2013 FORTE 2004","author":"Wang Xin","year":"2004","unstructured":"Xin Wang, Ji Wang, and Zhi-Chang Qi. 2004. Automatic Generation of Run-Time Test Oracles for Distributed Real-Time Systems. In Formal Techniques for Networked and Distributed Systems \u2013 FORTE 2004, David de Frutos-Escrig and Manuel N\u00fa\u00f1ez (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 199\u2013212."},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3386252"},{"key":"e_1_3_1_60_2","doi-asserted-by":"crossref","unstructured":"Cody Watson Michele Tufano Kevin Moran Gabriele Bavota and Denys Poshyvanyk. 2020. On learning meaningful assert statements for unit test cases. In Proceedings of the ACM\/IEEE 42nd International Conference on Software Engineering. 1398\u20131409.","DOI":"10.1145\/3377811.3380429"},{"key":"e_1_3_1_61_2","doi-asserted-by":"crossref","unstructured":"Jason Wei Xuezhi Wang Dale Schuurmans Maarten Bosma brian ichter Fei Xia Ed Chi Quoc V Le and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems S. Koyejo S. Mohamed A. Agarwal D. Belgrave K. Cho and A. Oh (Eds.). Vol. 35. Curran Associates Inc. 24824\u201324837. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf","DOI":"10.52202\/068431-1800"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728960","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728960","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T13:46:00Z","timestamp":1787492760000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728960"}},"issued":{"date-parts":[[2025,6,22]]},"references-count":60,"journal-issue":{"issue":"ISSTA","published-print":{"date-parts":[[2025,6,22]]}},"alternative-id":["10.1145\/3728960"],"URL":"https:\/\/doi.org\/10.1145\/3728960","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2025,6,22]]}},{"indexed":{"date-parts":[[2026,9,10]],"date-time":"2026-09-10T23:13:37Z","timestamp":1789082017039,"version":"build-2803163510"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:00:00Z","timestamp":1720742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2024,7,12]]},"abstract":"<jats:p>\n                    Unit testing is widely recognized as an essential aspect of the software development process. Generating high-quality assertions automatically is one of the most important and challenging problems in automatic unit test generation. To generate high-quality assertions, deep-learning-based approaches have been proposed in recent years. For state-of-the-art\n                    <jats:bold>d<\/jats:bold>\n                    eep-\n                    <jats:bold>l<\/jats:bold>\n                    earning-based approaches for\n                    <jats:bold>a<\/jats:bold>\n                    ssertion\n                    <jats:bold>g<\/jats:bold>\n                    eneration (DLAGs), the focal method (i.e., the main method under test) for a unit test case plays an important role of being a required part of the input to these approaches. To use DLAGs in practice, there are two main ways to provide a focal method for these approaches: (1) manually providing a developer-intended focal method or (2) identifying a likely focal method from the given test prefix (i.e., complete unit test code excluding assertions) with test-to-code traceability techniques. However, the state-of-the-art DLAGs are all evaluated on the ATLAS dataset, where the focal method for a test case is assumed as the last non-JUnit-API method invoked in the complete unit test code (i.e., code from both the test prefix and assertion portion). There exist two issues of the existing empirical evaluations of DLAGs, causing inaccurate assessment of DLAGs toward adoption in practice. First, it is unclear whether the last method call before assertions (LCBA) technique can accurately reflect developer-intended focal methods. Second, when applying DLAGs in practice, the assertion portion of a unit test is not available as a part of the input to DLAGs (actually being the output of DLAGs); thus, the assumption made by the ATLAS dataset does not hold in practical scenarios of applying DLAGs. To address the first issue, we conduct a study of seven test-to-code traceability techniques in the scenario of assertion generation. We find that the LCBA technique is not the best among the seven techniques and can accurately identify focal methods with only 43.38% precision and 38.42% recall; thus, the LCBA technique cannot accurately reflect developer-intended focal methods, raising a concern on using the ATLAS dataset for evaluation. To address the second issue along with the concern raised by the preceding finding, we apply\n                    <jats:bold>all seven test-to-code traceability techniques<\/jats:bold>\n                    , respectively, to identify focal methods automatically from\n                    <jats:bold>only test prefixes<\/jats:bold>\n                    and construct a new dataset named ATLAS+ by replacing the existing focal methods in the ATLAS dataset with the focal methods identified by the seven traceability techniques, respectively. On a test set from new ATLAS+, we evaluate four state-of-the-art DLAGs trained on a training set from the ATLAS dataset. We find that all of the four DLAGs achieve lower accuracy on a test set in ATLAS+ than the corresponding test set in the ATLAS dataset, indicating that DLAGs should be (re)evaluated with a test set in ATLAS+, which better reflects practical scenarios of providing focal methods than the ATLAS dataset. In addition, we evaluate state-of-the-art DLAGs trained on training sets in ATLAS+. We find that using training sets in ATLAS+ helps effectively improve the accuracy of the ATLAS approach and T5 approach over these approaches trained using the corresponding training set from the ATLAS dataset.\n                  <\/jats:p>","DOI":"10.1145\/3660785","type":"journal-article","created":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T10:22:09Z","timestamp":1720779729000},"page":"1750-1771","source":"Crossref","is-referenced-by-count":13,"title":["An Empirical Study on Focal Methods in Deep-Learning-Based Approaches for Assertion Generation"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-1999-436X","authenticated-orcid":false,"given":"Yibo","family":"He","sequence":"first","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0798-1633","authenticated-orcid":false,"given":"Jiaming","family":"Huang","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3828-7612","authenticated-orcid":false,"given":"Hao","family":"Yu","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6731-216X","authenticated-orcid":false,"given":"Tao","family":"Xie","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,7,12]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2024. https:\/\/github.com\/Yibo-He\/Artifact-ATLAS_Plus-FSE24."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","unstructured":"Miltiadis Allamanis. 2019. The Adverse Effects of Code Duplication in Machine Learning Models of Code. In Proceedings of the 2019 ACM SIGPLAN International Symposium on New Ideas New Paradigms and Reflections on Programming and Software. 143\u2013153. https:\/\/doi.org\/10.1145\/3359591.3359735 10.1145\/3359591.3359735","DOI":"10.1145\/3359591.3359735"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","unstructured":"M. Moein Almasi Hadi Hemmati Gordon Fraser Andrea Arcuri and Jundefinednis Benefelds. 2017. An Industrial Evaluation of Unit Test Generation: Finding Real Faults in a Financial Application. In Proceedings of the 39th IEEE\/ACM International Conference on Software Engineering: Software Engineering in Practice Track. 263\u2013272. https:\/\/doi.org\/10.1109\/ICSE-SEIP.2017.27 10.1109\/ICSE-SEIP.2017.27","DOI":"10.1109\/ICSE-SEIP.2017.27"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","unstructured":"Pietro Braione Giovanni Denaro and Mauro Pezz\u00e8. 2016. JBSE: A Symbolic Executor for Java Programs with Complex Heap Inputs. In Proceedings of the 24th ACM SIGSOFT International Symposium on Foundations of Software Engineering. 1018\u20131022. https:\/\/doi.org\/10.1145\/2950290.2983940 10.1145\/2950290.2983940","DOI":"10.1145\/2950290.2983940"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","unstructured":"Ermira Daka and Gordon Fraser. 2014. A Survey on Unit Testing Practices and Problems. In Proceedings of the 25th IEEE International Symposium on Software Reliability Engineering. 201\u2013211. https:\/\/doi.org\/10.1109\/ISSRE.2014.11 10.1109\/ISSRE.2014.11","DOI":"10.1109\/ISSRE.2014.11"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","unstructured":"Elizabeth Dinella Gabriel Ryan Todd Mytkowicz and Shuvendu K. Lahiri. 2022. TOGA: A Neural Method for Test Oracle Generation. In Proceedings of the 44th IEEE\/ACM International Conference on Software Engineering. 2130\u20132141. https:\/\/doi.org\/10.1145\/3510003.3510141 10.1145\/3510003.3510141","DOI":"10.1145\/3510003.3510141"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","unstructured":"Gordon Fraser and Andrea Arcuri. 2011. EvoSuite: Automatic Test Suite Generation for Object-Oriented Software. In Proceedings of the 19th ACM SIGSOFT Symposium and the 13th European Conference on Foundations of Software Engineering. 416\u2013419. https:\/\/doi.org\/10.1145\/2025113.2025179 10.1145\/2025113.2025179","DOI":"10.1145\/2025113.2025179"},{"key":"e_1_3_1_9_2","volume-title":"Introduction to Modern Information Retrieval","author":"Gerard Salton","year":"1983","unstructured":"Salton Gerard and Michael J. McGill. 1983. Introduction to Modern Information Retrieval. McGraw-Hill College."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","unstructured":"Mohammad Ghafari Carlo Ghezzi and Konstantin Rubinov. 2015. Automatically Identifying Focal Methods Under Test in Unit Test Cases. In Proceedings of the 15th IEEE International Working Conference on Source Code Analysis and Manipulation. 61\u201370. https:\/\/doi.org\/10.1109\/SCAM.2015.7335402 10.1109\/SCAM.2015.7335402","DOI":"10.1109\/SCAM.2015.7335402"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","unstructured":"Foyzul Hassan and Xiaoyin Wang. 2018. HireBuild: An Automatic Approach to History-Driven Repair of Build Scripts. In Proceedings of the 40th IEEE\/ACM International Conference on Software Engineering. 1078\u20131089. https:\/\/doi.org\/10.1145\/3180155.3180181 10.1145\/3180155.3180181","DOI":"10.1145\/3180155.3180181"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","unstructured":"Tianxing He Shengcheng Yu Ziyuan Wang Jieqiong Li and Zhenyu Chen. 2019. From Data Quality to Model Quality: An Exploratory Study on Deep Learning. In Proceedings of the 11th Asia-Pacific Symposium on Internetware. 1\u20136. https:\/\/doi.org\/10.1145\/3361242.3361260 10.1145\/3361242.3361260","DOI":"10.1145\/3361242.3361260"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","unstructured":"Ren\u00e9 Just Darioush Jalali and Michael D. Ernst. 2014. Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs. In Proceedings of the 23rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 437\u2013440. https:\/\/doi.org\/10.1145\/2610384.2628055 10.1145\/2610384.2628055","DOI":"10.1145\/2610384.2628055"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","unstructured":"Manabu Kamimura and Gail C. Murphy. 2013. Towards Generating Human-Oriented Summaries of Unit Test Cases. In Proceedings of the 21st IEEE International Conference on Program Comprehension. 215\u2013218. https:\/\/doi.org\/10.1109\/ICPC.2013.6613851 10.1109\/ICPC.2013.6613851","DOI":"10.1109\/ICPC.2013.6613851"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","unstructured":"Mike Lewis Yinhan Liu Naman Goyal Marjan Ghazvininejad Abdelrahman Mohamed Omer Levy Ves Stoyanov and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation Translation and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 7871\u20137880. https:\/\/doi.org\/10.18653\/v1\/2020.acl-main.703 10.18653\/v1\/2020.acl-main.703","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","unstructured":"Zhongxin Liu Kui Liu Xin Xia and Xiaohu Yang. 2023. Towards More Realistic Evaluation for Neural Test Oracle Generation. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis. 589\u2013600. https:\/\/doi.org\/10.1145\/3597926.3598080 10.1145\/3597926.3598080","DOI":"10.1145\/3597926.3598080"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","unstructured":"Yiling Lou Zhenpeng Chen Yanbin Cao Dan Hao and Lu Zhang. 2020. Understanding Build Issue Resolution in Practice: Symptoms and Fix Patterns. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 617\u2013628. https:\/\/doi.org\/10.1145\/3368089.3409760 10.1145\/3368089.3409760","DOI":"10.1145\/3368089.3409760"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2022.3183297"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","unstructured":"Antonio Mastropaolo Simone Scalabrino Nathan Cooper David Nader Palacio Denys Poshyvanyk Rocco Oliveto and Gabriele Bavota. 2021. Studying the Usage of Text-To-Text Transfer Transformer to Support Code-Related Tasks. In Proceedings of the 43rd IEEE\/ACM International Conference on Software Engineering. 336\u2013347. https:\/\/doi.org\/10.1109\/ICSE43902.2021.00041 10.1109\/ICSE43902.2021.00041","DOI":"10.1109\/ICSE43902.2021.00041"},{"key":"e_1_3_1_20_2","unstructured":"Microsoft. 2022. TOGA. https:\/\/github.com\/microsoft\/toga."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","unstructured":"Aleksandar Milicevic Sasa Misailovic Darko Marinov and Sarfraz Khurshid. 2007. Korat: A Tool for Generating Structurally Complex Test Inputs. In Proceedings of the 29th IEEE\/ACM International Conference on Software Engineering. 771\u2013774. https:\/\/doi.org\/10.1109\/ICSE.2007.48 10.1109\/ICSE.2007.48","DOI":"10.1109\/ICSE.2007.48"},{"key":"e_1_3_1_22_2","unstructured":"Moses. 2017. Multi-BLEU. https:\/\/github.com\/moses-smt\/mosesdecoder\/blob\/master\/scripts\/generic\/multi-bleu.perl. Retrieved on March 28 2023."},{"key":"e_1_3_1_23_2","unstructured":"NVIDIA NGC. 2017. TensorFlow Container. https:\/\/catalog.ngc.nvidia.com\/orgs\/nvidia\/containers\/tensorflow. Retrieved on March 28 2023."},{"key":"e_1_3_1_24_2","unstructured":"OpenAI. 2022. ChatGPT. https:\/\/chat.openai.com."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","unstructured":"Carlos Pacheco and Michael D. Ernst. 2007. Randoop: Feedback-Directed Random Testing for Java. In Companion to the 22nd ACM SIGPLAN Conference on Object-Oriented Programming Systems and Applications Companion. 815\u2013816. https:\/\/doi.org\/10.1145\/1297846.1297902 10.1145\/1297846.1297902","DOI":"10.1145\/1297846.1297902"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TR.2014.2338254"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","unstructured":"Abdallah Qusef Gabriele Bavota Rocco Oliveto Andrea De Lucia and David Binkley. 2011. SCOTCH: Test-to-code Traceability using Slicing and Conceptual Coupling. In Proceedings of the 27th IEEE International Conference on Software Maintenance. 63\u201372. https:\/\/doi.org\/10.1109\/ICSM.2011.6080773 10.1109\/ICSM.2011.6080773","DOI":"10.1109\/ICSM.2011.6080773"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","unstructured":"Abdallah Qusef Rocco Oliveto and Andrea De Lucia. 2010. Recovering Traceability Links between Unit Tests and Classes under Test: An Improved Method. In Proceedings of the 26th IEEE International Conference on Software Maintenance. 1\u201310. https:\/\/doi.org\/10.1109\/ICSM.2010.5609581 10.1109\/ICSM.2010.5609581","DOI":"10.1109\/ICSM.2010.5609581"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.5555\/3455716.3455856"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","unstructured":"Bart Van Rompaey and Serge Demeyer. 2009. Establishing Traceability Links between Unit Test Cases and Units under Test. In Proceedings of the 7th European Conference on Software Maintenance and Reengineering. 209\u2013218. https:\/\/doi.org\/10.1109\/CSMR.2009.39 10.1109\/CSMR.2009.39","DOI":"10.1109\/CSMR.2009.39"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","unstructured":"Sina Shamshiri. 2015. Automated Unit Test Generation for Evolving Software. In Proceedings of the 10th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 1038\u20131041. https:\/\/doi.org\/10.1145\/2786805.2803196 10.1145\/2786805.2803196","DOI":"10.1145\/2786805.2803196"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3607183"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","unstructured":"Ilya Sutskever Oriol Vinyals and Quoc V. Le. 2014. Sequence to Sequence Learning with Neural Networks. In Proceedings of the 27th International Conference on Neural Information Processing Systems. 3104\u20133112. https:\/\/doi.org\/10.48550\/arXiv.1409.3215 10.48550\/arXiv.1409.3215","DOI":"10.48550\/arXiv.1409.3215"},{"key":"e_1_3_1_34_2","unstructured":"T. T. Tanimoto. 1958. An Elementary Mathematical Theory of Classification and Prediction. IBM Internal Report."},{"key":"e_1_3_1_35_2","unstructured":"Chris Thunes. 2019. Javalang. https:\/\/github.com\/c2nes\/javalang. Retrieved on March 28 2023."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","unstructured":"Michele Tufano Shao Kun Deng Neel Sundaresan and Alexey Svyatkovskiy. 2022. Methods2Test: A Dataset of Focal Methods Mapped to Test Cases. In Proceedings of the 19th IEEE\/ACM International Conference on Mining Software Repositories. 299\u2013303. https:\/\/doi.org\/10.1145\/3524842.3528009 10.1145\/3524842.3528009","DOI":"10.1145\/3524842.3528009"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","unstructured":"Michele Tufano Dawn Drain Alexey Svyatkovskiy and Neel Sundaresan. 2022. Generating Accurate Assert Statements for Unit Test Cases using Pretrained Transformers. In Proceedings of the 17th IEEE\/ACM International Conference on Automation of Software Test. 54\u201364. https:\/\/doi.org\/10.1145\/3524481.3527220 10.1145\/3524481.3527220","DOI":"10.1145\/3524481.3527220"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","unstructured":"Sinan Wang Ming Wen Yepang Liu Ying Wang and Rongxin Wu. 2021. Understanding and Facilitating the Co-Evolution of Production and Test Code. In Proceedings of the 28th IEEE International Conference on Software Analysis Evolution and Reengineering. 272\u2013283. https:\/\/doi.org\/10.1109\/SANER50967.2021.00033 10.1109\/SANER50967.2021.00033","DOI":"10.1109\/SANER50967.2021.00033"},{"key":"e_1_3_1_39_2","unstructured":"Cody Watson. 2019. ATLAS Appendix. https:\/\/sites.google.com\/view\/atlas-nmt\/home. RetrievedonMarch28 2023."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","unstructured":"Cody Watson Michele Tufano Kevin Moran Gabriele Bavota and Denys Poshyvanyk. 2020. On Learning Meaningful Assert Statements for Unit Test Cases. In Proceedings of the 42nd IEEE\/ACM International Conference on Software Engineering. 1398\u20131409. https:\/\/doi.org\/10.1145\/3377811.3380429 10.1145\/3377811.3380429","DOI":"10.1145\/3377811.3380429"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-021-10079-1"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","unstructured":"Robert White Jens Krinke and Raymond Tan. 2020. Establishing Multilevel Test-to-Code Traceability Links. In Proceedings of the 42nd IEEE\/ACM International Conference on Software Engineering. 861\u2013872. https:\/\/doi.org\/10.1145\/3377811.3380921 10.1145\/3377811.3380921","DOI":"10.1145\/3377811.3380921"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","unstructured":"Tao Xie. 2006. Augmenting Automatically Generated Unit-Test Suites with Regression Oracle Checking. In Proceedings of the 20th European Conference on Object-Oriented Programming. 380\u2013403. https:\/\/doi.org\/10.1007\/11785477_23 10.1007\/11785477_23","DOI":"10.1007\/11785477_23"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","unstructured":"Hao Yu Yiling Lou Ke Sun Dezhi Ran Tao Xie Dan Hao Ying Li Ge Li and Qianxiang Wang. 2022. Automated Assertion Generation via Information Retrieval and Its Integration with Deep Learning. In Proceedings of the 44th IEEE\/ACM International Conference on Software Engineering. 163\u2013174. https:\/\/doi.org\/10.1145\/3510003.3510149 10.1145\/3510003.3510149","DOI":"10.1145\/3510003.3510149"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","unstructured":"Benwen Zhang Emily Hill and James Clause. 2016. Towards Automatically Generating Descriptive Names for Unit Tests. In Proceedings of the 31st IEEE\/ACM International Conference on Automated Software Engineering. 625\u2013636. https:\/\/doi.org\/10.1145\/2970276.2970342 10.1145\/2970276.2970342","DOI":"10.1145\/2970276.2970342"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3660785","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3660785","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T07:59:21Z","timestamp":1770191961000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3660785"}},"issued":{"date-parts":[[2024,7,12]]},"references-count":44,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2024,7,12]]}},"alternative-id":["10.1145\/3660785"],"URL":"https:\/\/doi.org\/10.1145\/3660785","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2024,7,12]]}},{"indexed":{"date-parts":[[2026,9,12]],"date-time":"2026-09-12T01:21:24Z","timestamp":1789176084777,"version":"build-2803163510"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:00:00Z","timestamp":1720742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2024,7,12]]},"abstract":"<jats:p>\n                    We present\n                    <jats:sc>FeatMaker<\/jats:sc>\n                    , a novel technique that automatically generates state features to enhance the search strategy of symbolic execution. Search strategies, designed to address the well-known state-explosion problem, prioritize which program states to explore. These strategies typically depend on a \u201cstate feature\u201d that describes a specific property of program states, using this feature to score and rank them. Recently, search strategies employing multiple state features have shown superior performance over traditional strategies that use a single, generic feature. However, the process of designing these features remains largely manual. Moreover, manually crafting state features is both time-consuming and prone to yielding unsatisfactory results. The goal of this paper is to fully automate the process of generating state features for search strategies from scratch. The key idea is to leverage path-conditions, which are basic but vital information maintained by symbolic execution, as state features. A challenge arises when employing all path-conditions as state features, as it results in an excessive number of state features. To address this, we present a specialized algorithm that iteratively generates and refines state features based on data accumulated during symbolic execution. Experimental results on 15 open-source C programs show that\n                    <jats:sc>FeatMaker<\/jats:sc>\n                    significantly outperforms existing search strategies that rely on manually-designed features, both in terms of branch coverage and bug detection. Notably,\n                    <jats:sc>FeatMaker<\/jats:sc>\n                    achieved an average of 35.3% higher branch coverage than state-of-the-art strategies and discovered 15 unique bugs. Of these, six were detected exclusively by\n                    <jats:sc>FeatMaker<\/jats:sc>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3660815","type":"journal-article","created":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T10:22:09Z","timestamp":1720779729000},"page":"2447-2468","source":"Crossref","is-referenced-by-count":6,"title":["FeatMaker: Automated Feature Engineering for Search Strategy of Symbolic Execution"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-5502-7520","authenticated-orcid":false,"given":"Jaehan","family":"Yoon","sequence":"first","affiliation":[{"name":"Sungkyunkwan University, Suwon, South Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4697-8536","authenticated-orcid":false,"given":"Sooyoung","family":"Cha","sequence":"additional","affiliation":[{"name":"Sungkyunkwan University, Suwon, South Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,7,12]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2020.2967380"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.5555\/2042243.2042252"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","first-page":"594","DOI":"10.1007\/s10664-013-9249-9","article-title":"Parameter tuning or default values? An empirical investigation in search-based software engineering","author":"Arcuri Andrea","year":"2013","unstructured":"AndreaArcuri and GordonFraser. 2013. Parameter tuning or default values? An empirical investigation in search-based software engineering. Empirical Software Engineering 594\u2013623","journal-title":"Empirical Software Engineering"},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","first-page":"3075","DOI":"10.1016\/j.ins.2007.11.024","article-title":"Search based software testing of object-oriented containers","author":"Arcuri Andrea","year":"2008","unstructured":"AndreaArcuri and XinYao. 2008. Search based software testing of object-oriented containers. Information Sciences 3075\u20133095","journal-title":"Information Sciences"},{"key":"e_1_3_1_6_2","first-page":"53","article-title":"Symbolic search-based testing","author":"Baars Arthur","year":"2011","unstructured":"ArthurBaars, MarkHarman, YoussefHassoun, KiranLakhotia, PhilMcMinn, PaoloTonella, and TanjaVos. 2011. Symbolic search-based testing. 2011 26th IEEE\/ACM International Conference on Automated Software Engineering (ASE 2011) (ASE \u201911) 53\u201362","journal-title":"2011 26th IEEE\/ACM International Conference on Automated Software Engineering (ASE 2011) (ASE \u201911)"},{"key":"e_1_3_1_7_2","doi-asserted-by":"crossref","unstructured":"PeterBoonstoppel CristianCadar and DawsonEngler. 2008. RWset: Attacking path explosion in constraint-based test generation. International Conference on Tools and Algorithms for the Construction and Analysis of Systems (TACAS \u201908) 351\u2013366","DOI":"10.1007\/978-3-540-78800-3_27"},{"key":"e_1_3_1_8_2","first-page":"199","article-title":"Redundant State Detection for Dynamic Symbolic Execution","author":"Bugrara Suhabe","year":"2013","unstructured":"SuhabeBugrara and DawsonEngler. 2013. Redundant State Detection for Dynamic Symbolic Execution. Proceedings of the 2013 USENIX Conference on Annual Technical Conference (USENIX ATC'13) 199\u2013212","journal-title":"Proceedings of the 2013 USENIX Conference on Annual Technical Conference (USENIX ATC'13)"},{"key":"e_1_3_1_9_2","first-page":"443","article-title":"Heuristics for Scalable Dynamic Test Generation","author":"Burnim Jacob","year":"2008","unstructured":"JacobBurnim and KoushikSen. 2008. Heuristics for Scalable Dynamic Test Generation. Proceedings of 23rd IEEE\/ACM International Conference on Automated Software Engineering (ASE'08) 443\u2013446","journal-title":"Proceedings of 23rd IEEE\/ACM International Conference on Automated Software Engineering (ASE'08)"},{"key":"e_1_3_1_10_2","first-page":"209","article-title":"KLEE: Unassisted and Automatic Generation of High-coverage Tests for Complex Systems Programs","author":"Cadar Cristian","year":"2008","unstructured":"CristianCadar, DanielDunbar, and DawsonEngler. 2008. KLEE: Unassisted and Automatic Generation of High-coverage Tests for Complex Systems Programs. Proceedings of the 8th USENIX Conference on Operating Systems Design and Implementation (OSDI \u201908) 209\u2013224.","journal-title":"Proceedings of the 8th USENIX Conference on Operating Systems Design and Implementation (OSDI \u201908)"},{"key":"e_1_3_1_11_2","first-page":"2","article-title":"Execution Generated test-cases: How to Make Systems Code Crash Itself","author":"Cadar Cristian","year":"2005","unstructured":"CristianCadar and DawsonEngler. 2005. Execution Generated test-cases: How to Make Systems Code Crash Itself. Proceedings of the 12th International Conference on Model Checking Software (SPIN'05) 2\u201323","journal-title":"Proceedings of the 12th International Conference on Model Checking Software (SPIN'05)"},{"issue":"2","key":"e_1_3_1_12_2","first-page":"10:1","article-title":"EXE: Automatically Generating Inputs of Death","volume":"12","author":"Cadar Cristian","year":"2008","unstructured":"CristianCadar, VijayGanesh, Peter M.Pawlowski, David L.Dill, and Dawson R.Engler. 2008. EXE: Automatically Generating Inputs of Death. ACM Transactions on Information and System Security 12 2 10:1\u201310:38","journal-title":"ACM Transactions on Information and System Security"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2021.3101870"},{"key":"e_1_3_1_14_2","first-page":"1244","article-title":"Automatically Generating Search Heuristics for Concolic Testing","author":"Cha Sooyoung","year":"2018","unstructured":"SooyoungCha, SeongjoonHong, JunheeLee, and HakjooOh. 2018. Automatically Generating Search Heuristics for Concolic Testing. Proceedings of the 40th International Conference on Software Engineering (ICSE \u201918) 1244\u20131254","journal-title":"Proceedings of the 40th International Conference on Software Engineering (ICSE \u201918)"},{"key":"e_1_3_1_15_2","article-title":"Making Symbolic Execution Promising by Learning Aggressive State-Pruning Strategy","author":"Cha Sooyoung","year":"2020","unstructured":"Sooyoung Cha and HakjooOh. 2020. Making Symbolic Execution Promising by Learning Aggressive State-Pruning Strategy. The 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC\/FSE \u201920)","journal-title":"The 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC\/FSE \u201920)"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3338906.3341180"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2017.61"},{"key":"e_1_3_1_18_2","unstructured":"CREST. 2008. A concolic test generation tool for C. https:\/\/github.com\/jburnim\/crest"},{"key":"e_1_3_1_19_2","first-page":"337","article-title":"Z3: An efficient SMT solver","author":"De Moura Leonardo","year":"2008","unstructured":"LeonardoDe Moura and NikolajBj\u00f8rner. 2008. Z3: An efficient SMT solver. International Conference on Tools and Algorithms for the Construction and Analysis of Systems (TACAS \u201908) 337\u2013340","journal-title":"International Conference on Tools and Algorithms for the Construction and Analysis of Systems (TACAS \u201908)"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/C-M.1978.218136"},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1109\/ICST.2012.92","article-title":"The seed is strong: Seeding strategies in search-based software testing","author":"Fraser Gordon","year":"2012","unstructured":"GordonFraser and AndreaArcuri. 2012. The seed is strong: Seeding strategies in search-based software testing. 2012 IEEE Fifth International Conference on Software Testing, Verification and Validation (ICST \u201912) 121\u2013130","journal-title":"2012 IEEE Fifth International Conference on Software Testing, Verification and Validation (ICST \u201912)"},{"key":"e_1_3_1_22_2","first-page":"519","article-title":"A decision procedure for bit-vectors and arrays","author":"Ganesh Vijay","year":"2007","unstructured":"VijayGanesh and David LDill. 2007. A decision procedure for bit-vectors and arrays. International Conference on Computer Aided Verification (CAV \u201907) 519\u2013531","journal-title":"International Conference on Computer Aided Verification (CAV \u201907)"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","first-page":"213","DOI":"10.1145\/1065010.1065036","article-title":"DART: Directed Automated Random Testing","author":"Godefroid Patrice","year":"2005","unstructured":"PatriceGodefroid, NilsKlarlund, and KoushikSen. 2005. DART: Directed Automated Random Testing. Proceedings of the 2005 ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI \u201905) 213\u2013223","journal-title":"Proceedings of the 2005 ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI \u201905)"},{"key":"e_1_3_1_24_2","first-page":"151","article-title":"Automated Whitebox Fuzz Testing","author":"Godefroid Patrice","year":"2008","unstructured":"PatriceGodefroid, Michael YLevin, and David AMolnar. 2008. Automated Whitebox Fuzz Testing. Proceedings of the Symposium on Network and Distributed System Security (NDSS \u201908) 151\u2013166","journal-title":"Proceedings of the Symposium on Network and Distributed System Security (NDSS \u201908)"},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1109\/TSE.1977.231145","article-title":"Testing Programs with the Aid of a Compiler","author":"Hamlet R.G.","year":"1977","unstructured":"R.G.Hamlet. 1977. Testing Programs with the Aid of a Compiler. IEEE Transactions on Software Engineering 279\u2013290","journal-title":"IEEE Transactions on Software Engineering"},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","DOI":"10.1109\/ICST.2015.7102580","article-title":"Achievements, open problems and challenges for search based software testing","author":"Harman Mark","year":"2015","unstructured":"MarkHarman, YueJia, and YuanyuanZhang. 2015. Achievements, open problems and challenges for search based software testing. 2015 IEEE 8th International Conference on Software Testing, Verification and Validation (ICST \u201915)","journal-title":"2015 IEEE 8th International Conference on Software Testing, Verification and Validation (ICST \u201915)"},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","first-page":"833","DOI":"10.1016\/S0950-5849(01)00189-6","article-title":"Search-based software engineering","author":"Harman Mark","year":"2001","unstructured":"MarkHarman and Bryan FJones. 2001. Search-based software engineering. Information and Software Technology 833\u2013839","journal-title":"Information and Software Technology"},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2379776.2379787","article-title":"Search-based software engineering: Trends, techniques and applications","author":"Harman Mark","year":"2012","unstructured":"MarkHarman, S AfshinMansouri, and YuanyuanZhang. 2012. Search-based software engineering: Trends, techniques and applications. ACM Computing Surveys (CSUR) 1\u201361","journal-title":"ACM Computing Surveys (CSUR)"},{"key":"e_1_3_1_29_2","first-page":"1","article-title":"Search based software engineering: Techniques, taxonomy, tutorial","author":"Harman Mark","year":"2010","unstructured":"MarkHarman, PhilMcMinn, JerffesonTeixeira De Souza, and ShinYoo. 2010. Search based software engineering: Techniques, taxonomy, tutorial Empirical software engineering and verification 1\u201359","journal-title":"Empirical software engineering and verification"},{"key":"e_1_3_1_30_2","first-page":"2526","article-title":"Learning to Explore Paths for Symbolic Execution","author":"He Jingxuan","year":"2021","unstructured":"JingxuanHe, GishorSivanrupan, PetarTsankov, and MartinVechev. 2021. Learning to Explore Paths for Symbolic Execution. Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security (CCS \u201921) 2526\u20132540","journal-title":"Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security (CCS \u201921)"},{"key":"e_1_3_1_31_2","first-page":"48","article-title":"Boosting Concolic Testing via Interpolation","author":"Jaffar Joxan","year":"2013","unstructured":"JoxanJaffar, VijayaraghavanMurali, and Jorge A.Navas. 2013. Boosting Concolic Testing via Interpolation. Proceedings of the 9th Joint Meeting on Foundations of Software Engineering (ESEC\/FSE \u201913) 48\u201358","journal-title":"Proceedings of the 9th Joint Meeting on Foundations of Software Engineering (ESEC\/FSE \u201913)"},{"key":"e_1_3_1_32_2","doi-asserted-by":"crossref","first-page":"649","DOI":"10.1109\/TSE.2010.62","article-title":"An Analysis and Survey of the Development of Mutation Testing","author":"Jia Yue","year":"2011","unstructured":"YueJia and MarkHarman. 2011. An Analysis and Survey of the Development of Mutation Testing. IEEE Transactions on Software Engineering (TSE) 649-678","journal-title":"IEEE Transactions on Software Engineering (TSE)"},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1145\/800125.804034","article-title":"Approximation Algorithms for Combinatorial Problems","author":"Johnson David S.","year":"1973","unstructured":"David S.Johnson. 1973. Approximation Algorithms for Combinatorial Problems. Proceedings of the Fifth Annual ACM Symposium on Theory of Computing 38\u201349","journal-title":"Proceedings of the Fifth Annual ACM Symposium on Theory of Computing"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4684-2001-2_9"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/3477.764879"},{"key":"e_1_3_1_36_2","first-page":"19","article-title":"Steering Symbolic Execution to Less Traveled Paths","author":"Li You","year":"2013","unstructured":"YouLi, ZhendongSu, LinzhangWang, and XuandongLi. 2013. Steering Symbolic Execution to Less Traveled Paths. Proceedings of the 2013 ACM SIGPLAN International Conference on Object Oriented Programming Systems, Languages, and Applications (OOPSLA \u201913) 19\u201332","journal-title":"Proceedings of the 2013 ACM SIGPLAN International Conference on Object Oriented Programming Systems, Languages, and Applications (OOPSLA \u201913)"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177730491"},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"153","DOI":"10.1109\/ICSTW.2011.100","article-title":"Search-based software testing: Past, present and future","author":"McMinn Phil","year":"2011","unstructured":"PhilMcMinn. 2011. Search-based software testing: Past, present and future. 2011 IEEE Fourth International Conference on Software Testing, Verification and Validation Workshops 153\u2013163","journal-title":"2011 IEEE Fourth International Conference on Software Testing, Verification and Validation Workshops"},{"key":"e_1_3_1_39_2","unstructured":"OSDI'08_Coreutil_Experiments2008Coreutils experimentshttps:\/\/klee.github.io\/docs\/coreutils-experiments"},{"key":"e_1_3_1_40_2","first-page":"35:1","article-title":"CarFast: Achieving Higher Statement Coverage Faster","author":"Park Sangmin","year":"2012","unstructured":"SangminPark, B. M.Mainul Hossain, IshtiaqueHussain, ChristophCsallner, MarkGrechanik, KunalTaneja, ChenFu, and QingXie. 2012. CarFast: Achieving Higher Statement Coverage Faster. Proceedings of the ACM SIGSOFT 20th International Symposium on the Foundations of Software Engineering (FSE \u201912) 35:1\u201335:11","journal-title":"Proceedings of the ACM SIGSOFT 20th International Symposium on the Foundations of Software Engineering (FSE \u201912)"},{"key":"e_1_3_1_41_2","first-page":"263","article-title":"CUTE: A Concolic Unit Testing Engine for C","author":"Sen Koushik","year":"2005","unstructured":"KoushikSen, DarkoMarinov, and GulAgha. 2005. CUTE: A Concolic Unit Testing Engine for C. Proceedings of the 10th European Software Engineering Conference Held Jointly with 13th ACM SIGSOFT International Symposium on Foundations of Software Engineering (ESEC\/FSE \u201905) 263\u2013272","journal-title":"Proceedings of the 10th European Software Engineering Conference Held Jointly with 13th ACM SIGSOFT International Symposium on Foundations of Software Engineering (ESEC\/FSE \u201905)"},{"key":"e_1_3_1_42_2","article-title":"How We Get There: A Context-guided Search Strategy in Concolic Testing","author":"Seo Hyunmin","year":"2014","unstructured":"HyunminSeo and SunghunKim. 2014. How We Get There: A Context-guided Search Strategy in Concolic Testing. Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering (FSE \u201914)","journal-title":"Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering (FSE \u201914)"},{"key":"e_1_3_1_43_2","unstructured":"ParaDysE2022A tool for automaticallygenerating search strategy of symbolic executionhttps:\/\/github.com\/kupl\/dd-klee\/tree\/master\/paradyse"},{"key":"e_1_3_1_44_2","unstructured":"Gcov2021A tool for measuring coveragehttps:\/\/gcc.gnu.org\/onlinedocs\/gcc\/Gcov.html"},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","first-page":"350","DOI":"10.1145\/3180155.3180251","article-title":"Chopped Symbolic Execution","author":"Trabish David","year":"2018","unstructured":"DavidTrabish, AndreaMattavelli, NoamRinetzky, and CristianCadar. 2018. Chopped Symbolic Execution. Proceedings of the 40th International Conference on Software Engineering (ICSE \u201918) 350\u2013360","journal-title":"Proceedings of the 40th International Conference on Software Engineering (ICSE \u201918)"},{"key":"e_1_3_1_46_2","doi-asserted-by":"crossref","first-page":"291","DOI":"10.1145\/3180155.3180177","article-title":"Towards Optimal Concolic Testing","author":"Wang Xinyu","year":"2018","unstructured":"XinyuWang, JunSun, ZhenbangChen, PeixinZhang, JingyiWang, and YunLin. 2018. Towards Optimal Concolic Testing. Proceedings of the 40th International Conference on Software Engineering (ICSE \u201918) 291\u2013302","journal-title":"Proceedings of the 40th International Conference on Software Engineering (ICSE \u201918)"},{"key":"e_1_3_1_47_2","first-page":"359","article-title":"Fitness-guided path exploration in dynamic symbolic execution","author":"Xie Tao","year":"2009","unstructured":"TaoXie, NikolaiTillmann, Jonathande Halleux, and WolframSchulte. 2009. Fitness-guided path exploration in dynamic symbolic execution. 2009 IEEE\/IFIP International Conference on Dependable Systems Networks 359\u2013368","journal-title":"2009 IEEE\/IFIP International Conference on Dependable Systems Networks"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2017.2659751"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2013.6693070"},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","unstructured":"LuZhang Shan-ShanHou Jun-JueHu TaoXie and HongMei. 2010. Is operator-based mutant selection superior to random mutant selection?. Proceedings of the 32nd ACM\/IEEE International Conference on Software Engineering - Volume 1 (ICSE \u201910) 435\u2013444","DOI":"10.1145\/1806799.1806863"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3660815","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3660815","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T07:55:40Z","timestamp":1770191740000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3660815"}},"issued":{"date-parts":[[2024,7,12]]},"references-count":49,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2024,7,12]]}},"alternative-id":["10.1145\/3660815"],"URL":"https:\/\/doi.org\/10.1145\/3660815","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2024,7,12]]}},{"indexed":{"date-parts":[[2026,9,12]],"date-time":"2026-09-12T01:55:57Z","timestamp":1789178157710,"version":"build-2803163510"},"reference-count":27,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:00:00Z","timestamp":1750291200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/publication-rights-and-licensing-policy"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS-2120386"],"award-info":[{"award-number":["CNS-2120386"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,19]]},"abstract":"<jats:p>\n                    Although Large Language Models (LLMs) are highly proficient in understanding source code and descriptive texts, they have limitations in reasoning on dynamic program behaviors, such as execution trace and code coverage prediction, and runtime error prediction, which usually require actual program execution. To advance the ability of LLMs in predicting dynamic behaviors, we leverage the strengths of both approaches, Program Analysis (PA) and LLM, in building\n                    <jats:sc>PredEx<\/jats:sc>\n                    , a predictive executor for Python. Our principle is a\n                    <jats:italic toggle=\"yes\">blended analysis<\/jats:italic>\n                    between PA and LLM to use PA to guide the LLM in predicting execution traces. We break down the task of predictive execution into smaller sub-tasks and leverage the deterministic nature when an execution order can be deterministically decided. When it is not certain, we use predictive backward slicing per variable, i.e., slicing the prior trace to only the parts that affect each variable separately breaks up the valuation prediction into significantly simpler problems. Our empirical evaluation on real-world datasets shows that\n                    <jats:sc>PredEx<\/jats:sc>\n                    achieves 31.5\u201347.1% relatively higher accuracy in predicting full execution traces than the state-of-the-art models. It also produces 8.6\u201353.7% more correct execution trace prefixes than those baselines. In predicting next executed statements, its relative improvement over the baselines is 15.7\u2013102.3%. Finally, we show\n                    <jats:sc>PredEx<\/jats:sc>\n                    \u2019s usefulness in two tasks: static code coverage analysis and static prediction of run-time errors for (in)complete code.\n                  <\/jats:p>","DOI":"10.1145\/3729402","type":"journal-article","created":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T11:15:34Z","timestamp":1750331734000},"page":"2987-3008","source":"Crossref","is-referenced-by-count":2,"title":["Blended Analysis for Predictive Execution"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-0143-0677","authenticated-orcid":false,"given":"Yi","family":"Li","sequence":"first","affiliation":[{"name":"University of Texas at Dallas, Computer Science Department, Dallas, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-4474-2984","authenticated-orcid":false,"given":"Hridya","family":"Dhulipala","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Computer Science Department, Dallas, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8785-6319","authenticated-orcid":false,"given":"Aashish","family":"Yadavally","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Computer Science Department, Dallas, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-8457-8528","authenticated-orcid":false,"given":"Xiaokai","family":"Rong","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Computer Science Department, Dallas, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5777-7759","authenticated-orcid":false,"given":"Shaohua","family":"Wang","sequence":"additional","affiliation":[{"name":"Central University of Finance and Economics, Bejing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-7962-6090","authenticated-orcid":false,"given":"Tien N.","family":"Nguyen","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Computer Science Department, Dallas, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,19]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"[n. d.]. python-graphs. https:\/\/github.com\/google-research\/python-graphs"},{"key":"e_1_3_1_3_2","unstructured":"[n. d.]. Python Hunter howpublished = https:\/\/github.com\/ionelmc\/python-hunter note = Accessed: 07\/21\/2023"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/93548.93576"},{"key":"e_1_3_1_5_2","unstructured":"Islem Bouzenia Yangruibo Ding Kexin Pei Baishakhi Ray and Michael Pradel. 2023. TraceFixer: Execution Trace Driven Program Repair. arXiv:2304.12743 [cs.SE]"},{"key":"e_1_3_1_6_2","first-page":"209","article-title":"KLEE: Unassisted and Automatic Generation of High-Coverage Tests for Complex Systems Programs","author":"Cadar Cristian","year":"2008","unstructured":"Cristian Cadar, Daniel Dunbar, and Dawson Engler. 2008. KLEE: Unassisted and Automatic Generation of High-Coverage Tests for Complex Systems Programs. Proceedings of the 8th USENIX Conference on Operating Systems Design and Implementation (San Diego, California) (OSDI\u201908). 209\u2013224","journal-title":"Proceedings of the 8th USENIX Conference on Operating Systems Design and Implementation (San Diego, California) (OSDI\u201908)"},{"key":"e_1_3_1_7_2","unstructured":"ChatGPT [n.d.]. OpenAI.https:\/\/openai.com\/"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3650105.3652292"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3608140"},{"key":"e_1_3_1_10_2","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1109\/SP.2017.31","article-title":"Stack Overflow Considered Harmful? The Impact of Copy&Paste on Android Application Security","author":"Fischer Felix","year":"2017","unstructured":"Felix Fischer, Konstantin B\u00f6ttinger, Huang Xiao, Christian Stransky, Yasemin Acar, Michael Backes, and Sascha Fahl. 2017. Stack Overflow Considered Harmful? The Impact of Copy&Paste on Android Application Security. Proceedings of the 2017 IEEE Symposium on Security and Privacy. 121\u2013136","journal-title":"Proceedings of the 2017 IEEE Symposium on Security and Privacy"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/1065010.1065036"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.499"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Md Mahim Anjum Haque Wasi Uddin Ahmad Ismini Lourentzou and Chris Brown. 2023. FixEval: Execution-based Evaluation of Program Fixes for Programming Problems. arXiv:2206.07796 .[cs.SE] arXiv:https:\/\/arxiv.org\/abs\/2206.07796","DOI":"10.1109\/APR59189.2023.00009"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3485832.3488026"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.308"},{"key":"e_1_3_1_16_2","unstructured":"Llama. [n.d.]. Llama. https:\/\/llama.meta.com\/"},{"key":"e_1_3_1_17_2","article-title":"NExT: Teaching Large Language Models to Reason about Code Execution","author":"Ni Ansong","year":"2024","unstructured":"Ansong Ni, Miltiadis Allamanis, Arman Cohan, Yinlin Deng, Kensen Shi, Charles Sutton, and Pengcheng Yin. 2024. NExT: Teaching Large Language Models to Reason about Code Execution. Proceedings of the International Conference on Machine Learning (Vienna, Austria) (ICML\u201924). JMLR.org, Article 1540, 28 pages","journal-title":"Proceedings of the International Conference on Machine Learning"},{"key":"e_1_3_1_18_2","unstructured":"Predictive Execution. [n.d.]. Predictive Execution. https:\/\/github.com\/predictiveexecution\/predictive-execution"},{"key":"e_1_3_1_19_2","article-title":"CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks","author":"Puri Ruchir","year":"2021","unstructured":"Ruchir Puri, David Kung, Geert Janssen, Wei Zhang, Giacomo Domeniconi, Vladimir Zolotov, Julian T Dolby, Jie Chen, Mihir Choudhury, Lindsey Decker, Veronika Thost, Luca Buratti, Saurabh Pujar, Shyam Ramji, Ulrich Finkler, Susan Malaika, and Frederick Reiss. 2021. CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks. Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, J. Vanschoren and S. Yeung (Eds.). https:\/\/datasets-benchmarks-proceedings.neurips.cc\/paper_files\/paper\/2021\/file\/a5bfc9e07964f8dddeb95fc584cd965d-Paper-round2.pdf","journal-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, J. Vanschoren and S. Yeung (Eds.)"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2019.2900307"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/1081706.1081750"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3611643.3616254"},{"key":"e_1_3_1_23_2","unstructured":"Michele Tufano Shubham Chandel Anisha Agarwal Neel Sundaresan and Colin Clement. 2023. Predicting Code Coverage without Execution. arXiv:2307.13383 [cs.SE]"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.685"},{"key":"e_1_3_1_25_2","article-title":"Chain-of-Thought Prompting Elicits Reasoning in Large Language Models","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Proceedings of the International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS \u201922). Curran Associates Inc., Red Hook, NY, USA, Article 1800, 14 pages","journal-title":"Proceedings of the International Conference on Neural Information Processing Systems"},{"key":"e_1_3_1_26_2","first-page":"1556","article-title":"BugsInPy: A Database of Existing Bugs in Python Programs to Enable Controlled Testing and Debugging Studies","author":"Widyasari Ratnadira","year":"2020","unstructured":"Ratnadira Widyasari, Sheng Qin Sim, Camellia Lok, Haodi Qi, Jack Phan, Qijin Tay, Constance Tan, Fiona Wee, Jodie Ethelda Tan, Yuheng Yieh, et al.. 2020. BugsInPy: A Database of Existing Bugs in Python Programs to Enable Controlled Testing and Debugging Studies. Proceedings of the ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 1556\u20131560","journal-title":"Proceedings of the ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00209"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019.00046"}],"container-title":["Proceedings of the ACM on Software Engineering"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729402","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729402","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T10:59:11Z","timestamp":1787482751000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3729402"}},"issued":{"date-parts":[[2025,6,19]]},"references-count":27,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2025,6,19]]}},"alternative-id":["10.1145\/3729402"],"URL":"https:\/\/doi.org\/10.1145\/3729402","ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"published":{"date-parts":[[2025,6,19]]}}],"items-per-page":20,"query":{"start-index":0,"search-terms":null}}}