{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T02:38:20Z","timestamp":1784342300950,"version":"3.55.0"},"reference-count":74,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:00:00Z","timestamp":1720742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2024,7,12]]},"abstract":"<jats:p>\n                    Large language models have gained significant popularity and are often provided as a service (i.e., LLMaaS). Companies like OpenAI and Google provide online APIs of LLMs to allow downstream users to create innovative applications. Despite its popularity, LLM safety and quality assurance is a well-recognized concern in the real world, requiring extra efforts for testing these LLMs. Unfortunately, while end-to-end services like ChatGPT have garnered rising attention in terms of testing, the LLMaaS embeddings have comparatively received less scrutiny. We state the importance of testing and uncovering problematic individual embeddings without considering downstream applications. The abstraction and non-interpretability of embedded vectors, combined with the black-box inaccessibility of LLMaaS, make testing a challenging puzzle. This paper proposes C\n                    <jats:sc>ostello<\/jats:sc>\n                    , a black-box approach to reveal potential defects in abstract embedding vectors from LLMaaS by\n                    <jats:italic toggle=\"yes\">contrastive testing<\/jats:italic>\n                    . Our intuition is that high-quality LLMs can adequately capture the semantic relationships of the input texts and properly represent their relationships in the high-dimensional space. For the given interface of LLMaaS and seed inputs,\n                    <jats:sc>Costello<\/jats:sc>\n                    can automatically generate test suites and output words with potential problematic embeddings. The idea is to synthesize contrastive samples with guidance, including positive and negative samples, by mutating seed inputs. Our synthesis guide will leverage task-specific properties to control the mutation procedure and generate samples with known partial relationships in the high-dimensional space. Thus, we can compare the expected relationship (oracle) and embedding distance (output of LLMs) to locate potential buggy cases. We evaluate C\n                    <jats:sc>ostello<\/jats:sc>\n                    on 42 open-source (encoder-based) language models and two real-world commercial LLMaaS. Experimental results show that C\n                    <jats:sc>ostello<\/jats:sc>\n                    can effectively detect semantic violations, where more than 62% of violations on average result in erroneous behaviors (e.g., unfairness) of downstream applications.\n                  <\/jats:p>","DOI":"10.1145\/3643767","type":"journal-article","created":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T10:22:09Z","timestamp":1720779729000},"page":"906-928","source":"Crossref","is-referenced-by-count":3,"title":["COSTELLO: Contrastive Testing for Embedding-Based Large Language Model as a Service Embeddings"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0382-6401","authenticated-orcid":false,"given":"Weipeng","family":"Jiang","sequence":"first","affiliation":[{"name":"Xi'an Jiaotong University, Xi'an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5017-8016","authenticated-orcid":false,"given":"Juan","family":"Zhai","sequence":"additional","affiliation":[{"name":"University of Massachusetts Amherst, Amherst, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1551-8948","authenticated-orcid":false,"given":"Shiqing","family":"Ma","sequence":"additional","affiliation":[{"name":"University of Massachusetts Amherst, Amherst, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7010-6749","authenticated-orcid":false,"given":"Xiaoyu","family":"Zhang","sequence":"additional","affiliation":[{"name":"Xi'an Jiaotong University, Xi'an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6959-0569","authenticated-orcid":false,"given":"Chao","family":"Shen","sequence":"additional","affiliation":[{"name":"Xi'an Jiaotong University, Xi'an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,7,12]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2021. Embedding API-Build faster great AI algorithms for HR. https:\/\/hrflow.ai\/embedding\/\/."},{"key":"e_1_3_1_3_2","unstructured":"2022. Introducing Text and Code Embeddings in the OpenAI API. https:\/\/openai.com\/blog\/introducing-text-and-codeembeddings\/."},{"key":"e_1_3_1_4_2","unstructured":"2023. ChatGPT could be used for good but like many other AI models it's rife with racist and discriminatory bias. https:\/\/www.insider.com\/chatgpt-is-like-many-other-ai-models-rife-with-bias-2023-1."},{"key":"e_1_3_1_5_2","unstructured":"2023. DeepL. https:\/\/www.deepl.com\/."},{"key":"e_1_3_1_6_2","unstructured":"2023. Hugging Face. https:\/\/huggingface.co\/."},{"key":"e_1_3_1_7_2","unstructured":"2023. Introducing ChatGPT. https:\/\/openai.com\/blog\/chatgpt."},{"key":"e_1_3_1_8_2","unstructured":"2023. Tutorial: ChatGPT Over Your Data. https:\/\/blog.langchain.dev\/tutorial-chatgpt-over-your-data\/."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2021.3136169"},{"issue":"2020","key":"e_1_3_1_10_2","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877-1901.","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_1_11_2","unstructured":"Kangjie Chen Yuxian Meng Xiaofei Sun Shangwei Guo Tianwei Zhang Jiwei Li and Chun Fan. 2021. Badpre: Task-agnostic backdoor attacks to pre-trained nlp foundation models. arXiv preprint arXiv:2110.02467 (2021)."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE51524.2021.9678670"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","unstructured":"Songqiang Chen Shuo Jin and Xiaoyuan Xie. 2021. Validation on machine reading comprehension software without annotated labels: A property-based method. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 590-602. https:\/\/doi.org\/10.1145\/3468264.3468569 10.1145\/3468264.3468569","DOI":"10.1145\/3468264.3468569"},{"key":"e_1_3_1_14_2","first-page":"1597","volume-title":"International conference on machine learning","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning. PMLR, 1597-1607."},{"key":"e_1_3_1_15_2","unstructured":"Aakanksha Chowdhery Sharan Narang Jacob Devlin Maarten Bosma Gaurav Mishra Adam Roberts Paul Barham Hyung Won Chung Charles Sutton Sebastian Gehrmann et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311 (2022)."},{"key":"e_1_3_1_16_2","doi-asserted-by":"crossref","unstructured":"Yung-Sung Chuang Rumen Dangovski Hongyin Luo Yang Zhang Shiyu Chang Marin Solja\u010di\u0107 Shang-Wen Li Wen-tau Yih Yoon Kim and James Glass. 2022. DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings. arXiv preprint arXiv:2204.10298 (2022).","DOI":"10.18653\/v1\/2022.naacl-main.311"},{"key":"e_1_3_1_17_2","unstructured":"Gregory W Corder and Dale I Foreman. 2011. Nonparametric statistics for non-statisticians."},{"key":"e_1_3_1_18_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","unstructured":"Xiaoning Du Xiaofei Xie Yi Li Lei Ma Yang Liu and Jianjun Zhao. 2019. Deepstellar: Model-based quantitative analysis of stateful deep learning systems. In Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 477-487.","DOI":"10.1145\/3338906.3338954"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/MIS.2017.4531228"},{"key":"e_1_3_1_21_2","unstructured":"Andrea Esuli and Fabrizio Sebastiani. 2006. Sentiwordnet: A publicly available lexical resource for opinion mining. In Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC'06)."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11023-020-09548-1"},{"key":"e_1_3_1_23_2","unstructured":"Tianyu Gao Xingcheng Yao and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. arXiv preprint arXiv:2104.08821 (2021)."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","unstructured":"Xuanqi Gao Juan Zhai Shiqing Ma Chao Shen Yufei Chen and Qian Wang. 2022. FairNeuron: Improving Deep Neural Network Fairness with Adversary Games on Selective Neurons. arXiv preprint arXiv:2204.02567 (2022). https:\/\/doi.org\/10.1145\/3510003.3510087 10.1145\/3510003.3510087","DOI":"10.1145\/3510003.3510087"},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","unstructured":"Matt Gardner Yoav Artzi Victoria Basmova Jonathan Berant Ben Bogin Sihao Chen Pradeep Dasigi Dheeru Dua Yanai Elazar Ananth Gottumukkala et al. 2020. Evaluating Models' Local Decision Boundaries via Contrast Sets. arXiv preprint arXiv:2004.02709 (2020).","DOI":"10.18653\/v1\/2020.findings-emnlp.117"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377811.3380391"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01453-z"},{"key":"e_1_3_1_28_2","unstructured":"Beliz Gunel Jingfei Du Alexis Conneau and Ves Stoyanov. 2020. Supervised contrastive learning for pre-trained language model fine-tuning. arXiv preprint arXiv:2011.01403 (2020)."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","unstructured":"Jianmin Guo Yue Zhao Xueying Han Yu Jiang and Jiaguang Sun. 2019. Rnn-test: Adversarial testing framework for recurrent neural network systems. arXiv preprint arXiv:1911.06155 (2019). https:\/\/doi.org\/10.1109\/TSE.2021.3114353 10.1109\/TSE.2021.3114353","DOI":"10.1109\/TSE.2021.3114353"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416571"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","unstructured":"Shashij Gupta Pinjia He Clara Meister and Zhendong Su. 2020. Machine translation testing via pathological invariance. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 863-875. https:\/\/doi.org\/10.1145\/3368089.3409756 10.1145\/3368089.3409756","DOI":"10.1145\/3368089.3409756"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377811.3380339"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00047"},{"key":"e_1_3_1_34_2","unstructured":"R Devon Hjelm Alex Fedorov Samuel Lavoie-Marchildon Karan Grewal Phil Bachman Adam Trischler and Yoshua Bengio. 2018. Learning deep representations by mutual information estimation and maximization. arXiv preprint arXiv:1808.06670 (2018)."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","unstructured":"Pin Ji Yang Feng Jia Liu Zhihong Zhao and Zhenyu Chen. 2022. ASRTest: automated testing for deep-neural-network-driven speech recognition systems. In Proceedings of the 31st ACM SIGSOFT International Symposium on Software Testing and Analysis. 189-201. https:\/\/doi.org\/10.1145\/3533767.3534391 10.1145\/3533767.3534391","DOI":"10.1145\/3533767.3534391"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019.00108"},{"key":"e_1_3_1_37_2","unstructured":"Zhenzhong Lan Mingda Chen Sebastian Goodman Kevin Gimpel Piyush Sharma and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942 (2019)."},{"key":"e_1_3_1_38_2","unstructured":"Dong-Ho Lee Mahak Agarwal Akshen Kadakia Jay Pujara and Xiang Ren. 2021. Good examples make A faster learner: Simple demonstration-based learning for low-resource NER. arXiv preprint arXiv:2110.08454 (2021)."},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","unstructured":"Mike Lewis Yinhan Liu Naman Goyal Marjan Ghazvininejad Abdelrahman Mohamed Omer Levy Ves Stoyanov and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation translation and comprehension. arXiv preprint arXiv:1910.13461 (2019).","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_1_40_2","unstructured":"Pengfei Liu Weizhe Yuan Jinlan Fu Zhengbao Jiang Hiroaki Hayashi and Graham Neubig. 2021. Pre-train prompt and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586 (2021)."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","unstructured":"Shangqing Liu Bozhi Wu Xiaofei Xie Guozhu Meng and Yang Liu. 2023. ContraBERT: Enhancing Code Pre-Trained Models via Contrastive Learning. (2023) 2476-2487. https:\/\/doi.org\/10.1109\/ICSE48619.2023.00207 10.1109\/ICSE48619.2023.00207","DOI":"10.1109\/ICSE48619.2023.00207"},{"key":"e_1_3_1_42_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019)."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/SANER.2019.8668044"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","unstructured":"Lei Ma Felix Juefei-Xu Fuyuan Zhang Jiyuan Sun Minhui Xue Bo Li Chunyang Chen Ting Su Li Li Yang Liu et al. 2018. Deepgauge: Multi-granularity testing criteria for deep learning systems. In Proceedings of the 33rd ACM\/IEEE International Conference on Automated Software Engineering. 120-131. https:\/\/doi.org\/10.1145\/3238147.3238202 10.1145\/3238147.3238202","DOI":"10.1145\/3238147.3238202"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","unstructured":"Shiqing Ma Yingqi Liu Wen-Chuan Lee Xiangyu Zhang and Ananth Grama. 2018. MODE: automated neural network model debugging via state differential analysis and input selection. In Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 175-186. https:\/\/doi.org\/10.1145\/3236024.3236082 10.1145\/3236024.3236082","DOI":"10.1145\/3236024.3236082"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/219717.219748"},{"key":"e_1_3_1_47_2","first-page":"4901","volume-title":"International Conference on Machine Learning","author":"Odena Augustus","year":"2019","unstructured":"Augustus Odena, Catherine Olsson, David Andersen, and Ian Goodfellow 2019 Tensorfuzz: Debugging neural networks with coverage-guided fuzzing In International Conference on Machine Learning. PMLR, 4901\u20134911."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","unstructured":"Kexin Pei Yinzhi Cao Junfeng Yang and Suman Jana. 2017. Deepxplore: Automated whitebox testing of deep learning systems. In proceedings of the 26th Symposium on Operating Systems Principles. 1-18. https:\/\/doi.org\/10.1145\/3361566 10.1145\/3361566","DOI":"10.1145\/3361566"},{"key":"e_1_3_1_49_2","doi-asserted-by":"crossref","unstructured":"Matthew E Peters Sebastian Ruder and Noah A Smith. 2019. To tune or not to tune? adapting pretrained representations to diverse tasks. arXiv preprint arXiv:1903.05987 (2019).","DOI":"10.18653\/v1\/W19-4302"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2020.3038167"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3234150"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3561970"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","unstructured":"Marco Tulio Ribeiro Tongshuang Wu Carlos Guestrin and Sameer Singh. 2020. Beyond Accuracy: Behavioral Testing of NLP Models with CheckList. (July 2020) 4902-4912. https:\/\/doi.org\/10.18653\/v1\/2020.acl-main.442 10.18653\/v1\/2020.acl-main.442","DOI":"10.18653\/v1\/2020.acl-main.442"},{"key":"e_1_3_1_54_2","doi-asserted-by":"crossref","unstructured":"Alexis Ross Tongshuang Wu Hao Peng Matthew E Peters and Matt Gardner. 2021. Tailor: Generating and perturbing text with semantic controls. arXiv preprint arXiv:2107.07150 (2021).","DOI":"10.18653\/v1\/2022.acl-long.228"},{"key":"e_1_3_1_55_2","unstructured":"Ohad Rubin Jonathan Herzig and Jonathan Berant. 2021. Learning to retrieve prompts for in-context learning. arXiv preprint arXiv:2112.08633 (2021)."},{"key":"e_1_3_1_56_2","unstructured":"Victor Sanh Lysandre Debut Julien Chaumond and Thomas Wolf. 2019. DistilBERT a distilled version of BERT: smaller faster cheaper and lighter. arXiv preprint arXiv:1910.01108 (2019)."},{"key":"e_1_3_1_57_2","doi-asserted-by":"crossref","unstructured":"Arshdeep Sekhon Yangfeng Ji Matthew B Dwyer and Yanjun Qi. 2022. White-box Testing of NLP models with Mask Neuron Coverage. arXiv preprint arXiv:2205.05050 (2022).","DOI":"10.18653\/v1\/2022.findings-naacl.116"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","unstructured":"Qingchao Shen Junjie Chen Jie M Zhang Haoyu Wang Shuang Liu and Menghan Tian. 2022. Natural Test Generation for Precise Testing of Question Answering Software. In 37th IEEE\/ACM International Conference on Automated Software Engineering. 1-12. https:\/\/doi.org\/10.1145\/3551349.3556953 10.1145\/3551349.3556953","DOI":"10.1145\/3551349.3556953"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","unstructured":"Ensheng Shi Yanlin Wang Wenchao Gu Lun Du Hongyu Zhang Shi Han Dongmei Zhang and Hongbin Sun. 2023. CoCoSoDa: Effective Contrastive Learning for Code Search. (2023) 2198-2210. https:\/\/doi.org\/10.1109\/ICSE48619.2023.00185 10.1109\/ICSE48619.2023.00185","DOI":"10.1109\/ICSE48619.2023.00185"},{"key":"e_1_3_1_60_2","unstructured":"Tianxiang Sun Yunfan Shao Hong Qian Xuanjing Huang and Xipeng Qiu. 2022. Black-box tuning for language-model-as-a-service. arXiv preprint arXiv:2201.03514 (2022)."},{"key":"e_1_3_1_61_2","unstructured":"Yu Sun Shuohuan Wang Shikun Feng Siyu Ding Chao Pang Junyuan Shang Jiaxiang Liu Xuyi Chen Yanbin Zhao Yuxiang Lu et al. 2021. Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint arXiv:2107.02137 (2021)."},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","unstructured":"Youcheng Sun Min Wu Wenjie Ruan Xiaowei Huang Marta Kwiatkowska and Daniel Kroening. 2018. Concolic testing for deep neural networks. In Proceedings of the 33rd ACM\/IEEE International Conference on Automated Software Engineering. 109-119. https:\/\/doi.org\/10.1145\/3238147.3238172 10.1145\/3238147.3238172","DOI":"10.1145\/3238147.3238172"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","unstructured":"Zeyu Sun Jie M Zhang Mark Harman Mike Papadakis and Lu Zhang. 2020. Automatic testing and improvement of machine translation. In Proceedings of the ACM\/IEEE 42nd International Conference on Software Engineering. 974-985. https:\/\/doi.org\/10.1145\/3377811.3380420 10.1145\/3377811.3380420","DOI":"10.1145\/3377811.3380420"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510206"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58621-8_45"},{"key":"e_1_3_1_66_2","unstructured":"Jindong Wang Xixu Hu Wenxin Hou Hao Chen Runkai Zheng Yidong Wang Linyi Yang Haojun Huang Wei Ye Xiubo Geng et al. 2023. On the robustness of chatgpt: An adversarial and out-of-distribution perspective. arXiv preprint arXiv:2302.12095 (2023)."},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","unstructured":"Zan Wang Ming Yan Junjie Chen Shuang Liu and Dongdi Zhang. 2020. Deep learning library testing via effective model generation. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 788-799. https:\/\/doi.org\/10.1145\/3368089.3409761 10.1145\/3368089.3409761","DOI":"10.1145\/3368089.3409761"},{"key":"e_1_3_1_68_2","unstructured":"Tongshuang Wu Marco Tulio Ribeiro Jeffrey Heer and Daniel S Weld. 2021. Polyjuice: Automated general-purpose counterfactual generation. arXiv preprint arXiv:2101.00288 1 2 (2021)."},{"key":"e_1_3_1_69_2","unstructured":"Zhuofeng Wu Sinong Wang Jiatao Gu Madian Khabsa Fei Sun and Hao Ma. 2020. Clear: Contrastive learning for sentence representation. arXiv preprint arXiv:2012.15466 (2020)."},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","unstructured":"Danning Xie Yitong Li Mijung Kim Hung Viet Pham Lin Tan Xiangyu Zhang and Michael W Godfrey. 2022. Docter: Documentation-guided fuzzing for testing deep learning api functions. In Proceedings of the 31st ACM SIGSOFT International Symposium on Software Testing and Analysis. 176-188. https:\/\/doi.org\/10.1145\/3533767.3534220 10.1145\/3533767.3534220","DOI":"10.1145\/3533767.3534220"},{"key":"e_1_3_1_71_2","unstructured":"Xiaofei Xie Lei Ma Felix Juefei-Xu Hongxu Chen Minhui Xue Bo Li Yang Liu Jianjun Zhao Jianxiong Yin and Simon See. 2018. Deephunter: Hunting deep neural network defects via coverage-guided fuzzing. arXiv preprint arXiv:1809.01266 (2018)."},{"key":"e_1_3_1_72_2","unstructured":"Xiaofei Xie Lei Ma Haijun Wang Yuekang Li Yang Liu and Xiaohong Li. 2019. Diffchaser: Detecting disagreements for deep neural networks. International Joint Conferences on Artificial Intelligence Organization."},{"key":"e_1_3_1_73_2","unstructured":"Yuanmeng Yan Rumei Li Sirui Wang Fuzheng Zhang Wei Wu and Weiran Xu. 2021. Consert: A contrastive framework for self-supervised sentence representation transfer. arXiv preprint arXiv:2105.11741 (2021)."},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSME52107.2021.00073"},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","unstructured":"Peixin Zhang Jingyi Wang Jun Sun Guoliang Dong Xinyu Wang Xingen Wang Jin Song Dong and Ting Dai. 2020. White-box fairness testing through adversarial sampling. In Proceedings of the ACM\/IEEE 42nd International Conference on Software Engineering. 949-960. https:\/\/doi.org\/10.1145\/3377811.3380331 10.1145\/3377811.3380331","DOI":"10.1145\/3377811.3380331"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643767","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3643767","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T07:57:54Z","timestamp":1770191874000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643767"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,12]]},"references-count":74,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2024,7,12]]}},"alternative-id":["10.1145\/3643767"],"URL":"https:\/\/doi.org\/10.1145\/3643767","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,12]]}}}