{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T19:58:22Z","timestamp":1784836702623,"version":"3.55.0"},"reference-count":84,"publisher":"Association for Computing Machinery (ACM)","issue":"ISSTA","funder":[{"DOI":"10.13039\/501100000781","name":"European Research Council","doi-asserted-by":"publisher","award":["949014"],"award-info":[{"award-number":["949014"]}],"id":[{"id":"10.13039\/501100000781","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,22]]},"abstract":"<jats:p>Audio classification systems, powered by deep neural networks (DNNs), are integral to various applications that impact daily lives, like voice-activated assistants. Ensuring the accuracy of these systems is crucial since inaccuracies can lead to significant security issues and user mistrust. However, testing audio classifiers presents a significant challenge: the high manual labeling cost for annotating audio test inputs. Test input prioritization has emerged as a promising approach to mitigate this labeling cost issue. It prioritizes potentially misclassified tests, allowing for the early labeling of such critical inputs and making debugging more efficient. However, when applying existing test prioritization methods to audio-type test inputs, there are some limitations: 1) Coverage-based methods are less effective and efficient than confidence-based methods. 2) Confidence-based methods rely only on prediction probability vectors, ignoring the unique characteristics of audio-type data. 3) Mutation-based methods lack designed mutation operations for audio data, making them unsuitable for audio-type test inputs. To overcome these challenges, we propose AudioTest, a novel test prioritization approach specifically designed for audio-type test inputs. The core premise is that tests closer to misclassified samples are more likely to be misclassified. Based on the special characteristics of audio-type data, AudioTest generates four types of features: time-domain features, frequency-domain features, perceptual features, and output features. For each test, AudioTest concatenates its four types of features into a feature vector and applies a carefully designed feature transformation strategy to bring misclassified tests closer in space. AudioTest leverages a trained model to predict the probability of misclassification of each test based on its transformed vectors and ranks all the tests accordingly. We evaluate the performance of AudioTest utilizing 96 subjects, encompassing natural and noisy datasets. We employed two classical metrics, Percentage of Fault Detection (PFD) and Average Percentage of Fault Detected (APFD), for our evaluation. The results demonstrate that AudioTest outperforms all the compared test prioritization approaches in terms of both PFD and APFD. The average improvement of AudioTest compared to the baseline test prioritization methods ranges from 12.63% to 54.58% on natural datasets and from 12.71% to 40.48% on noisy datasets.<\/jats:p>","DOI":"10.1145\/3728907","type":"journal-article","created":{"date-parts":[[2025,6,22]],"date-time":"2025-06-22T10:52:56Z","timestamp":1750589576000},"page":"707-730","source":"Crossref","is-referenced-by-count":2,"title":["AudioTest: Prioritizing Audio Test Cases"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1390-0393","authenticated-orcid":false,"given":"Yinghua","family":"Li","sequence":"first","affiliation":[{"name":"University of Luxembourg, Luxembourg, Luxembourg"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4097-9543","authenticated-orcid":false,"given":"Xueqi","family":"Dang","sequence":"additional","affiliation":[{"name":"University of Luxembourg, Luxembourg, Luxembourg"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-7312-6273","authenticated-orcid":false,"given":"Wendk\u00fbuni C.","family":"Ou\u00e9draogo","sequence":"additional","affiliation":[{"name":"University of Luxembourg, Luxembourg, Luxembourg"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4052-475X","authenticated-orcid":false,"given":"Jacques","family":"Klein","sequence":"additional","affiliation":[{"name":"University of Luxembourg, Luxembourg, Luxembourg"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7270-9869","authenticated-orcid":false,"given":"Tegawend\u00e9 F.","family":"Bissyand\u00e9","sequence":"additional","affiliation":[{"name":"University of Luxembourg, Luxembourg, Luxembourg"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,22]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2022.3223444"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1985793.1985795"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2024.3420816"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3236024.3236053"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2018.2889771"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394112"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_2_1_8_1","doi-asserted-by":"crossref","unstructured":"Yafeng Chen Siqi Zheng Hui Wang Luyao Cheng Qian Chen and Jiajun Qi. 2023. An enhanced res2net with local and global feature fusion for speaker verification. arXiv preprint arXiv:2305.12838.","DOI":"10.21437\/Interspeech.2023-1294"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-024-10515-y"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3607191"},{"key":"e_1_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Nauman Dawalatabad Mirco Ravanelli Fran\u00e7ois Grondin Jenthe Thienpondt Brecht Desplanques and Hwidong Na. 2021. ECAPA-TDNN embeddings for speaker diarization. arXiv preprint arXiv:2104.01466.","DOI":"10.21437\/Interspeech.2021-941"},{"key":"e_1_2_1_12_1","volume-title":"Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification. arXiv preprint arXiv:2005.07143.","author":"Desplanques Brecht","year":"2020","unstructured":"Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck. 2020. Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification. arXiv preprint arXiv:2005.07143."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICST.2013.27"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1002\/stvr.1572"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3576040"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2004.831663"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/32.988497"},{"key":"e_1_2_1_18_1","volume-title":"Chroma feature analysis and synthesis. Resources of laboratory for the recognition and organization of speech and Audio-LabROSA, 5","author":"Ellis Dan","year":"2007","unstructured":"Dan Ellis. 2007. Chroma feature analysis and synthesis. Resources of laboratory for the recognition and organization of speech and Audio-LabROSA, 5 (2007)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3395363.3397357"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683158"},{"key":"e_1_2_1_21_1","volume-title":"Xavier Favory, Jordi Pons, and Xavier Serra.","author":"Fonseca Eduardo","year":"2018","unstructured":"Eduardo Fonseca, Manoj Plakal, Frederic Font, Daniel PW Ellis, Xavier Favory, Jordi Pons, and Xavier Serra. 2018. General-purpose tagging of freesound audio with audioset labels: Task description, dataset, and baseline. arXiv preprint arXiv:1807.09902."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-021-11610-8"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952261"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639584"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-443-23814-7.00004-3"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3675888.3676028"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2648584.2648589"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-15919-0_19"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2884781.2884791"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952132"},{"key":"e_1_2_1_31_1","volume-title":"meeting of the American Educational Research Association. 1.","author":"Hess Melinda R","year":"2004","unstructured":"Melinda R Hess and Jeffrey D Kromrey. 2004. Robust confidence intervals for effect sizes: A comparative study of Cohen\u2019sd and Cliff\u2019s delta under non-normality and heterogeneous variances. In annual meeting of the American Educational Research Association. 1."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3460319.3464825"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICST46399.2020.00018"},{"key":"e_1_2_1_35_1","unstructured":"jiaaro. 2024. pydub. https:\/\/github.com\/jiaaro\/pydub\/tree\/master"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2002.1035731"},{"key":"e_1_2_1_37_1","unstructured":"R Juzenaite. 2019. Security vulnerabilities of voice recognition technologies."},{"key":"e_1_2_1_38_1","unstructured":"Kaggle. 2024. Kaggle - The Machine Learning and Data Science Community. https:\/\/www.kaggle.com"},{"key":"e_1_2_1_39_1","volume-title":"Examples are not enough, learn to criticize! criticism for interpretability. Advances in neural information processing systems, 29","author":"Kim Been","year":"2016","unstructured":"Been Kim, Rajiv Khanna, and Oluwasanmi O Koyejo. 2016. Examples are not enough, learn to criticize! criticism for interpretability. Advances in neural information processing systems, 29 (2016)."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019.00108"},{"key":"e_1_2_1_41_1","doi-asserted-by":"crossref","unstructured":"Anssi Klapuri and Manuel Davy. 2007. Signal processing methods for music transcription.","DOI":"10.1007\/0-387-32845-9"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2020.3030497"},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 28643\u201328652","author":"Lee Taeckyung","year":"2024","unstructured":"Taeckyung Lee, Sorn Chottananurak, Taesik Gong, and Sung-Ju Lee. 2024. AETTA: Label-Free Accuracy Estimation for Test-Time Adaptation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 28643\u201328652."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3643676"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2007.38"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3338906.3338930"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0196391"},{"key":"e_1_2_1_48_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3397320","article-title":"Vocallock: Sensing vocal tract for passphrase-independent user authentication leveraging acoustic signals on smartphones","volume":"4","author":"Lu Li","year":"2020","unstructured":"Li Lu, Jiadi Yu, Yingying Chen, and Yan Wang. 2020. Vocallock: Sensing vocal tract for passphrase-independent user authentication leveraging acoustic signals on smartphones. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 4, 2 (2020), 1\u201324.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3238147.3238202"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE.2018.00021"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3417330"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2021.101329"},{"key":"e_1_2_1_53_1","volume-title":"Matt McVicar, Eric Battenberg, and Oriol Nieto.","author":"McFee Brian","year":"2015","unstructured":"Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. 2015. librosa: Audio and music signal analysis in python. In SciPy. 18\u201324."},{"key":"e_1_2_1_54_1","doi-asserted-by":"crossref","unstructured":"Patrick E McKnight and Julius Najab. 2010. Mann-Whitney U Test. The Corsini encyclopedia of psychology 1\u20131.","DOI":"10.1002\/9780470479216.corpsy0524"},{"key":"e_1_2_1_55_1","unstructured":"Cu Nguyen Paolo Tonella Tanja Vos Nelly Condori Bilha Mendelson Daniel Citron and Onn Shehory. 2014. Test prioritization based on change sensitivity: an industrial case study. Technical Report Series\/Department of Information and Computing Sciences Utrecht University."},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1155\/2021\/4832864"},{"key":"e_1_2_1_57_1","volume-title":"2012 IEEE 11th International Conference on Signal Processing. 3, 1629\u20131632","author":"Ning Taikang","year":"2012","unstructured":"Taikang Ning, James Ning, Nikolay Atanasov, and Kai-Sheng Hsieh. 2012. A fast heart sounds detection and heart murmur classification algorithm. In 2012 IEEE 11th International Conference on Signal Processing. 3, 1629\u20131632."},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE.2015.7381799"},{"key":"e_1_2_1_59_1","volume-title":"Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, and Luca Antiga. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32 (2019)."},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3361566"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE51524.2021.9678764"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409730"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00104"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","unstructured":"Lior Rokach and Oded Maimon. 2005. Decision trees. Data mining and knowledge discovery handbook 165\u2013192. https:\/\/doi.org\/10.1007\/0-387-25465-X_9 10.1007\/0-387-25465-X_9","DOI":"10.1007\/0-387-25465-X_9"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2655045"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1002\/stvr.1695"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461375"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377811.3380353"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598109"},{"key":"e_1_2_1_70_1","doi-asserted-by":"crossref","unstructured":"Deepak-George Thomas Matteo Biagiola Nargiz Humbatova Mohammad Wardat Gunel Jahangirova Hridesh Rajan and Paolo Tonella. 2024. muPRL: A Mutation Testing Pipeline for Deep Reinforcement Learning based on Real Faults. arXiv preprint arXiv:2408.15150.","DOI":"10.1109\/ICSE55347.2025.00036"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-012-9219-7"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSM.2006.74"},{"key":"e_1_2_1_73_1","article-title":"Visualizing data using t-SNE","volume":"9","author":"der Maaten Laurens Van","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE.. Journal of machine learning research, 9, 11 (2008), http:\/\/jmlr.org\/papers\/v9\/vandermaaten08a.html","journal-title":"Journal of machine learning research"},{"key":"e_1_2_1_74_1","doi-asserted-by":"crossref","unstructured":"Hui Wang Siqi Zheng Yafeng Chen Luyao Cheng and Qian Chen. 2023. Cam++: A fast and efficient network for speaker verification using context-aware masking. arXiv preprint arXiv:2303.00332.","DOI":"10.21437\/Interspeech.2023-1513"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00046"},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/QRS57517.2022.00074"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/3533767.3534375"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-89960-2_22"},{"key":"e_1_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2023.107331"},{"key":"e_1_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1002\/stv.430"},{"key":"e_1_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1145\/1572272.1572296"},{"key":"e_1_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01526"},{"key":"e_1_2_1_83_1","unstructured":"Zifeng Zhao Ding Pan Junyi Peng and Rongzhi Gu. 2022. Probing Deep Speaker Embeddings for Speaker-related Tasks. arXiv preprint arXiv:2212.07068 arxiv:2212.07068"},{"key":"e_1_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1145\/3544792"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728907","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,16]],"date-time":"2025-07-16T16:56:25Z","timestamp":1752684985000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728907"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,22]]},"references-count":84,"journal-issue":{"issue":"ISSTA","published-print":{"date-parts":[[2025,6,22]]}},"alternative-id":["10.1145\/3728907"],"URL":"https:\/\/doi.org\/10.1145\/3728907","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,22]]}}}