{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T16:54:37Z","timestamp":1783184077243,"version":"3.54.6"},"reference-count":36,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,12,11]],"date-time":"2024-12-11T00:00:00Z","timestamp":1733875200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62202064"],"award-info":[{"award-number":["62202064"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["92270204"],"award-info":[{"award-number":["92270204"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"CAS Project for Young Scientists in Basic Research","award":["YSBR-118"],"award-info":[{"award-number":["YSBR-118"]}]},{"DOI":"10.13039\/501100004739","name":"Youth Innovation Promotion Association CAS","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004739","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Priv. Secur."],"published-print":{"date-parts":[[2025,2,28]]},"abstract":"<jats:p>The attacker can generate adversarial examples (AEs) to stealthily mislead automatic speech recognition (ASR) models, raising significant concerns about the security of intelligent voice control (IVC) devices. Existing adversarial attacks mainly generate AEs to mislead ASR models to output specific target English commands (e.g., open the door). However, it remains unknown whether AEs can be used to issue commands in other languages to attack commercial black-box ASR models.<\/jats:p>\n          <jats:p>\n            In this article, taking Chinese phrases (e.g., \u652f\u4ed8\u5b9d\u4ed8\u6b3e) and \u201cChinese\u2013English code-switching\u201d phrases (e.g., \u5173\u95edGPS) as the target commands, we propose adversarial attacks for commercial multilingual ASR models. In particular, if a multilingual speech recognition model can recognize Chinese and English, we call it a Chinese\u2013English speech recognition model. In English, the meaning of \u201c\u652f\u4ed8\u5b9d\u4ed8\u6b3e\u201d and \u201c\u5173\u95edGPS\u201d are \u201cAlipay payment\u201d and \u201cturn off GPS\u201d, respectively. In detail, we generate transferable AEs based on the open-sourced conventional DataTang Mandarin ASR model. Given 55 target commands, the success rate for generating AEs of them is up to 96% and 80% for Aliyun ASR API and Tencentyun ASR API, respectively. Our AEs can trigger actual attack actions on voice assistants (e.g., Apple Siri, Xiaomi Xiaoaitongxue) or spread malicious messages through ASR API services, while the target commands in the AEs are inaudible to human beings.\n            <jats:xref ref-type=\"fn\">\n              <jats:sup>1<\/jats:sup>\n            <\/jats:xref>\n            Finally, by analyzing the spectrum differences between benign audio clips and AEs, we propose a general defense against adversarial audio attacks.\n          <\/jats:p>\n          <jats:p\/>","DOI":"10.1145\/3701725","type":"journal-article","created":{"date-parts":[[2024,11,7]],"date-time":"2024-11-07T10:59:20Z","timestamp":1730977160000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Adversarial Attack and Defense for Commercial Black-box Chinese-English Speech Recognition Systems"],"prefix":"10.1145","volume":"28","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-1866-5828","authenticated-orcid":false,"given":"Xuejing","family":"Yuan","sequence":"first","affiliation":[{"name":"School of Cyberspace Security, Beijing University of Posts and Telecommunications, Beijing, China and Key Laboratory of Cyberspace Security Defense (Institute of Information Engineering, Chinese Academy of Sciences, Beijing 100085), Beijing China and School of Cyber Security, University of Chinese Academy of Sciences, China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7269-0011","authenticated-orcid":false,"given":"Jiangshan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, China, Beijing, China and School of Cyber Security, University of Chinese Academy of Sciences, China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5624-2987","authenticated-orcid":false,"given":"Kai","family":"Chen","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, China, Beijing, China and School of Cyber Security, University of Chinese Academy of Sciences, China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-7521-5164","authenticated-orcid":false,"given":"Cheng'an","family":"Wei","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, China, Beijing, China and School of Cyber Security, University of Chinese Academy of Sciences, China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7111-4772","authenticated-orcid":false,"given":"Ruiyuan","family":"Li","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, China, Beijing, China and School of Cyber Security, University of Chinese Academy of Sciences, China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2300-6308","authenticated-orcid":false,"given":"Zhenkun","family":"Ma","sequence":"additional","affiliation":[{"name":"OPPO ZIWU Cyber Security Lab in Shenzhen, China, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-2676-0013","authenticated-orcid":false,"given":"Xinqi","family":"Ling","sequence":"additional","affiliation":[{"name":"OPPO ZIWU Cyber Security Lab in Shenzhen, China, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,12,11]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/SPW.2018.00009"},{"key":"e_1_3_3_3_2","first-page":"49","volume-title":"Proceedings of the USENIX Security Symposium","author":"Yuan Xuejing","year":"2018","unstructured":"Xuejing Yuan, Yuxuan Chen, Yue Zhao, Yunhui Long, Xiaokang Liu, Kai Chen, Shengzhi Zhang, Heqing Huang, Xiaofeng Wang, and Carl A. Gunter. 2018. CommanderSong: A systematic approach for practical adversarial voice recognition. In Proceedings of the USENIX Security Symposium. 49\u201364."},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/741"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.14722\/ndss.2019.23288"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3427228.3427276"},{"key":"e_1_3_3_7_2","first-page":"5231","volume-title":"Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research)","volume":"97","author":"Qin Yao","year":"2019","unstructured":"Yao Qin, Nicholas Carlini, Garrison Cottrell, Ian Goodfellow, and Colin Raffel. 2019. Imperceptible, robust, and targeted adversarial examples for automatic speech recognition. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research), Vol. 97. PMLR, 5231\u20135240."},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3320269.3384733"},{"key":"e_1_3_3_9_2","first-page":"2667","volume-title":"Proceedings of the USENIX Security Symposium","author":"Chen Yuxuan","year":"2020","unstructured":"Yuxuan Chen, Xuejing Yuan, Jiangshan Zhang, Yue Zhao, Shengzhi Zhang, Kai Chen, and XiaoFeng Wang. 2020. Devil\u2019s whisper: A general approach for physical adversarial attacks against commercial black-box speech recognition devices.. In Proceedings of the USENIX Security Symposium. 2667\u20132684."},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460120.3485383"},{"key":"e_1_3_3_11_2","unstructured":"2023. CMUSphinx. Retrieved Dec. 28 2023 from https:\/\/cmusphinx.github.io\/"},{"key":"e_1_3_3_12_2","unstructured":"2024. Kaldi. Retrieved Oct. 4 2024 from https:\/\/github.com\/kaldi-asr\/kaldi"},{"key":"e_1_3_3_13_2","first-page":"173","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Amodei Dario","year":"2016","unstructured":"Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et\u00a0al. 2016. Deep speech 2: End-to-end speech recognition in English and Mandarin. In Proceedings of the International Conference on Machine Learning. PMLR, 173\u2013182."},{"key":"e_1_3_3_14_2","first-page":"12449","article-title":"WAV2VEC 2.0: A framework for self-supervised learning of speech representations","author":"Baevski Alexei","year":"2020","unstructured":"Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. WAV2VEC 2.0: A framework for self-supervised learning of speech representations. In Proceedings of the 34th International Conference on Neural Information Processing Systems. 12449\u201312460.","journal-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems"},{"key":"e_1_3_3_15_2","unstructured":"2021. wav2vec2-large-xlsr-53-chinese-zh-cn. Retrieved from https:\/\/huggingface.co\/jonatasgrosman\/wav2vec2-large-xlsr-53-chinese-zh-cn"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.14722\/ndss.2020.23055"},{"key":"e_1_3_3_17_2","first-page":"672","volume-title":"Proceedings of the Asian Conference on Machine Learning","author":"Abdullah Hadi","year":"2021","unstructured":"Hadi Abdullah, Muhammad Sajidur Rahman, Christian Peeters, Cassidy Gibson, Washington Garcia, Vincent Bindschaedler, Thomas Shrimpton, and Patrick Traynor. 2021. Beyond \\(L_p\\) clipping: Equalization based psychoacoustic attacks against ASRs. In Proceedings of the Asian Conference on Machine Learning. PMLR, 672\u2013688."},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3548606.3560660"},{"key":"e_1_3_3_19_2","article-title":"Advreverb: Rethinking the stealthiness of audio adversarial examples to human perception","author":"Chen Meng","year":"2023","unstructured":"Meng Chen, Li Lu, Jiadi Yu, Zhongjie Ba, Feng Lin, and Kui Ren. 2023. Advreverb: Rethinking the stealthiness of audio adversarial examples to human perception. IEEE Transactions on Information Forensics and Security 19 (2023), 1948\u20131962. https:\/\/ieeexplore.ieee.org\/document\/10368060","journal-title":"IEEE Transactions on Information Forensics and Security"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510582"},{"key":"e_1_3_3_21_2","first-page":"1","volume-title":"Proceedings of the 32th USENIX Security Symposium (USENIX Security\u201923)","author":"Yu Zhiyuan","year":"2023","unstructured":"Zhiyuan Yu, Yuanhaur Chang, Ning Zhang, and Chaowei Xiao. 2023. SMACK: Semantically meaningful adversarial audio attack. In Proceedings of the 32th USENIX Security Symposium (USENIX Security\u201923). 1\u201318."},{"key":"e_1_3_3_22_2","first-page":"247","volume-title":"Proceedings of the 32nd USENIX Security Symposium (USENIX Security\u201923)","author":"Wu Xinghui","year":"2023","unstructured":"Xinghui Wu, Shiqing Ma, Chao Shen, Chenhao Lin, Qian Wang, Qi Li, and Yuan Rao. 2023. \\(\\lbrace\\) KENKU \\(\\rbrace\\) : Towards efficient and stealthy black-box adversarial attacks against \\(\\lbrace\\) ASR \\(\\rbrace\\) systems. In Proceedings of the 32nd USENIX Security Symposium (USENIX Security\u201923). 247\u2013264."},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3607199.3607240"},{"key":"e_1_3_3_24_2","first-page":"2273","volume-title":"Proceedings of the 30th USENIX Security Symposium (USENIX Security\u201921)","author":"Hussain Shehzeen","year":"2021","unstructured":"Shehzeen Hussain, Paarth Neekhara, Shlomo Dubnov, Julian McAuley, and Farinaz Koushanfar. 2021. \\(\\lbrace\\) WaveGuard \\(\\rbrace\\) : Understanding and mitigating audio adversarial examples. In Proceedings of the 30th USENIX Security Symposium (USENIX Security\u201921). 2273\u20132290."},{"key":"e_1_3_3_25_2","first-page":"2309","volume-title":"Proceedings of the 30th USENIX Security Symposium (USENIX Security\u201921)","author":"Eisenhofer Thorsten","year":"2021","unstructured":"Thorsten Eisenhofer, Lea Sch\u00f6nherr, Joel Frank, Lars Speckemeier, Dorothea Kolossa, and Thorsten Holz. 2021. Dompteur: Taming audio adversarial examples. In Proceedings of the 30th USENIX Security Symposium (USENIX Security\u201921). 2309\u20132326."},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR48806.2021.9413263"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2019.00019"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1186\/s42400-023-00177-6"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3527153"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3380991"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460120.3484755"},{"key":"e_1_3_3_32_2","unstructured":"2024. Aliyun ASR service. Retrieved from https:\/\/ai.aliyun.com\/nls\/asr"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/SP54263.2024.00056"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3122291"},{"key":"e_1_3_3_35_2","unstructured":"2019. Kaldi models. Retrieved Jun. 8 2019 from https:\/\/kaldi-asr.org\/models.html"},{"key":"e_1_3_3_36_2","unstructured":"2024. wiki. Mean Opinion Score. Retrieved Feb. 23 2024 from https:\/\/en.wikipedia.org\/wiki\/Mean_opinion_score"},{"key":"e_1_3_3_37_2","unstructured":"OpenSLR. 2024. Retrieved Oct 16 2024 from http:\/\/openslr.org\/resources.php"}],"container-title":["ACM Transactions on Privacy and Security"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3701725","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3701725","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:57:16Z","timestamp":1750298236000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3701725"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,11]]},"references-count":36,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,2,28]]}},"alternative-id":["10.1145\/3701725"],"URL":"https:\/\/doi.org\/10.1145\/3701725","relation":{},"ISSN":["2471-2566","2471-2574"],"issn-type":[{"value":"2471-2566","type":"print"},{"value":"2471-2574","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,11]]},"assertion":[{"value":"2023-10-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}