{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T16:54:39Z","timestamp":1783184079467,"version":"3.54.6"},"reference-count":165,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2022,5,19]],"date-time":"2022-05-19T00:00:00Z","timestamp":1652918400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Priv. Secur."],"published-print":{"date-parts":[[2022,8,31]]},"abstract":"<jats:p>With the wide use of Automatic Speech Recognition (ASR) in applications such as human machine interaction, simultaneous interpretation, audio transcription, and so on, its security protection becomes increasingly important. Although recent studies have brought to light the weaknesses of popular ASR systems that enable out-of-band signal attack, adversarial attack, and so on, and further proposed various remedies (signal smoothing, adversarial training, etc.), a systematic understanding of ASR security (both attacks and defenses) is still missing, especially on how realistic such threats are and how general existing protection could be. In this article, we present our systematization of knowledge for ASR security and provide a comprehensive taxonomy for existing work based on a modularized workflow. More importantly, we align the research in this domain with that on security in Image Recognition System (IRS), which has been extensively studied, using the domain knowledge in the latter to help understand where we stand in the former. Generally, both IRS and ASR are perceptual systems. Their similarities allow us to systematically study existing literature in ASR security based on the spectrum of attacks and defense solutions proposed for IRS, and pinpoint the directions of more advanced attacks and the directions potentially leading to more effective protection in ASR. In contrast, their differences, especially the complexity of ASR compared with IRS, help us learn unique challenges and opportunities in ASR security. Particularly, our experimental study shows that transfer attacks across ASR models are feasible, even in the absence of knowledge about models (even their types) and training data.<\/jats:p>","DOI":"10.1145\/3510582","type":"journal-article","created":{"date-parts":[[2022,3,29]],"date-time":"2022-03-29T11:39:29Z","timestamp":1648553969000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":23,"title":["SoK: A Modularized Approach to Study the Security of Automatic Speech Recognition Systems"],"prefix":"10.1145","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0000-5031","authenticated-orcid":false,"given":"Yuxuan","family":"Chen","sequence":"first","affiliation":[{"name":"School of Cyber Science and Technology, Shandong University, Qingdao, Shandong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiangshan","family":"Zhang","sequence":"additional","affiliation":[{"name":"SKLOIS, Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences, Haidian, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xuejing","family":"Yuan","sequence":"additional","affiliation":[{"name":"SKLOIS, Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences, Haidian, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shengzhi","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Metropolitan College, Boston University, Boston, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kai","family":"Chen","sequence":"additional","affiliation":[{"name":"SKLOIS, Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences, Haidian, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaofeng","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Informatics, Computing and Engineering, Indiana University Bloomington, Bloomington, IN, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shanqing","family":"Guo","sequence":"additional","affiliation":[{"name":"School of Cyber Science and Technology, Shandong University, Qingdao, Shandong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,5,19]]},"reference":[{"key":"e_1_3_2_2_2","article-title":"Universal adversarial audio perturbations","author":"Abdoli Sajjad","year":"2019","unstructured":"Sajjad Abdoli, Luiz G. Hafemann, Jerome Rony, Ismail Ben Ayed, Patrick Cardinal, and Alessandro L. Koerich. 2019. Universal adversarial audio perturbations. IEEE Trans. Pattern Anal. Mach. Intell. (2019).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_3_2","doi-asserted-by":"crossref","unstructured":"Hadi Abdullah Washington Garcia Christian Peeters Patrick Traynor Kevin R. B. Butler and Joseph Wilson. 2019. Practical hidden voice attacks against speech and speaker recognition systems. In Proceedings of the 26th Annual Network and Distributed System Security Symposium (NDSS) .","DOI":"10.14722\/ndss.2019.23362"},{"key":"e_1_3_2_4_2","article-title":"Hear \u201cNo Evil,\u201d see \u201cKenansville\u201d: Efficient and transferable black-box attacks on speech recognition and voice identification systems","author":"Abdullah Hadi","year":"2021","unstructured":"Hadi Abdullah, Muhammad Sajidur Rahman, Washington Garcia, Logan Blue, Kevin Warren, Anurag Swarnim Yadav, Tom Shrimpton, and Patrick Traynor. 2021. Hear \u201cNo Evil,\u201d see \u201cKenansville\u201d: Efficient and transferable black-box attacks on speech recognition and voice identification systems. In 42nd IEEE Symposium on Security and Privacy.","journal-title":"42nd IEEE Symposium on Security and Privacy"},{"key":"e_1_3_2_5_2","article-title":"Beyond  \\( L\\_p \\)  clipping: Equalization-based psychoacoustic attacks against ASRs","author":"Abdullah Hadi","year":"2021","unstructured":"Hadi Abdullah, Muhammad Sajidur Rahman, Christian Peeters, Cassidy Gibson, Washington Garcia, Vincent Bindschaedler, Thomas Shrimpton, and Patrick Traynor. 2021. Beyond \\( L\\_p \\) clipping: Equalization-based psychoacoustic attacks against ASRs. arXiv preprint arXiv:2110.13250 (2021).","journal-title":"arXiv preprint arXiv:2110.13250"},{"key":"e_1_3_2_6_2","article-title":"SoK: The faults in our ASRs: An overview of attacks against automatic speech recognition and speaker identification systems","author":"Abdullah Hadi","year":"2021","unstructured":"Hadi Abdullah, Kevin Warren, Vincent Bindschaedler, Nicolas Papernot, and Patrick Traynor. 2021. SoK: The faults in our ASRs: An overview of attacks against automatic speech recognition and speaker identification systems. In 42nd IEEE Symposium on Security and Privacy.","journal-title":"42nd IEEE Symposium on Security and Privacy"},{"key":"e_1_3_2_7_2","article-title":"Identifying audio adversarial examples via anomalous pattern detection","author":"Akinwande Victor","year":"2020","unstructured":"Victor Akinwande, Celia Cintas, Skyler Speakman, and Srihari Sridharan. 2020. Identifying audio adversarial examples via anomalous pattern detection. arXiv preprint arXiv:2002.05463 (2020).","journal-title":"arXiv preprint arXiv:2002.05463"},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1145\/3137003.3137014","volume-title":"1st ACM Workshop on the Internet of Safe Things","author":"Alanwar Amr","year":"2017","unstructured":"Amr Alanwar, Bharathan Balaji, Yuan Tian, Shuo Yang, and Mani Srivastava. 2017. EchoSafe: Sonar-based verifiable interaction with intelligent digital agents. In 1st ACM Workshop on the Internet of Safe Things. 38\u201343."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2017.2747626"},{"key":"e_1_3_2_10_2","article-title":"Did you hear that? Adversarial examples against automatic speech recognition","author":"Alzantot Moustafa","year":"2017","unstructured":"Moustafa Alzantot, Bharathan Balaji, and Mani Srivastava. 2017. Did you hear that? Adversarial examples against automatic speech recognition. In NIPS 2017 Machine Deception Workshop.","journal-title":"NIPS 2017 Machine Deception Workshop"},{"key":"e_1_3_2_11_2","first-page":"173","volume-title":"International Conference on Machine Learning","author":"Amodei Dario","year":"2016","unstructured":"Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et\u00a0al. 2016. Deep Speech 2: End-to-end speech recognition in English and Mandarin. In International Conference on Machine Learning. 173\u2013182."},{"key":"e_1_3_2_12_2","first-page":"284","volume-title":"International Conference on Machine Learning","author":"Athalye Anish","year":"2018","unstructured":"Anish Athalye, Logan Engstrom, Andrew Ilyas, and Kevin Kwok. 2018. Synthesizing robust adversarial examples. In International Conference on Machine Learning. PMLR, 284\u2013293."},{"key":"e_1_3_2_13_2","volume-title":"Speech Enhancement","author":"Benesty Jacob","year":"2006","unstructured":"Jacob Benesty, Shoji Makino, and Jingdong Chen. 2006. Speech Enhancement. Springer Science & Business Media."},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","unstructured":"Mary K. Bispham Ioannis Agrafiotis and Michael Goldsmith. 2019. Nonsense attacks on Google Assistant and missense attacks on Amazon Alexa. (2019).","DOI":"10.5220\/0007309500750087"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3196494.3196545"},{"key":"e_1_3_2_16_2","first-page":"1048","volume-title":"IEEE Symposium on Security and Privacy (SP)","author":"Bolton Connor","year":"2018","unstructured":"Connor Bolton, Sara Rampazzi, Chaohao Li, Andrew Kwong, Wenyuan Xu, and Kevin Fu. 2018. Blue Note: How intentional acoustic interference damages availability and integrity in hard disk drives and operating systems. In IEEE Symposium on Security and Privacy (SP). IEEE, 1048\u20131062."},{"key":"e_1_3_2_17_2","article-title":"Small input noise is enough to defend against query-based black-box attacks","author":"Byun Junyoung","year":"2021","unstructured":"Junyoung Byun, Hyojun Go, and Changick Kim. 2021. Small input noise is enough to defend against query-based black-box attacks. arXiv preprint arXiv:2101.04829 (2021).","journal-title":"arXiv preprint arXiv:2101.04829"},{"key":"e_1_3_2_18_2","first-page":"176","volume-title":"IEEE Symposium on Security and Privacy (SP)","author":"Cao Yulong","year":"2021","unstructured":"Yulong Cao, Ningfei Wang, Chaowei Xiao, Dawei Yang, Jin Fang, Ruigang Yang, Qi Alfred Chen, Mingyan Liu, and Bo Li. 2021. Invisible for both camera and lidar: Security of multi-sensor fusion based perception in autonomous driving under physical-world attacks. In IEEE Symposium on Security and Privacy (SP). IEEE, 176\u2013194."},{"key":"e_1_3_2_19_2","article-title":"Are you (Google) home? Detecting users\u2019 presence through traffic analysis of smart speakers","author":"Caputo D.","year":"2020","unstructured":"D. Caputo, L. Verderame, A. Merlo, A. Ranieri, and L. Caviglione. 2020. Are you (Google) home? Detecting users\u2019 presence through traffic analysis of smart speakers. ITASEC 2020.","journal-title":"ITASEC 2020"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/78.553476"},{"key":"e_1_3_2_21_2","first-page":"513","volume-title":"25th USENIX Security Symposium (USENIX Security\u201916)","author":"Carlini Nicholas","year":"2016","unstructured":"Nicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang, Micah Sherr, Clay Shields, David Wagner, and Wenchao Zhou. 2016. Hidden voice commands. In 25th USENIX Security Symposium (USENIX Security\u201916). 513\u2013530."},{"key":"e_1_3_2_22_2","first-page":"39","volume-title":"IEEE Symposium on Security and Privacy (SP)","author":"Carlini Nicholas","year":"2017","unstructured":"Nicholas Carlini and David Wagner. 2017. Towards evaluating the robustness of neural networks. In IEEE Symposium on Security and Privacy (SP). IEEE, 39\u201357."},{"key":"e_1_3_2_23_2","first-page":"1","volume-title":"IEEE Security and Privacy Workshops (SPW)","author":"Carlini Nicholas","year":"2018","unstructured":"Nicholas Carlini and David Wagner. 2018. Audio adversarial examples: Targeted attacks on speech-to-text. In IEEE Security and Privacy Workshops (SPW). IEEE, 1\u20137."},{"key":"e_1_3_2_24_2","unstructured":"Lucy Chai Thavishi Illandara and Zhongxia Yan. [n.d.]. Private speech adversaries. ([n. d.])."},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASP-DAC47756.2020.9045597"},{"key":"e_1_3_2_26_2","article-title":"Who is real Bob? Adversarial attacks on speaker recognition systems","author":"Chen Guangke","year":"2019","unstructured":"Guangke Chen, Sen Chen, Lingling Fan, Xiaoning Du, Zhe Zhao, Fu Song, and Yang Liu. 2019. Who is real Bob? Adversarial attacks on speaker recognition systems. arXiv preprint arXiv:1911.01840 (2019).","journal-title":"arXiv preprint arXiv:1911.01840"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2005.860851"},{"key":"e_1_3_2_28_2","first-page":"15","volume-title":"10th ACM Workshop on Artificial Intelligence and Security","author":"Chen Pin-Yu","year":"2017","unstructured":"Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh. 2017. Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. In 10th ACM Workshop on Artificial Intelligence and Security. 15\u201326."},{"key":"e_1_3_2_29_2","doi-asserted-by":"crossref","first-page":"30","DOI":"10.1145\/3385003.3410925","volume-title":"1st ACM Workshop on Security and Privacy on Artificial Intelligence","author":"Chen Steven","year":"2020","unstructured":"Steven Chen, Nicholas Carlini, and David Wagner. 2020. Stateful detection of black-box adversarial attacks. In 1st ACM Workshop on Security and Privacy on Artificial Intelligence. 30\u201339."},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS.2017.133"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/83.806630"},{"key":"e_1_3_2_32_2","article-title":"Metamorph: Injecting inaudible commands into over-the-air voice controlled systems","author":"Chen Tao","year":"2020","unstructured":"Tao Chen, Longfei Shangguan, Zhenjiang Li, and Kyle Jamieson. 2020. Metamorph: Injecting inaudible commands into over-the-air voice controlled systems. NDSS (2020).","journal-title":"NDSS"},{"key":"e_1_3_2_33_2","article-title":"Understanding the effectiveness of ultrasonic microphone jammer","author":"Chen Yuxin","year":"2019","unstructured":"Yuxin Chen, Huiying Li, Steven Nagels, Zhijing Li, Pedro Lopes, Ben Y. Zhao, and Haitao Zheng. 2019. Understanding the effectiveness of ultrasonic microphone jammer. arXiv preprint arXiv:1904.08490 (2019).","journal-title":"arXiv preprint arXiv:1904.08490"},{"key":"e_1_3_2_34_2","volume-title":"29th USENIX Security Symposium (USENIX Security\u201920)","author":"Chen Yuxuan","year":"2020","unstructured":"Yuxuan Chen, Xuejing Yuan, Jiangshan Zhang, Yue Zhao, Shengzhi Zhang, Kai Chen, and XiaoFeng Wang. 2020. Devil\u2019s Whisper: A general approach for physical adversarial attacks against commercial black-box speech recognition devices. In 29th USENIX Security Symposium (USENIX Security\u201920)."},{"key":"e_1_3_2_35_2","article-title":"Houdini: Fooling deep structured prediction models","author":"Cisse Moustapha","year":"2017","unstructured":"Moustapha Cisse, Yossi Adi, Natalia Neverova, and Joseph Keshet. 2017. Houdini: Fooling deep structured prediction models. NIPS (2017).","journal-title":"NIPS"},{"key":"e_1_3_2_36_2","article-title":"Wav2Letter: An end-to-end convNet-based speech recognition system","author":"Collobert Ronan","year":"2016","unstructured":"Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve. 2016. Wav2Letter: An end-to-end convNet-based speech recognition system. arXiv preprint arXiv:1609.03193 (2016).","journal-title":"arXiv preprint arXiv:1609.03193"},{"key":"e_1_3_2_37_2","first-page":"677","volume-title":"Joint European Conference on Machine Learning and Knowledge Discovery in Databases","author":"Das Nilaksh","year":"2018","unstructured":"Nilaksh Das, Madhuri Shanbhogue, Shang-Tse Chen, Li Chen, Michael E. Kounavis, and Duen Horng Chau. 2018. Adagio: Interactive experimentation with adversarial attack and defense for audio. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, 677\u2013681."},{"key":"e_1_3_2_38_2","first-page":"321","volume-title":"28th USENIX Security Symposium (USENIX Security\u201919)","author":"Demontis Ambra","year":"2019","unstructured":"Ambra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski, Battista Biggio, Alina Oprea, Cristina Nita-Rotaru, and Fabio Roli. 2019. Why do adversarial attacks transfer? Explaining transferability of evasion and poisoning attacks. In 28th USENIX Security Symposium (USENIX Security\u201919). 321\u2013338."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/NSSMIC.1993.373563"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00957"},{"key":"e_1_3_2_41_2","article-title":"SirenAttack: Generating adversarial audio for end-to-end acoustic systems","author":"Du Tianyu","year":"2020","unstructured":"Tianyu Du, Shouling Ji, Jinfeng Li, Qinchen Gu, Ting Wang, and Raheem Beyah. 2020. SirenAttack: Generating adversarial audio for end-to-end acoustic systems. ASIACCS (2020).","journal-title":"ASIACCS"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683793"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00175"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3117811.3117823"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3176402"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICMLA.2019.00167"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.5555\/3367471.3367693"},{"key":"e_1_3_2_48_2","article-title":"Crafting adversarial examples for speech paralinguistics applications","author":"Gong Yuan","year":"2018","unstructured":"Yuan Gong and Christian Poellabauer. 2018. Crafting adversarial examples for speech paralinguistics applications. DYnamic and Novel Advances in Machine Learning and Intelligent Cyber Security (DYNAMICS) Workshop.","journal-title":"DYnamic and Novel Advances in Machine Learning and Intelligent Cyber Security (DYNAMICS) Workshop"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCCN.2018.8487334"},{"key":"e_1_3_2_50_2","article-title":"Explaining and harnessing adversarial examples","author":"Goodfellow Ian J.","year":"2014","unstructured":"Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572 (2014).","journal-title":"arXiv preprint arXiv:1412.6572"},{"key":"e_1_3_2_51_2","article-title":"INOR\u2014An intelligent noise reduction method to defend against adversarial audio examples","author":"Guo Qingli","year":"2020","unstructured":"Qingli Guo, Jing Ye, Yiran Chen, Yu Hu, Yazhu Lan, Guohe Zhang, and Xiaowei Li. 2020. INOR\u2014An intelligent noise reduction method to defend against adversarial audio examples. Neurocomputing (2020).","journal-title":"Neurocomputing"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2985231"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3319535.3363264"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1002\/0471461288"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3300061.3345429"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00483"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM.2014.6847959"},{"key":"e_1_3_2_58_2","volume-title":"30th USENIX Security Symposium (USENIX Security\u201921)","author":"Hussain Shehzeen","year":"2021","unstructured":"Shehzeen Hussain, Paarth Neekhara, Shlomo Dubnov, Julian McAuley, and Farinaz Koushanfar. 2021. WaveGuard: Understanding and mitigating audio adversarial examples. In 30th USENIX Security Symposium (USENIX Security\u201921)."},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/2660267.2660295"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/TEMC.2015.2463089"},{"key":"e_1_3_2_61_2","first-page":"3208","article-title":"Adversarial black-box attacks on automatic speech recognition systems using multi-objective evolutionary optimization","author":"Khare Shreya","year":"2019","unstructured":"Shreya Khare, Rahul Aralikatte, and Senthil Mani. 2019. Adversarial black-box attacks on automatic speech recognition systems using multi-objective evolutionary optimization. In Interspeech Conference. 3208\u20133212.","journal-title":"Interspeech Conference"},{"key":"e_1_3_2_62_2","article-title":"Adversarial audio: A new information hiding method and backdoor for DNN-based speech recognition models","author":"Kong Yehao","year":"2019","unstructured":"Yehao Kong and Jiliang Zhang. 2019. Adversarial audio: A new information hiding method and backdoor for DNN-based speech recognition models. arXiv preprint arXiv:1904.03829 (2019).","journal-title":"arXiv preprint arXiv:1904.03829"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8462693"},{"key":"e_1_3_2_64_2","first-page":"33","volume-title":"27th USENIX Security Symposium (USENIX Security\u201918)","author":"Kumar Deepak","year":"2018","unstructured":"Deepak Kumar, Riccardo Paccagnella, Paul Murley, Eric Hennenfent, Joshua Mason, Adam Bates, and Michael Bailey. 2018. Skill squatting attacks on amazon alexa. In 27th USENIX Security Symposium (USENIX Security\u201918). 33\u201347."},{"key":"e_1_3_2_65_2","article-title":"Adversarial examples in the physical world","author":"Kurakin Alexey","year":"2016","unstructured":"Alexey Kurakin, Ian Goodfellow, and Samy Bengio. 2016. Adversarial examples in the physical world. arXiv preprint arXiv:1607.02533 (2016).","journal-title":"arXiv preprint arXiv:1607.02533"},{"key":"e_1_3_2_66_2","article-title":"Adversarial machine learning at scale","author":"Kurakin Alexey","year":"2016","unstructured":"Alexey Kurakin, Ian Goodfellow, and Samy Bengio. 2016. Adversarial machine learning at scale. arXiv preprint arXiv:1611.01236 (2016).","journal-title":"arXiv preprint arXiv:1611.01236"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2019.2925452"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1145\/3319535.3363246"},{"issue":"5","key":"e_1_3_2_69_2","first-page":"14","article-title":"LeNet-5, convolutional neural networks","volume":"20","author":"LeCun Yann","year":"2015","unstructured":"Yann LeCun et\u00a0al. 2015. LeNet-5, convolutional neural networks. 20, 5 (2015), 14. Retrieved from http:\/\/yann.lecun.com\/exdb\/lenet.","journal-title":"R"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3380991"},{"key":"e_1_3_2_71_2","article-title":"The insecurity of home digital voice assistants\u2013Amazon Alexa as a case study","author":"Lei Xinyu","year":"2017","unstructured":"Xinyu Lei, Guan-Hua Tu, Alex X. Liu, Kamran Ali, Chi-Yu Li, and Tian Xie. 2017. The insecurity of home digital voice assistants\u2013Amazon Alexa as a case study. arXiv preprint arXiv:1712.03327 (2017).","journal-title":"arXiv preprint arXiv:1712.03327"},{"key":"e_1_3_2_72_2","first-page":"11908","volume-title":"Conference on Advances in Neural Information Processing Systems","author":"Li Juncheng","year":"2019","unstructured":"Juncheng Li, Shuhui Qu, Xinjian Li, Joseph Szurley, J. Zico Kolter, and Florian Metze. 2019. Adversarial music: Real world audio adversary against wake-word detection system. In Conference on Advances in Neural Information Processing Systems. 11908\u201311918."},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053076"},{"key":"e_1_3_2_74_2","first-page":"9","volume-title":"21st International Workshop on Mobile Computing Systems and Applications","author":"Li Zhuohang","year":"2020","unstructured":"Zhuohang Li, Cong Shi, Yi Xie, Jian Liu, Bo Yuan, and Yingying Chen. 2020. Practical adversarial attacks against speaker recognition systems. In 21st International Workshop on Mobile Computing Systems and Applications. 9\u201314."},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/3372297.3423348"},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-41278-3_43"},{"key":"e_1_3_2_77_2","article-title":"Adversarial attacks on spoofing countermeasures of automatic speaker verification","author":"Liu Songxiang","year":"2019","unstructured":"Songxiang Liu, Haibin Wu, Hung-Yi Lee, and Helen Meng. 2019. Adversarial attacks on spoofing countermeasures of automatic speaker verification. ASRU (2019).","journal-title":"ASRU"},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5928"},{"key":"e_1_3_2_79_2","article-title":"Delving into transferable adversarial examples and black-box attacks","author":"Liu Yanpei","year":"2016","unstructured":"Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Song. 2016. Delving into transferable adversarial examples and black-box attacks. arXiv preprint arXiv:1611.02770 (2016).","journal-title":"arXiv preprint arXiv:1611.02770"},{"key":"e_1_3_2_80_2","article-title":"Towards deep learning models resistant to adversarial attacks","author":"Madry Aleksander","year":"2017","unstructured":"Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2017. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083 (2017).","journal-title":"arXiv preprint arXiv:1706.06083"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/3133956.3134057"},{"key":"e_1_3_2_82_2","first-page":"81","volume-title":"18th ACM International Symposium on Mobile Ad Hoc Networking and Computing","author":"Meng Yan","year":"2018","unstructured":"Yan Meng, Zichang Wang, Wei Zhang, Peilin Wu, Haojin Zhu, Xiaohui Liang, and Yao Liu. 2018. WiVo: Enhancing the security of voice control system via wireless signal in IoT environment. In 18th ACM International Symposium on Mobile Ad Hoc Networking and Computing. 81\u201390."},{"key":"e_1_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1145\/3321705.3329842"},{"key":"e_1_3_2_84_2","article-title":"Voice recognition algorithms using Mel Frequency Cepstral Coefficient (MFCC) and Dynamic Time Warping (DTW) techniques","author":"Muda Lindasalwa","year":"2010","unstructured":"Lindasalwa Muda, Mumtaj Begam, and Irraivan Elamvazuthi. 2010. Voice recognition algorithms using Mel Frequency Cepstral Coefficient (MFCC) and Dynamic Time Warping (DTW) techniques. arXiv preprint arXiv:1003.4083 (2010).","journal-title":"arXiv preprint arXiv:1003.4083"},{"key":"e_1_3_2_85_2","first-page":"161","volume-title":"10th ISCA Speech Synthesis Workshop","author":"Nakamura Taiki","unstructured":"Taiki Nakamura, Yuki Saito, Shinnosuke Takamichi, Yusuke Ijima, and Hiroshi Saruwatari. [n.d.]. V2S attack: Building DNN-based voice conversion from automatic speaker verification. In 10th ISCA Speech Synthesis Workshop. 161\u2013165."},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1145\/3196494.3196506"},{"key":"e_1_3_2_87_2","first-page":"481","article-title":"Universal adversarial perturbations for speech recognition systems","author":"Neekhara Paarth","year":"2019","unstructured":"Paarth Neekhara, Shehzeen Hussain, Prakhar Pandey, Shlomo Dubnov, Julian McAuley, and Farinaz Koushanfar. 2019. Universal adversarial perturbations for speech recognition systems. In Interspeech Conference. 481\u2013485.","journal-title":"Interspeech Conference"},{"key":"e_1_3_2_88_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00772"},{"key":"e_1_3_2_89_2","doi-asserted-by":"publisher","DOI":"10.1145\/3372297.3417253"},{"key":"e_1_3_2_90_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403241"},{"key":"e_1_3_2_91_2","first-page":"4579","volume-title":"Conference on Advances in Neural Information Processing Systems","author":"Pang Tianyu","year":"2018","unstructured":"Tianyu Pang, Chao Du, Yinpeng Dong, and Jun Zhu. 2018. Towards robust detection of adversarial examples. In Conference on Advances in Neural Information Processing Systems. 4579\u20134589."},{"key":"e_1_3_2_92_2","article-title":"Transferability in machine learning: From phenomena to black-box attacks using adversarial samples","author":"Papernot Nicolas","year":"2016","unstructured":"Nicolas Papernot, Patrick McDaniel, and Ian Goodfellow. 2016. Transferability in machine learning: From phenomena to black-box attacks using adversarial samples. arXiv preprint arXiv:1605.07277 (2016).","journal-title":"arXiv preprint arXiv:1605.07277"},{"key":"e_1_3_2_93_2","first-page":"506","volume-title":"ACM on Asia conference on Computer and Communications Security","author":"Papernot Nicolas","year":"2017","unstructured":"Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, Somesh Jha, Z. Berkay Celik, and Ananthram Swami. 2017. Practical black-box attacks against machine learning. In ACM on Asia conference on Computer and Communications Security. 506\u2013519."},{"key":"e_1_3_2_94_2","doi-asserted-by":"crossref","first-page":"372","DOI":"10.1109\/EuroSP.2016.36","volume-title":"IEEE European Symposium on Security and Privacy (EuroS&P)","author":"Papernot Nicolas","year":"2016","unstructured":"Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z. Berkay Celik, and Ananthram Swami. 2016. The limitations of deep learning in adversarial settings. In IEEE European Symposium on Security and Privacy (EuroS&P). IEEE, 372\u2013387."},{"key":"e_1_3_2_95_2","first-page":"582","volume-title":"IEEE Symposium on Security and Privacy (SP)","author":"Papernot Nicolas","year":"2016","unstructured":"Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami. 2016. Distillation as a defense to adversarial perturbations against deep neural networks. In IEEE Symposium on Security and Privacy (SP). IEEE, 582\u2013597."},{"key":"e_1_3_2_96_2","article-title":"Imperceptible, robust, and targeted adversarial examples for automatic speech recognition","author":"Qin Yao","year":"2019","unstructured":"Yao Qin, Nicholas Carlini, Ian Goodfellow, Garrison Cottrell, and Colin Raffel. 2019. Imperceptible, robust, and targeted adversarial examples for automatic speech recognition. ICML (2019).","journal-title":"ICML"},{"key":"e_1_3_2_97_2","volume-title":"29th USENIX Security Symposium (USENIX Security\u201920)","author":"Quiring Erwin","year":"2020","unstructured":"Erwin Quiring, David Klein, Daniel Arp, Martin Johns, and Konrad Rieck. 2020. Adversarial preprocessing: Understanding and preventing image-scaling attacks in machine learning. In 29th USENIX Security Symposium (USENIX Security\u201920)."},{"key":"e_1_3_2_98_2","first-page":"197","volume-title":"IEEE International Symposium on Signal Processing and Information Technology (ISSPIT)","author":"Rajaratnam Krishan","year":"2018","unstructured":"Krishan Rajaratnam and Jugal Kalita. 2018. Noise flooding for detecting audio adversarial examples against automatic speech recognition. In IEEE International Symposium on Signal Processing and Information Technology (ISSPIT). IEEE, 197\u2013201."},{"key":"e_1_3_2_99_2","first-page":"16","volume-title":"30th Conference on Computational Linguistics and Speech Processing (ROCLING\u201918)","author":"Rajaratnam Krishan","year":"2018","unstructured":"Krishan Rajaratnam, Kunal Shah, and Jugal Kalita. 2018. Isolated and ensemble audio preprocessing methods for detecting adversarial examples against automatic speech recognition. In 30th Conference on Computational Linguistics and Speech Processing (ROCLING\u201918). 16\u201330."},{"key":"e_1_3_2_100_2","article-title":"Improving the adversarial robustness and interpretability of deep neural networks by regularizing their input gradients","author":"Ross Andrew Slavin","year":"2017","unstructured":"Andrew Slavin Ross and Finale Doshi-Velez. 2017. Improving the adversarial robustness and interpretability of deep neural networks by regularizing their input gradients. arXiv preprint arXiv:1711.09404 (2017).","journal-title":"arXiv preprint arXiv:1711.09404"},{"key":"e_1_3_2_101_2","doi-asserted-by":"publisher","DOI":"10.1145\/3081333.3081366"},{"key":"e_1_3_2_102_2","first-page":"547","volume-title":"15th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201918)","author":"Roy Nirupam","year":"2018","unstructured":"Nirupam Roy, Sheng Shen, Haitham Hassanieh, and Romit Roy Choudhury. 2018. Inaudible voice commands: The long-range attack and defense. In 15th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201918). 547\u2013560."},{"key":"e_1_3_2_103_2","article-title":"Adversarial manipulation of deep representations","author":"Sabour Sara","year":"2015","unstructured":"Sara Sabour, Yanshuai Cao, Fartash Faghri, and David J. Fleet. 2015. Adversarial manipulation of deep representations. arXiv preprint arXiv:1511.05122 (2015).","journal-title":"arXiv preprint arXiv:1511.05122"},{"key":"e_1_3_2_104_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9054750"},{"key":"e_1_3_2_105_2","article-title":"Adversarial attacks against automatic speech recognition systems via psychoacoustic hiding","author":"Sch\u00f6nherr Lea","year":"2019","unstructured":"Lea Sch\u00f6nherr, Katharina Kohls, Steffen Zeiler, Thorsten Holz, and Dorothea Kolossa. 2019. Adversarial attacks against automatic speech recognition systems via psychoacoustic hiding. NDSS (2019).","journal-title":"NDSS"},{"key":"e_1_3_2_106_2","article-title":"Robust over-the-air adversarial examples against automatic speech recognition systems","author":"Sch\u00f6nherr Lea","year":"2019","unstructured":"Lea Sch\u00f6nherr, Steffen Zeiler, Thorsten Holz, and Dorothea Kolossa. 2019. Robust over-the-air adversarial examples against automatic speech recognition systems. arXiv preprint arXiv:1908.01551 (2019).","journal-title":"arXiv preprint arXiv:1908.01551"},{"key":"e_1_3_2_107_2","doi-asserted-by":"publisher","DOI":"10.1145\/2976749.2978392"},{"key":"e_1_3_2_108_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2918261"},{"key":"e_1_3_2_109_2","article-title":"Lingvo: A modular and scalable framework for sequence-to-sequence modeling","author":"Shen Jonathan","year":"2019","unstructured":"Jonathan Shen, Patrick Nguyen, Yonghui Wu, Zhifeng Chen, Mia X. Chen, Ye Jia, Anjuli Kannan, Tara Sainath, Yuan Cao, Chung-Cheng Chiu, et\u00a0al. 2019. Lingvo: A modular and scalable framework for sequence-to-sequence modeling. arXiv preprint arXiv:1902.08295 (2019).","journal-title":"arXiv preprint arXiv:1902.08295"},{"key":"e_1_3_2_110_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00668"},{"key":"e_1_3_2_111_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-66787-4_22"},{"key":"e_1_3_2_112_2","first-page":"881","volume-title":"24th USENIX Security Symposium (USENIX Security\u201915)","author":"Son Yunmok","year":"2015","unstructured":"Yunmok Son, Hocheol Shin, Dongkwan Kim, Youngseok Park, Juhwan Noh, Kibum Choi, Jungwoo Choi, and Yongdae Kim. 2015. Rocking drones with intentional sound noise on gyroscopic sensors. In 24th USENIX Security Symposium (USENIX Security\u201915). 881\u2013896."},{"key":"e_1_3_2_113_2","doi-asserted-by":"publisher","DOI":"10.1145\/3133956.3138836"},{"key":"e_1_3_2_114_2","doi-asserted-by":"publisher","DOI":"10.1109\/TEVC.2019.2890858"},{"key":"e_1_3_2_115_2","doi-asserted-by":"crossref","unstructured":"Vinod Subramanian Emmanouil Benetos and Mark B. Sandler. 2019. Robustness of adversarial attacks in sound event classification. (2019).","DOI":"10.33682\/sp9n-qk06"},{"key":"e_1_3_2_116_2","unstructured":"Takeshi Sugawara Benjamin Cyr Sara Rampazzi Daniel Genkin and Kevin Fu. [n.d.]. Light commands: Laser-Based audio injection attacks on voice-controllable systems. ([n.d.])."},{"key":"e_1_3_2_117_2","first-page":"877","volume-title":"29th USENIX Security Symposium (USENIX Security\u201920)","author":"Sun Jiachen","year":"2020","unstructured":"Jiachen Sun, Yulong Cao, Qi Alfred Chen, and Z. Morley Mao. 2020. Towards robust lidar-based perception in autonomous driving: General black-box adversarial sensor attack and countermeasures. In 29th USENIX Security Symposium (USENIX Security\u201920). 877\u2013894."},{"key":"e_1_3_2_118_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2019.2933146"},{"key":"e_1_3_2_119_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-1247"},{"key":"e_1_3_2_120_2","doi-asserted-by":"publisher","DOI":"10.1145\/2462456.2464437"},{"key":"e_1_3_2_121_2","article-title":"Intriguing properties of neural networks","author":"Szegedy Christian","year":"2013","unstructured":"Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199 (2013).","journal-title":"arXiv preprint arXiv:1312.6199"},{"key":"e_1_3_2_122_2","article-title":"Perceptual based adversarial audio attacks","author":"Szurley Joseph","year":"2019","unstructured":"Joseph Szurley and Zico J. Kolter. 2019. Perceptual based adversarial audio attacks. CoRR (2019).","journal-title":"CoRR"},{"key":"e_1_3_2_123_2","first-page":"115","volume-title":"IEEE 11th International Workshop on Computational Intelligence and Applications (IWCIA)","author":"Tamura Keiichi","year":"2019","unstructured":"Keiichi Tamura, Akitada Omagari, and Shuichi Hashida. 2019. Novel defense method against audio adversarial example for speech-to-text transcription neural networks. In IEEE 11th International Workshop on Computational Intelligence and Applications (IWCIA). IEEE, 115\u2013120."},{"key":"e_1_3_2_124_2","first-page":"15","volume-title":"IEEE Security and Privacy Workshops (SPW)","author":"Taori Rohan","year":"2019","unstructured":"Rohan Taori, Amog Kamsetty, Brenton Chu, and Nikita Vemuri. 2019. Targeted adversarial examples for black box audio systems. In IEEE Security and Privacy Workshops (SPW). IEEE, 15\u201320."},{"key":"e_1_3_2_125_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2953924"},{"key":"e_1_3_2_126_2","article-title":"Black-box attacks on automatic speaker verification using feedback-controlled voice conversion","author":"Tian Xiaohai","year":"2019","unstructured":"Xiaohai Tian, Rohan Kumar Das, and Haizhou Li. 2019. Black-box attacks on automatic speaker verification using feedback-controlled voice conversion. arXiv preprint arXiv:1909.07655 (2019).","journal-title":"arXiv preprint arXiv:1909.07655"},{"key":"e_1_3_2_127_2","article-title":"Ensemble adversarial training: Attacks and defenses","author":"Tram\u00e8r Florian","year":"2017","unstructured":"Florian Tram\u00e8r, Alexey Kurakin, Nicolas Papernot, Ian Goodfellow, Dan Boneh, and Patrick McDaniel. 2017. Ensemble adversarial training: Attacks and defenses. arXiv preprint arXiv:1705.07204 (2017).","journal-title":"arXiv preprint arXiv:1705.07204"},{"key":"e_1_3_2_128_2","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1109\/EuroSP.2017.42","volume-title":"IEEE European Symposium on Security and Privacy (EuroS&P)","author":"Trippel Timothy","year":"2017","unstructured":"Timothy Trippel, Ofir Weisse, Wenyuan Xu, Peter Honeyman, and Kevin Fu. 2017. WALNUT: Waging doubt on the integrity of MEMS accelerometers with acoustic injection attacks. In IEEE European Symposium on Security and Privacy (EuroS&P). IEEE, 3\u201318."},{"key":"e_1_3_2_129_2","first-page":"1545","volume-title":"27th USENIX Security Symposium (USENIX Security\u201918)","author":"Tu Yazhou","year":"2018","unstructured":"Yazhou Tu, Zhiqiang Lin, Insup Lee, and Xiali Hei. 2018. Injected and delivered: Fabricating implicit control over actuation systems by spoofing inertial sensors. In 27th USENIX Security Symposium (USENIX Security\u201918). 1545\u20131562."},{"key":"e_1_3_2_130_2","article-title":"Universal adversarial examples in speech command classification","author":"Vadillo Jon","year":"2019","unstructured":"Jon Vadillo and Roberto Santana. 2019. Universal adversarial examples in speech command classification. arXiv preprint arXiv:1911.10182 (2019).","journal-title":"arXiv preprint arXiv:1911.10182"},{"key":"e_1_3_2_131_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cose.2021.102495"},{"key":"e_1_3_2_132_2","volume-title":"9th USENIX Workshop on Offensive Technologies (WOOT\u201915)","author":"Vaidya Tavish","year":"2015","unstructured":"Tavish Vaidya, Yuankai Zhang, Micah Sherr, and Clay Shields. 2015. Cocaine noodles: Exploiting the gap between human and machine speech recognition. In 9th USENIX Workshop on Offensive Technologies (WOOT\u201915)."},{"key":"e_1_3_2_133_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2019.05.005"},{"key":"e_1_3_2_134_2","doi-asserted-by":"publisher","DOI":"10.1145\/3359789.3359830"},{"key":"e_1_3_2_135_2","doi-asserted-by":"publisher","DOI":"10.1145\/2973750.2973765"},{"key":"e_1_3_2_136_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683479"},{"key":"e_1_3_2_137_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683479"},{"issue":"4","key":"e_1_3_2_138_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3369811","article-title":"Secure your voice: An oral airflow-based continuous liveness detection for voice assistants","volume":"3","author":"Wang Yao","year":"2019","unstructured":"Yao Wang, Wandong Cai, Tao Gu, Wei Shao, Yannan Li, and Yong Yu. 2019. Secure your voice: An oral airflow-based continuous liveness detection for voice assistants. Proc. ACM Interact., Mob., Wear. Ubiq. Technol. 3, 4 (2019), 1\u201328.","journal-title":"Proc. ACM Interact., Mob., Wear. Ubiq. Technol."},{"key":"e_1_3_2_139_2","unstructured":"Lei Wu Zhanxing Zhu Cheng Tai and E. Weinan. 2018. Enhancing the transferability of adversarial examples with noise reduced gradient. (2018)."},{"key":"e_1_3_2_140_2","first-page":"1","volume-title":"IEEE International Symposium on Dynamic Spectrum Access Networks (DySPAN)","author":"Wu Yi","year":"2019","unstructured":"Yi Wu, Jian Liu, Yingying Chen, and Jerry Cheng. 2019. Semi-black-box attacks against speech recognition systems using adversarial samples. In IEEE International Symposium on Dynamic Spectrum Access Networks (DySPAN). IEEE, 1\u20135."},{"key":"e_1_3_2_141_2","first-page":"443","volume-title":"28th USENIX Security Symposium (USENIX Security\u201919)","author":"Xiao Qixue","year":"2019","unstructured":"Qixue Xiao, Yufei Chen, Chao Shen, Yu Chen, and Kang Li. 2019. Seeing is not believing: Camouflage attacks on image scaling algorithms. In 28th USENIX Security Symposium (USENIX Security\u201919). 443\u2013460."},{"key":"e_1_3_2_142_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00284"},{"key":"e_1_3_2_143_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2018.2860991"},{"key":"e_1_3_2_144_2","article-title":"Enabling fast and universal audio adversarial attack using generative model","author":"Xie Yi","year":"2020","unstructured":"Yi Xie, Zhuohang Li, Cong Shi, Jian Liu, Yingying Chen, and Bo Yuan. 2020. Enabling fast and universal audio adversarial attack using generative model. arXiv preprint arXiv:2004.12261 (2020).","journal-title":"arXiv preprint arXiv:2004.12261"},{"key":"e_1_3_2_145_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053747"},{"key":"e_1_3_2_146_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2018.2867917"},{"key":"e_1_3_2_147_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASP-DAC47756.2020.9045584"},{"key":"e_1_3_2_148_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/741"},{"key":"e_1_3_2_149_2","first-page":"480","volume-title":"IEEE Symposium on Security and Privacy (SP)","author":"Yan Chen","year":"2020","unstructured":"Chen Yan, Hocheol Shin, Connor Bolton, Wenyuan Xu, Yongdae Kim, and Kevin Fu. 2020. SoK: A minimalist approach to formalizing analog sensor security. In IEEE Symposium on Security and Privacy (SP). 480\u2013495."},{"issue":"8","key":"e_1_3_2_150_2","first-page":"109","article-title":"Can you trust autonomous vehicles: Contactless attacks against sensors of self-driving vehicle","volume":"24","author":"Yan Chen","year":"2016","unstructured":"Chen Yan, Wenyuan Xu, and Jianhao Liu. 2016. Can you trust autonomous vehicles: Contactless attacks against sensors of self-driving vehicle. DEF CON 24, 8 (2016), 109.","journal-title":"DEF CON"},{"key":"e_1_3_2_151_2","doi-asserted-by":"publisher","DOI":"10.1109\/TDSC.2019.2906165"},{"key":"e_1_3_2_152_2","unstructured":"Qiben Yan Kehai Liu Qin Zhou Hanqing Guo and Ning Zhang. [n.d.]. SurfingAttack: Interactive hidden attack on voice assistants using ultrasonic guided waves. ([n.d.])."},{"key":"e_1_3_2_153_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053288"},{"key":"e_1_3_2_154_2","unstructured":"Zhuolin Yang Pin Yu Chen Bo Li and Dawn Song. 2019. Characterizing audio adversarial examples using temporal dependency. In 7th International Conference on Learning Representations ."},{"key":"e_1_3_2_155_2","first-page":"882","volume-title":"8th International Conference on Ubiquitous and Future Networks (ICUFN)","author":"Young Park Joon","year":"2016","unstructured":"Park Joon Young, Jo Hyo Jin, Samuel Woo, and Dong Hoon Lee. 2016. BadVoice: Soundless voice-control replay attack on modern smartphones. In 8th International Conference on Ubiquitous and Future Networks (ICUFN). IEEE, 882\u2013887."},{"key":"e_1_3_2_156_2","doi-asserted-by":"publisher","DOI":"10.1109\/GLOCOM.2018.8647762"},{"key":"e_1_3_2_157_2","first-page":"49","volume-title":"27th USENIX Security Symposium (USENIX Security\u201918)","author":"Yuan Xuejing","year":"2018","unstructured":"Xuejing Yuan, Yuxuan Chen, Yue Zhao, Yunhui Long, Xiaokang Liu, Kai Chen, Shengzhi Zhang, Heqing Huang, Xiaofeng Wang, and Carl A. Gunter. 2018. CommanderSong: A systematic approach for practical adversarial voice recognition. In 27th USENIX Security Symposium (USENIX Security\u201918). 49\u201364."},{"key":"e_1_3_2_158_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2019.00019"},{"key":"e_1_3_2_159_2","doi-asserted-by":"publisher","DOI":"10.1145\/3133956.3134052"},{"key":"e_1_3_2_160_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2020\/438"},{"key":"e_1_3_2_161_2","first-page":"23","volume-title":"7th International Workshop on Security in Cloud Computing","author":"Zhang Jiajie","year":"2019","unstructured":"Jiajie Zhang, Bingsheng Zhang, and Bincheng Zhang. 2019. Defending adversarial attacks on cloud-aided automatic speech recognition systems. In 7th International Workshop on Security in Cloud Computing. 23\u201331."},{"key":"e_1_3_2_162_2","first-page":"1381","volume-title":"IEEE Symposium on Security and Privacy (SP)","author":"Zhang Nan","year":"2019","unstructured":"Nan Zhang, Xianghang Mi, Xuan Feng, XiaoFeng Wang, Yuan Tian, and Feng Qian. 2019. Dangerous skills: Understanding and mitigating security risks of voice-controlled third-party functions on virtual personal assistant systems. In IEEE Symposium on Security and Privacy (SP). IEEE, 1381\u20131396."},{"key":"e_1_3_2_163_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-30619-9_27"},{"key":"e_1_3_2_164_2","article-title":"Life after speech recognition: Fuzzing semantic misinterpretation for voice assistant applications.","author":"Zhang Yangyong","year":"2019","unstructured":"Yangyong Zhang, Lei Xu, Abner Mendoza, Guangliang Yang, Phakpoom Chinprutthiwong, and Guofei Gu. 2019. Life after speech recognition: Fuzzing semantic misinterpretation for voice assistant applications. NDSS.","journal-title":"NDSS"},{"key":"e_1_3_2_165_2","article-title":"Black-box adversarial attacks on commercial speech platforms with minimal information","author":"Zheng Baolin","year":"2021","unstructured":"Baolin Zheng, Peipei Jiang, Qian Wang, Qi Li, Chao Shen, Cong Wang, Yunjie Ge, Qingyang Teng, and Shenyi Zhang. 2021. Black-box adversarial attacks on commercial speech platforms with minimal information. arXiv preprint arXiv:2110.09714 (2021).","journal-title":"arXiv preprint arXiv:2110.09714"},{"key":"e_1_3_2_166_2","doi-asserted-by":"publisher","DOI":"10.1145\/3131672.3131689"}],"container-title":["ACM Transactions on Privacy and Security"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3510582","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3510582","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:12:19Z","timestamp":1750191139000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3510582"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,5,19]]},"references-count":165,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,8,31]]}},"alternative-id":["10.1145\/3510582"],"URL":"https:\/\/doi.org\/10.1145\/3510582","relation":{},"ISSN":["2471-2566","2471-2574"],"issn-type":[{"value":"2471-2566","type":"print"},{"value":"2471-2574","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,5,19]]},"assertion":[{"value":"2021-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-05-19","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}