{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T22:28:27Z","timestamp":1783204107308,"version":"3.54.6"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,9,27]],"date-time":"2023-09-27T00:00:00Z","timestamp":1695772800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2023,9,27]]},"abstract":"<jats:p>Intelligent audio systems are ubiquitous in our lives, such as speech command recognition and speaker recognition. However, it is shown that deep learning-based intelligent audio systems are vulnerable to adversarial attacks. In this paper, we propose a physical adversarial attack that exploits reverberation, a natural indoor acoustic effect, to realize imperceptible, fast, and targeted black-box attacks. Unlike existing attacks that constrain the magnitude of adversarial perturbations within a fixed radius, we generate reverberation-alike perturbations that blend naturally with the original voice sample 1. Additionally, we can generate more robust adversarial examples even under over-the-air propagation by considering distortions in the physical environment. Extensive experiments are conducted using two popular intelligent audio systems in various situations, such as different room sizes, distance, and ambient noises. The results show that Echo can invade into intelligent audio systems in both digital and physical over-the-air environment.<\/jats:p>","DOI":"10.1145\/3610874","type":"journal-article","created":{"date-parts":[[2023,9,27]],"date-time":"2023-09-27T15:45:03Z","timestamp":1695829503000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Echo"],"prefix":"10.1145","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4745-6526","authenticated-orcid":false,"given":"Meng","family":"Xue","sequence":"first","affiliation":[{"name":"Hong Kong University of Science and Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-5441-8114","authenticated-orcid":false,"given":"Kuang","family":"Peng","sequence":"additional","affiliation":[{"name":"Wuhan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2190-8117","authenticated-orcid":false,"given":"Xueluan","family":"Gong","sequence":"additional","affiliation":[{"name":"Wuhan University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9205-1881","authenticated-orcid":false,"given":"Qian","family":"Zhang","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1382-0679","authenticated-orcid":false,"given":"Yanjiao","family":"Chen","sequence":"additional","affiliation":[{"name":"Zhejiang University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7471-4355","authenticated-orcid":false,"given":"Routing","family":"Li","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,9,27]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Kevin RB Butler, and Joseph Wilson","author":"Abdullah Hadi","year":"2019","unstructured":"Hadi Abdullah, Washington Garcia, Christian Peeters, Patrick Traynor, Kevin RB Butler, and Joseph Wilson. 2019. Practical hidden voice attacks against speech and speaker recognition systems. arXiv preprint arXiv:1904.05734 (2019)."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP40001.2021.00009"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP40001.2021.00014"},{"key":"e_1_2_1_4_1","unstructured":"Last accessed. 2022. Android app which enables unlock of mobile phone via voice print. https:\/\/app.mi.com\/details?id=com.jie.lockscreen"},{"key":"e_1_2_1_5_1","unstructured":"Last accessed. 2022. Social software wechat adds voiceprint lock login function. https:\/\/kf.qq.com\/touch\/wxappfaq\/ 1208117b2mai141125YZjAra.html"},{"key":"e_1_2_1_6_1","unstructured":"Last accessed. 2022. Voice Commands. https:\/\/www.tesla.com\/support\/voice-commands"},{"key":"e_1_2_1_7_1","volume-title":"Did you hear that? adversarial examples against automatic speech recognition. arXiv preprint arXiv:1801.00554","author":"Alzantot Moustafa","year":"2018","unstructured":"Moustafa Alzantot, Bharathan Balaji, and Mani Srivastava. 2018. Did you hear that? adversarial examples against automatic speech recognition. arXiv preprint arXiv:1801.00554 (2018)."},{"key":"e_1_2_1_8_1","unstructured":"Karissa Bell. 2015. A smarter Siri learns to recognize the sound of your voice in iOS 9. https:\/\/mashable.com\/archive\/hey-siri-voice-recognition"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397332"},{"key":"e_1_2_1_10_1","volume-title":"Decision-based adversarial attacks: Reliable attacks against black-box machine learning models. arXiv preprint arXiv:1712.04248","author":"Brendel Wieland","year":"2017","unstructured":"Wieland Brendel, Jonas Rauber, and Matthias Bethge. 2017. Decision-based adversarial attacks: Reliable attacks against black-box machine learning models. arXiv preprint arXiv:1712.04248 (2017)."},{"key":"e_1_2_1_11_1","volume-title":"25th USENIX security symposium (USENIX security 16). 513--530.","author":"Carlini Nicholas","unstructured":"Nicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang, Micah Sherr, Clay Shields, David Wagner, and Wenchao Zhou. 2016. Hidden voice commands. In 25th USENIX security symposium (USENIX security 16). 513--530."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/SPW.2018.00009"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP40001.2021.00004"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.14722\/ndss.2020.23055"},{"key":"e_1_2_1_15_1","unstructured":"Yuxuan Chen Xuejing Yuan Jiangshan Zhang Yue Zhao Shengzhi Zhang Kai Chen and XiaoFeng Wang. 2020. Devil's whisper: A general approach for physical adversarial attacks against commercial black-box speech recognition devices. In {USENIX} Security Symposium ({USENIX} Security 20). 2667--2684."},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"J. S. Chung A. Nagrani and A. Zisserman. 2018. VoxCeleb2: Deep Speaker Recognition. In INTERSPEECH.","DOI":"10.21437\/Interspeech.2018-1929"},{"key":"e_1_2_1_17_1","volume-title":"Houdini: Fooling deep structured prediction models. arXiv preprint arXiv:1707.05373","author":"Cisse Moustapha","year":"2017","unstructured":"Moustapha Cisse, Yossi Adi, Natalia Neverova, and Joseph Keshet. 2017. Houdini: Fooling deep structured prediction models. arXiv preprint arXiv:1707.05373 (2017)."},{"key":"e_1_2_1_18_1","volume-title":"Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, 677--681","author":"Das Nilaksh","year":"2018","unstructured":"Nilaksh Das, Madhuri Shanbhogue, Shang-Tse Chen, Li Chen, Michael E Kounavis, and Duen Horng Chau. 2018. Adagio: Interactive experimentation with adversarial attack and defense for audio. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, 677--681."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1749-6632.2011.06027.x"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2009.4960478"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3320269.3384733"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3478091"},{"key":"e_1_2_1_23_1","volume-title":"Neuroevolution: from architectures to learning. Evolutionary intelligence 1, 1","author":"Floreano Dario","year":"2008","unstructured":"Dario Floreano, Peter D\u00fcrr, and Claudio Mattiussi. 2008. Neuroevolution: from architectures to learning. Evolutionary intelligence 1, 1 (2008), 47--62."},{"key":"e_1_2_1_24_1","volume-title":"Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (CMA-ES). Evolutionary computation 11, 1","author":"Hansen Nikolaus","year":"2003","unstructured":"Nikolaus Hansen, Sibylle D M\u00fcller, and Petros Koumoutsakos. 2003. Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (CMA-ES). Evolutionary computation 11, 1 (2003), 1--18."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_26_1","volume-title":"Fooling end-to-end speaker verification by adversarial examples. arXiv preprint. arXiv preprint arXiv:1801.03339","author":"Kreuk F","year":"2018","unstructured":"F Kreuk, Y Adi, M Cisse, and J Keshet. 2018. Fooling end-to-end speaker verification by adversarial examples. arXiv preprint. arXiv preprint arXiv:1801.03339 (2018)."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.2936367"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2009.2035038"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASPAA.2007.4392980"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3376897.3377856"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3372297.3423348"},{"key":"e_1_2_1_32_1","volume-title":"N.Y.","author":"Lloyd Llewelyn S.","unstructured":"Llewelyn S. Lloyd. 1937. Music and sound. Freeport, N.Y., Books for Libraries Press."},{"key":"e_1_2_1_33_1","volume-title":"Evolutionary algorithms and neural networks","author":"Mirjalili Seyedali","unstructured":"Seyedali Mirjalili. 2019. Genetic algorithm. In Evolutionary algorithms and neural networks. Springer, 43--55."},{"key":"e_1_2_1_34_1","doi-asserted-by":"crossref","unstructured":"Satoshi Nakamura Kazuo Hiyane Futoshi Asano Takanobu Nishiura and Takeshi Yamada. 2000. Acoustical sound database in real environments for sound scene understanding and hands-free speech recognition. (2000).","DOI":"10.21437\/Eurospeech.1999-568x"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"e_1_2_1_36_1","volume-title":"Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32 (2019)."},{"key":"e_1_2_1_37_1","doi-asserted-by":"crossref","unstructured":"Vijayaditya Peddinti Daniel Povey and Sanjeev Khudanpur. 2015. A time delay neural network architecture for efficient modeling of long temporal contexts. In Sixteenth annual conference of the international speech communication association.","DOI":"10.21437\/Interspeech.2015-647"},{"key":"e_1_2_1_38_1","volume-title":"Particle swarm optimization. Swarm intelligence 1, 1","author":"Poli Riccardo","year":"2007","unstructured":"Riccardo Poli, James Kennedy, and Tim Blackwell. 2007. Particle swarm optimization. Swarm intelligence 1, 1 (2007), 33--57."},{"key":"e_1_2_1_39_1","volume-title":"International conference on machine learning. PMLR, 5231--5240","author":"Qin Yao","year":"2019","unstructured":"Yao Qin, Nicholas Carlini, Garrison Cottrell, Ian Goodfellow, and Colin Raffel. 2019. Imperceptible, robust, and targeted adversarial examples for automatic speech recognition. In International conference on machine learning. PMLR, 5231--5240."},{"key":"e_1_2_1_40_1","volume-title":"Renato De Mori, and Yoshua Bengio","author":"Ravanelli Mirco","year":"2021","unstructured":"Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, Ju-Chieh Chou, Sung-Lin Yeh, Szu-Wei Fu, Chien-Feng Liao, Elena Rastorgueva, Fran\u00e7ois Grondin, William Aris, Hwidong Na, Yan Gao, Renato De Mori, and Yoshua Bengio. 2021. SpeechBrain: A General-Purpose Speech Toolkit. arXiv:2106.04624 [eess.AS] arXiv:2106.04624."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/1830483.1830589"},{"key":"e_1_2_1_42_1","doi-asserted-by":"crossref","unstructured":"Tara Sainath and Carolina Parada. 2015. Convolutional neural networks for small-footprint keyword spotting. (2015).","DOI":"10.21437\/Interspeech.2015-352"},{"key":"e_1_2_1_43_1","volume-title":"Adversarial attacks against automatic speech recognition systems via psychoacoustic hiding. arXiv preprint arXiv:1808.05665","author":"Sch\u00f6nherr Lea","year":"2018","unstructured":"Lea Sch\u00f6nherr, Katharina Kohls, Steffen Zeiler, Thorsten Holz, and Dorothea Kolossa. 2018. Adversarial attacks against automatic speech recognition systems via psychoacoustic hiding. arXiv preprint arXiv:1808.05665 (2018)."},{"key":"e_1_2_1_44_1","volume-title":"Spoken language technology workshop (slt)","author":"Shon Suwon","unstructured":"Suwon Shon, Hao Tang, and James Glass. 2018. Frame-level speaker embeddings for text-independent speaker recognition and analysis of end-to-end model. In Spoken language technology workshop (slt). IEEE, 1007--1013."},{"key":"e_1_2_1_45_1","unstructured":"SLR31. 2022. Mini LibriSpeech ASR corpus. https:\/\/www.openslr.org\/31\/"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461375"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/TEVC.2005.856210"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-018-0006-z"},{"key":"e_1_2_1_49_1","volume-title":"Wave-u-net: A multi-scale neural network for end-to-end audio source separation. arXiv preprint arXiv:1806.03185","author":"Stoller Daniel","year":"2018","unstructured":"Daniel Stoller, Sebastian Ewert, and Simon Dixon. 2018. Wave-u-net: A multi-scale neural network for end-to-end audio source separation. arXiv preprint arXiv:1806.03185 (2018)."},{"key":"e_1_2_1_50_1","volume-title":"International symposium on signal processing and intelligent recognition systems. Springer, 190--201","author":"Tulshan Amrita S","year":"2018","unstructured":"Amrita S Tulshan and Sudhir Namdeorao Dhage. 2018. Survey on virtual assistant: Google assistant, siri, cortana, alexa. In International symposium on signal processing and intelligent recognition systems. Springer, 190--201."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.2908264"},{"key":"e_1_2_1_52_1","volume-title":"Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition. ArXiv e-prints (April","author":"Warden P.","year":"2018","unstructured":"P. Warden. 2018. Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition. ArXiv e-prints (April 2018). arXiv:1804.03209 [cs.CL] https:\/\/arxiv.org\/abs\/1804.03209"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2638566"},{"key":"e_1_2_1_54_1","unstructured":"Wiki. 2022. Reverberation. https:\/\/en.wikipedia.org\/wiki\/Reverberation"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683120"},{"key":"e_1_2_1_56_1","volume-title":"Enabling fast and universal audio adversarial attack using generative model. arXiv preprint arXiv:2004.12261","author":"Xie Yi","year":"2020","unstructured":"Yi Xie, Zhuohang Li, Cong Shi, Jian Liu, Yingying Chen, and Bo Yuan. 2020. Enabling fast and universal audio adversarial attack using generative model. arXiv preprint arXiv:2004.12261 (2020)."},{"key":"e_1_2_1_57_1","volume-title":"Robust audio adversarial example for a physical attack. arXiv preprint arXiv:1810.11793","author":"Yakura Hiromu","year":"2018","unstructured":"Hiromu Yakura and Jun Sakuma. 2018. Robust audio adversarial example for a physical attack. arXiv preprint arXiv:1810.11793 (2018)."},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","unstructured":"Junichi Yamagishi Christophe Veaux and Kirsten MacDonald. 2019. CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92). https:\/\/doi.org\/10.7488\/ds\/2645","DOI":"10.7488\/ds\/2645"},{"key":"e_1_2_1_59_1","volume-title":"Characterizing audio adversarial examples using temporal dependency. arXiv preprint arXiv:1809.10875","author":"Yang Zhuolin","year":"2018","unstructured":"Zhuolin Yang, Bo Li, Pin-Yu Chen, and Dawn Song. 2018. Characterizing audio adversarial examples using temporal dependency. arXiv preprint arXiv:1809.10875 (2018)."},{"key":"e_1_2_1_60_1","volume-title":"27th USENIX security symposium (USENIX security 18). 49--64.","author":"Yuan Xuejing","unstructured":"Xuejing Yuan, Yuxuan Chen, Yue Zhao, Yunhui Long, Xiaokang Liu, Kai Chen, Shengzhi Zhang, Heqing Huang, Xiaofeng Wang, and Carl A Gunter. 2018. {CommanderSong}: A Systematic Approach for Practical Adversarial Voice Recognition. In 27th USENIX security symposium (USENIX security 18). 49--64."},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133956.3134052"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3460120.3485383"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-1158"}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3610874","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3610874","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,28]],"date-time":"2025-07-28T16:27:40Z","timestamp":1753720060000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3610874"}},"subtitle":["Reverberation-based Fast Black-Box Adversarial Attacks on Intelligent Audio Systems"],"short-title":[],"issued":{"date-parts":[[2023,9,27]]},"references-count":63,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,9,27]]}},"alternative-id":["10.1145\/3610874"],"URL":"https:\/\/doi.org\/10.1145\/3610874","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,27]]},"assertion":[{"value":"2023-09-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}