{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T16:54:38Z","timestamp":1783184078097,"version":"3.54.6"},"reference-count":27,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2025,1,27]],"date-time":"2025-01-27T00:00:00Z","timestamp":1737936000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2018YFB0803400"],"award-info":[{"award-number":["2018YFB0803400"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2021ZD0110400"],"award-info":[{"award-number":["2021ZD0110400"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"China National Natural Science Foundation","doi-asserted-by":"crossref","award":["62132018"],"award-info":[{"award-number":["62132018"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key Research Program of Frontier Sciences","award":["QYZDY-SSW-JSC002"],"award-info":[{"award-number":["QYZDY-SSW-JSC002"]}]},{"name":"The University Synergy Innovation Program of Anhui Province","award":["GXXT-2019-024"],"award-info":[{"award-number":["GXXT-2019-024"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Sen. Netw."],"published-print":{"date-parts":[[2025,1,31]]},"abstract":"<jats:p>Intelligent voice systems are widely utilized to control smart home applications, which raises significant privacy and security concerns. Recent studies have revealed their vulnerability to adversarial attacks, replay attacks, and so on. However, these attacks rely on the victim\u2019s voice data. In our work, we investigate a stealthy and command-independent attack that does not necessitate collecting victims\u2019 voices. Our proposed attack, IUAC, misleads the voice system to go against the victim\u2019s will, regardless of the commands delivered. Our core concept is to train highly robust attack commands through the construction of diverse data, rendering the user\u2019s commands negligible. To achieve stealthy attacks, we leverage a high-frequency carrier to construct an inaudible universal adversarial command. Extensive experiments conducted with real-world datasets demonstrate that our attack system attains an average attack success rate of 96% while resisting environmental interference. Moreover, our attack success rate against real-world voice systems is 4.52\u00d7 higher than the state-of-the-art. Finally, we propose an effective defense mechanism and provide experimental tests to validate its efficacy.<\/jats:p>","DOI":"10.1145\/3698238","type":"journal-article","created":{"date-parts":[[2024,9,30]],"date-time":"2024-09-30T10:09:39Z","timestamp":1727690979000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["IUAC: Inaudible Universal Adversarial Attacks Against Smart Speakers"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-4980-6263","authenticated-orcid":false,"given":"Haifeng","family":"Sun","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8492-3990","authenticated-orcid":false,"given":"Haohua","family":"Du","sequence":"additional","affiliation":[{"name":"Beihang University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9730-606X","authenticated-orcid":false,"given":"Xiaojing","family":"Yu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3340-8585","authenticated-orcid":false,"given":"Jiahui","family":"Hou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1004-8588","authenticated-orcid":false,"given":"Lan","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6070-6625","authenticated-orcid":false,"given":"Xiangyang","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,1,27]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Sajjad Abdoli Luiz G. Hafemann J\u00e9r\u00f4me Rony Ismail Ben Ayed Patrick Cardinal and Alessandro L. Koerich. 2019. Universal adversarial audio perturbations. CoRR abs\/1908.03173 (2019). Retrieved from http:\/\/arxiv.org\/abs\/1908.03173"},{"key":"e_1_3_1_3_2","doi-asserted-by":"crossref","unstructured":"Matthew B. Hoy. 2018. Alexa Siri Cortana and more: An introduction to voice assistants. Medical Reference Services Quarterly 37 1 (2018) 81\u201388.","DOI":"10.1080\/02763869.2018.1404391"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","unstructured":"Andreas Kaplan and Michael Haenlein. 2019. Siri Siri in my hand: Who\u2019s the fairest in the land? On the Interpretations Illustrations and Implications of Artificial Intelligence. Business Horizons 62 1 (2019) 15\u201325.","DOI":"10.1016\/j.bushor.2018.08.004"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/3241094.3241135"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/SPW.2018.00009"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/SP40001.2021.00004"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","unstructured":"Joon Son Chung Jaesung Huh Seongkyu Mun Minjae Lee Hee-Soo Heo Soyeon Choe Chiheon Ham Sunghwan Jung Bong-Jin Lee and Icksang Han. 2020. In defence of metric learning for speaker recognition. In Interspeech 2020 21st Annual Conference of the International Speech Communication Association Virtual Event Shanghai China 25-29 October 2020 ISCA 2977\u20132981. DOI:10.21437\/Interspeech.2020-1064","DOI":"10.21437\/Interspeech.2020-1064"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/2666620.2666623"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143891"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.23915\/distill.00008"},{"key":"e_1_3_1_12_2","unstructured":"Awni Y. Hannun Carl Case Jared Casper Bryan Catanzaro Greg Diamos Erich Elsen Ryan Prenger Sanjeev Satheesh Shubho Sengupta Adam Coates and Andrew Y. Ng. 2014. Deep Speech: Scaling up end-to-end speech recognition. CoRR abs\/1412.5567 (2014). Retrieved from http:\/\/arxiv.org\/abs\/1412.5567"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Yichong Leng Xu Tan Linchen Zhu Jin Xu Renqian Luo Linquan Liu Tao Qin Xiangyang Li Edward Lin and Tie-Yan Liu. 2021. Fastcorrect: Fast error correction with edit alignment for automatic speech recognition. Advances in Neural Information Processing Systems 34 (2021) 21708\u201321719.","DOI":"10.18653\/v1\/2021.findings-emnlp.367"},{"key":"e_1_3_1_14_2","first-page":"2455","volume-title":"Proceedings of the 32nd USENIX Security Symposium","author":"Li Xinfeng","year":"2023","unstructured":"Xinfeng Li, Xiaoyu Ji, Chen Yan, Chaohao Li, Yichen Li, Zhenning Zhang, and Wenyuan Xu. 2023. Learning normality is enough: a software-based mitigation against inaudible voice attacks. In Proceedings of the 32nd USENIX Security Symposium. 2455\u20132472."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3376897.3377856"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3372297.3423348"},{"key":"e_1_3_1_17_2","first-page":"81","volume-title":"Proceedings of the 18th ACM International Symposium on Mobile Ad Hoc Networking and Computing","author":"Meng Yan","year":"2018","unstructured":"Yan Meng, Zichang Wang, Wei Zhang, Peilin Wu, Haojin Zhu, Xiaohui Liang, and Yao Liu. 2018. Wivo: Enhancing the security of voice control system via wireless signal in iot environment. In Proceedings of the 18th ACM International Symposium on Mobile Ad Hoc Networking and Computing. 81\u201390."},{"key":"e_1_3_1_18_2","article-title":"Project Deepspeech","year":"2017","unstructured":"Mozilla. 2017. Project Deepspeech. Retrieved from https:\/\/github.com\/mozilla\/DeepSpeech","journal-title":"https:\/\/github.com\/mozilla\/DeepSpeech"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","unstructured":"Nicolas Papernot Patrick McDaniel Ian Goodfellow Somesh Jha Z Berkay Celik and Ananthram Swami. 2017. Practical black-box attacks against machine learning. In Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security. 506\u2013519.","DOI":"10.1145\/3052973.3053009"},{"key":"e_1_3_1_21_2","first-page":"5231","volume-title":"Proceedings of the International Conference on Machine Learning, PMLR","author":"Qin Yao","year":"2019","unstructured":"Yao Qin, Nicholas Carlini, Garrison Cottrell, Ian Goodfellow, and Colin Raffel. 2019. Imperceptible, robust, and targeted adversarial examples for automatic speech recognition. In Proceedings of the International Conference on Machine Learning, PMLR. 5231\u20135240."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3081333.3081366"},{"key":"e_1_3_1_23_2","first-page":"547","volume-title":"Proceedings of the 15th USENIX Symposium on Networked Systems Design and Implementation","author":"Roy Nirupam","year":"2018","unstructured":"Nirupam Roy, Sheng Shen, Haitham Hassanieh, and Romit Roy Choudhury. 2018. Inaudible Voice Commands: The \\(\\lbrace\\) Long-Range \\(\\rbrace\\) Attack and Defense. In Proceedings of the 15th USENIX Symposium on Networked Systems Design and Implementation. 547\u2013560."},{"key":"e_1_3_1_24_2","volume-title":"Proceedings of the 9th USENIX Workshop on Offensive Technologies","author":"Vaidya Tavish","year":"2015","unstructured":"Tavish Vaidya, Yuankai Zhang, Micah Sherr, and Clay Shields. 2015. Cocaine noodles: exploiting the gap between human and machine speech recognition. In Proceedings of the 9th USENIX Workshop on Offensive Technologies."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3372297.3417254"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053747"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3133956.3134052"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM41043.2020.9155483"}],"container-title":["ACM Transactions on Sensor Networks"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3698238","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3698238","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:17:18Z","timestamp":1750295838000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3698238"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,27]]},"references-count":27,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1,31]]}},"alternative-id":["10.1145\/3698238"],"URL":"https:\/\/doi.org\/10.1145\/3698238","relation":{},"ISSN":["1550-4859","1550-4867"],"issn-type":[{"value":"1550-4859","type":"print"},{"value":"1550-4867","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,27]]},"assertion":[{"value":"2023-12-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-24","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}