{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,4]],"date-time":"2025-11-04T11:12:49Z","timestamp":1762254769109,"version":"3.41.2"},"reference-count":34,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,2,15]],"date-time":"2024-02-15T00:00:00Z","timestamp":1707955200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Comput. Sci."],"abstract":"<jats:p>This study explores the vulnerability of robot dialogue systems' automatic speech recognition (ASR) module to adversarial music attacks. Specifically, we explore music as a natural camouflage for such attacks. We propose a novel method to hide ghost speech commands in a music clip by slightly perturbing its raw waveform. We apply our attack on an industry-popular ASR model, namely the time-delay neural network (TDNN), widely used for speech and speaker recognition. Our experiment demonstrates that adversarial music crafted by our attack can easily mislead industry-level TDNN models into picking up ghost commands with high success rates. However, it sounds no different from the original music to the human ear. This reveals a serious threat by adversarial music to robot dialogue systems, calling for effective defenses against such stealthy attacks.<\/jats:p>","DOI":"10.3389\/fcomp.2024.1355975","type":"journal-article","created":{"date-parts":[[2024,2,15]],"date-time":"2024-02-15T04:26:13Z","timestamp":1707971173000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Phantom in the opera: adversarial music attack for robot dialogue system"],"prefix":"10.3389","volume":"6","author":[{"given":"Sheng","family":"Li","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiyi","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Cao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2024,2,15]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","DOI":"10.14722\/ndss.2019.23362","article-title":"Practical hidden voice attacks against speech and speaker recognition systems","author":"Abdullah","year":"2019","journal-title":"arXiv:1904.05734"},{"key":"B2","article-title":"\u201cDid you hear that? Adversarial examples against automatic speech recognition,\u201d","author":"Alzantot","year":"2017","journal-title":"NIPS 2017 Machine Deception Workshop"},{"key":"B3","first-page":"513","article-title":"\u201cHidden voice commands,\u201d","author":"Carlini","year":"2016","journal-title":"25th $USENIX$ Security Symposium ($USENIX$ Security 16)"},{"key":"B4","doi-asserted-by":"publisher","DOI":"10.1109\/SPW.2018.00009","article-title":"Audio adversarial examples: targeted attacks on speech-to-text","author":"Carlini","year":"2018","journal-title":"abs\/1801.01944"},{"key":"B5","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7472621","article-title":"\u201cListen, attend and spell: a neural network for large vocabulary conversational speech recognition,\u201d","author":"Chan","year":"2016","journal-title":"Proceedings of IEEE-ICASSP"},{"key":"B6","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2023-1217","article-title":"\u201cBuild a SRE challenge system: lessons from VoxSRC 2022 and CNSRC 2022,\u201d","author":"Chen","year":"2023","journal-title":"Proceedings of INTERSPEECH"},{"key":"B7","first-page":"6977","article-title":"\u201cHoudini: fooling deep structured visual and speech recognition models with adversarial examples,\u201d","author":"Cisse","year":"2017","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"B8","doi-asserted-by":"publisher","first-page":"30","DOI":"10.1109\/TASL.2011.2134090","article-title":"Context dependent pre-trained deep neural networks for large vocabulary speech recognition","volume":"20","author":"Dahl","year":"2012","journal-title":"IEEE Trans. ASLP"},{"key":"B9","article-title":"\u201cLinear prediction-based dereverberation with advanced speech enhancement and recognition technologies for the reverb challenge,\u201d","author":"Delcroix","year":"2014","journal-title":"Proceedings of REVERB Challenge Workshop"},{"key":"B10","doi-asserted-by":"publisher","DOI":"10.1109\/ROMAN.2016.7745086","article-title":"\u201cErica: the Erato intelligent conversational android,\u201d","author":"Glas","year":"2016","journal-title":"International Symposium on Robot and Human Interactive Communication (RO-MAN)"},{"key":"B11","article-title":"Explaining and harnessing adversarial examples","author":"Goodfellow","year":"2014","journal-title":"arXiv preprint arXiv:1412.6572"},{"key":"B12","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143891","article-title":"\u201cConnectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,\u201d","author":"Graves","year":"2006","journal-title":"Proceedings of ICML"},{"key":"B13","article-title":"\u201cTowards end-to-end speech recognition with recurrent neural networks,\u201d","author":"Graves","year":"2014","journal-title":"Proceedings of ICML"},{"key":"B14","doi-asserted-by":"publisher","first-page":"236","DOI":"10.1109\/TASSP.1984.1164317","article-title":"Signal estimation from modified short-time Fourier transform","volume":"32","author":"Griffin","year":"1984","journal-title":"IEEE Trans. ASSP"},{"key":"B15","article-title":"\u201cA job interview dialogue system with autonomous android Erica,\u201d","author":"Inoue","year":"2019","journal-title":"Int'l Workshop Spoken Dialogue Systems (IWSDS)"},{"key":"B16","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2015-711","article-title":"\u201cAudio augmentation for speech recognition,\u201d","author":"Ko","year":"2015","journal-title":"Proceedings of INTERSPEECH"},{"key":"B17","article-title":"Towards deep learning models resistant to adversarial attacks","author":"Madry","year":"2017","journal-title":"arXiv preprint arXiv:1706.06083"},{"key":"B18","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-1979","article-title":"Efficient keyword spotting using time delay neural networks","author":"Myer","year":"2018","journal-title":"arXiv preprint arXiv:1807.04353"},{"key":"B19","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178964","article-title":"\u201cLibrispeech: an ASR corpus based on public domain audio books,\u201d","author":"Panayotov","year":"2015","journal-title":"Proceedings of IEEE-ICASSP"},{"key":"B20","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2015-647","article-title":"\u201cA time delay neural network architecture for efficient modeling of long temporal contexts,\u201d","author":"Peddinti","year":"2015","journal-title":"Proceedings of INTERSPEECH"},{"key":"B21","article-title":"\u201cThe Kaldi speech recognition toolkit,\u201d","author":"Povey","year":"2011","journal-title":"Proceedings of IEEE-ASRU"},{"key":"B22","article-title":"\u201cParallel training of deep neural networks with natural gradient and parameter averaging,\u201d","author":"Povey","year":"2015","journal-title":"Proceedings of ICLR Workshop"},{"key":"B23","first-page":"5231","article-title":"\u201cImperceptible, robust, and targeted adversarial examples for automatic speech recognition,\u201d","author":"Qin","year":"2019","journal-title":"Proceedings of the 36th International Conference on Machine Learning (ICML), Vol. 97"},{"key":"B24","doi-asserted-by":"publisher","first-page":"257","DOI":"10.1109\/5.18626","article-title":"A tutorial on hidden markov models and selected applications in speech recognition","volume":"77","author":"Rabiner","year":"1988","journal-title":"Proc. IEEE"},{"key":"B25","doi-asserted-by":"publisher","DOI":"10.14722\/ndss.2019.23288","article-title":"Adversarial attacks against automatic speech recognition systems via psychoacoustic hiding","author":"Sch\u00f6nherr","year":"2019"},{"key":"B26","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-480","article-title":"\u201cCompressed time delay neural network for small-footprint keyword spotting,\u201d","author":"Sun","year":"2017","journal-title":"Interspeech"},{"key":"B27","article-title":"\u201cAttention is all you need,\u201d","volume-title":"31st Conference on Neural Information Processing Systems (NIPS 2017)","author":"Vaswani","year":"2017"},{"key":"B28","doi-asserted-by":"publisher","first-page":"328","DOI":"10.1109\/29.21701","article-title":"Phoneme recognition using time-delay neural networks","volume":"37","author":"Waibel","year":"1989","journal-title":"IEEE\/ACM Trans. ASLP"},{"key":"B29","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10096626","article-title":"\u201cWespeaker: a research and production oriented speaker embedding learning toolkit,\u201d","author":"Wang","year":"2023","journal-title":"Proceedings of IEEE-ICASSP"},{"key":"B30","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-1456","article-title":"\u201cEspnet: end-to-end speech processing toolkit,\u201d","author":"Watanabe","year":"2018","journal-title":"Proceedings of INTERSPEECH"},{"key":"B31","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/741","article-title":"\u201cRobust audio adversarial example for a physical attack,\u201d","author":"Yakura","year":"2019","journal-title":"International Joint Conferences on Artificial Intelligence Organization (IJCAI)"},{"key":"B32","article-title":"\u201cThe HTK book version 3.4.1,\u201d","author":"Young","year":"2009","journal-title":"Tutorial Books"},{"key":"B33","first-page":"49","article-title":"\u201cCommandersong: a systematic approach for practical adversarial voice recognition,\u201d","author":"Yuan","year":"2018","journal-title":"USENIX"},{"key":"B34","doi-asserted-by":"publisher","DOI":"10.1145\/3133956.3134052","article-title":"\u201cDolphinattack: inaudible voice commands,\u201d","author":"Zhang","year":"2017","journal-title":"ACM CCS"}],"container-title":["Frontiers in Computer Science"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fcomp.2024.1355975\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,2,15]],"date-time":"2024-02-15T04:26:30Z","timestamp":1707971190000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fcomp.2024.1355975\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,15]]},"references-count":34,"alternative-id":["10.3389\/fcomp.2024.1355975"],"URL":"https:\/\/doi.org\/10.3389\/fcomp.2024.1355975","relation":{},"ISSN":["2624-9898"],"issn-type":[{"type":"electronic","value":"2624-9898"}],"subject":[],"published":{"date-parts":[[2024,2,15]]},"article-number":"1355975"}}