{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,9]],"date-time":"2025-11-09T02:06:37Z","timestamp":1762653997696,"version":"build-2065373602"},"reference-count":93,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,11,9]],"date-time":"2025-11-09T00:00:00Z","timestamp":1762646400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,11,9]],"date-time":"2025-11-09T00:00:00Z","timestamp":1762646400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Cybersecurity"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Speech recognition technology has brought revolutionary changes to our lives, but existing work has demonstrated the feasibility of using adversarial examples (AEs) to mislead speech recognition systems. Most existing adversarial attacks are designed for white-box or grey-box systems and they are ineffective against strict black-box scenarios where only the recognition results can be queried. The known black-box attack methods all add perturbations limited to the bounded\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$L_{p}$$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:msub>\n                            <mml:mi>L<\/mml:mi>\n                            <mml:mi>p<\/mml:mi>\n                          <\/mml:msub>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    neighborhood in the time domain and they neglected the relationship between the perturbations and the carrier audios. Therefore, although AEs they generated cannot be recognized by humans as the target attack commands, they sound very unnatural and unsmooth. The abnormal sense in auditory perception is easy to alarm the victim and can be used for defense. To address this limitation, we propose a novel adversarial attack against black-box systems. Unlike previous works adding perturbations in the time domain, we extract the effective adversarial feature in the frequency domain and modify the spectral energy distribution of the carrier electronic music in a smooth way to maintain the the carriers\u2019 original timbre. Meanwhile, we dynamically adjust the insert position and duration of the adversarial feature according to the music\u2019s rhythm. Our higher requirement for imperceptibility is that AEs should sound natural and smooth. We evaluated our method on eight commercial black-box speech recognition systems (including five digital and three physical systems). Our AEs can achieved the 100% attack success rate and have the outstanding imperceptibility compared to the state-of-the-art black-box attacks. Our AEs are on par with normal music in terms of auditory naturalness and smoothness in the digital world, which requires only about 500 queries. Compared to the existing works, only 7.9% think their AEs are normal, at least 52.1% could not distinguish between our AEs and normal music in the physical world. The demos and code are available at the open source repository\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/Cybersecurity-Electronic-Music-Assassin\/Electronic-Music-Assassin.git.\" ext-link-type=\"uri\">https:\/\/github.com\/Cybersecurity-Electronic-Music-Assassin\/Electronic-Music-Assassin.git.<\/jats:ext-link>\n                  <\/jats:p>","DOI":"10.1186\/s42400-025-00374-5","type":"journal-article","created":{"date-parts":[[2025,11,9]],"date-time":"2025-11-09T02:02:35Z","timestamp":1762653755000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Electronic music assassin: towards imperceptible physical adversarial attacks against black-box automatic speech recognitions"],"prefix":"10.1186","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7111-4772","authenticated-orcid":false,"given":"Ruiyuan","family":"Li","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,11,9]]},"reference":[{"key":"374_CR1","doi-asserted-by":"crossref","unstructured":"Aalto D, Malinen J, Vainio M (2018) Formants. In: Oxford Research Encyclopedia of Linguistics,","DOI":"10.1093\/acrefore\/9780199384655.013.419"},{"key":"374_CR2","doi-asserted-by":"publisher","unstructured":"Abdullah H, Garcia W, Peeters C, Traynor P, Butler KRB, Wilson J (2019) Practical hidden voice attacks against speech and speaker recognition systems. In: Proceedings 2019 Network and Distributed System Security Symposium https:\/\/doi.org\/10.14722\/ndss.2019.23362","DOI":"10.14722\/ndss.2019.23362"},{"issue":"8","key":"374_CR3","doi-asserted-by":"publisher","first-page":"1928","DOI":"10.3390\/electronics12081928","volume":"12","author":"SS Alchekov","year":"2023","unstructured":"Alchekov SS, Al-Absi MA, Al-Absi AA, Lee HJ (2023) Inaudible attack on ai speakers. Electronics 12(8):1928","journal-title":"Electronics"},{"key":"374_CR4","unstructured":"Aliyun (2023) Alibaba Speech Recognition. https:\/\/ai.aliyun.com\/nls"},{"key":"374_CR5","unstructured":"Amazon (2023a) Alexa Auto SDK. https:\/\/developer.amazon.com\/en-US\/alexa\/devices\/alexa-built-in\/developme nt-resources\/auto-sdk"},{"key":"374_CR6","unstructured":"Amazon (2023b) Amazon Alexa. https:\/\/developer.amazon.com\/en-US\/alexa"},{"key":"374_CR7","unstructured":"Amodei D, Ananthanarayanan S, Anubhai R, Bai J, Battenberg E, Case C, Casper J, Catanzaro B, Cheng Q, Chen G et al (2016) Deep speech 2: End-to-end speech recognition in english and mandarin. In: International Conference on Machine Learning, pp. 173\u2013182 PMLR"},{"key":"374_CR8","unstructured":"Bengio Y, Ducharme R, Vincent P (2000) A neural probabilistic language model. Advances in neural information processing systems 13"},{"key":"374_CR9","doi-asserted-by":"publisher","first-page":"204","DOI":"10.1016\/j.neuroimage.2014.07.005","volume":"101","author":"GM Bidelman","year":"2014","unstructured":"Bidelman GM, Grall J (2014) Functional organization for musical consonance and tonal pitch hierarchy in human auditory cortex. Neuroimage 101:204\u2013214","journal-title":"Neuroimage"},{"issue":"42","key":"374_CR10","doi-asserted-by":"publisher","first-page":"13165","DOI":"10.1523\/JNEUROSCI.3900-09.2009","volume":"29","author":"GM Bidelman","year":"2009","unstructured":"Bidelman GM, Krishnan A (2009) Neural correlates of consonance, dissonance, and the hierarchy of musical pitch in the human brainstem. J Neurosci 29(42):13165\u201313171","journal-title":"J Neurosci"},{"issue":"1","key":"374_CR11","doi-asserted-by":"publisher","first-page":"118","DOI":"10.1525\/mp.2017.35.1.118","volume":"35","author":"DL Bowling","year":"2017","unstructured":"Bowling DL, Hoeschele M, Gill KZ, Fitch WT (2017) The nature and nurture of musical consonance. Music Percept Interdiscip J 35(1):118\u2013121","journal-title":"Music Percept Interdiscip J"},{"key":"374_CR12","doi-asserted-by":"crossref","unstructured":"Cao H, Jiang H, Liu D, Wang R, Min G, Liu J, Lui Dustdar S JC, (2022) Liveprobe: exploring continuous voice liveness detection via phonemic energy response patterns. IEEE Internet Things J 10(8):7215\u20137228","DOI":"10.1109\/JIOT.2022.3228819"},{"key":"374_CR13","doi-asserted-by":"crossref","unstructured":"Carlini N, Wagner D (2017) Towards evaluating the robustness of neural networks. In: 2017 Ieee Symposium on Security and Privacy (sp), pp. 39\u201357 Ieee","DOI":"10.1109\/SP.2017.49"},{"key":"374_CR14","doi-asserted-by":"crossref","unstructured":"Carlini N, Wagner D (2018) Audio adversarial examples: Targeted attacks on speech-to-text. Cornell University - arXiv, Cornell University - arXiv","DOI":"10.1109\/SPW.2018.00009"},{"key":"374_CR15","unstructured":"Carlini N, Mishra P, Vaidya T, Zhang Y, Sherr M, Shields C, Wagner D, Zhou W (2016) Hidden voice commands. USENIX Security Symposium, USENIX Security Symposium"},{"key":"374_CR16","doi-asserted-by":"crossref","unstructured":"Chen G, Chenb S, Fan L, Du X, Zhao Z, Song F, Liu Y (2021) Who is real bob? adversarial attacks on speaker recognition systems. In: 2021 IEEE Symposium on Security and Privacy (SP), pp. 694\u2013711 . IEEE","DOI":"10.1109\/SP40001.2021.00004"},{"key":"374_CR17","doi-asserted-by":"publisher","unstructured":"Chen P-Y, Zhang H, Sharma Y, Yi J, Hsieh C-J (2017) Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. In: Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security https:\/\/doi.org\/10.1145\/3128572.3140448","DOI":"10.1145\/3128572.3140448"},{"key":"374_CR18","doi-asserted-by":"crossref","unstructured":"Chen T, Shangguan L, Li Z, Jamieson K (2020a) Metamorph: Injecting inaudible commands into over-the-air voice controlled systems. In: Network and Distributed Systems Security (NDSS) Symposium","DOI":"10.14722\/ndss.2020.23055"},{"key":"374_CR19","doi-asserted-by":"publisher","unstructured":"Chen T, Shangguan L, Li Z, Jamieson K (2020b) Metamorph: Injecting inaudible commands into over-the-air voice controlled systems. In: Proceedings 2020 Network and Distributed System Security Symposium https:\/\/doi.org\/10.14722\/ndss.2020.23055","DOI":"10.14722\/ndss.2020.23055"},{"key":"374_CR20","unstructured":"Chen Y, Yuan X, Zhang J, Zhao Y, Zhang S, Chen K, Wang X (2020c) $$\\{$$Devil\u2019s$$\\}$$ whisper: A general approach for physical adversarial attacks against commercial black-box speech recognition devices. In: 29th USENIX Security Symposium (USENIX Security 20), pp. 2667\u20132684"},{"key":"374_CR21","doi-asserted-by":"publisher","DOI":"10.1093\/acprof:oso\/9780199553792.001.0001","author":"D Clarke","year":"2011","unstructured":"Clarke D, Clarke E (2011) Music perception and musical consciousness. Music Conscious Philosoph. https:\/\/doi.org\/10.1093\/acprof:oso\/9780199553792.001.0001","journal-title":"Music Conscious Philosoph"},{"issue":"1","key":"374_CR22","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1177\/0305735600281002","volume":"28","author":"M Costa","year":"2000","unstructured":"Costa M, Ricci Bitti PE, Bonfiglioli L (2000) Psychological connotations of harmonic musical intervals. Psychol Music 28(1):4\u201322","journal-title":"Psychol Music"},{"key":"374_CR23","doi-asserted-by":"publisher","unstructured":"Diao W, Liu X, Zhou Z, Zhang K (2014) Your voice assistant is mine: How to abuse speakers to steal information and control your phone. In: Proceedings of the 4th ACM Workshop on Security and Privacy in Smartphones & Mobile Devices https:\/\/doi.org\/10.1145\/2666620.2666623","DOI":"10.1145\/2666620.2666623"},{"key":"374_CR24","doi-asserted-by":"publisher","unstructured":"Dong Y, Su H, Wu B, Li Z, Liu W, Zhang T, Zhu J (2019) Efficient decision-based black-box adversarial attacks on face recognition. In: 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) https:\/\/doi.org\/10.1109\/cvpr.2019.00790","DOI":"10.1109\/cvpr.2019.00790"},{"key":"374_CR25","doi-asserted-by":"publisher","unstructured":"Du T, Ji S, Li J, GuQ, Wang T, Beyah R (2020) Sirenattack: Generating adversarial audio for end-to-end acoustic systems. In: Proceedings of the 15th ACM Asia Conference on Computer and Communications Security https:\/\/doi.org\/10.1145\/3320269.3384733","DOI":"10.1145\/3320269.3384733"},{"issue":"3","key":"374_CR26","doi-asserted-by":"publisher","first-page":"361","DOI":"10.1016\/S0959-440X(96)80056-X","volume":"6","author":"SR Eddy","year":"1996","unstructured":"Eddy SR (1996) Hidden markov models. Curr Opin Struct Biol 6(3):361\u2013365","journal-title":"Curr Opin Struct Biol"},{"key":"374_CR27","doi-asserted-by":"crossref","unstructured":"Gao Z, Li Z, Wang J, Luo H, Shi X, Chen M, Li Y, Zuo L, Du Z, Xiao Z, Zhang S (2023) Funasr: A fundamental end-to-end speech recognition toolkit. In: INTERSPEECH","DOI":"10.21437\/Interspeech.2023-1428"},{"key":"374_CR28","unstructured":"Goodfellow IJ, Shlens J, Szegedy C (2014) Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572"},{"key":"374_CR29","unstructured":"Google (2023a) Google assistant, your own personal google. https:\/\/assistant.google.com\/"},{"key":"374_CR30","unstructured":"Google (2023b) Google Cloud Speech-to-Text Service. https:\/\/cloud.google.com\/speech-to-text"},{"key":"374_CR31","doi-asserted-by":"publisher","unstructured":"Graves A, Mohamed A-r, Hinton G (2013) Speech recognition with deep recurrent neural networks. In: 2013 IEEE International Conference on Acoustics, Speech and Signal Processing https:\/\/doi.org\/10.1109\/icassp.2013.6638947","DOI":"10.1109\/icassp.2013.6638947"},{"key":"374_CR32","doi-asserted-by":"publisher","unstructured":"Graves A, Fern\u00e1ndez S, Gomez F, Schmidhuber J (2006) Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In: Proceedings of the 23rd International Conference on Machine Learning - ICML \u201906 https:\/\/doi.org\/10.1145\/1143844.1143891","DOI":"10.1145\/1143844.1143891"},{"issue":"2","key":"374_CR33","doi-asserted-by":"publisher","first-page":"236","DOI":"10.1109\/TASSP.1984.1164317","volume":"32","author":"D Griffin","year":"1984","unstructured":"Griffin D, Lim J (1984) Signal estimation from modified short-time Fourier transform. IEEE Trans Acoust Speech Signal Process 32(2):236\u2013243","journal-title":"IEEE Trans Acoust Speech Signal Process"},{"key":"374_CR34","unstructured":"Hannun A, Case C, Casper J, Catanzaro B, Diamos G, Elsen E, Prenger R, Satheesh S, Sengupta S, Coates A et al (2014) Deep speech: Scaling up end-to-end speech recognition. arXiv preprint arXiv:1412.5567"},{"issue":"6","key":"374_CR35","doi-asserted-by":"publisher","first-page":"82","DOI":"10.1109\/MSP.2012.2205597","volume":"29","author":"G Hinton","year":"2012","unstructured":"Hinton G, Deng L, Yu D, Dahl GE, Mohamed A, Jaitly N, Senior A, Vanhoucke V, Nguyen P, Sainath TN et al (2012) Deep neural networks for acoustic modeling in speech recognition: the shared views of four research groups. IEEE Signal Process Mag 29(6):82\u201397","journal-title":"IEEE Signal Process Mag"},{"key":"374_CR36","doi-asserted-by":"publisher","first-page":"778","DOI":"10.1016\/j.proeng.2014.03.054","volume":"69","author":"S Husnjak","year":"2014","unstructured":"Husnjak S, Perakovic D (2014) Jovovic I Possibilities of using speech recognition systems of smart terminal devices in traffic environment. Proc Eng 69:778\u2013787","journal-title":"Proc Eng"},{"key":"374_CR37","unstructured":"iFlytek (2023) iFlytek Speech-to-Text. https:\/\/www.xfyun.cn\/services\/voicedictation"},{"key":"374_CR38","unstructured":"Inria N (2016) The cma evolution strategy: A tutorial"},{"issue":"S1","key":"374_CR39","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1121\/1.1995189","volume":"57","author":"F Itakura","year":"1975","unstructured":"Itakura F (1975) Line spectrum representation of linear predictor coefficients of speech signals. J Acoust Soc Am 57(S1):35\u201335","journal-title":"J Acoust Soc Am"},{"key":"374_CR40","doi-asserted-by":"publisher","unstructured":"Izbassarova A, Duisembay A, James AP (2020) Speech Recognition Application Using Deep Learning Neural Network, pp. 69\u201379. https:\/\/doi.org\/10.1007\/978-3-030-14524-8_5","DOI":"10.1007\/978-3-030-14524-8_5"},{"key":"374_CR41","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4471-0249-6","volume-title":"Interval analysis","author":"L Jaulin","year":"2001","unstructured":"Jaulin L, Kieffer M, Didrit O, Walter E, Jaulin L, Kieffer M, Didrit O, Walter \u00c9 (2001) Interval analysis. Springer"},{"key":"374_CR42","doi-asserted-by":"publisher","DOI":"10.1109\/temc.2015.2463089","author":"C Kasmi","year":"2015","unstructured":"Kasmi C, Lopes Esteves J (2015) Iemi threats for information security: remote command injection on modern smartphones. IEEE Trans Electromagn Compat. https:\/\/doi.org\/10.1109\/temc.2015.2463089","journal-title":"IEEE Trans Electromagn Compat"},{"key":"374_CR43","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9781139165372","volume-title":"An introduction to harmonic analysis","author":"Y Katznelson","year":"2004","unstructured":"Katznelson Y (2004) An introduction to harmonic analysis. Cambridge University Press"},{"key":"374_CR44","doi-asserted-by":"publisher","unstructured":"Kim C, Misra A, Chin K, Hughes T, Narayanan A, Sainath T.N, Bacchiani M (2017) Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in google home. In: Interspeech 2017. https:\/\/doi.org\/10.21437\/interspeech.2017-1510","DOI":"10.21437\/interspeech.2017-1510"},{"key":"374_CR45","unstructured":"Kingma, D, Ba J (2014) Adam: A method for stochastic optimization. arXiv: Learning,arXiv: Learning"},{"key":"374_CR46","unstructured":"Kumar D, Paccagnella R, Murley P, Hennenfent E, Mason J, Bates A, Bailey M (2018) Skill squatting attacks on amazon alexa. USENIX Security Symposium, USENIX Security Symposium"},{"key":"374_CR47","doi-asserted-by":"publisher","unstructured":"Kumatani K, Panchapagesan S, Wu M, Kim M, Strom N, Tiwari G, Mandai A (2017) Direct modeling of raw audio with dnns for wake word detection. In: 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . https:\/\/doi.org\/10.1109\/asru.2017.8268943","DOI":"10.1109\/asru.2017.8268943"},{"key":"374_CR48","unstructured":"Li J, Qu S, Li X, Szurley J, Kolter JZ, Metze F (2019) Adversarial music: Real world audio adversary against wake-word detection system. Neural Information Processing Systems, Neural Information Processing Systems"},{"key":"374_CR49","unstructured":"Meng Y, Li J, Pillari M, Deopujari A, Brennan L, Shamsie H, Zhu H, Tian Y (2022) Your microphone array retains your identity: A robust voice liveness detection system for smart speakers. In: 31st USENIX Security Symposium (USENIX Security 22), pp. 1077\u20131094"},{"key":"374_CR50","unstructured":"Microsoft (2023) Microsoft Azure Speech Service. https:\/\/learn.microsoft.com\/zh-cn\/azure\/ai-services\/speech-service\/"},{"key":"374_CR51","doi-asserted-by":"crossref","unstructured":"Moosavi-Dezfooli S-M, Fawzi A, Frossard P (2016) Deepfool: a simple and accurate method to fool deep neural networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2574\u20132582","DOI":"10.1109\/CVPR.2016.282"},{"key":"374_CR52","unstructured":"Muda L, Begam M, Elamvazuthi I (2010) Voice recognition algorithms using mel frequency cepstral coefficient (mfcc) and dynamic time warping (dtw) techniques. arXiv preprint arXiv:1003.4083"},{"issue":"11","key":"374_CR53","doi-asserted-by":"publisher","first-page":"476","DOI":"10.1016\/j.cub.2010.03.044","volume":"20","author":"CJ Plack","year":"2010","unstructured":"Plack CJ (2010) Musical consonance: the importance of harmonicity. Curr Biol 20(11):476\u2013478","journal-title":"Curr Biol"},{"key":"374_CR54","unstructured":"Povey D, Ghoshal A, Boulianne G, Burget L, Glembek O, Goel N, Hannemann M, Motlicek P, Qian Y, Schwarz P, Silovsky J, Stemmer G, Vesely K (2011) The kaldi speech recognition toolkit. IEEE Automatic Speech Recognition and Understanding Workshop, IEEE Automatic Speech Recognition and Understanding Workshop"},{"key":"374_CR55","unstructured":"Prolific (2023) Prolific - Quickly find research participants you can trust. https:\/\/www.prolific.com\/"},{"key":"374_CR56","unstructured":"Qin Y, Carlini N, Goodfellow I, Cottrell G, Raffel C (2019) Imperceptible, robust, and targeted adversarial examples for automatic speech recognition. Cornell University - arXiv, Cornell University - arXiv"},{"key":"374_CR57","doi-asserted-by":"publisher","unstructured":"Rao K, Sak H, Prabhavalkar R (2017) Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer. In: 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . https:\/\/doi.org\/10.1109\/asru.2017.8268935","DOI":"10.1109\/asru.2017.8268935"},{"key":"374_CR58","doi-asserted-by":"publisher","first-page":"659","DOI":"10.1007\/978-0-387-73003-5_196","volume-title":"Encyclopedia of Biometrics","author":"Douglas Reynolds","year":"2009","unstructured":"Reynolds Douglas (2009) Gaussian Mixture Models. In: Li Stan Z., Jain Anil (eds) Encyclopedia of Biometrics. Springer US, Boston, MA, pp 659\u2013663. https:\/\/doi.org\/10.1007\/978-0-387-73003-5_196"},{"key":"374_CR59","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1016\/B978-012213564-4\/50006-8","volume-title":"The psychology of music","author":"JC Risset","year":"1999","unstructured":"Risset J-C, Wessel DL (1999) Exploration of timbre by analysis and synthesis. In: The psychology of music. Elsevier, pp 113\u2013169"},{"key":"374_CR60","doi-asserted-by":"publisher","unstructured":"Roy N, Hassanieh H, Roy\u00a0Choudhury R (2017) Backdoor: Making microphones hear inaudible sounds. In: Proceedings of the 15th Annual International Conference on Mobile Systems, Applications, and Services . https:\/\/doi.org\/10.1145\/3081333.3081366","DOI":"10.1145\/3081333.3081366"},{"key":"374_CR61","doi-asserted-by":"crossref","unstructured":"Samizade S, Tan Z-H, Shen C, Guan X (2020) Adversarial example detection by classification for deep speech recognition. In: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 3102\u20133106. IEEE, ???","DOI":"10.1109\/ICASSP40776.2020.9054750"},{"key":"374_CR62","doi-asserted-by":"crossref","unstructured":"Sch\u00f6nherr L, Kohls K, Zeiler S, Holz T, Kolossa D (2018) Adversarial attacks against automatic speech recognition systems via psychoacoustic hiding. Cornell University - arXiv, Cornell University - arXiv","DOI":"10.14722\/ndss.2019.23288"},{"key":"374_CR63","doi-asserted-by":"crossref","unstructured":"Sch\u00f6nherr L, Zeiler S, Holz T, Kolossa D (2019) Robust over-the-air adversarial examples against automatic speech recognition systems","DOI":"10.14722\/ndss.2019.23288"},{"key":"374_CR64","doi-asserted-by":"crossref","unstructured":"Sch\u00f6nherr L, Eisenhofer T, Zeiler S, Holz T, Kolossa D (2020) Imperio: Robust over-the-air adversarial examples for automatic speech recognition systems. In: Annual Computer Security Applications Conference, pp. 843\u2013855","DOI":"10.1145\/3427228.3427276"},{"issue":"1","key":"374_CR65","doi-asserted-by":"publisher","first-page":"153","DOI":"10.1016\/j.dsp.2007.12.004","volume":"19","author":"E Sejdi\u0107","year":"2009","unstructured":"Sejdi\u0107 E, Djurovi\u0107 I, Jiang J (2009) Time-frequency feature representation using energy concentration: an overview of recent advances. Digital Signal Process 19(1):153\u2013183","journal-title":"Digital Signal Process"},{"key":"374_CR66","unstructured":"Siri (2023) Apple Siri. https:\/\/www.apple.com\/siri\/"},{"key":"374_CR67","unstructured":"Sobel I (2014) History and definition of the sobel operator. Retrieved from the World Wide Web 1505"},{"key":"374_CR68","unstructured":"Statista(2023) Number of voice assistant users in the united states from 2022 to 2026. https:\/\/www.statista.c om\/statistics\/1299985\/voice-assistant-users-us"},{"issue":"29","key":"374_CR69","doi-asserted-by":"publisher","first-page":"1429","DOI":"10.1098\/rsif.2008.0143","volume":"5","author":"LL Stone","year":"2008","unstructured":"Stone LL (2008) Perception of musical consonance and dissonance: an outcome of neural synchronization. J R Soc Interface 5(29):1429\u20131434","journal-title":"J R Soc Interface"},{"key":"374_CR70","unstructured":"Sugawara T, Cyr B, Rampazzi S, Genkin D, Fu K (2020) Light commands: Laser-based audio injection attacks on voice-controllable systems. Cornell University - arXiv, Cornell University - arXiv"},{"key":"374_CR71","unstructured":"Szegedy C, Zaremba W, Sutskever I, Bruna J, Erhan D, Goodfellow I, Fergus R (2013) Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199"},{"issue":"2","key":"374_CR72","doi-asserted-by":"publisher","first-page":"345","DOI":"10.1037\/0033-2909.113.2.345","volume":"113","author":"AH Takeuchi","year":"1993","unstructured":"Takeuchi AH, Hulse SH (1993) Absolute pitch. Psychol Bull 113(2):345","journal-title":"Psychol Bull"},{"key":"374_CR73","doi-asserted-by":"publisher","unstructured":"Taori R, Kamsetty A, Chu B, Vemuri N (2019) Targeted adversarial examples for black box audio systems. In: 2019 IEEE Security and Privacy Workshops (SPW) . https:\/\/doi.org\/10.1109\/spw.2019.00016","DOI":"10.1109\/spw.2019.00016"},{"key":"374_CR74","unstructured":"Tencent (2023) Tencent Short Speech Recognition. https:\/\/cloud.tencent.com\/product\/asr"},{"issue":"3","key":"374_CR75","doi-asserted-by":"publisher","first-page":"276","DOI":"10.2307\/40285261","volume":"1","author":"E Terhardt","year":"1984","unstructured":"Terhardt E (1984) The concept of musical consonance: a link between music and psychoacoustics. Music Percept 1(3):276\u2013295","journal-title":"Music Percept"},{"key":"374_CR76","doi-asserted-by":"crossref","unstructured":"Vacher M, Fleury A, Portet F, Serignat J, Noury N (2009) Complete Sound and Speech Recognition System for Health Smart Homes: Application to the Recognition of Activities of Daily Living, pp. 603\u2013614. Springer","DOI":"10.5772\/7596"},{"key":"374_CR77","unstructured":"Vaidya T, Zhang Y, Sherr M, Shields C (2015) Cocaine noodles: exploiting the gap between human and machine speech recognition"},{"key":"374_CR78","doi-asserted-by":"crossref","unstructured":"Wang P(2020) Research and design of smart home speech recognition system based on deep learning. In: 2020 International Conference on Computer Vision, Image and Deep Learning (CVIDL)","DOI":"10.1109\/CVIDL51233.2020.00-98"},{"key":"374_CR79","doi-asserted-by":"publisher","DOI":"10.1109\/tifs.2020.3026543","author":"Q Wang","year":"2021","unstructured":"Wang Q, Zheng B, Li Q, Shen C, Ba Z (2021) Towards query-efficient adversarial attacks against automatic speech recognition systems. IEEE Trans Inform Forensics Sec. https:\/\/doi.org\/10.1109\/tifs.2020.3026543","journal-title":"IEEE Trans Inform Forensics Sec"},{"key":"374_CR80","doi-asserted-by":"publisher","first-page":"896","DOI":"10.1109\/tifs.2020.3026543","volume":"16","author":"Q Wang","year":"2021","unstructured":"Wang Q, Zheng B, Li Q, Shen C, Ba Z (2021) Towards query-efficient adversarial attacks against automatic speech recognition systems. IEEE Trans Inf Forensics Secur 16:896\u2013908. https:\/\/doi.org\/10.1109\/tifs.2020.3026543","journal-title":"IEEE Trans Inf Forensics Secur"},{"key":"374_CR81","doi-asserted-by":"publisher","DOI":"10.1007\/bf00175354","author":"D Whitley","year":"1994","unstructured":"Whitley D (1994) A genetic algorithm tutorial. Stat Comput. https:\/\/doi.org\/10.1007\/bf00175354","journal-title":"Stat Comput"},{"key":"374_CR82","unstructured":"Wu X, Ma S, Shen C, Lin C, Wang Q, Li Q, Rao Y (2023) $$\\{$$KENKU$$\\}$$: Towards efficient and stealthy black-box adversarial attacks against $$\\{$$ASR$$\\}$$ systems. In: 32nd USENIX Security Symposium (USENIX Security 23), pp. 247\u2013264"},{"key":"374_CR83","unstructured":"Xia Q, Chen Q, Xu S (2023) $$\\{$$Near-Ultrasound$$\\}$$ inaudible trojan (nuit): Exploiting your speaker to attack your microphone. In: 32nd USENIX Security Symposium (USENIX Security 23), pp. 4589\u20134606"},{"key":"374_CR84","doi-asserted-by":"crossref","unstructured":"Yakura H, Sakuma J (2018) Robust audio adversarial example for a physical attack. arXiv preprint arXiv:1810.11793","DOI":"10.24963\/ijcai.2019\/741"},{"key":"374_CR85","doi-asserted-by":"publisher","unstructured":"Yakura H, Sakuma J (2019) Robust audio adversarial example for a physical attack. In: Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence . https:\/\/doi.org\/10.24963\/ijcai.2019\/741","DOI":"10.24963\/ijcai.2019\/741"},{"key":"374_CR86","unstructured":"Yang Z, Li B, Chen P-Y, Song D (2018) Characterizing audio adversarial examples using temporal dependency. arXiv preprint arXiv:1809.10875"},{"key":"374_CR87","unstructured":"Yu Z, Chang Y, Zhang N, Xiao C (2023) $$\\{$$SMACK$$\\}$$: Semantically meaningful adversarial audio attack. In: 32nd USENIX Security Symposium (USENIX Security 23), pp. 3799\u20133816"},{"key":"374_CR88","unstructured":"Yuan X, Chen Y, Zhao Y, Long Y, Liu X-K, Chen K, Zhang S, Huang H, Wang XF, Gunter C (2018) Commandersong: A systematic approach for practical adversarial voice recognition. Cornell University - arXiv, Cornell University - arXiv"},{"key":"374_CR89","doi-asserted-by":"crossref","unstructured":"Zeng Q, Su J, Fu C, Kayas G, Luo L, Du X, Tan CC, Wu J (2019) A multiversion programming inspired approach to detecting audio adversarial examples. In: 2019 49th annual IEEE\/IFIP international conference on dependable systems and networks (DSN), pp 39\u201351","DOI":"10.1109\/DSN.2019.00019"},{"key":"374_CR90","doi-asserted-by":"publisher","unstructured":"Zhang G, Yan C, Ji X, Zhang T, Zhang T, Xu W (2017) Dolphinattack: Inaudible voice commands. In: Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security . https:\/\/doi.org\/10.1145\/3133956.3134052","DOI":"10.1145\/3133956.3134052"},{"key":"374_CR91","doi-asserted-by":"publisher","unstructured":"Zhang N, Mi X, Feng X, Wang X, Tian Y, Qian F (2019) Dangerous skills: Understanding and mitigating security risks of voice-controlled third-party functions on virtual personal assistant systems. In: 2019 IEEE Symposium on Security and Privacy (SP) . https:\/\/doi.org\/10.1109\/sp.2019.00016","DOI":"10.1109\/sp.2019.00016"},{"key":"374_CR92","doi-asserted-by":"crossref","unstructured":"Zheng B, Jiang P, Wang Q, Li Q, Shen C, Wang C, Ge Y, Teng Q, Zhang S (2021) Black-box adversarial attacks on commercial speech platforms with minimal information. In: Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security, pp. 86\u2013107","DOI":"10.1145\/3460120.3485383"},{"key":"374_CR93","unstructured":"Zhuang J, Tang T, Ding Y, Tatikonda S, Dvornek N, Papademetris X, Duncan J (2020) Adabelief optimizer: Adapting stepsizes by the belief in observed gradients. Cornell University - arXiv, Cornell University - arXiv"}],"container-title":["Cybersecurity"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s42400-025-00374-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s42400-025-00374-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s42400-025-00374-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,9]],"date-time":"2025-11-09T02:02:46Z","timestamp":1762653766000},"score":1,"resource":{"primary":{"URL":"https:\/\/cybersecurity.springeropen.com\/articles\/10.1186\/s42400-025-00374-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,9]]},"references-count":93,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["374"],"URL":"https:\/\/doi.org\/10.1186\/s42400-025-00374-5","relation":{},"ISSN":["2523-3246"],"issn-type":[{"value":"2523-3246","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,9]]},"assertion":[{"value":"27 December 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 February 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 November 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"83"}}