{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,9]],"date-time":"2026-04-09T14:44:19Z","timestamp":1775745859094,"version":"3.50.1"},"reference-count":27,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2022,3,20]],"date-time":"2022-03-20T00:00:00Z","timestamp":1647734400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61902235"],"award-info":[{"award-number":["61902235"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Shanghai Chenguang Program","award":["19CG46"],"award-info":[{"award-number":["19CG46"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Benefiting from the rapid development of computer hardware and big data, deep neural networks (DNNs) have been widely applied in commercial speaker recognition systems, achieving a kind of symmetry between \u201cmachine-learning-as-a-service\u201d providers and consumers. However, this symmetry is threatened by attackers whose goal is to illegally steal and use the service. It is necessary to protect these DNN models from symmetry breaking, i.e., intellectual property (IP) infringement, which motivated the authors to present a black-box watermarking method for IP protection of the speaker recognition model in this paper. The proposed method enables verification of the ownership of the target marked model by querying the model with a set of carefully crafted trigger audio samples, without knowing the internal details of the model. To achieve this goal, the proposed method marks the host model by training it with normal audio samples and carefully crafted trigger audio samples. The trigger audio samples are constructed by adding a trigger signal in the frequency domain of normal audio samples, which enables the trigger audio samples to not only resist against malicious attack but also avoid introducing noticeable distortion. In order to not impair the performance of the speaker recognition model on its original task, a new label is assigned to all the trigger audio samples. The experimental results show that the proposed black-box DNN watermarking method can not only reliably protect the intellectual property of the speaker recognition model but also maintain the performance of the speaker recognition model on its original task, which verifies the superiority and maintains the symmetry between \u201cmachine-learning-as-a-service\u201d providers and consumers.<\/jats:p>","DOI":"10.3390\/sym14030619","type":"journal-article","created":{"date-parts":[[2022,3,20]],"date-time":"2022-03-20T21:37:17Z","timestamp":1647812237000},"page":"619","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["Protecting the Intellectual Property of Speaker Recognition Model by Black-Box Watermarking in the Frequency Domain"],"prefix":"10.3390","volume":"14","author":[{"given":"Yumin","family":"Wang","sequence":"first","affiliation":[{"name":"School of Communication and Information Engineering, Shanghai University, Shanghai 200444, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1599-7232","authenticated-orcid":false,"given":"Hanzhou","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Communication and Information Engineering, Shanghai University, Shanghai 200444, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,3,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_2","first-page":"1097","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Xiong, W., Wu, L., Alleva, F., Droppo, J., Huang, X., and Stolcke, A. (2018, January 15\u201320). The Microsoft 2017 conversational speech recognition system. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461870"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Bahdanau, D., Chorowski, J., Serdyuk, D., Brakel, P., and Bengio, Y. (2016, January 20\u201325). End-to-end attention-based large vocabulary speech recognition. Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472618"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1109\/MCI.2018.2840738","article-title":"Recent trends in deep learning based natural language processing","volume":"13","author":"Young","year":"2018","journal-title":"IEEE Comput. Intell. Mag."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Uchida, Y., Nagai, Y., Sakazawa, S., and Satoh, S. (2017, January 6\u20139). Embedding watermarks into deep neural networks. Proceedings of the ICMR\u201917: International Conference on Multimedia Retrieval, Bucharest, Romania.","DOI":"10.1145\/3078971.3078974"},{"key":"ref_7","unstructured":"Rouhani, B.D., Chen, H., and Koushanfar, F. (2019, January 13\u201317). Deepsigns: An end-to-end watermarking framework for ownership protection of deep neural ne works. Proceedings of the ASPLOS\u201919: Architectural Support for Programming Languages and Operating Systems, Providence, RI, USA."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Wang, T., and Kerschbaum, F. (2019). Robust and undetectable white-box watermarks for deep neural networks. arXiv.","DOI":"10.1109\/ICASSP.2019.8682202"},{"key":"ref_9","unstructured":"Adi, Y., Baum, C., Cisse, M., Pinkas, B., and Keshet, J. (2018, January 15\u201317). Turning your weakness into a strength: Watermarking deep neural networks by backdooring. Proceedings of the 27th USENIX Security Symposium (USENIX Security 18), Baltimore, MD, USA."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhang, J., Gu, Z., Jang, J., Wu, H., Stoecklin, M.P., Huang, H., and Molloy, I. (2018, January 4). Protecting intellectual property of deep neural networks with watermarking. Proceedings of the ASIA CCS\u201918: ACM Asia Conference on Computer and Communications Security, Incheon, Korea.","DOI":"10.1145\/3196494.3196550"},{"key":"ref_11","unstructured":"Chen, H., Rouhani, B.D., and Koushanfar, F. (2019). Blackmarks: Blackbox multibit watermarking for deep neural networks. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Guo, J., and Potkonjak, M. (2018, January 5\u20138). Watermarking deep neural networks for embedded systems. Proceedings of the 2018 IEEE\/ACM International Conference on Computer-Aided Design (ICCAD), San Diego, CA, USA.","DOI":"10.1145\/3240765.3240862"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, J., Wu, H., Zhang, X., and Yao, Y. (2020). Watermarking in deep neural networks via error back-propagation. Electronic Imaging, Media Watermarking, Security and Forensics, Society for Imaging Science and Technology.","DOI":"10.2352\/ISSN.2470-1173.2020.4.MWSF-022"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zhao, X., Yao, Y., Wu, H., and Zhang, X. (2021, January 7\u201310). Structural watermarking to deep neural networks via network channel pruning. Proceedings of the 2021 IEEE International Workshop on Information Forensics and Security (WIFS), Montpellier, France.","DOI":"10.1109\/WIFS53200.2021.9648376"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Li, Y., Wang, H., and Barni, M. (2021). A survey of deep neural network watermarking techniques. arXiv.","DOI":"10.1016\/j.neucom.2021.07.051"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhao, X., Wu, H., and Zhang, X. (2021, January 28\u201329). Watermarking graph neural networks by random graphs. Proceedings of the 2021 9th International Symposium on Digital Forensics and Security (ISDFS), Elazig, Turkey.","DOI":"10.1109\/ISDFS52919.2021.9486352"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"2591","DOI":"10.1109\/TCSVT.2020.3030671","article-title":"Watermarking neural networks with watermarked images","volume":"31","author":"Wu","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Boenisch, F. (2020). A survey on model watermarking neural networks. arXiv.","DOI":"10.3389\/fdata.2021.729663"},{"key":"ref_19","first-page":"2013","article-title":"A study on spatial and transform domain watermarking techniques","volume":"71","author":"Dabas","year":"2013","journal-title":"Int. J. Comput. Appl."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Kong, Y., and Zhang, J. (2019). Adversarial audio: A new information hiding method and backdoor for dnn-based speech recognition models. arXiv.","DOI":"10.21437\/Interspeech.2020-1294"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Li, M., Zhong, Q., Zhang, L.Y., Du, Y., Zhang, J., and Xiang, Y. (January, January 29). Protecting the intellectual property of deep neural networks with watermarking: The frequency domain approach. Proceedings of the 2020 IEEE 19th International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom), Guangzhou, China.","DOI":"10.1109\/TrustCom50675.2020.00062"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zong, Q., and Guo, W. (2012, January 21\u201323). A speech information hiding algorithm based on the energy difference between the frequency band. Proceedings of the 2012 2nd International Conference on Consumer Electronics, Communications and Networks (CECNet), Yichang, China.","DOI":"10.1109\/CECNet.2012.6201965"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"3343","DOI":"10.1007\/s11042-016-3934-9","article-title":"Blind audio watermarking algorithm based on DCT, linear regression and standard deviation","volume":"76","author":"Jeyhoon","year":"2017","journal-title":"Multimed. Tools Appl."},{"key":"ref_24","first-page":"462","article-title":"Protecting IP of Deep Neural Networks with Watermarking: A New Label Helps","volume":"12085","author":"Zhong","year":"2020","journal-title":"Adv. Knowl. Discov. Data Min."},{"key":"ref_25","unstructured":"Garofolo, J.S., Lamel, L.F., Fisher, W.M., Fiscus, J.G., Pallett, D.S., and Dahlgren, N.L. (2022, January 10). DARPA TIMIT Acoustic Phonetic Continuous Speech Corpus CDROM, Available online: https:\/\/nvlpubs.nist.gov\/nistpubs\/Legacy\/IR\/nistir4930.pdf."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Ravanelli, M., and Bengio, Y. (2018, January 18\u201321). Speaker recognition from raw waveform with sincnet. Proceedings of the 2018 IEEE Spoken Language Technology Workshop (SLT), Athens, Greece.","DOI":"10.1109\/SLT.2018.8639585"},{"key":"ref_27","unstructured":"Maas, A.L., Hannun, A.Y., and Ng, A.Y. (2013, January 16\u201321). Rectifier nonlinearities improve neural network acoustic models. Proceedings of the 30th International Conference on International Conference on Machine Learning, Atlanta, GA, USA."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/14\/3\/619\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:39:49Z","timestamp":1760135989000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/14\/3\/619"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,20]]},"references-count":27,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2022,3]]}},"alternative-id":["sym14030619"],"URL":"https:\/\/doi.org\/10.3390\/sym14030619","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,3,20]]}}}