{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T20:08:53Z","timestamp":1774642133764,"version":"3.50.1"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2024,3,8]],"date-time":"2024-03-08T00:00:00Z","timestamp":1709856000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61972127, 61972142, U22A2030"],"award-info":[{"award-number":["61972127, 61972142, U22A2030"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,6,30]]},"abstract":"<jats:p>\n            As an important part of the text-to-speech (TTS) system, vocoders convert acoustic features into speech waveforms. The difference in vocoders is key to producing different types of forged speech in the TTS system. With the rapid development of general adversarial networks (GANs), an increasing number of GAN vocoders have been proposed. Detectors often encounter vocoders of unknown types, which leads to a decline in the generalization performance of models. However, existing studies lack research on detection generalization based on GAN vocoders. To solve this problem, this study proposes vocoder detection of spoofed speech based on GAN fingerprints and domain generalization. The framework can widen the distance between real speech and forged speech in feature space, improving the detection model\u2019s performance. Specifically, we utilize a fingerprint extractor based on an autoencoder to extract GAN fingerprints from vocoders. We then weight them to the forged speech for subsequent classification to learn the forged speech features with high differentiation. Subsequently, domain generalization is used to further improve the generalization ability of the model for unseen forgery types. We achieve domain generalization using domain-adversarial learning and asymmetric triplet loss to learn a better generalized feature space in which real speech is compact and forged speech synthesized by different vocoders is dispersed. Finally, to optimize the training process, curriculum learning is used to dynamically adjust the contributions of the samples with different difficulties in the training process. Experimental results show that the proposed method achieves the most advanced detection results among four GAN vocoders. The code is available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/multimedia-infomation-security\/GAN-Vocoder-detection\">https:\/\/github.com\/multimedia-infomation-security\/GAN-Vocoder-detection<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3630751","type":"journal-article","created":{"date-parts":[[2023,10,28]],"date-time":"2023-10-28T18:26:32Z","timestamp":1698517592000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Vocoder Detection of Spoofing Speech Based on GAN Fingerprints and Domain Generalization"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-4616-3424","authenticated-orcid":false,"given":"Fan","family":"Li","sequence":"first","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1163-7926","authenticated-orcid":false,"given":"Yanxiang","family":"Chen","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3040-974X","authenticated-orcid":false,"given":"Haiyang","family":"Liu","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3606-4375","authenticated-orcid":false,"given":"Zuxing","family":"Zhao","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1965-7670","authenticated-orcid":false,"given":"Yuanzhi","family":"Yao","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9131-0578","authenticated-orcid":false,"given":"Xin","family":"Liao","sequence":"additional","affiliation":[{"name":"Hunan University, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,3,8]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"4308","article-title":"SuperLoss: A generic loss for robust curriculum learning","volume":"33","author":"Castells Thibault","year":"2020","unstructured":"Thibault Castells, Philippe Weinzaepfel, and Jerome Revaud. 2020. SuperLoss: A generic loss for robust curriculum learning. Advances in Neural Information Processing Systems 33 (2020), 4308\u20134319.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_3_2","doi-asserted-by":"crossref","first-page":"102","DOI":"10.21437\/Interspeech.2017-1085","volume-title":"Interspeech","author":"Chen Zhuxin","year":"2017","unstructured":"Zhuxin Chen, Zhifeng Xie, Weibin Zhang, and Xiangmin Xu. 2017. ResNet and model fusion for automatic spoofing detection. In Interspeech. 102\u2013106."},{"key":"e_1_3_1_4_2","first-page":"1180","volume-title":"International Conference on Machine Learning","author":"Ganin Yaroslav","year":"2015","unstructured":"Yaroslav Ganin and Victor Lempitsky. 2015. Unsupervised domain adaptation by backpropagation. In International Conference on Machine Learning. PMLR, 1180\u20131189."},{"key":"e_1_3_1_5_2","first-page":"1311","volume-title":"International Conference on Machine Learning","author":"Graves Alex","year":"2017","unstructured":"Alex Graves, Marc G. Bellemare, Jacob Menick, Remi Munos, and Koray Kavukcuoglu. 2017. Automated curriculum learning for neural networks. In International Conference on Machine Learning. PMLR, 1311\u20131320."},{"key":"e_1_3_1_6_2","volume-title":"2006 IEEE International Symposium on Circuits and Systems (ISCAS\u201906)","author":"Han Wei","year":"2006","unstructured":"Wei Han, Cheong-Fat Chan, Chiu-Sing Choy, and Kong-Pang Pun. 2006. An efficient MFCC extraction method in speech recognition. In 2006 IEEE International Symposium on Circuits and Systems (ISCAS\u201906). IEEE, 4\u2013pp."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2021.3089437"},{"key":"e_1_3_1_8_2","first-page":"17022","article-title":"HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis","volume":"33","author":"Kong Jungil","year":"2020","unstructured":"Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in Neural Information Processing Systems 33 (2020), 17022\u201317033.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10772-020-09785-w"},{"key":"e_1_3_1_10_2","article-title":"MelGAN: Generative adversarial networks for conditional waveform synthesis","volume":"32","author":"Kumar Kundan","year":"2019","unstructured":"Kundan Kumar, Rithesh Kumar, Thibault De Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre De Brebisson, Yoshua Bengio, and Aaron C. Courville. 2019. MelGAN: Generative adversarial networks for conditional waveform synthesis. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","first-page":"82","DOI":"10.21437\/Interspeech.2017-360","volume-title":"Interspeech","author":"Lavrentyeva Galina","year":"2017","unstructured":"Galina Lavrentyeva, Sergey Novoselov, Egor Malykh, Alexander Kozlov, Oleg Kudashev, and Vadim Shchemelinin. 2017. Audio replay attack detection with deep learning frameworks. In Interspeech. 82\u201386."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1904.05576"},{"key":"e_1_3_1_13_2","first-page":"1116","volume-title":"Interspeech","author":"Lei Zhenchun","year":"2020","unstructured":"Zhenchun Lei, Yingen Yang, Changhong Liu, and Jihua Ye. 2020. Siamese convolutional neural network using Gaussian probability feature for spoofing speech detection. In Interspeech. 1116\u20131120."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1806.09276"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00566"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv:2107.08803"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3511808.3557574"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2023.3262419"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/MIPR.2019.00103"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9746059"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9413605"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1609.03499"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2017.2682788"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2017.2684705"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01026"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW59228.2023.00097"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv:2107.12710"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2005.10393"},{"key":"e_1_3_1_29_2","first-page":"283","volume-title":"Odyssey","author":"Todisco Massimiliano","year":"2016","unstructured":"Massimiliano Todisco, H\u00e9ctor Delgado, and Nicholas W. D. Evans. 2016. A new feature for automatic speaker verification anti-spoofing: Constant Q Cepstral coefficients. In Odyssey, Vol. 2016. 283\u2013290."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-2289"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.316"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv:2212.13466"},{"key":"e_1_3_1_33_2","first-page":"2243","volume-title":"Interspeech","year":"2016","unstructured":"Wenfu Wang, Shuang Xu, and Bo Xu. 2016. First step towards end-to-end parametric TTS synthesis: Generating spectral parameters with neural attention. In Interspeech. 2243\u20132247."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv:2103.11326"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1703.10135"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3467707.3467736"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053795"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3552466.3556525"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00765"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2021.3076358"},{"key":"e_1_3_1_42_2","volume-title":"Proc. Interspeech","author":"Zhang Yuxiang","year":"2021","unstructured":"Yuxiang Zhang, Wenchao Wang, and Pengyuan Zhang. 2021. The effect of silence and dual-band fusion in anti-spoofing system. In Proc. Interspeech."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3630751","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3630751","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:36:33Z","timestamp":1750178193000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3630751"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,8]]},"references-count":41,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2024,6,30]]}},"alternative-id":["10.1145\/3630751"],"URL":"https:\/\/doi.org\/10.1145\/3630751","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,8]]},"assertion":[{"value":"2023-05-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-21","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}