{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:58:21Z","timestamp":1750309101878,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":27,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,12,6]],"date-time":"2023-12-06T00:00:00Z","timestamp":1701820800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"JSPS KAKENHI Grant No.","award":["23K11227, 22H03595, 23H03402"],"award-info":[{"award-number":["23K11227, 22H03595, 23H03402"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,12,6]]},"DOI":"10.1145\/3595916.3626366","type":"proceedings-article","created":{"date-parts":[[2024,1,1]],"date-time":"2024-01-01T16:34:41Z","timestamp":1704126881000},"page":"1-5","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Reprogramming Self-supervised Learning-based Speech Representations for Speaker Anonymization"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5458-7025","authenticated-orcid":false,"given":"Xiaojiao","family":"Chen","sequence":"first","affiliation":[{"name":"Xinjiang University, CN"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7636-3797","authenticated-orcid":false,"given":"Sheng","family":"Li","sequence":"additional","affiliation":[{"name":"National Institute of Information and Communications Technology (NICT), JP"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4997-3850","authenticated-orcid":false,"given":"Jiyi","family":"Li","sequence":"additional","affiliation":[{"name":"University of Yamanashi, JP"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6604-0951","authenticated-orcid":false,"given":"Hao","family":"Huang","sequence":"additional","affiliation":[{"name":"Xinjiang University, CN"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6424-8633","authenticated-orcid":false,"given":"Yang","family":"Cao","sequence":"additional","affiliation":[{"name":"Hokkaido University, JP"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4076-7479","authenticated-orcid":false,"given":"Liang","family":"He","sequence":"additional","affiliation":[{"name":"Tsinghua University, CN"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,1]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in neural information processing systems 33","author":"Baevski Alexei","year":"2020","unstructured":"Alexei Baevski , Yuhao Zhou , Abdelrahman Mohamed , and Michael Auli . 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in neural information processing systems 33 ( 2020 ), 12449\u201312460. Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in neural information processing systems 33 (2020), 12449\u201312460."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2022.3188113"},{"key":"e_1_3_2_1_3_1","volume-title":"Adversarial Reprogramming of Neural Networks. In International Conference on Learning Representations.","author":"Elsayed F","year":"2019","unstructured":"Gamaleldin\u00a0 F Elsayed , Ian Goodfellow , and Jascha Sohl-Dickstein . 2019 . Adversarial Reprogramming of Neural Networks. In International Conference on Learning Representations. Gamaleldin\u00a0F Elsayed, Ian Goodfellow, and Jascha Sohl-Dickstein. 2019. Adversarial Reprogramming of Neural Networks. In International Conference on Learning Representations."},{"key":"e_1_3_2_1_4_1","volume-title":"2019. Speaker anonymization using x-vector and neural waveform models. arXiv preprint arXiv:1905.13561","author":"Fuming F.","year":"2019","unstructured":"F. Fuming and 2019. Speaker anonymization using x-vector and neural waveform models. arXiv preprint arXiv:1905.13561 ( 2019 ). F. Fuming and et al.2019. Speaker anonymization using x-vector and neural waveform models. arXiv preprint arXiv:1905.13561 (2019)."},{"key":"e_1_3_2_1_5_1","unstructured":"Y. Ganin E. Ustinova H. Ajakan P. Germain H. Larochelle F. Laviolette M. Marchand and V. Lempitsky. 2016. Domain-adversarial training of neural networks. The journal of machine learning research 17 1 (2016) 2096\u20132030.  Y. Ganin E. Ustinova H. Ajakan P. Germain H. Larochelle F. Laviolette M. Marchand and V. Lempitsky. 2016. Domain-adversarial training of neural networks. The journal of machine learning research 17 1 (2016) 2096\u20132030."},{"volume-title":"Deep learning","author":"Goodfellow I.","key":"e_1_3_2_1_6_1","unstructured":"I. Goodfellow , Y. Bengio , and A. Courville . 2016. Deep learning . MIT press . I. Goodfellow, Y. Bengio, and A.Courville. 2016. Deep learning. MIT press."},{"volume-title":"2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 5500\u20135504","author":"Hashimoto K.","key":"e_1_3_2_1_7_1","unstructured":"K. Hashimoto , J. Yamagishi , and I. Echizen . 2016. Privacy-preserving sound to degrade automatic speaker verification performance . In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 5500\u20135504 . K. Hashimoto, J. Yamagishi, and I. Echizen. 2016. Privacy-preserving sound to degrade automatic speaker verification performance. In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 5500\u20135504."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3122291"},{"key":"e_1_3_2_1_9_1","volume-title":"From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition. arXiv e-prints","author":"Huck\u00a0Yang Chao-Han","year":"2023","unstructured":"Chao-Han Huck\u00a0Yang , Bo Li , Yu Zhang , Nanxin Chen , Rohit Prabhavalkar , Tara\u00a0 N Sainath , and Trevor Strohman . 2023. From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition. arXiv e-prints ( 2023 ), arXiv\u20132301. Chao-Han Huck\u00a0Yang, Bo Li, Yu Zhang, Nanxin Chen, Rohit Prabhavalkar, Tara\u00a0N Sainath, and Trevor Strohman. 2023. From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition. arXiv e-prints (2023), arXiv\u20132301."},{"key":"e_1_3_2_1_10_1","volume-title":"Graz","author":"Ioffe Sergey","year":"2006","unstructured":"Sergey Ioffe . 2006 . Probabilistic linear discriminant analysis. In Computer Vision\u2013ECCV 2006: 9th European Conference on Computer Vision , Graz , Austria, May 7-13, 2006, Proceedings, Part IV 9. Springer Berlin Heidelberg, 531\u2013542. Sergey Ioffe. 2006. Probabilistic linear discriminant analysis. In Computer Vision\u2013ECCV 2006: 9th European Conference on Computer Vision, Graz, Austria, May 7-13, 2006, Proceedings, Part IV 9. Springer Berlin Heidelberg, 531\u2013542."},{"key":"e_1_3_2_1_11_1","first-page":"17022","article-title":"Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis","volume":"33","author":"Kong J.","year":"2020","unstructured":"J. Kong , J. Kim , and J. Bae . 2020 . Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis . Advances in Neural Information Processing Systems 33 (2020), 17022 \u2013 17033 . J. Kong, J. Kim, and J. Bae. 2020. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in Neural Information Processing Systems 33 (2020), 17022\u201317033.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"crossref","unstructured":"C.\u00a0O. Mawalim K. Galajit J. Karnjana and M. Unoki. 2020. X-Vector Singular Value Modification and Statistical-Based Decomposition with Ensemble Regression Modeling for Speaker Anonymization System.. In Interspeech. 1703\u20131707.  C.\u00a0O. Mawalim K. Galajit J. Karnjana and M. Unoki. 2020. X-Vector Singular Value Modification and Statistical-Based Decomposition with Ensemble Regression Modeling for Speaker Anonymization System.. In Interspeech. 1703\u20131707.","DOI":"10.21437\/Interspeech.2020-1887"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2022.101351"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Daniel Povey Gaofeng Cheng Yiming Wang Ke Li Hainan Xu Mahsa Yarmohammadi and Sanjeev Khudanpur. 2018. Semi-orthogonal low-rank matrix factorization for deep neural networks.. In Interspeech. 3743\u20133747.  Daniel Povey Gaofeng Cheng Yiming Wang Ke Li Hainan Xu Mahsa Yarmohammadi and Sanjeev Khudanpur. 2018. Semi-orthogonal low-rank matrix factorization for deep neural networks.. In Interspeech. 3743\u20133747.","DOI":"10.21437\/Interspeech.2018-1417"},{"key":"e_1_3_2_1_15_1","volume-title":"IEEE 2011 workshop on automatic speech recognition and understanding. IEEE Signal Processing Society.","author":"Povey Daniel","year":"2011","unstructured":"Daniel Povey , Arnab Ghoshal , Gilles Boulianne , Lukas Burget , Ondrej Glembek , Nagendra Goel , Mirko Hannemann , Petr Motlicek , Yanmin Qian , Petr Schwarz , 2011 . The Kaldi speech recognition toolkit . In IEEE 2011 workshop on automatic speech recognition and understanding. IEEE Signal Processing Society. Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, 2011. The Kaldi speech recognition toolkit. In IEEE 2011 workshop on automatic speech recognition and understanding. IEEE Signal Processing Society."},{"key":"e_1_3_2_1_16_1","first-page":"2804","article-title":"Forensic and automatic speaker recognition system","volume":"8","author":"Satyanand S.","year":"2018","unstructured":"S. Satyanand . 2018 . Forensic and automatic speaker recognition system . International Journal of Electrical and Computer Engineering 8 , 5 (2018), 2804 . S. Satyanand. 2018. Forensic and automatic speaker recognition system. International Journal of Electrical and Computer Engineering 8, 5 (2018), 2804.","journal-title":"International Journal of Electrical and Computer Engineering"},{"key":"e_1_3_2_1_17_1","volume-title":"wav2vec: Unsupervised pre-training for speech recognition. arXiv preprint arXiv:1904.05862","author":"Schneider Steffen","year":"2019","unstructured":"Steffen Schneider , Alexei Baevski , Ronan Collobert , and Michael Auli . 2019. wav2vec: Unsupervised pre-training for speech recognition. arXiv preprint arXiv:1904.05862 ( 2019 ). Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019. wav2vec: Unsupervised pre-training for speech recognition. arXiv preprint arXiv:1904.05862 (2019)."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461375"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2022.3190741"},{"key":"e_1_3_2_1_20_1","volume-title":"Design Choices for X-vector Based Speaker Anonymization. In INTERSPEECH","author":"Srivastava B.","year":"2020","unstructured":"B. Srivastava , N. Tomashenko , X. Wang , E. Vincent , J. Yamagishi , M. Maouche , A. Bellet , and M. Tommasi . 2020 . Design Choices for X-vector Based Speaker Anonymization. In INTERSPEECH 2020 . B. Srivastava, N. Tomashenko, X. Wang, E. Vincent, J.Yamagishi, M. Maouche, A. Bellet, and M. Tommasi. 2020. Design Choices for X-vector Based Speaker Anonymization. In INTERSPEECH 2020."},{"key":"e_1_3_2_1_21_1","volume-title":"INTERSPEECH","author":"Tomashenko N.","year":"2020","unstructured":"N. Tomashenko , B. Srivastava , X. Wang , E. Vincent , A. Nautsch , J. Yamagishi , N. Evans , J. Patino , J. Bonastre , P. No\u00e9, and et al.2020. Introducing the VoicePrivacy initiative . In INTERSPEECH 2020 . N. Tomashenko, B. Srivastava, X. Wang, E. Vincent, A. Nautsch, J. Yamagishi, N. Evans, J. Patino, J. Bonastre, P. No\u00e9, and et al.2020. Introducing the VoicePrivacy initiative. In INTERSPEECH 2020."},{"key":"e_1_3_2_1_22_1","unstructured":"N. Tomashenko X. Wang X. Miao H. Nourtel P. Champion M. Todisco E. Vincent N. Evans J. Yamagishi and J. Bonastre. 2022. The VoicePrivacy 2022 Challenge Evaluation Plan. arXiv preprint arXiv:2203.12468 (2022).  N. Tomashenko X. Wang X. Miao H. Nourtel P. Champion M. Todisco E. Vincent N. Evans J. Yamagishi and J. Bonastre. 2022. The VoicePrivacy 2022 Challenge Evaluation Plan. arXiv preprint arXiv:2203.12468 (2022)."},{"key":"e_1_3_2_1_23_1","unstructured":"H. Turner G. Lovisotto and I. Martinovic. 2020. Speaker anonymization with distribution-preserving x-vector generation for the VoicePrivacy Challenge 2020. arXiv preprint arXiv:2010.13457 (2020).  H. Turner G. Lovisotto and I. Martinovic. 2020. Speaker anonymization with distribution-preserving x-vector generation for the VoicePrivacy Challenge 2020. arXiv preprint arXiv:2010.13457 (2020)."},{"key":"e_1_3_2_1_24_1","volume-title":"Neural harmonic-plus-noise waveform model with trainable maximum voice frequency for text-to-speech synthesis. arXiv preprint arXiv:1908.10256","author":"Wang Xin","year":"2019","unstructured":"Xin Wang and Junichi Yamagishi . 2019. Neural harmonic-plus-noise waveform model with trainable maximum voice frequency for text-to-speech synthesis. arXiv preprint arXiv:1908.10256 ( 2019 ). Xin Wang and Junichi Yamagishi. 2019. Neural harmonic-plus-noise waveform model with trainable maximum voice frequency for text-to-speech synthesis. arXiv preprint arXiv:1908.10256 (2019)."},{"key":"e_1_3_2_1_25_1","volume-title":"International Conference on Machine Learning. PMLR, 11808\u201311819","author":"Yang Han\u00a0Huck","year":"2021","unstructured":"Chao- Han\u00a0Huck Yang , Yun-Yun Tsai , and Pin-Yu Chen . 2021 . Voice2series: Reprogramming acoustic models for time series classification . In International Conference on Machine Learning. PMLR, 11808\u201311819 . Chao-Han\u00a0Huck Yang, Yun-Yun Tsai, and Pin-Yu Chen. 2021. Voice2series: Reprogramming acoustic models for time series classification. In International Conference on Machine Learning. PMLR, 11808\u201311819."},{"key":"e_1_3_2_1_26_1","volume-title":"Superb: Speech processing universal performance benchmark. arXiv preprint arXiv:2105.01051","author":"Chi Po-Han","year":"2021","unstructured":"Shu-wen Yang, Po-Han Chi , Yung-Sung Chuang , Cheng- I\u00a0Jeff Lai , Kushal Lakhotia , Yist\u00a0 Y Lin , Andy\u00a0 T Liu , Jiatong Shi , Xuankai Chang , Guan-Ting Lin , 2021 . Superb: Speech processing universal performance benchmark. arXiv preprint arXiv:2105.01051 (2021). Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I\u00a0Jeff Lai, Kushal Lakhotia, Yist\u00a0Y Lin, Andy\u00a0T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, 2021. Superb: Speech processing universal performance benchmark. arXiv preprint arXiv:2105.01051 (2021)."},{"key":"e_1_3_2_1_27_1","volume-title":"A study of low-resource speech commands recognition based on adversarial reprogramming. arXiv e-prints","author":"Yen Hao","year":"2021","unstructured":"Hao Yen , Pin-Jui Ku , Chao-Han Huck\u00a0Yang , Hu Hu , Sabato\u00a0Marco Siniscalchi , Pin-Yu Chen , and Yu Tsao . 2021. A study of low-resource speech commands recognition based on adversarial reprogramming. arXiv e-prints ( 2021 ), arXiv\u20132110. Hao Yen, Pin-Jui Ku, Chao-Han Huck\u00a0Yang, Hu Hu, Sabato\u00a0Marco Siniscalchi, Pin-Yu Chen, and Yu Tsao. 2021. A study of low-resource speech commands recognition based on adversarial reprogramming. arXiv e-prints (2021), arXiv\u20132110."}],"event":{"name":"MMAsia '23: ACM Multimedia Asia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Tainan Taiwan","acronym":"MMAsia '23"},"container-title":["ACM Multimedia Asia 2023"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3595916.3626366","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3595916.3626366","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:48:40Z","timestamp":1750286920000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3595916.3626366"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,6]]},"references-count":27,"alternative-id":["10.1145\/3595916.3626366","10.1145\/3595916"],"URL":"https:\/\/doi.org\/10.1145\/3595916.3626366","relation":{},"subject":[],"published":{"date-parts":[[2023,12,6]]},"assertion":[{"value":"2024-01-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}