{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:15:09Z","timestamp":1750220109751,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":25,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,6,23]],"date-time":"2022-06-23T00:00:00Z","timestamp":1655942400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,6,23]]},"DOI":"10.1145\/3548636.3548644","type":"proceedings-article","created":{"date-parts":[[2022,8,23]],"date-time":"2022-08-23T16:09:33Z","timestamp":1661270973000},"page":"53-58","source":"Crossref","is-referenced-by-count":0,"title":["Wav2sv: End-to-end Speaker Embeddings Learning from Raw Waveforms based on Metric Learning for Speaker Verification"],"prefix":"10.1145","author":[{"given":"Zhiqing","family":"Chen","sequence":"first","affiliation":[{"name":"School of Electronic and Computer Engineering, Peking University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yifan","family":"Pan","sequence":"additional","affiliation":[{"name":"School of Electronic and Computer Engineering, Peking University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haoran","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Electronic and Computer Engineering, Peking University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuesheng","family":"Zhu","sequence":"additional","affiliation":[{"name":"School of Electronic and Computer Engineering, Peking University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,8,23]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Joon Son Chung Jaesung Huh Seongkyu Mun Minjae Lee Hee-Soo Heo Soyeon Choe Chiheon Ham Sunghwan Jung Bong-Jin Lee and Icksang Han. 2020. In Defence of Metric Learning for Speaker Recognition. In Interspeech 2020 ISCA 2977\u20132981. DOI:https:\/\/doi.org\/10.21437\/Interspeech.2020-1064  Joon Son Chung Jaesung Huh Seongkyu Mun Minjae Lee Hee-Soo Heo Soyeon Choe Chiheon Ham Sunghwan Jung Bong-Jin Lee and Icksang Han. 2020. In Defence of Metric Learning for Speaker Recognition. In Interspeech 2020 ISCA 2977\u20132981. DOI:https:\/\/doi.org\/10.21437\/Interspeech.2020-1064","DOI":"10.21437\/Interspeech.2020-1064"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2010.2064307"},{"key":"e_1_3_2_1_3_1","volume-title":"ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification. Interspeech 2020 (October","author":"Desplanques Brecht","year":"2020","unstructured":"Brecht Desplanques , Jenthe Thienpondt , and Kris Demuynck . 2020. ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification. Interspeech 2020 (October 2020 ), 3830\u20133834. DOI:https:\/\/doi.org\/10.21437\/Interspeech.2020-2650 Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck. 2020. ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification. Interspeech 2020 (October 2020), 3830\u20133834. DOI:https:\/\/doi.org\/10.21437\/Interspeech.2020-2650"},{"key":"e_1_3_2_1_4_1","volume-title":"CN-Celeb: A Challenging Chinese Speaker Recognition Dataset. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 7604\u20137608","author":"Fan Y.","year":"2020","unstructured":"Y. Fan , J.W. Kang , L.T. Li , K.C. Li , H.L. Chen , S.T. Cheng , P.Y. Zhang , Z.Y. Zhou , Y.Q. Cai , and D. Wang . 2020 . CN-Celeb: A Challenging Chinese Speaker Recognition Dataset. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 7604\u20137608 . DOI:https:\/\/doi.org\/10.1109\/ICASSP40776. 2020 .9054017 Y. Fan, J.W. Kang, L.T. Li, K.C. Li, H.L. Chen, S.T. Cheng, P.Y. Zhang, Z.Y. Zhou, Y.Q. Cai, and D. Wang. 2020. CN-Celeb: A Challenging Chinese Speaker Recognition Dataset. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 7604\u20137608. DOI:https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9054017"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2938758"},{"key":"e_1_3_2_1_6_1","volume-title":"A deep neural network for short-segment speaker recognition. ArXiv Prepr. ArXiv190710420","author":"Hajavi Amirhossein","year":"2019","unstructured":"Amirhossein Hajavi and Ali Etemad . 2019. A deep neural network for short-segment speaker recognition. ArXiv Prepr. ArXiv190710420 ( 2019 ). Amirhossein Hajavi and Ali Etemad. 2019. A deep neural network for short-segment speaker recognition. ArXiv Prepr. ArXiv190710420 (2019)."},{"key":"e_1_3_2_1_7_1","volume-title":"Retrieved","author":"Hajibabaei Mahdi","year":"2018","unstructured":"Mahdi Hajibabaei and Dengxin Dai . 2018 . Unified Hypersphere Embedding for Speaker Recognition. ArXiv180708312 Cs Eess (July 2018) . Retrieved February 12, 2022 from http:\/\/arxiv.org\/abs\/1807.08312 Mahdi Hajibabaei and Dengxin Dai. 2018. Unified Hypersphere Embedding for Speaker Recognition. ArXiv180708312 Cs Eess (July 2018). Retrieved February 12, 2022 from http:\/\/arxiv.org\/abs\/1807.08312"},{"key":"e_1_3_2_1_8_1","volume-title":"Retrieved","author":"He Kaiming","year":"2016","unstructured":"Kaiming He , Xiangyu Zhang , Shaoqing Ren , and Jian Sun . 2016 . Deep Residual Learning for Image Recognition. 770\u2013778 . Retrieved July 14, 2021 from https:\/\/openaccess.thecvf.com\/content_cvpr_2016\/html\/He_Deep_Residual_Learning_CVPR_2016_paper.html Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. 770\u2013778. Retrieved July 14, 2021 from https:\/\/openaccess.thecvf.com\/content_cvpr_2016\/html\/He_Deep_Residual_Learning_CVPR_2016_paper.html"},{"key":"e_1_3_2_1_9_1","volume-title":"Retrieved","author":"Hu Jie","year":"2018","unstructured":"Jie Hu , Li Shen , and Gang Sun . 2018 . Squeeze-and-Excitation Networks. 7132\u20137141 . Retrieved August 12, 2021 from https:\/\/openaccess.thecvf.com\/content_cvpr_2018\/html\/Hu_Squeeze-and-Excitation_Networks_CVPR_2018_paper Jie Hu, Li Shen, and Gang Sun. 2018. Squeeze-and-Excitation Networks. 7132\u20137141. Retrieved August 12, 2021 from https:\/\/openaccess.thecvf.com\/content_cvpr_2018\/html\/Hu_Squeeze-and-Excitation_Networks_CVPR_2018_paper"},{"key":"e_1_3_2_1_10_1","volume-title":"RawNet: Advanced end-to-end deep neural network using raw waveforms for text-independent speaker verification. ArXiv190408104 Cs Eess (July","author":"Heo Hee-Soo","year":"2019","unstructured":"Jee-weon Jung, Hee-Soo Heo , Ju-ho Kim, Hye-jin Shim, and Ha-Jin Yu. 2019. RawNet: Advanced end-to-end deep neural network using raw waveforms for text-independent speaker verification. ArXiv190408104 Cs Eess (July 2019 ). Retrieved July 7, 2021 from http:\/\/arxiv.org\/abs\/1904.08104 Jee-weon Jung, Hee-Soo Heo, Ju-ho Kim, Hye-jin Shim, and Ha-Jin Yu. 2019. RawNet: Advanced end-to-end deep neural network using raw waveforms for text-independent speaker verification. ArXiv190408104 Cs Eess (July 2019). Retrieved July 7, 2021 from http:\/\/arxiv.org\/abs\/1904.08104"},{"key":"e_1_3_2_1_11_1","volume-title":"19th Annual Conference of the International Speech Communication Association (interspeech 2018","author":"Jung Jee-Weon","year":"2018","unstructured":"Jee-Weon Jung , Hee-Soo Heo , Il-Ho Yang , Hye-Jin Shim , and Ha-Jin Yu . 2018 . Avoiding Speaker Overfitting in End-to-End DNNs using Raw Waveform for Text-Independent Speaker Verification . In 19th Annual Conference of the International Speech Communication Association (interspeech 2018 ), Vols 1-6: Speech Research for Emerging Markets in Multilingual Societies. Isca-Int Speech Communication Assoc, Baixas, 3583\u20133587. DOI:https:\/\/doi.org\/10.21437\/Interspeech. 2018-1608 Jee-Weon Jung, Hee-Soo Heo, Il-Ho Yang, Hye-Jin Shim, and Ha-Jin Yu. 2018. Avoiding Speaker Overfitting in End-to-End DNNs using Raw Waveform for Text-Independent Speaker Verification. In 19th Annual Conference of the International Speech Communication Association (interspeech 2018), Vols 1-6: Speech Research for Emerging Markets in Multilingual Societies. Isca-Int Speech Communication Assoc, Baixas, 3583\u20133587. DOI:https:\/\/doi.org\/10.21437\/Interspeech.2018-1608"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Jee-weon Jung Seung-bin Kim Hye-jin Shim Ju-ho Kim and Ha-Jin Yu. 2020. Improved RawNet with Feature Map Scaling for Text-independent Speaker Verification using Raw Waveforms. (2020).  Jee-weon Jung Seung-bin Kim Hye-jin Shim Ju-ho Kim and Ha-Jin Yu. 2020. Improved RawNet with Feature Map Scaling for Text-independent Speaker Verification using Raw Waveforms. (2020).","DOI":"10.21437\/Interspeech.2020-1011"},{"key":"e_1_3_2_1_13_1","unstructured":"Wei-Wei Lin and Man-Wai Mak. 2020. Wav2Spk: A Simple DNN Architecture for Learning Speaker Embeddings from Waveforms. In INTERSPEECH 3211\u20133215.  Wei-Wei Lin and Man-Wai Mak. 2020. Wav2Spk: A Simple DNN Architecture for Learning Speaker Embeddings from Waveforms. In INTERSPEECH 3211\u20133215."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Mitchell McLaren Luciana Ferrer Diego Castan and Aaron Lawson. 2016. The speakers in the wild (SITW) speaker recognition database. In Interspeech 818\u2013822.  Mitchell McLaren Luciana Ferrer Diego Castan and Aaron Lawson. 2016. The speakers in the wild (SITW) speaker recognition database. In Interspeech 818\u2013822.","DOI":"10.21437\/Interspeech.2016-1129"},{"key":"e_1_3_2_1_15_1","volume-title":"Towards Directly Modeling Raw Speech Signal for Speaker Verification Using CNNS. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 4884\u20134888","author":"Muckenhirn Hannah","year":"2018","unstructured":"Hannah Muckenhirn , Mathew Magimai .-Doss, and S\u00e9bastien Marcell . 2018 . Towards Directly Modeling Raw Speech Signal for Speaker Verification Using CNNS. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 4884\u20134888 . DOI:https:\/\/doi.org\/10.1109\/ICASSP.2018.8462165 Hannah Muckenhirn, Mathew Magimai.-Doss, and S\u00e9bastien Marcell. 2018. Towards Directly Modeling Raw Speech Signal for Speaker Verification Using CNNS. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 4884\u20134888. DOI:https:\/\/doi.org\/10.1109\/ICASSP.2018.8462165"},{"key":"e_1_3_2_1_16_1","volume-title":"Weidi Xie, and Andrew Zisserman.","author":"Nagrani Arsha","year":"2020","unstructured":"Arsha Nagrani , Joon Son Chung , Weidi Xie, and Andrew Zisserman. 2020 . Voxceleb : Large-scale speaker verification in the wild. Comput. Speech Lang . 60, (March 2020), 101027. DOI:https:\/\/doi.org\/10.1016\/j.csl.2019.101027 Arsha Nagrani, Joon Son Chung, Weidi Xie, and Andrew Zisserman. 2020. Voxceleb: Large-scale speaker verification in the wild. Comput. Speech Lang. 60, (March 2020), 101027. DOI:https:\/\/doi.org\/10.1016\/j.csl.2019.101027"},{"key":"e_1_3_2_1_17_1","volume-title":"Joon Son Chung, and Andrew Zisserman","author":"Nagrani Arsha","year":"2017","unstructured":"Arsha Nagrani , Joon Son Chung, and Andrew Zisserman . 2017 . VoxCeleb: A Large-Scale Speaker Identification Dataset. In Interspeech 2017, ISCA , 2616\u20132620. DOI:https:\/\/doi.org\/10.21437\/Interspeech.2017-950 Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. 2017. VoxCeleb: A Large-Scale Speaker Identification Dataset. In Interspeech 2017, ISCA, 2616\u20132620. DOI:https:\/\/doi.org\/10.21437\/Interspeech.2017-950"},{"key":"e_1_3_2_1_18_1","volume-title":"Attentive Statistics Pooling for Deep Speaker Embedding. Interspeech 2018 (September","author":"Okabe Koji","year":"2018","unstructured":"Koji Okabe , Takafumi Koshinaka , and Koichi Shinoda . 2018. Attentive Statistics Pooling for Deep Speaker Embedding. Interspeech 2018 (September 2018 ), 2252\u20132256. DOI:https:\/\/doi.org\/10.21437\/Interspeech.2018-993 Koji Okabe, Takafumi Koshinaka, and Koichi Shinoda. 2018. Attentive Statistics Pooling for Deep Speaker Embedding. Interspeech 2018 (September 2018), 2252\u20132256. DOI:https:\/\/doi.org\/10.21437\/Interspeech.2018-993"},{"key":"e_1_3_2_1_19_1","volume-title":"Revisiting SincNet: An Evaluation of Feature and Network Hyperparameters for Speaker Recognition. In 2020 28th European Signal Processing Conference (EUSIPCO), 1\u20135. DOI:https:\/\/doi.org\/10","author":"Dan Onea\u021b","year":"2021","unstructured":"Dan Onea\u021b \u0103 \u0103, Lucian Georgescu , Horia Cucu , Drago\u015f Burileanu , and Corneliu Burileanu . 2021 . Revisiting SincNet: An Evaluation of Feature and Network Hyperparameters for Speaker Recognition. In 2020 28th European Signal Processing Conference (EUSIPCO), 1\u20135. DOI:https:\/\/doi.org\/10 .23919\/Eusipco47968.2020.9287794 Dan Onea\u021b \u0103 \u0103, Lucian Georgescu, Horia Cucu, Drago\u015f Burileanu, and Corneliu Burileanu. 2021. Revisiting SincNet: An Evaluation of Feature and Network Hyperparameters for Speaker Recognition. In 2020 28th European Signal Processing Conference (EUSIPCO), 1\u20135. DOI:https:\/\/doi.org\/10.23919\/Eusipco47968.2020.9287794"},{"key":"e_1_3_2_1_20_1","volume-title":"2018 IEEE Spoken Language Technology Workshop (SLT), 1021\u20131028","author":"Ravanelli Mirco","year":"2018","unstructured":"Mirco Ravanelli and Yoshua Bengio . 2018 . Speaker Recognition from Raw Waveform with SincNet . In 2018 IEEE Spoken Language Technology Workshop (SLT), 1021\u20131028 . DOI:https:\/\/doi.org\/10.1109\/SLT.2018.8639585 Mirco Ravanelli and Yoshua Bengio. 2018. Speaker Recognition from Raw Waveform with SincNet. In 2018 IEEE Spoken Language Technology Workshop (SLT), 1021\u20131028. DOI:https:\/\/doi.org\/10.1109\/SLT.2018.8639585"},{"key":"e_1_3_2_1_21_1","volume-title":"Retrieved","author":"Schneider Steffen","year":"2019","unstructured":"Steffen Schneider , Alexei Baevski , Ronan Collobert , and Michael Auli . 2019 . wav2vec: Unsupervised Pre-training for Speech Recognition. ArXiv190405862 Cs (September 2019) . Retrieved August 16, 2021 from http:\/\/arxiv.org\/abs\/1904.05862 Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019. wav2vec: Unsupervised Pre-training for Speech Recognition. ArXiv190405862 Cs (September 2019). Retrieved August 16, 2021 from http:\/\/arxiv.org\/abs\/1904.05862"},{"key":"e_1_3_2_1_22_1","volume-title":"Frame-Level Speaker Embeddings for Text-Independent Speaker Recognition and Analysis of End-to-End Model. In 2018 IEEE Spoken Language Technology Workshop (SLT), 1007\u20131013","author":"Shon Suwon","year":"2018","unstructured":"Suwon Shon , Hao Tang , and James Glass . 2018 . Frame-Level Speaker Embeddings for Text-Independent Speaker Recognition and Analysis of End-to-End Model. In 2018 IEEE Spoken Language Technology Workshop (SLT), 1007\u20131013 . DOI:https:\/\/doi.org\/10.1109\/SLT.2018.8639622 Suwon Shon, Hao Tang, and James Glass. 2018. Frame-Level Speaker Embeddings for Text-Independent Speaker Recognition and Analysis of End-to-End Model. In 2018 IEEE Spoken Language Technology Workshop (SLT), 1007\u20131013. DOI:https:\/\/doi.org\/10.1109\/SLT.2018.8639622"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"crossref","unstructured":"David Snyder Daniel Garcia-Romero Daniel Povey and Sanjeev Khudanpur. 2017. Deep Neural Network Embeddings for Text-Independent Speaker Verification. In Interspeech 999\u20131003.  David Snyder Daniel Garcia-Romero Daniel Povey and Sanjeev Khudanpur. 2017. Deep Neural Network Embeddings for Text-Independent Speaker Verification. In Interspeech 999\u20131003.","DOI":"10.21437\/Interspeech.2017-620"},{"key":"e_1_3_2_1_24_1","volume-title":"X-Vectors: Robust DNN Embeddings for Speaker Recognition. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5329\u20135333","author":"Snyder David","year":"2018","unstructured":"David Snyder , Daniel Garcia-Romero , Gregory Sell , Daniel Povey , and Sanjeev Khudanpur . 2018 . X-Vectors: Robust DNN Embeddings for Speaker Recognition. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5329\u20135333 . DOI:https:\/\/doi.org\/10.1109\/ICASSP.2018.8461375 David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur. 2018. X-Vectors: Robust DNN Embeddings for Speaker Recognition. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5329\u20135333. DOI:https:\/\/doi.org\/10.1109\/ICASSP.2018.8461375"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Ge Zhu and Z. Duan. 2021. Y-Vector: Multiscale Waveform Encoder for Speaker Embedding. In Interspeech.  Ge Zhu and Z. Duan. 2021. Y-Vector: Multiscale Waveform Encoder for Speaker Embedding. In Interspeech.","DOI":"10.21437\/Interspeech.2021-1707"}],"event":{"name":"ITCC 2022: 2022 4th International Conference on Information Technology and Computer Communications","acronym":"ITCC 2022","location":"Guangzhou China"},"container-title":["2022 4th International Conference on Information Technology and Computer Communications (ITCC)"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3548636.3548644","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3548636.3548644","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:10:39Z","timestamp":1750183839000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3548636.3548644"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,23]]},"references-count":25,"alternative-id":["10.1145\/3548636.3548644","10.1145\/3548636"],"URL":"https:\/\/doi.org\/10.1145\/3548636.3548644","relation":{},"subject":[],"published":{"date-parts":[[2022,6,23]]}}}