{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,28]],"date-time":"2026-05-28T00:11:38Z","timestamp":1779927098264,"version":"3.53.1"},"reference-count":33,"publisher":"MDPI AG","issue":"23","license":[{"start":{"date-parts":[[2020,11,27]],"date-time":"2020-11-27T00:00:00Z","timestamp":1606435200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61901003"],"award-info":[{"award-number":["61901003"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["51904297"],"award-info":[{"award-number":["51904297"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004608","name":"Natural Science Foundation of Jiangsu Province","doi-asserted-by":"publisher","award":["BK20190623"],"award-info":[{"award-number":["BK20190623"]}],"id":[{"id":"10.13039\/501100004608","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003995","name":"Natural Science Foundation of Anhui Province","doi-asserted-by":"publisher","award":["1908085QF255"],"award-info":[{"award-number":["1908085QF255"]}],"id":[{"id":"10.13039\/501100003995","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Automatic speaker verification provides a flexible and effective way for biometric authentication. Previous deep learning-based methods have demonstrated promising results, whereas a few problems still require better solutions. In prior works examining speaker discriminative neural networks, the speaker representation of the target speaker is regarded as a fixed one when comparing with utterances from different speakers, and the joint information between enrollment and evaluation utterances is ignored. In this paper, we propose to combine CNN-based feature learning with a bidirectional attention mechanism to achieve better performance with only one enrollment utterance. The evaluation-enrollment joint information is exploited to provide interactive features through bidirectional attention. In addition, we introduce one individual cost function to identify the phonetic contents, which contributes to calculating the attention score more specifically. These interactive features are complementary to the constant ones, which are extracted from individual speakers separately and do not vary with the evaluation utterances. The proposed method archived a competitive equal error rate of 6.26% on the internal \u201cDAN DAN NI HAO\u201d benchmark dataset with 1250 utterances and outperformed various baseline methods, including the traditional i-vector\/PLDA, d-vector, self-attention, and sequence-to-sequence attention models.<\/jats:p>","DOI":"10.3390\/s20236784","type":"journal-article","created":{"date-parts":[[2020,11,27]],"date-time":"2020-11-27T09:16:49Z","timestamp":1606468609000},"page":"6784","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":12,"title":["Bidirectional Attention for Text-Dependent Speaker Verification"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4796-9444","authenticated-orcid":false,"given":"Xin","family":"Fang","sequence":"first","affiliation":[{"name":"School of Information Science and Technology, University of Science and Technology of China, Hefei 230022, China"},{"name":"iFLYTEK Research, iFLYTEK Co., Ltd., Hefei 230088, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2872-2155","authenticated-orcid":false,"given":"Tian","family":"Gao","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, University of Science and Technology of China, Hefei 230022, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7322-5735","authenticated-orcid":false,"given":"Liang","family":"Zou","sequence":"additional","affiliation":[{"name":"School of Information and Electrical Control Engineering, China University of Mining and Technology, Xuzhou 221116, China"},{"name":"School of Electronics and Information Engineering, Anhui University, Hefei 236601, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7853-5273","authenticated-orcid":false,"given":"Zhenhua","family":"Ling","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, University of Science and Technology of China, Hefei 230022, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,11,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1421","DOI":"10.1109\/TASLP.2017.2694708","article-title":"HMM-based phrase-independent i-vector extractor for text-dependent speaker verification","volume":"25","author":"Zeinali","year":"2017","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Kang, W.H., and Kim, N.S. (2019). Adversarially Learned Total Variability Embedding for Speaker Recognition with Random Digit Strings. Sensors, 19.","DOI":"10.3390\/s19214709"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"101078","DOI":"10.1016\/j.csl.2020.101078","article-title":"Optimization of the area under the roc curve using neural network supervectors for text-dependent speaker verification","volume":"63","author":"Mingote","year":"2020","journal-title":"Comput. Speech Lang."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Machado, T.J., Vieira Filho, J., and de Oliveira, M.A. (2019). Forensic Speaker Verification Using Ordinary Least Squares. Sensors, 19.","DOI":"10.3390\/s19204385"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1016\/j.specom.2014.03.001","article-title":"Text-dependent speaker verification: Classifiers, databases and RSR2015","volume":"60","author":"Larcher","year":"2014","journal-title":"Speech Commun."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zhong, J., Hu, W., Soong, F.K., and Meng, H. (2017, January 20\u201324). DNN i-Vector Speaker Verification with Short, Text-Constrained Test Utterances. Proceedings of the 18th Annual Conference of the International Speech Communication Association (Interspeech), Stockholm, Sweden.","DOI":"10.21437\/Interspeech.2017-1036"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Yu, M., Li, N., Yu, C., Cui, J., and Yu, D. (2019, January 12\u201317). Seq2Seq Attentional Siamese Neural Networks for Text-dependent Speaker Verification. Proceedings of the ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8682676"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Liu, Y., He, L., Tian, Y., Chen, Z., Liu, J., and Johnson, M.T. (2017, January 16\u201320). Comparison of multiple features and modeling methods for text-dependent speaker verification. Proceedings of the 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Okinawa, Japan.","DOI":"10.1109\/ASRU.2017.8268995"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"788","DOI":"10.1109\/TASL.2010.2064307","article-title":"Front-end factor analysis for speaker verification","volume":"19","author":"Dehak","year":"2010","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.specom.2015.07.003","article-title":"Deep feature for text-dependent speaker verification","volume":"73","author":"Liu","year":"2015","journal-title":"Speech Commun."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Yoon, S.H., Jeon, J.J., and Yu, H.J. (2020). Regularized Within-Class Precision Matrix Based PLDA in Text-Dependent Speaker Verification. Appl. Sci., 10.","DOI":"10.3390\/app10186571"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1016\/j.csl.2017.04.005","article-title":"Text-dependent speaker verification based on i-vectors, neural networks and hidden Markov models","volume":"46","author":"Zeinali","year":"2017","journal-title":"Comput. Speech Lang."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1670","DOI":"10.1109\/LSP.2018.2870726","article-title":"SNR-invariant multitask deep neural networks for robust speaker verification","volume":"25","author":"Yao","year":"2018","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"684","DOI":"10.1109\/JSTSP.2016.2647199","article-title":"An investigation of deep-learning frameworks for speaker verification antispoofing","volume":"11","author":"Zhang","year":"2017","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Garcia-Romero, D., and McCree, A. (2015, January 6\u201310). Insights into deep neural networks for speaker recognition. Proceedings of the Sixteenth Annual Conference of the International Speech Communication Association, Dresden, Germany.","DOI":"10.21437\/Interspeech.2015-298"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Dey, S., Madikeri, S.R., and Motlicek, P. (2018, January 2\u20136). End-to-end Text-dependent Speaker Verification Using Novel Distance Measures. Proceedings of the 19th Annual Conference of the International Speech Communication Association (Interspeech), Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-2300"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhang, S.X., Chen, Z., Zhao, Y., Li, J., and Gong, Y. (2016, January 13\u201316). End-to-end attention based text-dependent speaker verification. Proceedings of the 2016 IEEE Spoken Language Technology Workshop (SLT), San Diego, CA, USA.","DOI":"10.1109\/SLT.2016.7846261"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1016\/j.neucom.2019.08.046","article-title":"Self-attention based speaker recognition using Cluster-Range Loss","volume":"368","author":"Bian","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_19","unstructured":"Chowdhury, F.R.R., Wang, Q., Moreno, I.L., and Wan, L. (2018, January 15\u201320). Attention-based models for text-dependent speaker verification. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1016\/j.patcog.2019.01.006","article-title":"Wider or deeper: Revisiting the resnet model for visual recognition","volume":"90","author":"Wu","year":"2019","journal-title":"Pattern Recognit."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Sun, Y., Cheng, C., Zhang, Y., Zhang, C., Zheng, L., Wang, Z., and Wei, Y. (2020). Circle loss: A unified perspective of pair similarity optimization. arXiv.","DOI":"10.1109\/CVPR42600.2020.00643"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"2095","DOI":"10.1109\/TASL.2007.902758","article-title":"Modeling prosodic features with joint factor analysis for speaker verification","volume":"15","author":"Dehak","year":"2007","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"980","DOI":"10.1109\/TASL.2008.925147","article-title":"A study of interspeaker variability in speaker verification","volume":"16","author":"Kenny","year":"2008","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"308","DOI":"10.1109\/LSP.2006.870086","article-title":"Support vector machines using GMM supervectors for speaker verification","volume":"13","author":"Campbell","year":"2006","journal-title":"IEEE Signal Process Lett."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Heigold, G., Moreno, I., Bengio, S., and Shazeer, N. (2016, January 20\u201325). End-to-end text-dependent speaker verification. Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472652"},{"key":"ref_26","unstructured":"Li, C., Ma, X., Jiang, B., Li, X., Zhang, X., Liu, X., Cao, Y., Kannan, A., and Zhu, Z. (2017). Deep speaker: An end-to-end neural speaker embedding system. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Fang, X., Zou, L., Li, J., Sun, L., and Ling, Z.H. (2019, January 12\u201317). Channel adversarial training for cross-channel text-independent speaker recognition. Proceedings of the ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8682327"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"20160101","DOI":"10.1098\/rstb.2016.0101","article-title":"Modelling auditory attention","volume":"372","author":"Kaya","year":"2017","journal-title":"Philos. Trans. R. Soc. Biol. Sci."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"E3286","DOI":"10.1073\/pnas.1721226115","article-title":"Sensorineural hearing loss degrades behavioral and physiological measures of human spatial selective auditory attention","volume":"115","author":"Dai","year":"2018","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Yang, J., and Yang, G. (2018). Modified convolutional neural network based on dropout and the stochastic gradient descent optimizer. Algorithms, 11.","DOI":"10.3390\/a11030028"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"101027","DOI":"10.1016\/j.csl.2019.101027","article-title":"Voxceleb: Large-scale speaker verification in the wild","volume":"60","author":"Nagrani","year":"2020","journal-title":"Comput. Speech Lang."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"2195","DOI":"10.1109\/TASLP.2020.3009494","article-title":"Tandem assessment of spoofing countermeasures and automatic speaker verification: Fundamentals","volume":"28","author":"Kinnunen","year":"2020","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1016\/j.specom.2009.08.009","article-title":"An overview of text-independent speaker recognition: From features to supervectors","volume":"52","author":"Kinnunen","year":"2010","journal-title":"Speech Commun."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/23\/6784\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:38:28Z","timestamp":1760179108000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/23\/6784"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,11,27]]},"references-count":33,"journal-issue":{"issue":"23","published-online":{"date-parts":[[2020,12]]}},"alternative-id":["s20236784"],"URL":"https:\/\/doi.org\/10.3390\/s20236784","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,11,27]]}}}