{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T16:44:11Z","timestamp":1778085851387,"version":"3.51.4"},"reference-count":46,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2023,1,20]],"date-time":"2023-01-20T00:00:00Z","timestamp":1674172800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["61972324"],"award-info":[{"award-number":["61972324"]}]},{"name":"National Natural Science Foundation of China","award":["2021YFS0313"],"award-info":[{"award-number":["2021YFS0313"]}]},{"name":"National Natural Science Foundation of China","award":["2021YFG0133"],"award-info":[{"award-number":["2021YFG0133"]}]},{"name":"Sichuan Science and Technology Program","award":["61972324"],"award-info":[{"award-number":["61972324"]}]},{"name":"Sichuan Science and Technology Program","award":["2021YFS0313"],"award-info":[{"award-number":["2021YFS0313"]}]},{"name":"Sichuan Science and Technology Program","award":["2021YFG0133"],"award-info":[{"award-number":["2021YFG0133"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In speaker recognition tasks, convolutional neural network (CNN)-based approaches have shown significant success. Modeling the long-term contexts and efficiently aggregating the information are two challenges in speaker recognition, and they have a critical impact on system performance. Previous research has addressed these issues by introducing deeper, wider, and more complex network architectures and aggregation methods. However, it is difficult to significantly improve the performance with these approaches because they also have trouble fully utilizing global information, channel information, and time-frequency information. To address the above issues, we propose a lighter and more efficient CNN-based end-to-end speaker recognition architecture, ResSKNet-SSDP. ResSKNet-SSDP consists of a residual selective kernel network (ResSKNet) and self-attentive standard deviation pooling (SSDP). ResSKNet can capture long-term contexts, neighboring information, and global information, thus extracting a more informative frame-level. SSDP can capture short- and long-term changes in frame-level features, aggregating the variable-length frame-level features into fixed-length, more distinctive utterance-level features. Extensive comparison experiments were performed on two popular public speaker recognition datasets, Voxceleb and CN-Celeb, with current state-of-the-art speaker recognition systems and achieved the lowest EER\/DCF of 2.33%\/0.2298, 2.44%\/0.2559, 4.10%\/0.3502, and 12.28%\/0.5051. Compared with the lightest x-vector, our designed ResSKNet-SSDP has 3.1 M fewer parameters and 31.6 ms less inference time, but 35.1% better performance. The results show that ResSKNet-SSDP significantly outperforms the current state-of-the-art speaker recognition architectures on all test sets and is an end-to-end architecture with fewer parameters and higher efficiency for applications in realistic situations. The ablation experiments further show that our proposed approaches also provide significant improvements over previous methods.<\/jats:p>","DOI":"10.3390\/s23031203","type":"journal-article","created":{"date-parts":[[2023,1,20]],"date-time":"2023-01-20T06:52:41Z","timestamp":1674197561000},"page":"1203","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["ResSKNet-SSDP: Effective and Light End-To-End Architecture for Speaker Recognition"],"prefix":"10.3390","volume":"23","author":[{"given":"Fei","family":"Deng","sequence":"first","affiliation":[{"name":"College of Computer Science and Cyber Security (Oxford Brookes College), Chengdu University of Technology, Chengdu 610059, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2090-7158","authenticated-orcid":false,"given":"Lihong","family":"Deng","sequence":"additional","affiliation":[{"name":"College of Computer Science and Cyber Security (Oxford Brookes College), Chengdu University of Technology, Chengdu 610059, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peifan","family":"Jiang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Cyber Security (Oxford Brookes College), Chengdu University of Technology, Chengdu 610059, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gexiang","family":"Zhang","sequence":"additional","affiliation":[{"name":"Artificial Intelligence Research Center, Chengdu University of Technology, Chengdu 610059, China"},{"name":"School of Control Engineering, Chengdu University of Information Engineering, Chengdu 610059, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiang","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Control Engineering, Chengdu University of Information Engineering, Chengdu 610059, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,1,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Leonardis, A., Bischof, H., and Pinz, A. (2006, January 7\u201313). Probabilistic Linear Discriminant Analysis. Proceedings of the European Conference on Computer Vision 2006, Graz, Austria. Lecture Notes in Computer Science.","DOI":"10.1007\/11744023"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"788","DOI":"10.1109\/TASL.2010.2064307","article-title":"Front-End Factor Analysis for Speaker Verification","volume":"19","author":"Dehak","year":"2011","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Cai, W., Chen, J., and Li, M. (2018, January 26\u201329). Exploring the Encoding Layer and Loss Function in End-to-End Speaker and Language Recognition System. Proceedings of the Speaker and Language Recognition Workshop (Odyssey 2018), Les Sables d\u2019Olonne, France.","DOI":"10.21437\/Odyssey.2018-11"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Variani, E., Lei, X., McDermott, E., Moreno, I.L., and Gonzalez-Dominguez, J. (2014, January 4\u20139). Deep Neural Networks for Small Footprint Text-Dependent Speaker Verification. Proceedings of the 2014 IEEE International Conference on Acoustics, Speech, and Signal Processing, Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6854363"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_6","unstructured":"Rahman Chowdhury, F.R., Wang, Q., Moreno, I.L., and Wan, L. (2018, January 15\u201320). Attention-Based Models for Text-Dependent Speaker Verification. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Calgary, AB, Canada."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Snyder, D., Garcia-Romero, D., Sell, G., Povey, D., and Khudanpur, S. (2018, January 15\u201320). X-Vectors: Robust DNN Embeddings for Speaker Recognition. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461375"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Kanagasundaram, A., Sridharan, S., Sriram, G., Prachi, S., and Fookes, C. (2019, January 15\u201319). A Study of X-Vector Based Speaker Recognition on Short Utterances. Proceedings of the 20th Annual Conference of the International Speech Communication Association, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-1891"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Snyder, D., Garcia-Romero, D., Sell, G., McCree, A., Povey, D., and Khudanpur, S. (2019, January 12\u201317). Speaker Recognition for Multi-Speaker Conversations Using X-Vectors. Proceedings of the 2019 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683760"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Povey, D., Cheng, G., Wang, Y., Li, K., Xu, H., Yarmohammadi, M., and Khudanpur, S. (2018, January 2\u20136). Semi-Orthogonal Low-Rank Matrix Factorization for Deep Neural Networks. Proceedings of the Interspeech 2018, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-1417"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Yu, Y.-Q., and Li, W.-J. (2020, January 25\u201329). Densely Connected Time Delay Neural Network for Speaker Verification. Proceedings of the Interspeech 2020, Shanghai, China.","DOI":"10.21437\/Interspeech.2020-1275"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Desplanques, B., Thienpondt, J., and Demuynck, K. (2020, January 25\u201329). ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification. Proceedings of the Interspeech 2020, Shanghai, China.","DOI":"10.21437\/Interspeech.2020-2650"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1016\/j.neucom.2019.08.046","article-title":"Self-attention-based speaker recognition using cluster-range loss","volume":"368","author":"Bian","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_14","unstructured":"Heo, H.S., Lee, B.J., Huh, J., and Chung, J.S. (2020). Clova Baseline System for the Voxceleb Speaker Recognition Challenge 2020. arXiv."},{"key":"ref_15","unstructured":"Yao, W., Chen, S., Cui, J., and Lou, Y. (2020). Multi-Stream Convolutional Neural Network with Frequency Selection for Robust Speaker Verification. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhou, T., Zhao, Y., and Wu, J. (2021, January 19\u201322). ResNeXt and Res2Net Structures for Speaker Verification. Proceedings of the 2021 IEEE Spoken Language Technology Workshop (SLT), Shenzhen, China.","DOI":"10.1109\/SLT48900.2021.9383531"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"2011","DOI":"10.1109\/TPAMI.2019.2913372","article-title":"Squeeze-and-Excitation Networks","volume":"42","author":"Hu","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Li, X., Wang, W., Hu, X., and Yang, J. (2019, January 15\u201320). Selective Kernel Networks. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00060"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"101027","DOI":"10.1016\/j.csl.2019.101027","article-title":"Voxceleb: Large-scale speaker verification in the wild","volume":"60","author":"Nagrani","year":"2020","journal-title":"Comput. Speech Lang."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wang, M., Feng, D., Su, T., and Chen, M. (2022). Attention-Based Temporal-Frequency Aggregation for Speaker Verification. Sensors, 22.","DOI":"10.3390\/s22062147"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Chung, J.S., Huh, J., and Mun, S. (2020, January 2\u20135). Delving into VoxCeleb: Environment Invariant Speaker Recognition. Proceedings of the Odyssey 2020: The Speaker and Language Recognition Workshop, Tokyo, Japan.","DOI":"10.21437\/Odyssey.2020-49"},{"key":"ref_22","unstructured":"Kye, S.M., Chung, J.S., and Kim, H. (2021, January 19-22). Supervised Attention for Speaker Recognition. Proceedings of the 2021 IEEE Spoken Language Technology Workshop (SLT), Shenzhen, China."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Okabe, K., Koshinaka, T., and Shinoda, K. (2018, January 2\u20136). Attentive Statistics Pooling for Deep Speaker Embedding. Proceedings of the Interspeech 2018, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-993"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Georges, M., Huang, J., and Bocklet, T. (2020, January 25\u201329). Compact speaker embedding: Lrx-vector. Proceedings of the Interspeech 2020, Shanghai, China.","DOI":"10.21437\/Interspeech.2020-2106"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"15839","DOI":"10.1109\/JSEN.2020.3022536","article-title":"Ultra-Lightweight Mutual Authentication in the Vehicle Based on Smart Contract Blockchain: Case of MITM Attack","volume":"21","author":"Razmjouei","year":"2021","journal-title":"IEEE Sens. J."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Sahba, A., Sahba, R., Rad, P., and Jamshidi, M. (2019, January 10\u201312). Optimized IoT Based Decision Making For Autonomous Vehicles in Intersections. Proceedings of the 2019 IEEE 10th Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON), New York, NY, USA.","DOI":"10.1109\/UEMCON47517.2019.8992978"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"102520","DOI":"10.1016\/j.adhoc.2021.102520","article-title":"IoT-based data-driven fault allocation in microgrids using advanced \u00b5PMUs","volume":"119","author":"Nikkhah","year":"2021","journal-title":"Ad Hoc Netw."},{"key":"ref_28","first-page":"1437","article-title":"NetVLAD: CNN Architecture for Weakly Supervised Place Recognition","volume":"40","author":"Gronat","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"926","DOI":"10.1109\/LSP.2018.2822810","article-title":"Additive Margin Softmax for Face Verification","volume":"25","author":"Wang","year":"2018","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W., and Hu, Q. (2020, January 13\u201319). ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01155"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Lee, J., and Nam, J. (2017). Multi-Level and Multi-Scale Feature Aggregation Using Sample-Level Deep Convolutional Neural Networks for Music Classification. arXiv.","DOI":"10.1109\/LSP.2017.2713830"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Gao, Z., Song, Y., McLoughlin, I., Li, P., Jiang, Y., and Dai, L.-R. (, January 15\u201319). Improving Aggregation and Loss Function for Better Embedding Learning in End-to-End Speaker Verification System. Proceedings of the Interspeech, 2019, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-1489"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"201","DOI":"10.1016\/j.neunet.2021.03.014","article-title":"D-MONA: A dilated mixed-order non-local attention network for speaker and language recognition","volume":"139","author":"Miao","year":"2021","journal-title":"Neural Netw."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"77","DOI":"10.1016\/j.specom.2022.01.002","article-title":"CN-Celeb: Multi-genre speaker recognition","volume":"137","author":"Li","year":"2022","journal-title":"Speech Commun."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Park, D.S., Chan, W., Zhang, Y., Chiu, C.-C., Zoph, B., Cubuk, E.D., and Le, Q.V. (2019, January 15\u201319). Specaugment: A simple data augmentation method for automatic speech recognition. Proceedings of the Interspeech 2019, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-2680"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1853","DOI":"10.1109\/TASLP.2022.3178225","article-title":"Neonatal Bowel Sound Detection Using Convolutional Neural Network and Laplace Hidden Semi-Markov Model","volume":"30","author":"Sitaula","year":"2022","journal-title":"IEEE ACM Trans. Audio Speech Lang. Process."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Burne, L., Sitaula, C., Priyadarshi, A., Tracy, M., Kavehei, O., Hinder, M., Withana, A., McEwan, A., and Marzbanrad, F. (2022). Ensemble Approach on Deep and Handcrafted Features for Neonatal Bowel Sound Detection. IEEE J. Biomed. Health Inform.","DOI":"10.1109\/JBHI.2022.3217559"},{"key":"ref_38","unstructured":"Kingma, D., and Ba, J. (2014, January 14\u201316). Adam: A Method for Stochastic Optimization. Proceedings of the 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Sadjadi, S.O., Kheyrkhah, T., Tong, A., Greenberg, C., Reynolds, D., Singer, E., Mason, L., and Hernandez-Cordero, J. (2017, January 20\u201324). The 2016 Nist Speaker Recognition Evaluation. Proceedings of the Interspeech 2017: Conference of the International Speech Communication Association, Stockholm, Sweden.","DOI":"10.21437\/Interspeech.2017-458"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Wang, R., Ao, J., Zhou, L., Liu, S., Wei, Z., Ko, T., Li, Q., and Zhang, Y. (2021). Multi-View Self-Attention Based Transformer for Speaker Recognition. arXiv.","DOI":"10.1109\/ICASSP43922.2022.9746639"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Jung, J.W., Kim, Y.J., Heo, H.S., Lee, B.J., Kwon, Y., and Chung, J.S. (2022). Pushing the Limits of Raw Waveform Speaker Recognition. arXiv.","DOI":"10.21437\/Interspeech.2022-126"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Wei, Y., Du, J., Liu, H., and Wang, Q. (2022, January 18\u201322). CTFALite: Lightweight Channel-specific Temporal and Frequency Attention Mechanism for Enhancing the Speaker Embedding Extractor. Proceedings of the Interspeech 2022, Incheon, Korea.","DOI":"10.21437\/Interspeech.2022-10288"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Zhu, G., Jiang, F., and Duan, Z. (September, January 30). Y-Vector: Multiscale Waveform Encoder for Speaker Embedding. Proceedings of the Interspeech 2021, Brno, Czech Republic.","DOI":"10.21437\/Interspeech.2021-1707"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"404","DOI":"10.1109\/TASLP.2021.3134566","article-title":"S-Vectors and TESA: Speaker Embeddings and a Speaker Authenticator Based on Transformer Encoder","volume":"30","author":"Mary","year":"2020","journal-title":"IEEE ACM Trans. Audio Speech Lang. Process."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Li, J., Liu, W., and Lee, T. (2022, January 18\u201322). EDITnet: A Lightweight Network for Unsupervised Domain Adaptation in Speaker Verification. Proceedings of the Interspeech 2022, Incheon, Korea.","DOI":"10.21437\/Interspeech.2022-967"},{"key":"ref_46","first-page":"2579","article-title":"Visualizing Data using t-SNE","volume":"9","author":"Hinton","year":"2008","journal-title":"J. Mach. Learn. Res."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/3\/1203\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:12:00Z","timestamp":1760119920000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/3\/1203"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,20]]},"references-count":46,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,2]]}},"alternative-id":["s23031203"],"URL":"https:\/\/doi.org\/10.3390\/s23031203","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,20]]}}}