{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:20:57Z","timestamp":1760235657562,"version":"build-2065373602"},"reference-count":24,"publisher":"MDPI AG","issue":"18","license":[{"start":{"date-parts":[[2021,9,17]],"date-time":"2021-09-17T00:00:00Z","timestamp":1631836800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Two important tasks in many e-commerce applications are identity verification of the user accessing the system and determining the level of rights that the user has for accessing and manipulating system\u2019s resources. The performance of these tasks is directly dependent on the certainty of establishing the identity of the user. The main research focus of this paper is user identity verification approach based on voice recognition techniques. The paper presents research results connected to the usage of open-source speaker recognition technologies in e-commerce applications with an emphasis on evaluating the performance of the algorithms they use. Four open-source speaker recognition solutions (SPEAR, MARF, ALIZE, and HTK) have been evaluated in cases of mismatched conditions during training and recognition phases. In practice, mismatched conditions are influenced by various lengths of spoken sentences, different types of recording devices, and the usage of different languages in training and recognition phases. All tests conducted in this research were performed in laboratory conditions using the specially designed framework for multimodal biometrics. The obtained results show consistency with the findings of recent research which proves that i-vectors and solutions based on probabilistic linear discriminant analysis (PLDA) continue to be the dominant speaker recognition approaches for text-independent tasks.<\/jats:p>","DOI":"10.3390\/s21186231","type":"journal-article","created":{"date-parts":[[2021,9,22]],"date-time":"2021-09-22T03:47:35Z","timestamp":1632282455000},"page":"6231","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Evaluating the Performance of Speaker Recognition Solutions in E-Commerce Applications"],"prefix":"10.3390","volume":"21","author":[{"given":"Olja","family":"Kr\u010dadinac","sequence":"first","affiliation":[{"name":"Department for Information Technology, Faculty of Organizational Sciences, University of Belgrade, 11000 Belgrade, Serbia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Uro\u0161","family":"\u0160o\u0161evi\u0107","sequence":"additional","affiliation":[{"name":"Department for Information Technology, Faculty of Organizational Sciences, University of Belgrade, 11000 Belgrade, Serbia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Du\u0161an","family":"Star\u010devi\u0107","sequence":"additional","affiliation":[{"name":"Department for Information Technology, Faculty of Organizational Sciences, University of Belgrade, 11000 Belgrade, Serbia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,9,17]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"75","DOI":"10.1109\/MSP.2015.2462851","article-title":"Speaker Recognition by Machines and Humans: A Tutorial Review","volume":"32","author":"Hansen","year":"2015","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Jain, A.K., Flynn, P., and Ross, A.A. (2008). Introduction to Multibiometrics. Handbook of Biometrics, Springer.","DOI":"10.1007\/978-0-387-71041-9"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"720","DOI":"10.15837\/ijccc.2016.5.2683","article-title":"Continuous Distribution Approximation and Thresholds Optimization in Serial Multi-Modal Biometric Systems","volume":"11","year":"2016","journal-title":"Int. J. Comput. Commun. Control"},{"key":"ref_4","unstructured":"Kounoudes, A., Kekatos, V., and Mavromoustakos, S. (2006, January 24\u201328). Voice biometric authentication for enhancing Internet service security. Proceedings of the 2nd International Conference on Information & Communication Technologies, Damascus, Syria."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Grudin, J., and Jacques, R. (2019, January 4\u20139). Chatbots, humbots, and the quest for artificial general intelligence. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, Glasgow, Scotland.","DOI":"10.1145\/3290605.3300439"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1437","DOI":"10.1109\/5.628714","article-title":"Speaker recognition: A tutorial","volume":"85","author":"Campbell","year":"1997","journal-title":"Proc. IEEE"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1109\/89.365379","article-title":"Robust text-independent speaker identification using Gaussian mixture speaker models","volume":"3","author":"Reynolds","year":"1995","journal-title":"IEEE Trans. Speech Audio Process."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Rao, K.S., and Sarkar, S. (2014). Robust Speaker Recognition in Noisy Environments, Springer International Publishing.","DOI":"10.1007\/978-3-319-07130-5"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Ma, B., Meng, H.M., and Mak, M.W. (2007, January 15\u201320). Effects of device mismatch, language mismatch and environmental mismatch on speaker verification. Proceedings of the International Conference on Acoustics, Speech and Signal Processing-ICASSP\u201907, Honolulu, HI, USA.","DOI":"10.1109\/ICASSP.2007.366909"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"58","DOI":"10.1016\/j.specom.2017.09.004","article-title":"Modelling and compensation for language mismatch in speaker verification","volume":"96","author":"Misra","year":"2018","journal-title":"Speech Commun."},{"key":"ref_11","first-page":"1558","article-title":"Interoperability Framework for Multimodal Biometry: Open Source in Action","volume":"18","year":"2012","journal-title":"J. Univers. Comput. Sci."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1016\/j.specom.2009.08.009","article-title":"An overview of text-independent speaker recognition: From features to supervectors","volume":"52","author":"Kinnunen","year":"2010","journal-title":"Speech Commun."},{"key":"ref_13","unstructured":"Bekli, Z., and Ouda, W. (2018). A Performance Measurement of a Speaker Verification System Based on a Variance in Data Collection for Gaussian Mixture Model and Universal Background Model. [Master\u2019s Thesis, Malm\u00f6 Universitet]."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Richardson, F., Reynolds, D., and Dehak, N. (2015). A unified deep neural network for speaker and language recognition. arXiv.","DOI":"10.21437\/Interspeech.2015-299"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"101026","DOI":"10.1016\/j.csl.2019.101026","article-title":"State-of-the-art speaker recognition with neural network embeddings in NIST SRE18 and speakers in the wild evaluations","volume":"60","author":"Villalba","year":"2020","journal-title":"Comput. Speech Lang."},{"key":"ref_16","unstructured":"Pappagari, R., Cho, J., Moro-Velazquez, L., and Dehak, N. (2021, June 15). Using State of the Art Speaker Recognition and Natural Language Processing Technologies to Detect Alzheimer\u2019s Disease and Assess its Severity. Available online: https:\/\/www.researchgate.net\/profile\/Laureano-Moro-Velazquez\/publication\/346425054_Using_State_of_the_Art_Speaker_Recognition_and_Natural_Language_Processing_Technologies_to_Detect_Alzheimer\u2019s_Disease_and_Assess_its_Severity\/links\/60196d3a299bf1cc2698ff8e\/Using-State-of-the-Art-Speaker-Recognition-and-Natural-Language-Processing-Technologies-to-Detect-Alzheimers-Disease-and-Assess-its-Severity.pdf."},{"key":"ref_17","unstructured":"(2021, June 15). IDVoice Official Website. Available online: https:\/\/www.idrnd.ai\/text-dependent-voice-verification\/."},{"key":"ref_18","unstructured":"(2021, June 15). VoiSentry Official Website. Available online: https:\/\/www.aculab.com\/voice-biometrics\/."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Nainan, S., and Kulkarni, V. (2020). Enhancement in speaker recognition for optimized speech features using GMM, SVM and 1-D CNN. Int. J. Speech Technol., 1\u201314.","DOI":"10.1007\/s10772-020-09771-2"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Khoury, E., Shafey, L.E., and Marcel, S. (2014, January 4\u20139). Spear: An open source toolbox for speaker recognition based on Bob. Proceedings of the International Conference on Acoustic, Speech and Signal Processing (ICASSP), Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6853879"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Sobh, T. (2008). Introducing MARF: A Modular Audio Recognition Framework and its Applications for Scientific and Software Engineering Research. Advances in Computer and Information Sciences and Engineering, Springer.","DOI":"10.1007\/978-1-4020-8741-7"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Larcher, A., Bonastre, J.F., Fauve, B., Lee, B., Levy, C., Li, H., Mason, J., and Parfait, J.Y. (2013, January 25). ALIZE 3.0-Open Source Toolkit for State-of-the-Art Speaker Recognition. Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH 2013), Lyon, France.","DOI":"10.21437\/Interspeech.2013-634"},{"key":"ref_23","unstructured":"Young, S., Evermann, G., Gales, M., Hain, T., Kershaw, D., Liu, X., Moore, G., Odell, J., Ollason, D., and Povey, D. (2009). The HTK Book, Cambridge University Engineering Department."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Alam, J., Bhattacharya, G., and Kenny, P. (2018, January 26\u201329). Speaker Verification in Mismatched Conditions with Frustratingly Easy Domain Adaptation. Proceedings of the Odyssey 2018 The Speaker and Language Recognition Workshop, Les Sables d\u2019Olonne, France.","DOI":"10.21437\/Odyssey.2018-25"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/18\/6231\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:01:10Z","timestamp":1760166070000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/18\/6231"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,17]]},"references-count":24,"journal-issue":{"issue":"18","published-online":{"date-parts":[[2021,9]]}},"alternative-id":["s21186231"],"URL":"https:\/\/doi.org\/10.3390\/s21186231","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2021,9,17]]}}}