{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,15]],"date-time":"2025-07-15T03:34:13Z","timestamp":1752550453581,"version":"3.41.0"},"publisher-location":"New York, New York, USA","reference-count":40,"publisher":"ACM Press","license":[{"start":{"date-parts":[[2018,1,1]],"date-time":"2018-01-01T00:00:00Z","timestamp":1514764800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2018]]},"DOI":"10.1145\/3287921.3287954","type":"proceedings-article","created":{"date-parts":[[2018,12,13]],"date-time":"2018-12-13T15:45:16Z","timestamp":1544715916000},"page":"177-184","source":"Crossref","is-referenced-by-count":5,"title":["Vietnamese Speaker Authentication Using Deep Models"],"prefix":"10.1145","author":[{"given":"Son T.","family":"Nguyen","sequence":"first","affiliation":[{"name":"Posts and Telecommunications, Institute of Technology, Hanoi, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Viet D.","family":"Lai","sequence":"additional","affiliation":[{"name":"University of Oregon, Eugene, Oregon, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Quyen","family":"Dam-Ba","sequence":"additional","affiliation":[{"name":"Posts and Telecommunications, Institute of Technology, Hanoi, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anh","family":"Nguyen-Xuan","sequence":"additional","affiliation":[{"name":"Posts and Telecommunications, Institute of Technology, Hanoi, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cuong","family":"Pham","sequence":"additional","affiliation":[{"name":"Posts and Telecommunications, Institute of Technology, Hanoi, Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","reference":[{"key":"key-10.1145\/3287921.3287954-1","doi-asserted-by":"crossref","unstructured":"W. M. Campbell, J. P. Campbell, T. P. Gleason, D. A. Reynolds, and W. Shen. 2007. Speaker Verification Using Support Vector Machines and High-Level Features. IEEE Transactions on Audio, Speech, and Language Processing 15, 7 (Sept 2007), 2085--2094. https:\/\/doi.org\/10.1109\/TASL.2007.902874","DOI":"10.1109\/TASL.2007.902874"},{"key":"key-10.1145\/3287921.3287954-2","unstructured":"Shi-Huang Chen and Yu-Ren Luo. 2009. Speaker verification using MFCC and support vector machine. In Proceedings of the International MultiConference of Engineers and Computer Scientists, Vol. 1. 18--20."},{"key":"key-10.1145\/3287921.3287954-3","unstructured":"Namrata Dave. 2013. Feature extraction methods LPC, PLP and MFCC in speech recognition. International journal for advance research in engineering and technology 1, 6 (2013), 1--4."},{"key":"key-10.1145\/3287921.3287954-4","doi-asserted-by":"crossref","unstructured":"Nguyen Ngoc Diep, Cuong Pham, and Tu Minh Phuong. 2016. An Orientation Histogram Based Approach for Fall Detection Using Wearable Sensors. In PRICAI 2016: Trends in Artificial Intelligence - 14th Pacific Rim International Conference on Artificial Intelligence, Phuket, Thailand, August 22-26, 2016, Proceedings. 354--366. https:\/\/doi.org\/10.1007\/978-3-319-42911-3_30","DOI":"10.1007\/978-3-319-42911-3_30"},{"key":"key-10.1145\/3287921.3287954-5","unstructured":"Zhenhao Ge, Ananth N Iyer, Srinath Cheluvaraja, Ram Sundaram, and Aravind Ganapathiraju. 2017. Neural network based speaker classification and verification systems with enhanced features. In Intelligent Systems Conference (IntelliSys), 2017. IEEE, 1089--1094."},{"key":"key-10.1145\/3287921.3287954-6","doi-asserted-by":"crossref","unstructured":"Alex Graves, Santiago Fern&#225;ndez, Faustino Gomez, and J&#252;rgen Schmidhuber. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd international conference on Machine learning. ACM, 369--376.","DOI":"10.1145\/1143844.1143891"},{"key":"key-10.1145\/3287921.3287954-7","unstructured":"Md Rashidul Hasan, Mustafa Jamil, MGRMS Rahman, et al. 2004. Speaker identification using mel frequency cepstral coefficients. variations 1, 4 (2004)."},{"key":"key-10.1145\/3287921.3287954-8","doi-asserted-by":"crossref","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition. 770--778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"key-10.1145\/3287921.3287954-9","doi-asserted-by":"crossref","unstructured":"M. M. Homayounpour and I. Rezaian. 2008. Robust Speaker Verification Based on Multi Stage Vector Quantization of MFCC Parameters on Narrow Bandwidth Channels. In 2008 10th International Conference on Advanced Communication Technology, Vol. 1. 336--340. https:\/\/doi.org\/10.1109\/ICACT.2008.4493773","DOI":"10.1109\/ICACT.2008.4493773"},{"key":"key-10.1145\/3287921.3287954-10","doi-asserted-by":"crossref","unstructured":"Nguyen Hong Quang, Loan Trinh Van, and Le The Dat. 2010. Automatic Speech Recognition for Vietnamese Using HTK System. (11 2010).","DOI":"10.1109\/RIVF.2010.5633587"},{"key":"key-10.1145\/3287921.3287954-11","unstructured":"H. Hussain, S. H. Salleh, C. M. Ting, A. K. Ariff, I. Kamarulafizam, and R. A. Suraya. 2011. Speaker Verification Using Gaussian Mixture Model (GMM). In 5th Kuala Lumpur International Conference on Biomedical Engineering 2011, Noor Azuan Abu Osman, Wan Abu Bakar Wan Abas, Ahmad Khairi Abdul Wahab, and Hua-Nong Ting (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 560--564."},{"key":"key-10.1145\/3287921.3287954-12","unstructured":"Chadawan Ittichaichareon, Siwat Suksri, and Thaweesak Yingthawornsuk. 2012. Speech recognition using MFCC. In International Conference on Computer Graphics, Simulation and Modeling (ICGSM'2012) July. 28--29."},{"key":"key-10.1145\/3287921.3287954-13","doi-asserted-by":"crossref","unstructured":"S. S. Jagtap and D. G. Bhalke. 2015. Speaker verification using Gaussian Mixture Model. In 2015 International Conference on Pervasive Computing (ICPC). 1--5. https:\/\/doi.org\/10.1109\/PERVASIVE.2015.7087080","DOI":"10.1109\/PERVASIVE.2015.7087080"},{"key":"key-10.1145\/3287921.3287954-14","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097--1105."},{"key":"key-10.1145\/3287921.3287954-15","doi-asserted-by":"crossref","unstructured":"Yun Lei, Luciana Ferrer, Mitchell McLaren, and Nicolas Scheffer. 2014. A deep neural network speaker verification system targeting microphone speech. In Fifteenth Annual Conference of the International Speech Communication Association.","DOI":"10.21437\/Interspeech.2014-171"},{"key":"key-10.1145\/3287921.3287954-16","unstructured":"Chris Lengerich and Awni Hannun. 2016. An End-to-End Architecture for Keyword Spotting and Voice Activity Detection. (11 2016)."},{"key":"key-10.1145\/3287921.3287954-17","unstructured":"Hong Lu, A.J. Brush, Bodhi Priyantha, Amy Karlson, and Jie Liu. 2011. SpeakerSense: Energy Efficient Unobtrusive Speaker Identification on Mobile Phones. https:\/\/www.microsoft.com\/en-us\/research\/publication\/speakersense-energy-efficient-unobtrusive-speaker-identification-on-mobile-phones\/"},{"key":"key-10.1145\/3287921.3287954-18","unstructured":"Jorge Martinez, Hector Perez, Enrique Escamilla, and Masahisa Mabo Suzuki. 2012. Speaker recognition using Mel frequency Cepstral Coefficients (MFCC) and Vector quantization (VQ) techniques. In Electrical Communications and Computers (CONIELECOMP), 2012 22nd International Conference on. IEEE, 248--251."},{"key":"key-10.1145\/3287921.3287954-19","doi-asserted-by":"crossref","unstructured":"A. Mezghani and D. O'Shaughnessy. 2005. Speaker verification using a new representation based on a combination of MFCC and formants. In Canadian Conference on Electrical and Computer Engineering, 2005. 1461--1464. https:\/\/doi.org\/10.1109\/CCECE.2005.1557255","DOI":"10.1109\/CCECE.2005.1557255"},{"key":"key-10.1145\/3287921.3287954-20","unstructured":"Lindasalwa Muda, Mumtaj Begam, and Irraivan Elamvazuthi. 2010. Voice recognition algorithms using mel frequency cepstral coefficient (MFCC) and dynamic time warping (DTW) techniques. arXiv preprint arXiv:1003.4083 (2010)."},{"key":"key-10.1145\/3287921.3287954-21","doi-asserted-by":"crossref","unstructured":"Guiwen Ou and Dengfeng Ke. 2004. Text-independent speaker verification based on relation of MFCC components. In 2004 International Symposium on Chinese Spoken Language Processing. 57--60. https:\/\/doi.org\/10.1109\/CHINSL.2004.1409585","DOI":"10.1109\/CHINSL.2004.1409585"},{"key":"key-10.1145\/3287921.3287954-22","unstructured":"RV Pawar, PP Kajave, and SN Mali. 2005. Speaker Identification using Neural Networks.. In IEC (Prague). 429--433."},{"key":"key-10.1145\/3287921.3287954-23","doi-asserted-by":"crossref","unstructured":"Cuong Pham. 2015. MobiRAR: Real-Time Human Activity Recognition Using Mobile Devices. In 2015 Seventh International Conference on Knowledge and Systems Engineering, KSE 2015, Ho Chi Minh City, Vietnam, October 8-10, 2015. 144--149. https:\/\/doi.org\/10.1109\/KSE.2015.43","DOI":"10.1109\/KSE.2015.43"},{"key":"key-10.1145\/3287921.3287954-24","doi-asserted-by":"crossref","unstructured":"Cuong Pham. 2016. MobiCough: Real-Time Cough Detection and Monitoring Using Low-Cost Mobile Devices. In Intelligent Information and Database Systems - 8th Asian Conference, ACIIDS 2016, Da Nang, Vietnam, March 14-16, 2016, Proceedings, Part I. 300--309. https:\/\/doi.org\/10.1007\/978-3-662-49381-6_29","DOI":"10.1007\/978-3-662-49381-6_29"},{"key":"key-10.1145\/3287921.3287954-25","unstructured":"Cuong Pham, Nguyen Ngoc Diep, and Tu Minh Phuong. 2013. A Wearable Sensor based Approach to Real-Time Fall Detection and Fine-Grained Activity Recognition. J. Mobile Multimedia 9, 1&2 (2013), 15--26. http:\/\/www.rintonpress.com\/journals\/jmm\/abstractsJmm9-12.html"},{"key":"key-10.1145\/3287921.3287954-26","unstructured":"Cuong Pham, Nguyen N. Diep, and Tu M. Phuong. 2017. e-Shoes: Smart shoes for unobtrusive human activity recognition. In KSE, 2017. IEEE, 269--274."},{"key":"key-10.1145\/3287921.3287954-27","unstructured":"Cuong Pham, Clare Hooper, Stephen Lindsay, D. Jackson, J. Shearer, J. Wagner, C. Ladha, K. Ladha, T. Ploetz, Patrick Olivier, et al. 2012. The ambient kitchen: a pervasive sensing environment for situated services. (2012)."},{"key":"key-10.1145\/3287921.3287954-28","doi-asserted-by":"crossref","unstructured":"Cuong Pham and Patrick Olivier. 2009. Slice&dice: Recognizing food preparation activities using embedded accelerometers. In AmI. Springer, 34--43.","DOI":"10.1007\/978-3-642-05408-2_4"},{"key":"key-10.1145\/3287921.3287954-29","doi-asserted-by":"crossref","unstructured":"Cuong Pham and Tu Minh Phuong. 2013. Real-Time Fall Detection and Activity Recognition Using Low-Cost Wearable Sensors. In Computational Science and Its Applications - ICCSA 2013 - 13th International Conference, Ho Chi Minh City, Vietnam, June 24-27, 2013, Proceedings, Part I. 673--682. https:\/\/doi.org\/10.1007\/978-3-642-39637-3_53","DOI":"10.1007\/978-3-642-39637-3_53"},{"key":"key-10.1145\/3287921.3287954-30","doi-asserted-by":"crossref","unstructured":"Cuong Pham and Nguyen Thi Thanh Thuy. 2016. Real-Time Traffic Activity Detection Using Mobile Devices. In Proceedings of the 10th International Conference on Ubiquitous Information Management and Communication, IMCOM 2016, Danang, Vietnam, January 4-6, 2016. 64:1--64:7. https:\/\/doi.org\/10.1145\/2857546.2857611","DOI":"10.1145\/2857546.2857611"},{"key":"key-10.1145\/3287921.3287954-31","unstructured":"Son Thanh Phan, Thang Tat Vu, and Mai Chi Luong. 2013. Extracting MFCC, F0 feature in Vietnamese HMM-based speech synthesis. International Journal of Electronics and Computer Science Engineering 2, 1 (2013), 46--52."},{"key":"key-10.1145\/3287921.3287954-32","unstructured":"Parminder Singh Ravneet Singh, Navpreet Kaur. 2017. Speech Based Biometric System Using GFCC Features. Imperial Journal of Interdisciplinary Research 3, 6 (2017), 1156--1160."},{"key":"key-10.1145\/3287921.3287954-33","unstructured":"Douglas Reynolds. 2015. Gaussian mixture models. Encyclopedia of biometrics (2015), 827--832."},{"key":"key-10.1145\/3287921.3287954-34","unstructured":"Douglas A Reynolds, Thomas F Quatieri, and Robert B Dunn. 2000. Speaker verification using adapted Gaussian mixture models. Digital signal processing 10, 1-3 (2000), 19--41."},{"key":"key-10.1145\/3287921.3287954-35","unstructured":"Yang Shao and DeLiang Wang. 2008. Robust speaker identification using auditory features and computational auditory scene analysis. In Acoustics, Speech and Signal Processing, 2008. ICASSP 2008. IEEE International Conference on. IEEE, 1589--1592."},{"key":"key-10.1145\/3287921.3287954-36","doi-asserted-by":"crossref","unstructured":"Visalakshmi Suresh, Paul D. Ezhilchelvan, Paul Watson, Cuong Pham, Daniel Jackson, and Patrick Olivier. 2011. Distributed event processing for activity recognition. In Proceedings of the Fifth ACM International Conference on Distributed Event-Based Systems, DEBS 2011, New York, NY, USA, July 11-15, 2011. 371--372. https:\/\/doi.org\/10.1145\/2002259.2002315","DOI":"10.1145\/2002259.2002315"},{"key":"key-10.1145\/3287921.3287954-37","doi-asserted-by":"crossref","unstructured":"Amirsina Torfi, Jeremy Dawson, and Nasser M Nasrabadi. 2017. Text-independent speaker verification using 3d convolutional neural networks. arXiv preprint arXiv:1705.09422 (2017).","DOI":"10.1109\/ICME.2018.8486441"},{"key":"key-10.1145\/3287921.3287954-38","doi-asserted-by":"crossref","unstructured":"Navneet Upadhyay and Abhijit Karmakar. 2015. Speech enhancement using spectral subtraction-type algorithms: A comparison and simulation study. Procedia Computer Science 54 (2015), 574--584.","DOI":"10.1016\/j.procs.2015.06.066"},{"key":"key-10.1145\/3287921.3287954-39","unstructured":"Vincent Wan. 2003. Speaker verification using support vector machines. Ph.D. Dissertation. Citeseer."},{"key":"key-10.1145\/3287921.3287954-40","doi-asserted-by":"crossref","unstructured":"Xiaojia Zhao and DeLiang Wang. 2013. Analyzing noise robustness of MFCC and GFCC features in speaker identification. In Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on. IEEE, 7204--7208.","DOI":"10.1109\/ICASSP.2013.6639061"}],"event":{"number":"9","sponsor":["SOICT, School of Information and Communication Technology - HUST","NAFOSTED, The National Foundation for Science and Technology Development"],"acronym":"SoICT 2018","name":"the Ninth International Symposium","start":{"date-parts":[[2018,12,6]]},"location":"Danang City, Viet Nam","end":{"date-parts":[[2018,12,7]]}},"container-title":["Proceedings of the Ninth International Symposium on Information and Communication Technology - SoICT 2018"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3287921.3287954","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/dl.acm.org\/ft_gateway.cfm?id=3287954&ftid=2025940&dwn=1","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:57:54Z","timestamp":1750208274000},"score":1,"resource":{"primary":{"URL":"http:\/\/dl.acm.org\/citation.cfm?doid=3287921.3287954"}},"subtitle":[],"proceedings-subject":"Information and Communication Technology","short-title":[],"issued":{"date-parts":[[2018]]},"references-count":40,"URL":"https:\/\/doi.org\/10.1145\/3287921.3287954","relation":{},"subject":[],"published":{"date-parts":[[2018]]}}}