{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T02:03:23Z","timestamp":1782525803629,"version":"3.54.5"},"reference-count":145,"publisher":"Springer Science and Business Media LLC","issue":"21-23","license":[{"start":{"date-parts":[[2021,8,4]],"date-time":"2021-08-04T00:00:00Z","timestamp":1628035200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,8,4]],"date-time":"2021-08-04T00:00:00Z","timestamp":1628035200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"published-print":{"date-parts":[[2021,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The emergence of biometric technology provides enhanced security compared to the traditional identification and authentication techniques that were less efficient and secure. Despite the advantages brought by biometric technology, the existing biometric systems such as Automatic Speaker Verification (ASV) systems are weak against presentation attacks. A presentation attack is a spoofing attack launched to subvert an ASV system to gain access to the system. Though numerous Presentation Attack Detection (PAD) systems were reported in the literature, a systematic survey that describes the current state of research and application is unavailable. This paper presents a systematic analysis of the state-of-the-art voice PAD systems to promote further advancement in this area. The objectives of this paper are two folds: (i) to understand the nature of recent work on PAD systems, and (ii) to identify areas that require additional research. From the survey, a taxonomy of voice PAD and the trend analysis of recent work on PAD systems were built and presented, whereby the recent and relevant articles including articles from Interspeech and ICASSP Conferences, mostly indexed by Scopus, published between 2015 and 2021 were considered. A total of 172 articles were surveyed in this work. The findings of this survey present the limitation of recent works, which include spoof-type dependent PAD. Consequently, the future direction of work on voice PAD for interested researchers is established. The findings of this survey present the limitation of recent works, which include spoof-type dependent PAD. Consequently, the future direction of work on voice PAD for interested researchers is established.<\/jats:p>","DOI":"10.1007\/s11042-021-11235-x","type":"journal-article","created":{"date-parts":[[2021,8,4]],"date-time":"2021-08-04T17:05:58Z","timestamp":1628096758000},"page":"32725-32762","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":30,"title":["A survey on presentation attack detection for automatic speaker verification systems: State-of-the-art, taxonomy, issues and future direction"],"prefix":"10.1007","volume":"80","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7204-7305","authenticated-orcid":false,"given":"Choon Beng","family":"Tan","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0431-8967","authenticated-orcid":false,"given":"Mohd Hanafi Ahmad","family":"Hijazi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1078-8885","authenticated-orcid":false,"given":"Norazlina","family":"Khamis","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0627-5630","authenticated-orcid":false,"given":"Puteri Nor Ellyza binti","family":"Nohuddin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6881-7039","authenticated-orcid":false,"given":"Zuraini","family":"Zainol","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1026-6649","authenticated-orcid":false,"given":"Frans","family":"Coenen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4388-020X","authenticated-orcid":false,"given":"Abdullah","family":"Gani","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,8,4]]},"reference":[{"key":"11235_CR1","doi-asserted-by":"publisher","unstructured":"Abozaid A, Haggag A, Kasban H, Eltokhy M (2018) Multimodal biometric scheme for human authentication technique based on voice and face recognition fusion. Multimedia Tools and Applications. https:\/\/doi.org\/10.1007\/s11042-018-7012-3","DOI":"10.1007\/s11042-018-7012-3"},{"key":"11235_CR2","doi-asserted-by":"crossref","unstructured":"Adel M, Afify M, Gaballah A (2018) Text-Independent Speaker Verification Based on Deep Neural Networks and Segmental Dynamic Time Warping. 2018 IEEE Spoken Language Technology Workshop (SLT), pp 1001\u20131006, 1806.09932","DOI":"10.1109\/SLT.2018.8639574"},{"key":"11235_CR3","doi-asserted-by":"publisher","first-page":"101105","DOI":"10.1016\/j.csl.2020.101105","volume":"64","author":"M Adiban","year":"2020","unstructured":"Adiban M, Sameti H, Shehnepoor S (2020) Replay spoofing countermeasure using autoencoder and siamese networks on ASVspoof 2019 challenge. Computer Speech & Language 64:101105. https:\/\/doi.org\/10.1016\/j.csl.2020.101105","journal-title":"Computer Speech & Language"},{"key":"11235_CR4","unstructured":"Admuthe SS, Ghugardare S (2015) Survey paper on automatic speaker recognition systems. In: International conference on multimedia, computer graphics, and broadcasting international conference on signal processing, image processing, and pattern recognition, vol 4, pp 10895\u201310898"},{"key":"11235_CR5","doi-asserted-by":"publisher","unstructured":"Al-Ali AKH, Senadji B, Naik GR (2017) Enhanced forensic speaker verification using multi-run ICA in the presence of environmental noise and reverberation conditions. In: 2017 IEEE International conference on signal and image processing applications (ICSIPA), IEEE, pp 174\u2013179. https:\/\/doi.org\/10.1109\/ICSIPA.2017.8120601","DOI":"10.1109\/ICSIPA.2017.8120601"},{"key":"11235_CR6","unstructured":"ASVspoof (2019) ASVspoof 2019 Automatic Speaker Verification Spoofing and Countermeasures Challenge. https:\/\/www.asvspoof.org\/"},{"key":"11235_CR7","unstructured":"ASVspoof consortium (2019) ASVspoof 2019: Automatic Speaker Verification Spoofing and Countermeasures Challenge Evaluation Plan. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH 2019, pp 1\u201319"},{"issue":"1","key":"11235_CR8","doi-asserted-by":"publisher","first-page":"42","DOI":"10.1006\/dspr.1999.0360","volume":"10","author":"R Auckenthaler","year":"2000","unstructured":"Auckenthaler R, Carey M, Lloyd-Thomas H (2000) Score normalization for text-independent speaker verification systems. Digital Signal Processing: A Review Journal 10(1):42\u201354. https:\/\/doi.org\/10.1006\/dspr.1999.0360","journal-title":"Digital Signal Processing: A Review Journal"},{"key":"11235_CR9","doi-asserted-by":"publisher","first-page":"101132","DOI":"10.1016\/j.csl.2020.101132","volume":"65","author":"R Baumann","year":"2021","unstructured":"Baumann R, Malik KM, Javed A, Ball A, Kujawa B, Malik H (2021) Voice spoofing detection corpus for single and multi-order audio replays. Computer Speech & Language 65:101132. https:\/\/doi.org\/10.1016\/j.csl.2020.101132","journal-title":"Computer Speech & Language"},{"key":"11235_CR10","doi-asserted-by":"publisher","unstructured":"Billal K, Abdelhakim D (2017) A new speaker verification algorithm based on identification results. In: 2017 5Th international conference on electrical engineering - boumerdes (ICEE-B), IEEE, pp 1\u20136. https:\/\/doi.org\/10.1109\/ICEE-B.2017.8192139","DOI":"10.1109\/ICEE-B.2017.8192139"},{"key":"11235_CR11","unstructured":"Biometrics TF (2008) Biometrics Glossary (BG). https:\/\/www.hsdl.org\/?view&did=32101"},{"key":"11235_CR12","unstructured":"Biometrics Institute (2017) Types of Biometrics. https:\/\/www.biometricsinstitute.org\/types-of-biometrics"},{"key":"11235_CR13","doi-asserted-by":"publisher","unstructured":"Bonifaco H, Guzman KR, Jara JN, Jasareno AD, Zabala AC, Prado SV, Buenaventura CS (2017) Comparative analysis of filipino-based rhinolalia aperta speech using mel frequency cepstral analysis and Perceptual Linear Prediction. In: 2017IEEE 9Th international conference on humanoid, nanotechnology, information technology, communication and control, environment and management (HNICEM), IEEE, pp 1\u20136. https:\/\/doi.org\/10.1109\/HNICEM.2017.8269507","DOI":"10.1109\/HNICEM.2017.8269507"},{"key":"11235_CR14","doi-asserted-by":"publisher","unstructured":"Cai W, Cai D, Liu W, Li G, Li M (2017) Countermeasures for automatic speaker verification replay spoofing attack : on data augmentation, feature representation, classification and fusion. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 17\u201321. https:\/\/doi.org\/10.21437\/Interspeech.2017-906","DOI":"10.21437\/Interspeech.2017-906"},{"key":"11235_CR15","doi-asserted-by":"publisher","unstructured":"Chen Z, Xie Z, Zhang W, Xu X (2017) Resnet and Model Fusion for Automatic Spoofing Detection. In: Interspeech 2017, ISCA, ISCA, pp 102\u2013106. https:\/\/doi.org\/10.21437\/Interspeech.2017-1085","DOI":"10.21437\/Interspeech.2017-1085"},{"key":"11235_CR16","doi-asserted-by":"publisher","unstructured":"Chen Z, Zhang W, Xie Z, Xu X, Chen D (2018) Recurrent neural networks for automatic replay spoofing attack detection. In: 2018 IEEE International conference on acoustics, speech and signal processing (ICASSP), IEEE, pp 2052\u20132056. https:\/\/doi.org\/10.1109\/ICASSP.2018.8462644","DOI":"10.1109\/ICASSP.2018.8462644"},{"key":"11235_CR17","doi-asserted-by":"publisher","unstructured":"Chettri B, Sturm BL (2018) A deeper look at gaussian mixture model based Anti-Spoofing systems. In: ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, IEEE, pp 5159\u20135163. https:\/\/doi.org\/10.1109\/ICASSP.2018.8461467","DOI":"10.1109\/ICASSP.2018.8461467"},{"key":"11235_CR18","doi-asserted-by":"publisher","unstructured":"Chettri B, Mishra S, Sturm BL, Benetos E (2018) Analysing The Predictions Of a CNN-based Replay Spoofing Detection System. In: 2018 IEEE Spoken Language Technology Workshop (SLT), IEEE, pp 92\u201397. https:\/\/doi.org\/10.1109\/SLT.2018.8639666","DOI":"10.1109\/SLT.2018.8639666"},{"key":"11235_CR19","doi-asserted-by":"publisher","first-page":"3018","DOI":"10.1109\/TASLP.2020.3036777","volume":"28","author":"B Chettri","year":"2020","unstructured":"Chettri B, Benetos E, Sturm BLT (2020) Dataset Artefacts in Anti-Spoofing systems: A Case Study on the ASVspoof 2017 Benchmark. IEEE\/ACM Transactions on Audio, Speech, and Language Processing 28:3018\u20133028. https:\/\/doi.org\/10.1109\/TASLP.2020.3036777","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"11235_CR20","doi-asserted-by":"publisher","unstructured":"Das RK, Yang J, Li H (2019) Long range acoustic features for spoofed speech detection. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 1058\u20131062. https:\/\/doi.org\/10.21437\/Interspeech.2019-1887","DOI":"10.21437\/Interspeech.2019-1887"},{"issue":"4","key":"11235_CR21","doi-asserted-by":"publisher","first-page":"788","DOI":"10.1109\/TASL.2010.2064307","volume":"19","author":"N Dehak","year":"2011","unstructured":"Dehak N, Kenny PJ, Dehak R, Dumouchel P, Ouellet P (2011) Front-End Factor analysis for speaker verification. IEEE Transactions on Audio, Speech and Language Processing 19(4):788\u2013798","journal-title":"IEEE Transactions on Audio, Speech and Language Processing"},{"key":"11235_CR22","doi-asserted-by":"publisher","unstructured":"Delgado H, Todisco M, Sahidullah M, Evans N, Kinnunen T, Lee KA, Yamagishi J (2018) ASVspoof 2017 Version 2.0: meta-data analysis and baseline enhancements. In: Odyssey 2018 - The Speaker and Language Recognition Workshop, pp 296\u2013303. https:\/\/doi.org\/10.21437\/odyssey.2018-42","DOI":"10.21437\/odyssey.2018-42"},{"issue":"4","key":"11235_CR23","doi-asserted-by":"publisher","first-page":"671","DOI":"10.1109\/JSTSP.2017.2673807","volume":"11","author":"C Demiroglu","year":"2017","unstructured":"Demiroglu C, Buyuk O, Khodabakhsh A, Maia R (2017) Postprocessing synthetic speech with a complex cepstrum vocoder for spoofing phase-based synthetic speech detectors. IEEE J Select Top Signal Process 11(4):671\u2013683. https:\/\/doi.org\/10.1109\/JSTSP.2017.2673807","journal-title":"IEEE J Select Top Signal Process"},{"key":"11235_CR24","doi-asserted-by":"publisher","unstructured":"Dey S, Koshinaka T, Motlicek P, Madikeri S (2018) DNN based speaker embedding using content information for Text-Dependent speaker verification. In: 2018 IEEE International conference on acoustics, speech and signal processing (ICASSP), IEEE, pp 5344\u20135348. https:\/\/doi.org\/10.1109\/ICASSP.2018.8461389","DOI":"10.1109\/ICASSP.2018.8461389"},{"key":"11235_CR25","doi-asserted-by":"publisher","unstructured":"Dinkel H, Chen N, Qian Y, Yu K (2017) End-to-end spoofing detection with raw waveform CLDNNS. In: 2017 IEEE International conference on acoustics, speech and signal processing (ICASSP), IEEE, pp 4860\u20134864. https:\/\/doi.org\/10.1109\/ICASSP.2017.7953080","DOI":"10.1109\/ICASSP.2017.7953080"},{"key":"11235_CR26","doi-asserted-by":"publisher","unstructured":"Dua M, Jain C, Kumar S (2021) LSTM and CNN based ensemble approach for spoof detection task in automatic speaker verification systems. Journal of Ambient Intelligence and Humanized Computing. https:\/\/doi.org\/10.1007\/s12652-021-02960-0","DOI":"10.1007\/s12652-021-02960-0"},{"key":"11235_CR27","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-27733-7","volume-title":"Encyclopedia of biometrics","author":"N Evans","year":"2009","unstructured":"Evans N, Alegre F, Wu Z, Kinnunen T (2009) Encyclopedia of biometrics. Springer, Boston. https:\/\/doi.org\/10.1007\/978-3-642-27733-7"},{"key":"11235_CR28","doi-asserted-by":"publisher","unstructured":"Evans N, Kinnunen T, Yamagishi J, Wu Z, Alegre F, Leon PD (2014) Speaker Recognition Anti- Spoofing. Handbook of Biometric Anti-Spoofing pp 125\u2013146. https:\/\/doi.org\/10.1007\/978-1-4471-6524-8","DOI":"10.1007\/978-1-4471-6524-8"},{"key":"11235_CR29","doi-asserted-by":"crossref","unstructured":"Gomez-alanis A, Peinado AM, Gonzalez JA, Gomez AM (2018) A Deep Identity Representation for Noise Robust Spoofing Detection. In: Interspeech 2018, September, pp 676\u2013680","DOI":"10.21437\/Interspeech.2018-1909"},{"key":"11235_CR30","doi-asserted-by":"publisher","first-page":"108530","DOI":"10.1109\/ACCESS.2020.3000641","volume":"8","author":"A Gomez-Alanis","year":"2020","unstructured":"Gomez-Alanis A, Gonzalez-Lopez JA, Peinado AM (2020) A Kernel Density Estimation Based Loss Function and its Application to ASV-Spoofing Detection. IEEE Access 8:108530\u2013108543. https:\/\/doi.org\/10.1109\/ACCESS.2020.3000641","journal-title":"IEEE Access"},{"key":"11235_CR31","doi-asserted-by":"publisher","first-page":"1579","DOI":"10.1109\/TIFS.2020.3039045","volume":"16","author":"A Gomez-Alanis","year":"2021","unstructured":"Gomez-Alanis A, Gonzalez-Lopez JA, Dubagunta SP, Peinado AM, Magimai-Doss M (2021) On joint optimization of automatic speaker verification and Anti-Spoofing in the embedding space. IEEE Trans Inform Forensics Secur 16:1579\u20131593. https:\/\/doi.org\/10.1109\/TIFS.2020.3039045","journal-title":"IEEE Trans Inform Forensics Secur"},{"key":"11235_CR32","doi-asserted-by":"publisher","unstructured":"Goncalves AR, Violato RP, Korshunov P, Marcel S, Simoes FO (2017) On the generalization of fused systems in voice presentation attack detection. In: 2017 International conference of the biometrics special interest group (BIOSIG), IEEE, pp 1\u20135. https:\/\/doi.org\/10.23919\/BIOSIG.2017.8053516","DOI":"10.23919\/BIOSIG.2017.8053516"},{"key":"11235_CR33","doi-asserted-by":"publisher","unstructured":"Gong Y, Yang J, Huber J, MacKnight M, Poellabauer C (2019) REMASC: Realistic replay attack corpus for voice controlled systems. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 2355\u20132359. https:\/\/doi.org\/10.21437\/Interspeech.2019-1541, arXiv:1904.03365v2","DOI":"10.21437\/Interspeech.2019-1541"},{"key":"11235_CR34","doi-asserted-by":"publisher","first-page":"920","DOI":"10.1109\/LSP.2020.2996908","volume":"27","author":"Y Gong","year":"2020","unstructured":"Gong Y, Yang J, Poellabauer C (2020) Detecting replay attacks using multi-channel audio: A neural network-based method. IEEE Signal Process Lett 27:920\u2013924. https:\/\/doi.org\/10.1109\/LSP.2020.2996908, 2003.08225","journal-title":"IEEE Signal Process Lett"},{"key":"11235_CR35","doi-asserted-by":"publisher","unstructured":"Hanilci C (2017) Speaker verification anti-spoofing using linear prediction residual phase features. In: 2017 25th European Signal Processing Conference (EUSIPCO), IEEE, pp 96\u2013100 . https:\/\/doi.org\/10.23919\/EUSIPCO.2017.8081176","DOI":"10.23919\/EUSIPCO.2017.8081176"},{"key":"11235_CR36","doi-asserted-by":"publisher","first-page":"171","DOI":"10.1016\/j.dsp.2017.10.010","volume":"72","author":"C Hanil\u00e7i","year":"2018","unstructured":"Hanil\u00e7i C (2018) Data selection for i-vector based automatic speaker verification anti-spoofing. Digital Signal Process Rev J 72:171\u2013180. https:\/\/doi.org\/10.1016\/j.dsp.2017.10.010","journal-title":"Digital Signal Process Rev J"},{"key":"11235_CR37","unstructured":"Hanil\u00e7i C (2018) Features and classifiers for replay spoofing attack detection. In: 2017 10Th international conference on electrical and electronics engineering, ELECO 2017, pp 1187\u20131191"},{"key":"11235_CR38","doi-asserted-by":"crossref","unstructured":"Hanil\u00e7i C, Kinnunen T, Sahidullah M, Sizov A (2015) Classifiers for synthetic speech detection: A comparison. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 2057\u20132061","DOI":"10.21437\/Interspeech.2015-466"},{"issue":"10","key":"11235_CR39","doi-asserted-by":"publisher","first-page":"3037","DOI":"10.1166\/asl.2015.6490","volume":"21","author":"H Haviluddin","year":"2015","unstructured":"Haviluddin H, Alfred R, Obit J, Hijazi MHA, Ibrahim AAA (2015) A performance comparison of statistical and machine learning techniques in learning time series data. Adv Sci Lett 21(10):3037\u20133041. https:\/\/doi.org\/10.1166\/asl.2015.6490","journal-title":"Adv Sci Lett"},{"key":"11235_CR40","unstructured":"Heigold G, Moreno I, Bengio S, Shazeer N (2018) End-to-End text-dependent speaker verification. In: Acoustics, speech, and signal processing (ICASSP), International Conference, pp 3\u20137"},{"key":"11235_CR41","doi-asserted-by":"publisher","unstructured":"Hemavathi R, Kumaraswamy R (2021) Voice conversion spoofing detection by exploring artifacts estimates. Multimedia Tools and Applications . https:\/\/doi.org\/10.1007\/s11042-020-10212-0","DOI":"10.1007\/s11042-020-10212-0"},{"issue":"2","key":"11235_CR42","doi-asserted-by":"publisher","first-page":"1172","DOI":"10.1166\/asl.2018.10710","volume":"24","author":"MHA Hijazi","year":"2018","unstructured":"Hijazi MHA, Beng TC, Mountstephens J, Yuto L, Nisar K (2018) Malware Classification Using Ensemble Classifiers. Advanced Sci Lett 24 (2):1172\u20131176. https:\/\/doi.org\/10.1166\/asl.2018.10710","journal-title":"Advanced Sci Lett"},{"key":"11235_CR43","doi-asserted-by":"publisher","first-page":"377","DOI":"10.1016\/j.csl.2019.05.007","volume":"58","author":"I Himawan","year":"2019","unstructured":"Himawan I, Villavicencio F, Sridharan S, Fookes C (2019) Deep domain adaptation for anti-spoofing in speaker verification systems. Computer Speech and Language 58:377\u2013402. https:\/\/doi.org\/10.1016\/j.csl.2019.05.007","journal-title":"Computer Speech and Language"},{"key":"11235_CR44","doi-asserted-by":"publisher","unstructured":"Huang T, Wang H, Chen Y, He P (2020) GRU-SVM Model for Synthetic Speech Detection. In: Digital Forensics and Watermarking, pp 115\u2013125. https:\/\/doi.org\/10.1007\/978-3-030-43575-2","DOI":"10.1007\/978-3-030-43575-2"},{"key":"11235_CR45","unstructured":"Idiap Dataset Distribution Portal (2015) The AVspoof Database. https:\/\/www.idiap.ch\/dataset\/avspoof"},{"key":"11235_CR46","doi-asserted-by":"publisher","unstructured":"Jiang X, Wang S, Xiang X, Qian Y (2018) Integrating online i-vector into GMM-UBM for text-dependent speaker verification. In: Proceedings - 9th Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2017, pp 1628\u20131632. https:\/\/doi.org\/10.1109\/APSIPA.2017.8282293","DOI":"10.1109\/APSIPA.2017.8282293"},{"key":"11235_CR47","doi-asserted-by":"publisher","unstructured":"Jin M, Yoo CD (2010) Speaker verification and identification. Behavioral Biometrics for Human Identification, pp 264\u2013289. https:\/\/doi.org\/10.4018\/978-1-60566-725-6.ch013","DOI":"10.4018\/978-1-60566-725-6.ch013"},{"key":"11235_CR48","doi-asserted-by":"publisher","unstructured":"Kamble MR, Patil HA (2018) Novel energy separation based frequency modulation features for spoofed speech classification. In: 2017 9th International Conference on Advances in Pattern Recognition, ICAPR 2017, IEEE, pp 326\u2013331. https:\/\/doi.org\/10.1109\/ICAPR.2017.8593041","DOI":"10.1109\/ICAPR.2017.8593041"},{"key":"11235_CR49","doi-asserted-by":"publisher","unstructured":"Kamble MR, Sailor HB, Patil HA, Li H (2019) Advances in anti-spoofing: From the perspective of ASVspoof challenges. APSIPA Transactions on Signal and Information Processing 9. https:\/\/doi.org\/10.1017\/ATSIP.2019.21","DOI":"10.1017\/ATSIP.2019.21"},{"key":"11235_CR50","unstructured":"Kinnunen T, Evans N, Yamagishi J, Lee KA, Todisco M (2017) ASVSpoof 2017 : Automatic speaker verification spoofing and countermeasures challenge evaluation plan. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, pp 1\u20136"},{"key":"11235_CR51","doi-asserted-by":"publisher","unstructured":"Kinnunen T, Sahidullah M, Delgado H, Todisco M, Evans N, Yamagishi J, Lee KA (2017) The ASVspoof 2017 challenge: Assessing the limits of replay spoofing attack detection. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, pp 2\u20136. https:\/\/doi.org\/10.21437\/Interspeech.2017-1111","DOI":"10.21437\/Interspeech.2017-1111"},{"key":"11235_CR52","doi-asserted-by":"publisher","unstructured":"Kinnunen T, Sahidullah M, Falcone M, Costantini L, Hautam\u00e4ki R G, Thomsen D, Sarkar A, Tan ZH, Delgado H, Todisco M, Evans N, Hautam\u00e4ki V, Lee KA (2017) RedDots replayed: A new replay spoofing attack corpus for text-dependent speaker verification research. In: ICASSP IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, pp 5395\u20135399. https:\/\/doi.org\/10.1109\/ICASSP.2017.7953187","DOI":"10.1109\/ICASSP.2017.7953187"},{"key":"11235_CR53","doi-asserted-by":"publisher","unstructured":"Kinnunen T, Lee KA, Delgado H, Evans N, Todisco M, Sahidullah M, Yamagishi J, Reynolds DA (2018) t-DCF: A detection cost function for the tandem assessment of spoofing countermeasures and automatic speaker verification. In: Odyssey 2018 The Speaker and Language Recognition Workshop, pp 312\u2013319. https:\/\/doi.org\/10.21437\/odyssey.2018-44, 1804.09618","DOI":"10.21437\/odyssey.2018-44"},{"key":"11235_CR54","doi-asserted-by":"crossref","unstructured":"Korshunov P, Marcel S (2016) Cross-database evaluation of audio-based spoofing detection systems. In: INTERSPEECH 2016, pp 1705\u20131709","DOI":"10.21437\/Interspeech.2016-1326"},{"issue":"4","key":"11235_CR55","doi-asserted-by":"publisher","first-page":"695","DOI":"10.1109\/JSTSP.2017.2692389","volume":"11","author":"P Korshunov","year":"2017","unstructured":"Korshunov P, Marcel S (2017) Impact of score fusion on voice biometrics and presentation attack detection in cross-database evaluations. IEEE J Select Top Signal Process 11(4):695\u2013705. https:\/\/doi.org\/10.1109\/JSTSP.2017.2692389","journal-title":"IEEE J Select Top Signal Process"},{"key":"11235_CR56","unstructured":"Korshunov P, Marcel S (2017) Presentation attack detection in voice biometrics. In: Vielhauer C (ed)"},{"key":"11235_CR57","unstructured":"Kotta H, Patil AT, Acharya R, Patil HA (2020) Subband channel selection using teo for replay spoof detection in voice assistants. In: 2020 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), pp 538\u2013542"},{"issue":"1","key":"11235_CR58","doi-asserted-by":"publisher","first-page":"193","DOI":"10.1007\/s10772-020-09785-w","volume":"24","author":"AK Kumar","year":"2021","unstructured":"Kumar AK, Paul D, Pal M, Sahidullah M, Saha G (2021) Speech frame selection for spoofing detection with an application to partially spoofed audio-data. Int J Speech Technol 24(1):193\u2013203. https:\/\/doi.org\/10.1007\/s10772-020-09785-w","journal-title":"Int J Speech Technol"},{"key":"11235_CR59","doi-asserted-by":"publisher","unstructured":"Lai CI, Chen N, Villalba J, Dehak N (2019) ASSERT: Anti-spoofing with squeeze-excitation and residual networks. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 1013\u20131017. https:\/\/doi.org\/10.21437\/Interspeech.2019-1794, 1904.01120","DOI":"10.21437\/Interspeech.2019-1794"},{"key":"11235_CR60","doi-asserted-by":"publisher","unstructured":"Lavrentyeva G, Novoselov S, Malykh E, Kozlov A, Kudashev O, Shchemelinin V (2017) Audio Replay Attack Detection with Deep Learning Frameworks. In: Interspeech 2017, ISCA, ISCA, vol 2017-Augus, pp 82\u201386. https:\/\/doi.org\/10.21437\/Interspeech.2017-360","DOI":"10.21437\/Interspeech.2017-360"},{"key":"11235_CR61","doi-asserted-by":"crossref","unstructured":"Lee KA, Larcher A, Wang G, Kenny P, Br\u00fcmmer N, Van Leeuwen D, Aronowitz H, Kockmann M, Vaquero C, Ma B, Li H, Stafylakis T, Alam J, Swart A, Perez J (2015) The RedDots data collection for speaker recognition. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 2996\u20133000","DOI":"10.21437\/Interspeech.2015-95"},{"key":"11235_CR62","doi-asserted-by":"publisher","unstructured":"Lei Z, Yang Y, Liu C, Ye J (2020) Siamese convolutional neural network using gaussian probability feature for spoofing speech detection. In : Interspeech 2020, ISCA, ISCA, pp 1116\u20131120. https:\/\/doi.org\/10.21437\/Interspeech.2020-2723","DOI":"10.21437\/Interspeech.2020-2723"},{"key":"11235_CR63","doi-asserted-by":"publisher","first-page":"7907","DOI":"10.1109\/ACCESS.2020.2964048","volume":"8","author":"J Li","year":"2020","unstructured":"Li J, Sun M, Zhang X, Wang Y (2020) Joint decision of Anti-Spoofing and automatic speaker verification by Multi-Task learning with contrastive loss. IEEE Access 8:7907\u20137915. https:\/\/doi.org\/10.1109\/ACCESS.2020.2964048","journal-title":"IEEE Access"},{"key":"11235_CR64","doi-asserted-by":"publisher","unstructured":"Li L, Chen Y, Shi Y, Tang Z, Wang D (2017) Deep speaker feature learning for text-independent speaker verification. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 1542\u20131546. https:\/\/doi.org\/10.21437\/Interspeech.2017-452, 1705.03670","DOI":"10.21437\/Interspeech.2017-452"},{"key":"11235_CR65","unstructured":"Li SZ, Zhang D, Ma C, Shum HY, Chang E (2003) Learning to boost GMM based speaker verification. In: EUROSPEECH 2003 - 8th European Conference on Speech Communication and Technology, pp 1677\u20131680"},{"issue":"5","key":"11235_CR66","doi-asserted-by":"publisher","first-page":"982","DOI":"10.1109\/JSTSP.2020.2999828","volume":"14","author":"KM Malik","year":"2020","unstructured":"Malik KM, Javed A, Malik H, Irtaza A (2020) A light-weight replay detection framework for voice controlled IoT devices. IEEE J Select Top Signal Process 14(5):982\u2013996. https:\/\/doi.org\/10.1109\/JSTSP.2020.2999828","journal-title":"IEEE J Select Top Signal Process"},{"issue":"8","key":"11235_CR67","doi-asserted-by":"publisher","first-page":"2581","DOI":"10.1007\/s00521-017-2848-4","volume":"30","author":"AA Mallouh","year":"2018","unstructured":"Mallouh AA, Qawaqneh Z, Barkana BD (2018) New transformed features generated by deep bottleneck extractor and a GMM\u2013UBM classifier for speaker age and gender classification. Neural Comput Applic 30(8):2581\u20132593. https:\/\/doi.org\/10.1007\/s00521-017-2848-4","journal-title":"Neural Comput Applic"},{"key":"11235_CR68","unstructured":"Mariethoz J, Bengio S (2006) Can a Professional Imitator Fool a GMM-Based Speaker Verification System? Tech. rep. LIDIAP"},{"key":"11235_CR69","unstructured":"Markowitz J, Markowitz J, Road NS (2008) Speaker identification and verification (SIV ) applications and markets. Tech. rep., VoiceXML"},{"key":"11235_CR70","doi-asserted-by":"publisher","unstructured":"Mat\u011bjka P, Novotn\u00fd O, Plchot O, Burget L, S\u00e1nchez MD, C\u011brnock\u00fd JH Analysis of score normalization in multilingual speaker recognition. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 1567\u20131571. https:\/\/doi.org\/10.21437\/Interspeech.2017-803","DOI":"10.21437\/Interspeech.2017-803"},{"issue":"3","key":"11235_CR71","doi-asserted-by":"publisher","first-page":"7","DOI":"10.1016\/S0969-4765(17)30055-3","volume":"2017","author":"F Mather","year":"2017","unstructured":"Mather F (2017) From Scotland Yard to touchless authentication \u2013 fingerprinting makes its mark. Biometric Technology Today 2017(3):7\u20139. https:\/\/doi.org\/10.1016\/S0969-4765(17)30055-3","journal-title":"Biometric Technology Today"},{"key":"11235_CR72","doi-asserted-by":"publisher","unstructured":"Matic M, Stefanovic I, Radosavac U, Vidakovic M (2017) Challenges of integrating smart home automation with cloud based voice recognition systems. In: 2017 IEEE 7Th international conference on consumer electronics - berlin (ICCE-Berlin), IEEE, pp 248\u2013249. https:\/\/doi.org\/10.1109\/ICCE-Berlin.2017.8210640","DOI":"10.1109\/ICCE-Berlin.2017.8210640"},{"key":"11235_CR73","unstructured":"Mayhew S (2015) History of Biometrics. https:\/\/www.biometricupdate.com\/201802\/history-of-biometrics-2"},{"issue":"11","key":"11235_CR74","doi-asserted-by":"publisher","first-page":"1875","DOI":"10.1162\/jocn_a_00427","volume":"25","author":"C McGettigan","year":"2013","unstructured":"McGettigan C, Eisner F, Agnew ZK, Manly T, Wisbey D, Scott SK (2013) T\u2019ain\u2019t What You Say, It\u2019s the Way That You Say It \u2014Left Insula and Inferior Frontal Cortex Work in Interaction with Superior Temporal Regions to Control the Performance of Vocal Impersonations. Journal of Cognitive Neuroscience 25(11):1875\u20131886. 1511.04103","journal-title":"Journal of Cognitive Neuroscience"},{"issue":"November 2018","key":"11235_CR75","doi-asserted-by":"publisher","first-page":"103311","DOI":"10.1016\/j.jbi.2019.103311","volume":"100","author":"N Mehta","year":"2019","unstructured":"Mehta N, Pandit A, Shukla S (2019) Transforming healthcare with big data analytics and artificial intelligence: a systematic mapping study. J Biomed Inform 100(November 2018):103311. https:\/\/doi.org\/10.1016\/j.jbi.2019.103311","journal-title":"J Biomed Inform"},{"key":"11235_CR76","doi-asserted-by":"publisher","unstructured":"Mekonnen BW, Derebssa Dufera B (2015) Noise robust speaker verification using GMM-UBM multi-condition training. In: IEEE AFRICON Conference, IEEE, pp 1\u20135. https:\/\/doi.org\/10.1109\/AFRCON.2015.7331916","DOI":"10.1109\/AFRCON.2015.7331916"},{"key":"11235_CR77","doi-asserted-by":"publisher","unstructured":"Mishra J, Singh M, Pati D (2018) Processing linear prediction residual signal to counter replay attacks. In: 2018 International conference on signal processing and communications (SPCOM), IEEE, pp 95\u201399. https:\/\/doi.org\/10.1109\/SPCOM.2018.8724390","DOI":"10.1109\/SPCOM.2018.8724390"},{"key":"11235_CR78","doi-asserted-by":"publisher","unstructured":"Monteiro J, Alam J, Falk TH (2020) An ensemble based approach for generalized detection of spoofing attacks to automatic speaker recognizers. In: ICASSP 2020 - 2020 IEEE International conference on acoustics, speech and signal processing (ICASSP), IEEE, pp 6599\u20136603. https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9054558","DOI":"10.1109\/ICASSP40776.2020.9054558"},{"key":"11235_CR79","doi-asserted-by":"publisher","first-page":"101096","DOI":"10.1016\/j.csl.2020.101096","volume":"63","author":"J Monteiro","year":"2020","unstructured":"Monteiro J, Alam J, Falk TH (2020) Generalized end-to-end detection of spoofing attacks to automatic speaker recognizers. Computer Speech & Language 63:101096. https:\/\/doi.org\/10.1016\/j.csl.2020.101096","journal-title":"Computer Speech & Language"},{"key":"11235_CR80","doi-asserted-by":"publisher","unstructured":"Muckenhirn H, Magimai-Doss M, Marcel S (2018) End-to-End convolutional neural network-based voice presentation attack detection. In: IEEE International Joint Conference on Biometrics, IJCB 2017, vol 2018-Janua, pp 335\u2013341. https:\/\/doi.org\/10.1109\/BTAS.2017.8272715","DOI":"10.1109\/BTAS.2017.8272715"},{"issue":"4","key":"11235_CR81","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1109\/MCOM.2018.1700790","volume":"56","author":"G Muhammad","year":"2018","unstructured":"Muhammad G, Alhamid MF, Alsulaiman M, Gupta B (2018) Edge computing with cloud for voice disorder assessment and treatment. IEEE Commun Mag 56(4):60\u201365. https:\/\/doi.org\/10.1109\/MCOM.2018.1700790","journal-title":"IEEE Commun Mag"},{"key":"11235_CR82","doi-asserted-by":"publisher","unstructured":"Nagarsheth P, Khoury E, Patil K, Garland M (2017) Replay attack detection using DNN for channel discrimination. In: Interspeech 2017, ISCA, ISCA, pp 97\u2013101. https:\/\/doi.org\/10.21437\/Interspeech.2017-1377","DOI":"10.21437\/Interspeech.2017-1377"},{"key":"11235_CR83","doi-asserted-by":"publisher","unstructured":"Neelima M, Santiprabha I (2020) Mimicry voice detection using convolutional neural networks. In: 2020 International conference on smart electronics and communication (ICOSEC), IEEE, pp 314\u2013318. https:\/\/doi.org\/10.1109\/ICOSEC49089.2020.9215407","DOI":"10.1109\/ICOSEC49089.2020.9215407"},{"key":"11235_CR84","doi-asserted-by":"publisher","first-page":"214","DOI":"10.1016\/j.asoc.2015.01.036","volume":"30","author":"M Pal","year":"2015","unstructured":"Pal M, Saha G (2015) On robustness of speech based biometric systems against voice conversion attack. Appl Soft Comput J 30:214\u2013228. https:\/\/doi.org\/10.1016\/j.asoc.2015.01.036","journal-title":"Appl Soft Comput J"},{"key":"11235_CR85","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1016\/j.csl.2017.10.001","volume":"48","author":"M Pal","year":"2018","unstructured":"Pal M, Paul D, Saha G (2018) Synthetic speech detection using fundamental frequency variation and spectral features. Computer Speech and Language 48:31\u201350. https:\/\/doi.org\/10.1016\/j.csl.2017.10.001","journal-title":"Computer Speech and Language"},{"key":"11235_CR86","doi-asserted-by":"publisher","unstructured":"Parasu P, Epps J, Sriskandaraja K, Suthokumar G (2020) Investigating Light-ResNet architecture for spoofing detection under mismatched conditions. In: Interspeech 2020, ISCA, ISCA, pp 1111\u20131115. https:\/\/doi.org\/10.21437\/Interspeech.2020-2039","DOI":"10.21437\/Interspeech.2020-2039"},{"key":"11235_CR87","doi-asserted-by":"crossref","unstructured":"Patel TB, Patil HA (2015) Combining evidences from mel cepstral, cochlear filter cepstral and instantaneous frequency features for detection of natural vs. spoofed speech . In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, ISCA, pp 2062\u20132066","DOI":"10.21437\/Interspeech.2015-467"},{"key":"11235_CR88","doi-asserted-by":"publisher","unstructured":"Patil HA, Kamble MR (2018) A survey on replay attack detection for automatic speaker verification (ASV) system. In: 2018 Asia-pacific signal and information processing association annual summit and conference, APSIPA ASC, IEEE, pp 1047\u20131053. https:\/\/doi.org\/10.23919\/APSIPA.2018.8659666","DOI":"10.23919\/APSIPA.2018.8659666"},{"issue":"4","key":"11235_CR89","doi-asserted-by":"publisher","first-page":"605","DOI":"10.1109\/JSTSP.2017.2684705","volume":"11","author":"D Paul","year":"2017","unstructured":"Paul D, Pal M, Saha G (2017) Spectral features for synthetic speech detection. IEEE J Select Top Signal Process 11(4):605\u2013617. https:\/\/doi.org\/10.1109\/JSTSP.2017.2684705","journal-title":"IEEE J Select Top Signal Process"},{"key":"11235_CR90","doi-asserted-by":"crossref","unstructured":"Paull D, Saha G (2017) Generalization of Spoofing Countermeasures: A Case Study with ASVspoof 2015 And BTAS 2016 Corpora. In: ICASSP2017, pp 2047\u20132051","DOI":"10.1109\/ICASSP.2017.7952516"},{"key":"11235_CR91","doi-asserted-by":"publisher","first-page":"109","DOI":"10.1016\/j.cviu.2016.03.013","volume":"150","author":"X Peng","year":"2015","unstructured":"Peng X, Wang L, Wang X, Qiao Y (2015) Bag of visual words and fusion methods for action recognition: Comprehensive study and good practice. Comput Vis Image Underst 150:109\u2013125. https:\/\/doi.org\/10.1016\/j.cviu.2016.03.013","journal-title":"Comput Vis Image Underst"},{"key":"11235_CR92","doi-asserted-by":"publisher","unstructured":"Prajapati GP, Kamble MR, Patil HA (2021) Energy separation based features for replay spoof detection for voice assistant. In: 2020 28Th european signal processing conference (EUSIPCO), IEEE, pp 386\u2013390. https:\/\/doi.org\/10.23919\/Eusipco47968.2020.9287577","DOI":"10.23919\/Eusipco47968.2020.9287577"},{"key":"11235_CR93","doi-asserted-by":"publisher","unstructured":"Rahmeni R, Aicha AB, Ayed YB (2020) Speech spoofing detection using SVM and ELM technique with acoustic features. In: 2020 5Th international conference on advanced technologies for signal and image processing (ATSIP), IEEE, pp 1\u20134. https:\/\/doi.org\/10.1109\/ATSIP49331.2020.9231799","DOI":"10.1109\/ATSIP49331.2020.9231799"},{"issue":"4","key":"11235_CR94","first-page":"709","volume":"3","author":"JB Ramgire","year":"2016","unstructured":"Ramgire JB, Jagdale PSM (2016) A survey on speaker recognition with various feature extraction and classification techniques. Int Res J Eng Technol (IRJET) 3(4):709\u2013712","journal-title":"Int Res J Eng Technol (IRJET)"},{"issue":"1-2","key":"11235_CR95","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1016\/0167-6393(95)00009-D","volume":"17","author":"DA Reynolds","year":"1995","unstructured":"Reynolds DA (1995) Speaker identification and verification using Gaussian mixture speaker models. Speech Comm 17(1-2):91\u2013108. https:\/\/doi.org\/10.1016\/0167-6393(95)00009-D","journal-title":"Speech Comm"},{"key":"11235_CR96","doi-asserted-by":"crossref","unstructured":"Reynolds DA (2009) Gaussian mixture models. In: Encyclopedia of biometrics. Springer, Boston, pp 659\u2013663","DOI":"10.1007\/978-0-387-73003-5_196"},{"issue":"1","key":"11235_CR97","doi-asserted-by":"publisher","first-page":"72","DOI":"10.1109\/89.365379","volume":"3","author":"DA Reynolds","year":"1995","unstructured":"Reynolds DA, Rose R (1995) Robust text-independent speaker identification using Gaussian mixture speaker models. IEEE Trans Speech Audio Process 3 (1):72\u201383. https:\/\/doi.org\/10.1109\/89.365379","journal-title":"IEEE Trans Speech Audio Process"},{"key":"11235_CR98","unstructured":"Ross A, Jain AK, Nandakumar K (2006) Score level fusion. In: Handbook of multibiometrics. Kluwer Academic Publishers, Boston, pp 91\u2013142"},{"issue":"2","key":"11235_CR99","doi-asserted-by":"publisher","first-page":"872","DOI":"10.1007\/s00034-020-01501-y","volume":"40","author":"S Rupesh Kumar","year":"2021","unstructured":"Rupesh Kumar S, Bharathi B (2021) A novel approach towards generalization of countermeasure for spoofing attack on ASV systems. Circuits, Systems, and Signal Processing 40(2):872\u2013889. https:\/\/doi.org\/10.1007\/s00034-020-01501-y","journal-title":"Circuits, Systems, and Signal Processing"},{"issue":"5","key":"11235_CR100","first-page":"2276","volume":"13","author":"T Sabhanayagam","year":"2018","unstructured":"Sabhanayagam T, Prasanna Venkatesan V, Senthamaraikannan K (2018) A comprehensive survey on various biometric systems. Int J Appl Eng Res 13(5):2276\u20132297","journal-title":"Int J Appl Eng Res"},{"key":"11235_CR101","doi-asserted-by":"publisher","unstructured":"Safavi S, Gan H, Mporas I (2017) Improving speaker verification performance under spoofing attacks by fusion of different operational modes. In: Proceedings - 2017 IEEE 13th International Colloquium on Signal Processing and its Applications, CSPA 2017, pp 219\u2013223. https:\/\/doi.org\/10.1109\/CSPA.2017.8064954","DOI":"10.1109\/CSPA.2017.8064954"},{"key":"11235_CR102","doi-asserted-by":"crossref","unstructured":"Sahidullah M, Kinnunen T, Hanil\u00e7i C (2015) A comparison of features for synthetic speech detection. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 2087\u20132091","DOI":"10.21437\/Interspeech.2015-472"},{"key":"11235_CR103","doi-asserted-by":"crossref","unstructured":"Sahidullah M, Delgado H, Todisco M, Kinnunen T, Evans N, Yamagishi J, Lee KA (2019) Introduction to voice presentation attack detection and recent advances. Springer International Publishing, pp 321\u2013361","DOI":"10.1007\/978-3-319-92627-8_15"},{"key":"11235_CR104","doi-asserted-by":"publisher","unstructured":"Sailor HB, Kamble MR, Patil HA (2017) Unsupervised representation learning using convolutional restricted boltzmann machine for spoof speech detection. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 2601\u20132605. https:\/\/doi.org\/10.21437\/Interspeech.2017-1393","DOI":"10.21437\/Interspeech.2017-1393"},{"issue":"4","key":"11235_CR105","doi-asserted-by":"publisher","first-page":"810","DOI":"10.1109\/TIFS.2015.2398812","volume":"10","author":"J Sanchez","year":"2015","unstructured":"Sanchez J, Saratxaga I, Hernaez I, Navas E, Erro D, Raitio T (2015) Toward a universal synthetic speech spoofing detection using phase information. IEEE Transactions on Information Forensics and Security 10(4):810\u2013820. https:\/\/doi.org\/10.1109\/TIFS.2015.2398812","journal-title":"IEEE Transactions on Information Forensics and Security"},{"key":"11235_CR106","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1016\/j.specom.2016.04.001","volume":"81","author":"I Saratxaga","year":"2016","unstructured":"Saratxaga I, Sanchez J, Wu Z, Hernaez I, Navas E (2016) Synthetic speech detection using phase information. Speech Comm 81:31\u201341. https:\/\/doi.org\/10.1016\/j.specom.2016.04.001","journal-title":"Speech Comm"},{"key":"11235_CR107","doi-asserted-by":"publisher","unstructured":"Sarkar AK, Tan ZH (2016) Text dependent speaker verification using un-supervised HMM-UBM and Temporal GMM-UBM. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, vol 08-12-Sept, pp 425\u2013429. https:\/\/doi.org\/10.21437\/Interspeech.2016-362","DOI":"10.21437\/Interspeech.2016-362"},{"key":"11235_CR108","doi-asserted-by":"publisher","unstructured":"Sarria-Paja M, Senoussaoui M, O\u2019Shaughnessy D, Falk TH (2016) Feature mapping, score-, and feature-level fusion for improved normal and whispered speech speaker verification. In: 2016 IEEE International conference on acoustics, speech and signal processing (ICASSP), IEEE, pp 5480\u20135484. https:\/\/doi.org\/10.1109\/ICASSP.2016.7472725","DOI":"10.1109\/ICASSP.2016.7472725"},{"issue":"5","key":"11235_CR109","first-page":"1581","volume":"2","author":"V Sharma","year":"2013","unstructured":"Sharma V, Bansal PK (2013) A review on speaker recognition approaches and challenges. Int J Eng Res Technol 2(5):1581\u20131588","journal-title":"Int J Eng Res Technol"},{"key":"11235_CR110","doi-asserted-by":"publisher","unstructured":"Hj Shim, Heo HS, Jw Jung, Yu HJ (2020) Self-Supervised Pre-Training With acoustic configurations for replay spoofing detection. In: Interspeech 2020, ISCA, ISCA, pp 1091\u20131095. https:\/\/doi.org\/10.21437\/Interspeech.2020-1345","DOI":"10.21437\/Interspeech.2020-1345"},{"key":"11235_CR111","unstructured":"Simmons D (2017) BBC fools HSBC voice recognition security system. https:\/\/www.bbc.com\/news\/technology-39965545"},{"issue":"2","key":"11235_CR112","doi-asserted-by":"publisher","first-page":"313","DOI":"10.1007\/s10772-019-09604-x","volume":"22","author":"M Singh","year":"2019","unstructured":"Singh M, Pati D (2019) Combining evidences from Hilbert envelope and residual phase for detecting replay attacks. Int J Speech Technol 22(2):313\u2013326. https:\/\/doi.org\/10.1007\/s10772-019-09604-x","journal-title":"Int J Speech Technol"},{"issue":"4","key":"11235_CR113","doi-asserted-by":"publisher","first-page":"282","DOI":"10.1049\/iet-bmt.2016.0126","volume":"6","author":"R Singh","year":"2017","unstructured":"Singh R, Jim\u00e9nez A (2017) Voice disguise by mimicry: deriving statistical articulometric evidence to evaluate claimed impersonation. IET Biometrics 6(4):282\u2013289. https:\/\/doi.org\/10.1049\/iet-bmt.2016.0126","journal-title":"IET Biometrics"},{"key":"11235_CR114","doi-asserted-by":"publisher","unstructured":"Sinitca AM, Efimchik NV, Shalugin ED, Toropov VA, Simonchik K (2020) Voice antispoofing system vulnerabilities research. In: 2020 IEEE Conference of russian young researchers in electrical and electronic engineering (EIConRus), IEEE, pp 505\u2013508. https:\/\/doi.org\/10.1109\/EIConRus49466.2020.9039393","DOI":"10.1109\/EIConRus49466.2020.9039393"},{"key":"11235_CR115","doi-asserted-by":"publisher","unstructured":"Snyder D, Garcia-Romero D, Povey D, Khudanpur S (2017) Deep Neural Network Embeddings for Text-Independent Speaker Verification. In: Interspeech 2017, ISCA, ISCA, vol 2017-Augus, pp 999\u20131003. https:\/\/doi.org\/10.21437\/Interspeech.2017-620","DOI":"10.21437\/Interspeech.2017-620"},{"key":"11235_CR116","doi-asserted-by":"publisher","unstructured":"Snyder D, Garcia-Romero D, Sell G, Povey D, Khudanpur S (2018) x-vectors: robust DNN embeddings for speaker recognition. In: 2018 IEEE International conference on acoustics, speech and signal processing (ICASSP), IEEE, pp 5329\u20135333. https:\/\/doi.org\/10.1109\/ICASSP.2018.8461375","DOI":"10.1109\/ICASSP.2018.8461375"},{"issue":"3","key":"11235_CR117","doi-asserted-by":"publisher","first-page":"1592","DOI":"10.21817\/ijet\/2017\/v9i3\/170903513","volume":"9","author":"S Sujiya","year":"2017","unstructured":"Sujiya S, Chandra E (2017) A review on speaker recognition. Int J Eng Technol 9(3):1592\u20131598. https:\/\/doi.org\/10.21817\/ijet\/2017\/v9i3\/170903513","journal-title":"Int J Eng Technol"},{"issue":"12","key":"11235_CR118","doi-asserted-by":"publisher","first-page":"2437","DOI":"10.1016\/j.patcog.2004.12.013","volume":"38","author":"QS Sun","year":"2005","unstructured":"Sun QS, Zeng SG, Liu Y, Heng PA, Xia DS (2005) A new method of feature fusion and its application in image recognition. Pattern Recogn 38(12):2437\u20132448. https:\/\/doi.org\/10.1016\/j.patcog.2004.12.013","journal-title":"Pattern Recogn"},{"key":"11235_CR119","doi-asserted-by":"publisher","unstructured":"Suthokumar G, Sriskandaraja K, Sethu V, Wijenayake C, Ambikairajah E (2018) An Investigation about the Scalability of the Spoofing Detection System. In: 2018 IEEE 9th International Conference on Information and Automation for Sustainability, ICIAfS 2018, IEEE, pp 1\u20135. https:\/\/doi.org\/10.1109\/ICIAFS.2018.8913369","DOI":"10.1109\/ICIAFS.2018.8913369"},{"key":"11235_CR120","doi-asserted-by":"publisher","unstructured":"Suthokumar G, Sethu V, Sriskandaraja K, Ambikairajah E (2020) Adversarial Multi-Task learning for speaker normalization in replay detection. In: ICASSP 2020 - 2020 IEEE International conference on acoustics, speech and signal processing (ICASSP), IEEE, pp 6609\u20136613. https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9054322","DOI":"10.1109\/ICASSP40776.2020.9054322"},{"key":"11235_CR121","doi-asserted-by":"publisher","unstructured":"Tieran Z, Jiqing H, Guibin Z (2018) Deep neural network based discriminative training for i-vector\/PLDA speaker verification. In: 2018 IEEE International conference on acoustics, speech and signal processing (ICASSP), IEEE, pp 5354\u20135358. https:\/\/doi.org\/10.1109\/ICASSP.2018.8461344","DOI":"10.1109\/ICASSP.2018.8461344"},{"key":"11235_CR122","doi-asserted-by":"publisher","unstructured":"Todisco M, Delgado H, Evans N (2016) A new feature for automatic speaker verification anti-spoofing: Constant Q Cepstral coefficients. In: Odyssey 2016, pp 283\u2013290. https:\/\/doi.org\/10.21437\/odyssey.2016-41","DOI":"10.21437\/odyssey.2016-41"},{"issue":"September 2017","key":"11235_CR123","doi-asserted-by":"publisher","first-page":"516","DOI":"10.1016\/j.csl.2017.01.001","volume":"45","author":"M Todisco","year":"2017","unstructured":"Todisco M, Delgado H, Evans N (2017) Constant Q cepstral coefficients: A spoofing countermeasure for automatic speaker verification. Computer Speech and Language 45(September 2017):516\u2013535. https:\/\/doi.org\/10.1016\/j.csl.2017.01.001","journal-title":"Computer Speech and Language"},{"key":"11235_CR124","doi-asserted-by":"crossref","unstructured":"Todisco M, Wang X, Vestman V, Nautsch A, Yamagishi J, Evans N, Kinnunen T, Lee KA (2019) ASVspoof 2019 : Future horizons in spoofed and fake audio detection. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH 2019, pp 3\u20137. arXiv:1904.05441v2","DOI":"10.21437\/Interspeech.2019-2249"},{"key":"11235_CR125","doi-asserted-by":"publisher","unstructured":"Tsai WH, Lin JC, Ma CH, Liao YF (2016) Speaker identification for personalized smart TVs. In: 2016 IEEE International Conference on Consumer Electronics-Taiwan, ICCE-TW 2016, IEEE, pp 1\u20132. https:\/\/doi.org\/10.1109\/ICCE-TW.2016.7521051","DOI":"10.1109\/ICCE-TW.2016.7521051"},{"key":"11235_CR126","doi-asserted-by":"publisher","unstructured":"Valin JM, Skoglund J (2019) LPCNET: improving neural speech synthesis through linear prediction. In: ICASSP 2019 - 2019 IEEE International conference on acoustics, speech and signal processing (ICASSP), IEEE, pp 5891\u20135895. https:\/\/doi.org\/10.1109\/ICASSP.2019.8682804","DOI":"10.1109\/ICASSP.2019.8682804"},{"key":"11235_CR127","doi-asserted-by":"crossref","unstructured":"Villalba J, Miguel A, Ortega A, Lleida E (2015) Spoofing detection with DNN and one-class SVM for the ASVspoof 2015 challenge. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 2067\u20132071","DOI":"10.21437\/Interspeech.2015-468"},{"key":"11235_CR128","unstructured":"Vishi K, Mavroeidis V (2018) An evaluation of score level fusion approaches for fingerprint and finger-vein biometrics. arXiv:abs\/1805.1:1--11, 1805.10666"},{"key":"11235_CR129","doi-asserted-by":"publisher","unstructured":"Wang D, Li L, Tang Z, Zheng TF (2018) Deep speaker verification: Do we need end to end? In: Proceedings - 9th Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2017, pp 177\u2013181. https:\/\/doi.org\/10.1109\/APSIPA.2017.8282024, arXiv:1706.07859v1","DOI":"10.1109\/APSIPA.2017.8282024"},{"key":"11235_CR130","doi-asserted-by":"crossref","unstructured":"Wang L, Yoshida Y, Kawakami Y, Nakagawa S (2015) Relative phase information for detecting human speech and spoofed speech. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 2092\u20132096","DOI":"10.21437\/Interspeech.2015-473"},{"key":"11235_CR131","doi-asserted-by":"publisher","unstructured":"Wang X, Yamagishi J, Todisco M, Delgado H, Nautsch A, Evans N, Sahidullah M, Vestman V, Kinnunen T, Lee KA, Juvela L, Alku P, Peng YH, Hwang HT, Tsao Y, Wang HM, Maguer SL, Becker M, Henderson F, Clark R, Zhang Y, Wang Q, Jia Y, Onuma K, Mushika K, Kaneda T, Jiang Y, Liu LJ, Wu YC, Huang WC, Toda T, Tanaka K, Kameoka H, Steiner I, Matrouf D, Bonastre JF, Govender A, Ronanki S, Zhang JX, Ling ZH (2019) ASVspoof 2019: a large-scale public database of synthetic, converted and replayed speech. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH 2019, pp 1\u201324. https:\/\/doi.org\/10.1016\/j.csl.2020.101114, 1911.01601","DOI":"10.1016\/j.csl.2020.101114"},{"key":"11235_CR132","unstructured":"Wang Z, Cui S, Kang X, Sun W, Li Z (2020) Densely connected convolutional network for audio spoofing detection. In: 2020 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), pp 1352\u20131360"},{"key":"11235_CR133","doi-asserted-by":"publisher","unstructured":"Lin W (2015) An improved GMM-based clustering algorithm for efficient speaker identification. in: 2015 4th international conference on computer science and network technology ICCSNT), IEEE, pp 1490\u20131493. https:\/\/doi.org\/10.1109\/ICCSNT.2015.7491011","DOI":"10.1109\/ICCSNT.2015.7491011"},{"key":"11235_CR134","doi-asserted-by":"publisher","unstructured":"Wijethunga R, Matheesha D, Noman AA, De Silva K, Tissera M, Rupasinghe L (2020) Deepfake audio detection: a deep learning based solution for group conversations. In: 2020 2Nd international conference on advancements in computing (ICAC), IEEE, pp 192\u2013197. https:\/\/doi.org\/10.1109\/ICAC51239.2020.9357161","DOI":"10.1109\/ICAC51239.2020.9357161"},{"key":"11235_CR135","doi-asserted-by":"publisher","unstructured":"Wu Z, Li H (2016) On the study of replay and voice conversion attacks to text-dependent speaker verification. Multimedia Tools and Applications. pp 5311\u20135327. https:\/\/doi.org\/10.1007\/s11042-015-3080-9","DOI":"10.1007\/s11042-015-3080-9"},{"key":"11235_CR136","doi-asserted-by":"publisher","first-page":"130","DOI":"10.1016\/j.specom.2014.10.005","volume":"66","author":"Z Wu","year":"2015","unstructured":"Wu Z, Evans N, Kinnunen T, Yamagishi J, Alegre F, Li H (2015) Spoofing and countermeasures for speaker verification: a survey. Speech Comm 66:130\u2013153. https:\/\/doi.org\/10.1016\/j.specom.2014.10.005","journal-title":"Speech Comm"},{"key":"11235_CR137","doi-asserted-by":"crossref","unstructured":"Wu Z, Kinnunen T, Evans N, Yamagishi J, Hanil\u00e7i C, Sahidullah M, Sizov A (2015) ASVSpoof 2015: Automatic speaker verification spoofing and countermeasures challenge evaluation plan. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH 2015, pp 2037\u20132041","DOI":"10.21437\/Interspeech.2015-462"},{"issue":"4","key":"11235_CR138","doi-asserted-by":"publisher","first-page":"588","DOI":"10.1109\/JSTSP.2017.2671435","volume":"11","author":"Z Wu","year":"2017","unstructured":"Wu Z, Yamagishi J, Kinnunen T, Hanil\u00e7i C, Sahidullah M, Sizov A, Evans N, Todisco M, Delgado H (2017) ASVspoof 2015: The first automatic speaker verification spoofing and countermeasures challenge. IEEE J Select Top Signal Process 11(4):588\u2013604. https:\/\/doi.org\/10.1109\/JSTSP.2017.2671435","journal-title":"IEEE J Select Top Signal Process"},{"issue":"4","key":"11235_CR139","doi-asserted-by":"publisher","first-page":"588","DOI":"10.1109\/JSTSP.2017.2671435","volume":"11","author":"Z Wu","year":"2017","unstructured":"Wu Z, Yamagishi J, Kinnunen T, Hanil\u00e7i C, Sahidullah M, Sizov A, Evans N, Todisco M, Delgado H (2017) ASVSpoof: The automatic speaker verification spoofing and countermeasures challenge. IEEE J Select Top Signal Process 11(4):588\u2013604. https:\/\/doi.org\/10.1109\/JSTSP.2017.2671435","journal-title":"IEEE J Select Top Signal Process"},{"key":"11235_CR140","doi-asserted-by":"publisher","unstructured":"Wu Z, Das RK, Yang J, Li H (2020) Light convolutional neural network with feature genuinization for detection of synthetic speech attacks. In: Interspeech 2020, ISCA, ISCA, pp 1101\u20131105. https:\/\/doi.org\/10.21437\/Interspeech.2020-1810","DOI":"10.21437\/Interspeech.2020-1810"},{"issue":"6","key":"11235_CR141","doi-asserted-by":"publisher","first-page":"1369","DOI":"10.1016\/S0031-3203(02)00262-5","volume":"36","author":"J Yang","year":"2003","unstructured":"Yang J, Yang JY, Zhang D, Lu JF (2003) Feature fusion: Parallel strategy vs. serial strategy. Pattern Recogn 36(6):1369\u20131381. https:\/\/doi.org\/10.1016\/S0031-3203(02)00262-5","journal-title":"Pattern Recogn"},{"key":"11235_CR142","doi-asserted-by":"publisher","unstructured":"Yang J, Das RK, Li H (2019) Extended Constant-Q cepstral coefficients for detection of spoofing attacks. In: 2018 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2018 - Proceedings, APSIPA organization, pp 1024\u20131029. https:\/\/doi.org\/10.23919\/APSIPA.2018.8659537","DOI":"10.23919\/APSIPA.2018.8659537"},{"key":"11235_CR143","doi-asserted-by":"publisher","unstructured":"Ye Y, Lao L, Yan D, Lin L (2019) Detection of replay attack based on normalized constant q cepstral feature. In: 2019 IEEE 4Th international conference on cloud computing and big data analysis (ICCCBDA), IEEE, pp 407\u2013411. https:\/\/doi.org\/10.1109\/ICCCBDA.2019.8725688","DOI":"10.1109\/ICCCBDA.2019.8725688"},{"key":"11235_CR144","doi-asserted-by":"publisher","unstructured":"Zeinali H, Stafylakis T, Athanasopoulou G, Rohdin J, Gkinis I, Burget L, \u011arnock\u00fd J (2019) Detecting spoofing attacks using VGG and SINCNET: But-omilia submission to AsvSpoof 2019 challenge. In: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, pp 1073\u20131077. https:\/\/doi.org\/10.21437\/Interspeech.2019-2892, 1907.12908","DOI":"10.21437\/Interspeech.2019-2892"},{"key":"11235_CR145","doi-asserted-by":"publisher","unstructured":"Zhang C, Cheng J, Gu Y, Wang H, Ma J, Wang S, Xiao J (2020) Improving replay detection system with channel consistency DenseNeXt for the ASVspoof 2019 challenge. In: Interspeech 2020, ISCA, ISCA, pp 4596\u20134600. https:\/\/doi.org\/10.21437\/Interspeech.2020-1044","DOI":"10.21437\/Interspeech.2020-1044"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-021-11235-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-021-11235-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-021-11235-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,10,10]],"date-time":"2021-10-10T05:08:03Z","timestamp":1633842483000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-021-11235-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,4]]},"references-count":145,"journal-issue":{"issue":"21-23","published-print":{"date-parts":[[2021,9]]}},"alternative-id":["11235"],"URL":"https:\/\/doi.org\/10.1007\/s11042-021-11235-x","relation":{},"ISSN":["1380-7501","1573-7721"],"issn-type":[{"value":"1380-7501","type":"print"},{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,8,4]]},"assertion":[{"value":"25 June 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 July 2021","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 July 2021","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 August 2021","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"<!--Emphasis Type='Bold' removed-->Conflict of Interests"}}]}}