{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:42:06Z","timestamp":1760240526618,"version":"build-2065373602"},"reference-count":35,"publisher":"MDPI AG","issue":"14","license":[{"start":{"date-parts":[[2019,7,11]],"date-time":"2019-07-11T00:00:00Z","timestamp":1562803200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100003052","name":"Ministry of Trade, Industry and Energy","doi-asserted-by":"publisher","award":["10076583"],"award-info":[{"award-number":["10076583"]}],"id":[{"id":"10.13039\/501100003052","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","award":["NRF-2016R1C1B1015291"],"award-info":[{"award-number":["NRF-2016R1C1B1015291"]}],"id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Two main spatial cues that can be exploited for dual microphone voice activity detection (VAD) are the interchannel time difference (ITD) and the interchannel level difference (ILD). While both ITD and ILD provide information on the location of audio sources, they may be impaired in different manners by background noises and reverberation and therefore can have complementary information. Conventional approaches utilize the statistics from all frequencies with fixed weight, although the information from some time\u2013frequency bins may degrade the performance of VAD. In this letter, we propose a dual microphone VAD scheme based on the spatial cues in reliable frequency bins only, considering the sparsity of the speech signal in the time\u2013frequency domain. The reliability of each time\u2013frequency bin is determined by three conditions on signal energy, ILD, and ITD. ITD-based and ILD-based VADs and statistics are evaluated using the information from selected frequency bins and then combined to produce the final VAD results. Experimental results show that the proposed frequency selective approach enhances the performances of VAD in realistic environments.<\/jats:p>","DOI":"10.3390\/s19143056","type":"journal-article","created":{"date-parts":[[2019,7,11]],"date-time":"2019-07-11T11:28:28Z","timestamp":1562844508000},"page":"3056","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Dual Microphone Voice Activity Detection Based on Reliable Spatial Cues"],"prefix":"10.3390","volume":"19","author":[{"given":"Soojoong","family":"Hwang","sequence":"first","affiliation":[{"name":"School of Electrical Engineering and Computer Science, Gwangju Institute of Science and Technology, 123 Cheomdan-gwagiro, Buk-gu, Gwangju 61005, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu Gwang","family":"Jin","sequence":"additional","affiliation":[{"name":"AI Technology Unit, SK Telecom, 100 Eulji-ro, Jung-gu, Seoul 04551, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jong Won","family":"Shin","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering and Computer Science, Gwangju Institute of Science and Technology, 123 Cheomdan-gwagiro, Buk-gu, Gwangju 61005, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,7,11]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"6","DOI":"10.1109\/LSP.2015.2495102","article-title":"Speech Enhancement with Nonstationary Acoustic Noise Detection in Time Domain","volume":"23","author":"Tavares","year":"2016","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1601","DOI":"10.1109\/LSP.2017.2750979","article-title":"An Individualized Super-Gaussian Single Microphone Speech Enhancement for Hearing Aid Users With Smartphone as an Assistive Device","volume":"24","author":"Reddy","year":"2017","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_3","unstructured":"Meyer, J., Simmer, K.U., and Kammeyer, K.D. (1997, January 3). Comparison of one- and two-channel noise-estimation techniques. Proceedings of the 5th International Workshop on Acoustic Echo Control Noise Reduction, London, UK."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1133","DOI":"10.1109\/LSP.2017.2712646","article-title":"Robust Pitch Extraction Method for the HMM-Based Speech Synthesis System","volume":"24","author":"Reddy","year":"2017","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1745","DOI":"10.1109\/LSP.2018.2874155","article-title":"Traditional Machine Learning for Pitch Detection","volume":"25","author":"Drugman","year":"2018","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_6","unstructured":"(2019, July 11). TIA Document, PN-3292, Enhanced Variable Rate Codec, Speech Service Option 3 for Wide-Band Spectrum Digital Systems. Available online: https:\/\/www.3gpp2.org\/Public_html\/Specs\/C.S0014-A_v1.0_040426.pdf."},{"key":"ref_7","unstructured":"3GPP TS 26.104 (2014). ANSI-C Code for the Floating-Point Adaptive Multi-Rate (AMR) Speech Codec, 3GPP. Rev. 12.0.0."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1295","DOI":"10.1016\/j.patrec.2006.11.015","article-title":"Voice activity detection based on a family of parametric distributions","volume":"28","author":"Shin","year":"2007","journal-title":"Pattern Recognit. Lett."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1109\/LSP.2008.917027","article-title":"Voice activity detection based on conditional MAP criterion","volume":"15","author":"Shin","year":"2008","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1561","DOI":"10.1049\/el:20047090","article-title":"Voice activity detector employing generalized Gaussian distribution","volume":"40","author":"Chang","year":"2004","journal-title":"Electron. Lett."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"515","DOI":"10.1016\/j.csl.2009.02.003","article-title":"Voice activity detection based on statistical models and machine learning approaches","volume":"24","author":"Shin","year":"2010","journal-title":"Comput. Speech Lang."},{"key":"ref_12","unstructured":"Rabiner, L.R., and Sambur, M.R. (1977, January 9\u201311). Voiced-unvoiced-slience detection using Itakura LPC distance measure. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, Hartford, CT, USA."},{"key":"ref_13","unstructured":"Hoyt, J.D., and Wechsler, H. (1994, January 19\u201322). Detection of human speech in structured noise. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Adelaide, SA, Australia."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Junqua, J.C., Reaves, B., and Mark, B. (1991, January 24\u201326). A study of endpoint detection algorithms in adverse conditions: Incidence on a DTW and HMM recognize. Proceedings of the EUROSPEECH \u201991, Genova, Italy.","DOI":"10.21437\/Eurospeech.1991-313"},{"key":"ref_15","unstructured":"Haigh, J.A., and Mason, J.S. (1993, January 19\u201321). Robust voice activity detection using cepstral feature. Proceedings of the TENCON\u201993, Beijing, China."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"252","DOI":"10.1109\/LSP.2015.2495219","article-title":"Voice Activity Detection: Merging Source and Filter-based Information","volume":"23","author":"Drugman","year":"2016","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"271","DOI":"10.1016\/j.specom.2003.10.002","article-title":"Efficient voice activity detection algorithms using long-term speech information","volume":"42","author":"Segura","year":"2004","journal-title":"Speech Commun."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1119","DOI":"10.1109\/TSA.2005.853212","article-title":"An effective subband OSF-based VAD with noise reduction for robust speech recognition","volume":"13","author":"Segura","year":"2005","journal-title":"IEEE Trans. Speech Audio Process."},{"key":"ref_19","first-page":"288","article-title":"Performance analysis of voice activity detection algorithms for robust speech recognition","volume":"2","author":"Babu","year":"2011","journal-title":"TECHNIA Int. J. Comput. Sci. Commun. Technol."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s13634-015-0277-z","article-title":"Features for voice activity detection: A comparative analysis","volume":"2015","author":"Graf","year":"2015","journal-title":"EURASIP J. Adv. Signal Process."},{"key":"ref_21","unstructured":"Pencak, J., and Nelson, D. (1995, January 9\u201312). The NP speech activity detection algorithm. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, Detroit, MI, USA."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"697","DOI":"10.1109\/TASL.2012.2229986","article-title":"Deep belief network based voice activity detection","volume":"21","author":"Zhang","year":"2013","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"252","DOI":"10.1109\/TASLP.2015.2505415","article-title":"Boosting Contextual Information for Deep Neural Network Based Voice Activity Detection","volume":"24","author":"Zhang","year":"2016","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zazo, R., Sainath, T.N., Simko, G., and Parada, C. (2016). Feature Learning with Raw-Waveform CLDNNs for Voice Activity Detection. Proc. Interspeech, 3668\u20133672.","DOI":"10.21437\/Interspeech.2016-268"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1181","DOI":"10.1109\/LSP.2018.2811740","article-title":"Voice Activity Detection Using an Adaptive Context Attention Model","volume":"25","author":"Kim","year":"2018","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1290","DOI":"10.1109\/LSP.2018.2841653","article-title":"Speech Activity Detection in Naturalistic Audio Environments: Fearless Steps Apollo Corpus","volume":"25","author":"Kaushik","year":"2018","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Guo, Y., Li, K., Fu, Q., and Yan, Y. (2012, January 25\u201330). A two microphone based voice activity detection for distant talking speech in wide range of direction of arrival. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Kyoto, Japan.","DOI":"10.1109\/ICASSP.2012.6289018"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Jeub, M., Herglotz, C., Nelke, C., Beaugeant, C., and Vary, P. (2012, January 25\u201330). Noise reduction for dual-microphone mobile phones exploiting power level differences. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Kyoto, Japan.","DOI":"10.1109\/ICASSP.2012.6288223"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1069","DOI":"10.1109\/TASLP.2014.2313917","article-title":"Dual-microphone voice activity detection technique based on two-step power level difference ratio","volume":"22","author":"Choi","year":"2014","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1335","DOI":"10.1109\/LSP.2016.2597360","article-title":"Dual Microphone Voice Activity Detection Exploiting Interchannel Time and Level Difference","volume":"23","author":"Park","year":"2016","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"956","DOI":"10.1109\/LSP.2004.838200","article-title":"Estimation of Speech Presence Probability in the Field of Microphone Array","volume":"11","author":"Potamitis","year":"2004","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Lee, B., and Kalker, T. (2009, January 18\u201321). Multichannel voice activity detection with spherically invariant sparse distributions. Proceedings of the 2009 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, New Paltz, NY, USA.","DOI":"10.1109\/ASPAA.2009.5346523"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"320","DOI":"10.1109\/TASSP.1976.1162830","article-title":"The generalized correlation method for estimation of time delay","volume":"24","author":"Knapp","year":"1976","journal-title":"IEEE Trans. Acoust. Speech Signal Process."},{"key":"ref_34","unstructured":"Bishop, C.M. (2006). Pattern Recognition and Machine Learning, Springer."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"247","DOI":"10.1016\/0167-6393(93)90095-3","article-title":"Assessment for automatic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition system","volume":"12","author":"Varga","year":"1993","journal-title":"Speech Commun."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/14\/3056\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:04:30Z","timestamp":1760187870000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/14\/3056"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,7,11]]},"references-count":35,"journal-issue":{"issue":"14","published-online":{"date-parts":[[2019,7]]}},"alternative-id":["s19143056"],"URL":"https:\/\/doi.org\/10.3390\/s19143056","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2019,7,11]]}}}