{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:56:17Z","timestamp":1760237777440,"version":"build-2065373602"},"reference-count":41,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2020,6,22]],"date-time":"2020-06-22T00:00:00Z","timestamp":1592784000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Ministry of Education","award":["Allied Advanced Intelligent Biomedical Research Center (A2IBRC)"],"award-info":[{"award-number":["Allied Advanced Intelligent Biomedical Research Center (A2IBRC)"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In the context of assisted human, identifying and enhancing non-stationary speech targets speech in various noise environments, such as a cocktail party, is an important issue for real-time speech separation. Previous studies mostly used microphone signal processing to perform target speech separation and analysis, such as feature recognition through a large amount of training data and supervised machine learning. The method was suitable for stationary noise suppression, but relatively limited for non-stationary noise and difficult to meet the real-time processing requirement. In this study, we propose a real-time speech separation method based on an approach that combines an optical camera and a microphone array. The method was divided into two stages. Stage 1 used computer vision technology with the camera to detect and identify interest targets and evaluate source angles and distance. Stage 2 used beamforming technology with microphone array to enhance and separate the target speech sound. The asynchronous update function was utilized to integrate the beamforming control and speech processing to reduce the effect of the processing delay. The experimental results show that the noise reduction in various stationary and non-stationary noise environments were 6.1 dB and 5.2 dB respectively. The response time of speech processing was less than 10ms, which meets the requirements of a real-time system. The proposed method has high potential to be applied in auxiliary listening systems or machine language processing like intelligent personal assistant.<\/jats:p>","DOI":"10.3390\/s20123527","type":"journal-article","created":{"date-parts":[[2020,6,23]],"date-time":"2020-06-23T09:05:33Z","timestamp":1592903133000},"page":"3527","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["A Real-Time Speech Separation Method Based on Camera and Microphone Array Sensors Fusion Approach"],"prefix":"10.3390","volume":"20","author":[{"given":"Ching-Feng","family":"Liu","sequence":"first","affiliation":[{"name":"Department of Electrical Engineering, Southern Taiwan University of Science and Technology, Tainan 71005, Taiwan"},{"name":"Department of Otolaryngology, Head and Neck Surgery, Chi Mei Medical Center, Tainan 71004, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei-Siang","family":"Ciou","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, Southern Taiwan University of Science and Technology, Tainan 71005, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peng-Ting","family":"Chen","sequence":"additional","affiliation":[{"name":"Department of Biomedical Engineering &amp; Medical Device Innovation Center, National Cheng Kung University, Tainan 70105, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6087-4610","authenticated-orcid":false,"given":"Yi-Chun","family":"Du","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, Southern Taiwan University of Science and Technology, Tainan 71005, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,6,22]]},"reference":[{"key":"ref_1","unstructured":"World Health Organization (2019). Deafness and Hearing Loss, World Health Organization."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"63809","DOI":"10.1109\/ACCESS.2019.2916723","article-title":"Design of Novel Field Programmable Gate Array-Based Hearing Aid","volume":"7","author":"Lin","year":"2019","journal-title":"IEEE Access"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"80","DOI":"10.1044\/2016_AJA-16-0033","article-title":"The acceptable noise level and the pure-tone audiogram","volume":"26","author":"Olsen","year":"2017","journal-title":"Am. J. Audiol."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.3109\/14992027.2015.1072773","article-title":"Background sounds and hearing-aid users: A scoping review","volume":"55","author":"Gygi","year":"2016","journal-title":"Int. J. Audiol."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1121\/1.2920958","article-title":"Computational auditory scene analysis: Principles, algorithms and applications","volume":"124","author":"Wang","year":"2008","journal-title":"Acoust. Soc. Am. J."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Kwan, C., Yin, J., Ayhan, B., Chu, S., Liu, X., Puckett, K., and Sityar, I. (2008, January 1\u20138). Speech Separation Algorithms for Multiple Speaker Environments. Proceedings of the 2008 IEEE International Joint Conference on Neural Networks (IEEE World Congress on Computational Intelligence), Hong Kong, China.","DOI":"10.1109\/IJCNN.2008.4634018"},{"key":"ref_7","unstructured":"Johnson, D.H., and Dudgeon, D.E. (1993). Array Signal Processing: Concepts and Techniques, PTR Prentice Hall."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"197","DOI":"10.1007\/s13272-019-00383-4","article-title":"A review of acoustic imaging methods using phased microphone arrays","volume":"10","author":"Sijtsma","year":"2019","journal-title":"CEAS Aeronaut. J."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Tu, Y.H., Du, J., Zhou, N., and Lee, C.H. (2018, January 12\u201315). Online LSTM-Based Iterative Mask Estimation for Multi-Channel Speech Enhancement and ASR. Proceedings of the 2018 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Honolulu, HI, USA.","DOI":"10.23919\/APSIPA.2018.8659564"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"S3","DOI":"10.1080\/14992027.2016.1256504","article-title":"Functionality of hearing aids: State-of-the-art and future model-based solutions","volume":"57","author":"Kollmeier","year":"2018","journal-title":"Int. J. Audiol."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"114837","DOI":"10.1109\/ACCESS.2019.2932315","article-title":"Beamforming Designs Robust to Propagation Model Estimation Errors for Binaural Hearing Aids","volume":"7","author":"Bouchard","year":"2019","journal-title":"IEEE Access"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Sun, Z., Li, Y., Jiang, H., Chen, F., and Wang, Z. (2018, January 17\u201319). A MVDR-MWF Combined Algorithm for Binaural Hearing Aid System. Proceedings of the 2018 IEEE Biomedical Circuits and Systems Conference (BioCAS), Cleveland, OH, USA.","DOI":"10.1109\/BIOCAS.2018.8584798"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1349","DOI":"10.1109\/TASLP.2019.2918400","article-title":"Methods of Extending a Generalised Sidelobe Canceller with External Microphones","volume":"27","author":"Ali","year":"2019","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1900","DOI":"10.1121\/1.4978058","article-title":"Generalized sidelobe canceller beamforming method for ultrasound imaging","volume":"141","author":"Wang","year":"2017","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"515","DOI":"10.1109\/TASLP.2017.2782491","article-title":"Binaural speaker localization integrated into an adaptive beamformer for hearing aids","volume":"26","author":"Zohourian","year":"2017","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"131","DOI":"10.3766\/jaaa.17090","article-title":"An Evaluation of Hearing Aid Beamforming Microphone Arrays in a Noisy Laboratory Setting","volume":"30","author":"Picou","year":"2019","journal-title":"J. Am. Acad. Audiol."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Dillon, H. (2008). Hearing Aids, Hodder Arnold.","DOI":"10.1201\/b15118-293"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"7306902","DOI":"10.1155\/2018\/7306902","article-title":"Fast estimation method of space-time two-dimensional positioning parameters based on Hadamard product","volume":"2018","author":"Li","year":"2018","journal-title":"Int. J. Antennas Propag."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Mojiri, M., and Ghadiri-Modarres, M.A. (2016, January 10\u201312). Composition of Digital All-Pass Lattice Filter and Gradient Adaptive Filter for Amplitude and Delay Estimation of a Sinusoid with Unknown Frequency. Proceedings of the 2016 24th Iranian Conference on Electrical Engineering (ICEE), Shiraz, Iran.","DOI":"10.1109\/IranianCEE.2016.7585383"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Erdogan, H., Hershey, J.R., Watanabe, S., and Le Roux, J. (2015, January 19\u201324). Phase-Sensitive and Recognition-Boosted Speech Separation Using Deep Recurrent Neural Networks. Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, Australia.","DOI":"10.1109\/ICASSP.2015.7178061"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Feng, W., Guan, N., Li, Y., Zhang, X., and Luo, Z. (2017, January 14\u201319). Audio visual speech recognition with multimodal recurrent neural networks. Proceedings of the 2017 International Joint Conference on Neural Networks (IJCNN), Anchorage, AK, USA.","DOI":"10.1109\/IJCNN.2017.7965918"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., and Rubinstein, M. (2018). Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. arXiv.","DOI":"10.1145\/3197517.3201357"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1163","DOI":"10.1109\/JBHI.2018.2836180","article-title":"Development of Novel Hearing Aids by Using Image Recognition Technology","volume":"23","author":"Lin","year":"2018","journal-title":"IEEE J. Biomed. Health Inform."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"601","DOI":"10.1097\/AUD.0b013e3181734ef2","article-title":"Tolerable hearing aid delays. V. Estimation of limits for open canal fittings","volume":"29","author":"Stone","year":"2008","journal-title":"Ear Hear."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1080\/14992027.2017.1367848","article-title":"Tolerable delay for speech production and perception: Effects of hearing ability and experience with hearing aids","volume":"57","author":"Goehring","year":"2018","journal-title":"Int. J. Audiol."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1474","DOI":"10.4067\/S0717-95022012000400033","article-title":"Head circumference in Canadian male adults: Development of a normalized chart","volume":"30","author":"Nguyen","year":"2012","journal-title":"Int. J. Morphol."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Dentamaro, G., Cardellicchio, A., and Guaragnella, C. (2016, January 4\u20138). Real Time Artificial Auditory Systems for Cluttered Environments. Proceedings of the 2016 23rd International Conference on Pattern Recognition (ICPR), Canc\u00fan, Mexico.","DOI":"10.1109\/ICPR.2016.7899968"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"3220","DOI":"10.1109\/TVT.2016.2593697","article-title":"Three-dimensional positioning for LTE systems","volume":"66","author":"Chen","year":"2016","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Eriksson, E., D\u00e1n, G., and Fodor, V. (2014, January 26\u201328). Real-Time Distributed Visual Feature Extraction from Video in Sensor Networks. Proceedings of the 2014 IEEE International Conference on Distributed Computing in Sensor Systems, Marina Del Rey, CA, USA.","DOI":"10.1109\/DCOSS.2014.30"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Eriksson, E., Pacifici, V., and D\u00e1n, G. (July, January 29). Efficient Distribution of Visual Processing Tasks in Multi-Camera Visual Sensor Networks. Proceedings of the 2015 IEEE International Conference on Multimedia Expo Workshops (ICMEW), Turin, Italy.","DOI":"10.1109\/ICMEW.2015.7169817"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"164053","DOI":"10.1155\/2014\/164053","article-title":"Minimum variance distortionless response beamformer with enhanced nulling level control via dynamic mutated artificial immune system","volume":"2014","author":"Kiong","year":"2014","journal-title":"Sci. World J."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"357","DOI":"10.1260\/147547207783359459","article-title":"CLEAN based on spatial source coherence","volume":"6","author":"Sijtsma","year":"2007","journal-title":"Int. J. Aeroacoust."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"3976","DOI":"10.1016\/j.cub.2018.10.042","article-title":"Rapid transformation from auditory to linguistic representations of continuous speech","volume":"28","author":"Brodbeck","year":"2018","journal-title":"Curr. Biol."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Duan, Z., Mysore, G.J., and Smaragdis, P. (2012, January 9\u201313). Speech enhancement by online non-negative spectrogram decomposition in nonstationary noise environments. Proceedings of the Thirteenth Annual Conference of the International Speech Communication Association, Portland, OR, USA.","DOI":"10.21437\/Interspeech.2012-181"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Duan, Z., Mysore, G.J., and Smaragdis, P. (2012). Online PLCA for real-time semi-supervised source separation. International Conference on Latent Variable Analysis and Signal Separation, Springer.","DOI":"10.1007\/978-3-642-28551-6_5"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Xiao, X., Zhao, S., Jones, D.L., Chng, E.S., and Li, H. (2017, January 5\u20139). On time-frequency mask estimation for MVDR beamforming with application in robust speech recognition. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7952756"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Araki, S., Ito, N., Delcroix, M., Ogawa, A., Kinoshita, K., Higuchi, T., and Nakatani, T. (2017, January 1\u20133). Online meeting recognition in noisy environments with time-frequency mask based MVDR beamforming. Proceedings of the 2017 Hands-free Speech Communications and Microphone Arrays (HSCMA), San Francisco, CA, USA.","DOI":"10.1109\/HSCMA.2017.7895553"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Pfeifenberger, L., Z\u00f6hrer, M., and Pernkopf, F. (2017, January 5\u20139). DNN-based speech mask estimation for eigenvector beamforming. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7952119"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Meng, Z., Watanabe, S., Hershey, J.R., and Erdogan, H. (2017, January 5\u20139). Deep long short-term memory adaptive beamforming networks for multichannel robust speech recognition. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7952160"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"771","DOI":"10.1016\/j.chb.2005.12.012","article-title":"Designing for augmented attention: Towards a framework for attentive user interfaces","volume":"22","author":"Vertegaal","year":"2006","journal-title":"Comput. Hum. Behav."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1145\/2601097.2601119","article-title":"The Visual Microphone: Passive Recovery of Sound from Video","volume":"33","author":"Davis","year":"2014","journal-title":"ACM Trans. Graph."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/12\/3527\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:41:06Z","timestamp":1760175666000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/12\/3527"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,6,22]]},"references-count":41,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2020,6]]}},"alternative-id":["s20123527"],"URL":"https:\/\/doi.org\/10.3390\/s20123527","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2020,6,22]]}}}