{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T02:36:04Z","timestamp":1760150164336,"version":"build-2065373602"},"reference-count":50,"publisher":"MDPI AG","issue":"21","license":[{"start":{"date-parts":[[2023,10,27]],"date-time":"2023-10-27T00:00:00Z","timestamp":1698364800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Key Research and Development Projects in Hainan Province","award":["ZDYF2022GXJS001","ZR2020MF011","61971253","ZDKJ202017"],"award-info":[{"award-number":["ZDYF2022GXJS001","ZR2020MF011","61971253","ZDKJ202017"]}]},{"name":"Shandong Provincial Natural Science Foundation","award":["ZDYF2022GXJS001","ZR2020MF011","61971253","ZDKJ202017"],"award-info":[{"award-number":["ZDYF2022GXJS001","ZR2020MF011","61971253","ZDKJ202017"]}]},{"name":"National Natural Science Foundation of China","award":["ZDYF2022GXJS001","ZR2020MF011","61971253","ZDKJ202017"],"award-info":[{"award-number":["ZDYF2022GXJS001","ZR2020MF011","61971253","ZDKJ202017"]}]},{"name":"Hainan Province of China","award":["ZDYF2022GXJS001","ZR2020MF011","61971253","ZDKJ202017"],"award-info":[{"award-number":["ZDYF2022GXJS001","ZR2020MF011","61971253","ZDKJ202017"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The cocktail party problem can be more effectively addressed by leveraging the speaker\u2019s visual and audio information. This paper proposes a method to improve the audio\u2019s separation using two visual cues: facial features and lip movement. Firstly, residual connections are introduced in the audio separation module to extract detailed features. Secondly, considering the video stream contains information other than the face, which has a minimal correlation with the audio, an attention mechanism is employed in the face module to focus on crucial information. Then, the loss function considers the audio-visual similarity to take advantage of the relationship between audio and visual completely. Experimental results on the public VoxCeleb2 dataset show that the proposed model significantly enhanced SDR, PSEQ, and STOI, especially 4 dB improvements in SDR.<\/jats:p>","DOI":"10.3390\/s23218770","type":"journal-article","created":{"date-parts":[[2023,10,27]],"date-time":"2023-10-27T11:50:18Z","timestamp":1698407418000},"page":"8770","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["A Facial Feature and Lip Movement Enhanced Audio-Visual Speech Separation Model"],"prefix":"10.3390","volume":"23","author":[{"given":"Guizhu","family":"Li","sequence":"first","affiliation":[{"name":"College of Electronic Engineering, Ocean University of China, Qingdao 266100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Min","family":"Fu","sequence":"additional","affiliation":[{"name":"College of Electronic Engineering, Ocean University of China, Qingdao 266100, China"},{"name":"Sanya Oceanography Institution, Ocean University of China, Sanya 572024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mengnan","family":"Sun","sequence":"additional","affiliation":[{"name":"College of Electronic Engineering, Ocean University of China, Qingdao 266100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuefeng","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Automation and Electronic Engineering, Qingdao University of Science and Technology, Qingdao 266061, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bing","family":"Zheng","sequence":"additional","affiliation":[{"name":"College of Electronic Engineering, Ocean University of China, Qingdao 266100, China"},{"name":"Sanya Oceanography Institution, Ocean University of China, Sanya 572024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,10,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"975","DOI":"10.1121\/1.1907229","article-title":"Some experiments on the recognition of speech, with one and with two ears","volume":"25","author":"Cherry","year":"1953","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1109\/TNN.2007.913988","article-title":"Computational Auditory Scene Analysis: Principles, Algorithms, and Applications","volume":"19","author":"Wang","year":"2008","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"297","DOI":"10.1006\/csla.1994.1016","article-title":"Computational auditory scene analysis","volume":"8","author":"Brown","year":"1994","journal-title":"Comput. Speech Lang."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TASL.2006.876726","article-title":"Convolutive speech bases and their application to supervised speech separation","volume":"15","author":"Smaragdis","year":"2007","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"45","DOI":"10.1016\/j.csl.2008.11.001","article-title":"Superhuman multi-talker speech recognition: A graphical modeling approach","volume":"24","author":"Hershey","year":"2010","journal-title":"Comput. Speech Lang."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Hershey, J.R., Chen, Z., Le Roux, J., and Watanabe, S. (2016, January 20\u201325). Deep clustering: Discriminative embeddings for segmentation and separation. Proceedings of the 41th International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7471631"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Chen, Z., Luo, Y., and Mesgarani, N. (2017). Deep Attractor Network for Single-microphone Speech Separation. arXiv.","DOI":"10.1109\/ICASSP.2017.7952155"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Luo, Y., Chen, Z., and Mesgarani, N. (2018). Speaker-Independent Speech Separation with Deep Attractor Network. arXiv.","DOI":"10.1109\/TASLP.2018.2795749"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Han, C., Luo, Y., and Mesgarani, N. (2019, January 12\u201317). Online Deep Attractor Network for Real-time Single-channel Speech Separation. Proceedings of the ICASSP 2019\u20142019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8682884"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Isik, Y., Roux, J.L., Chen, Z., Watanabe, S., and Hershey, J.R. (2016, January 8\u201312). Single-Channel Multi-Speaker Separation Using Deep Clustering. Proceedings of the Interspeech, San Francisco, CA, USA.","DOI":"10.21437\/Interspeech.2016-1176"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"2136","DOI":"10.1109\/TASLP.2015.2468583","article-title":"Joint optimization of masks and deep recurrent neural networks for monaural source separation","volume":"23","author":"Huang","year":"2015","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_12","unstructured":"Shi, x., Chen, Z., Wang, H., Yeung, D.Y., Wong, W., and Woo, W. (2015). Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting. arXiv."},{"key":"ref_13","unstructured":"Stoller, D., Ewert, S., and Dixon, S. (2018). Wave-U-Net: A Multi-Scale Neural Network for End-to-End Audio Source Separation. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"4705","DOI":"10.1121\/1.4986931","article-title":"Long short-term memory for speaker generalization in supervised speech separation","volume":"141","author":"Chen","year":"2017","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Luo, Y., and Mesgarani, N. (2017). Tasnet: Time-domain audio separation network for real-time, single-channel speech separation. arXiv.","DOI":"10.1109\/ICASSP.2018.8462116"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Luo, Y., and Mesgarani, N. (2018). Conv-TasNet Surpassing Ideal TimeFrequency Magnitude_Masking for Speech Separation. arXiv.","DOI":"10.1109\/TASLP.2019.2915167"},{"key":"ref_17","unstructured":"Arango-S\u00e1nchez, J.A., and Arias-Londo\u00f1o, J.D. (2022). An enhanced Conv-TasNet model for speech separation using a speaker distancebased loss function. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Chung, J.S., Senior, A., Vinyals, O., and Zisserman, A. (2016). Lip reading sentences in the wild. arXiv.","DOI":"10.1109\/CVPR.2017.367"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1272","DOI":"10.1126\/science.283.5406.1272","article-title":"Communication Goes Multimodal","volume":"283","author":"Partan","year":"1999","journal-title":"Science"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1417","DOI":"10.1523\/JNEUROSCI.3675-12.2013","article-title":"Visual input enhances selective speech envelope tracking in auditory cortex at a \u201ccocktail party\u201d","volume":"33","author":"Golumbic","year":"2013","journal-title":"J. Neurosci."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1038\/nature11020","article-title":"Selective cortical representation of attended speaker in multi-talker speech perception","volume":"485","author":"Mesgarani","year":"2012","journal-title":"Nature"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3197517.3201357","article-title":"Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation","volume":"37","author":"Ephrat","year":"2018","journal-title":"ACM Trans. Graph."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Liu, Y., and Wei, Y. (2021, January 22\u201324). Multi-Modal Speech Separation Based on Two-Stage Feature Fusion. Proceedings of the 2021 IEEE 6th International Conference on Signal and Image Processing (ICSIP), Nanjing, China.","DOI":"10.1109\/ICSIP52628.2021.9688674"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wu, J., Xu, Y., Zhang, S., Chen, L., Yu, M., Xie, L., and Yu, D. (2019, January 14\u201318). Time domain audio visual speech separation. Proceedings of the 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Singapore.","DOI":"10.1109\/ASRU46091.2019.9003983"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1315","DOI":"10.1109\/LSP.2018.2853566","article-title":"Listen and Look: Audio-Visual Matching Assisted Speech Source Separation","volume":"25","author":"Lu","year":"2015","journal-title":"IEEE Signal Processing Lett."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Ito, K., Yamamoto, M., and Nagamatsu, K. (2011, January 6\u201311). Audio-visual speech enhancement method conditioned in the lip motion and speaker-discriminative embeddings. Proceedings of the ICASSP 2021\u20142021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9414133"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Deng, Y., and Wei, Y. (2022, January 5). Vision-Guided Speaker Embedding Based Speech Separation. Proceedings of the 2022 15th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI), Beijing, China.","DOI":"10.1109\/CISP-BMEI56279.2022.9980110"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Makishima, N., Ihori, M., Takashima, A., Tanaka, T., Orihashi, S., and Masumura, R. (2011, January 6\u201311). Audio Visual Speech Separation Using Cross-Modal Correspondence Loss. Proceedings of the ICASSP 2021\u20142021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9413491"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Li, C., and Qian, Y. (2020, January 4\u20138). Deep audio-visual speech separation with attention mechanism. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9054180"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Xiao, X., Lian, S., Luo, Z., and Li, S. (2018, January 19\u201321). Weighted Res-UNet for High-quality Retina Vessel Segmentation. Proceedings of the 2018 9th International Conference on Information Technology in Medicine and Education, Hangzhou, China.","DOI":"10.1109\/ITME.2018.00080"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015). U-Net: Convolutional Networks for Biomedical Image Segmentation. arXiv.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2015). Deep residual learning for image recognition. arXiv.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhou, X., Lin, M., and Sun, J. (2017). ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices. arXiv.","DOI":"10.1109\/CVPR.2018.00716"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Ma, N., Zhang, X., Zheng, H., and Sun, J. (2018). ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design. arXiv.","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"ref_35","unstructured":"Bai, S., Kolter, J.Z., and Koltun, V. (2018). An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1702","DOI":"10.1109\/TASLP.2018.2842159","article-title":"Supervised speech separation based on deep learning: An overview","volume":"26","author":"Wang","year":"2018","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"2236","DOI":"10.1121\/1.1610463","article-title":"Speech segregation based on sound localization","volume":"114","author":"Roman","year":"2003","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"1486","DOI":"10.1016\/j.specom.2006.09.003","article-title":"Binary and ratio timefrequency masks for robust speech recognition","volume":"48","author":"Srinivasan","year":"2006","journal-title":"Speech Commun."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"483","DOI":"10.1109\/TASLP.2015.2512042","article-title":"Complex ratio masking for monaural speech separation","volume":"24","author":"Williamson","year":"2016","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_40","unstructured":"Mangalam, K., and Salzamann, M. (2018). On Compressing U-net Using Knowledge Distillation. arXiv."},{"key":"ref_41","unstructured":"Ioffe, S., and Szegedy, C. (2015). Batch Normalization-Accelerating Deep Network Training by Reducing Internal Covariate Shift. arXiv."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J., and Kweon, I.S. (2018). CBAM: Convolutional Block Attention Module. arXiv.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., Albanie, S., Sun, G., and Wu, E. (2017). Squeeze-and-Excitation Networks. arXiv.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Gao, R., and Grauman, K. (2021, January 20\u201325). Visualvoice: Audio-visual speech separation with cross-modal consistency. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01524"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Martinez, B., Ma, P., Petridis, S., and Pantic, M. (2020, January 4\u20138). Lipreading using temporal convolutional networks. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053841"},{"key":"ref_46","unstructured":"Barron, J.T. (2017). A General and Adaptive Robust Loss Function. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Schroff, F., Kalenichenko, D., and Philbin, J. (2015). FaceNet: A Unified Embedding for Face Recognition and Clustering. arXiv.","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Chung, J.S., Nagrani, A., and Zisserman, A. (2018). VoxCeleb2: Deep Speaker Recognition. arXiv.","DOI":"10.21437\/Interspeech.2018-1929"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"2125","DOI":"10.1109\/TASL.2011.2114881","article-title":"An algorithm for intelligibility prediction of time-frequencyweighted noisy speech","volume":"19","author":"Taal","year":"2011","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_50","unstructured":"Rix, A.W., Beerends, J.G., Hollier, M.P., and Hekstra, A.P. (2001, January 7\u201311). Perceptual evaluation of speech quality (pesq)\u2014A new method for speech quality assessment of telephone networks and codecs. Proceedings of the 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing, Salt Lake City, UT, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/21\/8770\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:13:04Z","timestamp":1760130784000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/21\/8770"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,27]]},"references-count":50,"journal-issue":{"issue":"21","published-online":{"date-parts":[[2023,11]]}},"alternative-id":["s23218770"],"URL":"https:\/\/doi.org\/10.3390\/s23218770","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2023,10,27]]}}}