{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T00:08:13Z","timestamp":1755907693020,"version":"3.44.0"},"reference-count":67,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2024,5,13]],"date-time":"2024-05-13T00:00:00Z","timestamp":1715558400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Fund of China","doi-asserted-by":"crossref","award":["62102245"],"award-info":[{"award-number":["62102245"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2024,5,13]]},"abstract":"<jats:p>Speech enhancement on mobile devices is a very challenging task due to the complex environmental noises. Recent works using lip-induced ultrasound signals for speech enhancement open up new possibilities to solve such a problem. However, these multi-modal methods cannot be used in many scenarios where ultrasound-based lip sensing is unreliable or completely absent. In this paper, we propose a novel paradigm that can exploit the prior learned ultrasound knowledge for multi-modal speech enhancement only with the audio input and an additional pre-enrollment speaker embedding. We design a memory network to store the ultrasound memory and learn the interrelationship between the audio and ultrasound modality. During inference, the memory network is able to recall the ultrasound representations from audio input to achieve multi-modal speech enhancement without needing real ultrasound signals. Moreover, we introduce a speaker embedding module to further boost the enhancement performance as well as avoid the degradation of the recalling when the noise level is high. We adopt an end-to-end multi-task manner to train the proposed framework and perform extensive evaluations on the collected dataset. The results show that our method yields comparable performance with audio-ultrasound methods and significantly outperforms the audio-only methods.<\/jats:p>","DOI":"10.1145\/3659598","type":"journal-article","created":{"date-parts":[[2024,5,15]],"date-time":"2024-05-15T12:20:41Z","timestamp":1715775641000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Sensing to Hear through Memory"],"prefix":"10.1145","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7708-8694","authenticated-orcid":false,"given":"Qian","family":"Zhang","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-2347-1796","authenticated-orcid":false,"given":"Ke","family":"Liu","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8444-1636","authenticated-orcid":false,"given":"Dong","family":"Wang","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,5,15]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"The Conversation: Deep Audio-Visual Speech Enhancement. In INTERSPEECH.","author":"Afouras T.","year":"2018","unstructured":"T. Afouras, J. S. Chung, and A. Zisserman. 2018. The Conversation: Deep Audio-Visual Speech Enhancement. In INTERSPEECH."},{"key":"e_1_2_1_2_1","unstructured":"Ltd Beijing DataTang Technology Co. [n.d.]. aidatatang_200zh a free Chinese Mandarin speech corpus. https:\/\/www.datatang.com"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1979.1163209"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-1617"},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","unstructured":"J. S. Chung A. Nagrani and A. Zisserman. 2018. VoxCeleb2: Deep Speaker Recognition. In INTERSPEECH.","DOI":"10.21437\/Interspeech.2018-1929"},{"key":"e_1_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Alexandre Defossez Gabriel Synnaeve and Yossi Adi. 2020. Real Time Speech Enhancement in the Waveform Domain. In Interspeech.","DOI":"10.21437\/Interspeech.2020-2409"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-2650"},{"volume-title":"Adjunct Proceedings of the 2019 ACM International Joint Conference on Pervasive and Ubiquitous Computing and Proceedings of the 2019 ACM International Symposium on Wearable Computers","author":"Ding Feng","key":"e_1_2_1_8_1","unstructured":"Feng Ding, Dong Wang, Qian Zhang, and Run Zhao. 2019. ASSV: handwritten signature verification using acoustic signals. In Adjunct Proceedings of the 2019 ACM International Joint Conference on Pervasive and Ubiquitous Computing and Proceedings of the 2019 ACM International Symposium on Wearable Computers. Association for Computing Machinery, New York, NY, USA, 274--277."},{"key":"e_1_2_1_9_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3550303","article-title":"UltraSpeech: Speech Enhancement by Interaction between Ultrasound and Speech","volume":"6","author":"Ding Han","year":"2022","unstructured":"Han Ding, Yizhan Wang, Hao Li, Cui Zhao, Ge Wang, Wei Xi, and Jizhong Zhao. 2022. UltraSpeech: Speech Enhancement by Interaction between Ultrasound and Speech. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 6, 3 (2022), 1--25.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1984.1164453"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/89.397090"},{"key":"e_1_2_1_12_1","volume-title":"Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. arXiv preprint arXiv:1804.03619","author":"Ephrat Ariel","year":"2018","unstructured":"Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein. 2018. Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. arXiv preprint arXiv:1804.03619 (2018)."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201357"},{"key":"e_1_2_1_14_1","unstructured":"Sefik Emre Eskimez Takuya Yoshioka Huaming Wang Xiaofei Wang Zhuo Chen and Xuedong Huang. 2021. Personalized Speech Enhancement: New Models and Comprehensive Evaluation. arXiv:2110.09625"},{"key":"e_1_2_1_15_1","volume-title":"Proceedings of the 20th ACM Conference on Embedded Networked Sensor Systems (SenSys '22)","author":"Fu Yongjian","year":"2023","unstructured":"Yongjian Fu, Shuning Wang, Linghui Zhong, Lili Chen, Ju Ren, and Yaoxue Zhang. 2023. SVoice: Enabling Voice Communication in Silence via Acoustic Sensing on Commodity Devices. In Proceedings of the 20th ACM Conference on Embedded Networked Sensor Systems (SenSys '22). Association for Computing Machinery, New York, NY, USA, 622--636."},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"John Garofolo L Lamel W Fisher Jonathan Fiscus D Pallett and Nancy Dahlgren. 1993. Darpa Timit Acoustic-Phonetic Continuous Speech Corpus CD-ROM TIMIT.","DOI":"10.6028\/NIST.IR.4930"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2020.113582"},{"key":"e_1_2_1_18_1","doi-asserted-by":"crossref","unstructured":"Xiang Hao Xiangdong Su Radu Horaud and Xiaofei Li. 2021. Fullsubnet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement. In ICASSP 2021 - 2021 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). 6633--6637.","DOI":"10.1109\/ICASSP39728.2021.9414177"},{"volume-title":"Visual Speech Enhancement Without A Real Visual Stream. In 2021 IEEE Winter Conference on Applications of Computer Vision (WACV). 1925--1934","author":"Hegde Sindhu B","key":"e_1_2_1_19_1","unstructured":"Sindhu B Hegde, K R Prajwal, Rudrabha Mukhopadhyay, Vinay Namboodiri, and C.V. Jawahar. 2021. Visual Speech Enhancement Without A Real Visual Stream. In 2021 IEEE Winter Conference on Applications of Computer Vision (WACV). 1925--1934."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7471631"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/2209820.2210675"},{"key":"e_1_2_1_22_1","volume-title":"DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement. ArXiv abs\/2008.00264","author":"Hu Yanxin","year":"2020","unstructured":"Yanxin Hu, Yun Liu, Shubo Lv, Mengtao Xing, Shimin Zhang, Yihui Fu, Jian Wu, Bihong Zhang, and Lei Xie. 2020. DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement. ArXiv abs\/2008.00264 (2020)."},{"key":"e_1_2_1_23_1","volume-title":"Oriental COCOSDA","author":"Bengu Wu Hao Zheng Xingyu Na","year":"2017","unstructured":"Xingyu Na Bengu Wu Hao Zheng Hui Bu, Jiayu Du. 2017. AIShell-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline. In Oriental COCOSDA 2017. Submitted."},{"volume-title":"Proceedings of the 2022 Conference of the North American","author":"Xintong Li Renjie Zheng Junkun Chen","key":"e_1_2_1_24_1","unstructured":"Junkun Chen Xintong Li Renjie Zheng Yuxin Huang Xiaojie Chen Enlei Gong Zeyu Chen Xiaoguang Hu dianhai yu Yanjun Ma Liang Huang Hui Zhang, Tian Yuan. 2022. PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Demonstrations. Association for Computational Linguistics."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00036"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i1.20003"},{"key":"e_1_2_1_27_1","volume-title":"On information and sufficiency. The annals of mathematical statistics 22, 1","author":"Kullback Solomon","year":"1951","unstructured":"Solomon Kullback and Richard A Leibler. 1951. On information and sufficiency. The annals of mathematical statistics 22, 1 (1951), 79--86."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2003.808544"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3517253"},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the 20th ACM Conference on Embedded Networked Sensor Systems (SenSys '22)","author":"Li Dong","year":"2023","unstructured":"Dong Li, Jialin Liu, Sunghoon Ivan Lee, and Jie Xiong. 2023. Room-Scale Hand Gesture Recognition Using Smart Speakers. In Proceedings of the 20th ACM Conference on Embedded Networked Sensor Systems (SenSys '22). Association for Computing Machinery, New York, NY, USA, 462--475."},{"key":"e_1_2_1_31_1","first-page":"1","article-title":"BlinkListener: \" Listen\" to Your Eye Blink Using Your Smartphone","volume":"5","author":"Liu Jialin","year":"2021","unstructured":"Jialin Liu, Dong Li, Lei Wang, and Jie Xiong. 2021. BlinkListener: \" Listen\" to Your Eye Blink Using Your Smartphone. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 5, 2 (2021), 1--27.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_1_32_1","volume-title":"LipPass: Lip Reading-based User Authentication on Smartphones Leveraging Acoustic Signals. In IEEE INFOCOM 2018 - IEEE Conference on Computer Communications. 1466--1474","author":"Lu Li","year":"2018","unstructured":"Li Lu, Jiadi Yu, Yingying Chen, Hongbo Liu, Yanmin Zhu, Yunfei Liu, and Minglu Li. 2018. LipPass: Lip Reading-based User Authentication on Smartphones Leveraging Acoustic Signals. In IEEE INFOCOM 2018 - IEEE Conference on Computer Communications. 1466--1474."},{"key":"e_1_2_1_33_1","volume-title":"Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation","author":"Luo Yi","year":"2019","unstructured":"Yi Luo and Nima Mesgarani. 2019. Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation. IEEE\/ACM Transactions on Audio, Speech, and Language Processing 27 (Aug 2019), 1256--1266."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.archoralbio.2010.02.008"},{"key":"e_1_2_1_35_1","volume-title":"Hearing lips and seeing voices. Nature 264, 5588","author":"McGurk Harry","year":"1976","unstructured":"Harry McGurk and John MacDonald. 1976. Hearing lips and seeing voices. Nature 264, 5588 (1976), 746--748."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1147"},{"key":"e_1_2_1_37_1","doi-asserted-by":"crossref","unstructured":"A. Nagrani J. S. Chung and A. Zisserman. 2017. VoxCeleb: a large-scale speaker identification dataset. In INTERSPEECH.","DOI":"10.21437\/Interspeech.2017-950"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3412382.3458261"},{"key":"e_1_2_1_39_1","volume-title":"A Fully Convolutional Neural Network for Speech Enhancement. ArXiv abs\/1609.07132","author":"Park Se Rim","year":"2016","unstructured":"Se Rim Park and Jinwon Lee. 2016. A Fully Convolutional Neural Network for Speech Enhancement. ArXiv abs\/1609.07132 (2016)."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.283.5406.1272"},{"key":"e_1_2_1_41_1","volume-title":"SEGAN: Speech Enhancement Generative Adversarial Network. CoRR abs\/1703.09452","author":"Pascual Santiago","year":"2017","unstructured":"Santiago Pascual, Antonio Bonafonte, and Joan Serr\u00e0. 2017. SEGAN: Speech Enhancement Generative Adversarial Network. CoRR abs\/1703.09452 (2017). arXiv:1703.09452"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1010273"},{"key":"e_1_2_1_43_1","first-page":"1","article-title":"Mandibular movements in speech phrases---A syllabic quasiregular continuous oscillation","volume":"16","author":"Per Lindblad Stig Karlsson","year":"1991","unstructured":"Stig Karlsson Per Lindblad and Eva Heller. 1991. Mandibular movements in speech phrases---A syllabic quasiregular continuous oscillation. Scandinavian Journal of Logopedics and Phoniatrics 16, 1-2 (1991), 36--42.","journal-title":"Scandinavian Journal of Logopedics and Phoniatrics"},{"volume-title":"Proceedings of the 28th ACM International Conference on Multimedia. Association for Computing Machinery","author":"Prajwal K R","key":"e_1_2_1_44_1","unstructured":"K R Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, and C.V. Jawahar. 2020. A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild. In Proceedings of the 28th ACM International Conference on Multimedia. Association for Computing Machinery, New York, NY, USA, 484--492."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445687"},{"key":"e_1_2_1_46_1","volume-title":"2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.01CH37221)","volume":"2","author":"Rix A.W.","unstructured":"A.W. Rix, J.G. Beerends, M.P. Hollier, and A.P. Hekstra. 2001. Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs. In 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.01CH37221), Vol. 2. 749-752 vol.2."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447993.3448626"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2010.5495701"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2005.858005"},{"key":"e_1_2_1_50_1","unstructured":"Volcengine. 2023. Volcengine ASR. https:\/\/www.volcengine.com\/product\/asr."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3191771"},{"key":"e_1_2_1_52_1","volume-title":"Generalized End-to-End Loss for Speaker Verification. 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Wan Li","year":"2017","unstructured":"Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez-Moreno. 2017. Generalized End-to-End Loss for Speaker Verification. 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2017), 4879--4883."},{"key":"e_1_2_1_53_1","volume-title":"THCHS-30: A Free Chinese Speech Corpus. ArXiv abs\/1512.01882","author":"Wang Dong","year":"2015","unstructured":"Dong Wang and Xuewei Zhang. 2015. THCHS-30: A Free Chinese Speech Corpus. ArXiv abs\/1512.01882 (2015)."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00252"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-1101"},{"key":"e_1_2_1_56_1","volume-title":"Proceedings of the 22nd Annual International Conference on Mobile Computing and Networking. 82--94","author":"Wang Wei","year":"2016","unstructured":"Wei Wang, Alex X Liu, and Ke Sun. 2016. Device-free gesture tracking using acoustic signals. In Proceedings of the 22nd Annual International Conference on Mobile Computing and Networking. 82--94."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2014.2352935"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2019.2944058"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2014.2364452"},{"key":"e_1_2_1_60_1","volume-title":"PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement Network. In AAAI Conference on Artificial Intelligence.","author":"Yin Dacheng","year":"2019","unstructured":"Dacheng Yin, Chong Luo, Zhiwei Xiong, and Wenjun Zeng. 2019. PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement Network. In AAAI Conference on Artificial Intelligence."},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3081333.3081356"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3659614"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448087"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/3494990"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3381008"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3488544"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2019.2922820"}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3659598","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3659598","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T17:04:40Z","timestamp":1755882280000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3659598"}},"subtitle":["Ultrasound Speech Enhancement without Real Ultrasound Signals"],"short-title":[],"issued":{"date-parts":[[2024,5,13]]},"references-count":67,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,5,13]]}},"alternative-id":["10.1145\/3659598"],"URL":"https:\/\/doi.org\/10.1145\/3659598","relation":{},"ISSN":["2474-9567"],"issn-type":[{"type":"electronic","value":"2474-9567"}],"subject":[],"published":{"date-parts":[[2024,5,13]]},"assertion":[{"value":"2024-05-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}