{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,16]],"date-time":"2026-04-16T21:02:35Z","timestamp":1776373355673,"version":"3.51.2"},"reference-count":47,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"name":"Fundamental and Interdisciplinary Disciplines Breakthrough Plan of the Ministry of Education of China","award":["No. JYB2025XDXM901"],"award-info":[{"award-number":["No. JYB2025XDXM901"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["No. 62272223"],"award-info":[{"award-number":["No. 62272223"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100013058","name":"Key Research and Development Project of Jiangsu Province","doi-asserted-by":"crossref","award":["No. BE2016120"],"award-info":[{"award-number":["No. BE2016120"]}],"id":[{"id":"10.13039\/501100013058","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["No. 62472299"],"award-info":[{"award-number":["No. 62472299"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["No. U22A2031"],"award-info":[{"award-number":["No. U22A2031"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100013058","name":"Key Research and Development Project of Jiangsu Province","doi-asserted-by":"crossref","award":["No. BE2015154"],"award-info":[{"award-number":["No. BE2015154"]}],"id":[{"id":"10.13039\/501100013058","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["No. 2023YFB4502400"],"award-info":[{"award-number":["No. 2023YFB4502400"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2026,3,16]]},"abstract":"<jats:p>Eavesdropping on target speech in enclosed and private rooms, such as secret gatherings or private parties, has become a growing research trend and poses one of the most significant threats to personal privacy leakage. Previous studies often made various assumptions about the target and scenarios of eavesdropping, such as the presence of only a single speaker, a very slow speech rate, and clear separation of words, all of which greatly limited the applicability of eavesdropping systems. In this paper, we propose an acoustic eavesdropping system, mmMPS, based on commercial off-the-shelf millimeter-wave (COTS mmWave) radar, which is capable of eavesdropping on specific target speech in scenarios where multiple people speak simultaneously, without requiring prior knowledge of background voices, speech rate, or intensity. The system utilizes a single COTS mmWave radar to capture faint signals caused by the vibration of objects due to human speech, combines these with our proposed denoising algorithm to recover the speech signal, and then uses our recognition model along with large audio model (LAM) to extract target keyword speech from the mixed speech signals of multiple speakers. Extensive experiments were conducted on a nearly 120 hours audio dataset and 60 hours radar dataset that we created using LAM and collected in experimental scenarios. The system is capable of covering various target users, speech rates, and speech intensities. The results show that in a scenario with 1\u20132 simultaneous speakers, the system achieves the average WER\/CER (Word\/Character Error Rate) of 6.46% (the lower, the better) in identifying 36 target keywords. In a scenario of 3\u20135 speakers, the average WER\/CER is 14.97%, and even in the scenario with up to 7 simultaneous speakers, the average WER\/CER remains below 30%. To the best of our knowledge, mmMPS is the first system to achieve high-accuracy eavesdropping on target speech in multi-speaker scenarios, significantly expanding the application scope of acoustic eavesdropping.<\/jats:p>","DOI":"10.1145\/3789687","type":"journal-article","created":{"date-parts":[[2026,3,16]],"date-time":"2026-03-16T17:51:14Z","timestamp":1773683474000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["LAM-assisted Acoustic Eavesdropping in Multi-speaker Scenarios via Commercial mmWave Radar"],"prefix":"10.1145","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-2463-9586","authenticated-orcid":false,"given":"Guodong","family":"Liu","sequence":"first","affiliation":[{"name":"Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0091-0931","authenticated-orcid":false,"given":"Lei","family":"Wang","sequence":"additional","affiliation":[{"name":"Soochow University, Suzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-1263-6559","authenticated-orcid":false,"given":"Minjun","family":"Jiang","sequence":"additional","affiliation":[{"name":"Soochow University, Suzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-4767-5922","authenticated-orcid":false,"given":"Qianran","family":"Qiao","sequence":"additional","affiliation":[{"name":"Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-3674-7841","authenticated-orcid":false,"given":"Keran","family":"Li","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, Jiangsu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0545-8187","authenticated-orcid":false,"given":"Haipeng","family":"Dai","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6934-1685","authenticated-orcid":false,"given":"Guihai","family":"Chen","sequence":"additional","affiliation":[{"name":"Computer Science, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,3,16]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-023-14649-x"},{"key":"e_1_2_1_2_1","volume-title":"Jamie Ryan Kiros, and Geoffrey E Hinton","author":"Ba Jimmy Lei","year":"2016","unstructured":"Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. arXiv preprint arXiv:1607.06450 (2016)."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.14722\/ndss.2020.24076"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP46214.2022.9833568"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jfranklin.2023.11.038"},{"key":"e_1_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Abe Davis Michael Rubinstein Neal Wadhwa Gautham J Mysore Fredo Durand and William T Freeman. 2014. The visual microphone: Passive recovery of sound from video. (2014).","DOI":"10.1145\/2601097.2601119"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.408467"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM53939.2023.10229095"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3422622"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143891"},{"key":"e_1_2_1_11_1","volume-title":"CSI2Dig: Recovering Digit Content from Smartphone Loudspeakers Using Channel State Information. arXiv preprint arXiv:2504.14812","author":"Gu Yangyang","year":"2025","unstructured":"Yangyang Gu, Xianglong Li, Haolin Wu, Jing Chen, Kun He, Ruiying Du, and Cong Wu. 2025. CSI2Dig: Recovering Digit Content from Smartphone Loudspeakers Using Channel State Information. arXiv preprint arXiv:2504.14812 (2025)."},{"key":"e_1_2_1_12_1","volume-title":"Conformer: Convolution-augmented transformer for speech recognition. arXiv preprint arXiv:2005.08100","author":"Gulati Anmol","year":"2020","unstructured":"Anmol Gulati, James Qin, Chung-Cheng Chiu, et al. 2020. Conformer: Convolution-augmented transformer for speech recognition. arXiv preprint arXiv:2005.08100 (2020)."},{"key":"e_1_2_1_13_1","volume-title":"Denoising diffusion probabilistic models. Advances in neural information processing systems 33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840\u20136851."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2022.3226690"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3603165.3607440"},{"key":"e_1_2_1_16_1","first-page":"28708","article-title":"Masked autoencoders that listen","volume":"35","author":"Huang Poyao","year":"2022","unstructured":"Poyao Huang, Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, and Christoph Feichtenhofer. 2022. Masked autoencoders that listen. Advances in Neural Information Processing Systems 35 (2022), 28708\u201328720.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_17_1","unstructured":"Xuedong Huang Alex Acero Hsiao-Wuen Hon and Raj Reddy. 2001. Spoken Language Processing: A guide to theory algorithm and system development. Prentice hall PTR."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2013.04.014"},{"key":"e_1_2_1_19_1","unstructured":"Diederik P Kingma Max Welling et al. 2013. Auto-encoding variational bayes."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2019.00008"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME52920.2022.9859720"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1002\/admt.202301006"},{"key":"e_1_2_1_23_1","volume-title":"Zero-shot voice conversion with diffusion transformers. arXiv preprint arXiv.2411:09943","author":"Liu Songting","year":"2024","unstructured":"Songting Liu. 2024. Zero-shot voice conversion with diffusion transformers. arXiv preprint arXiv.2411:09943 (2024)."},{"key":"e_1_2_1_24_1","volume-title":"Separate and diffuse: Using a pretrained diffusion model for improving source separation. arXiv preprint arXiv:2301.10752","author":"Lutati Shahar","year":"2023","unstructured":"Shahar Lutati, Eliya Nachmani, and Lior Wolf. 2023. Separate and diffuse: Using a pretrained diffusion model for improving source separation. arXiv preprint arXiv:2301.10752 (2023)."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3287058"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2023.101869"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the 23rd USENIX Security Symposium. 1053\u20131067","author":"Michalevsky Yan","year":"2014","unstructured":"Yan Michalevsky, Dan Boneh, and Gabi Nakibly. 2014. Gyrophone: Recognizing speech from gyroscope signals. In Proceedings of the 23rd USENIX Security Symposium. 1053\u20131067."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.391370"},{"key":"e_1_2_1_29_1","volume-title":"Lamphone: Real-time passive sound recovery from light bulb vibrations. Cryptology ePrint Archive","author":"Nassi Ben","year":"2020","unstructured":"Ben Nassi, Yaron Pirutin, Adi Shamir, Yuval Elovici, and Boris Zadov. 2020. Lamphone: Real-time passive sound recovery from light bulb vibrations. Cryptology ePrint Archive (2020)."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"e_1_2_1_32_1","volume-title":"Theory and applications of digital speech processing","author":"Rabiner Lawrence","unstructured":"Lawrence Rabiner and Ronald Schafer. 2010. Theory and applications of digital speech processing. Prentice Hall Press."},{"key":"e_1_2_1_33_1","volume-title":"The chirp z-transform algorithm","author":"Rabiner L","year":"2003","unstructured":"L Rabiner, R W Schafer, and C Rader. 2003. The chirp z-transform algorithm. IEEE transactions on audio and electroacoustics 17, 2 (2003), 86\u201392."},{"key":"e_1_2_1_34_1","volume-title":"International conference on machine learning. 28492\u201328518","author":"Radford Alec","year":"2023","unstructured":"Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning. 28492\u201328518."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3384419.3430781"},{"key":"e_1_2_1_36_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_1_37_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3534592","article-title":"Wavesdropper: Through-wall word detection of human speech via commercial mmwave devices","volume":"6","author":"Wang Chao","year":"2022","unstructured":"Chao Wang, Feng Lin, Zhongjie Ba, Fan Zhang, Wenyao Xu, and Kui Ren. 2022. Wavesdropper: Through-wall word detection of human speech via commercial mmwave devices. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 6, 2 (2022), 1\u201326.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/9780470043387"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2639108.2639112"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1002\/advs.202105056"},{"key":"e_1_2_1_41_1","unstructured":"Xinsheng Wang Mingqi Jiang Ziyang Ma et al. 2025. Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens. arXiv preprint arXiv:2503.01710 (2025)."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2789168.2790119"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2742647.2742658"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3678577"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3610873"},{"key":"e_1_2_1_47_1","first-page":"49842","article-title":"Unipc: A unified predictor-corrector framework for fast sampling of diffusion models","volume":"36","author":"Zhao Wenliang","year":"2023","unstructured":"Wenliang Zhao, Lujia Bai, Yongming Rao, Jie Zhou, and Jiwen Lu. 2023. Unipc: A unified predictor-corrector framework for fast sampling of diffusion models. Advances in Neural Information Processing Systems 36 (2023), 49842\u201349869.","journal-title":"Advances in Neural Information Processing Systems"}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3789687","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,16]],"date-time":"2026-04-16T19:59:23Z","timestamp":1776369563000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3789687"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,16]]},"references-count":47,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,3,16]]}},"alternative-id":["10.1145\/3789687"],"URL":"https:\/\/doi.org\/10.1145\/3789687","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,16]]},"assertion":[{"value":"2026-03-16","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}