{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T15:11:25Z","timestamp":1784301085948,"version":"3.55.0"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,12,19]],"date-time":"2023-12-19T00:00:00Z","timestamp":1702944000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Singapore Ministry of Education (MOE) Academic Research Fund (AcRF) Tier 1 Grant","award":["21-SIS-SMU-036"],"award-info":[{"award-number":["21-SIS-SMU-036"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2023,12,19]]},"abstract":"<jats:p>Wireless earbuds have been gaining increasing popularity and using them to make phone calls or issue voice commands requires the earbud microphones to pick up human speech. When the speaker is in a noisy environment, speech quality degrades significantly and requires speech enhancement (SE). In this paper, we present ClearSpeech, a novel deep-learning-based SE system designed for wireless earbuds. Specifically, by jointly using the earbud's in-ear and out-ear microphones, we devised a suite of techniques to effectively fuse the two signals and enhance the magnitude and phase of the speech spectrogram. We built an earbud prototype to evaluate ClearSpeech under various settings with data collected from 20 subjects. Our results suggest that ClearSpeech can improve the SE performance significantly compared to conventional approaches using the out-ear microphone only. We also show that ClearSpeech can process user speech in real-time on smartphones.<\/jats:p>","DOI":"10.1145\/3631409","type":"journal-article","created":{"date-parts":[[2024,1,12]],"date-time":"2024-01-12T12:52:04Z","timestamp":1705063924000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["ClearSpeech"],"prefix":"10.1145","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3824-234X","authenticated-orcid":false,"given":"Dong","family":"Ma","sequence":"first","affiliation":[{"name":"Singapore Management University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3806-1493","authenticated-orcid":false,"given":"Ting","family":"Dang","sequence":"additional","affiliation":[{"name":"Nokia Bell Labs, Cambridge, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3690-0321","authenticated-orcid":false,"given":"Ming","family":"Ding","sequence":"additional","affiliation":[{"name":"Data61, CSIRO, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6289-9902","authenticated-orcid":false,"given":"Rajesh","family":"Balan","sequence":"additional","affiliation":[{"name":"Singapore Management University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,1,12]]},"reference":[{"key":"e_1_2_2_1_1","volume-title":"Bela Mini Board. https:\/\/learn.bela.io\/products\/bela-boards\/bela-mini\/. (Accessed on","year":"2022","unstructured":"Online. Bela Mini Board. https:\/\/learn.bela.io\/products\/bela-boards\/bela-mini\/. (Accessed on Dec 4, 2022)."},{"key":"e_1_2_2_2_1","volume-title":"https:\/\/www.cuidevices.com\/product\/resource\/cmc-4015-40l100.pdf. (Accessed on","year":"2022","unstructured":"Online. Microphone. https:\/\/www.cuidevices.com\/product\/resource\/cmc-4015-40l100.pdf. (Accessed on Dec 4, 2022)."},{"key":"e_1_2_2_3_1","volume-title":"https:\/\/pytorch.org\/docs\/stable\/quantization.html. (Accessed on","author":"Quantization Pytorch","year":"2022","unstructured":"Online. Pytorch Quantization. https:\/\/pytorch.org\/docs\/stable\/quantization.html. (Accessed on Dec 4, 2022)."},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/APSIPA.2016.7820886"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1979.1163209"},{"key":"e_1_2_2_6_1","volume-title":"Motion-resilient heart rate monitoring with in-ear microphones. arXiv preprint arXiv:2108.09393","author":"Butkow Kayla-Jade","year":"2021","unstructured":"Kayla-Jade Butkow, Ting Dang, Andrea Ferlini, Dong Ma, and Cecilia Mascolo. 2021. Motion-resilient heart rate monitoring with in-ear microphones. arXiv preprint arXiv:2108.09393 (2021)."},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3498361.3538933"},{"key":"e_1_2_2_8_1","volume-title":"Inter-Subnet: Speech Enhancement with Subband Interaction. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1--5.","author":"Chen Jun","year":"2023","unstructured":"Jun Chen, Wei Rao, Zilin Wang, Jiuxin Lin, Zhiyong Wu, Yannan Wang, Shidong Shang, and Helen Meng. 2023. Inter-Subnet: Speech Enhancement with Subband Interaction. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1--5."},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9747888"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201357"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9746962"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447993.3483240"},{"key":"e_1_2_2_13_1","first-page":"1","article-title":"EarEcho: Using ear canal echo for wearable authentication","volume":"3","author":"Gao Yang","year":"2019","unstructured":"Yang Gao, Wei Wang, Vir V Phoha, Wei Sun, and Zhanpeng Jin. 2019. EarEcho: Using ear canal echo for wearable authentication. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 3, 3 (2019), 1--24.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.6028\/NIST.IR.4930"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952261"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAP.1982.1142739"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1515\/cdbme-2018-0072"},{"key":"e_1_2_2_18_1","first-page":"1","article-title":"EarCommand: \"Hearing\" your silent speech commands in ear","volume":"6","author":"Jin Yincheng","year":"2022","unstructured":"Yincheng Jin, Yang Gao, Xuhai Xu, Seokmin Choi, Jiyang Li, Feng Liu, Zhengxiong Li, and Zhanpeng Jin. 2022. EarCommand: \"Hearing\" your silent speech commands in ear. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 6, 2 (2022), 1--28.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8462649"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2013-130"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9054266"},{"key":"e_1_2_2_22_1","volume-title":"Conv-tasnet: Surpassing ideal time--frequency magnitude masking for speech separation","author":"Luo Yi","year":"2019","unstructured":"Yi Luo and Nima Mesgarani. 2019. Conv-tasnet: Surpassing ideal time--frequency magnitude masking for speech separation. IEEE\/ACM transactions on audio, speech, and language processing 27, 8 (2019), 1256--1266."},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458864.3467680"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/TBME.2017.2720463"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3287058"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3066303"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2013.2270369"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1055\/s-0028-1089925"},{"key":"e_1_2_2_29_1","unstructured":"Vinod Nair and Geoffrey E Hinton. 2010. Rectified linear units improve restricted boltzmann machines. In Icml."},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.clinph.2006.10.008"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_2_2_32_1","volume-title":"1996 IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings","volume":"2","author":"Pascal","unstructured":"Pascal Scalart et al. 1996. Speech enhancement based on a priori signal to noise estimation. In 1996 IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings, Vol. 2. IEEE, 629--632."},{"key":"e_1_2_2_33_1","volume-title":"Audio Engineering Society Conference: 2019 AES International Conference on Headphone Technology. Audio Engineering Society.","author":"Schlieper Roman","year":"2019","unstructured":"Roman Schlieper, Song Li, Stephan Preihs, and J\u00fcrgen Peissig. 2019. The relationship between the acoustic impedance of headphones and the occlusion effect. In Audio Engineering Society Conference: 2019 AES International Conference on Headphone Technology. Audio Engineering Society."},{"key":"e_1_2_2_34_1","volume-title":"LTE - The UMTS Long Term Evolution: From Theory to Practice","author":"Sesia Stefania","unstructured":"Stefania Sesia, Issam Toufik, and Matthew Baker. 2011. LTE - The UMTS Long Term Evolution: From Theory to Practice. John Wiley & Sons Ltd."},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3539490.3539600"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447993.3448626"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7953221"},{"key":"e_1_2_2_38_1","volume-title":"Beamforming: A versatile approach to spatial filtering","author":"Van Veen Barry D","year":"1988","unstructured":"Barry D Van Veen and Kevin M Buckley. 1988. Beamforming: A versatile approach to spatial filtering. IEEE assp magazine 5, 2 (1988), 4--24."},{"key":"e_1_2_2_39_1","doi-asserted-by":"crossref","unstructured":"Panqu Wang Pengfei Chen Ye Yuan Ding Liu Zehua Huang Xiaodi Hou and Garrison Cottrell. 2018. Understanding convolution for semantic segmentation. In 2018 IEEE winter conference on applications of computer vision (WACV). Ieee 1451--1460.","DOI":"10.1109\/WACV.2018.00163"},{"key":"e_1_2_2_40_1","volume-title":"Complex ratio masking for monaural speech separation","author":"Williamson Donald S","year":"2015","unstructured":"Donald S Williamson, Yuxuan Wang, and DeLiang Wang. 2015. Complex ratio masking for monaural speech separation. IEEE\/ACM transactions on audio, speech, and language processing 24, 3 (2015), 483--492."},{"key":"e_1_2_2_41_1","volume-title":"An experimental study on speech enhancement based on deep neural networks","author":"Xu Yong","year":"2013","unstructured":"Yong Xu, Jun Du, Li-Rong Dai, and Chin-Hui Lee. 2013. An experimental study on speech enhancement based on deep neural networks. IEEE Signal processing letters 21, 1 (2013), 65--68."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6489"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.14203\/jet.v21.19-26"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3494990"}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3631409","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3631409","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,27]],"date-time":"2025-08-27T16:57:55Z","timestamp":1756313875000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3631409"}},"subtitle":["Improving Voice Quality of Earbuds Using Both In-Ear and Out-Ear Microphones"],"short-title":[],"issued":{"date-parts":[[2023,12,19]]},"references-count":44,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,12,19]]}},"alternative-id":["10.1145\/3631409"],"URL":"https:\/\/doi.org\/10.1145\/3631409","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,19]]},"assertion":[{"value":"2024-01-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}