{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T14:25:22Z","timestamp":1780755922043,"version":"3.54.1"},"reference-count":53,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2022,12,20]],"date-time":"2022-12-20T00:00:00Z","timestamp":1671494400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"the Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62132010"],"award-info":[{"award-number":["62132010"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"the Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62002198"],"award-info":[{"award-number":["62002198"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Tsinghua University Initiative Scientific Research Program","award":["62132010"],"award-info":[{"award-number":["62132010"]}]},{"name":"Tsinghua University Initiative Scientific Research Program","award":["62002198"],"award-info":[{"award-number":["62002198"]}]},{"name":"Beijing Key Lab of Networked Multimedia","award":["62132010"],"award-info":[{"award-number":["62132010"]}]},{"name":"Beijing Key Lab of Networked Multimedia","award":["62002198"],"award-info":[{"award-number":["62002198"]}]},{"name":"the Institute for Guo Qiang, Tsinghua University","award":["62132010"],"award-info":[{"award-number":["62132010"]}]},{"name":"the Institute for Guo Qiang, Tsinghua University","award":["62002198"],"award-info":[{"award-number":["62002198"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Voice communication using an air-conduction microphone in noisy environments suffers from the degradation of speech audibility. Bone-conduction microphones (BCM) are robust against ambient noises but suffer from limited effective bandwidth due to their sensing mechanism. Although existing audio super-resolution algorithms can recover the high-frequency loss to achieve high-fidelity audio, they require considerably more computational resources than is available in low-power hearable devices. This paper proposes the first-ever real-time on-chip speech audio super-resolution system for BCM. To accomplish this, we built and compared a series of lightweight audio super-resolution deep-learning models. Among all these models, ATS-UNet was the most cost-efficient because the proposed novel Audio Temporal Shift Module (ATSM) reduces the network\u2019s dimensionality while maintaining sufficient temporal features from speech audio. Then, we quantized and deployed the ATS-UNet to low-end ARM micro-controller units for a real-time embedded prototype. The evaluation results show that our system achieved real-time inference speed on Cortex-M7 and higher quality compared with the baseline audio super-resolution method. Finally, we conducted a user study with ten experts and ten amateur listeners to evaluate our method\u2019s effectiveness to human ears. Both groups perceived a significantly higher speech quality with our method when compared to the solutions with the original BCM or air-conduction microphone with cutting-edge noise-reduction algorithms.<\/jats:p>","DOI":"10.3390\/s23010035","type":"journal-article","created":{"date-parts":[[2022,12,21]],"date-time":"2022-12-21T02:31:33Z","timestamp":1671589893000},"page":"35","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["Enabling Real-Time On-Chip Audio Super Resolution for Bone-Conduction Microphones"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2244-0184","authenticated-orcid":false,"given":"Yuang","family":"Li","sequence":"first","affiliation":[{"name":"Key Laboratory of Pervasive Computing, Ministry of Education, Department of Commputer Science and Technology, Tsinghua University, Beijing 100084, China"},{"name":"Department of Engineering, University of Cambridge, Cambridge CB2 1TN, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4249-8893","authenticated-orcid":false,"given":"Yuntao","family":"Wang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Pervasive Computing, Ministry of Education, Department of Commputer Science and Technology, Tsinghua University, Beijing 100084, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xin","family":"Liu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Paul G. Allen School of Computer, University of Washington, Seattle, WA 98195, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuanchun","family":"Shi","sequence":"additional","affiliation":[{"name":"Key Laboratory of Pervasive Computing, Ministry of Education, Department of Commputer Science and Technology, Tsinghua University, Beijing 100084, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shwetak","family":"Patel","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Paul G. Allen School of Computer, University of Washington, Seattle, WA 98195, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shao-Fu","family":"Shih","sequence":"additional","affiliation":[{"name":"Google Inc., Mountain View, CA 94043, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,12,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1408","DOI":"10.1109\/PROC.1969.7278","article-title":"High-resolution frequency-wavenumber spectrum analysis","volume":"57","author":"Capon","year":"1969","journal-title":"Proc. IEEE"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"636","DOI":"10.1109\/T-C.1972.223567","article-title":"Generalized Wiener Filtering Computation Techniques","volume":"C-21","author":"Pratt","year":"1972","journal-title":"IEEE Trans. Comput."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"113","DOI":"10.1109\/TASSP.1979.1163209","article-title":"Suppression of acoustic noise in speech using spectral subtraction","volume":"27","author":"Boll","year":"1979","journal-title":"IEEE Trans. Acoust. Speech Signal Process."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Park, S.R., and Lee, J.W. (2017, January 20\u201324). A Fully Convolutional Neural Network for Speech Enhancement. Proceedings of the Interspeech 2017, Stockholm, Sweden.","DOI":"10.21437\/Interspeech.2017-1465"},{"key":"ref_5","unstructured":"Macartney, C., and Weyde, T. (2018). Improved speech enhancement with the wave-u-net. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Shimamura, T., and Tamiya, T. (2005, January 7\u201310). A reconstruction filter for bone-conducted speech. Proceedings of the 48th Midwest Symposium on Circuits and Systems, 2005, Covington, KY, USA.","DOI":"10.1109\/MWSCAS.2005.1594483"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"495","DOI":"10.1016\/j.apergo.2010.09.004","article-title":"The effect of bone conduction microphone locations on speech intelligibility and sound quality","volume":"42","author":"McBride","year":"2011","journal-title":"Appl. Ergon."},{"key":"ref_8","unstructured":"Kuleshov, V., Enam, S.Z., and Ermon, S. (2017, January 24\u201326). Audio super-resolution using neural nets. Proceedings of the ICLR (Workshop Track), Toulon, France."},{"key":"ref_9","unstructured":"Birnbaum, S., Kuleshov, V., Enam, Z., Koh, P.W.W., and Ermon, S. (2019, January 8\u201314). Temporal FiLM: Capturing Long-Range Sequence Dependencies with Feature-Wise Modulations. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, USA."},{"key":"ref_10","unstructured":"Kim, S., and Sathe, V. (2019). Bandwidth extension on raw audio via generative adversarial networks. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Hao, X., Xu, C., Hou, N., Xie, L., Chng, E.S., and Li, H. (2020, January 4\u20138). Time-Domain Neural Network Approach for Speech Bandwidth Extension. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9054551"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Kegler, M., Beckmann, P., and Cernak, M. (2020, January 25\u201329). Deep Speech Inpainting of Time-Frequency Masks. Proceedings of the Interspeech 2020, Shanghai, China.","DOI":"10.21437\/Interspeech.2020-1532"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1006","DOI":"10.1109\/LSP.2014.2379648","article-title":"Can we Automatically Transform Speech Recorded on Common Consumer Devices in Real-World Environments into Professional Production Quality Speech?\u2014A Dataset, Insights, and Challenges","volume":"22","author":"Mysore","year":"2015","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Su, J., Jin, Z., and Finkelstein, A. (2020, January 25\u201329). HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial Networks. Proceedings of the Interspeech 2020, Shanghai, China.","DOI":"10.21437\/Interspeech.2020-2143"},{"key":"ref_15","unstructured":"International Telecommunication Union (2001). Perceptual Evaluation of Speech Quality (PESQ): An Objective Method for End-To-End Speech Quality Assessment of Narrow-Band Telephone Networks and Speech Codecs, ITU-T Publications."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"380","DOI":"10.1109\/TASSP.1976.1162849","article-title":"Distance measures for speech processing","volume":"24","author":"Gray","year":"1976","journal-title":"IEEE Trans. Acoust. Speech Signal Process."},{"key":"ref_17","first-page":"238","article-title":"About Multichannel Speech Signal Extraction and Separation Techniques","volume":"03","author":"Hidri","year":"2012","journal-title":"J. Signal Inf. Process."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1109","DOI":"10.1109\/TASSP.1984.1164453","article-title":"Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator","volume":"32","author":"Ephraim","year":"1984","journal-title":"IEEE Trans. Acoust. Speech Signal Process."},{"key":"ref_19","unstructured":"Scalart, P., and Filho, J. (1996, January 9). Speech enhancement based on a priori signal to noise estimation. Proceedings of the 1996 IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings, Atlanta, GA, USA."},{"key":"ref_20","unstructured":"Shin, H.S., Kang, H.G., and Fingscheidt, T. (2012, January 26\u201328). Survey of speech enhancement supported by a bone conduction microphone. Proceedings of the Speech Communication; 10. ITG Symposium, Braunschweig, Germany."},{"key":"ref_21","unstructured":"Liu, Z., Zhang, Z., Acero, A., Droppo, J., and Huang, X. (October, January 29). Direct filtering for air- and bone-conductive microphones. Proceedings of the IEEE sixth Workshop on Multimedia Signal Processing, Siena, Italy."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Lee, C.H., Rao, B.D., and Garudadri, H. (2018, January 2\u20136). Bone-Conduction Sensor Assisted Noise Estimation for Improved Speech Enhancement. Proceedings of the Interspeech 2018, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-1046"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Takada, M., Seki, S., and Toda, T. (2018, January 12\u201315). Self-Produced Speech Enhancement and Suppression Method using Air- and Body-Conductive Microphones. Proceedings of the 2018 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Honolulu, HI, USA.","DOI":"10.23919\/APSIPA.2018.8659663"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Chen, Y., Ma, Y., and Liu, H. (2020). A Real-Time Dual-Microphone Speech Enhancement Algorithm Assisted by Bone Conduction Sensor. Sensors, 20.","DOI":"10.3390\/s20185050"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1035","DOI":"10.1109\/LSP.2020.3000968","article-title":"Time-Domain Multi-Modal Bone\/Air Conducted Speech Enhancement","volume":"27","author":"Yu","year":"2020","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Shimamura, T., Mamiya, J., and Tamiya, T. (2006, January 27\u201330). Improving Bone-Conducted Speech Quality via Neural Network. Proceedings of the 2006 IEEE International Symposium on Signal Processing and Information Technology, Vancouver, BC, Canada.","DOI":"10.1109\/ISSPIT.2006.270876"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Rahman, M.S., and Shimamura, T. (2011, January 7\u201310). Intelligibility enhancement of bone conducted speech by an analysis-synthesis method. Proceedings of the 2011 IEEE 54th International Midwest Symposium on Circuits and Systems (MWSCAS), Seoul, Republic of Korea.","DOI":"10.1109\/MWSCAS.2011.6026374"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1321","DOI":"10.1121\/1.4976051","article-title":"In-ear microphone speech quality enhancement via adaptive filtering and artificial bandwidth extension","volume":"141","author":"Bouserhal","year":"2017","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"62638","DOI":"10.1109\/ACCESS.2018.2873728","article-title":"A Novel Encoder-Decoder Model via NS-LSTM Used for Bone-Conducted Speech Enhancement","volume":"6","author":"Shan","year":"2018","journal-title":"IEEE Access"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"106","DOI":"10.1016\/j.specom.2018.06.002","article-title":"Bone-conducted speech enhancement using deep denoising autoencoder","volume":"104","author":"Liu","year":"2018","journal-title":"Speech Commun."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Hussain, T., Tsao, Y., Siniscalchi, S.M., Wang, J.C., Wang, H.M., and Liao, W.H. (2021). Bone-Conducted Speech Enhancement Using Hierarchical Extreme Learning Machine. Lecture Notes in Electrical Engineering, Springer.","DOI":"10.1007\/978-981-15-9323-9_14"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zheng, C., Yang, J., Zhang, X., Sun, M., and Yao, K. (2019, January 18\u201321). Improving the Spectra Recovering of Bone-Conducted Speech via Structural SIMilarity Loss Function. Proceedings of the 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Lanzhou, China.","DOI":"10.1109\/APSIPAASC47483.2019.9023226"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Li, K., and Lee, C.H. (2015, January 19\u201324). A deep neural network approach to speech bandwidth expansion. Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, QLD, Australia.","DOI":"10.1109\/ICASSP.2015.7178801"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"347","DOI":"10.1109\/JSTSP.2019.2909077","article-title":"Adversarial Training for Speech Super-Resolution","volume":"13","author":"Eskimez","year":"2019","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Lim, T.Y., Yeh, R.A., Xu, Y., Do, M.N., and Hasegawa-Johnson, M. (2018, January 15\u201320). Time-Frequency Networks for Audio Super-Resolution. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8462049"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Feng, B., Jin, Z., Su, J., and Finkelstein, A. (2019, January 12\u201317). Learning Bandwidth Expansion Using Perceptually-motivated Loss. Proceedings of the ICASSP 2019\u20142019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8682367"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Wang, M., Wu, Z., Kang, S., Wu, X., Jia, J., Su, D., Yu, D., and Meng, H. (2018, January 26\u201329). Speech Super-Resolution Using Parallel WaveNet. Proceedings of the 2018 11th International Symposium on Chinese Spoken Language Processing (ISCSLP), Taipei, Taiwan.","DOI":"10.1109\/ISCSLP.2018.8706637"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Gupta, A., Shillingford, B., Assael, Y., and Walters, T.C. (2019, January 20\u201323). Speech Bandwidth Extension with Wavenet. Proceedings of the 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), New Paltz, NY, USA.","DOI":"10.1109\/WASPAA.2019.8937169"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Li, Y., Tagliasacchi, M., Rybakov, O., Ungureanu, V., and Roblek, D. (2021, January 6\u201311). Real-Time Speech Frequency Bandwidth Extension. Proceedings of the ICASSP 2021\u20142021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9413439"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Li, X., Chebiyyam, V., and Kirchhoff, K. (2019, January 15\u201319). Speech Audio Super-Resolution for Speech Recognition. Proceedings of the Interspeech 2019, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-3043"},{"key":"ref_41","unstructured":"Kumar, R., Kumar, K., Anand, V., Bengio, Y., and Courville, A. (2020). NU-GAN: High resolution neural upsampling with GAN. arXiv."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1145\/3448104","article-title":"SplitSR: An end-to-end approach to super-resolution on mobile devices","volume":"5","author":"Liu","year":"2021","journal-title":"Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Lee, R., Venieris, S.I., Dudziak, L., Bhattacharya, S., and Lane, N.D. (2019, January 21\u201325). MobiSR: Efficient on-device super-resolution through heterogeneous mobile processors. Proceedings of the The 25th Annual International Conference on Mobile Computing and Networking, Los Cabos, Mexico.","DOI":"10.1145\/3300061.3345455"},{"key":"ref_44","unstructured":"Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., and Isard, M. (2016, January 2\u20134). Tensorflow: A system for large-scale machine learning. Proceedings of the 12th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 16), Savannah, GA, USA."},{"key":"ref_45","unstructured":"Lai, L., Suda, N., and Chandra, V. (2018). Cmsis-nn: Efficient neural network kernels for arm cortex-m cpus. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Lin, J., Gan, C., and Han, S. (November, January 27). TSM: Temporal Shift Module for Efficient Video Understanding. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00718"},{"key":"ref_47","unstructured":"Abbas, S., Mosbah, M., and Zemmari, A. (, January August). ITU-T Recommendation G. 114, \u201cOne way transmission time\u201d. Proceedings of the International Conference on Dynamics in Logistics 2007 (LDIC 2007), Bremen, Germany. Available online: https:\/\/www.itu.int\/rec\/T-REC-G.114-200305-I\/en."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"1558","DOI":"10.1109\/PROC.1977.10770","article-title":"A unified approach to short-time Fourier analysis and synthesis","volume":"65","author":"Allen","year":"1977","journal-title":"Proc. IEEE"},{"key":"ref_49","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"208","DOI":"10.1121\/1.1901999","article-title":"A Scale for the Measurement of the Psychological Magnitude Pitch","volume":"8","author":"Volkmann","year":"1937","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"2371","DOI":"10.1121\/10.0002279","article-title":"Acoustic effects of medical, cloth, and transparent face masks on speech signals","volume":"148","author":"Corey","year":"2020","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Ochiai, T., Delcroix, M., Kinoshita, K., Ogawa, A., and Nakatani, T. (2019, January 12\u201317). A Unified Framework for Neural Speech Separation and Extraction. Proceedings of the ICASSP 2019\u20142019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683448"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Drakopoulos, G., Pikramenos, G., Spyrou, E., and Perantonis, S. (2019, January 18\u201320). Emotion Recognition from Speech: A Survey. Proceedings of the 15th International Conference on Web Information Systems and Technologies, Vienna, Austria.","DOI":"10.5220\/0008495004320439"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/1\/35\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:44:58Z","timestamp":1760147098000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/1\/35"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,20]]},"references-count":53,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,1]]}},"alternative-id":["s23010035"],"URL":"https:\/\/doi.org\/10.3390\/s23010035","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,20]]}}}