{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,24]],"date-time":"2026-03-24T15:55:30Z","timestamp":1774367730973,"version":"3.50.1"},"reference-count":41,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2023,3,9]],"date-time":"2023-03-09T00:00:00Z","timestamp":1678320000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China Joint Fund Key Project","award":["U2003207"],"award-info":[{"award-number":["U2003207"]}]},{"name":"National Natural Science Foundation of China Joint Fund Key Project","award":["61902064"],"award-info":[{"award-number":["61902064"]}]},{"name":"National Natural Science Foundation of China Joint Fund Key Project","award":["1908085MF209"],"award-info":[{"award-number":["1908085MF209"]}]},{"name":"National Natural Science Foundation of China Joint Fund Key Project","award":["KJ2018A0018"],"award-info":[{"award-number":["KJ2018A0018"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U2003207"],"award-info":[{"award-number":["U2003207"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61902064"],"award-info":[{"award-number":["61902064"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["1908085MF209"],"award-info":[{"award-number":["1908085MF209"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["KJ2018A0018"],"award-info":[{"award-number":["KJ2018A0018"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Natural Science Foundation of Anhui Province","award":["U2003207"],"award-info":[{"award-number":["U2003207"]}]},{"name":"Natural Science Foundation of Anhui Province","award":["61902064"],"award-info":[{"award-number":["61902064"]}]},{"name":"Natural Science Foundation of Anhui Province","award":["1908085MF209"],"award-info":[{"award-number":["1908085MF209"]}]},{"name":"Natural Science Foundation of Anhui Province","award":["KJ2018A0018"],"award-info":[{"award-number":["KJ2018A0018"]}]},{"name":"Key Projects of Natural Science Foundation of Anhui Province Universities","award":["U2003207"],"award-info":[{"award-number":["U2003207"]}]},{"name":"Key Projects of Natural Science Foundation of Anhui Province Universities","award":["61902064"],"award-info":[{"award-number":["61902064"]}]},{"name":"Key Projects of Natural Science Foundation of Anhui Province Universities","award":["1908085MF209"],"award-info":[{"award-number":["1908085MF209"]}]},{"name":"Key Projects of Natural Science Foundation of Anhui Province Universities","award":["KJ2018A0018"],"award-info":[{"award-number":["KJ2018A0018"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Because of the acoustic characteristics of bone-conducted (BC) speech, BC speech can be enhanced to better communicate in a complex environment with high noise. Existing BC speech enhancement models have weak spectral recovery capability for the high-frequency part of BC speech and have poor enhancement and robustness for the speaker-independent BC speech datasets. To improve the enhancement effect of BC speech for speaker-independent speech enhancement, we use a GANs method to establish the feature mapping between BC and air-conducted (AC) speech to recover the missing components of BC speech. In addition, the method adds the training of the spectral distance constraint model and, finally, uses the enhanced model completed by the training to reconstruct the BC speech. The experimental results show that this method is superior to the comparison methods such as CycleGAN, BLSTM, GMM, and StarGAN in terms of speaker-independent BC speech enhancement and can obtain higher subjective and objective evaluation results of enhanced BC speech.<\/jats:p>","DOI":"10.3390\/a16030153","type":"journal-article","created":{"date-parts":[[2023,3,10]],"date-time":"2023-03-10T01:31:41Z","timestamp":1678411901000},"page":"153","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Speaker-Independent Spectral Enhancement for Bone-Conducted Speech"],"prefix":"10.3390","volume":"16","author":[{"given":"Liangliang","family":"Cheng","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Anhui University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yunfeng","family":"Dou","sequence":"additional","affiliation":[{"name":"Anhui Finance & Trade Vocational College, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jian","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Anhui University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huabin","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Anhui University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liang","family":"Tao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Anhui University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"574","DOI":"10.1016\/j.procs.2015.06.066","article-title":"Speech Enhancement using Spectral Subtraction-type Algorithms: A Comparison and Simulation Study","volume":"54","author":"Upadhyay","year":"2015","journal-title":"Procedia Comput. Sci."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Duong, H., Nguyen, Q.C., Nguyen, C., Tran, T., and Duong, N.Q. (2015, January 3\u20134). Speech enhancement based on nonnegative matrix factorization with mixed group sparsity constraint. Proceedings of the Sixth International Symposium on Information and Communication Technology, Hue, Vietnam.","DOI":"10.1145\/2833258.2833276"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Wang, K., He, B., and Zhu, W. (2021, January 6\u201311). TSTNN: Two-Stage Transformer Based Neural Network for Speech Enhancement in the Time Domain. Proceedings of the ICASSP 2021\u20142021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9413740"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Hu, Y., Liu, Y., Lv, S., Xing, M., Zhang, S., Fu, Y., Wu, J., Zhang, B., and Xie, L. (2020). DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement. arXiv.","DOI":"10.21437\/Interspeech.2020-2537"},{"key":"ref_5","unstructured":"Shin, H.S., Kang, H.G., and Fingscheidt, T. (2012, January 26\u201328). Survey of Speech Enhancement Supported by a Bone Conduction Microphone. Proceedings of the Speech Communication; 10. ITG Symposium; Proceedings of VDE, Braunschweig, Germany."},{"key":"ref_6","unstructured":"tat Vu, T., Unoki, M., and Akagi, M. (2008, January 4\u20136). An LP-based blind model for restoring bone-conducted speech. Proceedings of the Communications and Electronics, 2008. ICCE 2008. Second International Conference on IEEE, Hoi An, Vietnam."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2001","DOI":"10.1587\/transfun.E102.A.2001","article-title":"Spectra Restoration of Bone-Conducted Speech via Attention-Based Contextual Information and Spectro-Temporal Structure Constraint","volume":"102","author":"Zheng","year":"2019","journal-title":"IEICE Trans. Fundam. Electron. Commun. Comput. Sci."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"106","DOI":"10.1016\/j.specom.2018.06.002","article-title":"Bone-conducted speech enhancement using deep denoising autoencoder","volume":"104","author":"Liu","year":"2018","journal-title":"Speech Commun."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Huang, B., Gong, Y., Sun, J., and Shen, Y. (2017, January 16\u201319). A wearable bone-conducted speech enhancement system for strong background noises. Proceedings of the 2017 18th International Conference on Electronic Packaging Technology (ICEPT), Harbin, China.","DOI":"10.1109\/ICEPT.2017.8046759"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Watanabe, D., Sugiura, Y., Shimamura, T., and Makinae, H. (2017, January 6\u20139). Speech enhancement for bone-conducted speech based on low-order cepstrum restoration. Proceedings of the 2017 International Symposium on Intelligent Signal Processing and Communication Systems (ISPACS), Xiamen, China.","DOI":"10.1109\/ISPACS.2017.8266475"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"465","DOI":"10.1016\/j.specom.2010.12.003","article-title":"The importance of phase in speech enhancement","volume":"53","author":"Paliwal","year":"2011","journal-title":"Speech Commun."},{"key":"ref_12","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014). Neural Information Processing Systems, MIT Press."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, X., Li, Y., Zhang, H., and Shan, Y. (2021, January 20\u201325). Towards Real-World Blind Face Restoration with Generative Facial Prior. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00905"},{"key":"ref_14","unstructured":"Tian, Y., Ren, J., Chai, M., Olszewski, K., Peng, X., Metaxas, D.N., and Tulyakov, S. (2021). A Good Image Generator Is What You Need for High-Resolution Video Synthesis. arXiv."},{"key":"ref_15","unstructured":"Binkowski, M., Donahue, J., Dieleman, S., Clark, A., Elsen, E., Casagrande, N., Cobo, L.C., and Simonyan, K. (2019). High Fidelity Speech Synthesis with Adversarial Networks. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Pan, Q., Zhou, J., Gao, T., and Tao, L. (2020, January 12\u201315). Bone-Conducted Speech to Air-Conducted Speech Conversion Based on CycleConsistent Adversarial Networks. Proceedings of the 2020 IEEE 3rd International Conference on Information Communication and Signal Processing (ICICSP), Shanghai, China.","DOI":"10.1109\/ICICSP50920.2020.9232121"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Cho, K., Merrienboer, B.V., G\u00fcl\u00e7ehre, \u00c7., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014, January 25\u201329). Learning Phrase Representations using RNN Encoder\u2014Decoder for Statistical Machine Translation. Proceedings of the Conference on Empirical Methods in Natural Language Processing 2014, Doha, Qatar.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Shimamura, T., and Tamiya, T. (2005, January 7\u201310). A reconstruction filter for bone-conducted speech. Proceedings of the 48th Midwest Symposium on Circuits and Systems, Covington, KY, USA.","DOI":"10.1109\/MWSCAS.2005.1594483"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1109\/LSP.2003.808549","article-title":"Combining standard and throat microphones for robust speech recognition","volume":"10","author":"Graciarena","year":"2003","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TASL.2013.2274696","article-title":"Body Conducted Speech Enhancement by Equalization and Signal Fusion","volume":"21","author":"Dekens","year":"2013","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Kondo, K., Fujita, T., and Nakagawa, K. (2006, January 28\u201330). On Equalization of Bone Conducted Speech for Improved Speech Quality. Proceedings of the 2006 IEEE International Symposium on Signal Processing and Information Technology, Vancouer, BC, Canada.","DOI":"10.1109\/ISSPIT.2006.270839"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1321","DOI":"10.1121\/1.4976051","article-title":"In-ear microphone speech quality enhancement via adaptive filtering and artificial bandwidth extension","volume":"1413","author":"Bouserhal","year":"2017","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Rahman, M.S., and Shimamura, T. (2011, January 7\u201310). Intelligibility enhancement of bone conducted speech by an analysis-synthesis method. Proceedings of the 2011 IEEE 54th International Midwest Symposium on Circuits and Systems (MWSCAS), Seoul, Republic of Korea.","DOI":"10.1109\/MWSCAS.2011.6026374"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1877","DOI":"10.1587\/transinf.2015EDP7457","article-title":"WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications","volume":"99","author":"Morise","year":"2016","journal-title":"IEICE Trans. Inf. Syst."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhang, M., Sisman, B., Rallabandi, S.S., Li, H., and Zhao, L. (2018, January 12\u201315). Error Reduction Network for DBLSTM-based Voice Conversion. Proceedings of the 2018 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Honolulu, HI, USA.","DOI":"10.23919\/APSIPA.2018.8659543"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Liu, K., Zhang, J., and Yan, Y. (2007, January 24\u201327). High Quality Voice Conversion through Phoneme-Based Linear Mapping Functions with STRAIGHT for Mandarin. Proceedings of the Fourth International Conference on Fuzzy Systems and Knowledge Discovery (FSKD 2007), Haikou, China.","DOI":"10.1109\/FSKD.2007.347"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Ohtani, Y., Toda, T., Saruwatari, H., and Shikano, K. (2006, January 17\u201321). Maximum likelihood voice conversion based on GMM with STRAIGHT mixed excitation. Proceedings of the 9th International Conference on Spoken Language Processing (ICSLP), Pittsburgh, PA, USA.","DOI":"10.21437\/Interspeech.2006-582"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Mao, X., Li, Q., Xie, H., Lau, R.Y., Wang, Z., and Smolley, S.P. (2017, January 22\u201329). Least Squares Generative Adversarial Networks. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.304"},{"key":"ref_29","unstructured":"Radford, A., Metz, L., and Chintala, S. (2015). Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks. arXiv."},{"key":"ref_30","unstructured":"Dauphin, Y., Fan, A., Auli, M., and Grangier, D. (2016, January 19\u201324). Language Modeling with Gated Convolutional Networks. Proceedings of the International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_31","unstructured":"Ulyanov, D., Vedaldi, A., and Lempitsky, V.S. (2016). Instance Normalization: The Missing Ingredient for Fast Stylization. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Shi, W., Caballero, J., Husz\u00e1r, F., Totz, J., Aitken, A.P., Bishop, R., Rueckert, D., and Wang, Z. (2016, January 27\u201330). Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.207"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"62638","DOI":"10.1109\/ACCESS.2018.2873728","article-title":"A Novel Encoder-Decoder Model via NS-LSTM Used for Bone-Conducted Speech Enhancement","volume":"6","author":"Shan","year":"2018","journal-title":"IEEE Access"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"2505","DOI":"10.1109\/TASL.2012.2205241","article-title":"Statistical Voice Conversion Techniques for Body-Conducted Unvoiced Speech Enhancement","volume":"20","author":"Toda","year":"2012","journal-title":"IEEE Trans. Audio, Speech, Lang. Process."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Zheng, C., Zhang, X., Sun, M., Yang, J., and Xing, Y. (2018, January 7\u201310). A Novel Throat Microphone Speech Enhancement Framework Based on Deep BLSTM Recurrent Neural Networks. Proceedings of the 2018 IEEE 4th International Conference on Computer and Communications (ICCC), Chengdu, China.","DOI":"10.1109\/CompComm.2018.8780872"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Kameoka, H., Kaneko, T., Tanaka, K., and Hojo, N. (2018, January 18\u201321). StarGAN-VC: Non-parallel many-to-many Voice Conversion Using Star Generative Adversarial Networks. Proceedings of the 2018 IEEE Spoken Language Technology Workshop (SLT), Athens, Greece.","DOI":"10.1109\/SLT.2018.8639535"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Taal, C.H., Hendriks, R.C., Heusdens, R., and Jensen, J.R. (2010, January 14\u201319). A short-time objective intelligibility measure for time-frequency weighted noisy speech. Proceedings of the 2010 IEEE International Conference on Acoustics, Speech and Signal Processing, Dallas, TX, USA.","DOI":"10.1109\/ICASSP.2010.5495701"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"1924","DOI":"10.1109\/TASL.2006.883177","article-title":"P.563\u2014The ITU-T Standard for Single-Ended Speech Quality Assessment","volume":"14","author":"Malfait","year":"2006","journal-title":"IEEE Trans. Audio, Speech, Lang. Process."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"380","DOI":"10.1109\/TASSP.1976.1162849","article-title":"Distance measures for speech processing","volume":"24","author":"Gray","year":"1976","journal-title":"IEEE Trans. Acoust. Speech Signal Process."},{"key":"ref_40","unstructured":"Kraljevski, I., Chungurski, S., Stojanovic, I., and Arsenovski, S. (2010, January 23\u201325). Synthesized Speech Quality Evaluation Using ITU-T P.563. Proceedings of the 18th Telecommunications Forum 2010, Belgrade, Serbia."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"213","DOI":"10.1007\/s00530-014-0446-1","article-title":"Mean opinion score (MOS) revisited: Methods and applications, limitations and alternatives","volume":"22","author":"Streijl","year":"2016","journal-title":"Multimedia Systems"}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/16\/3\/153\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:51:19Z","timestamp":1760122279000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/16\/3\/153"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,9]]},"references-count":41,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["a16030153"],"URL":"https:\/\/doi.org\/10.3390\/a16030153","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,9]]}}}