{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,2]],"date-time":"2026-01-02T07:37:09Z","timestamp":1767339429987,"version":"build-2065373602"},"reference-count":31,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2022,2,20]],"date-time":"2022-02-20T00:00:00Z","timestamp":1645315200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>A conversion method based on the inversion of Mel frequency cepstral coefficient (MFCC) features was proposed to convert whispered speech into normal speech. First, the MFCC features of whispered speech and normal speech were extracted and a matching relation between the MFCC feature parameters of whispered speech and normal speech was developed through the Gaussian mixture model (GMM). Then, the MFCC feature parameters of normal speech corresponding to whispered speech were obtained based on the GMM and, finally, whispered speech was converted into normal speech through the inversion of MFCC features. The experimental results showed that the cepstral distortion (CD) of the normal speech converted by the proposed method was 21% less than that of the normal speech converted by the linear predictive coefficient (LPC) features, the mean opinion score (MOS) was 3.56, and a satisfactory outcome in both intelligibility and sound quality was achieved.<\/jats:p>","DOI":"10.3390\/a15020068","type":"journal-article","created":{"date-parts":[[2022,2,21]],"date-time":"2022-02-21T08:14:46Z","timestamp":1645431286000},"page":"68","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Whispered Speech Conversion Based on the Inversion of Mel Frequency Cepstral Coefficient Features"],"prefix":"10.3390","volume":"15","author":[{"given":"Qiang","family":"Zhu","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Hefei Normal University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7987-4883","authenticated-orcid":false,"given":"Zhong","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Hefei Normal University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yunfeng","family":"Dou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Anhui University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jian","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Anhui University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,2,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"2313","DOI":"10.1109\/TASLP.2017.2738559","article-title":"Whispered speech recognition using deep denoising autoencoder and inverse filtering","volume":"25","year":"2017","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_2","first-page":"5235","article-title":"Recognizing emotions from whispered speech based on acoustic feature transfer learning","volume":"5","author":"Deng","year":"2017","journal-title":"IEEE Access"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1191","DOI":"10.1109\/TASE.2015.2467311","article-title":"Speaker identification with whispered speech for the access control system","volume":"12","author":"Wang","year":"2015","journal-title":"IEEE Trans. Autom. Sci. Eng."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"927","DOI":"10.1109\/TASLP.2021.3053388","article-title":"Analysis and calibration of Lombard effect and whisper for speaker recognition","volume":"29","author":"Kelly","year":"2021","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Raeesy, Z., Gillespie, K., Ma, C., Drugman, T., Gu, J., Maas, R., Rastrow, A., and Hoffmeister, B. (2018, January 18\u201321). Lstm-based whisper detection. Proceedings of the 2018 IEEE Spoken Language Technology Workshop (SLT), Athens, Greece.","DOI":"10.1109\/SLT.2018.8639614"},{"key":"ref_6","first-page":"11","article-title":"Whispered speech recognition using hidden markov models and support vector machines","volume":"15","year":"2018","journal-title":"Acta Polytech. Hung."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"4299","DOI":"10.1109\/ACCESS.2016.2591442","article-title":"Exploitation of phase-based features for whispered speech emotion recognition","volume":"4","author":"Deng","year":"2016","journal-title":"IEEE Access"},{"key":"ref_8","first-page":"1047","article-title":"Timbre features for speaker identification of whispering speech: Selection of optimal audio descriptors","volume":"43","author":"Sardar","year":"2021","journal-title":"Int. J. Comput. Appl."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"4002","DOI":"10.1121\/10.0002952","article-title":"Acoustic differences between voiced and whispered speech in gender diverse speakers","volume":"148","author":"Houle","year":"2020","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"134","DOI":"10.1016\/j.specom.2011.07.007","article-title":"Speaking-aid systems using GMM-based voice conversion for electrolaryngeal speech","volume":"54","author":"Nakamura","year":"2012","journal-title":"Speech Commun."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"130495","DOI":"10.1109\/ACCESS.2019.2940700","article-title":"Whisper to normal speech conversion using sequence-to-sequence mapping model with auditory attention","volume":"7","author":"Lian","year":"2019","journal-title":"IEEE Access"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Huang, C., Tao, X.Y., Tao, L., Zhou, J., and Bin Wang, H. (2012, January 14\u201317). Reconstruction of whisper in Chinese by modified MELP. Proceedings of the 2012 7th International Conference on Computer Science & Education (ICCSE), Melbourne, VIC, Australia.","DOI":"10.1109\/ICCSE.2012.6295089"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Li, J., McLoughlin, I.V., and Song, Y. (2014, January 12\u201314). Reconstruction of pitch for whisper-to-speech conversion of Chinese. Proceedings of the 9th International Symposium on Chinese Spoken Language Processing, Singapore.","DOI":"10.1109\/ISCSLP.2014.6936709"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1016\/j.jvoice.2006.08.012","article-title":"Acoustic analysis of consonants in whispered speech","volume":"22","year":"2008","journal-title":"J. Voice"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"889","DOI":"10.1109\/TASLP.2020.2971417","article-title":"Glottal flow synthesis for whisper-to-speech conversion","volume":"28","author":"Perrotin","year":"2020","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_16","unstructured":"Sharifzadeh, H.R., Mcloughlin, I.V., and Ahmadi, F. (2014, January 12\u201314). Regeneration of Speech in Speech-Loss Patients. Proceedings of the 13th International Conference on Biomedical Engineering, Singapore."},{"key":"ref_17","first-page":"44","article-title":"Study on the conversion of Chinese whispered speech into normal speech","volume":"12","author":"Fan","year":"2005","journal-title":"Audio Eng."},{"key":"ref_18","first-page":"69","article-title":"Phonological segmentation of whispered speech based on the entropy function","volume":"1","author":"Li","year":"2005","journal-title":"Acta Acustica"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"2222","DOI":"10.1109\/TASL.2007.907344","article-title":"Speech Conversion Based on Maximum-Likelihood Estimation of Spectral Parameter Trajectory","volume":"15","author":"Toda","year":"2007","journal-title":"Audio Speech Lang. Process. IEEE Trans."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Janke, M., Wand, M., Heistermann, T., Schultz, T., and Prahallad, K. (2014, January 4\u20139). Fundamental frequency generation for whisper-to-audible speech conversion. Proceedings of the 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6854066"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1097","DOI":"10.1109\/TASL.2008.2001109","article-title":"Speaker identification using instantaneous frequencies","volume":"16","author":"Grimaldi","year":"2008","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"2505","DOI":"10.1109\/TASL.2012.2205241","article-title":"Statistical voice conversion techniques for body-conducted unvoiced speech enhancement","volume":"20","author":"Toda","year":"2012","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1781","DOI":"10.1049\/el.2014.1645","article-title":"Whisper-to-speech conversion using restricted Boltzmann machine arrays","volume":"50","author":"Li","year":"2014","journal-title":"Electron. Lett."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Chen, X., Yu, Y., and Zhao, H. (2014, January 23\u201325). F0 prediction from linear predictive cepstral coefficient. Proceedings of the 2014 Sixth International Conference on Wireless Communications and Signal Processing (WCSP), Hefei, China.","DOI":"10.1109\/WCSP.2014.6992061"},{"key":"ref_25","first-page":"610","article-title":"Low bit-rate speech coding through quantization of mel-frequency cepstral coefficients","volume":"20","author":"Boucheron","year":"2011","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_26","unstructured":"Wenbin, J., Rendong, Y., and Peilin, L. (2014, January 14\u201318). Speech reconstruction for MFCC-based low bit-rate speech coding. Proceedings of the 2014 IEEE International Conference on Multimedia and Expo Workshops (ICMEW), Chengdu, China."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1007\/s10772-014-9257-1","article-title":"Recognition of isolated words using Zernike and MFCC features for audio visual speech recognition","volume":"18","author":"Borde","year":"2015","journal-title":"Int. J. Speech Technol."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1013","DOI":"10.1109\/TNNLS.2012.2197412","article-title":"L1\/2 regularization: A thresholding representation theory and a fast solver","volume":"23","author":"Xu","year":"2012","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"974","DOI":"10.1109\/TCBB.2017.2665557","article-title":"Regularized non-negative matrix factorization for identifying differentially expressed genes and clustering samples: A survey","volume":"15","author":"Liu","year":"2017","journal-title":"IEEE\/ACM Trans. Comput. Biol. Bioinform."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"45","DOI":"10.1109\/MCAS.2016.2583681","article-title":"Recent developments in speech enhancement in the short-time Fourier transform domain","volume":"16","author":"Parchami","year":"2016","journal-title":"IEEE Circuits Syst. Mag."},{"key":"ref_31","first-page":"329","article-title":"A linear prediction algorithm in low bit rate speech coding improved by multi-band excitation model","volume":"26","author":"Yang","year":"2001","journal-title":"Acta Acust."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/15\/2\/68\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:23:31Z","timestamp":1760135011000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/15\/2\/68"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,20]]},"references-count":31,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2022,2]]}},"alternative-id":["a15020068"],"URL":"https:\/\/doi.org\/10.3390\/a15020068","relation":{},"ISSN":["1999-4893"],"issn-type":[{"type":"electronic","value":"1999-4893"}],"subject":[],"published":{"date-parts":[[2022,2,20]]}}}