{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,24]],"date-time":"2026-02-24T19:30:46Z","timestamp":1771961446316,"version":"3.50.1"},"reference-count":34,"publisher":"MDPI AG","issue":"16","license":[{"start":{"date-parts":[[2023,8,20]],"date-time":"2023-08-20T00:00:00Z","timestamp":1692489600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"China State Shipbuilding Corporation (CSSC) Guangxi Shipbuilding and Offshore Engineering Technology Collaboration Project","award":["ZCGXJSB20226300222-06"],"award-info":[{"award-number":["ZCGXJSB20226300222-06"]}]},{"name":"China State Shipbuilding Corporation (CSSC) Guangxi Shipbuilding and Offshore Engineering Technology Collaboration Project","award":["2018"],"award-info":[{"award-number":["2018"]}]},{"name":"Guangxi Zhuang Autonomous Region of China","award":["ZCGXJSB20226300222-06"],"award-info":[{"award-number":["ZCGXJSB20226300222-06"]}]},{"name":"Guangxi Zhuang Autonomous Region of China","award":["2018"],"award-info":[{"award-number":["2018"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In recent years, deep learning-based speech synthesis has attracted a lot of attention from the machine learning and speech communities. In this paper, we propose Mixture-TTS, a non-autoregressive speech synthesis model based on mixture alignment mechanism. Mixture-TTS aims to optimize the alignment information between text sequences and mel-spectrogram. Mixture-TTS uses a linguistic encoder based on soft phoneme-level alignment and hard word-level alignment approaches, which explicitly extract word-level semantic information, and introduce pitch and energy predictors to optimally predict the rhythmic information of the audio. Specifically, Mixture-TTS introduces a post-net based on a five-layer 1D convolution network to optimize the reconfiguration capability of the mel-spectrogram. We connect the output of the decoder to the post-net through the residual network. The mel-spectrogram is converted into the final audio by the HiFi-GAN vocoder. We evaluate the performance of the Mixture-TTS on the AISHELL3 and LJSpeech datasets. Experimental results show that Mixture-TTS is somewhat better in alignment information between the text sequences and mel-spectrogram, and is able to achieve high-quality audio. The ablation studies demonstrate that the structure of Mixture-TTS is effective.<\/jats:p>","DOI":"10.3390\/s23167283","type":"journal-article","created":{"date-parts":[[2023,8,21]],"date-time":"2023-08-21T01:49:34Z","timestamp":1692582574000},"page":"7283","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Research on Speech Synthesis Based on Mixture Alignment Mechanism"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0778-6144","authenticated-orcid":false,"given":"Yan","family":"Deng","sequence":"first","affiliation":[{"name":"School of Computer, Electronics and Information, Guangxi University, Nanning 530004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4951-6337","authenticated-orcid":false,"given":"Ning","family":"Wu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Beibu Gulf Offshore Engineering Equipment and Technology, Beibu Gulf University, Qinzhou 535011, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengjun","family":"Qiu","sequence":"additional","affiliation":[{"name":"College of Mechanical Naval Architecture and Ocean Engineering, Beibu Gulf University, Qinzhou 535011, China"},{"name":"Guangxi Key Laboratory of Ocean Engineering Equipment and Technology, Qinzhou 535011, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9950-684X","authenticated-orcid":false,"given":"Yan","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer, Electronics and Information, Guangxi University, Nanning 530004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xueshan","family":"Gao","sequence":"additional","affiliation":[{"name":"College of Mechanical Naval Architecture and Ocean Engineering, Beibu Gulf University, Qinzhou 535011, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,8,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Yu, J., Xu, Z., He, X., Wang, J., Liu, B., Feng, R., Zhu, S., Wang, W., and Li, J. (2023). DIA-TTS: Deep-Inherited Attention-Based Text-to-Speech Synthesizer. Entropy, 25.","DOI":"10.2139\/ssrn.4257520"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"57674","DOI":"10.1109\/ACCESS.2023.3283772","article-title":"MixGAN-TTS: Efficient and Stable Speech Synthesis Based on Diffusion Model","volume":"11","author":"Deng","year":"2023","journal-title":"IEEE Access"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"30929","DOI":"10.1109\/ACCESS.2023.3260844","article-title":"Advancements in Arabic Text-to-Speech Systems: A 22-Year Literature Review","volume":"11","author":"Chemnad","year":"2023","journal-title":"IEEE Access"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Kim, Y., Kim, J., Hong, J., and Seok, J. (2023). The Tacotron-Based Signal Synthesis Method for Active Sonar. Sensors, 23.","DOI":"10.3390\/s23010028"},{"key":"ref_5","unstructured":"Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. (2016). Wavenet: A generative model for raw audio. arXiv."},{"key":"ref_6","first-page":"17022","article-title":"Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis","volume":"33","author":"Kong","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Prenger, R., Valle, R., and Catanzaro, B. (2019, January 12\u201317). Waveglow: A Flow-Based Generative Network for Speech Synthesis. Proceedings of the ICASSP 2019\u20142019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683143"},{"key":"ref_8","unstructured":"Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.Y. (2020). Fastspeech 2: Fast and high-quality end-to-end text to speech. arXiv."},{"key":"ref_9","unstructured":"Kim, J., Kong, J., and Son, J. (2021, January 18\u201324). Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Lim, D., Jung, S., and Kim, E. (2022). JETS: Jointly training FastSpeech2 and HiFi-GAN for end to end text to speech. arXiv.","DOI":"10.21437\/Interspeech.2022-10294"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R.J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., and Bengio, S. (2017). Tacotron: Towards end-to-end speech synthesis. arXiv.","DOI":"10.21437\/Interspeech.2017-1452"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Shen, J., Pang, R., Weiss, R.J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., and Skerrv-Ryan, R. (2018, January 15\u201320). Natural tts Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461368"},{"key":"ref_13","unstructured":"Li, N., Liu, S., Liu, Y., Zhao, S., and Liu, M. (February, January 27). Neural Speech Synthesis with Transformer Network. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_14","first-page":"3165","article-title":"FastSpeech: Fast, robust and controllable text to speech","volume":"32","author":"Ren","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_15","first-page":"8067","article-title":"Glow-tts: A generative flow for text-to-speech via monotonic alignment search","volume":"33","author":"Kim","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Miao, C., Liang, S., Chen, M., Ma, J., Wang, S., and Xiao, J. (2020, January 4\u20138). Flow-tts: A Non-Autoregressive Network for Text to Speech Based on Flow. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9054484"},{"key":"ref_17","first-page":"963","article-title":"PortaSpeech: Portable and high-quality generative text-to-speech","volume":"34","author":"Ren","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_18","unstructured":"Lee, Y., Shin, J., and Jung, K. (2022, January 29). Bidirectional variational inference for non-autoregressive text-to-speech. Proceedings of the International Conference on Learning Representations 2022, Online meeting."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yang, J., Bae, J.-S., Bak, T., Kim, Y., and Cho, H.-Y. (2021). Ganspeech: Adversarial training for high-fidelity multi-speaker speech synthesis. arXiv.","DOI":"10.21437\/Interspeech.2021-971"},{"key":"ref_20","unstructured":"Chen, N., Zhang, Y., Zen, H., Weiss, R.J., Norouzi, M., and Chan, W. (2020). Wavegrad: Estimating gradients for waveform generation. arXiv."},{"key":"ref_21","unstructured":"Popov, V., Vovk, I., Gogoryan, V., Sadekova, T., and Kudinov, M. (2021, January 18\u201324). Gradtts: A Diffusion Probabilistic Model for Text-to-Speech. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_22","unstructured":"Liu, J., Li, C., Ren, Y., Chen, F., and Zhao, Z. (March, January 22). Diffsinger: Singing Voice Synthesis via Shallow Diffusion Mechanism. Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada."},{"key":"ref_23","unstructured":"Liu, S., Su, D., and Yu, D. (2022). DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Shi, Y., Bu, H., Xu, X., Zhang, S., and Li, M. (2020). AISHELL-3: A multispeaker Mandarin TTS corpus and the baselines. arXiv.","DOI":"10.21437\/Interspeech.2021-755"},{"key":"ref_25","unstructured":"Ito, K., and Johnson, L. The LJ Speech Dataset, Available online: https:\/\/keithito.com\/LJ-Speech-Dataset\/."},{"key":"ref_26","first-page":"5998","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Tachibana, H., Uenoyama, K., and Aihara, S. (2018, January 15\u201320). Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461829"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Cui, Y., Che, W., Liu, T., Qin, B., Wang, S., and Hu, G. (2020). Revisiting pre-trained models for Chinese natural language processing. arXiv.","DOI":"10.18653\/v1\/2020.findings-emnlp.58"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","article-title":"Image quality assessment: From error visibility to structural similarity","volume":"13","author":"Wang","year":"2004","journal-title":"IEEE Trans. Image Process."},{"key":"ref_31","unstructured":"Kubichek, R. (1993, January 19\u201321). Mel-Cepstral Distance Measure for Objective Speech Quality Assessment. Proceedings of the IEEE Pacific Rim Conference on Communications Computers and Signal Processing, Victoria, BC, Canada."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"M\u00fcller, M. (2007). Information Retrieval for Music and Motion, Springer.","DOI":"10.1007\/978-3-540-74048-3"},{"key":"ref_33","unstructured":"Chu, M., and Peng, H. (2006). Objective Measure for Estimating Mean Opinion Score of Synthesized Speech. (7,024,362), U.S. Patent."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Loizou, P.C. (2011). Speech quality assessment. Chapter of Information Retrieval for Music and Motion. Multimed. Anal. Process. Commun., 623\u2013654.","DOI":"10.1007\/978-3-642-19551-8_23"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/16\/7283\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:37:57Z","timestamp":1760128677000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/16\/7283"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,20]]},"references-count":34,"journal-issue":{"issue":"16","published-online":{"date-parts":[[2023,8]]}},"alternative-id":["s23167283"],"URL":"https:\/\/doi.org\/10.3390\/s23167283","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,20]]}}}