{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,23]],"date-time":"2026-01-23T09:55:39Z","timestamp":1769162139677,"version":"3.49.0"},"reference-count":32,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2020,1,9]],"date-time":"2020-01-09T00:00:00Z","timestamp":1578528000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"the National Key R&D Program of China","award":["2017YFB1002202"],"award-info":[{"award-number":["2017YFB1002202"]}]},{"DOI":"10.13039\/501100001809","name":"the National Nature Science Foundation of China","doi-asserted-by":"crossref","award":["61871358, U1636201"],"award-info":[{"award-number":["61871358, U1636201"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"the Key Science and Technology Project of Anhui Province","award":["17030901005"],"award-info":[{"award-number":["17030901005"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2020,5,31]]},"abstract":"<jats:p>A method of learning and modeling unit embeddings using deep neutral networks (DNNs) is presented in this article for unit-selection-based Mandarin speech synthesis. Here, a unit embedding is defined as a fixed-length embedding vector for a phone-sized unit candidate in a corpus. Modeling phone-sized embedding vectors instead of frame-sized acoustic features can better measure the long-term dependencies among consecutive units in an utterance. First, a DNN with an embedding layer is built to learn the embedding vectors of all unit candidates in the corpus from scratch. In order to enable the extracted embedding vectors to carry both acoustic and linguistic information of unit candidates, a multitarget learning strategy is designed for the DNN. Its optional prediction targets include frame-level acoustic features, unit durations, monophone and tone identifiers, and context classes. Then, another two DNNs are constructed to map linguistic features toward the extracted embedding vectors. One of them employs the unit vectors of preceding phones besides the linguistic features of current phone as its input. At synthesis time, the distances between the unit vectors predicted by these two DNNs and the ones derived from unit candidates are used as a part of the target cost and a part of the concatenation cost, respectively. Our experiments on a Mandarin speech synthesis corpus demonstrate that learning and modeling unit embeddings improve the naturalness of hidden Markov model (HMM)-based unit selection speech synthesis. Furthermore, integrating multiple targets for learning unit embeddings achieves better performance than using only acoustic targets according to our subjective evaluation results.<\/jats:p>","DOI":"10.1145\/3372244","type":"journal-article","created":{"date-parts":[[2020,4,4]],"date-time":"2020-04-04T03:08:03Z","timestamp":1585969683000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Learning and Modeling Unit Embeddings Using Deep Neural Networks for Unit-Selection-Based Mandarin Speech Synthesis"],"prefix":"10.1145","volume":"19","author":[{"given":"Xiao","family":"Zhou","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, Hefei, Anhui, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhen-Hua","family":"Ling","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, Anhui, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Li-Rong","family":"Dai","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, Anhui, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,1,9]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Black and Nick Campbell","author":"Alan","year":"1995"},{"key":"e_1_2_1_2_1","volume-title":"MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems. arXiv preprint arXiv:1512.01274","author":"Chen Tianqi","year":"2015"},{"key":"e_1_2_1_3_1","volume-title":"9th ISCA Speech Synthesis Workshop (SSW9 \u201916)","author":"Den Oord Aaron Van","year":"2016"},{"key":"e_1_2_1_4_1","volume-title":"IEEE International Conference on Acoustics, Speech, and Signal Processing, 1996 (ICASSP \u201996), Conference Proceedings.","volume":"1","author":"Andrew"},{"key":"e_1_2_1_5_1","volume-title":"Blizzard Challenge Workshop.","author":"Jiang Y.","year":"2018"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-6393(98)00085-5"},{"key":"e_1_2_1_7_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2014"},{"key":"e_1_2_1_8_1","first-page":"301","article-title":"The relationship between light tone and syntactic structure of modern Chinese","volume":"7","author":"Lin Tao","year":"1962","journal-title":"Studies of the Chinese Language"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2014.2359987"},{"key":"e_1_2_1_10_1","volume-title":"9th International Conference on Spoken Language Processing.","author":"Ling Zhen-Hua","year":"2006"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2007.367302"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11265-018-1336-0"},{"key":"e_1_2_1_13_1","volume-title":"Blizzard Challenge Workshop.","author":"Liu Li-Juan"},{"key":"e_1_2_1_14_1","first-page":"2579","article-title":"Visualizing data using t-SNE","author":"van der Maaten Laurens","year":"2008","journal-title":"Journal of Machine Learning Research 9"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7472658"},{"key":"e_1_2_1_16_1","volume-title":"Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781","author":"Mikolov Tomas","year":"2013"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-00810-9_3"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-428"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2012.2221460"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1978.1163055"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/1367985.1367993"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461368"},{"key":"e_1_2_1_23_1","first-page":"79","article-title":"MDL-based context-dependent subword modeling for speech recognition","volume":"21","author":"Shinoda Koichi","year":"2000","journal-title":"Acoustical Science and Technology"},{"key":"e_1_2_1_24_1","doi-asserted-by":"crossref","unstructured":"Akira Tamamori Tomoki Hayashi Kazuhiro Kobayashi Kazuya Takeda and Tomoki Toda. 2017. Speaker-dependent WaveNet vocoder. In Interspeech. 1118--1122.  Akira Tamamori Tomoki Hayashi Kazuhiro Kobayashi Kazuya Takeda and Tomoki Toda. 2017. Speaker-dependent WaveNet vocoder. In Interspeech. 1118--1122.","DOI":"10.21437\/Interspeech.2017-314"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2000.861820"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-1107"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178814"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2014.04.002"},{"key":"e_1_2_1_29_1","doi-asserted-by":"crossref","unstructured":"T. Yoshimura K. Tokuda T. Masuko T. Kobayashi and T. Kitamura. 1999. Simultaneous modeling of spectrum pitch and duration in HMM-based speech synthesis. In EUROSPEECH. 2347--2350.  T. Yoshimura K. Tokuda T. Masuko T. Kobayashi and T. Kitamura. 1999. Simultaneous modeling of spectrum pitch and duration in HMM-based speech synthesis. In EUROSPEECH. 2347--2350.","DOI":"10.21437\/Eurospeech.1999-513"},{"key":"e_1_2_1_30_1","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201913)","author":"Zen Heiga","year":"2013"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2009.04.004"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-1198"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3372244","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3372244","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:32:10Z","timestamp":1750195930000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3372244"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,1,9]]},"references-count":32,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2020,5,31]]}},"alternative-id":["10.1145\/3372244"],"URL":"https:\/\/doi.org\/10.1145\/3372244","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,1,9]]},"assertion":[{"value":"2019-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-01-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}