{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T22:29:17Z","timestamp":1781216957933,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":61,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,18]],"date-time":"2021-10-18T00:00:00Z","timestamp":1634515200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Digital Futures","award":["AAIS"],"award-info":[{"award-number":["AAIS"]}]},{"DOI":"10.13039\/501100004472","name":"Riksbankens Jubileumsfond","doi-asserted-by":"publisher","award":["P20-0298 (CAPTivating)"],"award-info":[{"award-number":["P20-0298 (CAPTivating)"]}],"id":[{"id":"10.13039\/501100004472","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Swedish Research Council","award":["2019-05003 (Connected), 2018-05409 (StyleBot)"],"award-info":[{"award-number":["2019-05003 (Connected), 2018-05409 (StyleBot)"]}]},{"name":"Wallenberg AI, Autonomous Systems and Software Program (WASP)"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,18]]},"DOI":"10.1145\/3462244.3479914","type":"proceedings-article","created":{"date-parts":[[2021,10,15]],"date-time":"2021-10-15T15:01:58Z","timestamp":1634310118000},"page":"177-185","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":16,"title":["Integrated Speech and Gesture Synthesis"],"prefix":"10.1145","author":[{"given":"Siyang","family":"Wang","sequence":"first","affiliation":[{"name":"KTH Royal Institute of Technology, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Simon","family":"Alexanderson","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Joakim","family":"Gustafson","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jonas","family":"Beskow","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gustav Eje","family":"Henter","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"\u00c9va","family":"Sz\u00e9kely","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,10,18]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.170"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.13946"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3383652.3423874"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9054106"},{"key":"e_1_3_2_2_5_1","unstructured":"Samy Bengio Oriol Vinyals Navdeep Jaitly and Noam Shazeer. 2015. Scheduled sampling for sequence prediction with recurrent neural networks. arXiv preprint arXiv:1506.03099(2015).  Samy Bengio Oriol Vinyals Navdeep Jaitly and Noam Shazeer. 2015. Scheduled sampling for sequence prediction with recurrent neural networks. arXiv preprint arXiv:1506.03099(2015)."},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-40415-3_12"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/192161.192272"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/383259.383315"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-21996-7_17"},{"key":"e_1_3_2_2_11_1","volume-title":"Proc. IVA. 93\u201398","unstructured":"Ylva. Ferstl and Rachel. McDonnell. 2018. Investigating the use of recurrent motion modelling for speech gesture generation . In Proc. IVA. 93\u201398 . https:\/\/trinityspeechgesture.scss.tcd.ie Ylva. Ferstl and Rachel. McDonnell. 2018. Investigating the use of recurrent motion modelling for speech gesture generation. In Proc. IVA. 93\u201398. https:\/\/trinityspeechgesture.scss.tcd.ie"},{"key":"e_1_3_2_2_12_1","first-page":"1","article-title":"Multi-objective adversarial gesture generation","volume":"3","author":"Ferstl Ylva","year":"2019","unstructured":"Ylva Ferstl , Michael Neff , and Rachel McDonnell . 2019 . Multi-objective adversarial gesture generation . In Proc. MIG. 3 : 1 \u2013 3 :10. Ylva Ferstl, Michael Neff, and Rachel McDonnell. 2019. Multi-objective adversarial gesture generation. In Proc. MIG. 3:1\u20133:10.","journal-title":"Proc. MIG."},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cag.2020.04.007"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00361"},{"key":"e_1_3_2_2_15_1","volume-title":"Proc. NIPS. 2672\u20132680","author":"Goodfellow J.","year":"2014","unstructured":"Ian\u00a0 J. Goodfellow , Jean Pouget-Abadie , Mehdi Mirza , Bing Xu , David Warde-Farley , Sherjil Ozair , Aaron Courville , and Yoshua Bengio . 2014 . Generative adversarial networks . In Proc. NIPS. 2672\u20132680 . Ian\u00a0J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial networks. In Proc. NIPS. 2672\u20132680."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1080\/10867651.1998.10487493"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"crossref","unstructured":"Ikhsanul Habibie Weipeng Xu Dushyant Mehta Lingjie Liu Hans-Peter Seidel Gerard Pons-Moll Mohamed Elgharib and Christian Theobalt. 2021. Learning Speech-driven 3D Conversational Gestures from Video. arXiv preprint arXiv:2102.06837(2021).  Ikhsanul Habibie Weipeng Xu Dushyant Mehta Lingjie Liu Hans-Peter Seidel Gerard Pons-Moll Mohamed Elgharib and Christian Theobalt. 2021. Learning Speech-driven 3D Conversational Gestures from Video. arXiv preprint arXiv:2102.06837(2021).","DOI":"10.1145\/3472306.3478335"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3267851.3267878"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417836"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2018.2856281"},{"key":"e_1_3_2_2_21_1","unstructured":"ITU-R BS.1534-3. 2015. Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems. Standard. ITU. https:\/\/www.itu.int\/rec\/R-REC-BS.1534-3-201510-I  ITU-R BS.1534-3. 2015. Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems. Standard. ITU. https:\/\/www.itu.int\/rec\/R-REC-BS.1534-3-201510-I"},{"key":"e_1_3_2_2_22_1","unstructured":"ITU-T P.800. 1996. Methods for Subjective Determination of Transmission Quality. Standard. ITU. https:\/\/www.itu.int\/rec\/T-REC-P.800-199608-I  ITU-T P.800. 1996. Methods for Subjective Determination of Transmission Quality. Standard. ITU. https:\/\/www.itu.int\/rec\/T-REC-P.800-199608-I"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3462244.3479957"},{"key":"e_1_3_2_2_24_1","volume-title":"Proc. NeurIPS. 8067\u20138077","author":"Kim Jaehyeon","year":"2020","unstructured":"Jaehyeon Kim , Sungwon Kim , Jungil Kong , and Sungroh Yoon . 2020 . Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search . In Proc. NeurIPS. 8067\u20138077 . Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon. 2020. Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search. In Proc. NeurIPS. 8067\u20138077."},{"key":"e_1_3_2_2_25_1","volume-title":"Proc. ICML. 3370\u20133378","author":"Kim Sungwon","year":"2019","unstructured":"Sungwon Kim , Sang-Gil Lee , Jongyoon Song , Jaehyeon Kim , and Sungroh Yoon . 2019 . FloWaveNet: A generative flow for raw audio . In Proc. ICML. 3370\u20133378 . Sungwon Kim, Sang-Gil Lee, Jongyoon Song, Jaehyeon Kim, and Sungroh Yoon. 2019. FloWaveNet: A generative flow for raw audio. In Proc. ICML. 3370\u20133378."},{"key":"e_1_3_2_2_26_1","volume-title":"Kingma and Prafulla Dhariwal","author":"P.","year":"2018","unstructured":"Durk\u00a0 P. Kingma and Prafulla Dhariwal . 2018 . Glow : Generative flow with invertible 1x1 convolutions. In Proc. NeurIPS. 10236\u201310245. Durk\u00a0P. Kingma and Prafulla Dhariwal. 2018. Glow: Generative flow with invertible 1x1 convolutions. In Proc. NeurIPS. 10236\u201310245."},{"key":"e_1_3_2_2_27_1","volume-title":"Gesture generation by imitation: From human behavior to computer character animation","author":"Kipp Michael","unstructured":"Michael Kipp . 2005. Gesture generation by imitation: From human behavior to computer character animation . Universal-Publishers . Michael Kipp. 2005. Gesture generation by imitation: From human behavior to computer character animation. Universal-Publishers."},{"key":"e_1_3_2_2_28_1","volume-title":"Proc. NeurIPS. 17022\u201317033","author":"Kong Jungil","year":"2020","unstructured":"Jungil Kong , Jaehyeon Kim , and Jaekyoung Bae . 2020 . HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis . In Proc. NeurIPS. 17022\u201317033 . Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis. In Proc. NeurIPS. 17022\u201317033."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/1071195.1071199"},{"key":"e_1_3_2_2_30_1","volume-title":"Proc. GENEA Workshop. https:\/\/doi.org\/10","author":"Korzun Vladislav","year":"2020","unstructured":"Vladislav Korzun , Ilya Dimov , and Andrey Zharkov . 2020 . The FineMotion entry to the GENEA Challenge 2020 . In Proc. GENEA Workshop. https:\/\/doi.org\/10 .5281\/zenodo.4088608 Vladislav Korzun, Ilya Dimov, and Andrey Zharkov. 2020. The FineMotion entry to the GENEA Challenge 2020. In Proc. GENEA Workshop. https:\/\/doi.org\/10.5281\/zenodo.4088608"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1080\/10447318.2021.1883883"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3382507.3418815"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397481.3450692"},{"key":"e_1_3_2_2_34_1","volume-title":"Multimodal analysis of the predictability of hand-gesture properties. arXiv preprint","author":"Kucherenko Taras","year":"2021","unstructured":"Taras Kucherenko , Rajmund Nagy , Michael Neff , Hedvig Kjellstr\u00f6m , and Gustav\u00a0Eje Henter . 2021. Multimodal analysis of the predictability of hand-gesture properties. arXiv preprint ( 2021 ). Taras Kucherenko, Rajmund Nagy, Michael Neff, Hedvig Kjellstr\u00f6m, and Gustav\u00a0Eje Henter. 2021. Multimodal analysis of the predictability of hand-gesture properties. arXiv preprint (2021)."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414660"},{"key":"e_1_3_2_2_36_1","volume-title":"Proc","author":"Luo Pengcheng","unstructured":"Pengcheng Luo , Victor Ng-Thow-Hing , and Michael Neff . 2013. An examination of whether people prefer agents whose gestures mimic their own . In Proc . IVA. Springer , 229\u2013238. Pengcheng Luo, Victor Ng-Thow-Hing, and Michael Neff. 2013. An examination of whether people prefer agents whose gestures mimic their own. In Proc. IVA. Springer, 229\u2013238."},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485895.2485900"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9054484"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/1330511.1330516"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2010.5654322"},{"key":"e_1_3_2_2_41_1","volume-title":"Proc. ICML. 7706\u20137716","author":"Ping Wei","year":"2020","unstructured":"Wei Ping , Kainan Peng , Kexin Zhao , and Zhao Song . 2020 . WaveFlow: A compact flow-based model for raw audio . In Proc. ICML. 7706\u20137716 . Wei Ping, Kainan Peng, Kexin Zhao, and Zhao Song. 2020. WaveFlow: A compact flow-based model for raw audio. In Proc. ICML. 7706\u20137716."},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683143"},{"key":"e_1_3_2_2_43_1","volume-title":"Language models are unsupervised multitask learners. OpenAI blog","author":"Radford Alec","year":"2019","unstructured":"Alec Radford , Jeffrey Wu , Rewon Child , David Luan , Dario Amodei , and Ilya Sutskever . 2019. Language models are unsupervised multitask learners. OpenAI blog ( 2019 ). https:\/\/cdn.openai.com\/better-language-models\/language_models_are_unsupervised_multitask_learners.pdf Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog (2019). https:\/\/cdn.openai.com\/better-language-models\/language_models_are_unsupervised_multitask_learners.pdf"},{"key":"e_1_3_2_2_44_1","volume-title":"Proc. Interspeech. 1586\u20131590","author":"Ribeiro Manuel\u00a0Sam","unstructured":"Manuel\u00a0Sam Ribeiro , Junichi Yamagishi , and Robert A . \u00a0J. Clark. 2015. A perceptual investigation of wavelet-based decomposition of f0 for text-to-speech synthesis . In Proc. Interspeech. 1586\u20131590 . Manuel\u00a0Sam Ribeiro, Junichi Yamagishi, and Robert A.\u00a0J. Clark. 2015. A perceptual investigation of wavelet-based decomposition of f0 for text-to-speech synthesis. In Proc. Interspeech. 1586\u20131590."},{"key":"e_1_3_2_2_45_1","volume-title":"Non-attentive Tacotron: Robust and controllable neural TTS synthesis including unsupervised duration modeling. arXiv preprint arXiv:2010.04301(2020).","author":"Shen Jonathan","year":"2020","unstructured":"Jonathan Shen , Ye Jia , Mike Chrzanowski , Yu Zhang , Isaac Elias , Heiga Zen , and Yonghui Wu . 2020 . Non-attentive Tacotron: Robust and controllable neural TTS synthesis including unsupervised duration modeling. arXiv preprint arXiv:2010.04301(2020). Jonathan Shen, Ye Jia, Mike Chrzanowski, Yu Zhang, Isaac Elias, Heiga Zen, and Yonghui Wu. 2020. Non-attentive Tacotron: Robust and controllable neural TTS synthesis including unsupervised duration modeling. arXiv preprint arXiv:2010.04301(2020)."},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461368"},{"key":"e_1_3_2_2_47_1","volume-title":"Proc. LREC. 6368\u20136374","author":"Sz\u00e9kely \u00c9va","year":"2020","unstructured":"\u00c9va Sz\u00e9kely , Jens Edlund , and Joakim Gustafson . 2020 . Augmented Prompt Selection for Evaluation of Spontaneous Speech Synthesis . In Proc. LREC. 6368\u20136374 . \u00c9va Sz\u00e9kely, Jens Edlund, and Joakim Gustafson. 2020. Augmented Prompt Selection for Evaluation of Spontaneous Speech Synthesis. In Proc. LREC. 6368\u20136374."},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-2836"},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9054107"},{"key":"e_1_3_2_2_50_1","volume-title":"Proc. SSW. 245\u2013250","author":"Henter Gustav\u00a0Eje","year":"2019","unstructured":"\u00c9va. Sz\u00e9kely, Gustav\u00a0Eje Henter , Jonas Beskow , and Joakim Gustafson . 2019 . How to train your fillers: uh and um in spontaneous speech synthesis . In Proc. SSW. 245\u2013250 . \u00c9va. Sz\u00e9kely, Gustav\u00a0Eje Henter, Jonas Beskow, and Joakim Gustafson. 2019. How to train your fillers: uh and um in spontaneous speech synthesis. In Proc. SSW. 245\u2013250."},{"key":"e_1_3_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683846"},{"key":"e_1_3_2_2_52_1","volume-title":"Proc. ICLR.","author":"Valle Rafael","year":"2021","unstructured":"Rafael Valle , Kevin\u00a0 J. Shih , Ryan Prenger , and Bryan Catanzaro . 2021 . Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis . In Proc. ICLR. Rafael Valle, Kevin\u00a0J. Shih, Ryan Prenger, and Bryan Catanzaro. 2021. Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis. In Proc. ICLR."},{"key":"e_1_3_2_2_53_1","unstructured":"A\u00e4ron van\u00a0den Oord Sander Dieleman Heiga Zen Karen Simonyan Oriol Vinyals Alex Graves Nal Kalchbrenner 2016. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499(2016).  A\u00e4ron van\u00a0den Oord Sander Dieleman Heiga Zen Karen Simonyan Oriol Vinyals Alex Graves Nal Kalchbrenner 2016. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499(2016)."},{"key":"e_1_3_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2013.09.008"},{"key":"e_1_3_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2016-134"},{"key":"e_1_3_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-1452"},{"key":"e_1_3_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.21437\/SSW.2019-39"},{"key":"e_1_3_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2014.19"},{"key":"e_1_3_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403331"},{"key":"e_1_3_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417838"},{"key":"e_1_3_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8793720"},{"key":"e_1_3_2_2_62_1","unstructured":"Chengzhu Yu Heng Lu Na Hu Meng Yu Chao Weng Kun Xu Peng Liu Deyi Tuo Shiyin Kang Guangzhi Lei 2019. DurIAN: Duration informed attention network for multimodal synthesis. arXiv preprint arXiv:1909.01700(2019).  Chengzhu Yu Heng Lu Na Hu Meng Yu Chao Weng Kun Xu Peng Liu Deyi Tuo Shiyin Kang Guangzhi Lei 2019. DurIAN: Duration informed attention network for multimodal synthesis. arXiv preprint arXiv:1909.01700(2019)."}],"event":{"name":"ICMI '21: INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION","location":"Montr\u00e9al QC Canada","acronym":"ICMI '21","sponsor":["SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["Proceedings of the 2021 International Conference on Multimodal Interaction"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3462244.3479914","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3462244.3479914","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:48:54Z","timestamp":1750193334000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3462244.3479914"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,18]]},"references-count":61,"alternative-id":["10.1145\/3462244.3479914","10.1145\/3462244"],"URL":"https:\/\/doi.org\/10.1145\/3462244.3479914","relation":{},"subject":[],"published":{"date-parts":[[2021,10,18]]},"assertion":[{"value":"2021-10-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}