{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T06:06:35Z","timestamp":1784268395949,"version":"3.55.0"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2020,11,27]],"date-time":"2020-11-27T00:00:00Z","timestamp":1606435200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Korea government","award":["2017-0-00162"],"award-info":[{"award-number":["2017-0-00162"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2020,12,31]]},"abstract":"<jats:p>For human-like agents, including virtual avatars and social robots, making proper gestures while speaking is crucial in human-agent interaction. Co-speech gestures enhance interaction experiences and make the agents look alive. However, it is difficult to generate human-like gestures due to the lack of understanding of how people gesture. Data-driven approaches attempt to learn gesticulation skills from human demonstrations, but the ambiguous and individual nature of gestures hinders learning. In this paper, we present an automatic gesture generation model that uses the multimodal context of speech text, audio, and speaker identity to reliably generate gestures. By incorporating a multimodal context and an adversarial training scheme, the proposed model outputs gestures that are human-like and that match with speech content and rhythm. We also introduce a new quantitative evaluation metric for gesture generation models. Experiments with the introduced metric and subjective human evaluation showed that the proposed gesture generation model is better than existing end-to-end generation models. We further confirm that our model is able to work with synthesized audio in a scenario where contexts are constrained, and show that different gesture styles can be generated for the same speech by specifying different speaker identities in the style embedding space that is learned from videos of various speakers. All the code and data is available at https:\/\/github.com\/ai4r\/Gesture-Generation-from-Trimodal-Context.<\/jats:p>","DOI":"10.1145\/3414685.3417838","type":"journal-article","created":{"date-parts":[[2020,11,27]],"date-time":"2020-11-27T21:51:05Z","timestamp":1606513865000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":294,"title":["Speech gesture generation from the trimodal context of text, audio, and speaker identity"],"prefix":"10.1145","volume":"39","author":[{"given":"Youngwoo","family":"Yoon","sequence":"first","affiliation":[{"name":"KAIST"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bok","family":"Cha","sequence":"additional","affiliation":[{"name":"University of Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Joo-Haeng","family":"Lee","sequence":"additional","affiliation":[{"name":"University of Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Minsu","family":"Jang","sequence":"additional","affiliation":[{"name":"ETRI"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jaeyeon","family":"Lee","sequence":"additional","affiliation":[{"name":"ETRI"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jaehong","family":"Kim","sequence":"additional","affiliation":[{"name":"ETRI"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Geehyuk","family":"Lee","sequence":"additional","affiliation":[{"name":"KAIST"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,11,27]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2019.00084"},{"key":"e_1_2_2_2_1","volume-title":"Taras Kucherenko, and Jonas Beskow.","author":"Alexanderson Simon","year":"2020","unstructured":"Simon Alexanderson , Gustav Eje Henter , Taras Kucherenko, and Jonas Beskow. 2020 . Style-Controllable Speech-Driven Gesture Synthesis Using Normalising Flows. In Computer Graphics Forum, Vol. 39 . Wiley Online Library , 487--496. Simon Alexanderson, Gustav Eje Henter, Taras Kucherenko, and Jonas Beskow. 2020. Style-Controllable Speech-Driven Gesture Synthesis Using Normalising Flows. In Computer Graphics Forum, Vol. 39. Wiley Online Library, 487--496."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2755566"},{"key":"e_1_2_2_4_1","volume-title":"International Conference on Learning Representations.","author":"Bahdanau Dzmitry","year":"2015","unstructured":"Dzmitry Bahdanau , Kyunghyun Cho , and Yoshua Bengio . 2015 . Neural Machine Translation by Jointly Learning to Align and Translate . International Conference on Learning Representations. Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural Machine Translation by Jointly Learning to Align and Translate. International Conference on Learning Representations."},{"key":"e_1_2_2_5_1","volume-title":"An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. arXiv:1803.01271","author":"Bai Shaojie","year":"2018","unstructured":"Shaojie Bai , J. Zico Kolter , and Vladlen Koltun . 2018. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. arXiv:1803.01271 ( 2018 ). Shaojie Bai, J. Zico Kolter, and Vladlen Koltun. 2018. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. arXiv:1803.01271 (2018)."},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2798607"},{"key":"e_1_2_2_7_1","volume-title":"Proceedings of the 2nd Workshop on Gesture and Speech in Interaction.","author":"Bergmann Kirsten","year":"2011","unstructured":"Kirsten Bergmann , Volkan Aksu , and Stefan Kopp . 2011 . The Relation of Speech and Gestures: Temporal Synchrony Follows Semantic Synchrony . In Proceedings of the 2nd Workshop on Gesture and Speech in Interaction. Kirsten Bergmann, Volkan Aksu, and Stefan Kopp. 2011. The Relation of Speech and Gestures: Temporal Synchrony Follows Semantic Synchrony. In Proceedings of the 2nd Workshop on Gesture and Speech in Interaction."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00051"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2018.10.009"},{"key":"e_1_2_2_10_1","volume-title":"The Effects of Robot-Performed Co-Verbal Gesture on Listener Behaviour. In IEEE-RAS International Conference on Humanoid Robots. IEEE, 458--465","author":"Bremner Paul","year":"2011","unstructured":"Paul Bremner , Anthony G Pipe , Chris Melhuish , Mike Fraser , and Sriram Subramanian . 2011 . The Effects of Robot-Performed Co-Verbal Gesture on Listener Behaviour. In IEEE-RAS International Conference on Humanoid Robots. IEEE, 458--465 . Paul Bremner, Anthony G Pipe, Chris Melhuish, Mike Fraser, and Sriram Subramanian. 2011. The Effects of Robot-Performed Co-Verbal Gesture on Listener Behaviour. In IEEE-RAS International Conference on Humanoid Robots. IEEE, 458--465."},{"key":"e_1_2_2_11_1","volume-title":"Human communication research 17, 1","author":"Burgoon Judee K","year":"1990","unstructured":"Judee K Burgoon , Thomas Birk , and Michael Pfau . 1990. Nonverbal Behaviors , Persuasion, and Credibility. Human communication research 17, 1 ( 1990 ), 140--169. Judee K Burgoon, Thomas Birk, and Michael Pfau. 1990. Nonverbal Behaviors, Persuasion, and Credibility. Human communication research 17, 1 (1990), 140--169."},{"key":"e_1_2_2_12_1","volume-title":"Hannes H\u00f6gni Vilhj\u00e1lmsson, and Timothy Bickmore","author":"Cassell Justine","year":"2004","unstructured":"Justine Cassell , Hannes H\u00f6gni Vilhj\u00e1lmsson, and Timothy Bickmore . 2004 . BEAT: the Behavior Expression Animation Toolkit. In Life-Like Characters. Springer , 163--185. Justine Cassell, Hannes H\u00f6gni Vilhj\u00e1lmsson, and Timothy Bickmore. 2004. BEAT: the Behavior Expression Animation Toolkit. In Life-Like Characters. Springer, 163--185."},{"key":"e_1_2_2_13_1","volume-title":"Predicting Co-verbal Gestures: A Deep and Temporal Modeling Approach. In ACM International Conference on Intelligent Virtual Agents. Springer, 152--166","author":"Chiu Chung-Cheng","year":"2015","unstructured":"Chung-Cheng Chiu , Louis-Philippe Morency , and Stacy Marsella . 2015 . Predicting Co-verbal Gestures: A Deep and Temporal Modeling Approach. In ACM International Conference on Intelligent Virtual Agents. Springer, 152--166 . Chung-Cheng Chiu, Louis-Philippe Morency, and Stacy Marsella. 2015. Predicting Co-verbal Gestures: A Deep and Temporal Modeling Approach. In ACM International Conference on Intelligent Virtual Agents. Springer, 152--166."},{"key":"e_1_2_2_14_1","unstructured":"Kyunghyun Cho Bart van Merri\u00ebnboer Caglar Gulcehre Dzmitry Bahdanau Fethi Bougares Holger Schwenk and Yoshua Bengio. 2014. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. In Empirical Methods in Natural Language Processing. 1724--1734.  Kyunghyun Cho Bart van Merri\u00ebnboer Caglar Gulcehre Dzmitry Bahdanau Fethi Bougares Holger Schwenk and Yoshua Bengio. 2014. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. In Empirical Methods in Natural Language Processing. 1724--1734."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1037\/a0036281"},{"key":"e_1_2_2_16_1","volume-title":"Extensions of Gaussian Processes for Ranking: Semi-supervised and Active Learning. Learning to Rank","author":"Chu Wei","year":"2005","unstructured":"Wei Chu and Zoubin Ghahramani . 2005. Extensions of Gaussian Processes for Ranking: Semi-supervised and Active Learning. Learning to Rank ( 2005 ), 29. Wei Chu and Zoubin Ghahramani. 2005. Extensions of Gaussian Processes for Ranking: Semi-supervised and Active Learning. Learning to Rank (2005), 29."},{"key":"e_1_2_2_17_1","volume-title":"Why Rate When You Could Compare? Using the \"EloChoice\" Package to Assess Pairwise Comparisons of Perceived Physical Strength. PloS one 13, 1","author":"Clark Andrew P","year":"2018","unstructured":"Andrew P Clark , Kate L Howard , Andy T Woods , Ian S Penton-Voak , and Christof Neumann . 2018. Why Rate When You Could Compare? Using the \"EloChoice\" Package to Assess Pairwise Comparisons of Perceived Physical Strength. PloS one 13, 1 ( 2018 ). Andrew P Clark, Kate L Howard, Andy T Woods, Ian S Penton-Voak, and Christof Neumann. 2018. Why Rate When You Could Compare? Using the \"EloChoice\" Package to Assess Pairwise Comparisons of Perceived Physical Strength. PloS one 13, 1 (2018)."},{"key":"e_1_2_2_18_1","doi-asserted-by":"crossref","unstructured":"Ylva Ferstl Michael Neff and Rachel McDonnell. 2019. Multi-Objective Adversarial Gesture Generation. In Motion Interaction and Games. 1--10.  Ylva Ferstl Michael Neff and Rachel McDonnell. 2019. Multi-Objective Adversarial Gesture Generation. In Motion Interaction and Games. 1--10.","DOI":"10.1145\/3359566.3360053"},{"key":"e_1_2_2_19_1","volume-title":"AAAI Conference on Artificial Intelligence.","author":"Fu Peng","year":"2018","unstructured":"Peng Fu , Zheng Lin , Fengcheng Yuan , Weiping Wang , and Dan Meng . 2018 . Learning Sentiment-Specific Word Embedding via Global Sentiment Representation . In AAAI Conference on Artificial Intelligence. Peng Fu, Zheng Lin, Fengcheng Yuan, Weiping Wang, and Dan Meng. 2018. Learning Sentiment-Specific Word Embedding via Global Sentiment Representation. In AAAI Conference on Artificial Intelligence."},{"key":"e_1_2_2_20_1","volume-title":"Learning Individual Styles of Conversational Gesture. In IEEE Conference on Computer Vision and Pattern Recognition. 3497--3506","author":"Ginosar Shiry","year":"2019","unstructured":"Shiry Ginosar , Amir Bar , Gefen Kohavi , Caroline Chan , Andrew Owens , and Jitendra Malik . 2019 . Learning Individual Styles of Conversational Gesture. In IEEE Conference on Computer Vision and Pattern Recognition. 3497--3506 . Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik. 2019. Learning Individual Styles of Conversational Gesture. In IEEE Conference on Computer Vision and Pattern Recognition. 3497--3506."},{"key":"e_1_2_2_21_1","unstructured":"Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative Adversarial Nets. In Advances in Neural Information Processing Systems. 2672--2680.  Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative Adversarial Nets. In Advances in Neural Information Processing Systems. 2672--2680."},{"key":"e_1_2_2_22_1","unstructured":"Google. 2018. Google Cloud Text-to-Speech. https:\/\/cloud.google.com\/text-to-speech Accessed: 2020-03-01.  Google. 2018. Google Cloud Text-to-Speech. https:\/\/cloud.google.com\/text-to-speech Accessed: 2020-03-01."},{"key":"e_1_2_2_23_1","unstructured":"Martin Heusel Hubert Ramsauer Thomas Unterthiner Bernhard Nessler and Sepp Hochreiter. 2017. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Advances in Neural Information Processing Systems. 6626--6637.  Martin Heusel Hubert Ramsauer Thomas Unterthiner Bernhard Nessler and Sepp Hochreiter. 2017. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Advances in Neural Information Processing Systems. 6626--6637."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1075\/gest.12.1.04hos"},{"key":"e_1_2_2_25_1","volume-title":"Learning-Based Modeling of Multimodal Behaviors for Humanlike Robots. In ACM\/IEEE International Conference on Human-Robot Interaction. ACM, 57--64","author":"Huang Chien-Ming","year":"2014","unstructured":"Chien-Ming Huang and Bilge Mutlu . 2014 . Learning-Based Modeling of Multimodal Behaviors for Humanlike Robots. In ACM\/IEEE International Conference on Human-Robot Interaction. ACM, 57--64 . Chien-Ming Huang and Bilge Mutlu. 2014. Learning-Based Modeling of Multimodal Behaviors for Humanlike Robots. In ACM\/IEEE International Conference on Human-Robot Interaction. ACM, 57--64."},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177703732"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.248"},{"key":"e_1_2_2_28_1","volume-title":"International Conference on Learning Representations.","author":"Jahanian Ali","year":"2020","unstructured":"Ali Jahanian , Lucy Chai , and Phillip Isola . 2020 . On the \"Steerability\" of Generative Adversarial Networks . In International Conference on Learning Representations. Ali Jahanian, Lucy Chai, and Phillip Isola. 2020. On the \"Steerability\" of Generative Adversarial Networks. In International Conference on Learning Representations."},{"key":"e_1_2_2_29_1","volume-title":"Towards Social Artificial Intelligence: Nonverbal Social Signal Prediction in A Triadic Interaction. In IEEE Conference on Computer Vision and Pattern Recognition. 10873--10883","author":"Joo Hanbyul","year":"2019","unstructured":"Hanbyul Joo , Tomas Simon , Mina Cikara , and Yaser Sheikh . 2019 . Towards Social Artificial Intelligence: Nonverbal Social Signal Prediction in A Triadic Interaction. In IEEE Conference on Computer Vision and Pattern Recognition. 10873--10883 . Hanbyul Joo, Tomas Simon, Mina Cikara, and Yaser Sheikh. 2019. Towards Social Artificial Intelligence: Nonverbal Social Signal Prediction in A Triadic Interaction. In IEEE Conference on Computer Vision and Pattern Recognition. 10873--10883."},{"key":"e_1_2_2_30_1","volume-title":"Fr\u00e9chet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms. arXiv preprint arXiv:1812.08466","author":"Kilgour Kevin","year":"2018","unstructured":"Kevin Kilgour , Mauricio Zuluaga , Dominik Roblek , and Matthew Sharifi . 2018. Fr\u00e9chet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms. arXiv preprint arXiv:1812.08466 ( 2018 ). Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi. 2018. Fr\u00e9chet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms. arXiv preprint arXiv:1812.08466 (2018)."},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA40945.2020.9196948"},{"key":"e_1_2_2_32_1","volume-title":"Auto-Encoding Variational Bayes. In International Conference on Learning Representations.","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Max Welling . 2014 . Auto-Encoding Variational Bayes. In International Conference on Learning Representations. Diederik P Kingma and Max Welling. 2014. Auto-Encoding Variational Bayes. In International Conference on Learning Representations."},{"key":"e_1_2_2_33_1","volume-title":"Gesture Generation by Imitation: From Human Behavior to Computer Character Animation","author":"Kipp Michael","unstructured":"Michael Kipp . 2005. Gesture Generation by Imitation: From Human Behavior to Computer Character Animation . Universal-Publishers . Michael Kipp. 2005. Gesture Generation by Imitation: From Human Behavior to Computer Character Animation. Universal-Publishers."},{"key":"e_1_2_2_34_1","volume-title":"How Representational Gestures Help Speaking. Language and gesture 1","author":"Kita Sotaro","year":"2000","unstructured":"Sotaro Kita . 2000. How Representational Gestures Help Speaking. Language and gesture 1 ( 2000 ), 162--185. Sotaro Kita. 2000. How Representational Gestures Help Speaking. Language and gesture 1 (2000), 162--185."},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/11821830_17"},{"key":"e_1_2_2_36_1","volume-title":"Analyzing Input and Output Representations for Speech-Driven Gesture Generation. In ACM International Conference on Intelligent Virtual Agents. 97--104","author":"Kucherenko Taras","year":"2019","unstructured":"Taras Kucherenko , Dai Hasegawa , Gustav Eje Henter , Naoshi Kaneko , and Hedvig Kjellstr\u00f6m . 2019 . Analyzing Input and Output Representations for Speech-Driven Gesture Generation. In ACM International Conference on Intelligent Virtual Agents. 97--104 . Taras Kucherenko, Dai Hasegawa, Gustav Eje Henter, Naoshi Kaneko, and Hedvig Kjellstr\u00f6m. 2019. Analyzing Input and Output Representations for Speech-Driven Gesture Generation. In ACM International Conference on Intelligent Virtual Agents. 97--104."},{"key":"e_1_2_2_37_1","volume-title":"Gesticulator: A Framework for Semantically-Aware Speech-Driven Gesture Generation. In ACM International Conference on Multimodal Interaction.","author":"Kucherenko Taras","year":"2020","unstructured":"Taras Kucherenko , Patrik Jonell , Sanne van Waveren , Gustav Eje Henter , Simon Alexanderson , Iolanda Leite , and Hedvig Kjellstr\u00f6m . 2020 . Gesticulator: A Framework for Semantically-Aware Speech-Driven Gesture Generation. In ACM International Conference on Multimodal Interaction. Taras Kucherenko, Patrik Jonell, Sanne van Waveren, Gustav Eje Henter, Simon Alexanderson, Iolanda Leite, and Hedvig Kjellstr\u00f6m. 2020. Gesticulator: A Framework for Semantically-Aware Speech-Driven Gesture Generation. In ACM International Conference on Multimodal Interaction."},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/1778765.1778861"},{"key":"e_1_2_2_39_1","volume-title":"Virtual Character Performance From Speech. In ACM SIGGRAPH\/Eurographics Symposium on Computer Animation. 25--35","author":"Marsella Stacy","year":"2013","unstructured":"Stacy Marsella , Yuyu Xu , Margaux Lhommet , Andrew Feng , Stefan Scherer , and Ari Shapiro . 2013 . Virtual Character Performance From Speech. In ACM SIGGRAPH\/Eurographics Symposium on Computer Animation. 25--35 . Stacy Marsella, Yuyu Xu, Margaux Lhommet, Andrew Feng, Stefan Scherer, and Ari Shapiro. 2013. Virtual Character Performance From Speech. In ACM SIGGRAPH\/Eurographics Symposium on Computer Animation. 25--35."},{"key":"e_1_2_2_40_1","volume-title":"UMAP: Uniform Manifold Approximation and Projection. Journal of Open Source Software 3","author":"McInnes Leland","year":"2018","unstructured":"Leland McInnes , John Healy , Nathaniel Saul , and Lukas Gro\u00dfberger . 2018 . UMAP: Uniform Manifold Approximation and Projection. Journal of Open Source Software 3 (2018). Leland McInnes, John Healy, Nathaniel Saul, and Lukas Gro\u00dfberger. 2018. UMAP: Uniform Manifold Approximation and Projection. Journal of Open Source Software 3 (2018)."},{"key":"e_1_2_2_41_1","volume-title":"Hand and Mind: What Gestures Reveal About Thought","author":"McNeill David","unstructured":"David McNeill . 1992. Hand and Mind: What Gestures Reveal About Thought . University of Chicago press. David McNeill. 1992. Hand and Mind: What Gestures Reveal About Thought. University of Chicago press."},{"key":"e_1_2_2_42_1","unstructured":"David McNeill. 2008. Gesture and Thought. University of Chicago press.  David McNeill. 2008. Gesture and Thought. University of Chicago press."},{"key":"e_1_2_2_43_1","unstructured":"Alberto Menache. 2000. Understanding Motion Capture for Computer Animation and Video Games. Morgan Kaufmann.  Alberto Menache. 2000. Understanding Motion Capture for Computer Animation and Video Games. Morgan Kaufmann."},{"key":"e_1_2_2_44_1","unstructured":"Tomas Mikolov Ilya Sutskever Kai Chen Greg S Corrado and Jeff Dean. 2013. Distributed Representations of Words and Phrases and their Compositionality. In Advances in Neural Information Processing Systems. 3111--3119.  Tomas Mikolov Ilya Sutskever Kai Chen Greg S Corrado and Jeff Dean. 2013. Distributed Representations of Words and Phrases and their Compositionality. In Advances in Neural Information Processing Systems. 3111--3119."},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/219717.219748"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/1330511.1330516"},{"key":"e_1_2_2_47_1","volume-title":"Gentle: A Forced Aligner. https:\/\/lowerquality.com\/gentle\/ Accessed: 2020-01-06.","author":"Ochshorn Robert","year":"2016","unstructured":"Robert Ochshorn and Max Hawkins . 2016 . Gentle: A Forced Aligner. https:\/\/lowerquality.com\/gentle\/ Accessed: 2020-01-06. Robert Ochshorn and Max Hawkins. 2016. Gentle: A Forced Aligner. https:\/\/lowerquality.com\/gentle\/ Accessed: 2020-01-06."},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00794"},{"key":"e_1_2_2_49_1","doi-asserted-by":"crossref","unstructured":"Jeffrey Pennington Richard Socher and Christopher Manning. 2014. GloVe: Global Vectors for Word Representation. In Empirical Methods in Natural Language Processing. 1532--1543.  Jeffrey Pennington Richard Socher and Christopher Manning. 2014. GloVe: Global Vectors for Word Representation. In Empirical Methods in Natural Language Processing. 1532--1543.","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_2_2_50_1","volume-title":"Proceedings of the International Conference on Machine Learning","volume":"32","author":"Rezende Danilo Jimenez","year":"2014","unstructured":"Danilo Jimenez Rezende , Shakir Mohamed , and Daan Wierstra . 2014 . Stochastic back-propagation and approximate inference in deep generative models . In Proceedings of the International Conference on Machine Learning , Vol. 32 . 1278--1286. Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. 2014. Stochastic back-propagation and approximate inference in deep generative models. In Proceedings of the International Conference on Machine Learning, Vol. 32. 1278--1286."},{"key":"e_1_2_2_51_1","volume-title":"Multimodal Continuous Turn-Taking Prediction Using Multiscale RNNs. In ACM International Conference on Multimodal Interaction. ACM, 186--190","author":"Roddy Matthew","year":"2018","unstructured":"Matthew Roddy , Gabriel Skantze , and Naomi Harte . 2018 . Multimodal Continuous Turn-Taking Prediction Using Multiscale RNNs. In ACM International Conference on Multimodal Interaction. ACM, 186--190 . Matthew Roddy, Gabriel Skantze, and Naomi Harte. 2018. Multimodal Continuous Turn-Taking Prediction Using Multiscale RNNs. In ACM International Conference on Multimodal Interaction. ACM, 186--190."},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2019.04.005"},{"key":"e_1_2_2_54_1","unstructured":"Tim Salimans Ian Goodfellow Wojciech Zaremba Vicki Cheung Alec Radford and Xi Chen. 2016. Improved Techniques for Training GANs. In Advances in Neural Information Processing Systems. 2234--2242.  Tim Salimans Ian Goodfellow Wojciech Zaremba Vicki Cheung Alec Radford and Xi Chen. 2016. Improved Techniques for Training GANs. In Advances in Neural Information Processing Systems. 2234--2242."},{"key":"e_1_2_2_55_1","volume-title":"NAOqi API Documentation","author":"Softbank Robotics","unstructured":"Robotics Softbank . 2018. NAOqi API Documentation . http:\/\/doc.aldebaran.com\/2-5\/index_dev_guide.html Accessed: 2020-01-06. Robotics Softbank. 2018. NAOqi API Documentation. http:\/\/doc.aldebaran.com\/2-5\/index_dev_guide.html Accessed: 2020-01-06."},{"key":"e_1_2_2_56_1","volume-title":"International Conference on Learning Representations Workshop.","author":"Unterthiner Thomas","year":"2019","unstructured":"Thomas Unterthiner , Sjoerd van Steenkiste , Karol Kurach , Rapha\u00ebl Marinier , Marcin Michalski , and Sylvain Gelly . 2019 . FVD: A new Metric for Video Generation . In International Conference on Learning Representations Workshop. Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Rapha\u00ebl Marinier, Marcin Michalski, and Sylvain Gelly. 2019. FVD: A new Metric for Video Generation. In International Conference on Learning Representations Workshop."},{"key":"e_1_2_2_57_1","volume-title":"Special Iss.","author":"Wagner Petra","year":"2014","unstructured":"Petra Wagner , Zofia Malisz , and Stefan Kopp . 2014. Gesture and Speech in Interaction: An Overview. Speech Communication 57 , Special Iss. ( 2014 ). Petra Wagner, Zofia Malisz, and Stefan Kopp. 2014. Gesture and Speech in Interaction: An Overview. Speech Communication 57, Special Iss. (2014)."},{"key":"e_1_2_2_58_1","volume-title":"Hand Gestures and Verbal Acknowledgments Improve Human-Robot Rapport. In International Conference on Social Robotics. Springer, 334--344","author":"Wilson Jason R","year":"2017","unstructured":"Jason R Wilson , Nah Young Lee , Annie Saechao , Sharon Hershenson , Matthias Scheutz , and Linda Tickle-Degnen . 2017 . Hand Gestures and Verbal Acknowledgments Improve Human-Robot Rapport. In International Conference on Social Robotics. Springer, 334--344 . Jason R Wilson, Nah Young Lee, Annie Saechao, Sharon Hershenson, Matthias Scheutz, and Linda Tickle-Degnen. 2017. Hand Gestures and Verbal Acknowledgments Improve Human-Robot Rapport. In International Conference on Social Robotics. Springer, 334--344."},{"key":"e_1_2_2_59_1","volume-title":"Diversity-Sensitive Conditional Generative Adversarial Networks. In International Conference on Learning Representations.","author":"Yang Dingdong","year":"2019","unstructured":"Dingdong Yang , Seunghoon Hong , Yunseok Jang , Tianchen Zhao , and Honglak Lee . 2019 . Diversity-Sensitive Conditional Generative Adversarial Networks. In International Conference on Learning Representations. Dingdong Yang, Seunghoon Hong, Yunseok Jang, Tianchen Zhao, and Honglak Lee. 2019. Diversity-Sensitive Conditional Generative Adversarial Networks. In International Conference on Learning Representations."},{"key":"e_1_2_2_60_1","volume-title":"Robots Learn Social Skills: End-to-End Learning of Co-Speech Gesture Generation for Humanoid Robots. In International Conference on Robotics and Automation. IEEE, 4303--4309","author":"Yoon Youngwoo","year":"2019","unstructured":"Youngwoo Yoon , Woo-Ri Ko , Minsu Jang , Jaeyeon Lee , Jaehong Kim , and Geehyuk Lee . 2019 . Robots Learn Social Skills: End-to-End Learning of Co-Speech Gesture Generation for Humanoid Robots. In International Conference on Robotics and Automation. IEEE, 4303--4309 . Youngwoo Yoon, Woo-Ri Ko, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee. 2019. Robots Learn Social Skills: End-to-End Learning of Co-Speech Gesture Generation for Humanoid Robots. In International Conference on Robotics and Automation. IEEE, 4303--4309."},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00068"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3414685.3417838","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3414685.3417838","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:03:15Z","timestamp":1750197795000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3414685.3417838"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,11,27]]},"references-count":60,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2020,12,31]]}},"alternative-id":["10.1145\/3414685.3417838"],"URL":"https:\/\/doi.org\/10.1145\/3414685.3417838","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,11,27]]},"assertion":[{"value":"2020-11-27","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}