{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T14:54:19Z","timestamp":1784904859323,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":48,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,21]],"date-time":"2020-10-21T00:00:00Z","timestamp":1603238400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004063","name":"Knut och Alice Wallenbergs Stiftelse","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100004063","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Vetenskapsr\u00e5det","award":["2018-05409"],"award-info":[{"award-number":["2018-05409"]}]},{"name":"Stiftelsen f\u00f6r Strategisk Forskning","award":["RIT15-0107"],"award-info":[{"award-number":["RIT15-0107"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,21]]},"DOI":"10.1145\/3382507.3418815","type":"proceedings-article","created":{"date-parts":[[2020,10,22]],"date-time":"2020-10-22T10:04:35Z","timestamp":1603361075000},"page":"242-250","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":160,"title":["Gesticulator: A framework for semantically-aware speech-driven gesture generation"],"prefix":"10.1145","author":[{"given":"Taras","family":"Kucherenko","sequence":"first","affiliation":[{"name":"KTH Royal Institute of Technology, Stockholm, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Patrik","family":"Jonell","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Stockholm, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sanne","family":"van Waveren","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Stockholm, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gustav Eje","family":"Henter","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Stockholm, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Simon","family":"Alexandersson","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Stockholm, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Iolanda","family":"Leite","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Stockholm, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hedvig","family":"Kjellstr\u00f6m","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Stockholm, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,10,22]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3340555.3353725"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.13946"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/566654.566606"},{"key":"e_1_3_2_2_4_1","volume-title":"Proceedings of the 2nd Workshop on Gesture and Speech in Interaction (GeSpIn","author":"Bergmann Kirsten","year":"2011","unstructured":"Kirsten Bergmann , Volkan Aksu , and Stefan Kopp . 2011 . The relation of speech and gestures: Temporal synchrony follows semantic synchrony . In Proceedings of the 2nd Workshop on Gesture and Speech in Interaction (GeSpIn 2011). Kirsten Bergmann, Volkan Aksu, and Stefan Kopp. 2011. The relation of speech and gestures: Temporal synchrony follows semantic synchrony. In Proceedings of the 2nd Workshop on Gesture and Speech in Interaction (GeSpIn 2011)."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/975817.975842"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/383259.383315"},{"key":"e_1_3_2_2_7_1","volume-title":"Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems. International Foundation for Autonomous Agents and Multiagent Systems, 702--710","author":"Castillo Gabriel","year":"2019","unstructured":"Gabriel Castillo and Michael Neff . 2019 . What do we express without knowing?: Emotion in Gesture . In Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems. International Foundation for Autonomous Agents and Multiagent Systems, 702--710 . Gabriel Castillo and Michael Neff. 2019. What do we express without knowing?: Emotion in Gesture. In Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems. International Foundation for Autonomous Agents and Multiagent Systems, 702--710."},{"key":"e_1_3_2_2_8_1","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Chen Xi","year":"2017","unstructured":"Xi Chen , Diederik P. Kingma , Tim Salimans , Yan Duan , Prafulla Dhariwal , John Schulman , Ilya Sutskever , and Pieter Abbeel . 2017 . Variational lossy autoencoder . In Proceedings of the International Conference on Learning Representations. Xi Chen, Diederik P. Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Ilya Sutskever, and Pieter Abbeel. 2017. Variational lossy autoencoder. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-21996-7_17"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1037\/xlm0000168"},{"key":"e_1_3_2_2_11_1","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . BERT: Pre-training of deep bidirectional transformers for language understanding . Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (2018)."},{"key":"e_1_3_2_2_12_1","volume-title":"Workshop on Multimodal Corpora","volume":"6","author":"Ferr\u00e9 Ga\u00eblle","year":"2010","unstructured":"Ga\u00eblle Ferr\u00e9 . 2010 . Timing relationships between speech and co-verbal gestures in spontaneous French. In Language Resources and Evaluation , Workshop on Multimodal Corpora , Vol. 6 . 86--91. Ga\u00eblle Ferr\u00e9. 2010. Timing relationships between speech and co-verbal gestures in spontaneous French. In Language Resources and Evaluation, Workshop on Multimodal Corpora, Vol. 6. 86--91."},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3267851.3267898"},{"key":"e_1_3_2_2_14_1","volume-title":"Adversarial gesture generation with realistic gesture phasing. Computers & Graphics","author":"Ferstl Ylva","year":"2020","unstructured":"Ylva Ferstl , Michael Neff , and Rachel McDonnell . 2020. Adversarial gesture generation with realistic gesture phasing. Computers & Graphics ( 2020 ). Ylva Ferstl, Michael Neff, and Rachel McDonnell. 2020. Adversarial gesture generation with realistic gesture phasing. Computers & Graphics (2020)."},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00361"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1080\/10867651.1998.10487493"},{"key":"e_1_3_2_2_17_1","volume-title":"When speech stops, gesture stops: Evidence from developmental and crosslinguistic comparisons. Frontiers in Psychology","author":"Graziano Maria","year":"2018","unstructured":"Maria Graziano and Marianne Gullberg . 2018. When speech stops, gesture stops: Evidence from developmental and crosslinguistic comparisons. Frontiers in Psychology ( 2018 ). Maria Graziano and Marianne Gullberg. 2018. When speech stops, gesture stops: Evidence from developmental and crosslinguistic comparisons. Frontiers in Psychology (2018)."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1111\/lang.12376"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3267851.3267878"},{"key":"e_1_3_2_2_20_1","volume-title":"MoGlow: Probabilistic and controllable motion synthesis using normalising flows. arXiv preprint arXiv:1905.06598","author":"Henter Gustav E.","year":"2019","unstructured":"Gustav E. Henter , Simon Alexanderson , and Jonas Beskow . 2019. MoGlow: Probabilistic and controllable motion synthesis using normalising flows. arXiv preprint arXiv:1905.06598 ( 2019 ). Gustav E. Henter, Simon Alexanderson, and Jonas Beskow. 2019. MoGlow: Probabilistic and controllable motion synthesis using normalising flows. arXiv preprint arXiv:1905.06598 (2019)."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2157689.2157694"},{"key":"e_1_3_2_2_22_1","volume-title":"A speech-driven hand gesture generation method and evaluation in android robots","author":"Ishi Carlos T.","year":"2018","unstructured":"Carlos T. Ishi , Daichi Machiyashiki , Ryusuke Mikata , and Hiroshi Ishiguro . 2018. A speech-driven hand gesture generation method and evaluation in android robots . IEEE Robotics and Automation Letters ( 2018 ). Carlos T. Ishi, Daichi Machiyashiki, Ryusuke Mikata, and Hiroshi Ishiguro. 2018. A speech-driven hand gesture generation method and evaluation in android robots. IEEE Robotics and Automation Letters (2018)."},{"key":"e_1_3_2_2_23_1","first-page":"11","article-title":"Hand, mouth and brain. The dynamic emergence of speech and gesture","volume":"6","author":"Iverson Jana M.","year":"1999","unstructured":"Jana M. Iverson and Esther Thelen . 1999 . Hand, mouth and brain. The dynamic emergence of speech and gesture . Journal of Consciousness Studies 6 , 11 -- 12 (1999), 19--40. Jana M. Iverson and Esther Thelen. 1999. Hand, mouth and brain. The dynamic emergence of speech and gesture. Journal of Consciousness Studies 6, 11--12 (1999), 19--40.","journal-title":"Journal of Consciousness Studies"},{"key":"e_1_3_2_2_24_1","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Kingma Diederik P","year":"2015","unstructured":"Diederik P Kingma and Jimmy Ba . 2015 . Adam: A method for stochastic optimization . In Proceedings of the International Conference on Learning Representations. Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_2_25_1","volume-title":"Proceedings of the Conference of the International Society for Gesture Studies.","author":"Kopp Stefan","year":"2007","unstructured":"Stefan Kopp , Hannes Rieser , Ipke Wachsmuth , Kirsten Bergmann , and Andy L\u00fccking . 2007 . Speech-gesture alignment . In Proceedings of the Conference of the International Society for Gesture Studies. Stefan Kopp, Hannes Rieser, Ipke Wachsmuth, Kirsten Bergmann, and Andy L\u00fccking. 2007. Speech-gesture alignment. In Proceedings of the Conference of the International Society for Gesture Studies."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/566654.566605"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3242969.3264970"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3308532.3329472"},{"key":"e_1_3_2_2_29_1","unstructured":"Gilwoo Lee Zhiwei Deng Shugao Ma Takaaki Shiratori Siddhartha S Srinivasa and Yaser Sheikh. 2019. Talking With Hands 16.2 M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and Synthesis.. In ICCV. 763--772.  Gilwoo Lee Zhiwei Deng Shugao Ma Takaaki Shiratori Siddhartha S Srinivasa and Yaser Sheikh. 2019. Talking With Hands 16.2 M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and Synthesis.. In ICCV. 763--772."},{"key":"e_1_3_2_2_30_1","volume-title":"Tune: A research platform for distributed model selection and training. arXiv preprint arXiv:1807.05118","author":"Liaw Richard","year":"2018","unstructured":"Richard Liaw , Eric Liang , Robert Nishihara , Philipp Moritz , Joseph E. Gonzalez , and Ion Stoica . 2018 . Tune: A research platform for distributed model selection and training. arXiv preprint arXiv:1807.05118 (2018). Richard Liaw, Eric Liang, Robert Nishihara, Philipp Moritz, Joseph E. Gonzalez, and Ion Stoica. 2018. Tune: A research platform for distributed model selection and training. arXiv preprint arXiv:1807.05118 (2018)."},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1515\/lp-2012-0006"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CIG.2018.8490399"},{"key":"e_1_3_2_2_33_1","volume-title":"Hand and Mind: What Gestures Reveal about Thought","author":"McNeill David","unstructured":"David McNeill . 1992. Hand and Mind: What Gestures Reveal about Thought . University of Chicago Press. David McNeill. 1992. Hand and Mind: What Gestures Reveal about Thought. University of Chicago Press."},{"key":"e_1_3_2_2_34_1","volume-title":"WordNet: a lexical database for English. Commun. ACM","author":"Miller George A.","year":"1995","unstructured":"George A. Miller . 1995. WordNet: a lexical database for English. Commun. ACM ( 1995 ). George A. Miller. 1995. WordNet: a lexical database for English. Commun. ACM (1995)."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-93846-2_41"},{"key":"e_1_3_2_2_36_1","volume-title":"Spatial control of arm movements. Experimental brain research 42, 2","author":"Morasso Pietro","year":"1981","unstructured":"Pietro Morasso . 1981. Spatial control of arm movements. Experimental brain research 42, 2 ( 1981 ), 223--227. Pietro Morasso. 1981. Spatial control of arm movements. Experimental brain research 42, 2 (1981), 223--227."},{"key":"e_1_3_2_2_37_1","volume-title":"Gesture modeling and animation based on a probabilistic re-creation of speaker style. ACM Transactions on Graphics","author":"Neff Michael","year":"2008","unstructured":"Michael Neff , Michael Kipp , Irene Albrecht , and Hans-Peter Seidel . 2008. Gesture modeling and animation based on a probabilistic re-creation of speaker style. ACM Transactions on Graphics ( 2008 ). Michael Neff, Michael Kipp, Irene Albrecht, and Hans-Peter Seidel. 2008. Gesture modeling and animation based on a probabilistic re-creation of speaker style. ACM Transactions on Graphics (2008)."},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11671"},{"key":"e_1_3_2_2_39_1","volume-title":"Proceedings of the Gesture and Speech in Interaction Workshop.","author":"Pouw Wim","unstructured":"Wim Pouw and James A. Dixon . 2019. Quantifying gesture-speech synchrony . In Proceedings of the Gesture and Speech in Interaction Workshop. Wim Pouw and James A. Dixon. 2019. Quantifying gesture-speech synchrony. In Proceedings of the Gesture and Speech in Interaction Workshop."},{"key":"e_1_3_2_2_40_1","volume-title":"Dixon","author":"Pouw Wim","year":"2019","unstructured":"Wim Pouw , Steven J. Harrison , and James A . Dixon . 2019 . Gesture--speech physics: The biomechanical basis for the emergence of gesture--speech synchrony. Journal of Experimental Psychology: General ( 2019). Wim Pouw, Steven J. Harrison, and James A. Dixon. 2019. Gesture--speech physics: The biomechanical basis for the emergence of gesture--speech synchrony. Journal of Experimental Psychology: General (2019)."},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.2196\/iproc.6065"},{"key":"e_1_3_2_2_42_1","volume-title":"Speech-driven animation with meaningful behaviors. Speech Communication","author":"Sadoughi Najmeh","year":"2019","unstructured":"Najmeh Sadoughi and Carlos Busso . 2019. Speech-driven animation with meaningful behaviors. Speech Communication ( 2019 ). Najmeh Sadoughi and Carlos Busso. 2019. Speech-driven animation with meaningful behaviors. Speech Communication (2019)."},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/ROMAN.2011.6005285"},{"key":"e_1_3_2_2_44_1","volume-title":"Samer Al Moubayed, and Bj\u00f6rn Granstr\u00f6m","author":"Salvi Giampiero","year":"2009","unstructured":"Giampiero Salvi , Jonas Beskow , Samer Al Moubayed, and Bj\u00f6rn Granstr\u00f6m . 2009 . SynFace: Speech-driven facial animation for virtual speech-reading support. Journal on Audio, Speech, and Music Processing ( 2009). Giampiero Salvi, Jonas Beskow, Samer Al Moubayed, and Bj\u00f6rn Granstr\u00f6m. 2009. SynFace: Speech-driven facial animation for virtual speech-reading support. Journal on Audio, Speech, and Music Processing (2009)."},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.5555\/1167629.1167639"},{"key":"e_1_3_2_2_46_1","volume-title":"Formation and control of optimal trajectory in human multijoint arm movement. Biological cybernetics 61, 2","author":"Uno Yoji","year":"1989","unstructured":"Yoji Uno , Mitsuo Kawato , and Rika Suzuki . 1989. Formation and control of optimal trajectory in human multijoint arm movement. Biological cybernetics 61, 2 ( 1989 ), 89--101. Yoji Uno, Mitsuo Kawato, and Rika Suzuki. 1989. Formation and control of optimal trajectory in human multijoint arm movement. Biological cybernetics 61, 2 (1989), 89--101."},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.21437\/SSW.2016-33"},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8793720"}],"event":{"name":"ICMI '20: INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION","location":"Virtual Event Netherlands","acronym":"ICMI '20","sponsor":["SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["Proceedings of the 2020 International Conference on Multimodal Interaction"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3382507.3418815","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3382507.3418815","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:38:27Z","timestamp":1750199907000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3382507.3418815"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,21]]},"references-count":48,"alternative-id":["10.1145\/3382507.3418815","10.1145\/3382507"],"URL":"https:\/\/doi.org\/10.1145\/3382507.3418815","relation":{},"subject":[],"published":{"date-parts":[[2020,10,21]]},"assertion":[{"value":"2020-10-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}