{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T03:51:36Z","timestamp":1783482696368,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":45,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,10]],"date-time":"2021-10-10T00:00:00Z","timestamp":1633824000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,10]]},"DOI":"10.1145\/3472749.3474789","type":"proceedings-article","created":{"date-parts":[[2021,10,13]],"date-time":"2021-10-13T01:04:28Z","timestamp":1634087068000},"page":"826-840","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["SGToolkit: An Interactive Gesture Authoring Toolkit for Embodied Conversational Agents"],"prefix":"10.1145","author":[{"given":"Youngwoo","family":"Yoon","sequence":"first","affiliation":[{"name":"Electronics and Telecommunications Research Institute (ETRI), Korea, Republic of"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Keunwoo","family":"Park","sequence":"additional","affiliation":[{"name":"HCI Lab, School of Computing KAIST, Korea, Republic of"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Minsu","family":"Jang","sequence":"additional","affiliation":[{"name":"Electronics and Telecommunications Research Institute, Korea, Republic of"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jaehong","family":"Kim","sequence":"additional","affiliation":[{"name":"ETRI, Korea, Republic of"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Geehyuk","family":"Lee","sequence":"additional","affiliation":[{"name":"HCI Lab School of Computing, KAIST, Korea, Republic of"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,10,12]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58523-5_15"},{"key":"e_1_3_2_2_2_1","volume-title":"Computer Graphics Forum, Vol.\u00a039","author":"Alexanderson Simon","unstructured":"Simon Alexanderson , Gustav\u00a0Eje Henter , Taras Kucherenko , and Jonas Beskow . 2020. Style-Controllable Speech-Driven Gesture Synthesis Using Normalising Flows . In Computer Graphics Forum, Vol.\u00a039 . Wiley Online Library , 487\u2013496. Simon Alexanderson, Gustav\u00a0Eje Henter, Taras Kucherenko, and Jonas Beskow. 2020. Style-Controllable Speech-Driven Gesture Synthesis Using Normalising Flows. In Computer Graphics Forum, Vol.\u00a039. Wiley Online Library, 487\u2013496."},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00051"},{"key":"e_1_3_2_2_4_1","volume-title":"The Effects of Robot-Performed Co-Verbal Gesture on Listener Behaviour. In IEEE-RAS International Conference on Humanoid Robots. IEEE, 458\u2013465","author":"Bremner Paul","year":"2011","unstructured":"Paul Bremner , Anthony\u00a0 G Pipe , Chris Melhuish , Mike Fraser , and Sriram Subramanian . 2011 . The Effects of Robot-Performed Co-Verbal Gesture on Listener Behaviour. In IEEE-RAS International Conference on Humanoid Robots. IEEE, 458\u2013465 . Paul Bremner, Anthony\u00a0G Pipe, Chris Melhuish, Mike Fraser, and Sriram Subramanian. 2011. The Effects of Robot-Performed Co-Verbal Gesture on Listener Behaviour. In IEEE-RAS International Conference on Humanoid Robots. IEEE, 458\u2013465."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1468-2958.1990.tb00229.x"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"crossref","unstructured":"R Calvo S D\u2019Mello J Gratch A Kappas M Lhommet and SC Marsella. 2015. Expressing emotion through posture and gesture. The Oxford Handbook of Affective Computing(2015).  R Calvo S D\u2019Mello J Gratch A Kappas M Lhommet and SC Marsella. 2015. Expressing emotion through posture and gesture. The Oxford Handbook of Affective Computing(2015).","DOI":"10.1093\/oxfordhb\/9780199942237.013.039"},{"key":"e_1_3_2_2_7_1","volume-title":"Life-Like Characters","author":"Cassell Justine","unstructured":"Justine Cassell , Hannes\u00a0H\u00f6gni Vilhj\u00e1lmsson , and Timothy Bickmore . 2004. BEAT: the Behavior Expression Animation Toolkit . In Life-Like Characters . Springer , 163\u2013185. Justine Cassell, Hannes\u00a0H\u00f6gni Vilhj\u00e1lmsson, and Timothy Bickmore. 2004. BEAT: the Behavior Expression Animation Toolkit. In Life-Like Characters. Springer, 163\u2013185."},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Diane Chi Monica Costa Liwei Zhao and Norman Badler. 2000. The EMOTE model for effort and shape. In ACM SIGGRAPH. 173\u2013182.  Diane Chi Monica Costa Liwei Zhao and Norman Badler. 2000. The EMOTE model for effort and shape. In ACM SIGGRAPH. 173\u2013182.","DOI":"10.1145\/344779.352172"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"crossref","unstructured":"Ylva Ferstl Michael Neff and Rachel McDonnell. 2019. Multi-Objective Adversarial Gesture Generation. In Motion Interaction and Games. 1\u201310.  Ylva Ferstl Michael Neff and Rachel McDonnell. 2019. Multi-Objective Adversarial Gesture Generation. In Motion Interaction and Games. 1\u201310.","DOI":"10.1145\/3359566.3360053"},{"key":"e_1_3_2_2_10_1","volume-title":"Learning Individual Styles of Conversational Gesture. In IEEE Conference on Computer Vision and Pattern Recognition. 3497\u20133506","author":"Ginosar Shiry","year":"2019","unstructured":"Shiry Ginosar , Amir Bar , Gefen Kohavi , Caroline Chan , Andrew Owens , and Jitendra Malik . 2019 . Learning Individual Styles of Conversational Gesture. In IEEE Conference on Computer Vision and Pattern Recognition. 3497\u20133506 . Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik. 2019. Learning Individual Styles of Conversational Gesture. In IEEE Conference on Computer Vision and Pattern Recognition. 3497\u20133506."},{"key":"e_1_3_2_2_11_1","unstructured":"Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative Adversarial Nets. In Advances in Neural Information Processing Systems. 2672\u20132680.  Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative Adversarial Nets. In Advances in Neural Information Processing Systems. 2672\u20132680."},{"key":"e_1_3_2_2_12_1","unstructured":"Google. 2018. Google Cloud Text-to-Speech. https:\/\/cloud.google.com\/text-to-speech Accessed: 2021-01-01.  Google. 2018. Google Cloud Text-to-Speech. https:\/\/cloud.google.com\/text-to-speech Accessed: 2021-01-01."},{"key":"e_1_3_2_2_13_1","volume-title":"All Together Now: Introducing the Virtual Human Toolkit. In International Workshop on Intelligent Virtual Agents. Springer, 368\u2013381","author":"Hartholt Arno","year":"2013","unstructured":"Arno Hartholt , David Traum , Stacy\u00a0 C Marsella , Ari Shapiro , Giota Stratou , Anton Leuski , Louis-Philippe Morency , and Jonathan Gratch . 2013 . All Together Now: Introducing the Virtual Human Toolkit. In International Workshop on Intelligent Virtual Agents. Springer, 368\u2013381 . Arno Hartholt, David Traum, Stacy\u00a0C Marsella, Ari Shapiro, Giota Stratou, Anton Leuski, Louis-Philippe Morency, and Jonathan Gratch. 2013. All Together Now: Introducing the Virtual Human Toolkit. In International Workshop on Intelligent Virtual Agents. Springer, 368\u2013381."},{"key":"e_1_3_2_2_14_1","volume-title":"International Gesture Workshop. Springer, 188\u2013199","author":"Hartmann Bj\u00f6rn","year":"2005","unstructured":"Bj\u00f6rn Hartmann , Maurizio Mancini , and Catherine Pelachaud . 2005 . Implementing expressive gesture synthesis for embodied conversational agents . In International Gesture Workshop. Springer, 188\u2013199 . Bj\u00f6rn Hartmann, Maurizio Mancini, and Catherine Pelachaud. 2005. Implementing expressive gesture synthesis for embodied conversational agents. In International Gesture Workshop. Springer, 188\u2013199."},{"key":"e_1_3_2_2_15_1","volume-title":"Learning-Based Modeling of Multimodal Behaviors for Humanlike Robots. In ACM\/IEEE International Conference on Human-Robot Interaction. ACM, 57\u201364","author":"Huang Chien-Ming","year":"2014","unstructured":"Chien-Ming Huang and Bilge Mutlu . 2014 . Learning-Based Modeling of Multimodal Behaviors for Humanlike Robots. In ACM\/IEEE International Conference on Human-Robot Interaction. ACM, 57\u201364 . Chien-Ming Huang and Bilge Mutlu. 2014. Learning-Based Modeling of Multimodal Behaviors for Humanlike Robots. In ACM\/IEEE International Conference on Human-Robot Interaction. ACM, 57\u201364."},{"key":"e_1_3_2_2_16_1","unstructured":"Reallusion Inc.2016. CrazyTalk8. https:\/\/www.reallusion.com\/crazytalk\/ Accessed: 2021-01-01.  Reallusion Inc.2016. CrazyTalk8. https:\/\/www.reallusion.com\/crazytalk\/ Accessed: 2021-01-01."},{"key":"e_1_3_2_2_17_1","volume-title":"HEMVIP: Human evaluation of multiple videos in parallel. arxiv:2101.11898","author":"Jonell Patrik","year":"2021","unstructured":"Patrik Jonell , Youngwoo Yoon , Pieter Wolfert , Taras Kucherenko , and Gustav\u00a0Eje Henter . 2021 . HEMVIP: Human evaluation of multiple videos in parallel. arxiv:2101.11898 Patrik Jonell, Youngwoo Yoon, Pieter Wolfert, Taras Kucherenko, and Gustav\u00a0Eje Henter. 2021. HEMVIP: Human evaluation of multiple videos in parallel. arxiv:2101.11898"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073658"},{"key":"e_1_3_2_2_19_1","volume-title":"Gesture Generation by Imitation: From Human Behavior to Computer Character Animation","author":"Kipp Michael","unstructured":"Michael Kipp . 2005. Gesture Generation by Imitation: From Human Behavior to Computer Character Animation . Universal-Publishers . Michael Kipp. 2005. Gesture Generation by Imitation: From Human Behavior to Computer Character Animation. Universal-Publishers."},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ROMAN.2014.6926264"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/11821830_17"},{"key":"e_1_3_2_2_22_1","volume-title":"Analyzing Input and Output Representations for Speech-Driven Gesture Generation. In ACM International Conference on Intelligent Virtual Agents. 97\u2013104","author":"Kucherenko Taras","year":"2019","unstructured":"Taras Kucherenko , Dai Hasegawa , Gustav\u00a0Eje Henter , Naoshi Kaneko , and Hedvig Kjellstr\u00f6m . 2019 . Analyzing Input and Output Representations for Speech-Driven Gesture Generation. In ACM International Conference on Intelligent Virtual Agents. 97\u2013104 . Taras Kucherenko, Dai Hasegawa, Gustav\u00a0Eje Henter, Naoshi Kaneko, and Hedvig Kjellstr\u00f6m. 2019. Analyzing Input and Output Representations for Speech-Driven Gesture Generation. In ACM International Conference on Intelligent Virtual Agents. 97\u2013104."},{"key":"e_1_3_2_2_23_1","volume-title":"Gesticulator: A Framework for Semantically-Aware Speech-Driven Gesture Generation. In ACM International Conference on Multimodal Interaction.","author":"Kucherenko Taras","year":"2020","unstructured":"Taras Kucherenko , Patrik Jonell , Sanne van Waveren , Gustav\u00a0Eje Henter , Simon Alexanderson , Iolanda Leite , and Hedvig Kjellstr\u00f6m . 2020 . Gesticulator: A Framework for Semantically-Aware Speech-Driven Gesture Generation. In ACM International Conference on Multimodal Interaction. Taras Kucherenko, Patrik Jonell, Sanne van Waveren, Gustav\u00a0Eje Henter, Simon Alexanderson, Iolanda Leite, and Hedvig Kjellstr\u00f6m. 2020. Gesticulator: A Framework for Semantically-Aware Speech-Driven Gesture Generation. In ACM International Conference on Multimodal Interaction."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397481.3450692"},{"key":"e_1_3_2_2_25_1","unstructured":"Rudolf Laban and Lisa Ullmann. 1971. The mastery of movement.(1971).  Rudolf Laban and Lisa Ullmann. 1971. The mastery of movement.(1971)."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/11821830_20"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/1778765.1778861"},{"key":"e_1_3_2_2_28_1","volume-title":"Asian Conference on Computer Vision.","author":"Liao Miao","year":"2020","unstructured":"Miao Liao , Sibo Zhang , Peng Wang , Hao Zhu , Xinxin Zuo , and Ruigang Yang . 2020 . Speech2video synthesis with 3D skeleton regularization and expressive body poses . In Asian Conference on Computer Vision. Miao Liao, Sibo Zhang, Peng Wang, Hao Zhu, Xinxin Zuo, and Ruigang Yang. 2020. Speech2video synthesis with 3D skeleton regularization and expressive body poses. In Asian Conference on Computer Vision."},{"key":"e_1_3_2_2_29_1","volume-title":"Virtual Character Performance From Speech. In ACM SIGGRAPH\/Eurographics Symposium on Computer Animation. 25\u201335","author":"Marsella Stacy","year":"2013","unstructured":"Stacy Marsella , Yuyu Xu , Margaux Lhommet , Andrew Feng , Stefan Scherer , and Ari Shapiro . 2013 . Virtual Character Performance From Speech. In ACM SIGGRAPH\/Eurographics Symposium on Computer Animation. 25\u201335 . Stacy Marsella, Yuyu Xu, Margaux Lhommet, Andrew Feng, Stefan Scherer, and Ari Shapiro. 2013. Virtual Character Performance From Speech. In ACM SIGGRAPH\/Eurographics Symposium on Computer Animation. 25\u201335."},{"key":"e_1_3_2_2_30_1","volume-title":"Hand and Mind: What Gestures Reveal About Thought","author":"McNeill David","unstructured":"David McNeill . 1992. Hand and Mind: What Gestures Reveal About Thought . University of Chicago press. David McNeill. 1992. Hand and Mind: What Gestures Reveal About Thought. University of Chicago press."},{"key":"e_1_3_2_2_31_1","unstructured":"Alberto Menache. 2000. Understanding motion capture for computer animation and video games. Morgan kaufmann.  Alberto Menache. 2000. Understanding motion capture for computer animation and video games. Morgan kaufmann."},{"key":"e_1_3_2_2_32_1","volume-title":"js Essentials","author":"Moreau-Mathis Julien","unstructured":"Julien Moreau-Mathis . 2016. Babylon. js Essentials . Packt Publishing Ltd . Julien Moreau-Mathis. 2016. Babylon. js Essentials. Packt Publishing Ltd."},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01169"},{"key":"e_1_3_2_2_34_1","volume-title":"Gentle: A Forced Aligner. https:\/\/lowerquality.com\/gentle\/ Accessed: 2021-01-01.","author":"Ochshorn Robert","year":"2016","unstructured":"Robert Ochshorn and Max Hawkins . 2016 . Gentle: A Forced Aligner. https:\/\/lowerquality.com\/gentle\/ Accessed: 2021-01-01. Robert Ochshorn and Max Hawkins. 2016. Gentle: A Forced Aligner. https:\/\/lowerquality.com\/gentle\/ Accessed: 2021-01-01."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00794"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2559636.2559657"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2017.7989197"},{"key":"e_1_3_2_2_38_1","volume-title":"NAOqi API Documentation","author":"Softbank Robotics","unstructured":"Robotics Softbank . 2018. NAOqi API Documentation . http:\/\/doc.aldebaran.com\/2-5\/index_dev_guide.html Accessed: 2021-01-01. Robotics Softbank. 2018. NAOqi API Documentation. http:\/\/doc.aldebaran.com\/2-5\/index_dev_guide.html Accessed: 2021-01-01."},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00584"},{"key":"e_1_3_2_2_40_1","volume-title":"International Joint Conference on Autonomous Agents and Multiagent Systems. 151\u2013158","author":"Thiebaux Marcus","year":"2008","unstructured":"Marcus Thiebaux , Stacy Marsella , Andrew\u00a0 N Marshall , and Marcelo Kallmann . 2008 . Smartbody: Behavior realization for embodied conversational agents . In International Joint Conference on Autonomous Agents and Multiagent Systems. 151\u2013158 . Marcus Thiebaux, Stacy Marsella, Andrew\u00a0N Marshall, and Marcelo Kallmann. 2008. Smartbody: Behavior realization for embodied conversational agents. In International Joint Conference on Autonomous Agents and Multiagent Systems. 151\u2013158."},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58517-4_42"},{"key":"e_1_3_2_2_42_1","volume-title":"Method for the subjective assessment of intermediate quality level of audio systems","author":"Union International\u00a0Telecommunication","year":"2014","unstructured":"International\u00a0Telecommunication Union . 2014. Method for the subjective assessment of intermediate quality level of audio systems . International Telecommunication Union Radiocommunication Assembly ( 2014 ). International\u00a0Telecommunication Union. 2014. Method for the subjective assessment of intermediate quality level of audio systems. International Telecommunication Union Radiocommunication Assembly (2014)."},{"key":"e_1_3_2_2_43_1","volume-title":"Hand Gestures and Verbal Acknowledgments Improve Human-Robot Rapport. In International Conference on Social Robotics. Springer, 334\u2013344","author":"Wilson R","year":"2017","unstructured":"Jason\u00a0 R Wilson , Nah\u00a0Young Lee , Annie Saechao , Sharon Hershenson , Matthias Scheutz , and Linda Tickle-Degnen . 2017 . Hand Gestures and Verbal Acknowledgments Improve Human-Robot Rapport. In International Conference on Social Robotics. Springer, 334\u2013344 . Jason\u00a0R Wilson, Nah\u00a0Young Lee, Annie Saechao, Sharon Hershenson, Matthias Scheutz, and Linda Tickle-Degnen. 2017. Hand Gestures and Verbal Acknowledgments Improve Human-Robot Rapport. In International Conference on Social Robotics. Springer, 334\u2013344."},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417838"},{"key":"e_1_3_2_2_45_1","volume-title":"Robots Learn Social Skills: End-to-End Learning of Co-Speech Gesture Generation for Humanoid Robots. In International Conference on Robotics and Automation. IEEE, 4303\u20134309","author":"Yoon Youngwoo","year":"2019","unstructured":"Youngwoo Yoon , Woo-Ri Ko , Minsu Jang , Jaeyeon Lee , Jaehong Kim , and Geehyuk Lee . 2019 . Robots Learn Social Skills: End-to-End Learning of Co-Speech Gesture Generation for Humanoid Robots. In International Conference on Robotics and Automation. IEEE, 4303\u20134309 . Youngwoo Yoon, Woo-Ri Ko, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee. 2019. Robots Learn Social Skills: End-to-End Learning of Co-Speech Gesture Generation for Humanoid Robots. In International Conference on Robotics and Automation. IEEE, 4303\u20134309."}],"event":{"name":"UIST '21: The 34th Annual ACM Symposium on User Interface Software and Technology","location":"Virtual Event USA","acronym":"UIST '21","sponsor":["SIGGRAPH ACM Special Interest Group on Computer Graphics and Interactive Techniques","SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["The 34th Annual ACM Symposium on User Interface Software and Technology"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3472749.3474789","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3472749.3474789","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:17:09Z","timestamp":1750191429000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3472749.3474789"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,10]]},"references-count":45,"alternative-id":["10.1145\/3472749.3474789","10.1145\/3472749"],"URL":"https:\/\/doi.org\/10.1145\/3472749.3474789","relation":{},"subject":[],"published":{"date-parts":[[2021,10,10]]},"assertion":[{"value":"2021-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}