{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T11:38:03Z","timestamp":1786534683655,"version":"3.56.0"},"reference-count":41,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T00:00:00Z","timestamp":1786492800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T00:00:00Z","timestamp":1786492800000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100003977","name":"Israel Science Foundation","doi-asserted-by":"publisher","award":["1427\/25"],"award-info":[{"award-number":["1427\/25"]}],"id":[{"id":"10.13039\/501100003977","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>While recent advances in co\u2010speech gesture generation have achieved impressive rhythmic synchronization, synthesizing gestures that are both semantically meaningful and faithful to a speaker's unique non\u2010verbal style remains an open challenge. Semantic gestures, such as iconic shapes or deictic pointing, are statistically sparse, making them difficult to learn effectively within standard generative models. We present SiGnature, a framework for Stylized and Semantic Gesture generation that reconciles precise semantic control with high\u2010fidelity style preservation.<\/jats:p>\n                  <jats:p>Unlike prevalent methods that rely on entangled latent representations, SiGnature operates in an explicit joint\u2010rotation space. This design enables our core contribution, Joint Motion Integration (JMI), a training\u2010free inference mechanism capable of injecting any external motion sequence, particularly in\u2010the\u2010wild semantic gestures, directly into the diffusion process. JMI automatically identifies the specific \u201cactive joints\u201d conveying a semantic action and injects them into the generation, while relying on the diffusion backbone to synthesize the remaining body dynamics, including posture and flow, in accordance with the pre\u2010learned style of the target speaker. This allows for the plug\u2010and\u2010play integration of arbitrary motions, including complex semantic gestures, without retraining or introducing the \u201cFrankenstein\u201d artifacts typical of cut\u2010and\u2010paste methods. Extensive experiments and perceptual studies demonstrate that SiGnature offers superior semantic motion control while maintaining smooth and natural co\u2010speech gesture generation and preserving the distinct characteristics of the speaker, thereby outperforming state\u2010of\u2010the\u2010art baselines.<\/jats:p>","DOI":"10.1111\/cgf.70557","type":"journal-article","created":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T11:07:31Z","timestamp":1786532851000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["SiGnature: Explicit Motion Diffusion for Stylized Semantic Gesture Generation"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-6567-696X","authenticated-orcid":false,"given":"Adi","family":"Rosenthal","sequence":"first","affiliation":[{"name":"Reichman University  Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-8472-5727","authenticated-orcid":false,"given":"Tomer","family":"Koren","sequence":"additional","affiliation":[{"name":"Reichman University  Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-5615-9478","authenticated-orcid":false,"given":"Nadav","family":"Shaked","sequence":"additional","affiliation":[{"name":"Reichman University  Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2584-044X","authenticated-orcid":false,"given":"Doron","family":"Friedman","sequence":"additional","affiliation":[{"name":"Reichman University  Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7082-7845","authenticated-orcid":false,"given":"Ariel","family":"Shamir","sequence":"additional","affiliation":[{"name":"Reichman University  Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2026,8,12]]},"reference":[{"key":"e_1_2_8_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2019.00084"},{"key":"e_1_2_8_3_2","unstructured":"Ao Tenglong Zhang Zeyi andLiu Libin. \u201cGestureDiffu\u2010CLIP: Gesture Diffusion Model with CLIP Latents\u201d.ACM Trans. Graph. (). doi:10.1145\/35920973."},{"key":"e_1_2_8_4_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1002183"},{"key":"e_1_2_8_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/332051.332075"},{"key":"e_1_2_8_6_2","doi-asserted-by":"crossref","unstructured":"Chhatre Kiran Dan\u011b\u010dek Radek Athanasiou Nikos et al. \u201cAMUSE: Emotional Speech\u2010driven 3D Body Animation via Disentangled Latent Diffusion\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June2024 1942\u20131953. url:https:\/\/amuse.is.tue.mpg.de3.","DOI":"10.1109\/CVPR52733.2024.00190"},{"key":"e_1_2_8_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3680847"},{"key":"e_1_2_8_8_2","doi-asserted-by":"crossref","unstructured":"Cheng Qingrong Li Xu andFu Xinghui. \u201cSiggesture: Generalized co\u2010speech gesture synthesis via semantic injection with large\u2010scale pre\u2010training diffusion models\u201d.SIGGRAPH Asia 2024 Conference Papers.2024 1\u2013112 3.","DOI":"10.1145\/3680528.3687677"},{"key":"e_1_2_8_9_2","unstructured":"Chen Junming Liu Yunfei Wang Jianan et al. \u201cDiff\u2010SHEG: A Diffusion\u2010Based Approach for Real\u2010Time Speech\u2010driven Holistic 3D Expression and Gesture Generation\u201d.CVPR.20242."},{"key":"e_1_2_8_10_2","unstructured":"Chen Changan Zhang Juze Lakshmikanth Shrinidhi Kowshika et al. \u201cThe Language of Motion: Unifying Verbal and Non\u2010verbal Language of 3D Human Motion\u201d.arXiv.20242."},{"key":"e_1_2_8_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3610548.3618183"},{"key":"e_1_2_8_12_2","doi-asserted-by":"publisher","DOI":"10.1523\/JNEUROSCI.05-07-01688.1985"},{"key":"e_1_2_8_13_2","doi-asserted-by":"crossref","unstructured":"Ginosar Shiry Bar Amir Kohavi Gefen et al. \u201cLearning individual styles of conversational gesture\u201d.Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition.2019 3497\u201335062.","DOI":"10.1109\/CVPR.2019.00361"},{"key":"e_1_2_8_14_2","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.14734"},{"key":"e_1_2_8_14_3","unstructured":"eprint:https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.14734."},{"key":"e_1_2_8_14_4","unstructured":"url:https:\/\/onlinelibrary.wiley.com\/doi\/abs\/10.1111\/cgf.147343."},{"key":"e_1_2_8_15_2","first-page":"259","volume-title":"The development of social cognition and communication","author":"Goldin\u2010Meadow Susan","year":"2013"},{"key":"e_1_2_8_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3656374"},{"key":"e_1_2_8_17_2","doi-asserted-by":"crossref","unstructured":"Li Jing Kang Di Pei Wenjie et al. \u201cAudio2gestures: Generating diverse gestures from speech audio with conditional variational autoencoders\u201d.Proceedings of the IEEE\/CVF International Conference on Computer Vision.2021 11293\u2013113022 6.","DOI":"10.1109\/ICCV48922.2021.01110"},{"key":"e_1_2_8_18_2","doi-asserted-by":"crossref","unstructured":"Liu Pinxin Song Luchuan Huang Junhua et al. \u201cGestureLSM: Latent Shortcut based Co\u2010Speech Gesture Generation with Spatial\u2010Temporal Modeling\u201d.arXiv preprint arXiv:2501.18898(2025) 2 6.","DOI":"10.1109\/ICCV51701.2025.01017"},{"key":"e_1_2_8_19_2","doi-asserted-by":"crossref","unstructured":"Liu Xian Wu Qianyi Zhou Hang et al. \u201cLearning hierarchical cross\u2010modal association for co\u2010speech gesture generation\u201d.Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition.2022 10462\u2013104722.","DOI":"10.1109\/CVPR52688.2022.01021"},{"key":"e_1_2_8_20_2","doi-asserted-by":"crossref","unstructured":"Liu Haiyang Zhu Zihao Becherini Giorgio et al. \u201cEmage: Towards unified holistic co\u2010speech gesture generation via expressive masked audio gesture modeling\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 1144\u201311542 5 6.","DOI":"10.1109\/CVPR52733.2024.00115"},{"key":"e_1_2_8_21_2","first-page":"612","volume-title":"European conference on computer vision","author":"Liu Haiyang","year":"2022"},{"key":"e_1_2_8_22_2","doi-asserted-by":"crossref","unstructured":"Mughal M Hamza Dabral Rishabh Scholman Merel CJ et al. \u201cRetrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis\u201d.Proceedings of the Computer Vision and Pattern Recognition Conference.2025 16578\u2013165883.","DOI":"10.1109\/CVPR52734.2025.01545"},{"key":"e_1_2_8_23_2","doi-asserted-by":"crossref","unstructured":"Mahmood Naureen Ghorbani Nima Troje Nikolaus F et al. \u201cAMASS: Archive of motion capture as surface shapes\u201d.Proceedings of the IEEE\/CVF international conference on computer vision.2019 5442\u201354515.","DOI":"10.1109\/ICCV.2019.00554"},{"key":"e_1_2_8_24_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11671"},{"key":"e_1_2_8_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3407692"},{"key":"e_1_2_8_26_2","unstructured":"Shafir Yoni Tevet Guy Kapon Roy andBermano Amit Haim. \u201cHuman Motion Diffusion as a Generative Prior\u201d.The Twelfth International Conference on Learning Representations.20242 4 8."},{"key":"e_1_2_8_27_2","doi-asserted-by":"crossref","unstructured":"Siyao Li Yu Weijiang Gu Tianpei et al. \u201cBailando: 3d dance generation by actor\u2010critic gpt with choreographic memory\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2022 11050\u2013110596.","DOI":"10.1109\/CVPR52688.2022.01077"},{"key":"e_1_2_8_28_2","first-page":"358","volume-title":"European Conference on Computer Vision (ECCV)","author":"Tevet Guy","year":"2022"},{"key":"e_1_2_8_29_2","unstructured":"Tevet Guy Raab Sigal Gordon Brian et al. \u201cHuman Motion Diffusion Model\u201d.The Eleventh International Conference on Learning Representations.2023. url:https:\/\/openreview.net\/forum?id=SJ1kSyO2jwu2."},{"key":"e_1_2_8_30_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-0633"},{"key":"e_1_2_8_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417838"},{"key":"e_1_2_8_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417827"},{"key":"e_1_2_8_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8793720"},{"key":"e_1_2_8_34_2","unstructured":"Yi Hongwei Liang Hualin Liu Yifei et al. \u201cGenerating Holistic 3D Human Motion from Speech\u201d.CVPR.20232."},{"key":"e_1_2_8_35_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2023\/650"},{"key":"e_1_2_8_35_3","unstructured":"url:https:\/\/doi.org\/10.24963\/ijcai.2023\/6502 3."},{"key":"e_1_2_8_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3658134"},{"key":"e_1_2_8_37_2","doi-asserted-by":"crossref","unstructured":"Zhi Yihao Cun Xiaodong Chen Xuelin et al. \u201cLivelyspeaker: Towards semantic\u2010aware co\u2010speech gesture generation\u201d.Proceedings of the IEEE\/CVF international conference on computer vision.2023 20807\u2013208172.","DOI":"10.1109\/ICCV51070.2023.01902"},{"key":"e_1_2_8_38_2","doi-asserted-by":"crossref","unstructured":"Zhu Lingting Liu Xian Liu Xuanyu et al. \u201cTaming diffusion models for audio\u2010driven co\u2010speech gesture generation\u201d.Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition.2023 10544\u2013105532.","DOI":"10.1109\/CVPR52729.2023.01016"},{"key":"e_1_2_8_39_2","doi-asserted-by":"crossref","unstructured":"Zhang Xiangyue Li Jianfang Zhang Jiaxu et al. \u201cSemTalk: Holistic Co\u2010speech Motion Generation with Frame\u2010level Semantic Emphasis\u201d.Proceedings of the IEEE\/CVF International Conference on Computer Vision.2025 13761\u2013137713.","DOI":"10.1109\/ICCV51701.2025.01277"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70557","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70557","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70557","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T11:07:38Z","timestamp":1786532858000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70557"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,8,12]]},"references-count":41,"alternative-id":["10.1111\/cgf.70557"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70557","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,8,12]]},"assertion":[{"value":"2026-08-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70557"}}