{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,15]],"date-time":"2026-04-15T10:55:48Z","timestamp":1776250548362,"version":"3.50.1"},"reference-count":55,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,4,15]],"date-time":"2026-04-15T00:00:00Z","timestamp":1776211200000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,4,15]],"date-time":"2026-04-15T00:00:00Z","timestamp":1776211200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100003977","name":"Israel Science Foundation","doi-asserted-by":"publisher","award":["1427\/25"],"award-info":[{"award-number":["1427\/25"]}],"id":[{"id":"10.13039\/501100003977","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>In this work we extend speaker\u2010centric audio\u2010driven gesture synthesis toward a unified conversational model that jointly captures both speaking and listening behaviors. Existing speaker\u2010centric models effectively generate gestures aligned with speech but overlook the bidirectional dynamics that characterize natural dialogue. To address this limitation, we propose the Conversational Gesture Model (CGM), a cross\u2010attention\u2010based model capable of synthesizing gestures conditioned on interlocutor conversational cues such as gestures, tone, and textual semantics. By leveraging cross\u2010attention mechanisms, the model fuses interlocutor audio and text features with character gesture encodings, enabling a single system to seamlessly alternate between speaking and listening roles of the same character. Hence, our model enables a single system to act as both speaker and listener, capturing the fluid role shifts and mutual influence inherent in conversation. Experiments demonstrate that this approach preserves the quality of speaker\u2010driven gestures while significantly improving the realism, coherence, and responsiveness of full conversational interactions.<\/jats:p>","DOI":"10.1111\/cgf.70412","type":"journal-article","created":{"date-parts":[[2026,4,15]],"date-time":"2026-04-15T10:07:25Z","timestamp":1776247645000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Conversational Gesture Model (CGM): Extending Speaker\u2010Centric Audio\u2010Driven Motion Generation to Full Conversation Gestures"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-8472-5727","authenticated-orcid":false,"given":"T.","family":"Koren","sequence":"first","affiliation":[{"name":"School of Computer Science Reichman University  Israel"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6567-696X","authenticated-orcid":false,"given":"A.","family":"Rosenthal","sequence":"additional","affiliation":[{"name":"School of Computer Science Reichman University  Israel"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2584-044X","authenticated-orcid":false,"given":"D.","family":"Friedman","sequence":"additional","affiliation":[{"name":"School of Communication Reichman University  Israel"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7082-7845","authenticated-orcid":false,"given":"A.","family":"Shamir","sequence":"additional","affiliation":[{"name":"School of Computer Science Reichman University  Israel"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,4,15]]},"reference":[{"key":"e_1_2_8_2_2","article-title":"Head movements as markers of conversational alignment: A systematic review","volume":"14","author":"Albuquerque Isadora","year":"2023","journal-title":"Frontiers in Psychology"},{"key":"e_1_2_8_3_2","doi-asserted-by":"crossref","unstructured":"Ahuja Chaitanya Ma Shih-Yun Morency LouisPhilippe andSheikh Yaser. \u201cTo react or not to react: End-to-end visual pose forecasting for personalized avatar during dyadic conversations\u201d.Proceedings of the 2019 International Conference on Multimodal Interaction (ICMI).2019 74\u2013842.","DOI":"10.1145\/3340555.3353725"},{"key":"e_1_2_8_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3592458"},{"key":"e_1_2_8_5_2","doi-asserted-by":"crossref","unstructured":"Ao Tenglong Zhang Zeyi andLiu Libin. \u201cGestureDiffuCLIP: Gesture diffusion model with CLIP latents\u201d.ACM Transactions on Graphics(2023). doi:10.1145\/35920972.","DOI":"10.1145\/3592097"},{"key":"e_1_2_8_6_2","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1162\/tacl_a_00051","article-title":"Enriching word vectors with subword information","volume":"5","author":"Bojanowski Piotr","year":"2017","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"e_1_2_8_7_2","unstructured":"Bruderlin AlainandJun Sung Yong. \u201cConversational gestures and head movements synthesis for embodied agents\u201d.Proceedings of the 2008 ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games.2008 139\u20131478."},{"key":"e_1_2_8_8_2","doi-asserted-by":"crossref","unstructured":"Chhatre Kiran Dan\u011b\u010dek Radek Athanasiou Nikos et al. \u201cAMUSE: Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June2024 1942\u20131953. url:https:\/\/amuse.is.tue.mpg.de2.","DOI":"10.1109\/CVPR52733.2024.00190"},{"key":"e_1_2_8_9_2","doi-asserted-by":"crossref","unstructured":"Chen Bohong Li Yumeng Ding Yao-Xiang et al. \u201cEnabling synergistic full-body control in prompt-based co-speech motion generation\u201d.Proceedings of the 32nd ACM International Conference on Multimedia.2024 6774\u201367832 3 7 8.","DOI":"10.1145\/3664647.3680847"},{"key":"e_1_2_8_10_2","unstructured":"Chen Junming Liu Yunfei Wang Jianan et al. \u201cDiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation\u201d.CVPR.20242."},{"key":"e_1_2_8_11_2","doi-asserted-by":"crossref","unstructured":"Dan\u011b\u010dek Radek Chhatre Kiran Tripathi Shashank et al. \u201cEmotional speech-driven animation with content-emotion disentanglement\u201d.Proceedings of the 31st ACM International Conference on Multimedia (MM '23).2023. doi:10.1145\/3610548.36181832.","DOI":"10.1145\/3610548.3618183"},{"key":"e_1_2_8_12_2","doi-asserted-by":"crossref","unstructured":"Deichler Anna Mehta Shivam Alexanderson Simon andBeskow Jonas. \u201cDiffusion-Based Co-Speech Gesture Generation Using Joint Text and Audio Representation\u201d.Proceedings of the ACM International Conference on Multimodal Interaction (ICMI '23).2023. doi:10.1145\/3577190.36161173.","DOI":"10.1145\/3577190.3616117"},{"key":"e_1_2_8_13_2","unstructured":"Dhariwal PrafullaandNichol Alex. \u201cDiffusion models beat GANs on image synthesis\u201d.Advances in Neural Information Processing Systems (NeurIPS).20212."},{"key":"e_1_2_8_14_2","doi-asserted-by":"crossref","unstructured":"Girshick Ross. \u201cFast R-CNN\u201d.Proceedings of the IEEE International Conference on Computer Vision (ICCV).2015 1440\u201314485.","DOI":"10.1109\/ICCV.2015.169"},{"issue":"5","key":"e_1_2_8_15_2","doi-asserted-by":"crossref","first-page":"579","DOI":"10.1002\/cav.267","article-title":"Responsive listening behavior","volume":"19","author":"Gillies Michael","year":"2008","journal-title":"Computer Animation and Virtual Worlds"},{"key":"e_1_2_8_16_2","first-page":"241","volume-title":"International Conference on Intelligent Virtual Agents","author":"Heylen Dirk","year":"2006"},{"key":"e_1_2_8_17_2","unstructured":"Heusel Martin Ramsauer Hubert Unterthiner Thomas et al. \u201cGANs trained by a two time-scale update rule converge to a Nash equilibrium\u201d.Advances in Neural Information Processing Systems (NeurIPS).20178."},{"key":"e_1_2_8_18_2","doi-asserted-by":"crossref","unstructured":"He Kaiming Zhang Xiangyu Ren Shaoqing andSun Jian. \u201cDeep residual learning for image recognition\u201d.Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).2016 770\u20137784.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_8_19_2","doi-asserted-by":"crossref","unstructured":"Jonell Patrik Kucherenko Taras Henter Gustav Eje andBeskow Jonas. \u201cLet's face it: Probabilistic multi-modal interlocutor-aware generation of facial gestures in dyadic settings\u201d.Proceedings of the 20th ACM International Conference on Intelligent Virtual Agents (IVA).2020 1\u201382.","DOI":"10.1145\/3383652.3423911"},{"key":"e_1_2_8_20_2","doi-asserted-by":"crossref","unstructured":"Joo Hanbyul Simon Tomas Cikara Mina andSheikh Yaser. \u201cTowards social artificial intelligence: Nonverbal social signal prediction in a triadic interaction\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).2019 10873\u2013108832.","DOI":"10.1109\/CVPR.2019.01113"},{"key":"e_1_2_8_21_2","doi-asserted-by":"crossref","unstructured":"Kucherenko Taras Nagy Rich\u00e1rd Yoon Youngwoo et al. \u201cThe GENEA challenge2023: A large-scale evaluation of gesture generation models in monadic and dyadic settings\u201d.Proceedings of the 25th International Conference on Multimodal Interaction (ICMI).2023 792\u20138013.","DOI":"10.1145\/3577190.3616120"},{"key":"e_1_2_8_22_2","doi-asserted-by":"crossref","unstructured":"Lee Gwan Deng Zijian Ma Shih-Yun et al. \u201cTalking with hands 16.2M: A large-scale dataset of synchronized body-finger motion and audio for conversational motion analysis and synthesis\u201d.Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV).2019 763\u20137727 10.","DOI":"10.1109\/ICCV.2019.00085"},{"key":"e_1_2_8_23_2","doi-asserted-by":"crossref","unstructured":"Liu Jiahao Wang Xuran Fu Xiaoyu et al. \u201cMFR-Net: Multi-faceted responsive listening head generation via denoising diffusion model\u201d.Proceedings of the 31st ACM International Conference on Multimedia (MM '23).2023 6734\u201367432.","DOI":"10.1145\/3581783.3612123"},{"key":"e_1_2_8_24_2","unstructured":"Li Siyao Yu Weijiang Gu Tianpei et al. \u201cBailando: 3D dance generation via actor-critic GPT with choreographic memory\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).20228."},{"key":"e_1_2_8_25_2","doi-asserted-by":"crossref","unstructured":"Liu Haiyang Zhu Zihao Becherini Giorgio et al. \u201cEMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).2024 1144\u201311542 3 7 8 11.","DOI":"10.1109\/CVPR52733.2024.00115"},{"key":"e_1_2_8_26_2","unstructured":"Liu Haiyang Zhu Zihao Iwamoto Naoya et al. \u201cBEAT: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis\u201d.European Conference on Computer Vision (ECCV).20228."},{"issue":"7","key":"e_1_2_8_27_2","doi-asserted-by":"crossref","first-page":"855","DOI":"10.1016\/S0378-2166(99)00079-X","article-title":"Linguistic functions of head movements in the context of speech","volume":"32","author":"McClave Evelyn","year":"2000","journal-title":"Journal of Pragmatics"},{"key":"e_1_2_8_28_2","doi-asserted-by":"crossref","unstructured":"Mughal Muhammad Hamza Dabral Rishabh Habibie Ikhsanul et al. \u201cConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).2024 1388\u201313983.","DOI":"10.1109\/CVPR52733.2024.00138"},{"key":"e_1_2_8_29_2","doi-asserted-by":"crossref","unstructured":"Ng Eshrat Arjmand Joo Hanbyul Hu Liwen et al. \u201cLearning to listen: Modeling non-deterministic dyadic facial motion\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).2022 20395\u2013204052.","DOI":"10.1109\/CVPR52688.2022.01975"},{"key":"e_1_2_8_30_2","doi-asserted-by":"crossref","unstructured":"Ng Evonne Romero Javier Bagautdinov Timur et al. \u201cFrom Audio to Photoreal Embodiment: Synthesizing Humans in Conversations\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).2024 1144\u201311543 8.","DOI":"10.1109\/CVPR52733.2024.00101"},{"key":"e_1_2_8_31_2","unstructured":"Ng Jonathan Wang Jialin Xu Ruiqi et al. \u201cGesture-Dialogue: Co-speech gesture synthesis for multi-party conversations\u201d.Proceedings of the ACM International Conference on Multimodal Interaction (ICMI).20241."},{"key":"e_1_2_8_32_2","first-page":"4","volume-title":"Proceedings of the 2nd Workshop on Understanding Social Behavior in Dyadic and Small Group Interactions","author":"Palmero Carmen","year":"2022"},{"key":"e_1_2_8_33_2","unstructured":"Pavlakos Georgios Choutas Vasileios Ghorbani Nima et al. \u201cExpressive body capture: 3D hands face and body from a single image\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).20197."},{"key":"e_1_2_8_34_2","unstructured":"Qi Xingqun Wang Yatian Zhang Hengyuan et al. \u201cCo3Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion\u201d.International Conference on Learning Representations (ICLR).20253."},{"key":"e_1_2_8_35_2","volume-title":"ACM SIGGRAPH 2022 Conference Proceedings","author":"Shuai Qing","year":"2022"},{"key":"e_1_2_8_36_2","doi-asserted-by":"crossref","unstructured":"Song Shuran Spitale Michele Luo Chengzhu et al. \u201cREACT2023: The first multi-modal multiple appropriate facial reaction generation challenge\u201d.arXiv preprint arXiv:2306.06583(2023) 2.","DOI":"10.1145\/3581783.3612832"},{"key":"e_1_2_8_37_2","doi-asserted-by":"crossref","unstructured":"Song Li Yin Guoxian Jin Zhiwei et al. \u201cEmotional listener portrait: Realistic listener motion simulation in conversation\u201d.arXiv preprint arXiv:2310.00068(2023) 2.","DOI":"10.1109\/ICCV51070.2023.01905"},{"key":"e_1_2_8_38_2","doi-asserted-by":"crossref","unstructured":"Tuyen Tuyen Nguyen ThiandCeliktutan Ozgur. \u201cAgree or disagree? Generating body gestures from affective contextual cues during dyadic interactions\u201d.Proceedings of the 31st IEEE International Conference on Robot and Human Interactive Communication (RO-MAN).2022 1542\u201315473.","DOI":"10.1109\/RO-MAN53752.2022.9900760"},{"issue":"24","key":"e_1_2_8_39_2","doi-asserted-by":"crossref","first-page":"1552","DOI":"10.1080\/01691864.2023.2279595","article-title":"It takes two, not one: Context-aware nonverbal behaviour generation in dyadic interactions","volume":"37","author":"Tuyen Tuyen Nguyen Thi","year":"2023","journal-title":"Advanced Robotics"},{"key":"e_1_2_8_40_2","unstructured":"Tuyen Tuyen Nguyen ThiandCeliktutan Ozgur. \u201cIt Takes Two: Context-Aware Nonverbal Behaviour Generation for HumanHuman Interaction\u201d.arXiv preprint arXiv:2510.10206(2025) 3."},{"key":"e_1_2_8_41_2","first-page":"358","volume-title":"European Conference on Computer Vision (ECCV)","author":"Tevet Guy","year":"2022"},{"key":"e_1_2_8_42_2","unstructured":"Tevet Guy Raab Sigal Gordon Brian et al. \u201cHuman motion diffusion model\u201d.International Conference on Learning Representations (ICLR).20235."},{"key":"e_1_2_8_43_2","unstructured":"Van den Oord Aaron Vinyals Oriol andKavukcuoglu Koray. \u201cNeural discrete representation learning\u201d.Advances in Neural Information Processing Systems (NeurIPS).20173."},{"key":"e_1_2_8_44_2","unstructured":"Wightman Ross.PyTorch Image Models.https:\/\/github.com\/rwightman\/pytorch-image-models. Accessed: 2025-09-26.20194."},{"key":"e_1_2_8_45_2","doi-asserted-by":"crossref","unstructured":"Xue Haiwei Fan Yanbo Wang Xuan andWu Zhiyong. \u201cEcho: Enhancing Conversational Behavior Generation via Hierarchical Semantic Comprehension with Large Language Models\u201d.Proceedings of SIGGRAPH Asia 2025.20253.","DOI":"10.1145\/3757377.3763998"},{"key":"e_1_2_8_46_2","unstructured":"Xu Sirui Li Dongting Zhang Yucheng et al. \u201cInterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).20253."},{"key":"e_1_2_8_47_2","doi-asserted-by":"crossref","unstructured":"Yang Sicheng Wu Zhiyong Li Minglei et al. \u201cDiffuseStyleGesture: Stylized audio-driven co-speech gesture generation with diffusion models\u201d.Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence (IJCAI-23).2023 5860\u20135868. doi:10.24963\/ijcai.2023\/6502 8.","DOI":"10.24963\/ijcai.2023\/650"},{"key":"e_1_2_8_48_2","doi-asserted-by":"crossref","unstructured":"Yang Sicheng Wang Zilin Wu Zhiyong et al. \u201cUnifiedGesture: A unified gesture synthesis model for multiple skeletons\u201d.Proceedings of the 31st ACM International Conference on Multimedia (MM '23).2023 1033\u20131044. doi:10.1145\/3581783.36125032.","DOI":"10.1145\/3581783.3612503"},{"key":"e_1_2_8_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3577190.3616114"},{"key":"e_1_2_8_49_3","unstructured":"url:https:\/\/doi.org\/10.1145\/3577190.36161141 8 9."},{"key":"e_1_2_8_50_2","doi-asserted-by":"crossref","unstructured":"Zhang Zeyi Ao Tenglong Zhang Yuyao et al. \u201cSemantic gesticulator: Semantics-aware co-speech gesture synthesis\u201d.arXiv preprint arXiv:2405.09814(2024) 2.","DOI":"10.1145\/3658134"},{"key":"e_1_2_8_51_2","first-page":"124","volume-title":"European Conference on Computer Vision (ECCV)","author":"Zhou Mingyuan","year":"2022"},{"key":"e_1_2_8_52_2","unstructured":"Zhang Yifan Chen Yuwei Ling Hongwei et al. \u201cConversational gesture synthesis with neural motion fields\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).20241."},{"key":"e_1_2_8_53_2","unstructured":"Zhang Minghao Ferreira Rafael Talman Arthur et al. \u201cGestures are in the eye of the beholder: Contrastive motion learning from listener's gaze\u201d.International Conference on Learning Representations (ICLR).20243."},{"key":"e_1_2_8_54_2","doi-asserted-by":"crossref","unstructured":"Zhu Lingting Liu Xian Liu Xuanyu et al. \u201cTaming diffusion models for audio-driven co-speech gesture generation\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).2023 10544\u2013105532.","DOI":"10.1109\/CVPR52729.2023.01016"},{"key":"e_1_2_8_55_2","unstructured":"Zhang Jianrong Zhang Yangsong Cun Xiaodong et al. \u201cT2M-GPT: Generating human motion from textual descriptions with discrete representations\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).20233."}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70412","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70412","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70412","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,15]],"date-time":"2026-04-15T10:07:59Z","timestamp":1776247679000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70412"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,15]]},"references-count":55,"alternative-id":["10.1111\/cgf.70412"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70412","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,15]]},"assertion":[{"value":"2026-04-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70412"}}