{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T15:06:35Z","timestamp":1781363195151,"version":"3.54.1"},"reference-count":127,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,7,31]],"date-time":"2025-07-31T00:00:00Z","timestamp":1753920000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100010665","name":"H2020 Marie Sk\u0142odowska-Curie Actions","doi-asserted-by":"publisher","award":["860768"],"award-info":[{"award-number":["860768"]}],"id":[{"id":"10.13039\/100010665","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Comput. Sci."],"abstract":"<jats:sec><jats:title>Introduction<\/jats:title><jats:p>Social interactions incorporate various nonverbal signals to convey emotions alongside speech, including facial expressions and body gestures. Generative models have demonstrated promising results in creating full-body nonverbal animations synchronized with speech; however, evaluations using statistical metrics in 2D settings fail to fully capture user-perceived emotions, limiting our understanding of the effectiveness of these models.<\/jats:p><\/jats:sec><jats:sec><jats:title>Methods<\/jats:title><jats:p>To address this, we evaluate emotional 3D animation generative models within an immersive Virtual Reality (VR) environment, emphasizing user\u2014centric metrics-emotional arousal realism, naturalness, enjoyment, diversity, and interaction quality\u2014in a real-time human-agent interaction scenario. Through a user study (<jats:italic>N<\/jats:italic> = 48), we systematically examine perceived emotional quality for three state-of-the-art speech-driven 3D animation methods across two specific emotions: happiness (high arousal) and neutral (mid arousal). Additionally, we compare these generative models against real human expressions obtained via a reconstruction-based method to assess both their strengths and limitations and how closely they replicate real human facial and body expressions.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>Our results demonstrate that methods explicitly modeling emotions lead to higher recognition accuracy compared to those focusing solely on speech-driven synchrony. Users rated the realism and naturalness of happy animations significantly higher than those of neutral animations, highlighting the limitations of current generative models in handling subtle emotional states.<\/jats:p><\/jats:sec><jats:sec><jats:title>Discussion<\/jats:title><jats:p>Generative models underperformed compared to reconstruction-based methods in facial expression quality, and all methods received relatively low ratings for animation enjoyment and interaction quality, emphasizing the importance of incorporating user-centric evaluations into generative model development. Finally, participants positively recognized animation diversity across all generative models.<\/jats:p><\/jats:sec>","DOI":"10.3389\/fcomp.2025.1598099","type":"journal-article","created":{"date-parts":[[2025,7,31]],"date-time":"2025-07-31T05:39:13Z","timestamp":1753940353000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Evaluation of generative models for emotional 3D animation generation in VR"],"prefix":"10.3389","volume":"7","author":[{"given":"Kiran","family":"Chhatre","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Renan","family":"Guarese","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrii","family":"Matviienko","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christopher","family":"Peters","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2025,7,31]]},"reference":[{"key":"B1","unstructured":"Meet Your Soul Machines AI Assistants"},{"key":"B2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3592458","article-title":"Listen, denoise, action! Audio-driven motion synthesis with diffusion models","volume":"42","author":"Alexanderson","year":"2023","journal-title":"ACM Trans. Graph"},{"key":"B3","doi-asserted-by":"publisher","first-page":"1619","DOI":"10.1016\/j.neuropsychologia.2013.03.022","article-title":"Brain function overlaps when people observe emblems, speech, and grasping","volume":"51","author":"Andric","year":"2013","journal-title":"Neuropsychologia"},{"key":"B4","first-page":"12449","article-title":"\u201cwav2vec 2.0: A framework for self-supervised learning of speech representations,\u201d","author":"Baevski","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B5","doi-asserted-by":"publisher","first-page":"107047","DOI":"10.1016\/j.chb.2021.107047","article-title":"Psychological benefits of using social virtual reality platforms during the covid-19 pandemic: the role of social and spatial presence","volume":"127","author":"Barreda-\u00c1ngeles","year":"2021","journal-title":"Comput. Human Behav"},{"key":"B6","doi-asserted-by":"publisher","first-page":"178","DOI":"10.1016\/j.neuropsychologia.2005.05.007","article-title":"Speech and gesture share the same communication system","volume":"44","author":"Bernardis","year":"2006","journal-title":"Neuropsychologia"},{"key":"B7","doi-asserted-by":"publisher","first-page":"456","DOI":"10.1162\/105474603322761270","article-title":"Toward a more robust theory and measure of social presence: review and suggested criteria","volume":"12","author":"Biocca","year":"2003","journal-title":"Presence: Teleoperators and Virtual Environments"},{"key":"B8","unstructured":"Bolkart\n              T.\n            \n          \n          Tf flame: Tensorflow Framework for the Flame 3d Head Model\n          \n          2013"},{"key":"B9","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1037\/0278-7393.18.2.379","article-title":"Remembering pictures: pleasure and arousal in memory","volume":"18","author":"Bradley","year":"1992","journal-title":"J. Exp. Psychol. Learn. Memory Cognit"},{"key":"B10","doi-asserted-by":"crossref","DOI":"10.1145\/3611659.3615695","article-title":"\u201cDialogues for one: Single-user content creation using immersive record and replay,\u201d","volume-title":"Proceedings of the 29th ACM Symposium on Virtual Reality Software and Technology","author":"Brandst\u00e4tter","year":"2023"},{"key":"B11","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2112.02418","article-title":"YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone","author":"Casanova","year":"2021","journal-title":"arXiv"},{"key":"B12","doi-asserted-by":"publisher","first-page":"70","DOI":"10.1145\/332051.332075","article-title":"Embodied conversational interface agents","volume":"43","author":"Cassell","year":"2000","journal-title":"Commun. ACM"},{"key":"B13","doi-asserted-by":"crossref","DOI":"10.1145\/383259.383315","article-title":"\u201cBeat: the behavior expression animation toolkit,\u201d","volume-title":"Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques","author":"Cassell","year":"2001"},{"key":"B14","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2502.13133","article-title":"Av-flow: Transforming text to audio-visual human-like interactions","author":"Chatziagapi","year":"2025","journal-title":"arXiv"},{"key":"B15","first-page":"1942","article-title":"\u201cAMUSE: Emotional speech-driven 3D body animation via disentangled latent diffusion,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Chhatre","year":"2024"},{"key":"B16","doi-asserted-by":"crossref","DOI":"10.1145\/3722564.3728374","article-title":"\u201cEvaluating speech and video models forface-body congruence,\u201d","volume-title":"Companion Proceedings of the ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games","author":"Chhatre","year":"2025"},{"key":"B17","doi-asserted-by":"publisher","first-page":"1059","DOI":"10.1016\/j.psychres.2018.04.066","article-title":"The racially diverse affective expression (radiate) face stimulus set","volume":"270","author":"Conley","year":"2018","journal-title":"Psychiatry Res"},{"key":"B18","unstructured":"33501025\n          Embodied ai Characters for Virtual Worlds\n          \n          2025"},{"key":"B19","doi-asserted-by":"publisher","first-page":"5432c","DOI":"10.31234\/osf.io\/5432c","article-title":"Universals and diversity in gesture: research past, present, and future","volume":"18","author":"Cooperrider","year":"2020","journal-title":"Gesture"},{"key":"B20","first-page":"10101","article-title":"\u201cCapture, learning, and synthesis of 3d speaking styles,\u201d","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019","author":"Cudeiro","year":"2019"},{"key":"B21","doi-asserted-by":"crossref","DOI":"10.1145\/3610548.3618183","volume-title":"Emotional Speech-Driven Animation With Content-Emotion Disentanglement","author":"Dan\u0115\u010dek","year":""},{"key":"B22","doi-asserted-by":"crossref","DOI":"10.1145\/3610548.3618183","volume-title":"Emotional Speech-Driven Animation With Content-Emotion Disentanglement","author":"Dan\u0115\u010dek","year":""},{"key":"B23","doi-asserted-by":"publisher","first-page":"2063","DOI":"10.3389\/fpsyg.2019.02063","article-title":"Language, gesture, and emotional communication: an embodied view of social interaction","volume":"10","author":"De Stefani","year":"2019","journal-title":"Front. Psychol"},{"key":"B24","article-title":"\u201cGesture evaluation in virtual reality,\u201d","volume-title":"GENEA: Generation and Evaluation of Non-verbal behavior for Embodied Agents Workshop 2024","author":"Deichler","year":"2024"},{"key":"B25","doi-asserted-by":"crossref","DOI":"10.1145\/3656650.3656691","article-title":"\u201cA user study on the relationship between empathy and facial-based emotion simulation in virtual reality,\u201d","volume-title":"Proceedings of the 2024 International Conference on Advanced Visual Interfaces","author":"Della Greca","year":"2024"},{"key":"B26","doi-asserted-by":"publisher","first-page":"100288","DOI":"10.1016\/j.chbr.2023.100288","article-title":"Uncanny valley effect: a qualitative synthesis of empirical research to assess the suitability of using virtual faces in psychological research","volume":"10","author":"Di Natale","year":"2023","journal-title":"Comp. Human Behav. Reports"},{"key":"B27","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2010.11929","article-title":"An image is worth 16x16 words: transformers for image recognition at scale","author":"Dosovitskiy","year":"2020","journal-title":"arXiv"},{"key":"B28","doi-asserted-by":"publisher","first-page":"384","DOI":"10.1037\/\/0003-066X.48.4.384","article-title":"Facial expression and emotion","volume":"48","author":"Ekman","year":"1993","journal-title":"Am. Psychol"},{"key":"B29","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01821","article-title":"Faceformer: Speech-driven 3d facial animation with transformers","author":"Fan","year":"2021","journal-title":"arXiv"},{"key":"B30","first-page":"18749","article-title":"\u201cFaceformer: Speech-driven 3d facial animation with transformers,\u201d","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022","author":"Fan","year":"2022"},{"key":"B31","doi-asserted-by":"crossref","DOI":"10.1109\/3DV53792.2021.00088","article-title":"\u201cCollaborative regression of expressive bodies using moderation,\u201d","volume-title":"International Conference on 3D Vision (3DV)","author":"Feng","year":""},{"key":"B32","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3450626.3459936","article-title":"Learning an animatable detailed 3D face model from in-the-wild images","volume":"40","author":"Feng","year":"","journal-title":"ACM Trans. Graph"},{"key":"B33","doi-asserted-by":"publisher","first-page":"981400","DOI":"10.3389\/frvir.2022.981400","article-title":"Expressiveness of real-time motion captured avatars influences perceived animation realism and perceived quality of social interaction in virtual reality","volume":"3","author":"Fraser","year":"2022","journal-title":"Front. Virtual Reality"},{"key":"B34","doi-asserted-by":"publisher","first-page":"1059","DOI":"10.1162\/jocn.2006.18.7.1059","article-title":"Repetitive transcranial magnetic stimulation of broca's area affects verbal responses to gesture observation","volume":"18","author":"Gentilucci","year":"2006","journal-title":"J. Cogn. Neurosci"},{"key":"B35","doi-asserted-by":"publisher","first-page":"206","DOI":"10.1111\/cgf.14734","article-title":"Zeroeggs: Zero-shot example-based gesture generation from speech","volume":"42","author":"Ghorbani","year":"2023","journal-title":"Comp. Graphics Forum"},{"key":"B36","first-page":"3497","article-title":"\u201cLearning individual styles of conversational gesture,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Ginosar","year":"2019"},{"key":"B37","unstructured":"Virtual humans and persuasion: The effects of agency and behavioral realism\n          \n          1\n          22\n          \n            \n              Guadagno\n              R. E.\n            \n            \n              Blascovich\n              J.\n            \n            \n              Bailenson\n              J. N.\n            \n            \n              McCall\n              C. A.\n            \n          \n          Media Psychol\n          10\n          2007"},{"key":"B38","doi-asserted-by":"publisher","first-page":"52","DOI":"10.1016\/j.neulet.2004.09.011","article-title":"Communicating hands: Erps elicited by meaningful symbolic hand postures","volume":"372","author":"Gunter","year":"2004","journal-title":"Neurosci. Lett"},{"key":"B39","doi-asserted-by":"crossref","DOI":"10.1145\/3528233.3530750","article-title":"\u201cA motion matching-based framework for controllable gesture synthesis from speech,\u201d","volume-title":"International Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)","author":"Habibie","year":"2022"},{"key":"B40","doi-asserted-by":"crossref","first-page":"929","DOI":"10.1145\/1240624.1240764","article-title":"\u201cExpressing emotion in text-based communication,\u201d","volume-title":"Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI '07","author":"Hancock","year":"2007"},{"key":"B41","doi-asserted-by":"publisher","first-page":"3073","DOI":"10.1007\/s00429-018-1674-5","article-title":"Spatial-temporal dynamics of gesture-speech integration: a simultaneous eeg-fmri study","volume":"223","author":"He","year":"2018","journal-title":"Brain Struct. Funct"},{"key":"B42","unstructured":"Video for Everyone\n          \n          2025"},{"key":"B43","doi-asserted-by":"publisher","first-page":"110663","DOI":"10.1016\/j.isci.2024.110663","article-title":"Impact of social context on human facial and gestural emotion expressions","volume":"27","author":"Heesen","year":"2024","journal-title":"iScience"},{"key":"B44","doi-asserted-by":"publisher","first-page":"163","DOI":"10.1162\/pres_a_00324","article-title":"Effect of behavioral realism on social interactions inside collaborative virtual environments","volume":"27","author":"Herrera","year":"2018","journal-title":"Presence: Virtual Augment. Reality"},{"key":"B45","doi-asserted-by":"publisher","first-page":"100447","DOI":"10.1016\/j.ssaho.2023.100447","article-title":"Social emotional interaction in collaborative learning: why it matters and how can we measure it?","volume":"7","author":"Huang","year":"2023","journal-title":"Social Sci Humanit Open"},{"key":"B46","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1007\/978-3-030-85613-7_11","article-title":"\u201cGesture interaction in virtual reality: A low-cost machine learning system and a qualitative assessment of effectiveness of selected gestures vs. gaze and controller interaction,\u201d","volume-title":"Human-Computer Interaction INTERACT 2021: 18th IFIP TC 13 International Conference","author":"Huesser","year":"2021"},{"key":"B47","unstructured":"The Leading AI Engine for Games\n          \n          2025"},{"key":"B48","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3072959.3073658","article-title":"Audio-driven facial animation by joint end-to-end learning of pose and emotion","volume":"36","author":"Karras","year":"2017","journal-title":"ACM Trans. Graph"},{"key":"B49","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2106.06103","article-title":"Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech","author":"Kim","year":"2021","journal-title":"arXiv"},{"key":"B50","doi-asserted-by":"crossref","DOI":"10.1145\/2822013.2822035","article-title":"\u201cAnimation realism affects perceived character appeal of a self-virtual face,\u201d","volume-title":"Proceedings of the 8th ACM SIGGRAPH Conference on Motion in Games","author":"Kokkinara","year":"2015"},{"key":"B51","doi-asserted-by":"crossref","first-page":"205","DOI":"10.1007\/11821830_17","article-title":"\u201cTowards a common framework for multimodal generation: the behavior markup language,\u201d","author":"Kopp","year":"2006","journal-title":"Intelligent Virtual Agents"},{"key":"B52","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1038\/s41598-020-76672-4","article-title":"Facial expressions contribute more than body movements to conversational outcomes in avatar-mediated virtual environments","volume":"10","author":"Kruzic","year":"2020","journal-title":"Sci. Rep"},{"key":"B53","first-page":"792","article-title":"\u201cThe genea challenge 2023: A large-scale evaluation of gesture generation models in monadic and dyadic settings,\u201d","volume-title":"Proceedings of the 25th International Conference on Multimodal Interaction, ICMI '23","author":"Kucherenko","year":"2023"},{"key":"B54","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.113","article-title":"\u201cTemporal convolutional networks for action segmentation and detection,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Lea","year":"2017"},{"key":"B55","volume-title":"Audio2gestures: Generating Diverse Gestures from Audio","author":"Li","year":"2023"},{"key":"B56","author":"Li","year":"2021","journal-title":"Ai Choreographer: Music Conditioned 3D Dance Generation With AIST"},{"key":"B57","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3130800.3130813","article-title":"Learning a model of facial shape and expression from 4D scans","volume":"194","author":"Li","year":"2017","journal-title":"ACM Trans. Graph"},{"key":"B58","author":"Liu","year":"2024","journal-title":"Emotion Detection Through Body Gesture and Face"},{"key":"B59","first-page":"3764","article-title":"\u201cDisco: Disentangled implicit content and rhythm learning for diverse co-speech gestures synthesis,\u201d","author":"Liu","year":"","journal-title":"Proceedings of the 30th ACM International Conference on Multimedia"},{"key":"B60","first-page":"1144","article-title":"\u201cEmage: towards unified holistic co-speech gesture generation via expressive masked audio gesture modeling,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Liu","year":"2024"},{"key":"B61","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-031-20071-7_36","article-title":"\u201cBeat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis,\u201d","volume-title":"European Conference on Computer Vision","author":"Liu","year":""},{"key":"B62","unstructured":"Lombard\n              M.\n            \n            \n              Ditton\n              T. B.\n            \n            \n              Weinstein\n              L.\n            \n          \n          Measuring Presence: The Temple Presence Inventory\n          \n          2009"},{"key":"B63","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TAFFC.2025.3573878","article-title":"A review of human emotion synthesis based on generative technology","volume":"2025","author":"Ma","year":"2025","journal-title":"IEEE Trans. Affect. Comput"},{"key":"B64","volume-title":"Evaluating the Quality of a Synthesized Motion with the Frchet Motion Distance","author":"Maiorca","year":"2022"},{"key":"B65","doi-asserted-by":"crossref","DOI":"10.1145\/3613905.3650773","article-title":"\u201cFrom 2d-screens to vr: Exploring the effect of immersion on the plausibility of virtual humans,\u201d","volume-title":"Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA '24","author":"Mal","year":"2024"},{"key":"B66","first-page":"31","author":"Marinetti","year":"2011","journal-title":"Emotions in Social Interactions: Unfolding Emotional Experience"},{"key":"B67","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1145\/2485895.2485900","article-title":"\u201cVirtual character performance from speech,\u201d","volume-title":"Proceedings of the 12th ACM SIGGRAPH\/Eurographics Symposium on Computer Animation, SCA '13","author":"Marsella","year":"2013"},{"key":"B68","article-title":"\u201cHand and mind1,\u201d","author":"McNeill","year":"1992","journal-title":"Advances in Visual Semiotics"},{"key":"B69","unstructured":"Microsoft Mesh\n          \n          2025"},{"key":"B70","first-page":"1388","article-title":"\u201cConvofusion: Multi-modal conversational diffusion for co-speech gesture synthesis,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Mughal","year":"2024"},{"key":"B71","unstructured":"NVIDIA ACE\n          \n          2025"},{"key":"B72","doi-asserted-by":"publisher","first-page":"srep13899","DOI":"10.1038\/srep13899","article-title":"Conversations between self and self as sigmund freuda virtual body ownership paradigm for self counselling","volume":"5","author":"Osimo","year":"2015","journal-title":"Sci. Rep"},{"key":"B73","doi-asserted-by":"publisher","first-page":"20130296","DOI":"10.1098\/rstb.2013.0296","article-title":"Hearing and seeing meaning in speech and gesture: Insights from brain and behavior","volume":"369","author":"\u00d6zy\u00fcrek","year":"2014","journal-title":"Philosoph. Trans. Royal Soc. B: Biol. Sci"},{"key":"B74","doi-asserted-by":"publisher","first-page":"395","DOI":"10.1111\/bjop.12290","article-title":"Why and how to use virtual reality to study human social interaction: the challenges of exploring a new research landscape","volume":"109","author":"Pan","year":"2018","journal-title":"Br. J. Psychol"},{"key":"B75","doi-asserted-by":"publisher","first-page":"162","DOI":"10.1016\/j.cag.2023.01.001","article-title":"Avoiding virtual humans in a constrained environment: Exploration of novel behavioral measures","volume":"110","author":"Patotskaya","year":"2023","journal-title":"Comput. Graph"},{"key":"B76","doi-asserted-by":"publisher","first-page":"1388","DOI":"10.1177\/17456916221148142","article-title":"Four misconceptions about nonverbal communication","volume":"18","author":"Patterson","year":"2023","journal-title":"Persp. Psychol. Sci"},{"key":"B77","first-page":"10975","article-title":"\u201cExpressive body capture: 3D hands, face, and body from a single image,\u201d","volume-title":"Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Pavlakos","year":"2019"},{"key":"B78","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV48922.2021.01080","article-title":"\u201cAction-conditioned 3D human motion synthesis with transformer VAE,\u201d","volume-title":"International Conference on Computer Vision (ICCV)","author":"Petrovich","year":"2021"},{"key":"B79","doi-asserted-by":"crossref","DOI":"10.1109\/CVPRW.2017.287","article-title":"\u201cSpeech-driven 3D facial animation with implicit emotional awareness: a deep learning approach,\u201d","volume-title":"2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2017","author":"Pham","year":""},{"key":"B80","doi-asserted-by":"publisher","DOI":"10.1145\/3242969.3243017","article-title":"End-to-end learning for 3d facial animation from raw waveforms of speech","author":"Pham","year":"","journal-title":"arXiv"},{"key":"B81","unstructured":"AI Voice Generator & Text to Speech AI Voice Platform\n          \n          2025"},{"key":"B82","doi-asserted-by":"publisher","first-page":"344","DOI":"10.1511\/2001.28.344","article-title":"The nature of emotions: Human emotions have deep evolutionary roots, a fact that may explain their complexity and provide tools for clinical practice","volume":"89","author":"Plutchik","year":"2001","journal-title":"Am. Scient"},{"key":"B83","doi-asserted-by":"crossref","DOI":"10.1007\/1-4020-3051-7_1","volume-title":"Greta. A Believable Embodied Conversational Agent","author":"Poggi","year":"2005"},{"key":"B84","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR52729.2023.00448","volume-title":"Diverse 3D Hand Gesture Prediction from Body Dynamics by Bilateral Hand Disentanglement","author":"Qi","year":"2023"},{"key":"B85","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2007.08501","article-title":"Accelerating 3D deep learning with pytorch3d","author":"Ravi","year":"2020","journal-title":"arXiv"},{"key":"B86","volume-title":"The Media Equation: How People Treat Computers, Television, and New Media Like Real People and PLA","author":"Reeves","year":"1996"},{"key":"B87","unstructured":"The AI Companion Who Cares\n          \n          2025"},{"key":"B88","doi-asserted-by":"publisher","first-page":"1173","DOI":"10.1109\/ICCV48922.2021.00121","article-title":"\u201cMeshtalk: 3D face animation from speech using cross-modality disentanglement,\u201d","author":"Richard","year":"2021"},{"key":"B89","doi-asserted-by":"publisher","first-page":"188","DOI":"10.1016\/S0166-2236(98)01260-0","article-title":"Language within our grasp","volume":"21","author":"Rizzolatti","year":"1998","journal-title":"Trends Neurosci"},{"key":"B90","doi-asserted-by":"publisher","first-page":"750729","DOI":"10.3389\/frvir.2021.750729","article-title":"Realistic motion avatars are the future for social interaction in virtual reality","volume":"2","author":"Rogers","year":"2021","journal-title":"Front. Virtual Reality"},{"key":"B91","first-page":"10684","article-title":"\u201cHigh-resolution image synthesis with latent diffusion models,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Rombach","year":"2022"},{"key":"B92","doi-asserted-by":"publisher","first-page":"1641","DOI":"10.1016\/j.chb.2010.06.012","article-title":"\u201cIt doesn't matter what you are!\u201d explaining social effects of agents and avatars","volume":"26","author":"Rosenthal-von der P\u00fctten","year":"2010","journal-title":"Comput. Human Behav"},{"key":"B93","first-page":"215","article-title":"\u201cBeyond replication: augmenting social behaviors in multi-user virtual realities,\u201d","volume-title":"2018 IEEE Conference on Virtual Reality and 3D User Interfaces (VR)","author":"Roth","year":""},{"key":"B94","first-page":"103","article-title":"\u201cEffects of hybrid and synthetic social gaze in avatar-mediated interactions,\u201d","volume-title":"2018 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct)","author":"Roth","year":""},{"key":"B95","doi-asserted-by":"publisher","first-page":"1161","DOI":"10.1037\/h0077714","article-title":"A circumplex model of affect","volume":"39","author":"Russell","year":"1980","journal-title":"J. Pers. Soc. Psychol"},{"key":"B96","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1037\/h0054570","article-title":"Three dimensions of emotion","volume":"61","author":"Schlosberg","year":"1954","journal-title":"Psychol. Rev"},{"key":"B97","doi-asserted-by":"publisher","first-page":"387","DOI":"10.22363\/2313-2272-2022-22-2-387-403","article-title":"Non-verbal signs of personality: Communicative meanings of facial expressions","volume":"22","author":"Sharkov","year":"2022","journal-title":"RUDN J. Sociol"},{"key":"B98","doi-asserted-by":"crossref","DOI":"10.1145\/3173574.3173863","article-title":"\u201cCommunication behavior in embodied virtual reality,\u201d","volume-title":"Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI '18","author":"Smith","year":"2018"},{"key":"B99","doi-asserted-by":"publisher","first-page":"980","DOI":"10.1038\/s41598-024-60980-0","article-title":"The nonverbal expression of guilt in healthy adults","volume":"14","author":"Stewart","year":"2024","journal-title":"Sci. Rep"},{"key":"B100","unstructured":"Turn Text to Video, in Minutes\n          \n          2025"},{"key":"B101","volume-title":"Integrating Facial, Gesture, and Posture Emotion Expression for a 3D Virtual Agent","author":"Tan","year":"2009"},{"key":"B102","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3072959.3073699","article-title":"A deep learning approach for generalized speech animation","volume":"93","author":"Taylor","year":"2017","journal-title":"ACM Trans. Graph"},{"key":"B103","article-title":"\u201cSmartbody: behavior realization for embodied conversational agents,\u201d","volume-title":"Adaptive Agents and Multi-Agent Systems","author":"Thi\u00e9baux","year":"2008"},{"key":"B104","first-page":"11","author":"Thomas","year":"2022"},{"key":"B105","unstructured":"Unreal Engine 5.4 Documentation\n          \n          2025"},{"key":"B106","unstructured":"Blender Foundation Documentation\n          \n          2025"},{"key":"B107","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1016\/j.cobeha.2015.01.001","article-title":"The nonverbal communication of emotions","volume":"3","author":"Tracy","year":"2015","journal-title":"Curr. Opini. Behav. Sci. Soc. Behav"},{"key":"B108","unstructured":"Documentation\n          \n          2025"},{"key":"B109","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3478513.3480570","article-title":"Transflower","volume":"40","author":"Valle-P\u00e9rez","year":"2021","journal-title":"ACM Trans. Graph"},{"key":"B110","unstructured":"Neural discrete representation learning\n          \n          6309\n          6318\n          \n            \n              Van Den Oord\n              A.\n            \n            \n              Vinyals\n              O.\n            \n            \n              Kavukcuoglu\n              K.\n            \n          \n          Adv. Neural Inf. Process. Syst\n          30\n          2017"},{"key":"B111","doi-asserted-by":"crossref","DOI":"10.1145\/3719160.3737643","article-title":"\u201cTowards enhancing industrial training through conversational AI,\u201d","author":"Vasiliu","year":"2025"},{"key":"B112","article-title":"\u201cAttention is all you need,\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani","year":"2017"},{"key":"B113","doi-asserted-by":"crossref","first-page":"219","DOI":"10.1037\/0033-2909.98.2.219","article-title":"Toward a consensual structure of mood","volume":"982","author":"Watson","year":"1985","journal-title":"Psychol. Bullet"},{"key":"B114","first-page":"143","author":"Wobbrock","year":"2011","journal-title":"The Aligned Rank Transform for Nonparametric Factorial Analyses Using Only Anova Procedures"},{"key":"B115","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01229","article-title":"Codetalker: Speech-driven 3d facial animation with discrete motion prior","author":"Xing","year":"2023","journal-title":"arXiv"},{"key":"B116","article-title":"\u201cDiffusestylegesture: Stylized audio-driven co-speech gesture generation with diffusion models,\u201d","author":"Yang","year":"","journal-title":"Proceedings of the 32nd International Joint Conference on Artificial Intelligence, IJCAI 2023"},{"key":"B117","first-page":"2321","article-title":"\u201cQpgesture: Quantization-based and phase-guided motion matching for natural speech-driven gesture generation,\u201d","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR","author":"Yang","year":""},{"key":"B118","first-page":"469","article-title":"\u201cGenerating holistic 3D human motion from speech,\u201d","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Yi","year":"2023"},{"key":"B119","author":"Yin","year":"2023","journal-title":"Emog: Synthesizing Emotive Co-Speech 3D Gesture With Diffusion Model"},{"key":"B120","doi-asserted-by":"publisher","first-page":"2245","DOI":"10.1109\/TVCG.2022.3150507","article-title":"The one-man-crowd: Single user generation of crowd motions using virtual reality","volume":"28","author":"Yin","year":"2022","journal-title":"IEEE Trans. Vis. Comput. Graph"},{"key":"B121","doi-asserted-by":"publisher","first-page":"2785","DOI":"10.1109\/TVCG.2024.3372038","article-title":"With or without you: Effect of contextual and responsive crowds on vr-based crowd motion capture","volume":"30","author":"Yin","year":"2024","journal-title":"IEEE Trans. Vis. Comput. Graph"},{"key":"B122","doi-asserted-by":"publisher","first-page":"6","DOI":"10.1145\/3414685.3417838","article-title":"Speech gesture generation from the trimodal context of text, audio, and speaker identity","volume":"39","author":"Yoon","year":"2020","journal-title":"ACM Trans. Graph"},{"key":"B123","doi-asserted-by":"crossref","first-page":"940","DOI":"10.1109\/ISMAR59233.2023.00110","article-title":"\u201cSupporting co-presence in populated virtual environments by actor takeover of animated characters,\u201d","volume-title":"2023 IEEE International Symposium on Mixed and Augmented Reality (ISMAR)","author":"Zhang","year":"2023"},{"key":"B124","doi-asserted-by":"publisher","first-page":"1891","DOI":"10.1523\/JNEUROSCI.1748-17.2017","article-title":"Transcranial magnetic stimulation over left inferior frontal and posterior temporal cortex disrupts gesture-speech integration","volume":"38","author":"Zhao","year":"2018","journal-title":"J. Neurosci"},{"key":"B125","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1145\/3197517.3201292","article-title":"Visemenet: Audio-driven animator-centric speech animation","volume":"37","author":"Zhou","year":"2018","journal-title":"ACM Trans. Graph"},{"key":"B126","doi-asserted-by":"publisher","first-page":"1","DOI":"10.3389\/fpsyg.2023.1199537","article-title":"Incongruent gestures slow the processing of facial expressions in university students with social anxiety","volume":"14","author":"Zhu","year":"2023","journal-title":"Front. Psychol"},{"key":"B127","doi-asserted-by":"publisher","first-page":"1681","DOI":"10.1109\/TVCG.2018.2794638","article-title":"The effect of realistic appearance of virtual characters in immersive environments - does the character's personality play a role?","volume":"24","author":"Zibrek","year":"2018","journal-title":"IEEE Trans. Visualizat. Comp. Graph"}],"container-title":["Frontiers in Computer Science"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fcomp.2025.1598099\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,31]],"date-time":"2025-07-31T05:39:27Z","timestamp":1753940367000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fcomp.2025.1598099\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,31]]},"references-count":127,"alternative-id":["10.3389\/fcomp.2025.1598099"],"URL":"https:\/\/doi.org\/10.3389\/fcomp.2025.1598099","relation":{},"ISSN":["2624-9898"],"issn-type":[{"value":"2624-9898","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,31]]},"article-number":"1598099"}}