{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T08:18:22Z","timestamp":1783066702076,"version":"3.54.6"},"reference-count":76,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T00:00:00Z","timestamp":1783036800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2026,7,3]]},"abstract":"<jats:p>Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed required for interactive settings. Conversely, existing online methods enable real-time synthesis but often sacrifice controllability or struggle with complex text semantics and long-horizon goals due to limited context windows. In this work, we introduce ARDY, a streaming generation framework that bridges this gap by enabling high-fidelity motion generation controllable via online text prompts and flexible kinematic constraints. ARDY employs a hybrid representation that combines explicit root features with a latent body embedding, balancing precise trajectory control with efficient generative learning. We propose a two-stage autoregressive transformer denoiser that features variable history context and supports conditioning on flexible, long-horizon kinematic constraints. By training on a large-scale motion capture dataset and being directly conditioned on text labels and kinematic constraints sampled from ground truth poses, ARDY natively learns controllable generation that supports online prompting and flexible long-horizon goals. Extensive evaluations on the HumanML3D benchmark and the large-scale, high-fidelity Bones Rigplay dataset demonstrate ARDY's high motion quality and constraint adherence, validating the efficacy of our key architectural decisions. Finally, we demonstrate the method's practical versatility through an interactive demo featuring dynamic text control, diverse keyframe pose constraints, path following, and interactive locomotion control via mouse and keyboard. Supplementary video results, code, and model releases can be found at https:\/\/research.nvidia.com\/labs\/sil\/projects\/ardy\/.<\/jats:p>","DOI":"10.1145\/3811284","type":"journal-article","created":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:05:51Z","timestamp":1783062351000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation"],"prefix":"10.1145","volume":"45","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7278-3329","authenticated-orcid":false,"given":"Kaifeng","family":"Zhao","sequence":"first","affiliation":[{"name":"ETH Z\u00fcrich, Zurich, Switzerland"},{"name":"NVIDIA, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0859-1170","authenticated-orcid":false,"given":"Mathis","family":"Petrovich","sequence":"additional","affiliation":[{"name":"NVIDIA, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-0293-337X","authenticated-orcid":false,"given":"Haotian","family":"Zhang","sequence":"additional","affiliation":[{"name":"NVIDIA, Santa Clara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2006-0660","authenticated-orcid":false,"given":"Tingwu","family":"Wang","sequence":"additional","affiliation":[{"name":"NVIDIA, Santa Clara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1015-4770","authenticated-orcid":false,"given":"Siyu","family":"Tang","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2256-5507","authenticated-orcid":false,"given":"Davis","family":"Rempe","sequence":"additional","affiliation":[{"name":"NVIDIA, Santa Clara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,3]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"AI@Meta. 2024. Llama 3 Model Card. (2024). https:\/\/github.com\/meta-llama\/llama3\/blob\/main\/MODEL_CARD.md"},{"key":"e_1_2_2_2_1","doi-asserted-by":"crossref","unstructured":"German Barquero Sergio Escalera and Cristina Palmero. 2024. Seamless Human Motion Composition with Blended Positional Encodings. (2024).","DOI":"10.1109\/CVPR52733.2024.00051"},{"key":"e_1_2_2_3_1","volume-title":"First Conference on Language Modeling. https:\/\/openreview.net\/forum?id=IW1PR7vEBf","author":"BehnamGhader Parishad","year":"2024","unstructured":"Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders. In First Conference on Language Modeling. https:\/\/openreview.net\/forum?id=IW1PR7vEBf"},{"key":"e_1_2_2_4_1","volume-title":"AI Datasets for Machine Learning and Motion Capture. https:\/\/bones.studio\/datasets. Accessed","author":"Studio Bones","year":"2026","unstructured":"Bones Studio. 2026. AI Datasets for Machine Learning and Motion Capture. https:\/\/bones.studio\/datasets. Accessed: 2026."},{"key":"e_1_2_2_5_1","unstructured":"Zhi Cen Huaijin Pi Sida Peng Qing Shuai Yujun Shen Hujun Bao Xiaowei Zhou and Ruizhen Hu. 2025. Ready-to-React: Online Reaction Policy for Two-Character Interaction Generation. In ICLR."},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3641519.3657440"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01726"},{"key":"e_1_2_2_8_1","volume-title":"Motionlcm: Real-time controllable motion generation via latent consistency model. In ECCV. 390\u2013408.","author":"Dai Wenxun","year":"2025","unstructured":"Wenxun Dai, Ling-Hao Chen, Jingbo Wang, Jinpeng Liu, Bo Dai, and Yansong Tang. 2025. Motionlcm: Real-time controllable motion generation via latent consistency model. In ECCV. 390\u2013408."},{"key":"e_1_2_2_9_1","volume-title":"International Conference on Machine Learning","author":"Everett Katie","year":"2024","unstructured":"Katie Everett, Lechao Xiao, Mitchell Wortsman, Alexander A Alemi, Roman Novak, Peter J Liu, Izzeddin Gur, Jascha Sohl-Dickstein, Leslie Pack Kaelbling, Jaehoon Lee, et al. 2024. Scaling exponents across parameterizations and optimizers. International Conference on Machine Learning (2024)."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01239"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.494"},{"key":"e_1_2_2_12_1","volume-title":"Mean Flows for One-step Generative Modeling. In The Thirty-ninth Annual Conference on Neural Information Processing Systems.","author":"Geng Zhengyang","year":"2025","unstructured":"Zhengyang Geng, Mingyang Deng, Xingjian Bai, J Zico Kolter, and Kaiming He. 2025. Mean Flows for One-step Generative Modeling. In The Thirty-ninth Annual Conference on Neural Information Processing Systems."},{"key":"e_1_2_2_13_1","volume-title":"Fast R-CNN. In International Conference on Computer Vision (ICCV).","author":"Girshick Ross","year":"2015","unstructured":"Ross Girshick. 2015. Fast R-CNN. In International Conference on Computer Vision (ICCV)."},{"key":"e_1_2_2_14_1","volume-title":"Control Operators for Interactive Character Animation. ACM Transactions on Graphics (TOG)","author":"Gou Ruiyu","year":"2025","unstructured":"Ruiyu Gou, Michiel van de Panne, and Daniel Holden. 2025. Control Operators for Interactive Character Animation. ACM Transactions on Graphics (TOG) (2025)."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00186"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00509"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413635"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01118"},{"key":"e_1_2_2_19_1","volume-title":"Asap: Aligning simulation and real-world physics for learning agile humanoid whole-body skills. arXiv preprint arXiv:2502.01143","author":"He Tairan","year":"2025","unstructured":"Tairan He, Jiawei Gao, Wenli Xiao, Yuanhang Zhang, Zi Wang, Jiashun Wang, Zhengyi Luo, Guanqi He, Nikhil Sobanbab, Chaoyi Pan, et al. 2025. Asap: Aligning simulation and real-world physics for learning agile humanoid whole-body skills. arXiv preprint arXiv:2502.01143 (2025)."},{"key":"e_1_2_2_20_1","volume-title":"Denoising diffusion probabilistic models. Advances in neural information processing systems 33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840\u20136851."},{"key":"e_1_2_2_21_1","volume-title":"Classifier-Free Diffusion Guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications.","author":"Ho Jonathan","year":"2021","unstructured":"Jonathan Ho and Tim Salimans. 2021. Classifier-Free Diffusion Guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications."},{"key":"e_1_2_2_22_1","volume-title":"Learned motion matching. ACM Transactions on Graphics (ToG)","author":"Holden Daniel","year":"2020","unstructured":"Daniel Holden, Oussama Kanoun, Maksym Perepichka, and Tiberiu Popa. 2020. Learned motion matching. ACM Transactions on Graphics (ToG) (2020)."},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073663"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3731206"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.00948"},{"key":"e_1_2_2_26_1","volume-title":"Motiongpt: Human motion as a foreign language. Advances in Neural Information Processing Systems 36","author":"Jiang Biao","year":"2024","unstructured":"Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen. 2024a. Motiongpt: Human motion as a foreign language. Advances in Neural Information Processing Systems 36 (2024)."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3680528.3687595"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00133"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00205"},{"key":"e_1_2_2_30_1","volume-title":"Auto-Encoding Variational Bayes. In International Conference on Learning Representations.","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Max Welling. 2014. Auto-Encoding Variational Bayes. In International Conference on Learning Representations."},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV66043.2025.00027"},{"key":"e_1_2_2_32_1","volume-title":"Beyondmimic: From motion tracking to versatile humanoid control via guided diffusion. arXiv preprint arXiv:2508.08241","author":"Liao Qiayuan","year":"2025","unstructured":"Qiayuan Liao, Takara E Truong, Xiaoyu Huang, Guy Tevet, Koushil Sreenath, and C Karen Liu. 2025. Beyondmimic: From motion tracking to versatile humanoid control via guided diffusion. arXiv preprint arXiv:2508.08241 (2025)."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392422"},{"key":"e_1_2_2_34_1","volume-title":"Black","author":"Loper Matthew","year":"2015","unstructured":"Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. 2015. SMPL: A Skinned Multi-Person Linear Model. ACM Trans. Graphics (Proc. SIGGRAPH Asia) (2015)."},{"key":"e_1_2_2_35_1","volume-title":"Stabilizing and Scaling Continuous-time Consistency Models. In The Thirteenth International Conference on Learning Representations.","author":"Lu Cheng","year":"2025","unstructured":"Cheng Lu and Yang Song. 2025. Simplifying, Stabilizing and Scaling Continuous-time Consistency Models. In The Thirteenth International Conference on Learning Representations."},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02595"},{"key":"e_1_2_2_37_1","volume-title":"Universal humanoid motion representations for physics-based control. arXiv preprint arXiv:2310.04582","author":"Luo Zhengyi","year":"2023","unstructured":"Zhengyi Luo, Jinkun Cao, Josh Merel, Alexander Winkler, Jing Huang, Kris Kitani, and Weipeng Xu. 2023. Universal humanoid motion representations for physics-based control. arXiv preprint arXiv:2310.04582 (2023)."},{"key":"e_1_2_2_38_1","volume-title":"SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control. arXiv preprint arXiv:2511.07820","author":"Luo Zhengyi","year":"2025","unstructured":"Zhengyi Luo, Ye Yuan, Tingwu Wang, Chenran Li, Sirui Chen, Fernando Casta\u00f1eda, Zi-Ang Cao, Jiefeng Li, David Minor, Qingwei Ben, Xingye Da, Runyu Ding, Cyrus Hogg, Lina Song, Edy Lim, Eugene Jeong, Tairan He, Haoru Xue, Wenli Xiao, Zi Wang, Simon Yuen, Jan Kautz, Yan Chang, Umar Iqbal, Linxi Fan, and Yuke Zhu. 2025. SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control. arXiv preprint arXiv:2511.07820 (2025)."},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02594"},{"key":"e_1_2_2_40_1","volume-title":"Finite scalar quantization: Vq-vae made simple. arXiv preprint arXiv:2309.15505","author":"Mentzer Fabian","year":"2023","unstructured":"Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen. 2023. Finite scalar quantization: Vq-vae made simple. arXiv preprint arXiv:2309.15505 (2023)."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528223.3530110"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20047-2_28"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00870"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW63382.2024.00197"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.00928"},{"key":"e_1_2_2_46_1","volume-title":"European Conference on Computer Vision. Springer, 172\u2013190","author":"Pinyoanuntapong Ekkasit","year":"2024","unstructured":"Ekkasit Pinyoanuntapong, Muhammad Usama Saleem, Pu Wang, Minwoo Lee, Srijan Das, and Chen Chen. 2024a. Bamm: Bidirectional autoregressive motion model. In European Conference on Computer Vision. Springer, 172\u2013190."},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00153"},{"key":"e_1_2_2_48_1","volume-title":"The kit motion-language dataset. Big data 4, 4","author":"Plappert Matthias","year":"2016","unstructured":"Matthias Plappert, Christian Mandery, and Tamim Asfour. 2016. The kit motion-language dataset. Big data 4, 4 (2016), 236\u2013252."},{"key":"e_1_2_2_49_1","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans Ilya Sutskever et al. 2018. Improving language understanding by generative pre-training. (2018)."},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01129"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01322"},{"key":"e_1_2_2_52_1","unstructured":"Davis Rempe Mathis Petrovich Ye Yuan Haotian Zhang Xue Bin Peng Yifeng Jiang Tingwu Wang Umar Iqbal David Minor Michael de Ruyter Jiefeng Li Chen Tessler Edy Lim Eugene Jeong Sam Wu Ehsan Hassani Michael Huang Jin-Bey Yu Chaeyeon Chung Lina Song Olivier Dionne Jan Kautz Simon Yuen and Sanja Fidler. 2026. Kimodo: Scaling Controllable Human Motion Generation. arXiv:2603.15546 (2026)."},{"key":"e_1_2_2_53_1","volume-title":"InsActor: Instruction-driven Physics-based Characters. NeurIPS","author":"Ren Jiawei","year":"2023","unstructured":"Jiawei Ren, Mingyuan Zhang, Cunjun Yu, Xiao Ma, Liang Pan, and Ziwei Liu. 2023. InsActor: Instruction-driven Physics-based Characters. NeurIPS (2023)."},{"key":"e_1_2_2_54_1","volume-title":"Xue Bin Peng, and Michiel van de Panne","author":"Setareh Cohan","year":"2024","unstructured":"Cohan Setareh, Guy Tevet, Daniele Reda, Xue Bin Peng, and Michiel van de Panne. 2024. Flexible Motion In-betweening with Diffusion Models. (2024)."},{"key":"e_1_2_2_55_1","volume-title":"Interactive Character Control with Auto-Regressive Motion Diffusion Models. ACM Trans. Graph. 43 (jul","author":"Shi Yi","year":"2024","unstructured":"Yi Shi, Jingbo Wang, Xuekun Jiang, Bingkun Lin, Bo Dai, and Xue Bin Peng. 2024. Interactive Character Control with Auto-Regressive Motion Diffusion Models. ACM Trans. Graph. 43 (jul 2024)."},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528223.3530178"},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3355089.3356505"},{"key":"e_1_2_2_58_1","volume-title":"Modeling human motion using binary latent variables. Advances in neural information processing systems 19","author":"Taylor Graham W","year":"2006","unstructured":"Graham W Taylor, Geoffrey E Hinton, and Sam Roweis. 2006. Modeling human motion using binary latent variables. Advances in neural information processing systems 19 (2006)."},{"key":"e_1_2_2_59_1","volume-title":"Masked-Mimic: Unified Physics-Based Character Control Through Masked Motion Inpainting. ACM Transactions on Graphics (TOG)","author":"Tessler Chen","year":"2024","unstructured":"Chen Tessler, Yunrong Guo, Ofir Nabati, Gal Chechik, and Xue Bin Peng. 2024. Masked-Mimic: Unified Physics-Based Character Control Through Masked Motion Inpainting. ACM Transactions on Graphics (TOG) (2024)."},{"key":"e_1_2_2_60_1","volume-title":"The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=pZISppZSTv","author":"Tevet Guy","unstructured":"Guy Tevet, Sigal Raab, Setareh Cohan, Daniele Reda, Zhengyi Luo, Xue Bin Peng, Amit Haim Bermano, and Michiel van de Panne. 2025. CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=pZISppZSTv"},{"key":"e_1_2_2_61_1","volume-title":"Human Motion Diffusion Model. In The Eleventh International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=SJ1kSyO2jwu","author":"Tevet Guy","year":"2023","unstructured":"Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. 2023. Human Motion Diffusion Model. In The Eleventh International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=SJ1kSyO2jwu"},{"key":"e_1_2_2_62_1","volume-title":"European Conference on Computer Vision. Springer, 37\u201354","author":"Wan Weilin","year":"2024","unstructured":"Weilin Wan, Zhiyang Dou, Taku Komura, Wenping Wang, Dinesh Jayaraman, and Lingjie Liu. 2024. Tlcontrol: Trajectory and language control for human motion synthesis. In European Conference on Computer Vision. Springer, 37\u201354."},{"key":"e_1_2_2_63_1","volume-title":"UniPhys: Unified Planner and Controller with Diffusion for Flexible Physics-Based Character Control. arXiv preprint arXiv:2504.12540","author":"Wu Yan","year":"2025","unstructured":"Yan Wu, Korrawe Karunratanakul, Zhengyi Luo, and Siyu Tang. 2025. UniPhys: Unified Planner and Controller with Diffusion for Flexible Physics-Based Character Control. arXiv preprint arXiv:2504.12540 (2025)."},{"key":"e_1_2_2_64_1","volume-title":"MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent Space. arXiv preprint arXiv:2503.15451","author":"Xiao Lixing","year":"2025","unstructured":"Lixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan, Liang Pan, Yueer Zhou, Ziyong Feng, Xiaowei Zhou, Sida Peng, and Jingbo Wang. 2025. MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent Space. arXiv preprint arXiv:2503.15451 (2025)."},{"key":"e_1_2_2_65_1","volume-title":"The Twelfth International Conference on Learning Representations.","author":"Xie Yiming","year":"2024","unstructured":"Yiming Xie, Varun Jampani, Lei Zhong, Deqing Sun, and Huaizu Jiang. 2024. OmniControl: Control Any Joint at Any Time for Human Motion Generation. In The Twelfth International Conference on Learning Representations."},{"key":"e_1_2_2_66_1","volume-title":"Justin Kerr, Gina Wu, Rebecca Feng, Anthony Zhang, Jonas Kulhanek, Hongsuk Choi, Yi Ma, Matthew Tancik, and Angjoo Kanazawa.","author":"Yi Brent","year":"2025","unstructured":"Brent Yi, Chung Min Kim, Justin Kerr, Gina Wu, Rebecca Feng, Anthony Zhang, Jonas Kulhanek, Hongsuk Choi, Yi Ma, Matthew Tancik, and Angjoo Kanazawa. 2025. Viser: Imperative, web-based 3d visualization in python. arXiv preprint arXiv:2507.22885 (2025)."},{"key":"e_1_2_2_67_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Zhang Jianrong","year":"2023","unstructured":"Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hongwei Zhao, Hongtao Lu, and Xi Shen. 2023. T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_2_68_1","volume-title":"Motiondiffuse: Text-driven human motion generation with diffusion model","author":"Zhang Mingyuan","year":"2024","unstructured":"Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. 2024a. Motiondiffuse: Text-driven human motion generation with diffusion model. IEEE transactions on pattern analysis and machine intelligence 46, 6 (2024), 4115\u20134128."},{"key":"e_1_2_2_69_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV).","author":"Zhang Yan","unstructured":"Yan Zhang, Yao Feng, Alp\u00e1r Cseke, Nitin Saini, Nathan Bajandas, Nicolas Heron, and Michael J. Black. 2025. PRIMAL: Physically Reactive and Interactive Motor Model for Avatar Learning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)."},{"key":"e_1_2_2_70_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01983"},{"key":"e_1_2_2_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/3641519.3657515"},{"key":"e_1_2_2_72_1","volume-title":"DartControl: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion Control. In The Thirteenth International Conference on Learning Representations (ICLR).","author":"Zhao Kaifeng","year":"2025","unstructured":"Kaifeng Zhao, Gen Li, and Siyu Tang. 2025a. DartControl: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion Control. In The Thirteenth International Conference on Learning Representations (ICLR)."},{"key":"e_1_2_2_73_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01354"},{"key":"e_1_2_2_74_1","volume-title":"ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning. arXiv preprint arXiv:2510.05070","author":"Zhao Siheng","year":"2025","unstructured":"Siheng Zhao, Yanjie Ze, Yue Wang, C Karen Liu, Pieter Abbeel, Guanya Shi, and Rocky Duan. 2025b. ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning. arXiv preprint arXiv:2510.05070 (2025)."},{"key":"e_1_2_2_75_1","volume-title":"European Conference on Computer Vision. Springer, 18\u201338","author":"Zhou Wenyang","year":"2024","unstructured":"Wenyang Zhou, Zhiyang Dou, Zeyu Cao, Zhouyingcheng Liao, Jingbo Wang, Wenjia Wang, Yuan Liu, Taku Komura, Wenping Wang, and Lingjie Liu. 2024. Emdm: Efficient motion diffusion model for fast and high-quality motion generation. In European Conference on Computer Vision. Springer, 18\u201338."},{"key":"e_1_2_2_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00589"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","deposited":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:40:14Z","timestamp":1783064414000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811284"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,3]]},"references-count":76,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,7,3]]}},"alternative-id":["10.1145\/3811284"],"URL":"https:\/\/doi.org\/10.1145\/3811284","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,3]]},"assertion":[{"value":"2026-01-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}