{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T08:19:47Z","timestamp":1783066787569,"version":"3.54.6"},"reference-count":100,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T00:00:00Z","timestamp":1783036800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2026,7,3]]},"abstract":"<jats:p>\n                    Despite transformative advances in generative motion synthesis, real-time interactive motion control remains dominated by traditional techniques. In this work, we identify two key challenges in bridging research and production: 1)\n                    <jats:italic toggle=\"yes\">Real-time scalability<\/jats:italic>\n                    : Industry applications demand real-time generation of a vast repertoire of motion skills, while generative methods exhibit significant degradation in quality and scalability under real-time computation constraints, and 2)\n                    <jats:italic toggle=\"yes\">Integration<\/jats:italic>\n                    : Industry applications demand fine-grained multi-modal control involving velocity commands, style selection, and precise keyframes, a need largely unmet by existing text- or tag-driven models. Moreover, a systematic motion design interface for generative models remains absent. To overcome these limitations, we introduce MotionBricks: a large-scale, real-time generative framework with a two-fold solution. First, we propose a large-scale modular latent generative backbone tailored for robust real-time motion generation, effectively modeling a dataset of over 350,000 motion clips with a single model. Second, we introduce\n                    <jats:italic toggle=\"yes\">smart primitives<\/jats:italic>\n                    that provide a unified, robust, and intuitive interface for authoring both navigation and object interaction. Notably, MotionBricks applies to new downstream tasks in a zero-shot manner, where no fine-tuning or task-specific tagging is required. Applications can be designed in a plug-and-play manner like assembling bricks without expert animation knowledge, enabling an accessible interface for applications in animation and robotics. Quantitatively, we show that MotionBricks produces state-of-the-art motion quality on open-source and proprietary datasets of various scales, while also achieving a real-time throughput of 15,000 FPS with 2ms latency. We demonstrate the flexibility and robustness of MotionBricks in a complete production-level animation demo, covering navigation and object-scene interaction across various styles with a unified model. To showcase our framework's application beyond animation, we deploy MotionBricks on the Unitree G1 humanoid robot to demonstrate its flexibility and generalization for real-time robotic control.\n                  <\/jats:p>","DOI":"10.1145\/3811334","type":"journal-article","created":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:05:51Z","timestamp":1783062351000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["MotionBricks: Scalable Real-Time Motions with Modular Latent Generative Model and Smart Primitives"],"prefix":"10.1145","volume":"45","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2006-0660","authenticated-orcid":false,"given":"Tingwu","family":"Wang","sequence":"first","affiliation":[{"name":"NVIDIA, Santa Clara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-9569-1824","authenticated-orcid":false,"given":"Olivier","family":"Dionne","sequence":"additional","affiliation":[{"name":"NVIDIA, Montreal, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-1220-8337","authenticated-orcid":false,"given":"Michael","family":"De Ruyter","sequence":"additional","affiliation":[{"name":"NVIDIA, San Diego, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-6682-8652","authenticated-orcid":false,"given":"David","family":"Minor","sequence":"additional","affiliation":[{"name":"NVIDIA, Vancouver, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2256-5507","authenticated-orcid":false,"given":"Davis","family":"Rempe","sequence":"additional","affiliation":[{"name":"NVIDIA, Santa Clara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7278-3329","authenticated-orcid":false,"given":"Kaifeng","family":"Zhao","sequence":"additional","affiliation":[{"name":"NVIDIA, Z\u00fcrich, Switzerland"},{"name":"ETH Z\u00fcrich, Z\u00fcrich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0859-1170","authenticated-orcid":false,"given":"Mathis","family":"Petrovich","sequence":"additional","affiliation":[{"name":"NVIDIA, Z\u00fcrich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5316-6002","authenticated-orcid":false,"given":"Ye","family":"Yuan","sequence":"additional","affiliation":[{"name":"NVIDIA, Santa Clara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0231-7741","authenticated-orcid":false,"given":"Chenran","family":"Li","sequence":"additional","affiliation":[{"name":"NVIDIA, Santa Clara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1842-7622","authenticated-orcid":false,"given":"Zhengyi","family":"Luo","sequence":"additional","affiliation":[{"name":"NVIDIA, Santa Clara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-8956-7031","authenticated-orcid":false,"given":"Brian","family":"Robison","sequence":"additional","affiliation":[{"name":"NVIDIA, Moraga, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-9905-4409","authenticated-orcid":false,"given":"Xavier","family":"Blackwell","sequence":"additional","affiliation":[{"name":"NVIDIA, Santa Clara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1709-2578","authenticated-orcid":false,"given":"Bernardo","family":"Antoniazzi","sequence":"additional","affiliation":[{"name":"NVIDIA, Los Angeles, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3677-5655","authenticated-orcid":false,"given":"Xue Bing","family":"Peng","sequence":"additional","affiliation":[{"name":"NVIDIA, Vancouver, Canada"},{"name":"Simon Fraser University, Vancouver, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9198-2227","authenticated-orcid":false,"given":"Yuke","family":"Zhu","sequence":"additional","affiliation":[{"name":"NVIDIA, SANTA CLARA, USA"},{"name":"The University of Texas at Austin, SANTA CLARA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-9012-0656","authenticated-orcid":false,"given":"Simon","family":"Yuen","sequence":"additional","affiliation":[{"name":"NVIDIA, Los Angeles, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,3]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392462"},{"key":"e_1_2_2_2_1","volume-title":"Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.","author":"Achiam Josh","year":"2023","unstructured":"Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3592458"},{"key":"e_1_2_2_4_1","volume-title":"Retargeting matters: General motion retargeting for humanoid motion tracking. arXiv preprint arXiv:2510.02252","author":"Araujo Joao Pedro","year":"2025","unstructured":"Joao Pedro Araujo, Yanjie Ze, Pei Xu, Jiajun Wu, and C Karen Liu. 2025. Retargeting matters: General motion retargeting for humanoid motion tracking. arXiv preprint arXiv:2510.02252 (2025)."},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/566654.566606"},{"key":"e_1_2_2_6_1","volume-title":"Shadow of HyperPose: New Animation System. In ACM SIGGRAPH 2024 Talks. 1\u20132.","author":"Bereznyak Alex","year":"2024","unstructured":"Alex Bereznyak. 2024. Shadow of HyperPose: New Animation System. In ACM SIGGRAPH 2024 Talks. 1\u20132."},{"key":"e_1_2_2_7_1","unstructured":"Bones Studio. 2026. BONES-SEED: Skeletal Everyday Embodied Dataset. https:\/\/bones.studio\/datasets. Open-source motion dataset approximately 140k motion clips.."},{"key":"e_1_2_2_8_1","unstructured":"Michael Buttner. 2015. Motion Matching-The Road to Next-Gen Animation. In Nucl. ai Conference."},{"key":"e_1_2_2_9_1","volume-title":"Interactive 3D Graphics and Games (I3D)","author":"Buttner Michael","year":"2019","unstructured":"Michael Buttner. 2019. Machine Learning for Motion Synthesis and Character Control. In Interactive 3D Graphics and Games (I3D) 2019. https:\/\/www.youtube.com\/watch?v=zuvmQxcCOM4"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3641519.3657440"},{"key":"e_1_2_2_11_1","volume-title":"Hand-eye autonomous delivery: Learning humanoid navigation, locomotion and reaching. arXiv preprint arXiv:2508.03068","author":"Chen Sirui","year":"2025","unstructured":"Sirui Chen, Yufei Ye, Zi-ang Cao, Jennifer Lew, Pei Xu, and C Karen Liu. 2025b. Hand-eye autonomous delivery: Learning humanoid navigation, locomotion and reaching. arXiv preprint arXiv:2508.03068 (2025)."},{"key":"e_1_2_2_12_1","volume-title":"Xue Bin Peng, and Xiaolong Wang","author":"Chen Zixuan","year":"2025","unstructured":"Zixuan Chen, Mazeyu Ji, Xuxin Cheng, Xuanbin Peng, Xue Bin Peng, and Xiaolong Wang. 2025a. GMT: General Motion Tracking for Humanoid Whole-Body Control. arXiv preprint arXiv:2506.14770 (2025)."},{"key":"e_1_2_2_13_1","volume-title":"Proc. of GDC 2016","author":"Clavet Simon","year":"2016","unstructured":"Simon Clavet. 2016. Motion matching and the road to next-gen animation. Proc. of GDC 2016 (2016)."},{"key":"e_1_2_2_14_1","volume-title":"ACM SIGGRAPH 2024 Conference Papers. 1\u20139.","author":"Cohan Setareh","unstructured":"Setareh Cohan, Guy Tevet, Daniele Reda, Xue Bin Peng, and Michiel van de Panne. 2024. Flexible motion in-betweening with diffusion models. In ACM SIGGRAPH 2024 Conference Papers. 1\u20139."},{"key":"e_1_2_2_15_1","unstructured":"Boston Dynamics. 2025. Use Choreographer with Spot. https:\/\/support.bostondynamics.com\/s\/article\/Use-Choreographer-with-Spot-72036."},{"key":"e_1_2_2_16_1","unstructured":"Unreal Engine. 2025. Graphing in Animation Blueprints. https:\/\/dev.epicgames.com\/documentation\/en-us\/unreal-engine\/graphing-in-animation-blueprints-in-unreal-engine."},{"key":"e_1_2_2_17_1","unstructured":"Epic Games. 2025. Blend Spaces in Unreal Engine. https:\/\/dev.epicgames.com\/documentation\/en-us\/unreal-engine\/blend-spaces-in-unreal-engine. Accessed: 2026-01-22."},{"key":"e_1_2_2_18_1","volume-title":"Forty-first international conference on machine learning.","author":"Esser Patrick","year":"2024","unstructured":"Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M\u00fcller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. 2024. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first international conference on machine learning."},{"key":"e_1_2_2_19_1","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 12873\u201312883","author":"Esser Patrick","year":"2021","unstructured":"Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for highresolution image synthesis. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 12873\u201312883."},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.494"},{"key":"e_1_2_2_21_1","unstructured":"Epic Games. 2025. IK Rig Animation Retargeting in Unreal Engine. https:\/\/dev.epicgames.com\/documentation\/en-us\/unreal-engine\/ik-rig-animation-retargeting-in-unreal-engine."},{"key":"e_1_2_2_22_1","unstructured":"Google. 2025. Veo 3. https:\/\/deepmind.google\/technologies\/veo\/"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01239"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3763319"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00186"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00509"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19833-5_34"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i3.27973"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392480"},{"key":"e_1_2_2_30_1","volume-title":"VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation. arXiv preprint arXiv:2511.15200","author":"He Tairan","year":"2025","unstructured":"Tairan He, Zi Wang, Haoru Xue, Qingwei Ben, Zhengyi Luo, Wenli Xiao, Ye Yuan, Xingye Da, Fernando Casta\u00f1eda, Shankar Sastry, et al. 2025. VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation. arXiv preprint arXiv:2511.15200 (2025)."},{"key":"e_1_2_2_31_1","volume-title":"Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_32_1","unstructured":"Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. In Advances in Neural Information Processing Systems. 4565\u20134573."},{"key":"e_1_2_2_33_1","volume-title":"Denoising diffusion probabilistic models. Advances in neural information processing systems 33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840\u20136851."},{"key":"e_1_2_2_34_1","volume-title":"Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598","author":"Ho Jonathan","year":"2022","unstructured":"Jonathan Ho and Tim Salimans. 2022. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598 (2022)."},{"key":"e_1_2_2_35_1","first-page":"2","article-title":"Character control with neural networks and machine learning","volume":"1","author":"Holden Daniel","year":"2018","unstructured":"Daniel Holden. 2018. Character control with neural networks and machine learning. In Proc. of GDC, Vol. 1. 2.","journal-title":"Proc. of GDC"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392440"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073663"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925975"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00889"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-0880"},{"key":"e_1_2_2_41_1","volume-title":"PyRoki: A Modular Toolkit for Robot Kinematic Optimization. arXiv preprint arXiv:2505.03728","author":"Kim Chung Min","year":"2025","unstructured":"Chung Min Kim, Brent Yi, Hongsuk Choi, Yi Ma, Ken Goldberg, and Angjoo Kanazawa. 2025. PyRoki: A Modular Toolkit for Robot Kinematic Optimization. arXiv preprint arXiv:2505.03728 (2025)."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.108894"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ymssp.2020.107398"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/566654.566605"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015706.1015760"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/566570.566607"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/1866158.1866160"},{"key":"e_1_2_2_48_1","volume-title":"GENMO: A GENeralist Model for Human MOtion. arXiv preprint arXiv:2505.01425","author":"Li Jiefeng","year":"2025","unstructured":"Jiefeng Li, Jinkun Cao, Haotian Zhang, Davis Rempe, Jan Kautz, Umar Iqbal, and Ye Yuan. 2025. GENMO: A GENeralist Model for Human MOtion. arXiv preprint arXiv:2505.01425 (2025)."},{"key":"e_1_2_2_49_1","volume-title":"European Conference on Computer Vision. Springer, 54\u201372","author":"Li Jiaman","year":"2024","unstructured":"Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu, Xavier Puig, and C Karen Liu. 2024a. Controllable human-object interaction synthesis. In European Conference on Computer Vision. Springer, 54\u201372."},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3618333"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00178"},{"key":"e_1_2_2_52_1","volume-title":"Heli Ben-Hamu, Maximilian Nickel, and Matt Le.","author":"Lipman Yaron","year":"2022","unstructured":"Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2022. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747 (2022)."},{"key":"e_1_2_2_53_1","volume-title":"CHOICE: Coordinated human-object interaction in cluttered environments for pick-and-place actions. arXiv preprint arXiv:2412.06702","author":"Lu Jintao","year":"2024","unstructured":"Jintao Lu, He Zhang, Yuting Ye, Takaaki Shiratori, Sebastian Starke, and Taku Komura. 2024. CHOICE: Coordinated human-object interaction in cluttered environments for pick-and-place actions. arXiv preprint arXiv:2412.06702 (2024)."},{"key":"e_1_2_2_54_1","volume-title":"Open-magvit2: An open-source project toward democratizing auto-regressive visual generation. arXiv preprint arXiv:2409.04410","author":"Luo Zhuoyan","year":"2024","unstructured":"Zhuoyan Luo, Fengyuan Shi, Yixiao Ge, Yujiu Yang, Limin Wang, and Ying Shan. 2024. Open-magvit2: An open-source project toward democratizing auto-regressive visual generation. arXiv preprint arXiv:2409.04410 (2024)."},{"key":"e_1_2_2_55_1","volume-title":"Sonic: Supersizing motion tracking for natural humanoid whole-body control. arXiv preprint arXiv:2511.07820","author":"Luo Zhengyi","year":"2025","unstructured":"Zhengyi Luo, Ye Yuan, Tingwu Wang, Chenran Li, Sirui Chen, Fernando Casta\u00f1eda, Zi-Ang Cao, Jiefeng Li, David Minor, Qingwei Ben, et al. 2025. Sonic: Supersizing motion tracking for natural humanoid whole-body control. arXiv preprint arXiv:2511.07820 (2025)."},{"key":"e_1_2_2_56_1","volume-title":"Finite scalar quantization: Vq-vae made simple. arXiv preprint arXiv:2309.15505","author":"Mentzer Fabian","year":"2023","unstructured":"Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen. 2023. Finite scalar quantization: Vq-vae made simple. arXiv preprint arXiv:2309.15505 (2023)."},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2366145.2366172"},{"key":"e_1_2_2_58_1","volume-title":"From Generated Human Videos to Physically Plausible Robot Trajectories. arXiv preprint arXiv:2512.05094","author":"Ni James","year":"2025","unstructured":"James Ni, Zekai Wang, Wei Lin, Amir Bar, Yann LeCun, Trevor Darrell, Jitendra Malik, and Roei Herzig. 2025. From Generated Human Videos to Physically Plausible Robot Trajectories. arXiv preprint arXiv:2512.05094 (2025)."},{"key":"e_1_2_2_59_1","volume-title":"Sora: Creating Video From Text. https:\/\/openai.com\/sora","author":"AI.","year":"2024","unstructured":"OpenAI. 2024. Sora: Creating Video From Text. https:\/\/openai.com\/sora"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2023.3309107"},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201311"},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00870"},{"key":"e_1_2_2_63_1","volume-title":"Korrawe Karunratanakul, Pu Wang, Hongfei Xue, Chen Chen, Chuan Guo, Junli Cao, Jian Ren, and Sergey Tulyakov.","author":"Pinyoanuntapong Ekkasit","year":"2024","unstructured":"Ekkasit Pinyoanuntapong, Muhammad Usama Saleem, Korrawe Karunratanakul, Pu Wang, Hongfei Xue, Chen Chen, Chuan Guo, Junli Cao, Jian Ren, and Sergey Tulyakov. 2024a. Controlmm: Controllable masked motion generation. arXiv preprint arXiv:2410.10780 (2024)."},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00153"},{"key":"e_1_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3550454.3555454"},{"key":"e_1_2_2_66_1","volume-title":"Aaron Van den Oord, and Oriol Vinyals","author":"Razavi Ali","year":"2019","unstructured":"Ali Razavi, Aaron Van den Oord, and Oriol Vinyals. 2019. Generating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems 32 (2019)."},{"key":"e_1_2_2_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/1275808.1276510"},{"key":"e_1_2_2_69_1","volume-title":"Human motion diffusion as a generative prior. arXiv preprint arXiv:2303.01418","author":"Shafir Yonatan","year":"2023","unstructured":"Yonatan Shafir, Guy Tevet, Roy Kapon, and Amit H Bermano. 2023. Human motion diffusion as a generative prior. arXiv preprint arXiv:2303.01418 (2023)."},{"key":"e_1_2_2_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/3658140"},{"key":"e_1_2_2_71_1","volume-title":"Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502","author":"Song Jiaming","year":"2020","unstructured":"Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502 (2020)."},{"key":"e_1_2_2_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/3606921"},{"key":"e_1_2_2_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/3658209"},{"key":"e_1_2_2_74_1","doi-asserted-by":"publisher","DOI":"10.1145\/3355089.3356505"},{"key":"e_1_2_2_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392450"},{"key":"e_1_2_2_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459881"},{"key":"e_1_2_2_77_1","volume-title":"Autoregressive model beats diffusion: Llama for scalable image generation. arXiv preprint arXiv:2406.06525","author":"Sun Peize","year":"2024","unstructured":"Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan. 2024. Autoregressive model beats diffusion: Llama for scalable image generation. arXiv preprint arXiv:2406.06525 (2024)."},{"key":"e_1_2_2_78_1","volume-title":"Scaling Flow Matching Models for Text-To-Motion Generation. arXiv preprint arXiv:2512.23464","author":"Digital Human Team Tencent Hunyuan","year":"2025","unstructured":"Tencent Hunyuan 3D Digital Human Team. 2025. HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation. arXiv preprint arXiv:2512.23464 (2025)."},{"key":"e_1_2_2_79_1","volume-title":"The Thirteenth International Conference on Learning Representations.","author":"Tevet Guy","unstructured":"Guy Tevet, Sigal Raab, Setareh Cohan, Daniele Reda, Zhengyi Luo, Xue Bin Peng, Amit Haim Bermano, and Michiel van de Panne. 2025. CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control. In The Thirteenth International Conference on Learning Representations."},{"key":"e_1_2_2_80_1","volume-title":"Human motion diffusion model. arXiv preprint arXiv:2209.14916","author":"Tevet Guy","year":"2022","unstructured":"Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano. 2022. Human motion diffusion model. arXiv preprint arXiv:2209.14916 (2022)."},{"key":"e_1_2_2_81_1","volume-title":"Visual autoregressive modeling: Scalable image generation via next-scale prediction. Advances in neural information processing systems 37","author":"Tian Keyu","year":"2024","unstructured":"Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang. 2024. Visual autoregressive modeling: Scalable image generation via next-scale prediction. Advances in neural information processing systems 37 (2024), 84839\u201384865."},{"key":"e_1_2_2_82_1","unstructured":"Unitree. 2025. Unitree Boxing. https:\/\/www.unitree.com\/boxing. Accessed: 2024-06-30."},{"key":"e_1_2_2_83_1","unstructured":"Unity. 2025. Unity User Manual (6.2). https:\/\/docs.unity3d.com\/Manual\/class-BlendTree.html."},{"key":"e_1_2_2_84_1","unstructured":"Aaron Van Den Oord Oriol Vinyals et al. 2017. Neural discrete representation learning. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_85_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_86_1","volume-title":"UniCon: Universal Neural Controller For Physics-based Character Motion. arXiv preprint arXiv:2011.15119","author":"Wang Tingwu","year":"2020","unstructured":"Tingwu Wang, Yunrong Guo, Maria Shugrina, and Sanja Fidler. 2020. UniCon: Universal Neural Controller For Physics-based Character Motion. arXiv preprint arXiv:2011.15119 (2020)."},{"key":"e_1_2_2_87_1","unstructured":"Boran Wen Ye Lu Keyan Wan Sirui Wang Jiahong Zhou Junxuan Liang Xinpeng Liu Bang Xiao Dingbang Huang Ruiyang Liu et al. 2025. Efficient and Scalable Monocular Human-Object Interaction Motion Reconstruction. arXiv preprint arXiv:2512.00960 (2025)."},{"key":"e_1_2_2_88_1","volume-title":"Omniretarget: Interaction-preserving data generation for humanoid whole-body loco-manipulation and scene interaction. arXiv preprint arXiv:2509.26633","author":"Yang Lujie","year":"2025","unstructured":"Lujie Yang, Xiaoyu Huang, Zhen Wu, Angjoo Kanazawa, Pieter Abbeel, Carmelo Sferrazza, C Karen Liu, Rocky Duan, and Guanya Shi. 2025. Omniretarget: Interaction-preserving data generation for humanoid whole-body loco-manipulation and scene interaction. arXiv preprint arXiv:2509.26633 (2025)."},{"key":"e_1_2_2_89_1","doi-asserted-by":"crossref","unstructured":"Gwonjin Yi and Junghoon Jee. 2019. Search Space Reduction In Motion Matching by Trajectory Clustering. In SIGGRAPH Asia 2019 Posters. 1\u20132.","DOI":"10.1145\/3355056.3364558"},{"key":"e_1_2_2_90_1","unstructured":"Kangning Yin Weishuai Zeng Ke Fan Minyue Dai Zirui Wang Qiang Zhang Zheng Tian Jingbo Wang Jiangmiao Pang and Weinan Zhang. 2025. Uni-Tracker: Learning Universal Whole-Body Motion Tracker for Humanoid Robots. arXiv:2507.07356 [cs.RO] https:\/\/arxiv.org\/abs\/2507.07356"},{"key":"e_1_2_2_91_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01008"},{"key":"e_1_2_2_92_1","unstructured":"Lijun Yu Jos\u00e9 Lezama Nitesh B Gundavarapu Luca Versari Kihyuk Sohn David Minnen Yong Cheng Vighnesh Birodkar Agrim Gupta Xiuye Gu et al. 2023b. Language Model Beats Diffusion-Tokenizer is Key to Visual Generation. arXiv preprint arXiv:2310.05737 (2023)."},{"key":"e_1_2_2_93_1","volume-title":"Jiajun Wu, and C. Karen Liu.","author":"Ze Yanjie","year":"2025","unstructured":"Yanjie Ze, Jo\u00e3o Pedro Ara\u00fajo, Jiajun Wu, and C. Karen Liu. 2025. GMR: General Motion Retargeting. https:\/\/github.com\/YanjieZe\/GMR GitHub repository."},{"key":"e_1_2_2_94_1","volume-title":"Behavior Foundation Model for Humanoid Robots. arXiv preprint arXiv:2509.13780","author":"Zeng Weishuai","year":"2025","unstructured":"Weishuai Zeng, Shunlin Lu, Kangning Yin, Xiaojie Niu, Minyue Dai, Jingbo Wang, and Jiangmiao Pang. 2025. Behavior Foundation Model for Humanoid Robots. arXiv preprint arXiv:2509.13780 (2025)."},{"key":"e_1_2_2_95_1","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275090"},{"key":"e_1_2_2_96_1","volume-title":"Motiondiffuse: Text-driven human motion generation with diffusion model","author":"Zhang Mingyuan","year":"2024","unstructured":"Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. 2024a. Motiondiffuse: Text-driven human motion generation with diffusion model. IEEE transactions on pattern analysis and machine intelligence 46, 6 (2024), 4115\u20134128."},{"key":"e_1_2_2_97_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i7.28567"},{"key":"e_1_2_2_98_1","unstructured":"Zhikai Zhang Jun Guo Chao Chen Jilong Wang Chenghuai Lin Yunrui Lian Han Xue Zhenrong Wang Maoqi Liu Huaping Liu et al. 2025. Track Any Motions under Any Disturbances. arXiv preprint arXiv:2509.13833 (2025)."},{"key":"e_1_2_2_99_1","volume-title":"European Conference on Computer Vision. Springer, 18\u201338","author":"Zhou Wenyang","year":"2024","unstructured":"Wenyang Zhou, Zhiyang Dou, Zeyu Cao, Zhouyingcheng Liao, Jingbo Wang, Wenjia Wang, Yuan Liu, Taku Komura, Wenping Wang, and Lingjie Liu. 2024. Emdm: Efficient motion diffusion model for fast and high-quality motion generation. In European Conference on Computer Vision. Springer, 18\u201338."},{"key":"e_1_2_2_100_1","volume-title":"Hiking in the Wild: A Scalable Perceptive Parkour Framework for Humanoids. arXiv preprint arXiv:2601.07718","author":"Zhu Shaoting","year":"2026","unstructured":"Shaoting Zhu, Ziwen Zhuang, Mengjie Zhao, Kun-Ying Lee, and Hang Zhao. 2026. Hiking in the Wild: A Scalable Perceptive Parkour Framework for Humanoids. arXiv preprint arXiv:2601.07718 (2026)."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","deposited":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:50:36Z","timestamp":1783065036000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811334"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,3]]},"references-count":100,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,7,3]]}},"alternative-id":["10.1145\/3811334"],"URL":"https:\/\/doi.org\/10.1145\/3811334","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,3]]},"assertion":[{"value":"2026-01-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}