{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T08:16:46Z","timestamp":1783066606254,"version":"3.54.6"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T00:00:00Z","timestamp":1783036800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["W2431046"],"award-info":[{"award-number":["W2431046"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2025YFA1309603"],"award-info":[{"award-number":["2025YFA1309603"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Central Guided Local Science and Technology Foundation of China","award":["YDZX20253100001001"],"award-info":[{"award-number":["YDZX20253100001001"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2026,7,3]]},"abstract":"<jats:p>Recent advances in generative models have democratized the creation of high-quality static 3D assets, yet animating these meshes remains a labor-intensive bottleneck. Traditional pipelines fracture this process into sequential stages\u2014rigging, skinning, and motion synthesis\u2014ignoring the inherent coupling between morphological structure and motor function. To bridge this gap, we introduce ACT, a unified generative framework that reformulates rigging and animation not as independent tasks, but as complementary views of a single hyper-kinematic process.<\/jats:p>\n                  <jats:p>Our key insight is to model the joint distribution of skeletal topology and temporal motion within a shared latent space. ACT utilizes a Vision Language Model (VLM) to extract semantic topological priors from arbitrary meshes, which then condition a Diffusion Transformer (DiT) backbone. By treating static rest poses and dynamic trajectories as a unified sequence, our model employs a task-aware masking strategy to flexibly perform zero-shot rigging, text-guided motion generation, and motion completion within a single end-to-end architecture. Furthermore, a geometry-guided decoder ensures that surface deformations are tightly coupled with the generated kinematics. Extensive experiments demonstrate that ACT generalizes robustly to diverse, non-humanoid characters without retraining. By replacing brittle cascaded pipelines with a holistic prior, our method enables novel applications such as semantic-driven topology editing and generative in-betweening, offering a versatile and efficient solution for automating 3D character animation.<\/jats:p>","DOI":"10.1145\/3811392","type":"journal-article","created":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:05:51Z","timestamp":1783062351000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["ACT: A Unified Framework for Rigging and Animating Characters with Arbitrary Topologies"],"prefix":"10.1145","volume":"45","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-8490-2652","authenticated-orcid":false,"given":"Pengyu","family":"Long","sequence":"first","affiliation":[{"name":"ShanghaiTech University, Shanghai, China"},{"name":"ByteDance Games, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-3755-9967","authenticated-orcid":false,"given":"Weirui","family":"Wang","sequence":"additional","affiliation":[{"name":"ShanghaiTech University, Shanghai, China"},{"name":"Deemos Technology, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-8968-8341","authenticated-orcid":false,"given":"Qingcheng","family":"Zhao","sequence":"additional","affiliation":[{"name":"University of Toronto, Toronto, Canada"},{"name":"Deemos Technology, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7613-5920","authenticated-orcid":false,"given":"Xiaoyang","family":"Guo","sequence":"additional","affiliation":[{"name":"ByteDance Games, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7992-9581","authenticated-orcid":false,"given":"Xiaoyu","family":"Pan","sequence":"additional","affiliation":[{"name":"ByteDance Games, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4837-7152","authenticated-orcid":false,"given":"Qixuan","family":"Zhang","sequence":"additional","affiliation":[{"name":"ShanghaiTech University, Shanghai, China"},{"name":"Deemos Technology, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-6279-8501","authenticated-orcid":false,"given":"Jiaqing","family":"Zhou","sequence":"additional","affiliation":[{"name":"ByteDance Games, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0744-6454","authenticated-orcid":false,"given":"Tianlei","family":"Hu","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1189-1254","authenticated-orcid":false,"given":"Wei","family":"Yang","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8807-7787","authenticated-orcid":false,"given":"Lan","family":"Xu","sequence":"additional","affiliation":[{"name":"ShanghaiTech University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8580-0036","authenticated-orcid":false,"given":"Jingyi","family":"Yu","sequence":"additional","affiliation":[{"name":"ShanghaiTech University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,3]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1360612.1360643"},{"key":"e_1_2_1_2_1","volume-title":"Automatic rigging and animation of 3d characters. ACM Transactions on graphics (TOG) 26, 3","author":"Baran Ilya","year":"2007","unstructured":"Ilya Baran and Jovan Popovi\u0107. 2007. Automatic rigging and animation of 3d characters. ACM Transactions on graphics (TOG) 26, 3 (2007), 72\u2013es."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/SMI.2010.25"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01083"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3757377.3763811"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01726"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00037"},{"key":"e_1_2_1_8_1","unstructured":"Gheorghe Comanici Eric Bieber Mike Schaekermann Ice Pasupat Noveen Sachdeva Inderjit Dhillon Marcel Blistein Ori Ram Dan Zhang Evan Rosen et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning multimodality long context and next generation agentic capabilities. arXiv preprint arXiv:2507.06261 (2025)."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-1554"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01263"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3721238.3730743"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485895.2485919"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485895.2485919"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3771928"},{"key":"e_1_2_1_15_1","volume-title":"GaussianFlow: Splatting Gaussian Dynamics for 4D Content Creation. Transactions on Machine Learning Research","author":"Gao Quankai","year":"2025","unstructured":"Quankai Gao, Qiangeng Xu, Zhe Cao, Ben Mildenhall, Wenchao Ma, Le Chen, Danhang Tang, and Ulrich Neumann. 2025. GaussianFlow: Splatting Gaussian Dynamics for 4D Content Creation. Transactions on Machine Learning Research (2025)."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3721238.3730621"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00509"},{"key":"e_1_2_1_18_1","volume-title":"Denoising diffusion probabilistic models. Advances in neural information processing systems 33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840\u20136851."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2461913"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3757377.3763885"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964973"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-0880"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3999"},{"key":"e_1_2_1_24_1","volume-title":"The Twelfth International Conference on Learning Representations.","author":"Jiang Yanqin","year":"2024","unstructured":"Yanqin Jiang, Li Zhang, Jin Gao, Weiming Hu, and Yao Yao. 2024b. Consistent4D: Consistent 360\u00b0 Dynamic Object Generation from Monocular Video. In The Twelfth International Conference on Learning Representations."},{"key":"e_1_2_1_25_1","volume-title":"Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video. arXiv preprint arXiv:2601.05251","author":"Jiang Zeren","year":"2026","unstructured":"Zeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus, and Andrea Vedaldi. 2026. Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video. arXiv preprint arXiv:2601.05251 (2026)."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00205"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459852"},{"key":"e_1_2_1_28_1","volume-title":"Triposg: High-fidelity 3d shape synthesis using large-scale rectified flow models","author":"Li Yangguang","year":"2025","unstructured":"Yangguang Li, Zi-Xin Zou, Zexiang Liu, Dehu Wang, Yuan Liang, Zhipeng Yu, Xingchao Liu, Yuan-Chen Guo, Ding Liang, Wanli Ouyang, et al. 2025. Triposg: High-fidelity 3d shape synthesis using large-scale rectified flow models. IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-0674"},{"key":"e_1_2_1_30_1","volume-title":"The Thirty-eighth Annual Conference on Neural Information Processing Systems.","author":"Yuyang Yin HANWEN LIANG","year":"2024","unstructured":"HANWEN LIANG, Yuyang Yin, Dejia Xu, Zhangyang Wang, Konstantinos N Plataniotis, Yao Zhao, Yunchao Wei, et al. 2024. Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion Models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00426"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3731149"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3322969"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.00905"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818013"},{"key":"e_1_2_1_36_1","volume-title":"Decoupled Weight Decay Regularization. In International Conference on Learning Representations.","author":"Loshchilov Ilya","year":"2018","unstructured":"Ilya Loshchilov and Frank Hutter. 2018. Decoupled Weight Decay Regularization. In International Conference on Learning Representations."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.5555\/102313.102317"},{"key":"e_1_2_1_38_1","volume-title":"Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion Models. arXiv preprint arXiv:2503.15996","author":"San Mill\u00e1n Marc Bened\u00ed","year":"2025","unstructured":"Marc Bened\u00ed San Mill\u00e1n, Angela Dai, and Matthias Nie\u00dfner. 2025. Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion Models. arXiv preprint arXiv:2503.15996 (2025)."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01804"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01332"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3451262"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"e_1_2_1_43_1","volume-title":"The Eleventh International Conference on Learning Representations.","author":"Poole Ben","year":"2023","unstructured":"Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2023. DreamFusion: Text-to-3D using 2D Diffusion. In The Eleventh International Conference on Learning Representations."},{"key":"e_1_2_1_44_1","volume-title":"International conference on machine learning. PMLR, 8748\u20138763","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PMLR, 8748\u20138763."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-1810"},{"key":"e_1_2_1_46_1","volume-title":"The Thirty-ninth Annual Conference on Neural Information Processing Systems.","author":"Song Chaoyue","year":"2025","unstructured":"Chaoyue Song, Xiu Li, Fan Yang, Zhongcong Xu, Jiacheng Wei, Fayao Liu, Jiashi Feng, Guosheng Lin, and Jianfeng Zhang. 2025a. Puppeteer: Rig and Animate Your 3D Models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01491"},{"key":"e_1_2_1_48_1","volume-title":"The Twelfth International Conference on Learning Representations.","author":"Sun Jingxiang","year":"2024","unstructured":"Jingxiang Sun, Bo Zhang, Ruizhi Shao, Lizhen Wang, Wen Liu, Zhenda Xie, and Yebin Liu. 2024. DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior. In The Twelfth International Conference on Learning Representations."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01972"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8659.2012.03178.x"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20047-2_21"},{"key":"e_1_2_1_52_1","volume-title":"Human Motion Diffusion Model. In The Eleventh International Conference on Learning Representations.","author":"Tevet Guy","year":"2022","unstructured":"Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-Or, and Amit Haim Bermano. 2022b. Human Motion Diffusion Model. In The Eleventh International Conference on Learning Representations."},{"key":"e_1_2_1_53_1","unstructured":"Truebones Motion Animation Studios. 2022. Truebones. https:\/\/truebones.gumroad.com\/. Accessed: 2022-01-15."},{"key":"e_1_2_1_54_1","volume-title":"Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314","author":"Wan Team","year":"2025","unstructured":"Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. 2025. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314 (2025)."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01259"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392379"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3550469.3555390"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00525"},{"key":"e_1_2_1_59_1","volume-title":"Do transformers really perform badly for graph representation? Advances in neural information processing systems 34","author":"Ying Chengxuan","year":"2021","unstructured":"Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. 2021. Do transformers really perform badly for graph representation? Advances in neural information processing systems 34 (2021), 28877\u201328888."},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02592"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.00623"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3730930"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/3658146"},{"key":"e_1_2_1_64_1","volume-title":"Motiondiffuse: Text-driven human motion generation with diffusion model","author":"Zhang Mingyuan","year":"2024","unstructured":"Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. 2024a. Motiondiffuse: Text-driven human motion generation with diffusion model. IEEE transactions on pattern analysis and machine intelligence 46, 6 (2024), 4115\u20134128."},{"key":"e_1_2_1_65_1","volume-title":"Large Motion Model for Unified Multi-modal Motion Generation. In European Conference on Computer Vision. 397\u2013421","author":"Zhang Mingyuan","year":"2024","unstructured":"Mingyuan Zhang, Daisheng Jin, Chenyang Gu, Fangzhou Hong, Zhongang Cai, Jingfang Huang, Chongzhi Zhang, Xinying Guo, Lei Yang, Ying He, et al. 2024b. Large Motion Model for Unified Multi-modal Motion Generation. In European Conference on Computer Vision. 397\u2013421."},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00589"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","deposited":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:29:59Z","timestamp":1783063799000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811392"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,3]]},"references-count":66,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,7,3]]}},"alternative-id":["10.1145\/3811392"],"URL":"https:\/\/doi.org\/10.1145\/3811392","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,3]]},"assertion":[{"value":"2026-01-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}