{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T03:25:34Z","timestamp":1783740334489,"version":"3.55.0"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2025,8,1]]},"abstract":"<jats:p>\n                    We present\n                    <jats:bold>\n                      <jats:italic toggle=\"yes\">RigAnything<\/jats:italic>\n                    <\/jats:bold>\n                    , a novel autoregressive transformer-based model, which makes 3D assets rig-ready by probabilistically generating joints and skeleton topologies and assigning skinning weights in a template-free manner. Unlike most existing auto-rigging methods, which rely on predefined skeleton templates and are limited to specific categories like humanoid, RigAnything approaches the rigging problem in an autoregressive manner, iteratively predicting the next joint based on the global input shape and the previous prediction. While autoregressive models are typically used to generate sequential data, RigAnything extends its application to effectively learn and represent skeletons, which are inherently tree structures. To achieve this, we organize the joints in a breadth-first search (BFS) order, enabling the skeleton to be defined as a sequence of 3D locations and the parent index. Furthermore, our model improves the accuracy of position prediction by leveraging diffusion modeling, ensuring precise and consistent placement of joints within the hierarchy. This formulation allows the autoregressive model to efficiently capture both spatial and hierarchical relationships within the skeleton. Trained end-to-end on both RigNet and Objaverse datasets, RigAnything demonstrates state-of-the-art performance across diverse object types, including humanoids, quadrupeds, marine creatures, insects, and many more, surpassing prior methods in quality, robustness, generalizability, and efficiency. It achieves significantly faster performance than existing auto-rigging methods, completing rigging in under a few seconds per shape.\n                  <\/jats:p>","DOI":"10.1145\/3731149","type":"journal-article","created":{"date-parts":[[2025,7,27]],"date-time":"2025-07-27T04:02:22Z","timestamp":1753588942000},"page":"1-12","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["RigAnything: Template-Free Autoregressive Rigging for Diverse 3D Assets"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2882-2383","authenticated-orcid":false,"given":"Isabella","family":"Liu","sequence":"first","affiliation":[{"name":"University of California San Diego, La Jolla, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5537-1509","authenticated-orcid":false,"given":"Zhan","family":"Xu","sequence":"additional","affiliation":[{"name":"Adobe Research, San Jose, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2275-7288","authenticated-orcid":false,"given":"Wang","family":"Yifan","sequence":"additional","affiliation":[{"name":"Adobe Research, San Jose, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9755-9040","authenticated-orcid":false,"given":"Hao","family":"Tan","sequence":"additional","affiliation":[{"name":"Adobe Research, San Jose, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8487-8018","authenticated-orcid":false,"given":"Zexiang","family":"Xu","sequence":"additional","affiliation":[{"name":"Hillbot Inc., San Jose, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3150-778X","authenticated-orcid":false,"given":"Xiaolong","family":"Wang","sequence":"additional","affiliation":[{"name":"University of California San Diego, La Jolla, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1796-2682","authenticated-orcid":false,"given":"Hao","family":"Su","sequence":"additional","affiliation":[{"name":"University of California San Diego, La Jolla, USA"},{"name":"Hillbot Inc., La Jolla, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0692-2869","authenticated-orcid":false,"given":"Zifan","family":"Shi","sequence":"additional","affiliation":[{"name":"Adobe Research, San Jose, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,7,27]]},"reference":[{"key":"e_1_2_2_1_1","volume-title":"Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.","author":"Achiam Josh","year":"2023","unstructured":"Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)."},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00764"},{"key":"e_1_2_2_3_1","volume-title":"Automatic rigging and animation of 3d characters. ACM Transactions on graphics (TOG)","author":"Baran Ilya","year":"2007","unstructured":"Ilya Baran and Jovan Popovi\u0107. 2007. Automatic rigging and animation of 3d characters. ACM Transactions on graphics (TOG) (2007)."},{"key":"e_1_2_2_4_1","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. Advances in neural information processing systems (2020)."},{"key":"e_1_2_2_5_1","volume-title":"International conference on machine learning. PMLR, 1691\u20131703","author":"Chen Mark","year":"2020","unstructured":"Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. 2020. Generative pretraining from pixels. In International conference on machine learning. PMLR, 1691\u20131703."},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20062-5_6"},{"key":"e_1_2_2_7_1","volume-title":"HumanRig: Learning Automatic Rigging for Humanoid Character in a Large Scale Dataset. arXiv preprint arXiv:2412.02317","author":"Chu Zedong","year":"2024","unstructured":"Zedong Chu, Feng Xiong, Meiduo Liu, Jinzhi Zhang, Mingqi Shao, Zhaoxu Sun, Di Wang, and Mu Xu. 2024. HumanRig: Learning Automatic Rigging for Humanoid Character in a Large Scale Dataset. arXiv preprint arXiv:2412.02317 (2024)."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01263"},{"key":"e_1_2_2_9_1","volume-title":"Diffusion models beat gans on image synthesis. Advances in neural information processing systems 34","author":"Dhariwal Prafulla","year":"2021","unstructured":"Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion models beat gans on image synthesis. Advances in neural information processing systems 34 (2021), 8780\u20138794."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01268"},{"key":"e_1_2_2_11_1","volume-title":"Make-It-Animatable: An Efficient Framework for Authoring Animation-Ready 3D Characters. arXiv preprint arXiv:2411.18197","author":"Guo Zhiyang","year":"2024","unstructured":"Zhiyang Guo, Jinxu Xiang, Kai Ma, Wengang Zhou, Houqiang Li, and Ran Zhang. 2024. Make-It-Animatable: An Efficient Framework for Authoring Animation-Ready 3D Characters. arXiv preprint arXiv:2411.18197 (2024)."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i3.27973"},{"key":"e_1_2_2_13_1","volume-title":"Denoising diffusion probabilistic models. Advances in neural information processing systems 33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840\u20136851."},{"key":"e_1_2_2_14_1","volume-title":"Lrm: Large reconstruction model for single image to 3d. arXiv preprint arXiv:2311.04400","author":"Hong Yicong","year":"2023","unstructured":"Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. 2023. Lrm: Large reconstruction model for single image to 3d. arXiv preprint arXiv:2311.04400 (2023)."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW59228.2023.00270"},{"key":"e_1_2_2_16_1","volume-title":"Shap-e: Generating conditional 3d implicit functions. arXiv preprint arXiv:2305.02463","author":"Jun Heewoo","year":"2023","unstructured":"Heewoo Jun and Alex Nichol. 2023. Shap-e: Generating conditional 3d implicit functions. arXiv preprint arXiv:2305.02463 (2023)."},{"key":"e_1_2_2_17_1","volume-title":"Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model. arXiv preprint arXiv:2311.06214","author":"Li Jiahao","year":"2023","unstructured":"Jiahao Li, Hao Tan, Kai Zhang, Zexiang Xu, Fujun Luan, Yinghao Xu, Yicong Hong, Kalyan Sunkavalli, Greg Shakhnarovich, and Sai Bi. 2023. Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model. arXiv preprint arXiv:2311.06214 (2023)."},{"key":"e_1_2_2_18_1","volume-title":"Learning skeletal articulations with neural blend shapes. ACM Transactions on Graphics (TOG)","author":"Li Peizhuo","year":"2021","unstructured":"Peizhuo Li, Kfir Aberman, Rana Hanocka, Libin Liu, Olga Sorkine-Hornung, and Baoquan Chen. 2021. Learning skeletal articulations with neural blend shapes. ACM Transactions on Graphics (TOG) (2021)."},{"key":"e_1_2_2_19_1","volume-title":"Autoregressive Image Generation without Vector Quantization. arXiv preprint arXiv:2406.11838","author":"Li Tianhong","year":"2024","unstructured":"Tianhong Li, Yonglong Tian, He Li, Mingyang Deng, and Kaiming He. 2024. Autoregressive Image Generation without Vector Quantization. arXiv preprint arXiv:2406.11838 (2024)."},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392422"},{"key":"e_1_2_2_21_1","volume-title":"Dynamic Gaussians Mesh: Consistent Mesh Reconstruction from Monocular Videos. arXiv preprint arXiv:2404.12379","author":"Liu Isabella","year":"2024","unstructured":"Isabella Liu, Hao Su, and Xiaolong Wang. 2024a. Dynamic Gaussians Mesh: Consistent Mesh Reconstruction from Monocular Videos. arXiv preprint arXiv:2404.12379 (2024)."},{"key":"e_1_2_2_22_1","volume-title":"One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization. Advances in Neural Information Processing Systems","author":"Liu Minghua","year":"2024","unstructured":"Minghua Liu, Chao Xu, Haian Jin, Linghao Chen, Mukund Varma T, Zexiang Xu, and Hao Su. 2024b. One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization. Advances in Neural Information Processing Systems (2024)."},{"key":"e_1_2_2_23_1","volume-title":"Meshformer: High-quality mesh generation with 3d-guided reconstruction model. arXiv preprint arXiv:2408.10198","author":"Liu Minghua","year":"2024","unstructured":"Minghua Liu, Chong Zeng, Xinyue Wei, Ruoxi Shi, Linghao Chen, Chao Xu, Mengqi Zhang, Zhaoning Wang, Xiaoshuai Zhang, Isabella Liu, et al. 2024c. Meshformer: High-quality mesh generation with 3d-guided reconstruction model. arXiv preprint arXiv:2408.10198 (2024)."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00853"},{"key":"e_1_2_2_25_1","volume-title":"TARig: Adaptive template-aware neural rigging for humanoid characters. Computers & Graphics","author":"Ma Jing","year":"2023","unstructured":"Jing Ma and Dongliang Zhang. 2023. TARig: Adaptive template-aware neural rigging for humanoid characters. Computers & Graphics (2023)."},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00040"},{"key":"e_1_2_2_27_1","volume-title":"Point-e: A system for generating 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751","author":"Nichol Alex","year":"2022","unstructured":"Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. 2022. Point-e: A system for generating 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751 (2022)."},{"key":"e_1_2_2_28_1","volume-title":"International conference on machine learning. PMLR, 8162\u20138171","author":"Nichol Alexander Quinn","year":"2021","unstructured":"Alexander Quinn Nichol and Prafulla Dhariwal. 2021. Improved denoising diffusion probabilistic models. In International conference on machine learning. PMLR, 8162\u20138171."},{"key":"e_1_2_2_29_1","volume-title":"International conference on machine learning.","author":"Parmar Niki","year":"2018","unstructured":"Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018. Image transformer. In International conference on machine learning."},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"e_1_2_2_31_1","unstructured":"Xuelin Qian Yu Wang Simian Luo Yinda Zhang Ying Tai Zhenyu Zhang Chengjie Wang Xiangyang Xue Bo Zhao Tiejun Huang et al. 2024. Pushing Auto-regressive Models for 3D Shape Generation at Capacity and Scalability. arXiv preprint arXiv:2402.12225 (2024)."},{"key":"e_1_2_2_32_1","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei Ilya Sutskever et al. 2019. Language models are unsupervised multitask learners. OpenAI blog (2019)."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01129"},{"key":"e_1_2_2_34_1","volume-title":"Dreamgaussian4d: Generative 4d gaussian splatting. arXiv preprint arXiv:2312.17142","author":"Ren Jiawei","year":"2023","unstructured":"Jiawei Ren, Liang Pan, Jiaxiang Tang, Chi Zhang, Ang Cao, Gang Zeng, and Ziwei Liu. 2023. Dreamgaussian4d: Generative 4d gaussian splatting. arXiv preprint arXiv:2312.17142 (2023)."},{"key":"e_1_2_2_35_1","volume-title":"Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512","author":"Shi Yichun","year":"2023","unstructured":"Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. 2023. Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512 (2023)."},{"key":"e_1_2_2_36_1","unstructured":"Uriel Singer Shelly Sheynin Adam Polyak Oron Ashual Iurii Makarov Filippos Kokkinos Naman Goyal Andrea Vedaldi Devi Parikh Justin Johnson et al. 2023. Text-to-4d dynamic scene generation. arXiv preprint arXiv:2301.11280 (2023)."},{"key":"e_1_2_2_37_1","unstructured":"A Waswani N Shazeer N Parmar J Uszkoreit L Jones A Gomez L Kaiser and I Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems."},{"key":"e_1_2_2_38_1","unstructured":"Yinghao Xu Hao Tan Fujun Luan Sai Bi Peng Wang Jiahao Li Zifan Shi Kalyan Sunkavalli Gordon Wetzstein Zexiang Xu et al. 2023. Dmv3d: Denoising multiview diffusion using 3d large reconstruction model. arXiv preprint arXiv:2311.09217 (2023)."},{"key":"e_1_2_2_39_1","volume-title":"Rignet: Neural rigging for articulated characters. arXiv preprint arXiv:2005.00559","author":"Xu Zhan","year":"2020","unstructured":"Zhan Xu, Yang Zhou, Evangelos Kalogerakis, Chris Landreth, and Karan Singh. 2020. Rignet: Neural rigging for articulated characters. arXiv preprint arXiv:2005.00559 (2020)."},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00614"},{"key":"e_1_2_2_41_1","volume-title":"4dgen: Grounded 4d content generation with spatial-temporal consistency. arXiv preprint arXiv:2312.17225","author":"Yin Yuyang","year":"2023","unstructured":"Yuyang Yin, Dejia Xu, Zhangyang Wang, Yao Zhao, and Yunchao Wei. 2023. 4dgen: Grounded 4d content generation with spatial-temporal consistency. arXiv preprint arXiv:2312.17225 (2023)."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01415"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-72670-5_1"},{"key":"e_1_2_2_44_1","volume-title":"Animate124: Animating one image to 4d dynamic scene. arXiv preprint arXiv:2311.14603","author":"Zhao Yuyang","year":"2023","unstructured":"Yuyang Zhao, Zhiwen Yan, Enze Xie, Lanqing Hong, Zhenguo Li, and Gim Hee Lee. 2023. Animate124: Animating one image to 4d dynamic scene. arXiv preprint arXiv:2311.14603 (2023)."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3731149","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T17:58:02Z","timestamp":1774634282000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3731149"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,27]]},"references-count":44,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,8,1]]}},"alternative-id":["10.1145\/3731149"],"URL":"https:\/\/doi.org\/10.1145\/3731149","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,27]]},"assertion":[{"value":"2025-07-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}