{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T08:18:47Z","timestamp":1783066727666,"version":"3.54.6"},"reference-count":77,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T00:00:00Z","timestamp":1783036800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Hong Kong Research Grant Council","award":["17202422"],"award-info":[{"award-number":["17202422"]}]},{"name":"Hong Kong Research Grant Council","award":["17212923"],"award-info":[{"award-number":["17212923"]}]},{"name":"Hong Kong Research Grant Council","award":["17215025"],"award-info":[{"award-number":["17215025"]}]},{"name":"International (Hong Kong, Macao, and Taiwan) Collaborative R&D Project, Beijing Major Science and Technology Project","award":["Z251100007125016"],"award-info":[{"award-number":["Z251100007125016"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2026,7,3]]},"abstract":"<jats:p>\n                    Animatable 3D assets, defined as geometry equipped with an articulated skeleton and skinning weights, are fundamental to interactive graphics, embodied agents, and animation production. While recent 3D generative models can synthesize visually plausible shapes from images, the results are typically static. Obtaining usable rigs via post-hoc auto-rigging is brittle and often produces skeletons that are topologically inconsistent with the generated geometry. We present\n                    <jats:italic toggle=\"yes\">AniGen<\/jats:italic>\n                    , a unified framework that directly generates animate-ready 3D assets conditioned on a single image. Our key insight is to represent shape, skeleton, and skinning as mutually consistent\n                    <jats:italic toggle=\"yes\">S<\/jats:italic>\n                    <jats:sup>3<\/jats:sup>\n                    <jats:italic toggle=\"yes\">Fields<\/jats:italic>\n                    (Shape, Skeleton, Skin) defined over a shared spatial domain. To enable the robust learning of these fields, we introduce two technical innovations: (i) a\n                    <jats:italic toggle=\"yes\">confidence-decaying skeleton field<\/jats:italic>\n                    that explicitly handles the geometric ambiguity of bone prediction at Voronoi boundaries, and (ii) a\n                    <jats:italic toggle=\"yes\">dual skin feature field<\/jats:italic>\n                    that decouples skinning weights from specific joint counts, allowing a fixed-architecture network to predict rigs of arbitrary complexity. Built upon a two-stage flow-matching pipeline,\n                    <jats:italic toggle=\"yes\">AniGen<\/jats:italic>\n                    first synthesizes a sparse structural scaffold and then generates dense geometry and articulation in a structured latent space. Extensive experiments demonstrate that\n                    <jats:italic toggle=\"yes\">AniGen<\/jats:italic>\n                    substantially outperforms state-of-the-art sequential baselines in rig validity and animation quality, generalizing effectively to in-the-wild images across diverse categories including animals, humanoids, and machinery.\n                  <\/jats:p>\n                  <jats:p>\n                    <jats:bold>Homepage<\/jats:bold>\n                    : https:\/\/yihua7.github.io\/AniGen_web\/\n                  <\/jats:p>","DOI":"10.1145\/3811297","type":"journal-article","created":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:05:51Z","timestamp":1783062351000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["AniGen: Unified S\n                    <sup>3<\/sup>\n                    Fields for Animatable 3D Asset Generation"],"prefix":"10.1145","volume":"45","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2208-8280","authenticated-orcid":false,"given":"Yi-Hua","family":"Huang","sequence":"first","affiliation":[{"name":"The University of Hong Kong (HKU), Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2945-552X","authenticated-orcid":false,"given":"Zi-Xin","family":"Zou","sequence":"additional","affiliation":[{"name":"VAST, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-1954-952X","authenticated-orcid":false,"given":"Yuting","family":"He","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-0628-2851","authenticated-orcid":false,"given":"Chirui","family":"Chang","sequence":"additional","affiliation":[{"name":"The University of Hong Kong (HKU), Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-4483-2751","authenticated-orcid":false,"given":"Cheng-Feng","family":"Pu","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9318-4704","authenticated-orcid":false,"given":"Ziyi","family":"Yang","sequence":"additional","affiliation":[{"name":"The University of Hong Kong (HKU), Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6164-8343","authenticated-orcid":false,"given":"Yuan-Chen","family":"Guo","sequence":"additional","affiliation":[{"name":"VAST, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0416-4374","authenticated-orcid":false,"given":"Yan-Pei","family":"Cao","sequence":"additional","affiliation":[{"name":"VAST, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4285-1626","authenticated-orcid":false,"given":"Xiaojuan","family":"Qi","sequence":"additional","affiliation":[{"name":"The University of Hong Kong (HKU), Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,3]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"crossref","unstructured":"J. A. B\u00e6rentzen R. Abdrashitov and K. Singh. 2014. Interactive Shape Modeling using a Skeleton-Mesh Co-Representation. ACM Transactions on Graphics (proceedings of ACM SIGGRAPH) 33 4 (2014).","DOI":"10.1145\/2601097.2601226"},{"key":"e_1_2_2_2_1","volume-title":"Automatic rigging and animation of 3d characters. ACM Transactions on graphics (TOG) 26, 3","author":"Baran Ilya","year":"2007","unstructured":"Ilya Baran and Jovan Popovi\u0107. 2007. Automatic rigging and animation of 3d characters. ACM Transactions on graphics (TOG) 26, 3 (2007), 72\u2013es."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2366145.2366217"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3757377.3763922"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01515"},{"key":"e_1_2_2_6_1","volume-title":"Ultra3d: Efficient and high-fidelity 3d generation with part attention. arXiv preprint arXiv:2507.17745","author":"Chen Yiwen","year":"2025","unstructured":"Yiwen Chen, Zhihao Li, Yikai Wang, Hu Zhang, Qin Li, Chi Zhang, and Guosheng Lin. 2025b. Ultra3d: Efficient and high-fidelity 3d generation with part attention. arXiv preprint arXiv:2507.17745 (2025)."},{"key":"e_1_2_2_7_1","doi-asserted-by":"crossref","unstructured":"Peng Dai Feitong Tan Qiangeng Xu David Futschik Ruofei Du Sean Fanello XIAOJUAN QI and Yinda Zhang. 2025. SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix. In ICLR.","DOI":"10.1109\/TPAMI.2026.3692948"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-1554"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01263"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3721238.3730743"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2403"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3721238.3730621"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1307\/mmj\/1029003026"},{"key":"e_1_2_2_14_1","volume-title":"Make-It-Poseable: Feed-forward Latent Posing Model for 3D Humanoid Character Animation. arXiv preprint arXiv:2512.16767","author":"Guo Zhiyang","year":"2025","unstructured":"Zhiyang Guo, Ori Zhang, Jax Xiang, Alan Zhao, Wengang Zhou, and Houqiang Li. 2025. Make-It-Poseable: Feed-forward Latent Posing Model for 3D Humanoid Character Animation. arXiv preprint arXiv:2512.16767 (2025)."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01375"},{"key":"e_1_2_2_16_1","volume-title":"Denoising diffusion probabilistic models. Advances in neural information processing systems 33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840\u20136851."},{"key":"e_1_2_2_17_1","volume-title":"The Twelfth International Conference on Learning Representations.","author":"Hong Yicong","year":"2023","unstructured":"Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. 2023. LRM: Large Reconstruction Model for Single Image to 3D. In The Twelfth International Conference on Learning Representations."},{"key":"e_1_2_2_18_1","volume-title":"CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling. arXiv preprint arXiv:2510.20776","author":"Huang Binbin","year":"2025","unstructured":"Binbin Huang, Haobin Duan, Yiqun Zhao, Zibo Zhao, Yi Ma, and Shenghua Gao. 2025a. CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling. arXiv preprint arXiv:2510.20776 (2025)."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00404"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3757377.3763885"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00094"},{"key":"e_1_2_2_22_1","volume-title":"What uncertainties do we need in bayesian deep learning for computer vision? Advances in neural information processing systems 30","author":"Kendall Alex","year":"2017","unstructured":"Alex Kendall and Yarin Gal. 2017. What uncertainties do we need in bayesian deep learning for computer vision? Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3592433"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3757377.3763937"},{"key":"e_1_2_2_25_1","unstructured":"Diederik P Kingma and Max Welling. 2014. Auto-encoding variational bayes. In ICLR."},{"key":"e_1_2_2_26_1","unstructured":"Jiahao Li Hao Tan Kai Zhang Zexiang Xu Fujun Luan Yinghao Xu Yicong Hong Kalyan Sunkavalli Greg Shakhnarovich and Sai Bi. 2024. Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction Model. In ICLR."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459852"},{"key":"e_1_2_2_28_1","volume-title":"Particulate: Feed-Forward 3D Object Articulation. arXiv preprint arXiv:2512.11798","author":"Li Ruining","year":"2025","unstructured":"Ruining Li, Yuxin Yao, Chuanxia Zheng, Christian Rupprecht, Joan Lasenby, Shangzhe Wu, and Andrea Vedaldi. 2025a. Particulate: Feed-Forward 3D Object Articulation. arXiv preprint arXiv:2512.11798 (2025)."},{"key":"e_1_2_2_29_1","volume-title":"Triposg: High-fidelity 3d shape synthesis using large-scale rectified flow models. arXiv preprint arXiv:2502.06608","author":"Li Yangguang","year":"2025","unstructured":"Yangguang Li, Zi-Xin Zou, Zexiang Liu, Dehu Wang, Yuan Liang, Zhipeng Yu, Xingchao Liu, Yuan-Chen Guo, Ding Liang, Wanli Ouyang, et al. 2025c. Triposg: High-fidelity 3d shape synthesis using large-scale rectified flow models. arXiv preprint arXiv:2502.06608 (2025)."},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3763302"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3731149"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3355089.3356495"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00853"},{"key":"e_1_2_2_34_1","volume-title":"Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003","author":"Liu Xingchao","year":"2022","unstructured":"Xingchao Liu, Chengyue Gong, and Qiang Liu. 2022. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003 (2022)."},{"key":"e_1_2_2_35_1","unstructured":"Yuan Liu Cheng Lin Zijiao Zeng Xiaoxiao Long Lingjie Liu Taku Komura and Wenping Wang. 2024. SyncDreamer: Generating Multiview-consistent Images from a Single-view Image. In ICLR."},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00951"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cag.2023.05.018"},{"key":"e_1_2_2_39_1","series-title":"Series B. Biological Sciences 200, 1140","volume-title":"Representation and recognition of the spatial organization of three-dimensional shapes. Proceedings of the Royal Society of London","author":"Marr David","year":"1978","unstructured":"David Marr and Herbert Keith Nishihara. 1978. Representation and recognition of the spatial organization of three-dimensional shapes. Proceedings of the Royal Society of London. Series B. Biological Sciences 200, 1140 (1978), 269\u2013294."},{"key":"e_1_2_2_40_1","volume-title":"Gromov-Wasserstein distances and the metric approach to object matching. Foundations of computational mathematics 11, 4","author":"M\u00e9moli Facundo","year":"2011","unstructured":"Facundo M\u00e9moli. 2011. Gromov-Wasserstein distances and the metric approach to object matching. Foundations of computational mathematics 11, 4 (2011), 417\u2013487."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503250"},{"key":"e_1_2_2_42_1","unstructured":"Maxime Oquab Timoth\u00e9e Darcet Th\u00e9o Moutakanni Huy Vo Marc Szafraniec Vasil Khalidov Pierre Fernandez Daniel Haziza Francisco Massa Alaaeldin El-Nouby et al. 2024. DINOv2: Learning Robust Visual Features without Supervision. Transactions on Machine Learning Research Journal (2024)."},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528233.3530754"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"e_1_2_2_45_1","unstructured":"Ben Poole Ajay Jain Jonathan T Barron and Ben Mildenhall. 2023. DreamFusion: Text-to-3D using 2D Diffusion. In ICLR."},{"key":"e_1_2_2_46_1","volume-title":"International conference on machine learning. PmLR, 8748\u20138763","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PmLR, 8748\u20138763."},{"key":"e_1_2_2_47_1","volume-title":"Dreamgaussian4d: Generative 4d gaussian splatting. arXiv preprint arXiv:2312.17142","author":"Ren Jiawei","year":"2023","unstructured":"Jiawei Ren, Liang Pan, Jiaxiang Tang, Chi Zhang, Ang Cao, Gang Zeng, and Ziwei Liu. 2023. Dreamgaussian4d: Generative 4d gaussian splatting. arXiv preprint arXiv:2312.17142 (2023)."},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3592430"},{"key":"e_1_2_2_49_1","unstructured":"Yichun Shi Peng Wang Jianglong Ye Long Mai Kejie Li and Xiao Yang. 2023. MV-Dream: Multi-view Diffusion for 3D Generation. In ICLR."},{"key":"e_1_2_2_50_1","volume-title":"The Thirty-ninth Annual Conference on Neural Information Processing Systems.","author":"Song Chaoyue","year":"2025","unstructured":"Chaoyue Song, Xiu Li, Fan Yang, Zhongcong Xu, Jiacheng Wei, Fayao Liu, Jiashi Feng, Guosheng Lin, and Jianfeng Zhang. 2025a. Puppeteer: Rig and Animate Your 3D Models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems."},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01491"},{"key":"e_1_2_2_52_1","unstructured":"Jiaming Song Chenlin Meng and Stefano Ermon. 2021. Denoising Diffusion Implicit Models. In ICLR."},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01972"},{"key":"e_1_2_2_54_1","volume-title":"European Conference on Computer Vision. Springer, 1\u201318","author":"Tang Jiaxiang","year":"2024","unstructured":"Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. 2024. Lgm: Large multi-view gaussian model for high-resolution 3d content creation. In European Conference on Computer Vision. Springer, 1\u201318."},{"key":"e_1_2_2_55_1","volume-title":"TripoSR: Fast 3D Object Reconstruction from a Single Image. arXiv preprint arXiv:2403.02151","author":"Tochilkin Dmitry","year":"2024","unstructured":"Dmitry Tochilkin, David Pankratz, Zexiang Liu, Zixuan Huang,, Adam Letts, Yangguang Li, Ding Liang, Christian Laforte, Varun Jampani, and Yan-Pei Cao. 2024. TripoSR: Fast 3D Object Reconstruction from a Single Image. arXiv preprint arXiv:2403.02151 (2024)."},{"key":"e_1_2_2_56_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_57_1","volume-title":"Hmc: Hierarchical mesh coarsening for skeleton-free motion retargeting. arXiv preprint arXiv:2303.10941","author":"Wang Haoyu","year":"2023","unstructured":"Haoyu Wang, Shaoli Huang, Fang Zhao, Chun Yuan, and Ying Shan. 2023a. Hmc: Hierarchical mesh coarsening for skeleton-free motion retargeting. arXiv preprint arXiv:2303.10941 (2023)."},{"key":"e_1_2_2_58_1","volume-title":"Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation. Advances in neural information processing systems 36","author":"Wang Zhengyi","year":"2023","unstructured":"Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. 2023b. Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation. Advances in neural information processing systems 36 (2023), 8406\u20138441."},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02427"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3873"},{"key":"e_1_2_2_61_1","volume-title":"European Conference on Computer Vision. Springer, 361\u2013379","author":"Wu Zijie","year":"2024","unstructured":"Zijie Wu, Chaohui Yu, Yanqin Jiang, Chenjie Cao, Fan Wang, and Xiang Bai. 2024b. Sc4d: Sparse-controlled video-to-4d generation and motion transfer. In European Conference on Computer Vision. Springer, 361\u2013379."},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01259"},{"key":"e_1_2_2_63_1","volume-title":"Nicholas Jing Yuan, et al","author":"Xiang Jianfeng","year":"2025","unstructured":"Jianfeng Xiang, Xiaoxue Chen, Sicheng Xu, Ruicheng Wang, Zelong Lv, Yu Deng, Hongyuan Zhu, Yue Dong, Hao Zhao, Nicholas Jing Yuan, et al. 2025a. Native and Compact Structured Latents for 3D Generation. arXiv preprint arXiv:2512.14692 (2025)."},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02000"},{"key":"e_1_2_2_65_1","volume-title":"AnimaMimic: Imitating 3D Animation from Video Priors. arXiv preprint arXiv:2512.14133","author":"Xie Tianyi","year":"2025","unstructured":"Tianyi Xie, Yunuo Chen, Yaowei Guo, Yin Yang, Bolei Zhou, Demetri Terzopoulos, Ying Jiang, and Chenfanfu Jiang. 2025. AnimaMimic: Imitating 3D Animation from Video Priors. arXiv preprint arXiv:2512.14133 (2025)."},{"key":"e_1_2_2_66_1","unstructured":"Yinghao Xu Hao Tan Fujun Luan Sai Bi Peng Wang Jiahao Li Zifan Shi Kalyan Sunkavalli Gordon Wetzstein Zexiang Xu et al. 2024. DMV3D: Denoising Multiview Diffusion Using 3D Large Reconstruction Model. In ICLR."},{"key":"e_1_2_2_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392379"},{"key":"e_1_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2019.00041"},{"key":"e_1_2_2_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3550469.3555390"},{"key":"e_1_2_2_70_1","volume-title":"Latent denoising makes good visual tokenizers. arXiv preprint arXiv:2507.15856","author":"Yang Jiawei","year":"2025","unstructured":"Jiawei Yang, Tianhong Li, Lijie Fan, Yonglong Tian, and Yue Wang. 2025. Latent denoising makes good visual tokenizers. arXiv preprint arXiv:2507.15856 (2025)."},{"key":"e_1_2_2_71_1","volume-title":"Towards Scalable Pre-training of Visual Tokenizers for Generation. arXiv preprint arXiv:2512.13687","author":"Yao Jingfeng","year":"2025","unstructured":"Jingfeng Yao, Yuda Song, Yucong Zhou, and Xinggang Wang. 2025. Towards Scalable Pre-training of Visual Tokenizers for Generation. arXiv preprint arXiv:2512.13687 (2025)."},{"key":"e_1_2_2_72_1","unstructured":"Xin Yu Yuan-Chen Guo Yangguang Li Ding Liang Song-Hai Zhang and Xiaojuan Qi. 2023. Text-to-3D with Classifier Score Distillation. In ICLR."},{"key":"e_1_2_2_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/3618342"},{"key":"e_1_2_2_74_1","doi-asserted-by":"publisher","DOI":"10.1145\/3730930"},{"key":"e_1_2_2_75_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2024.3423426"},{"key":"e_1_2_2_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/3658146"},{"key":"e_1_2_2_77_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00983"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","deposited":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:48:27Z","timestamp":1783064907000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811297"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,3]]},"references-count":77,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,7,3]]}},"alternative-id":["10.1145\/3811297"],"URL":"https:\/\/doi.org\/10.1145\/3811297","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,3]]},"assertion":[{"value":"2026-01-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}