{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T06:04:14Z","timestamp":1784268254551,"version":"3.55.0"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2024,11,19]],"date-time":"2024-11-19T00:00:00Z","timestamp":1731974400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2024,12,19]]},"abstract":"<jats:p>\n            Neural implicit functions have brought impressive advances to the state-of-the-art of clothed human digitization from multiple or even single images. However, despite the progress, current arts still have difficulty generalizing to unseen images with complex cloth deformation and body poses. In this work, we present GarVerseLOD, a new dataset and framework that paves the way to achieving unprecedented robustness in high-fidelity 3D garment reconstruction from a single unconstrained image. Inspired by the recent success of large generative models, we believe that one key to addressing the generalization challenge lies in the quantity and quality of 3D garment data. Towards this end, GarVerseLOD collects 6,000 high-quality cloth models with fine-grained geometry details manually created by professional artists. In addition to the scale of training data, we observe that having disentangled granularities of geometry can play an important role in boosting the generalization capability and inference accuracy of the learned model. We hence craft GarVerseLOD as a hierarchical dataset with\n            <jats:italic>levels of details (LOD)<\/jats:italic>\n            , spanning from detail-free stylized shape to pose-blended garment with pixel-aligned details. This allows us to make this highly under-constrained problem tractable by factorizing the inference into easier tasks, each narrowed down with smaller searching space. To ensure GarVerseLOD can generalize well to in-the-wild images, we propose a novel labeling paradigm based on conditional diffusion models to generate extensive paired images for each garment model with high photorealism. We evaluate our method on a massive amount of in-the-wild images. Experimental results demonstrate that GarVerseLOD can generate standalone garment pieces with significantly better quality than prior approaches while being robust against a large variation of pose, illumination, occlusion, and deformation. Code and dataset are available at garverselod.github.io.\n          <\/jats:p>","DOI":"10.1145\/3687921","type":"journal-article","created":{"date-parts":[[2024,11,19]],"date-time":"2024-11-19T15:46:04Z","timestamp":1732031164000},"page":"1-12","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["GarVerseLOD: High-Fidelity 3D Garment Reconstruction from a Single In-the-Wild Image using a Dataset with Levels of Details"],"prefix":"10.1145","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3483-4236","authenticated-orcid":false,"given":"Zhongjin","family":"Luo","sequence":"first","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-4962-9217","authenticated-orcid":false,"given":"Haolin","family":"Liu","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-0604-7421","authenticated-orcid":false,"given":"Chenghong","family":"Li","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-3005-5007","authenticated-orcid":false,"given":"Wanghao","family":"Du","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-0558-8781","authenticated-orcid":false,"given":"Zirong","family":"Jin","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-5740-9997","authenticated-orcid":false,"given":"Wanhu","family":"Sun","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7023-6797","authenticated-orcid":false,"given":"Yinyu","family":"Nie","sequence":"additional","affiliation":[{"name":"Huawei Technologies Ltd., London, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3212-1072","authenticated-orcid":false,"given":"Weikai","family":"Chen","sequence":"additional","affiliation":[{"name":"Tencent America, Los Angeles, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0162-3296","authenticated-orcid":false,"given":"Xiaoguang","family":"Han","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,11,19]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00238"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2007.383165"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1186822.1073207"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58565-5_21"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00552"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00843"},{"key":"e_1_2_1_7_1","volume-title":"Neural-ABC: Neural Parametric Models for Articulated Body With Clothes","author":"Chen Honghu","year":"2024","unstructured":"Honghu Chen, Yuxin Yao, and Juyong Zhang. 2024. Neural-ABC: Neural Parametric Models for Articulated Body With Clothes. IEEE Transactions on Visualization and Computer Graphics (2024)."},{"key":"e_1_2_1_8_1","unstructured":"Paolo Cignoni Marco Callieri Massimiliano Corsini Matteo Dellepiane Fabio Ganovelli Guido Ranzuglia et al. 2008. Meshlab: an open-source mesh processing tool.. In Eurographics Italian chapter conference Vol. 2008. Salerno Italy 129--136."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01170"},{"key":"e_1_2_1_10_1","volume-title":"MeshUDF: Fast and Differentiable Meshing of Unsigned Distance Field Networks. In European Conference on Computer Vision.","author":"Guillard Benoit","year":"2022","unstructured":"Benoit Guillard, Federico Stella, and Pascal Fua. 2022. MeshUDF: Fast and Differentiable Meshing of Unsigned Distance Field Networks. In European Conference on Computer Vision."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00883"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3311970"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00510"},{"key":"e_1_2_1_14_1","volume-title":"Computer graphics forum","author":"Hasler Nils","unstructured":"Nils Hasler, Carsten Stoll, Martin Sunkel, Bodo Rosenhahn, and H-P Seidel. 2009. A statistical model of human pose and body shape. In Computer graphics forum, Vol. 28. Wiley Online Library, 337--346."},{"key":"e_1_2_1_15_1","volume-title":"Proceedings, Part XX 16","author":"Jiang Boyi","year":"2020","unstructured":"Boyi Jiang, Juyong Zhang, Yang Hong, Jinhao Luo, Ligang Liu, and Hujun Bao. 2020. Bcnet: Learning body and cloth shape from a single image. In Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XX 16. Springer, 18--35."},{"key":"e_1_2_1_16_1","volume-title":"Hifecap: Monocular high-fidelity and expressive capture of human performances. arXiv preprint arXiv:2210.05665","author":"Jiang Yue","year":"2022","unstructured":"Yue Jiang, Marc Habermann, Vladislav Golyanik, and Christian Theobalt. 2022. Hifecap: Monocular high-fidelity and expressive capture of human performances. arXiv preprint arXiv:2210.05665 (2022)."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.500"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 1586--1595","author":"Li Ren","year":"2024","unstructured":"Ren Li, Corentin Dumery, Beno\u00eet Guillard, and Pascal Fua. 2024a. Garment Recovery with Shape and Deformation Priors. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 1586--1595."},{"key":"e_1_2_1_19_1","volume-title":"Isp: Multi-layered garment draping with implicit sewing patterns. Advances in Neural Information Processing Systems 36","author":"Li Ren","year":"2024","unstructured":"Ren Li, Beno\u00eet Guillard, and Pascal Fua. 2024b. Isp: Multi-layered garment draping with implicit sewing patterns. Advances in Neural Information Processing Systems 36 (2024)."},{"key":"e_1_2_1_20_1","volume-title":"Proceedings, Part XXIII 16","author":"Li Ruilong","year":"2020","unstructured":"Ruilong Li, Yuliang Xiu, Shunsuke Saito, Zeng Huang, Kyle Olszewski, and Hao Li. 2020. Monocular real-time volumetric performance capture. In Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XXIII 16. Springer, 49--67."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV53792.2021.00047"},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision. 14485--14496","author":"Lin Siyou","year":"2023","unstructured":"Siyou Lin, Boyao Zhou, Zerong Zheng, Hongwen Zhang, and Yebin Liu. 2023. Leveraging intrinsic properties for non-rigid garment alignment. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 14485--14496."},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 20454--20464","author":"Liu Haolin","year":"2024","unstructured":"Haolin Liu, Chongjie Ye, Yinyu Nie, Yingfan He, and Xiaoguang Han. 2024. LASA: Instance Reconstruction from Real Scans using A Large-scale Aligned Shape Annotation Dataset. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 20454--20464."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818013"},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 12825--12835","author":"Luo Zhongjin","year":"2023","unstructured":"Zhongjin Luo, Shengcai Cai, Jinguo Dong, Ruibo Ming, Liangdong Qiu, Xiaohang Zhan, and Xiaoguang Han. 2023. RaBit: Parametric Modeling of 3D Biped Cartoon Characters with a Topological-consistent Dataset. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 12825--12835."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3472749.3474791"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00459"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20086-1_11"},{"key":"e_1_2_1_29_1","volume-title":"T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. arXiv preprint arXiv:2302.08453","author":"Mou Chong","year":"2023","unstructured":"Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, Ying Shan, and Xiaohu Qie. 2023. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. arXiv preprint arXiv:2302.08453 (2023)."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00461"},{"key":"e_1_2_1_31_1","volume-title":"Shape and Garment Style. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE.","author":"Patel Chaitanya","year":"2020","unstructured":"Chaitanya Patel, Zhouyingcheng Liao, and Gerard Pons-Moll. 2020. TailorNet: Predicting Clothing in 3D as a Function of Human Pose, Shape and Garment Style. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01123"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01405"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073711"},{"key":"e_1_2_1_35_1","volume-title":"Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988","author":"Poole Ben","year":"2022","unstructured":"Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2022. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988 (2022)."},{"key":"e_1_2_1_36_1","volume-title":"Accelerating 3d deep learning with pytorch3d. arXiv preprint arXiv:2007.08501","author":"Ravi Nikhila","year":"2020","unstructured":"Nikhila Ravi, Jeremy Reizenstein, David Novotny, Taylor Gordon, Wan-Yen Lo, Justin Johnson, and Georgia Gkioxari. 2020. Accelerating 3d deep learning with pytorch3d. arXiv preprint arXiv:2007.08501 (2020)."},{"key":"e_1_2_1_37_1","unstructured":"RenderPeople. 2018. In https:\/\/renderpeople.com\/3d-people."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00239"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00016"},{"key":"e_1_2_1_41_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 650--659","author":"Tan Feitong","year":"2020","unstructured":"Feitong Tan, Hao Zhu, Zhaopeng Cui, Siyu Zhu, Marc Pollefeys, and Ping Tan. 2020. Self-supervised human depth estimation from monocular videos. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 650--659."},{"key":"e_1_2_1_42_1","volume-title":"Proceedings, Part III 16","author":"Tiwari Garvita","year":"2020","unstructured":"Garvita Tiwari, Bharat Lal Bhatnagar, Tony Tung, and Gerard Pons-Moll. 2020. Sizer: A dataset and model for parsing 3d clothing and learning size sensitive 3d clothing. In Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part III 16. Springer, 1--18."},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 550--560","author":"Wang Wenbo","year":"2024","unstructured":"Wenbo Wang, Hsuan-I Ho, Chen Guo, Boxiang Rong, Artur Grigorev, Jie Song, Juan Jose Zarate, and Otmar Hilliges. 2024. 4D-DRESS: A 4D Dataset of Real-World Human Clothing With Semantic Annotations. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 550--560."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV50981.2020.00042"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00057"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01294"},{"key":"e_1_2_1_47_1","first-page":"1","article-title":"Monoperfcap: Human performance capture from monocular video","volume":"37","author":"Xu Weipeng","year":"2018","unstructured":"Weipeng Xu, Avishek Chatterjee, Michael Zollh\u00f6fer, Helge Rhodin, Dushyant Mehta, Hans-Peter Seidel, and Christian Theobalt. 2018. Monoperfcap: Human performance capture from monocular video. ACM Transactions on Graphics (ToG) 37, 2 (2018), 1--15.","journal-title":"ACM Transactions on Graphics (ToG)"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00785"},{"key":"e_1_2_1_49_1","volume-title":"DreamDissector: Learning Disentangled Text-to-3D Generation from 2D Diffusion Priors. arXiv preprint arXiv:2407.16260","author":"Yan Zizheng","year":"2024","unstructured":"Zizheng Yan, Jiapeng Zhou, Fanpeng Meng, Yushuang Wu, Lingteng Qiu, Zisheng Ye, Shuguang Cui, Guanying Chen, and Xiaoguang Han. 2024. DreamDissector: Learning Disentangled Text-to-3D Generation from 2D Diffusion Priors. arXiv preprint arXiv:2407.16260 (2024)."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3026479"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.582"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01125"},{"key":"e_1_2_1_53_1","doi-asserted-by":"crossref","unstructured":"Lvmin Zhang Anyi Rao and Maneesh Agrawala. 2023. Adding Conditional Control to Text-to-Image Diffusion Models.","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00783"},{"key":"e_1_2_1_55_1","volume-title":"Proceedings, Part I 16","author":"Zhu Heming","year":"2020","unstructured":"Heming Zhu, Yu Cao, Hang Jin, Weikai Chen, Dong Du, Zhangye Wang, Shuguang Cui, and Xiaoguang Han. 2020. Deep fashion3d: A dataset and benchmark for 3d garment reconstruction from single images. In Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part I 16. Springer, 512--530."},{"key":"e_1_2_1_56_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3845--3854","author":"Zhu Heming","year":"2022","unstructured":"Heming Zhu, Lingteng Qiu, Yuda Qiu, and Xiaoguang Han. 2022. Registering explicit to implicit: Towards high-fidelity garment mesh reconstruction from single images. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3845--3854."},{"key":"e_1_2_1_57_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 12847--12857","author":"Zou Xingxing","year":"2023","unstructured":"Xingxing Zou, Xintong Han, and Waikeung Wong. 2023. CLOTH4D: A Dataset for Clothed Human Reconstruction. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 12847--12857."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3687921","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3687921","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:09:57Z","timestamp":1750295397000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3687921"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,19]]},"references-count":57,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2024,12,19]]}},"alternative-id":["10.1145\/3687921"],"URL":"https:\/\/doi.org\/10.1145\/3687921","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,19]]},"assertion":[{"value":"2024-11-19","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}