{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T02:28:01Z","timestamp":1777516081056,"version":"3.51.4"},"reference-count":50,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2025,3,10]],"date-time":"2025-03-10T00:00:00Z","timestamp":1741564800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62372441, U22A2034"],"award-info":[{"award-number":["62372441, U22A2034"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100021171","name":"Guangdong Basic and Applied Basic Research Foundation","doi-asserted-by":"crossref","award":["2023A1515030268"],"award-info":[{"award-number":["2023A1515030268"]}],"id":[{"id":"10.13039\/501100021171","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Shenzhen Science and Technology Program","award":["RCYX20231211090127030, JSGG20220831105002004"],"award-info":[{"award-number":["RCYX20231211090127030, JSGG20220831105002004"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>Clothed human modeling plays a crucial role in multimedia research, with applications spanning virtual reality, gaming, and fashion design. The goal is to learn clothed human dynamics from observations and then generate humans with high-fidelity clothing details for motion animation. Despite tremendous advancements in clothing shape analysis by existing approaches, the community still faces challenges in generating convincing visual effects of cloth dynamics, maintaining temporally smooth clothing details, and handling diverse clothing patterns. To address these challenges, we introduce ClothDiffuse, a temporal diffusion model that seamlessly integrates three key components into this task\u2014temporal dynamics modeling, iterative refinement, and diversified generation. Our approach begins by using an encoder to extract high-level temporal features from input human body motions. These features are combined with a learnable pixel-aligned garment feature, serving as prior conditions for the shape decoder. The decoder then iteratively denoise Gaussian noise to produce clothing deformations over time on the input unclothed human bodies. To ensure that the results align with observations and adhere to physical plausibility for clothing shape inference, we propose two physics-inspired loss functions that preserve the intra-frame distances and inter-frame forces of clothing points. Additionally, the stochastic nature of the denoising process allows for the generation of diverse and plausible clothing shapes. Experiments show that our approach outperforms state-of-the-art methods in chamfer distance and visual effects, particularly for loose clothing such as dresses and skirts. Furthermore, our approach effectively adapts to out-of-domain clothing types and generates realistic clothes dynamics.<\/jats:p>","DOI":"10.1145\/3712011","type":"journal-article","created":{"date-parts":[[2025,1,13]],"date-time":"2025-01-13T15:37:23Z","timestamp":1736782643000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Generating High-Fidelity Clothed Human Dynamics with Temporal Diffusion"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9042-6069","authenticated-orcid":false,"given":"Shihao","family":"Zou","sequence":"first","affiliation":[{"name":"Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7095-1018","authenticated-orcid":false,"given":"Yuanlu","family":"Xu","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Redmond, Washington, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6350-3890","authenticated-orcid":false,"given":"Nikolaos","family":"Sarafianos","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Redmond, Washington, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-4991-185X","authenticated-orcid":false,"given":"Federica","family":"Bogo","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Redmond, Washington, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0824-0960","authenticated-orcid":false,"given":"Tony","family":"Tung","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Redmond, Washington, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3289-9714","authenticated-orcid":false,"given":"Weixin","family":"Si","sequence":"additional","affiliation":[{"name":"Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, Guangdong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3261-3533","authenticated-orcid":false,"given":"Li","family":"Cheng","sequence":"additional","affiliation":[{"name":"University of Alberta, Edmonton, Alberta, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,3,10]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00477"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3478513.3480479"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3550454.3555491"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01058"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00790"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01978"},{"key":"e_1_3_3_8_2","first-page":"1","article-title":"Full-body human motion reconstruction with sparse joint tracking using flexible sensors","volume":"20","author":"Chen Xiaowei","year":"2023","unstructured":"Xiaowei Chen, Xiao Jiang, Lishuang Zhan, Shihui Guo, Qunsheng Ruan, Guoliang Luo, Minghong Liao, and Yipeng Qin. 2023. Full-body human motion reconstruction with sparse joint tracking using flexible sensors. ACM Trans. Multimedia Comput. Commun. Appl. 20 (2023), 1\u201319.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01139"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00609"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01170"},{"key":"e_1_3_3_12_2","article-title":"Learning elementary structures for 3d shape generation and matching","author":"Deprelle Theo","year":"2019","unstructured":"Theo Deprelle, Thibault Groueix, Matthew Fisher, Vladimir Kim, Bryan Russell, and Mathieu Aubry. 2019. Learning elementary structures for 3d shape generation and matching. In NeurIPS.","journal-title":"NeurIPS"},{"key":"e_1_3_3_13_2","article-title":"PINA: Learning a personalized implicit neural avatar from a single RGB-D video sequence","author":"Dong Zijian","year":"2022","unstructured":"Zijian Dong, Chen Guo, Jie Song, Xu Chen, Andreas Geiger, and Otmar Hilliges. 2022. PINA: Learning a personalized implicit neural avatar from a single RGB-D video sequence. In CVPR.","journal-title":"CVPR"},{"key":"e_1_3_3_14_2","article-title":"A point set generation network for 3d object reconstruction from a single image","author":"Fan Haoqiang","year":"2017","unstructured":"Haoqiang Fan, Hao Su, and Leonidas J. Guibas. 2017. A point set generation network for 3d object reconstruction from a single image. In CVPR.","journal-title":"CVPR"},{"key":"e_1_3_3_15_2","article-title":"Hood: Hierarchical graphs for generalized modelling of clothing dynamics","author":"Grigorev Artur","year":"2023","unstructured":"Artur Grigorev, Michael J. Black, and Otmar Hilliges. 2023. Hood: Hierarchical graphs for generalized modelling of clothing dynamics. In CVPR.","journal-title":"CVPR"},{"key":"e_1_3_3_16_2","article-title":"Denoising diffusion probabilistic models","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. In NeurIPS.","journal-title":"NeurIPS"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.5555\/3586589.3586636"},{"key":"e_1_3_3_18_2","article-title":"BodyMap: Learning full-body dense correspondence map","author":"Ianina Anastasia","year":"2022","unstructured":"Anastasia Ianina, Nikolaos Sarafianos, Yuanlu Xu, Ignacio Rocco, and Tony Tung. 2022. BodyMap: Learning full-body dense correspondence map. In CVPR.","journal-title":"CVPR"},{"key":"e_1_3_3_19_2","article-title":"LoRD: Local 4D implicit representation for high-fidelity dynamic human modeling","author":"Jiang Boyan","year":"2022","unstructured":"Boyan Jiang, Xinlin Ren, Mingsong Dou, Xiangyang Xue, Yanwei Fu, and Yinda Zhang. 2022. LoRD: Local 4D implicit representation for high-fidelity dynamic human modeling. In ECCV.","journal-title":"ECCV"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/1275808.1276457"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3550454.35555"},{"key":"e_1_3_3_22_2","article-title":"Learning implicit templates for point-based clothed human modeling","author":"Lin Siyou","year":"2022","unstructured":"Siyou Lin, Hongwen Zhang, Zerong Zheng, Ruizhi Shao, and Yebin Liu. 2022. Learning implicit templates for point-based clothed human modeling. In ECCV.","journal-title":"ECCV"},{"key":"e_1_3_3_23_2","first-page":"1","article-title":"Neuroskinning: Automatic skin binding for production characters with deep graph networks","volume":"38","author":"Liu Lijuan","year":"2019","unstructured":"Lijuan Liu, Youyi Zheng, Di Tang, Yi Yuan, Changjie Fan, and Kun Zhou. 2019. Neuroskinning: Automatic skin binding for production characters with deep graph networks. ACM Trans. Graph. 38 (2019), 1\u201312.","journal-title":"ACM Trans. Graph."},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818013"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01582"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV57658.2022.00078"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00650"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01079"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01032"},{"key":"e_1_3_3_30_2","volume-title":"ECCV","author":"Mildenhall Ben","year":"2020","unstructured":"Ben Mildenhall, P. Pratul, Matthew Srinivasan, Jonathan T. Tancik, Ravi Barron, and Ren Ramamoorthi. 2020. Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. In ECCV."},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01251"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00025"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00739"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01123"},{"key":"e_1_3_3_35_2","article-title":"Dreamfusion: Text-to-3D using 2D diffusion","author":"Poole Ben","year":"2023","unstructured":"Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. 2023. Dreamfusion: Text-to-3D using 2D diffusion. In CVPR.","journal-title":"CVPR"},{"key":"e_1_3_3_36_2","article-title":"Dynamic point fields","author":"Prokudin Sergey","year":"2023","unstructured":"Sergey Prokudin, Qianli Ma, Maxime Raafat, Julien Valentin, and Siyu Tang. 2023. Dynamic point fields. In ICCV.","journal-title":"ICCV"},{"key":"e_1_3_3_37_2","unstructured":"Aditya Ramesh Prafulla Dhariwal Alex Nichol Casey Chu and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv:2204.06125. Retrieved from https:\/\/arxiv.org\/abs\/2204.06125"},{"key":"e_1_3_3_38_2","article-title":"High-resolution image synthesis with latent diffusion models","author":"Rombach Robin","year":"2022","unstructured":"Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj\u00f6rn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In CVPR.","journal-title":"CVPR"},{"key":"e_1_3_3_39_2","article-title":"Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3D human digitization","author":"Saito Shunsuke","year":"2020","unstructured":"Shunsuke Saito, Tomas Simon, Jason Saragih, and Hanbyul Joo. 2020. Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3D human digitization. In CVPR.","journal-title":"CVPR"},{"key":"e_1_3_3_40_2","article-title":"SCANimate: Weakly supervised learning of skinned clothed avatar networks","author":"Saito Shunsuke","year":"2021","unstructured":"Shunsuke Saito, Jinlong Yang, Qianli Ma, and Michael J. Black. 2021. SCANimate: Weakly supervised learning of skinned clothed avatar networks. In CVPR.","journal-title":"CVPR"},{"key":"e_1_3_3_41_2","article-title":"Snug: Self-supervised neural dynamic garments","author":"Santesteban Igor","year":"2022","unstructured":"Igor Santesteban, Miguel A. Otaduy, and Dan Casas. 2022. Snug: Self-supervised neural dynamic garments. In CVPR.","journal-title":"CVPR"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01541"},{"key":"e_1_3_3_43_2","article-title":"Diffustereo: High quality human reconstruction via diffusion-based stereo using sparse cameras","author":"Shao Ruizhi","year":"2022","unstructured":"Ruizhi Shao, Zerong Zheng, Hongwen Zhang, Jingxiang Sun, and Yebin Liu. 2022. Diffustereo: High quality human reconstruction via diffusion-based stereo using sparse cameras. In ECCV.","journal-title":"ECCV"},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3478513.3480545"},{"key":"e_1_3_3_45_2","article-title":"Icon: Implicit clothed humans obtained from normals","author":"Xiu Yuliang","year":"2022","unstructured":"Yuliang Xiu, Jinlong Yang, Dimitrios Tzionas, and Michael J. Black. 2022. Icon: Implicit clothed humans obtained from normals. In CVPR.","journal-title":"CVPR"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3567596"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626235"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3514248"},{"key":"e_1_3_3_49_2","article-title":"LION: Latent point diffusion models for 3D shape generation","author":"Zeng Xiaohui","year":"2022","unstructured":"Xiaohui Zeng, Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, and Karsten Kreis. 2022. LION: Latent point diffusion models for 3D shape generation. In NeurIPS.","journal-title":"NeurIPS"},{"key":"e_1_3_3_50_2","article-title":"CloSET: Modeling clothed humans on continuous surface with explicit template decomposition","author":"Zhang Hongwen","year":"2023","unstructured":"Hongwen Zhang, Siyou Lin, Ruizhi Shao, Yuxiang Zhang, Zerong Zheng, Han Huang, Yandong Guo, and Yebin Liu. 2023. CloSET: Modeling clothed humans on continuous surface with explicit template decomposition. In CVPR.","journal-title":"CVPR"},{"key":"e_1_3_3_51_2","article-title":"Structured local radiance fields for human avatar modeling","author":"Zheng Zerong","year":"2022","unstructured":"Zerong Zheng, Han Huang, Tao Yu, Hongwen Zhang, Yandong Guo, and Yebin Liu. 2022. Structured local radiance fields for human avatar modeling. In CVPR.","journal-title":"CVPR"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3712011","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3712011","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:10Z","timestamp":1750295890000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3712011"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,10]]},"references-count":50,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3712011"],"URL":"https:\/\/doi.org\/10.1145\/3712011","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,10]]},"assertion":[{"value":"2024-08-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-05","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}