{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T18:55:19Z","timestamp":1774637719325,"version":"3.50.1"},"reference-count":77,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2025,8,1]]},"abstract":"<jats:p>Generating photorealistic videos of digital humans in a controllable manner is crucial for a plethora of applications. Existing approaches either build on methods that employ template-based 3D representations or emerging video generation models but suffer from poor quality or limited consistency and identity preservation when generating individual or multiple digital humans. In this paper, we introduce a new interspatial attention (ISA) mechanism as a scalable building block for modern diffusion transformer (DiT)-based video generation models. ISA is a new type of cross attention that uses relative positional encodings tailored for the generation of human videos. Leveraging a custom-developed video variation autoencoder, we train a latent ISA-based diffusion model on a large corpus of video data. Our model achieves state-of-the-art performance for 4D human video synthesis, demonstrating remarkable motion consistency and identity preservation while providing precise control of the camera and body poses. Our code and model are publicly released at https:\/\/dsaurus.github.io\/isa4d\/.<\/jats:p>","DOI":"10.1145\/3731165","type":"journal-article","created":{"date-parts":[[2025,7,27]],"date-time":"2025-07-27T04:02:22Z","timestamp":1753588942000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Interspatial Attention for Efficient 4D Human Video Generation"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2188-1348","authenticated-orcid":false,"given":"Ruizhi","family":"Shao","sequence":"first","affiliation":[{"name":"Tsinghua University, Beijing, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2696-9664","authenticated-orcid":false,"given":"Yinghao","family":"Xu","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3801-6705","authenticated-orcid":false,"given":"Yujun","family":"Shen","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1417-1938","authenticated-orcid":false,"given":"Ceyuan","family":"Yang","sequence":"additional","affiliation":[{"name":"ByteDance Inc., Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6586-7775","authenticated-orcid":false,"given":"Yang","family":"Zheng","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3990-6873","authenticated-orcid":false,"given":"Changan","family":"Chen","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3215-0225","authenticated-orcid":false,"given":"Yebin","family":"Liu","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9243-6885","authenticated-orcid":false,"given":"Gordon","family":"Wetzstein","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,7,27]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00902"},{"key":"e_1_2_2_2_1","unstructured":"Niket Agarwal Arslan Ali Maciej Bala Yogesh Balaji Erik Barker Tiffany Cai Prithvijit Chattopadhyay Yongxin Chen Yin Cui Yifan Ding et al. 2025. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575 (2025)."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1447"},{"key":"e_1_2_2_4_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 5968\u20135976","author":"Bhunia Ankan Kumar","year":"2023","unstructured":"Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Jorma Laaksonen, Mubarak Shah, and Fahad Shahbaz Khan. 2023. Person image synthesis via denoising diffusion model. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 5968\u20135976."},{"key":"e_1_2_2_5_1","unstructured":"Carnegie Mellon University. 2014. CMU MoCap Dataset. http:\/\/mocap.cs.cmu.edu"},{"key":"e_1_2_2_6_1","volume-title":"Scaling Rectified Flow Transformers for High-Resolution Image Synthesis. In Forty-first International Conference on Machine Learning, ICML 2024","author":"Esser Patrick","year":"2024","unstructured":"Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M\u00fcller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rombach. 2024. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21\u201327, 2024. OpenReview.net."},{"key":"e_1_2_2_7_1","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 12873\u201312883","author":"Esser Patrick","year":"2021","unstructured":"Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for highresolution image synthesis. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 12873\u201312883."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19790-1_7"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01358"},{"key":"e_1_2_2_10_1","volume-title":"Generative adversarial nets. Advances in neural information processing systems 27","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. Advances in neural information processing systems 27 (2014)."},{"key":"e_1_2_2_11_1","volume-title":"Photorealistic video generation with diffusion models. arXiv preprint arXiv:2312.06662","author":"Gupta Agrim","year":"2023","unstructured":"Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, and Jos\u00e9 Lezama. 2023. Photorealistic video generation with diffusion models. arXiv preprint arXiv:2312.06662 (2023)."},{"key":"e_1_2_2_12_1","unstructured":"Hao He Yinghao Xu Yuwei Guo Gordon Wetzstein Bo Dai Hongsheng Li and Ceyuan Yang. 2024. CameraCtrl: Enabling Camera Control for Text-to-Video Generation. arXiv:2404.02101 [cs.CV]"},{"key":"e_1_2_2_13_1","volume-title":"Animate anyone: Consistent and controllable image-to-video synthesis for character animation. arXiv preprint arXiv:2311.17117","author":"Hu Li","year":"2023","unstructured":"Li Hu, Xin Gao, Peng Zhang, Ke Sun, Bang Zhang, and Liefeng Bo. 2023a. Animate anyone: Consistent and controllable image-to-video synthesis for character animation. arXiv preprint arXiv:2311.17117 (2023)."},{"key":"e_1_2_2_14_1","volume-title":"Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation. arXiv preprint arXiv:2311.17117","author":"Hu Li","year":"2023","unstructured":"Li Hu, Xin Gao, Peng Zhang, Ke Sun, Bang Zhang, and Liefeng Bo. 2023b. Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation. arXiv preprint arXiv:2311.17117 (2023)."},{"key":"e_1_2_2_15_1","volume-title":"Humannorm: Learning normal diffusion model for high-quality and realistic 3d human generation.","author":"Huang Xin","year":"2024","unstructured":"Xin Huang, Ruizhi Shao, Qi Zhang, Hongwen Zhang, Ying Feng, Yebin Liu, and Qing Wang. 2024b. Humannorm: Learning normal diffusion model for high-quality and realistic 3d human generation."},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02060"},{"key":"e_1_2_2_17_1","first-page":"48955","article-title":"Miradata: A large-scale video dataset with long durations and structured captions","volume":"37","author":"Ju Xuan","year":"2024","unstructured":"Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xintao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. 2024. Miradata: A large-scale video dataset with long durations and structured captions. Advances in Neural Information Processing Systems 37 (2024), 48955\u201348970.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_2_18_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision. 22680\u201322690","author":"Karras Johanna","year":"2023","unstructured":"Johanna Karras, Aleksander Holynski, Ting-Chun Wang, and Ira Kemelmacher-Shlizerman. 2023. Dreampose: Fashion video synthesis with stable diffusion. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 22680\u201322690."},{"key":"e_1_2_2_19_1","unstructured":"Will Kay Joao Carreira Karen Simonyan Brian Zhang Chloe Hillier Sudheendra Vijayanarasimhan Fabio Viola Tim Green Trevor Back Paul Natsev et al. 2017. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950 (2017)."},{"key":"e_1_2_2_20_1","volume-title":"Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114","author":"Kingma Diederik P","year":"2013","unstructured":"Diederik P Kingma. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013)."},{"key":"e_1_2_2_21_1","volume-title":"Mihai Fieraru, and Cristian Sminchisescu.","author":"Kolotouros Nikos","year":"2023","unstructured":"Nikos Kolotouros, Thiemo Alldieck, Andrei Zanfir, Eduard Gabriel Bazavan, Mihai Fieraru, and Cristian Sminchisescu. 2023. DreamHuman: Animatable 3D Avatars from Text. (2023)."},{"key":"e_1_2_2_22_1","volume-title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models. arXiv preprint arXiv:2412.03603","author":"Kong Weijie","year":"2024","unstructured":"Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, Kathrina Wu, Qin Lin, Aladdin Wang, Andong Wang, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Hongmei Wang, Jacob Song, Jiawang Bai, Jianbing Wu, Jinbao Xue, Joey Wang, Junkun Yuan, Kai Wang, Mengyang Liu, Pengyu Li, Shuai Li, Weiyan Wang, Wenqing Yu, Xinchi Deng, Yang Li, Yanxin Long, Yi Chen, Yutao Cui, Yuanbo Peng, Zhentao Yu, Zhiyu He, Zhiyong Xu, Zixiang Zhou, Zunnan Xu, Yangyu Tao, Qinglin Lu, Songtao Liu, Daquan Zhou, Hongfa Wang, Yong Yang, Di Wang, Yuhong Liu, Jie Jiang, and Caesar Zhong. 2024. HunyuanVideo: A Systematic Framework For Large Video Generative Models. arXiv preprint arXiv:2412.03603 (2024). https:\/\/arxiv.org\/abs\/2412.03603"},{"key":"e_1_2_2_23_1","volume-title":"Jamie Ryan Kiros, and Geoffrey E Hinton","author":"Ba Jimmy Lei","year":"2016","unstructured":"Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. ArXiv e-prints (2016), arXiv-1607."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3130800.3130813"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01864"},{"key":"e_1_2_2_26_1","unstructured":"Tingting Liao Hongwei Yi Yuliang Xiu Jiaxiang Tang Yangyi Huang Justus Thies and Michael J Black. 2023. TADA! Text to Animatable Digital Avatars. ArXiv (Aug 2023)."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3478513.3480528"},{"key":"e_1_2_2_28_1","volume-title":"HumanGaussian: Text-Driven 3D Human Generation with Gaussian Splatting. arXiv preprint arXiv:2311.17061","author":"Liu Xian","year":"2023","unstructured":"Xian Liu, Xiaohang Zhan, Jiaxiang Tang, Ying Shan, Gang Zeng, Dahua Lin, Xihui Liu, and Ziwei Liu. 2023. HumanGaussian: Text-Driven 3D Human Generation with Gaussian Splatting. arXiv preprint arXiv:2311.17061 (2023)."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3596711.3596800"},{"key":"e_1_2_2_30_1","unstructured":"LumaAI. 2024. Luma Dream Machine. https:\/\/lumalabs.ai\/dream-machine. Accessed: 2025-01-22."},{"key":"e_1_2_2_31_1","volume-title":"Open-magvit2: An open-source project toward democratizing auto-regressive visual generation. arXiv preprint arXiv:2409.04410","author":"Luo Zhuoyan","year":"2024","unstructured":"Zhuoyan Luo, Fengyuan Shi, Yixiao Ge, Yujiu Yang, Limin Wang, and Ying Shan. 2024. Open-magvit2: An open-source project toward democratizing auto-regressive visual generation. arXiv preprint arXiv:2409.04410 (2024)."},{"key":"e_1_2_2_32_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). 5441\u20135450","author":"Mahmood Naureen","unstructured":"Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. 2019. AMASS: Archive of Motion Capture as Surface Shapes. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). 5441\u20135450."},{"key":"e_1_2_2_33_1","volume-title":"Occupancy Networks: Learning 3D Reconstruction in Function Space. In CVPR.","author":"Mescheder Lars","year":"2019","unstructured":"Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. 2019. Occupancy Networks: Learning 3D Reconstruction in Function Space. In CVPR."},{"key":"e_1_2_2_34_1","volume-title":"Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV.","author":"Mildenhall Ben","year":"2020","unstructured":"Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2020. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV."},{"key":"e_1_2_2_35_1","volume-title":"Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784","author":"Mirza Mehdi","year":"2014","unstructured":"Mehdi Mirza and Simon Osindero. 2014. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784 (2014)."},{"key":"e_1_2_2_36_1","unstructured":"Mochi-Team. 2024. Mochi. https:\/\/github.com\/Mochi-Team\/mochi."},{"key":"e_1_2_2_37_1","volume-title":"Openvid-1m: A large-scale high-quality dataset for text-to-video generation. arXiv preprint arXiv:2407.02371","author":"Nan Kepan","year":"2024","unstructured":"Kepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan, Zhenheng Yang, Zhijie Chen, Xiang Li, Jian Yang, and Ying Tai. 2024. Openvid-1m: A large-scale high-quality dataset for text-to-video generation. arXiv preprint arXiv:2407.02371 (2024)."},{"key":"e_1_2_2_38_1","unstructured":"OpenAI. 2024. Video generation models as world simulators. https:\/\/openai.com\/index\/video-generation-models-as-world-simulators\/. Accessed: 2024-05-19."},{"key":"e_1_2_2_39_1","volume-title":"Deepsdf: Learning continuous signed distance functions for shape representation. In CVPR.","author":"Park Jeong Joon","year":"2019","unstructured":"Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. 2019. Deepsdf: Learning continuous signed distance functions for shape representation. In CVPR."},{"key":"e_1_2_2_40_1","unstructured":"Nikhila Ravi Valentin Gabeur Yuan-Ting Hu Ronghang Hu Chaitanya Ryali Tengyu Ma Haitham Khedr Roman R\u00e4dle Chloe Rolland Laura Gustafson et al. 2024. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714 (2024)."},{"key":"e_1_2_2_41_1","unstructured":"Fitsum Reda Jinwei Gu Xian Liu Songwei Ge Ting-Chun Wang Haoxiang Wang and Ming-Yu Liu. 2024. Cosmos-Tokenizer. https:\/\/github.com\/NVIDIA\/Cosmos-Tokenizer."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_2_2_43_1","unstructured":"RunwayAI. 2025. Runway Gen-3. https:\/\/runwayml.com\/. Accessed: 2025-01-22."},{"key":"e_1_2_2_44_1","volume-title":"Human4DiT: Free-view Human Video Generation with 4D Diffusion Transformer. arXiv preprint arXiv:2405.17405","author":"Shao Ruizhi","year":"2024","unstructured":"Ruizhi Shao, Youxin Pang, Zerong Zheng, Jingxiang Sun, and Yebin Liu. 2024. Human4DiT: Free-view Human Video Generation with 4D Diffusion Transformer. arXiv preprint arXiv:2405.17405 (2024)."},{"key":"e_1_2_2_45_1","volume-title":"Appearance and pose-conditioned human image generation using deformable gans","author":"Siarohin Aliaksandr","year":"2019","unstructured":"Aliaksandr Siarohin, St\u00e9phane Lathuili\u00e8re, Enver Sangineto, and Nicu Sebe. 2019a. Appearance and pose-conditioned human image generation using deformable gans. IEEE transactions on pattern analysis and machine intelligence 43, 4 (2019), 1156\u20131171."},{"key":"e_1_2_2_46_1","volume-title":"First order motion model for image animation. Advances in neural information processing systems 32","author":"Siarohin Aliaksandr","year":"2019","unstructured":"Aliaksandr Siarohin, St\u00e9phane Lathuili\u00e8re, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019b. First order motion model for image animation. Advances in neural information processing systems 32 (2019)."},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00359"},{"key":"e_1_2_2_48_1","volume-title":"Denoising Diffusion Implicit Models. arXiv:2010.02502 (October","author":"Song Jiaming","year":"2020","unstructured":"Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020. Denoising Diffusion Implicit Models. arXiv:2010.02502 (October 2020). https:\/\/arxiv.org\/abs\/2010.02502"},{"key":"e_1_2_2_49_1","volume-title":"Animate-x: Universal character image animation with enhanced motion representation. arXiv preprint arXiv:2410.10306","author":"Tan Shuai","year":"2024","unstructured":"Shuai Tan, Biao Gong, Xiang Wang, Shiwei Zhang, Dandan Zheng, Ruobing Zheng, Kecheng Zheng, Jingdong Chen, and Ming Yang. 2024. Animate-x: Universal character image animation with enhanced motion representation. arXiv preprint arXiv:2410.10306 (2024)."},{"key":"e_1_2_2_50_1","volume-title":"Deep Patch Visual Odometry. Advances in Neural Information Processing Systems","author":"Teed Zachary","year":"2023","unstructured":"Zachary Teed, Lahav Lipson, and Jia Deng. 2023. Deep Patch Visual Odometry. Advances in Neural Information Processing Systems (2023)."},{"key":"e_1_2_2_51_1","volume-title":"A good image generator is what you need for high-resolution video synthesis. arXiv preprint arXiv:2104.15069","author":"Tian Yu","year":"2021","unstructured":"Yu Tian, Jian Ren, Menglei Chai, Kyle Olszewski, Xi Peng, Dimitris N Metaxas, and Sergey Tulyakov. 2021. A good image generator is what you need for high-resolution video synthesis. arXiv preprint arXiv:2104.15069 (2021)."},{"key":"e_1_2_2_52_1","volume-title":"MusePose: a Pose-Driven Image-to-Video Framework for Virtual Human Generation. arxiv","author":"Tong Zhengyan","year":"2024","unstructured":"Zhengyan Tong, Chao Li, Zhaokang Chen, Bin Wu, and Wenjiang Zhou. 2024. MusePose: a Pose-Driven Image-to-Video Framework for Virtual Human Generation. arxiv (2024)."},{"key":"e_1_2_2_53_1","volume-title":"FVD: A new metric for video generation.","author":"Unterthiner Thomas","year":"2019","unstructured":"Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Rapha\u00ebl Marinier, Marcin Michalski, and Sylvain Gelly. 2019. FVD: A new metric for video generation. (2019)."},{"key":"e_1_2_2_54_1","unstructured":"Dani Valevski Yaniv Leviathan Moab Arar and Shlomi Fruchter. 2024. Diffusion Models Are Real-Time Game Engines. arXiv:2408.14837 [cs.LG] https:\/\/arxiv.org\/abs\/2408.14837"},{"key":"e_1_2_2_55_1","unstructured":"Aaron Van Den Oord Oriol Vinyals et al. 2017. Neural discrete representation learning. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_56_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_57_1","volume-title":"International Conference on Learning Representations.","author":"Villegas Ruben","year":"2022","unstructured":"Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. 2022. Phenaki: Variable length video generation from open domain textual descriptions. In International Conference on Learning Representations."},{"key":"e_1_2_2_58_1","volume-title":"Wan: Open and Advanced Large-Scale Video Generative Models. arXiv preprint arXiv:2503.20314","author":"Wang Ang","year":"2025","unstructured":"Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, et al. 2025. Wan: Open and Advanced Large-Scale Video Generative Models. arXiv preprint arXiv:2503.20314 (2025)."},{"key":"e_1_2_2_59_1","doi-asserted-by":"crossref","unstructured":"Qiuheng Wang Yukai Shi Jiarong Ou Rui Chen Ke Lin Jiahao Wang Boyuan Jiang Haotian Yang Mingwu Zheng Xin Tao et al. 2024. Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content. arXiv preprint arXiv:2410.08260 (2024).","DOI":"10.1109\/CVPR52734.2025.00789"},{"key":"e_1_2_2_60_1","volume-title":"Disco: Disentangled control for referring human dance generation in real world. arXiv e-prints","author":"Wang Tan","year":"2023","unstructured":"Tan Wang, Linjie Li, Kevin Lin, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, and Lijuan Wang. 2023. Disco: Disentangled control for referring human dance generation in real world. arXiv e-prints (2023), arXiv\u20132307."},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00991"},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00531"},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01573"},{"key":"e_1_2_2_64_1","volume-title":"Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou.","author":"Xu Zhongcong","year":"2024","unstructured":"Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou. 2024. MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model."},{"key":"e_1_2_2_65_1","volume-title":"Videogpt: Video generation using vq-vae and transformers. arXiv preprint arXiv:2104.10157","author":"Yan Wilson","year":"2021","unstructured":"Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. 2021. Videogpt: Video generation using vq-vae and transformers. arXiv preprint arXiv:2104.10157 (2021)."},{"key":"e_1_2_2_66_1","volume-title":"Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion. arXiv preprint arXiv:2402.03162","author":"Yang Shiyuan","year":"2024","unstructured":"Shiyuan Yang, Liang Hou, Haibin Huang, Chongyang Ma, Pengfei Wan, Di Zhang, Xiaodong Chen, and Jing Liao. 2024a. Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion. arXiv preprint arXiv:2402.03162 (2024)."},{"key":"e_1_2_2_67_1","volume-title":"Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072","author":"Yang Zhuoyi","year":"2024","unstructured":"Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. 2024b. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072 (2024)."},{"key":"e_1_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW60793.2023.00455"},{"key":"e_1_2_2_69_1","volume-title":"Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu.","author":"Yu Jiahui","year":"2021","unstructured":"Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. 2021. Vector-quantized image modeling with improved vqgan. arXiv preprint arXiv:2110.04627 (2021)."},{"key":"e_1_2_2_70_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01008"},{"key":"e_1_2_2_71_1","unstructured":"Lijun Yu Jos\u00e9 Lezama Nitesh B Gundavarapu Luca Versari Kihyuk Sohn David Minnen Yong Cheng Vighnesh Birodkar Agrim Gupta Xiuye Gu et al. 2023b. Language Model Beats Diffusion-Tokenizer is Key to Visual Generation. arXiv preprint arXiv:2310.05737 (2023)."},{"key":"e_1_2_2_72_1","doi-asserted-by":"crossref","unstructured":"Richard Zhang Phillip Isola Alexei A Efros Eli Shechtman and Oliver Wang. 2018. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In CVPR.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_2_2_73_1","volume-title":"CV-VAE: A Compatible Video VAE for Latent Generative Video Models. https:\/\/arxiv.org\/abs\/2405.20279","author":"Zhao Sijie","year":"2024","unstructured":"Sijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Muyao Niu, Xiaoyu Li, Wenbo Hu, and Ying Shan. 2024. CV-VAE: A Compatible Video VAE for Latent Generative Video Models. https:\/\/arxiv.org\/abs\/2405.20279 (2024)."},{"key":"e_1_2_2_74_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01818"},{"key":"e_1_2_2_75_1","unstructured":"Zangwei Zheng Xiangyu Peng Tianji Yang Chenhui Shen Shenggui Li Hongxin Liu Yukun Zhou Tianyi Li and Yang You. 2024. Open-Sora: Democratizing Efficient Video Production for All. https:\/\/github.com\/hpcaitech\/Open-Sora"},{"key":"e_1_2_2_76_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3592101","article-title":"Avatarrex: Real-time expressive full-body avatars","volume":"42","author":"Zheng Zerong","year":"2023","unstructured":"Zerong Zheng, Xiaochen Zhao, Hongwen Zhang, Boning Liu, and Yebin Liu. 2023b. Avatarrex: Real-time expressive full-body avatars. ACM Transactions on Graphics (TOG) 42, 4 (2023), 1\u201319.","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_2_77_1","volume-title":"European Conference on Computer Vision (ECCV).","author":"Zhu Shenhao","year":"2024","unstructured":"Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Yinghui Xu, Xun Cao, Yao Yao, Hao Zhu, and Siyu Zhu. 2024. Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance. In European Conference on Computer Vision (ECCV)."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3731165","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T17:55:41Z","timestamp":1774634141000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3731165"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,27]]},"references-count":77,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,8,1]]}},"alternative-id":["10.1145\/3731165"],"URL":"https:\/\/doi.org\/10.1145\/3731165","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,27]]},"assertion":[{"value":"2025-07-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}