{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T14:32:14Z","timestamp":1765031534954,"version":"3.46.0"},"reference-count":201,"publisher":"Association for Computing Machinery (ACM)","issue":"6","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>\n                    An image may convey a thousand words, but a video, composed of hundreds or thousands of image frames, tells a more intricate story. Despite significant progress in multimodal large language models (MLLMs), generating extended videos remains a formidable challenge. As of this writing, OpenAI\u2019s Sora\u00a0[\n                    <jats:xref ref-type=\"bibr\">1<\/jats:xref>\n                    ], the current state-of-the-art system, is still limited to producing videos of up to one minute in length. This limitation stems from the complexity of long video generation, which requires more than generative AI techniques for approximating density functions. Critical elements, such as planning, narrative construction, and spatiotemporal continuity, pose significant challenges. Integrating generative AI with a divide-and-conquer approach could improve scalability for longer videos while offering greater control. In this survey, we examine the current landscape of long video generation, covering foundational techniques such as GANs and diffusion models, video generation strategies, large-scale training datasets, quality metrics for evaluating long videos, and future research areas to address the limitations of existing video generation capabilities. We believe it would serve as a comprehensive foundation, offering extensive information to guide future advancements and research in the field of long video generation.\n                  <\/jats:p>","DOI":"10.1145\/3771724","type":"journal-article","created":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T11:29:56Z","timestamp":1760095796000},"page":"1-35","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Video is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-3382-9411","authenticated-orcid":false,"given":"Faraz","family":"Waseem","sequence":"first","affiliation":[{"name":"University of Reading","place":["Reading, United Kingdom of Great Britain and Northern Ireland"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-9394-343X","authenticated-orcid":false,"given":"Muhammad","family":"Shahzad","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Reading","place":["Reading, United Kingdom of Great Britain and Northern Ireland"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,12,6]]},"reference":[{"unstructured":"Sora Team. 2024. Video generation models as world simulators by Open A.I. (2024). Retrieved 1 July 2024 from https:\/\/openai.com\/index\/video-generation-models-as-world-simulators\/","key":"e_1_3_1_2_2"},{"unstructured":"Open A.I Team. 2022. Introducing ChatGPT by Open A.I. (2022). Retrieved 22 October 2024 from https:\/\/openai.com\/index\/chatgpt\/","key":"e_1_3_1_3_2"},{"unstructured":"Meta A.I Team. 2023. Meta LLama Models. (2023). Retrieved 21 September 2024 from https:\/\/www.llama.com\/","key":"e_1_3_1_4_2"},{"unstructured":"Google A.I Team. 2023. Google Gemini series. (2023). Retrieved 20 June 2024 from https:\/\/gemini.google.com\/","key":"e_1_3_1_5_2"},{"unstructured":"Claude A.I Team. 2023. Anthropic by Claude. (2023). Retrieved 8 June 2024 from https:\/\/www.anthropic.com\/claude","key":"e_1_3_1_6_2"},{"doi-asserted-by":"crossref","unstructured":"Mistral A.I Team. 2024. Mistral Large Model by Mistral. (2024). Retrieved 1 September 20 from https:\/\/mistral.ai\/news\/mistral-large\/","key":"e_1_3_1_7_2","DOI":"10.5840\/mr201814"},{"unstructured":"Aditya Ramesh Prafulla Dhariwal Alex Nichol Casey Chu and Mark Chen. 2022. Hierarchical text-conditional image generation with CLIP latents. ArXiv abs\/2204.06125 (2022). Retrieved from https:\/\/arxiv.org\/abs\/2204.06125","key":"e_1_3_1_8_2"},{"unstructured":"Patrick Esser Sumith Kulal Andreas Blattmann Rahim Entezari Jonas M\u00fcller Harry Saini Yam Levi Dominik Lorenz Axel Sauer Frederic Boesel Dustin Podell Tim Dockhorn Zion English and Robin Rombach. 2024. Scaling rectified flow transformers for high-resolution image synthesis. In ICML. Retrieved from https:\/\/openreview.net\/forum?id=FPnUhsQJ5B","key":"e_1_3_1_9_2"},{"unstructured":"Midjourney A.I Team. 2024. The Midjourney V5.2 model For Image generation. (2024). Retrieved 10 August 2024 from https:\/\/docs.midjourney.com\/docs\/model-version-5","key":"e_1_3_1_10_2"},{"key":"e_1_3_1_11_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Hong Wenyi","year":"2023","unstructured":"Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. 2023. CogVideo: Large-scale pretraining for text-to-video generation via transformers. In Proceedings of the 11th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=rB6TpjAuSRy"},{"unstructured":"Han Lin Abhay Zala Jaemin Cho and Mohit Bansal. 2024. VideoDirectorGPT: Consistent Multi-Scene Video Generation via LLM-Guided Planning. (2024). Retrieved from https:\/\/openreview.net\/forum?id=5PkgaUwiY0","key":"e_1_3_1_12_2"},{"key":"e_1_3_1_13_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Singer Uriel","year":"2023","unstructured":"Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et\u00a0al. 2023. Make-a-video: Text-to-video generation without text-video data. In Proceedings of the 11th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=nJfylDvgzlq"},{"unstructured":"Midjourney Team. 2023. Gen2 by Runway ML. (2023). Retrieved 20 June 2024 from https:\/\/runwayml.com\/research\/gen-2","key":"e_1_3_1_14_2"},{"key":"e_1_3_1_15_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Hong Wenyi","year":"2023","unstructured":"Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. 2023. CogVideo: Large-scale pretraining for text-to-video generation via transformers. In Proceedings of the 11th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=rB6TpjAuSRy"},{"key":"e_1_3_1_16_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Villegas Ruben","year":"2023","unstructured":"Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. 2023. Phenaki: Variable length video generation from open domain textual descriptions. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=vOEXS39nOF"},{"doi-asserted-by":"publisher","unstructured":"Fu-Yun Wang Wenshuo Chen Guanglu Song Han-Jia Ye Yu Liu and Hongsheng Li. 2023. Gen-L-Video: Multi-text to long video generation via temporal co-denoising. CoRR abs\/2305.18264 (2023). Retrieved from 10.48550\/arXiv.2305.18264","key":"e_1_3_1_17_2","DOI":"10.48550\/arXiv.2305.18264"},{"doi-asserted-by":"publisher","key":"e_1_3_1_18_2","DOI":"10.1109\/ICCV51070.2023.00387"},{"unstructured":"Midjourney Team. 2024. Gen-4 Alpha by MidJourney. (2024). Retrieved 23 March 2025 from https:\/\/runwayml.com\/research\/introducing-runway-gen-4","key":"e_1_3_1_19_2"},{"unstructured":"Chengxuan Li Di Huang Zeyu Lu Yang Xiao Qingqi Pei and Lei Bai. 2024. A survey on long video generation: Challenges methods and prospects. arXiv preprint arXiv:2403.16407 (2024).","key":"e_1_3_1_20_2"},{"doi-asserted-by":"crossref","unstructured":"Pengyuan Zhou Lin Wang Zhi Liu Yanbin Hao Pan Hui Sasu Tarkoma and Jussi Kangasharju. 2024. A survey on generative AI and LLM for video generation understanding and streaming. arXiv preprint arXiv:2404.16038 (2024).","key":"e_1_3_1_21_2","DOI":"10.36227\/techrxiv.171172801.19993069\/v1"},{"unstructured":"Yixin Liu Kai Zhang Yuan Li Zhiling Yan Chujie Gao Ruoxi Chen Zhengqing Yuan Yue Huang Hanchi Sun Jianfeng Gao and others. 2024. Sora: A review on background technology limitations and opportunities of large vision models. arXiv preprint arXiv:2402.17177 (2024).","key":"e_1_3_1_22_2"},{"key":"e_1_3_1_23_2","volume-title":"Proceedings of the 37th Conference on Neural Information Processing Systems","author":"Huang Hanzhuo","year":"2023","unstructured":"Hanzhuo Huang, Yufan Feng, Cheng Shi, Lan Xu, Jingyi Yu, and Sibei Yang. 2023. Free-bloom: Zero-shot text-to-video generator with LLM director and LDM animator. In Proceedings of the 37th Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=paa2OU5jN8"},{"unstructured":"Yu Lu Linchao Zhu Hehe Fan and Yi Yang. 2023. Flowzero: Zero-shot text-to-video synthesis with LLM-driven dynamic scene syntax. arXiv preprint arXiv:2311.15813 (2023).","key":"e_1_3_1_24_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_25_2","DOI":"10.1109\/CVPR52729.2023.02161"},{"unstructured":"Bowen Zhang Xiaofei Xie Haotian Lu Na Ma Tianlin Li and Qing Guo. 2024. Mavin: Multi-action video generation with diffusion models via transition video infilling. arXiv preprint arXiv:2405.18003 (2024).","key":"e_1_3_1_26_2"},{"key":"e_1_3_1_27_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Chen Xinyuan","year":"2024","unstructured":"Xinyuan Chen, Yaohui Wang, Lingjun Zhang, Shaobin Zhuang, Xin Ma, Jiashuo Yu, Yali Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. 2024. SEINE: Short-to-long video diffusion model for generative transition and prediction. In Proceedings of the 12th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=FNq3nIvP4F"},{"unstructured":"Alec Radford Jong Wook Kim Chris Hallacy Aditya Ramesh Gabriel Goh Sandhini Agarwal Girish Sastry Amanda Askell Pamela Mishkin Jack Clark and others. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning. PMLR 8748\u20138763.","key":"e_1_3_1_28_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_29_2","DOI":"10.1145\/3422622"},{"unstructured":"Alec Radford Luke Metz and Soumith Chintala. 2015. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434 (2015).","key":"e_1_3_1_30_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_31_2","DOI":"10.5555\/2969239.2969405"},{"doi-asserted-by":"publisher","key":"e_1_3_1_32_2","DOI":"10.1109\/ICCV.2017.629"},{"doi-asserted-by":"publisher","key":"e_1_3_1_33_2","DOI":"10.1109\/CVPR.2018.00143"},{"doi-asserted-by":"publisher","key":"e_1_3_1_34_2","DOI":"10.1109\/CVPR42600.2020.00813"},{"doi-asserted-by":"publisher","key":"e_1_3_1_35_2","DOI":"10.1109\/ICCV.2017.244"},{"doi-asserted-by":"publisher","key":"e_1_3_1_36_2","DOI":"10.1007\/978-3-030-58542-6_11"},{"doi-asserted-by":"publisher","key":"e_1_3_1_37_2","DOI":"10.1109\/CVPR.2017.632"},{"unstructured":"Michael Mathieu Camille Couprie and Yann LeCun. 2015. Deep multi-scale video prediction beyond mean square error. arXiv preprint arXiv:1511.05440 (2015).","key":"e_1_3_1_38_2"},{"unstructured":"Carl Vondrick Hamed Pirsiavash and Antonio Torralba. 2016. Generating videos with scene dynamics. Advances in Neural Information Processing Systems 29 (2016).","key":"e_1_3_1_39_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_40_2","DOI":"10.1109\/CVPR.2018.00251"},{"doi-asserted-by":"crossref","unstructured":"Yingwei Pan Zhaofan Qiu Ting Yao Houqiang Li and Tao Mei. 2017. To create what you tell: Generating videos from captions. In Proceedings of the 25th ACM International conference on Multimedia. 1789\u20131798.","key":"e_1_3_1_41_2","DOI":"10.1145\/3123266.3127905"},{"key":"e_1_3_1_42_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Yu Sihyun","year":"2022","unstructured":"Sihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim, Junho Kim, Jung-Woo Ha, and Jinwoo Shin. 2022. Generating videos with dynamics-aware implicit generative adversarial networks. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=Czsdv-S4-w9"},{"doi-asserted-by":"publisher","key":"e_1_3_1_43_2","DOI":"10.1109\/CVPR52688.2022.00361"},{"unstructured":"Standford tutorial. Auto-Encoders. Retrieved 10 July 2024 from http:\/\/ufldl.stanford.edu\/tutorial\/unsupervised\/Autoencoders\/","key":"e_1_3_1_44_2"},{"unstructured":"Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013).","key":"e_1_3_1_45_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_46_2","DOI":"10.1109\/CVPR52688.2022.01553"},{"doi-asserted-by":"publisher","key":"e_1_3_1_47_2","DOI":"10.1109\/CVPR52688.2022.01042"},{"unstructured":"Aaron Van Den Oord Oriol Vinyals and others. 2017. Neural discrete representation learning. Advances in Neural Information Processing Systems 30 (2017).","key":"e_1_3_1_48_2"},{"unstructured":"Wilson Yan Yunzhi Zhang Pieter Abbeel and Aravind Srinivas. 2021. VideoGPT: Video generation using VQ-VAE and transformers. arXiv preprint arXiv:2104.10157 (2021).","key":"e_1_3_1_49_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_50_2","DOI":"10.1109\/CVPR46437.2021.01268"},{"key":"e_1_3_1_51_2","first-page":"8821","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"Ramesh Aditya","year":"2021","unstructured":"Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. CLIP: Connecting vision and language with contrastive learning. In Proceedings of the 38th International Conference on Machine Learning. PMLR, 8821\u20138831. Retrieved from https:\/\/proceedings.mlr.press\/v139\/ramesh21a.html"},{"unstructured":"Shir Gur Sagie Benaim and Lior Wolf. 2020. Hierarchical patch VAE-GAN: Generating diverse videos from a single sample. In Advances in Neural Information Processing Systems H. Larochelle M. Ranzato R. Hadsell M. F. Balcan and H. Lin (Eds.). Curran Associates Inc. 16761\u201316772. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2020\/file\/c2f32522a84d5e6357e6abac087f1b0b-Paper.pdf","key":"e_1_3_1_52_2"},{"unstructured":"Faraz Waseem Rafael Perez Martinez and Chris Wu. 2022. Visual anomaly detection in video by variational autoencoder. arXiv preprint arXiv:2203.03872 (2022).","key":"e_1_3_1_53_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_54_2","DOI":"10.1109\/CVPR52733.2024.02145"},{"doi-asserted-by":"publisher","key":"e_1_3_1_55_2","DOI":"10.1109\/CVPR52729.2023.01008"},{"doi-asserted-by":"publisher","key":"e_1_3_1_56_2","DOI":"10.1109\/CVPR52729.2023.02235"},{"unstructured":"Quan Hoang Tu Dinh Nguyen Trung Le and Dinh Phung. 2017. Multi-generator generative adversarial nets. arXiv preprint arXiv:1708.02556 (2017).","key":"e_1_3_1_57_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_58_2","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_1_59_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et\u00a0al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=YicbFdNTTy"},{"key":"e_1_3_1_60_2","series-title":"Proceedings of Machine Learning Research","first-page":"8821","volume-title":"Proceedings of the 38th International Conference on Machine Learning.","volume":"139","author":"Ramesh Aditya","year":"2021","unstructured":"Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In Proceedings of the 38th International Conference on Machine Learning.Marina Meila and Tong Zhang (Eds.), Proceedings of Machine Learning Research, Vol. 139, PMLR, 8821\u20138831. Retrieved from https:\/\/proceedings.mlr.press\/v139\/ramesh21a.html"},{"unstructured":"Jason Tyler Rolfe. 2016. Discrete variational autoencoders. arXiv preprint arXiv:1609.02200 (2016).","key":"e_1_3_1_61_2"},{"key":"e_1_3_1_62_2","volume-title":"Proceedings of the 35th International Conference on Neural Information Processing Systems (NIPS \u201921)","author":"Ding Ming","year":"2024","unstructured":"Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. 2024. CogView: Mastering text-to-image generation via transformers. In Proceedings of the 35th International Conference on Neural Information Processing Systems (NIPS \u201921). Curran Associates Inc., Red Hook, NY, USA, Article 1516, 14 pages."},{"key":"e_1_3_1_63_2","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS \u201922)","author":"Ding Ming","year":"2024","unstructured":"Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. 2024. CogView2: Faster and better text-to-image generation via hierarchical transformers. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS \u201922). Curran Associates Inc., Red Hook, NY, USA, Article 1229, 13 pages."},{"doi-asserted-by":"publisher","key":"e_1_3_1_64_2","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"e_1_3_1_65_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Villegas Ruben","year":"2023","unstructured":"Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. 2023. Phenaki: Variable length video generation from open domain textual descriptions. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=vOEXS39nOF"},{"unstructured":"Adam Roberts Hyung Won Chung Gaurav Mishra Anselm Levskaya James Bradbury Daniel Andor Sharan Narang Brian Lester Colin Gaffney Afroz Mohiuddin and others. 2023. Scaling up models and data with t5x and seqio. Journal of Machine Learning Research 24 377 (2023) 1\u20138.","key":"e_1_3_1_66_2"},{"key":"e_1_3_1_67_2","volume-title":"Proceedings of the 38th Annual Conference on Neural Information Processing Systems (NeurIPS 2024)","author":"Zhu Hanxin","year":"2024","unstructured":"Hanxin Zhu, Tianyu He, Anni Tang, Junliang Guo, Zhibo Chen, and Jiang Bian. 2024. Compositional 3D-aware video generation with LLM director. In Proceedings of the 38th Annual Conference on Neural Information Processing Systems (NeurIPS 2024). Retrieved from https:\/\/nips.cc\/virtual\/2024\/poster\/93599Poster presentation."},{"key":"e_1_3_1_68_2","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS \u201920)","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et\u00a0al. 2020. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS \u201920). Curran Associates Inc., Red Hook, NY, USA, Article 159, 25 pages."},{"key":"e_1_3_1_69_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Lian Long","year":"2024","unstructured":"Long Lian, Baifeng Shi, Adam Yala, Trevor Darrell, and Boyi Li. 2024. LLM-grounded Video Diffusion Models. In Proceedings of the 12th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=exKHibougU"},{"doi-asserted-by":"publisher","key":"e_1_3_1_70_2","DOI":"10.1109\/CVPR52733.2024.00841"},{"unstructured":"Jascha Sohl-Dickstein Eric Weiss Niru Maheswaranathan and Surya Ganguli. 2015. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning. pmlr 2256\u20132265.","key":"e_1_3_1_71_2"},{"key":"e_1_3_1_72_2","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS \u201920)","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS \u201920). Curran Associates Inc., Red Hook, NY, USA, Article 574, 12 pages."},{"unstructured":"Yang Song and Stefano Ermon. 2019. Generative modeling by estimating gradients of the data distribution. Advances in Neural Information Processing Systems 32 (2019).","key":"e_1_3_1_73_2"},{"doi-asserted-by":"crossref","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Volume 1 (Long and Short Papers). 4171\u20134186.","key":"e_1_3_1_74_2","DOI":"10.18653\/v1\/N19-1423"},{"issue":"1","key":"e_1_3_1_75_2","first-page":"67","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21, 1, Article 140 (2020), 67 pages.","journal-title":"Journal of Machine Learning Research"},{"unstructured":"Open A.I Team. Video generation models as world simulators. Retrieved from https:\/\/openai.com\/index\/video-generation-models-as-world-simulators\/","key":"e_1_3_1_76_2"},{"unstructured":"Ming Ding Zhuoyi Yang Wenyi Hong Wendi Zheng Chang Zhou Da Yin Junyang Lin Xu Zou Zhou Shao Hongxia Yang and Jie Tang. 2021. CogView: Mastering text-to-image generation via transformers. In Advances in Neural Information Processing Systems Curran Associates Inc. 19822\u201319835. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2021\/file\/a4d92e2cd541fca87e4620aba658316d-Paper.pdf","key":"e_1_3_1_77_2"},{"key":"e_1_3_1_78_2","volume-title":"Proceedings of the 38th Annual Conference on Neural Information Processing Systems","author":"Tian Ye","year":"2024","unstructured":"Ye Tian, Ling Yang, Haotian Yang, Yuan Gao, Yufan Deng, Xintao Wang, Zhaochen Yu, Xin Tao, Pengfei Wan, Di ZHANG, et\u00a0al. 2024. VideoTetris: Towards compositional text-to-video generation. In Proceedings of the 38th Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=RPM7STrnVz"},{"key":"e_1_3_1_79_2","first-page":"16890","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"35","author":"Ding Ming","year":"2022","unstructured":"Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. 2022. CogView2: Faster and better text-to-image generation via hierarchical transformers. In Proceedings of the Advances in Neural Information Processing Systems. S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35, Curran Associates, Inc., 16890\u201316902. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/6baec7c4ba0a8734ccbd528a8090cb1f-Paper-Conference.pdf"},{"key":"e_1_3_1_80_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Liang Jian","year":"2022","unstructured":"Jian Liang, Chenfei Wu, Xiaowei Hu, Zhe Gan, Jianfeng Wang, Lijuan Wang, Zicheng Liu, Yuejian Fang, and Nan Duan. 2022. NUWA-infinity: Autoregressive over autoregressive generation for infinite visual synthesis. In Proceedings of the Advances in Neural Information Processing Systems. Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (Eds.). Retrieved from https:\/\/openreview.net\/forum?id=0Kv7cLhuhQT"},{"key":"e_1_3_1_81_2","volume-title":"Proceedings of the 41st International Conference on Machine Learning (ICML\u201924)","author":"Kondratyuk Dan","year":"2024","unstructured":"Dan Kondratyuk, Lijun Yu, Xiuye Gu, Jos\u00e9 Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, et\u00a0al. 2024. VideoPoet: A large language model for zero-shot video generation. In Proceedings of the 41st International Conference on Machine Learning (ICML\u201924). JMLR.org, Article 1005, 20 pages."},{"doi-asserted-by":"publisher","key":"e_1_3_1_82_2","DOI":"10.1109\/CVPRW63382.2024.00735"},{"doi-asserted-by":"publisher","key":"e_1_3_1_83_2","DOI":"10.1109\/CVPR52733.2024.00834"},{"unstructured":"Thomas Unterthiner Sjoerd van Steenkiste Karol Kurach Rapha\u00ebl Marinier Marcin Michalski and Sylvain Gelly. 2019. FVD: A new metric for video generation. In DGS@ICLR. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:198489709","key":"e_1_3_1_84_2"},{"unstructured":"Zongyi Li Shujie Hu Shujie Liu Long Zhou Jeongsoo Choi Lingwei Meng Xun Guo Jinyu Li Hefei Ling and Furu Wei. 2024. ARLON: Boosting diffusion transformers with autoregressive models for long video generation. arXiv preprint arXiv:2410.20502 (2024).","key":"e_1_3_1_85_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_86_2","DOI":"10.1109\/CVPR52688.2022.00361"},{"unstructured":"Vincent Sitzmann Julien Martel Alexander Bergman David Lindell and Gordon Wetzstein. 2020. Implicit neural representations with periodic activation functions. Advances in Neural Information Processing Systems 33 (2020) 7462\u20137473.","key":"e_1_3_1_87_2"},{"doi-asserted-by":"crossref","unstructured":"Ben Mildenhall Pratul P. Srinivasan Matthew Tancik Jonathan T. Barron Ravi Ramamoorthi and Ren Ng. 2021. NeRF: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM 65 1 (2021) 99\u2013106.","key":"e_1_3_1_88_2","DOI":"10.1145\/3503250"},{"doi-asserted-by":"publisher","key":"e_1_3_1_89_2","DOI":"10.1109\/CVPR52729.2023.01770"},{"doi-asserted-by":"publisher","key":"e_1_3_1_90_2","DOI":"10.1109\/CVPR52729.2023.02192"},{"volume-title":"Proceedings of the Submitted to The 13th International Conference on Learning Representations","year":"2024","unstructured":"Anonymous. 2024. StreamingT2V: Consistent, dynamic, and extendable long video generation from text. In Proceedings of the Submitted to The 13th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=26oSbRRpEYunder review.","key":"e_1_3_1_91_2"},{"unstructured":"Kaifeng Gao Jiaxin Shi Hanwang Zhang Chunping Wang and Jun Xiao. 2024. ViD-GPT: Introducing GPT-style autoregressive generation in video diffusion models. arXiv preprint arXiv:2406.10981 (2024).","key":"e_1_3_1_92_2"},{"unstructured":"Yichen Ouyang Hao Zhao Gaoang Wang and others. 2024. Flexifilm: Long video generation with flexible conditions. arXiv preprint arXiv:2404.18620 (2024).","key":"e_1_3_1_93_2"},{"unstructured":"Zhengqing Yuan Yixin Liu Yihan Cao Weixiang Sun Haolong Jia Ruoxi Chen Zhaoxu Li Bin Lin Li Yuan Lifang He and others. 2024. Mora: Enabling generalist video generation via a multi-agent framework. arXiv preprint arXiv:2403.13248 (2024).","key":"e_1_3_1_94_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_95_2","DOI":"10.1109\/I2CT61223.2024.10544050"},{"unstructured":"Zhifei Xie Daniel Tang Dingwei Tan Jacques Klein Tegawend F. Bissyand and Saad Ezzini. 2024. Dreamfactory: Pioneering multi-scene long video generation with a multi-agent framework. arXiv preprint arXiv:2408.11788 (2024).","key":"e_1_3_1_96_2"},{"unstructured":"Liu He Yizhi Song Hejun Huang Pinxin Liu Yunlong Tang Daniel Aliaga and Xin Zhou. 2024. Kubrick: Multimodal agent collaborations for synthetic video generation. arXiv preprint arXiv:2408.10453 (2024).","key":"e_1_3_1_97_2"},{"unstructured":"Jiuniu Wang Hangjie Yuan Dayou Chen Yingya Zhang Xiang Wang and Shiwei Zhang. 2023. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571 (2023).","key":"e_1_3_1_98_2"},{"key":"e_1_3_1_99_2","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923)","author":"Huang Hanzhuo","year":"2024","unstructured":"Hanzhuo Huang, Yufan Feng, Cheng Shi, Lan Xu, Jingyi Yu, and Sibei Yang. 2024. Free-bloom: Zero-shot text-to-video generator with LLM director and LDM animator. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923). Curran Associates Inc., Red Hook, NY, USA, Article 1138, 24 pages."},{"key":"e_1_3_1_100_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Song Jiaming","year":"2021","unstructured":"Jiaming Song, Chenlin Meng, and Stefano Ermon. 2021. Denoising diffusion implicit models. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=St1giarCHLP"},{"key":"e_1_3_1_101_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Lian Long","year":"2024","unstructured":"Long Lian, Baifeng Shi, Adam Yala, Trevor Darrell, and Boyi Li. 2024. LLM-grounded video diffusion models. In Proceedings of the 12th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=exKHibougU"},{"unstructured":"Zhengqing Yuan Yixin Liu Yihan Cao Weixiang Sun Haolong Jia Ruoxi Chen Zhaoxu Li Bin Lin Li Yuan Lifang He and others. 2024. Mora: Enabling generalist video generation via a multi-agent framework. arXiv preprint arXiv:2403.13248 (2024).","key":"e_1_3_1_102_2"},{"unstructured":"Yixin Liu Kai Zhang Yuan Li Zhiling Yan Chujie Gao Ruoxi Chen Zhengqing Yuan Yue Huang Hanchi Sun Jianfeng Gao and others. 2024. Sora: A review on background technology limitations and opportunities of large vision models. arXiv preprint arXiv:2402.17177 (2024).","key":"e_1_3_1_103_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_104_2","DOI":"10.1007\/978-3-031-72775-7_23"},{"doi-asserted-by":"publisher","key":"e_1_3_1_105_2","DOI":"10.1109\/ICIP49359.2023.10222725"},{"unstructured":"Siyang Zhang Harry Yang and Ser-Nam Lim. 2025. Videomerge: Towards training-free long video generation. arXiv preprint arXiv:2503.09926 (2025).","key":"e_1_3_1_106_2"},{"doi-asserted-by":"crossref","unstructured":"Gyeongrok Oh Jaehwan Jeong Sieun Kim Wonmin Byeon Jinkyu Kim Sungwoong Kim and Sangpil Kim. 2024. Mevg: Multi-event video generation with text-to-video models. In European Conference on Computer Vision. Springer 401\u2013418.","key":"e_1_3_1_107_2","DOI":"10.1007\/978-3-031-72775-7_23"},{"unstructured":"Weijie Kong Qi Tian Zijian Zhang Rox Min Zuozhuo Dai Jin Zhou Jiangfeng Xiong Xin Li Bo Wu Jianwei Zhang and others. 2024. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603 (2024).","key":"e_1_3_1_108_2"},{"key":"e_1_3_1_109_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Qiu Haonan","year":"2024","unstructured":"Haonan Qiu, Menghan Xia, Yong Zhang, Yingqing He, Xintao Wang, Ying Shan, and Ziwei Liu. 2024. FreeNoise: Tuning-free longer video diffusion via noise rescheduling. In Proceedings of the 12th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=ijoqFqSC7p"},{"key":"e_1_3_1_110_2","volume-title":"Proceedings of the 37th Conference on Neural Information Processing Systems","author":"Sun Mingzhen","year":"2023","unstructured":"Mingzhen Sun, Weining Wang, Zihan Qin, Jiahui Sun, Sihan Chen, and Jing Liu. 2023. GLOBER: Coherent non-autoregressive video generation via GLOBal guided video decodER. In Proceedings of the 37th Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=TRbklCR2ZW"},{"doi-asserted-by":"crossref","unstructured":"Shoufa Chen Chongjian Ge Yuqi Zhang Yida Zhang Fengda Zhu Hao Yang Hongxiang Hao Hui Wu Zhichao Lai Yifei Hu and others. 2025. Goku: Flow based video generative foundation models. In Proceedings of the Computer Vision and Pattern Recognition Conference. 23516\u201323527.","key":"e_1_3_1_111_2","DOI":"10.1109\/CVPR52734.2025.02190"},{"unstructured":"Rui Tian Qi Dai Jianmin Bao Kai Qiu Yifan Yang Chong Luo Zuxuan Wu and Yu-Gang Jiang. 2024. REDUCIO! Generating 1K Video within 16 Seconds using Extremely Compressed Motion Latents. arXiv preprint arXiv:2411.13552 (2024).","key":"e_1_3_1_112_2"},{"key":"e_1_3_1_113_2","article-title":"Mochi 1","author":"Team Genmo","year":"2024","unstructured":"Genmo Team. 2024. Mochi 1. Retrieved from https:\/\/github.com\/genmoai\/models. (2024).","journal-title":"Retrieved from"},{"key":"e_1_3_1_114_2","article-title":"MovieGen: A cast of media foundation models","author":"Team Meta Research","year":"2024","unstructured":"Meta Research Team. 2024. MovieGen: A cast of media foundation models. arXiv preprint (October2024). Retrieved from https:\/\/ai.meta.com\/research\/publications\/movie-gen-a-cast-of-media-foundation-models\/","journal-title":"arXiv preprint"},{"unstructured":"Rui Sun Yumin Zhang Tejal Shah Jiahao Sun Shuoying Zhang Wenqi Li Haoran Duan Bo Wei and Rajiv Ranjan. 2024. From sora what we can see: A survey of text-to-video generation. arXiv preprint arXiv:2405.10674 (2024).","key":"e_1_3_1_115_2"},{"unstructured":"NVIDIA Niket Agarwal Arslan Ali Maciej Bala Yogesh Balaji Erik Barker Tiffany Cai Prithvijit Chattopadhyay Yongxin Chen Yin Cui Yifan Ding Daniel Dworakowski Jiaojiao Fan Michele Fenzi Francesco Ferroni Sanja Fidler Dieter Fox Songwei Ge Yunhao Ge Jinwei Gu Siddharth Gururani Ethan He Jiahui Huang Jacob Huffman Pooya Jannaty Jingyi Jin Seung Wook Kim Gergely Kl\u00e1r Grace Lam Shiyi Lan Laura Leal-Taix\u00e9 Anqi Li Zhaoshuo Li Chen-Hsuan Lin Tsung-Yi Lin Huan Ling Ming-Yu Liu Xian Liu Alice Luo Qianli Ma Hanzi Mao Kaichun Mo Arsalan Mousavian Seungjun Nah Sriharsha Niverty David Page Despoina Paschalidou Zeeshan Patel Lindsey Pavao Morteza Ramezanali Fitsum Reda Xiaowei Ren Vasanth Rao Naik Sabavat Ed Schmerling Stella Shi Bartosz Stefaniak Shitao Tang Lyne Tchapmi Przemek Tredak Wei-Cheng Tseng Jibin Varghese Hao Wang Haoxiang Wang Heng Wang Ting-Chun Wang Fangyin Wei Xinyue Wei Jay Zhangjie Wu Jiashu Xu Wei Yang Lin Yen-Chen Xiaohui Zeng Yu Zeng Jing Zhang Qinsheng Zhang Yuxuan Zhang Qingqing Zhao and Artur Zolkowski. 2025. Cosmos world foundation model platform for physical AI. arXiv preprint arXiv:2501.03575 (2025). Retrieved from https:\/\/arxiv.org\/abs\/2501.03575","key":"e_1_3_1_116_2"},{"doi-asserted-by":"crossref","unstructured":"Chenfei Wu Jian Liang Lei Ji Fan Yang Yuejian Fang Daxin Jiang and Nan Duan. 2022. N\u00dcWA: Visual synthesis pre-training for neural visual world creation. In European Conference on Computer Vision. Springer 720\u2013736.","key":"e_1_3_1_117_2","DOI":"10.1007\/978-3-031-19787-1_41"},{"unstructured":"Younggyo Seo Kimin Lee Fangchen Liu Stephen James and Pieter Abbeel. 2022. HARP: Autoregressive latent video prediction with high-fidelity image generator. In 2022 IEEE International Conference on Image Processing (ICIP). IEEE 3943\u20133947.","key":"e_1_3_1_118_2"},{"doi-asserted-by":"crossref","unstructured":"Qihang Yu Mark Weber Xueqing Deng Xiaohui Shen Daniel Cremers and Liang-Chieh Chen. 2024. An image is worth 32 tokens for reconstruction and generation. Advances in Neural Information Processing Systems 37 (2024) 128940\u2013128966.","key":"e_1_3_1_119_2","DOI":"10.52202\/079017-4096"},{"doi-asserted-by":"crossref","unstructured":"Junke Wang Yi Jiang Zehuan Yuan Bingyue Peng Zuxuan Wu and Yu-Gang Jiang. 2024. OmniTokenizer: A joint image-video tokenizer for visual generation. Advances in Neural Information Processing Systems 37 (2024) 28281\u201328295.","key":"e_1_3_1_120_2","DOI":"10.52202\/079017-0887"},{"unstructured":"David Dehaene and R\u00e9my Brossard. 2021. Re-parameterizing VAEs for stability. arXiv preprint arXiv:2106.13739 (2021).","key":"e_1_3_1_121_2"},{"unstructured":"Fengyuan Shi Zhuoyan Luo Yixiao Ge Yujiu Yang Ying Shan and Limin Wang. 2025. Scalable image tokenization with index backpropagation quantization. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 16037\u201316046.","key":"e_1_3_1_122_2"},{"doi-asserted-by":"crossref","unstructured":"Linjie Li Yen-Chun Chen Yu Cheng Zhe Gan Licheng Yu and Jingjing Liu. 2020. HERO: Hierarchical encoder for video+ language omni-representation pre-training. arXiv preprint arXiv:2005.00200 (2020).","key":"e_1_3_1_123_2","DOI":"10.18653\/v1\/2020.emnlp-main.161"},{"unstructured":"Ziqin Zhou Yifan Yang Yuqing Yang Tianyu He Houwen Peng Kai Qiu Qi Dai Lili Qiu Chong Luo and Lingqiao Liu. 2025. HiTVideo: Hierarchical tokenizers for enhancing text-to-video generation with autoregressive large language models. arXiv preprint arXiv:2503.11513 (2025).","key":"e_1_3_1_124_2"},{"unstructured":"Dawit Mureja Argaw Xian Liu Joon Son Chung Ming-Yu Liu and Fitsum Reda. 2025. MambaVideo for discrete video tokenization with channel-split quantization. arXiv preprint arXiv:2507.04559 (2025).","key":"e_1_3_1_125_2"},{"doi-asserted-by":"crossref","unstructured":"Lijun Yu Yong Cheng Kihyuk Sohn Jos\u00e9 Lezama Han Zhang Huiwen Chang Alexander G Hauptmann Ming-Hsuan Yang Yuan Hao Irfan Essa and others. 2023. MAGVIT: Masked generative video transformer. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10459\u201310469.","key":"e_1_3_1_126_2","DOI":"10.1109\/CVPR52729.2023.01008"},{"key":"e_1_3_1_127_2","volume-title":"Proceedings of the 1st Workshop on Controllable Video Generation @ICML24","author":"Hong Susung","year":"2024","unstructured":"Susung Hong, Junyoung Seo, Heeseong Shin, Sunghwan Hong, and Seungryong Kim. 2024. Large language models are frame-level directors for zero-shot text-to-video generation. In Proceedings of the 1st Workshop on Controllable Video Generation @ICML24. Retrieved from https:\/\/openreview.net\/forum?id=VmOO0GsG0K"},{"doi-asserted-by":"publisher","key":"e_1_3_1_128_2","DOI":"10.1109\/CVPR52733.2024.00804"},{"doi-asserted-by":"crossref","unstructured":"Fuchen Long Zhaofan Qiu Ting Yao and Tao Mei. 2024. VideoStudio: Generating consistent-content and multi-scene videos. In European Conference on Computer Vision. Springer 468\u2013485.","key":"e_1_3_1_129_2","DOI":"10.1007\/978-3-031-73027-6_27"},{"doi-asserted-by":"crossref","unstructured":"Lvmin Zhang Anyi Rao and Maneesh Agrawala. 2023. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 3836\u20133847.","key":"e_1_3_1_130_2","DOI":"10.1109\/ICCV51070.2023.00355"},{"unstructured":"Zhihao Hu and Dong Xu. 2023. VideoControlNet: A motion-guided video-to-video translation framework by using diffusion model with controlnet. arXiv preprint arXiv:2307.14073 (2023).","key":"e_1_3_1_131_2"},{"unstructured":"David Junhao Zhang Dongxu Li Hung Le Mike Zheng Shou Caiming Xiong and Doyen Sahoo. 2024. Moonshot: Towards controllable video generation and editing with multimodal conditions. arXiv preprint arXiv:2401.01827 (2024).","key":"e_1_3_1_132_2"},{"unstructured":"Cong Wang Jiaxi Gu Panwen Hu Haoyu Zhao Yuanfan Guo Jianhua Han Hang Xu and Xiaodan Liang. 2024. EasyControl: Transfer controlnet to video diffusion for controllable generation and interpolation. arXiv preprint arXiv:2408.13005 (2024).","key":"e_1_3_1_133_2"},{"unstructured":"Tianjun Zhang Yi Zhang Vibhav Vineet Neel Joshi and Xin Wang. 2023. Controllable text-to-image generation with GPT-4. arXiv preprint arXiv:2305.18583 (2023).","key":"e_1_3_1_134_2"},{"doi-asserted-by":"crossref","unstructured":"Fuchen Long Zhaofan Qiu Ting Yao and Tao Mei. 2024. Videostudio: Generating consistent-content and multi-scene videos. In European Conference on Computer Vision. Springer 468\u2013485.","key":"e_1_3_1_135_2","DOI":"10.1007\/978-3-031-73027-6_27"},{"unstructured":"Dustin Podell Zion English Kyle Lacey Andreas Blattmann Tim Dockhorn Jonas M\u00fcller Joe Penna and Robin Rombach. 2023. SDXL: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952 (2023).","key":"e_1_3_1_136_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_137_2","DOI":"10.1109\/CVPR52733.2024.00639"},{"doi-asserted-by":"publisher","unstructured":"Y. Wang X. Chen X. Ma et\u00a0al. 2025. LaVie: High-quality video generation with cascaded latent diffusion models. Int J Comput Vis 133 (2025) 3059\u20133078. 10.1007\/s11263-024-02295-1","key":"e_1_3_1_138_2","DOI":"10.1007\/s11263-024-02295-1"},{"unstructured":"Khurram Soomro Amir Roshan Zamir and Mubarak Shah. 2012. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012).","key":"e_1_3_1_139_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_140_2","DOI":"10.1109\/CVPR.2017.502"},{"unstructured":"Sami Abu-El-Haija Nisarg Kothari Joonseok Lee Paul Natsev George Toderici Balakrishnan Varadarajan and Sudheendra Vijayanarasimhan. 2016. YouTube-8M: A large-scale video classification benchmark. arXiv preprint arXiv:1609.08675 (2016).","key":"e_1_3_1_141_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_142_2","DOI":"10.1109\/ICCV.2019.00272"},{"doi-asserted-by":"publisher","key":"e_1_3_1_143_2","DOI":"10.1109\/CVPR.2016.571"},{"doi-asserted-by":"publisher","key":"e_1_3_1_144_2","DOI":"10.1109\/ICCV48922.2021.00175"},{"doi-asserted-by":"publisher","key":"e_1_3_1_145_2","DOI":"10.18653\/v1\/P18-1238"},{"key":"e_1_3_1_146_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Wang Yi","year":"2024","unstructured":"Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, et\u00a0al. 2024. InternVid: A large-scale video-text dataset for multimodal understanding and generation. In Proceedings of the 12th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=MLBdiWu4Fw"},{"unstructured":"Muhammad Maaz Hanoona Rasheed Salman Khan and Fahad Shahbaz Khan. 2023. Video-ChatGPT: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424 (2023).","key":"e_1_3_1_147_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_148_2","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"e_1_3_1_149_2","volume-title":"Proceedings of the 38th Conference on Neural Information Processing Systems Datasets and Benchmarks Track","author":"Ju Xuan","year":"2024","unstructured":"Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xintao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. 2024. MiraData: A large-scale video dataset with long durations and structured captions. In Proceedings of the 38th Conference on Neural Information Processing Systems Datasets and Benchmarks Track. Retrieved from https:\/\/openreview.net\/forum?id=2myGfVgfva"},{"unstructured":"Joao Carreira Eric Noland Andras Banki-Horvath Chloe Hillier and Andrew Zisserman. 2018. A short note about kinetics-600. arXiv preprint arXiv:1808.01340 (2018).","key":"e_1_3_1_150_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_151_2","DOI":"10.1109\/CVPR52733.2024.01265"},{"unstructured":"Wenjing Wang Huan Yang Zixi Tuo Huiguo He Junchen Zhu Jianlong Fu and Jiaying Liu. 2023. VideoFactory: Swap attention in spatiotemporal diffusions for text-to-video generation. arXiv preprint arXiv:2305.10874.","key":"e_1_3_1_152_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_153_2","DOI":"10.52202\/079017-2096"},{"doi-asserted-by":"publisher","unstructured":"Kristen Grauman Andrew Westbury Eugene Byrne Vincent Cartillier Zachary Chavis Antonino Furnari Rohit Girdhar Jackson Hamburger Hao Jiang Devansh Kukreja Miao Liu Xingyu Liu Miguel Martin Tushar Nagarajan Ilija Radosavovic Santhosh Kumar Ramakrishnan Fiona Ryan Jayant Sharma Michael Wray Mengmeng Xu Eric Zhongcong Xu Chen Zhao Siddhant Bansal Dhruv Batra Sean Crane Tien Do Morrie Doulaty Akshay Erapalli Christoph Feichtenhofer Adriano Fragomeni Qichen Fu Abrham Gebreselasie Cristina Gonzalez James Hillis Xuhua Huang Yifei Huang Wenqi Jia Weslie Khoo Jachym Kolar Satwik Kottur Anurag Kumar Federico Landini Chao Li Yanghao Li Zhenqiang Li Karttikeya Mangalam Raghava Modhugu Jonathan Munro Tullie Murrell Takumi Nishiyasu Will Price Paola Ruiz Puentes Merey Ramazanova Leda Sari Kiran Somasundaram Audrey Southerland Yusuke Sugano Ruijie Tao Minh Vo Yuchen Wang Xindi Wu Takuma Yagi Ziwei Zhao Yunyi Zhu Pablo Arbelaez David Crandall Dima Damen Giovanni Maria Farinella Christian Fuegen Bernard Ghanem Vamsi Krishna C. V. Jawahar Hanbyul Joo Kris Kitani Haizhou Li Richard Newcombe Aude Oliva Hyun Soo Park James M. Rehg Yoichi Sato Jianbo Shi Mike Zheng Shou Antonio Torralba Lorenzo Torresani Mingfei Yan and Jitendra Malik. 2025. Ego4D: Around the world in 3 600 hours of egocentric video. IEEE Transactions on Pattern Analysis & Machine Intelligence 47 11 (November 2025) 9468\u20139509. DOI:10.1109\/TPAMI.2024.3381075","key":"e_1_3_1_154_2","DOI":"10.1109\/TPAMI.2024.3381075"},{"doi-asserted-by":"crossref","unstructured":"Dingyi Yang Chunru Zhan Ziheng Wang Biao Wang Tiezheng Ge Bo Zheng and Qin Jin. 2024. Synchronized video storytelling: Generating video narrations with structured storyline. arXiv preprint arXiv:2405.14040 (2024).","key":"e_1_3_1_155_2","DOI":"10.18653\/v1\/2024.acl-long.513"},{"doi-asserted-by":"crossref","unstructured":"Zhichao Zhang Wei Sun Li Xinyue Jun Jia Xiongkuo Min Zicheng Zhang Chunyi Li Zijian Chen Wang Puyi Sun Fengyu and others. 2025. Benchmarking multi-dimensional AIGC video quality assessment: A dataset and unified model. ACM Transactions on Multimedia Computing Communications and Applications 21 9 (2025) 1\u201324.","key":"e_1_3_1_156_2","DOI":"10.1145\/3749844"},{"doi-asserted-by":"publisher","key":"e_1_3_1_157_2","DOI":"10.5555\/3157096.3157346"},{"doi-asserted-by":"publisher","key":"e_1_3_1_158_2","DOI":"10.1109\/CVPR.2016.308"},{"doi-asserted-by":"publisher","key":"e_1_3_1_159_2","DOI":"10.5555\/3295222.3295408"},{"doi-asserted-by":"publisher","key":"e_1_3_1_160_2","DOI":"10.1109\/ICCV51070.2023.01843"},{"doi-asserted-by":"publisher","key":"e_1_3_1_161_2","DOI":"10.1007\/978-3-030-58536-5_24"},{"doi-asserted-by":"crossref","unstructured":"Jack Hessel Ari Holtzman Maxwell Forbes Ronan Le Bras and Yejin Choi. 2021. CLIPScore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718 (2021).","key":"e_1_3_1_162_2","DOI":"10.18653\/v1\/2021.emnlp-main.595"},{"unstructured":"Chenfei Wu Lun Huang Qianxi Zhang Binyang Li Lei Ji Fan Yang Guillermo Sapiro and Nan Duan. 2021. GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions. arXiv:2104.14806. Retrieved from https:\/\/arxiv.org\/abs\/2104.14806","key":"e_1_3_1_163_2"},{"doi-asserted-by":"crossref","unstructured":"Jialian Wu Jianfeng Wang Zhengyuan Yang Zhe Gan Zicheng Liu Junsong Yuan and Lijuan Wang. 2024. GRiT: A generative region-to-text transformer for object understanding. In European Conference on Computer Vision. Springer 207\u2013224.","key":"e_1_3_1_164_2","DOI":"10.1007\/978-3-031-72989-8_12"},{"key":"e_1_3_1_165_2","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923)","author":"Liu Yuanxin","year":"2024","unstructured":"Yuanxin Liu, Lei Li, Shuhuai Ren, Rundong Gao, Shicheng Li, Sishuo Chen, Xu Sun, and Lu Hou. 2024. FETV: A benchmark for fine-grained evaluation of open-domain text-to-video generation. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923). Curran Associates Inc., Red Hook, NY, USA, Article 2723, 36 pages."},{"doi-asserted-by":"publisher","key":"e_1_3_1_166_2","DOI":"10.1109\/CVPR52733.2024.02060"},{"doi-asserted-by":"crossref","unstructured":"Tengchuan Kou Xiaohong Liu Zicheng Zhang Chunyi Li Haoning Wu Xiongkuo Min Guangtao Zhai and Ning Liu. 2024. Subjective-aligned dataset and metric for text-to-video quality assessment. In Proceedings of the 32nd ACM International Conference on Multimedia. 7793\u20137802.","key":"e_1_3_1_167_2","DOI":"10.1145\/3664647.3680868"},{"key":"e_1_3_1_168_2","volume-title":"Proceedings of the 1st Workshop on Controllable Video Generation @ICML24","author":"Liu Jiahe","year":"2024","unstructured":"Jiahe Liu, Youran Qu, Qi Yan, Xiaohui Zeng, Lele Wang, and Renjie Liao. 2024. Fr\u00e9chet video motion distance: A metric for evaluating motion consistency in videos. In Proceedings of the 1st Workshop on Controllable Video Generation @ICML24. Retrieved from https:\/\/openreview.net\/forum?id=tTZ2eAhK9D"},{"doi-asserted-by":"crossref","unstructured":"Kaiyue Sun Kaiyi Huang Xian Liu Yue Wu Zihan Xu Zhenguo Li and Xihui Liu. 2025. T2V-compbench: A comprehensive benchmark for compositional text-to-video generation. In Proceedings of the Computer Vision and Pattern Recognition Conference. 8406\u20138416.","key":"e_1_3_1_169_2","DOI":"10.1109\/CVPR52734.2025.00787"},{"unstructured":"Jiaxuan Guo Yichun Li Shangzhe Wang Yinan Zhang Xihui Liu Yu Wang Hanyang Yang Jing Yang and Ziwei Liu. 2023. VideoPoet: A large language model for zero-shot video generation. arXiv:2312.14125. Retrieved from https:\/\/arxiv.org\/abs\/2312.14125","key":"e_1_3_1_170_2"},{"unstructured":"Uriel Singer Adam Polyak Thomas Hayes Xi Yin Junjie An Songyang Zhang Qiyuan Hu Oran Yang Omri Ashual Oran Gafni and others. 2022. Make-A-Video: Text-to-video generation without Text-Video Data. arXiv:2209.14792 (2022).","key":"e_1_3_1_171_2"},{"unstructured":"Yingqing He Tianyu Yang Yong Zhang Ying Shan and Qifeng Chen. 2022. Latent video diffusion models for high-fidelity long video generation. arXiv preprint arXiv:2211.13221 (2022).","key":"e_1_3_1_172_2"},{"unstructured":"Junnan Li Dongxu Li Caiming Xiong and Steven Hoi. 2023. VLogger: Generating Long Videos of Dynamic Human Activities. arXiv:2306.04308. Retrieved from https:\/\/arxiv.org\/abs\/2306.04308","key":"e_1_3_1_173_2"},{"unstructured":"Yinan He Yaohui Wang Ceyuan Yang Shangchen Zhou Xiangyu Zhang Xiaodong Yang Yu Qiao Dahua Lin and Ying Shan. 2023. MicroCinema: A divide-and-conquer approach for text-to-video generation. arXiv:2312.04889. Retrieved from https:\/\/arxiv.org\/abs\/2312.04889","key":"e_1_3_1_174_2"},{"unstructured":"Han Lin Abhay Zala Jaemin Cho and Mohit Bansal. 2023. VideoDirectorGPT: Consistent multi-scene video generation via LLM-guided planning. arXiv preprint arXiv:2309.15091 (2023).","key":"e_1_3_1_175_2"},{"unstructured":"Jonathan Ho Tim Salimans Alexey Gritsenko William Chan Mohammad Norouzi and David J. Fleet. 2022. Video diffusion models. Advances in Neural Information Processing Systems 35 (2022) 8633\u20138646.","key":"e_1_3_1_176_2"},{"unstructured":"Chenfei Wu Jian Liang Lei Ji Fan Yang Yuejian Fang Daxin Jiang and Nan Duan. 2021. NUWA: Visual synthesis pre-training for neural visual world creation. arXiv:2111.12417. Retrieved from https:\/\/arxiv.org\/abs\/2111.12417","key":"e_1_3_1_177_2"},{"doi-asserted-by":"crossref","unstructured":"Masaki Saito Eiichi Matsumoto and Shunta Saito. 2017. Temporal generative adversarial nets with singular value clipping. In Proceedings of the IEEE International Conference on Computer Vision. 2830\u20132839.","key":"e_1_3_1_178_2","DOI":"10.1109\/ICCV.2017.308"},{"key":"e_1_3_1_179_2","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Huang Qingqiu","year":"2021","unstructured":"Qingqiu Huang, Wentao Yu, Yuanze Xu, Yitong Wang, and Dacheng Zhang. 2021. LVT: Language-vision transformer for multi-modal video understanding. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)."},{"key":"e_1_3_1_180_2","volume-title":"Proceedings of the NeurIPS","author":"Vondrick Carl","year":"2016","unstructured":"Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. 2016. Generating videos with scene dynamics. In Proceedings of the NeurIPS."},{"doi-asserted-by":"crossref","unstructured":"Zhengxiong Luo Dayou Chen Yingya Zhang Yan Huang Liangsheng Wang Yujun Shen Deli Zhao Jinren Zhou and Tien-Ping Tan. 2023. VideoFusion: Decomposed diffusion models for high-quality video generation. 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201923). 10209\u201310218. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:257532642","key":"e_1_3_1_181_2","DOI":"10.1109\/CVPR52729.2023.00984"},{"unstructured":"Jonathan Ho Tim Salimans Alexey Gritsenko William Chan Mohammad Norouzi and David J. Fleet. 2022. Video diffusion models. Advances in Neural Information Processing Systems 35 (2022) 8633\u20138646.","key":"e_1_3_1_182_2"},{"unstructured":"Emiel Hoogeboom Jonathan Heek and Tim Salimans. 2023. simple diffusion: End-to-end diffusion for high resolution images. In International Conference on Machine Learning. PMLR 13213\u201313232.","key":"e_1_3_1_183_2"},{"unstructured":"Andreas Blattmann Robin Rombach Huan Ling Tim Dockhorn Seung Wook Kim Sanja Fidler and Karsten Kreis. 2023. PYoCo: Latent diffusion priors for zero-shot video editing. arXiv:2303.04734. Retrieved from https:\/\/arxiv.org\/abs\/2303.04734","key":"e_1_3_1_184_2"},{"unstructured":"Wenjing Wang Huan Yang Zixi Tuo Huiguo He Junchen Zhu Jianlong Fu and Jiaying Liu. 2025. Swap attention in spatiotemporal diffusions for text-to-video generation. International Journal of Computer Vision (2025). 1\u201319.","key":"e_1_3_1_185_2"},{"unstructured":"Jiuniu Wang Hangjie Yuan Dayou Chen Yingya Zhang Xiang Wang and Shiwei Zhang. 2023. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571 (2023).","key":"e_1_3_1_186_2"},{"unstructured":"Julio J. Vald\u00e9s and Alain B. Tchagang. 2023. Understanding the structure of qm7b and qm9 quantum mechanical datasets using unsupervised learning. arXiv preprint arXiv:2309.15130 (2023).","key":"e_1_3_1_187_2"},{"doi-asserted-by":"crossref","unstructured":"Andreas Blattmann Robin Rombach Huan Ling Tim Dockhorn Seung Wook Kim Sanja Fidler and Karsten Kreis. 2023. VideoLDM: High-resolution video generation with latent diffusion models. arXiv:2304.08818. Retrieved from https:\/\/arxiv.org\/abs\/2304.08818","key":"e_1_3_1_188_2","DOI":"10.1109\/CVPR52729.2023.02161"},{"unstructured":"Yaohui Wang Xinyuan Chen Xin Ma Shangchen Zhou Ziwei Huang Yu Wang Ceyuan Yang Yinan He Jiashuo Yu Yinan Yang et\u00a0al. 2023. InternVid: A large-scale video-text dataset for multimodal understanding and generation. arXiv:2307.06942. Retrieved from https:\/\/arxiv.org\/abs\/2307.06942","key":"e_1_3_1_189_2"},{"unstructured":"Yan Zeng Guoqiang Wei Jiani Zheng Jiaxin Zou Yang Wei Yuchen Zhang and Hang Li. 2023. Make Pixels Dance: High-Dynamic Video Generation. arXiv:2311.10982. Retrieved from https:\/\/arxiv.org\/abs\/2311.10982","key":"e_1_3_1_190_2"},{"doi-asserted-by":"crossref","unstructured":"Rohit Girdhar Mannat Singh Andrew Brown Quentin Duval Samaneh Azadi Sai Saketh Rambhatla Akbar Shah Xi Yin Devi Parikh and Ishan Misra. 2024. Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning. arXiv:2311.10709. Retrieved from https:\/\/arxiv.org\/abs\/2311.10709","key":"e_1_3_1_191_2","DOI":"10.1007\/978-3-031-73033-7_12"},{"unstructured":"Omer Bar-Tal Dolev Ofri-Amar Rafail Fridman Yoni Kasten and Tali Dekel. 2024. Lumiere: A Space-Time Diffusion Model for Video Generation. arXiv:2401.12945. Retrieved from https:\/\/arxiv.org\/abs\/2401.12945","key":"e_1_3_1_192_2"},{"unstructured":"Yujie Lu Xianjun Yang Xiujun Li Xin Eric Wang and William Yang Wang. 2023. LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation. arXiv:2305.11116. Retrieved from https:\/\/arxiv.org\/abs\/2305.11116","key":"e_1_3_1_193_2"},{"unstructured":"Yaohui Wang Xinyuan Chen Xin Ma Shangchen Zhou Ziqi Huang Yi Wang Ceyuan Yang Yinan He Jiashuo Yu Peiqing Yang et\u00a0al. 2023. LAVIE: High-quality video generation with cascaded latent diffusion models. arXiv:2309.15103. Retrieved from https:\/\/arxiv.org\/abs\/2309.15103","key":"e_1_3_1_194_2"},{"unstructured":"Jiuniu Wang Hangjie Yuan Dayou Chen Yingya Zhang Xiang Wang and Shiwei Zhang. 2023. Modelscope text-to-video technical report. arXiv:2308.06571. Retrieved from https:\/\/arxiv.org\/abs\/2308.06571","key":"e_1_3_1_195_2"},{"doi-asserted-by":"crossref","unstructured":"Haoxin Chen Yong Zhang Xiaodong Cun Menghan Xia Xintao Wang Chao Weng and Ying Shan. 2024. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models. arXiv:2401.09047. Retrieved from https:\/\/arxiv.org\/abs\/2401.09047","key":"e_1_3_1_196_2","DOI":"10.1109\/CVPR52733.2024.00698"},{"unstructured":"2024. Gen-3. Retrieved June 17 2024 from https:\/\/runwayml.com\/research\/introducing-gen-3-alpha","key":"e_1_3_1_197_2"},{"unstructured":"Yuwei Guo Ceyuan Yang Anyi Rao Yaohui Wang Y. Qiao Dahua Lin and Bo Dai. 2023. AnimateDiff: Animate your personalized text-to-image diffusion models without specific tuning. ArXiv abs\/2307.04725 (2023). Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:259501509","key":"e_1_3_1_198_2"},{"unstructured":"Xin Ma Yaohui Wang Gengyun Jia Xinyuan Chen Ziwei Liu Yuan-Fang Li Cunjian Chen and Yu Qiao. 2024. Latte: Latent diffusion transformer for video generation. arXiv:2401.03048. Retrieved from https:\/\/arxiv.org\/abs\/2401.03048","key":"e_1_3_1_199_2"},{"unstructured":"2023. Pika Labs. Retrieved September 25 2023 from https:\/\/www.pika.art\/","key":"e_1_3_1_200_2"},{"unstructured":"2024. Kling. Retrieved June 6 2024 from https:\/\/klingai.kuaishou.com\/","key":"e_1_3_1_201_2"},{"unstructured":"Zhuoyi Yang Jiayan Teng Wendi Zheng Ming Ding Shiyu Huang Jiazheng Xu Yuanming Yang Wenyi Hong Xiaohan Zhang Guanyu Feng et\u00a0al. 2024. CogVideoX: Text-to-video diffusion models with an expert transformer. arXiv:2408.06072. Retrieved from https:\/\/arxiv.org\/abs\/2408.06072","key":"e_1_3_1_202_2"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3771724","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T14:31:11Z","timestamp":1765031471000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3771724"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,6]]},"references-count":201,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3771724"],"URL":"https:\/\/doi.org\/10.1145\/3771724","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"type":"print","value":"0360-0300"},{"type":"electronic","value":"1557-7341"}],"subject":[],"published":{"date-parts":[[2025,12,6]]},"assertion":[{"value":"2024-12-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-29","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-06","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}