{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T02:22:03Z","timestamp":1785291723255,"version":"3.55.0"},"reference-count":46,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T00:00:00Z","timestamp":1782259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>Motion customization plays a pivotal role in video generation by preserving the original appearance and context while adhering to specific motion patterns. In contrast, video generation techniques often lack coherence and realism due to difficulties in capturing and transferring motion patterns. Building upon the Video Motion Customization (VMC) framework, we proposed a few-shot learning approach using our unified Multi-Head Temporal Attention (MHTA) module for motion customization in text-to-video diffusion models. This significantly reduces computational requirements while maintaining and improving motion quality. Our model provides a streamlined mechanism for motion distillation while maintaining separate self-, cross-, and temporal attention. Moreover, the temporal attention layer is adapted through a simplified mechanism with efficient Q\/K\/V projections, while maintaining fixed spatial self- and cross-attention. The model distills a ground-truth motion vector from consecutive frames to align the predicted and ground-truth motion. Our proposed MHTA model outperforms the baseline in video generation using motion customization while being significantly more resource-efficient. Moreover, our approach can easily be applied to generate conditional prompt-based videos in the gaming industry.<\/jats:p>","DOI":"10.1145\/3793552","type":"journal-article","created":{"date-parts":[[2026,1,27]],"date-time":"2026-01-27T13:35:32Z","timestamp":1769520932000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["CP-Diffusion: Conditional Prompt-Based Diffusion Models for Video Generation"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-4462-9863","authenticated-orcid":false,"given":"Muhammad","family":"Saeed","sequence":"first","affiliation":[{"name":"School of Electrical Engineering and Computer Science, University of Ottawa, Ottawa, Ontario, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8020-3590","authenticated-orcid":false,"given":"Mustaqeem","family":"Khan","sequence":"additional","affiliation":[{"name":"College of Information Technology, United Arab Emirates University, Al Ain, UAE"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-9846-8991","authenticated-orcid":false,"given":"Muhammad","family":"Saad","sequence":"additional","affiliation":[{"name":"Computer Vision, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, United Arab Emirates"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6274-2593","authenticated-orcid":false,"given":"Nasir","family":"Rahim","sequence":"additional","affiliation":[{"name":"School of Computing, Gachon University, Seongnam, Republic of Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6490-4648","authenticated-orcid":false,"given":"Wail","family":"Gueaieb","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering and Computer Science, University of Ottawa, Ottawa, Ontario, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7690-8547","authenticated-orcid":false,"given":"Abdulmotaleb","family":"El Saddik","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering and Computer Science, University of Ottawa, Ottawa, Ontario, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"issue":"2","key":"e_1_3_2_2_2","first-page":"1","article-title":"Generating virtual wire sculptural art from 3D models","volume":"18","author":"Aeh Chih-Kuo","year":"2022","unstructured":"Chih-Kuo Aeh, Thi-Ngoc-Hanh Le, Zhi-Ying Hou, and Tong-Yee Lee. 2022. Generating virtual wire sculptural art from 3D models. ACM Transactions on Multimedia Computing, Communications, and Applications 18, 2 (2022), 1\u201323.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_2_3_2","first-page":"10853","volume-title":"Proceedings of the 32nd ACM International Conference on Multimedia","author":"Aun Mingzhen","year":"2024","unstructured":"Mingzhen Aun, Weining Wang, Yanyuan Qiao, Jiahui Sun, Zihan Qin, Longteng Guo, Xinxin Zhu, and Jing Liu. 2024. MM-LDM: Multi-modal latent diffusion model for sounding video generation. In Proceedings of the 32nd ACM International Conference on Multimedia, 10853\u201310861."},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00175"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02161"},{"key":"e_1_3_2_6_2","unstructured":"Minwoo Byeon Beomhee Park Haecheon Kim Sungjun Lee Woonhyuk Baek and Saehoon Kim. 2022. COYO-700M: Image-text pair dataset. Retrieved from https:\/\/github.com\/kakaobrain\/coyo-dataset"},{"key":"e_1_3_2_7_2","first-page":"1","volume-title":"ACM Transactions on Multimedia Computing, Communications and Applications","volume":"21","author":"Cai Haoyu","year":"2025","unstructured":"Haoyu Cai, Wenqi Lou, Chao Wang, and Xuehai Zhou. 2025. Picasso: Analyzing prompt design for text-to-image generative diffusion models from a temporal-spatial perspective. ACM Transactions on Multimedia Computing, Communications and Applications 21, 11 (2025), 1\u201324."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3680637"},{"key":"e_1_3_2_9_2","unstructured":"Weifeng Chen Yatai Ji Jie Wu Hefeng Wu Xuefeng Xiao and Liang Lin. 2023. Control-a-video: Controllable text-to-video generation with diffusion models. arXiv:2305.13840. Retrieved from https:\/\/arxiv.org\/abs\/2305.13840"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00882"},{"issue":"9","key":"e_1_3_2_11_2","first-page":"1","article-title":"Text-driven video prediction","volume":"20","author":"Cong Xue","year":"2024","unstructured":"Xue Cong, Jingjing Chen, Bin Zhu, and Yu-Gang Jiang. 2024. Text-driven video prediction. ACM Transactions on Multimedia Computing, Communications and Applications 20, 9 (2024), 1\u201315.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"issue":"7","key":"e_1_3_2_12_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3713075","article-title":"Unleashing creativity in the metaverse: Generative AI and multimodal content","volume":"21","author":"El Saddik Abdulmotaleb","year":"2024","unstructured":"Abdulmotaleb El Saddik, Jamil Ahmad, Mustaqeem Khan, Saad Abouzahir, and Wail Gueaieb. 2024. Unleashing creativity in the metaverse: Generative AI and multimodal content. ACM Transactions on Multimedia Computing, Communications, and Applications 21, 7 (2024), 1\u201343.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00675"},{"issue":"2","key":"e_1_3_2_14_2","first-page":"1","article-title":"Moment is important: Language-based video moment retrieval via adversarial learning","volume":"18","author":"Feng Yawen","year":"2022","unstructured":"Yawen Feng, Da Cao, Shaofei Lu, Hanling Zhang, Jiao Xu, and Zheng Qin. 2022. Moment is important: Language-based video moment retrieval via adversarial learning. ACM Transactions on Multimedia Computing, Communications, and Applications 18, 2 (2022), 1\u201321.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.02096"},{"key":"e_1_3_2_16_2","first-page":"6840","article-title":"Denoising diffusion probabilistic models","volume":"33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 33, 6840\u20136851.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_17_2","first-page":"8633","article-title":"Video diffusion models","volume":"35","author":"Ho Jonathan","year":"2022","unstructured":"Jonathan Ho, Tim Salimans, Alexey Gritsenko, and David J. Fleet. 2022. Video diffusion models. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 35, 8633\u20138646.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"issue":"4","key":"e_1_3_2_18_2","first-page":"695","article-title":"Estimation of non-normalized statistical models by score matching","volume":"6","author":"Hyv\u00e4rinen Aapo","year":"2005","unstructured":"Aapo Hyv\u00e4rinen and Peter Dayan. 2005. Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research 6, 4 (2005), 695\u2013709.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00701"},{"issue":"2","key":"e_1_3_2_20_2","first-page":"1","article-title":"Automatic comic generation with stylistic multi-page layouts and emotion-driven text balloon generation","volume":"17","author":"Jang Xin","year":"2021","unstructured":"Xin Jang, Zongliang Ma, Letian Yu, Ying Cao, Baocai Yin, Xiaopeng Wei, Qiang Zhang, and Rynson W. H. Lau. 2021. Automatic comic generation with stylistic multi-page layouts and emotion-driven text balloon generation. ACM Transactions on Multimedia Computing, Communications, and Applications 17, 2 (2021), 1\u201319.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_2_21_2","first-page":"9212","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Jeong Hyeonho","year":"2024","unstructured":"Hyeonho Jeong, Geon Yeong Park, and Jong Chul Ye. 2024. VMC: Video motion customization using temporal attention adaption for text-to-video diffusion models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9212\u20139221."},{"key":"e_1_3_2_22_2","first-page":"22623","volume-title":"Proceedings of the 2023 IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Karras Johanna","year":"2023","unstructured":"Johanna Karras, Aleksander Holynski, Ting-Chun Wang, and Ira Kemelmacher-Shlizerman. 2023. DreamPose: Fashion image-to-video synthesis via stable diffusion. In Proceedings of the 2023 IEEE\/CVF International Conference on Computer Vision (ICCV). IEEE, 22623\u201322633."},{"key":"e_1_3_2_23_2","unstructured":"Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv:1711.05101. Retrieved from https:\/\/arxiv.org\/abs\/1711.05101"},{"key":"e_1_3_2_24_2","unstructured":"Jian Ma Junhao Liang and Haonan Lu. 2023. Subject-diffusion: Open domain personalized text-to-image generation without test-time fine-tuning. arXiv:2307.11410. Retrieved from https:\/\/arxiv.org\/abs\/2307.11410"},{"key":"e_1_3_2_25_2","unstructured":"Jordi Pont-Tuset Federico Perazzi Sergi Caelles and Luc Van Gool. 2017. The 2017 DAVIS challenge on video object segmentation. arXiv:1704.00675. Retrieved from https:\/\/arxiv.org\/abs\/1704.00675"},{"key":"e_1_3_2_26_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00624"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1833"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00816"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"5997","DOI":"10.1109\/CVPR52729.2023.00581","volume-title":"Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Singh Jaskirat","year":"2023","unstructured":"Jaskirat Singh, Stephen Gould, and Liang Zheng. 2023. High-fidelity guided image synthesis with latent diffusion models. In Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 5997\u20136006."},{"key":"e_1_3_2_32_2","first-page":"2256","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Sohl-Dickstein Jascha","year":"2015","unstructured":"Jascha Sohl-Dickstein, Eric Weiss, and Surya Ganguli. 2015. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the International Conference on Machine Learning. PMLR, 2256\u20132265."},{"key":"e_1_3_2_33_2","unstructured":"Jiaming Song Chenlin Meng and Stefano Ermon. 2020. Denoising diffusion implicit models. arXiv:2010.02502. Retrieved from https:\/\/arxiv.org\/abs\/2010.02502"},{"key":"e_1_3_2_34_2","article-title":"Generative modeling by estimating gradients of the data distribution","volume":"32","author":"Song Yang","year":"2019","unstructured":"Yang Song and Stefano Ermon. 2019. Generative modeling by estimating gradients of the data distribution. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 32.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_35_2","unstructured":"Yang Song Jascha Sohl-Dickstein Diederik P. Kingma and Ben Poole. 2020. Score-based generative modeling through stochastic differential equations. arXiv:2011.13456. Retrieved from https:\/\/arxiv.org\/abs\/2011.13456"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","first-page":"6561","DOI":"10.1109\/CVPR52733.2024.00627","volume-title":"Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Sou\u010dek Tom\u00e1\u0161","year":"2024","unstructured":"Tom\u00e1\u0161 Sou\u010dek, Dima Damen, Michael Wray, Ivan Laptev, and Josef Sivic. 2024. Genhowto: Learning to generate actions and state transformations from instructional videos. In Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 6561\u20136571."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681050"},{"key":"e_1_3_2_38_2","article-title":"VideoComposer: Compositional video synthesis with motion controllability","volume":"36","author":"Wang Xiang","year":"2024","unstructured":"Xiang Wang, Hangjie Yuan, Shiwei Zhang, Deli Zhao, and Jingren Zhou. 2024. VideoComposer: Compositional video synthesis with motion controllability. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_39_2","unstructured":"Yaohui Wang Xinyuan Chen Xin Ma Shangchen Zhou Ziqi Huang Yi Wang Ceyuan Yang Yinan He Jiashuo Yu Peiqing Yang et al. 2023. LAVIE: High-quality video generation with cascaded latent diffusion models. arXiv:2309.15103. Retrieved from https:\/\/arxiv.org\/abs\/2309.15103"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3691344"},{"key":"e_1_3_2_41_2","doi-asserted-by":"crossref","first-page":"180","DOI":"10.1007\/978-3-031-33380-4_14","volume-title":"Proceedings of the Pacific-Asia Conference on Knowledge Discovery and Data Mining","author":"Weber Tobias","year":"2023","unstructured":"Tobias Weber, Michael Ingrisch, Bernd Bischl, and David R\u00fcgamer. 2023. Cascaded latent diffusion models for high-resolution chest X-ray synthesis. In Proceedings of the Pacific-Asia Conference on Knowledge Discovery and Data Mining. Springer, 180\u2013191."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00701"},{"key":"e_1_3_2_43_2","unstructured":"Ruiqi Wu Liangyu Chen Tong Yang Chunle Guo Chongyi Li and Xiangyu Zhang. 2023. LAMP: Learn a motion pattern for few-shot-based video generation. arXiv:2310.10769. Retrieved from https:\/\/arxiv.org\/abs\/2310.10769"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626235"},{"key":"e_1_3_2_45_2","unstructured":"David Junhao Zhang Jay Zhangjie Wu and Mike Zheng Shou. 2023. Show-1: Marrying pixel and latent diffusion models for text-to-video generation. arXiv:2309.15818. Retrieved from https:\/\/arxiv.org\/abs\/2309.15818"},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","first-page":"7496","DOI":"10.1109\/CVPR52733.2024.00716","volume-title":"Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhangli Qilong","year":"2024","unstructured":"Qilong Zhangli, Jindong Jiang, Di Liu, Licheng Yu, Xiaoliang Dai, Ankit Ramchandani, Guan Pang, Dimitris N. Metaxas, and Praveen Krishnan. 2024. Layout-agnostic scene text image synthesis with diffusion models. In Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE Computer Society, 7496\u20137506."},{"key":"e_1_3_2_47_2","unstructured":"Rui Zhao Yuchao Gu Jay Zhangjie Wu and Mike Zheng Shou. 2023. MotionDirector: Motion customization of text-to-video diffusion models. arXiv:2310.08465. Retrieved from https:\/\/arxiv.org\/abs\/2310.08465"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3793552","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:46:57Z","timestamp":1782312417000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3793552"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":46,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3793552"],"URL":"https:\/\/doi.org\/10.1145\/3793552","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-06-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-17","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}