{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T13:48:05Z","timestamp":1782308885968,"version":"3.54.5"},"reference-count":172,"publisher":"Association for Computing Machinery (ACM)","issue":"13","license":[{"start":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T00:00:00Z","timestamp":1782259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272409"],"award-info":[{"award-number":["62272409"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,10,31]]},"abstract":"<jats:p>The burgeoning growth of video-to-music generation can be attributed to the ascendancy of multimodal generative models. However, there is a lack of literature that comprehensively combs through the work in this field. To fill this gap, this article presents a comprehensive review of video-to-music generation using deep generative AI techniques, focusing on three key components: conditioning input construction, conditioning mechanism, and music generation frameworks. We categorize existing approaches based on their designs for each component, clarifying the roles of different strategies. Preceding this, we provide a fine-grained categorization of video and music modalities, illustrating how different categories influence the design of components within the generation pipelines. Furthermore, we summarize available multimodal datasets and evaluation metrics while highlighting ongoing challenges in the field.<\/jats:p>","DOI":"10.1145\/3816020","type":"journal-article","created":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T11:36:18Z","timestamp":1782300978000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["A Comprehensive Survey on Generative AI for Video-to-Music Generation"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9908-136X","authenticated-orcid":false,"given":"Shulei","family":"Ji","sequence":"first","affiliation":[{"name":"Zhejiang University","place":["Hangzhou, China"]},{"name":"Innovation Center of Yangtze River Delta, Zhejiang University","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0760-6289","authenticated-orcid":false,"given":"Songruoyao","family":"Wu","sequence":"additional","affiliation":[{"name":"Zhejiang University","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5613-3262","authenticated-orcid":false,"given":"Zihao","family":"Wang","sequence":"additional","affiliation":[{"name":"Zhejiang University","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-3452-1641","authenticated-orcid":false,"given":"Shuyu","family":"Li","sequence":"additional","affiliation":[{"name":"Zhejiang University","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0778-2303","authenticated-orcid":false,"given":"Kejun","family":"Zhang","sequence":"additional","affiliation":[{"name":"Zhejiang University","place":["Hangzhou, China"]},{"name":"Innovation Center of Yangtze River Delta, Zhejiang University","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et\u00a0al. 2023. Gpt-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00785"},{"key":"e_1_3_3_4_2","unstructured":"Gunjan Aggarwal and Devi Parikh. 2021. Dance2music: Automatic dance-driven music generation. arXiv:2107.06252. Retrieved from https:\/\/arxiv.org\/abs\/2107.06252"},{"key":"e_1_3_3_5_2","unstructured":"Andrea Agostinelli Timo I. Denk Zal\u00e1n Borsos Jesse Engel Mauro Verzetti Antoine Caillon Qingqing Huang Aren Jansen Adam Roberts Marco Tagliasacchi et\u00a0al. 2023. Musiclm: Generating music from text. arXiv:2301.11325. Retrieved from https:\/\/arxiv.org\/abs\/2301.11325"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"e_1_3_3_7_2","first-page":"4","volume-title":"Proc. of the 16th Int. Conf. on Digital Audio Effects.","author":"B\u00f6ck Sebastian","year":"2013","unstructured":"Sebastian B\u00f6ck and Gerhard Widmer. 2013. Maximum filter vibrato suppression for onset detection. In Proc. of the 16th Int. Conf. on Digital Audio Effects. Citeseer, 4."},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2023.3288409"},{"key":"e_1_3_3_9_2","unstructured":"Zal\u00e1n Borsos Matt Sharifi Damien Vincent Eugene Kharitonov Neil Zeghidour and Marco Tagliasacchi. 2023. Soundstorm: Efficient parallel audio generation. arXiv:2305.09636. Retrieved from https:\/\/arxiv.org\/abs\/2305.09636"},{"key":"e_1_3_3_10_2","unstructured":"Jean-Pierre Briot Ga\u00ebtan Hadjeres and Fran\u00e7ois-David Pachet. 2017. Deep learning techniques for music generation\u2013a survey. arXiv:1709.01620. Retrieved from https:\/\/arxiv.org\/abs\/1709.01620"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2929257"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_3_13_2","unstructured":"Brandon Castellano. 2018. Pyscenedetect: Intelligent scene cut detection and video splitting tool. (2018)."},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460426.3463619"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3126686.3126723"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3009820"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-024-4231-5"},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10096514"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.52202\/075280-2066"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1285"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3672554"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1177\/20592043221117651"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201371"},{"key":"e_1_3_3_24_2","article-title":"High fidelity neural audio compression","volume":"2023","author":"D\u00e9fossez Alexandre","year":"2023","unstructured":"Alexandre D\u00e9fossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2023. High fidelity neural audio compression. Trans. Mach. Learn. Res. 2023 (2023).","journal-title":"Trans. Mach. Learn. Res."},{"key":"e_1_3_3_25_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers). Association for Computational Linguistics, 4171\u20134186."},{"key":"e_1_3_3_26_2","unstructured":"Prafulla Dhariwal Heewoo Jun Christine Payne Jong Wook Kim Alec Radford and Ilya Sutskever. 2020. Jukebox: A generative model for music. arXiv:2005.00341. Retrieved from https:\/\/arxiv.org\/abs\/2005.00341"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475195"},{"key":"e_1_3_3_28_2","first-page":"8000","volume-title":"Advances in Neural Information Processing Systems","author":"Dieleman Sander","year":"2018","unstructured":"Sander Dieleman, Aaron Van Den Oord, and Karen Simonyan. 2018. The challenge of realistic music generation: Modelling raw audio at scale. In Advances in Neural Information Processing Systems. 8000\u20138010."},{"key":"e_1_3_3_29_2","volume-title":"Proceedings of the 9th International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et\u00a0al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the 9th International Conference on Learning Representations. OpenReview.net."},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7953127"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49660.2025.10888461"},{"key":"e_1_3_3_32_2","unstructured":"Han Fang Pengfei Xiong Luhui Xu and Yu Chen. 2021. Clip2video: Mastering video-text retrieval via image clip. arXiv:2106.11097. Retrieved from https:\/\/arxiv.org\/abs\/2106.11097"},{"key":"e_1_3_3_33_2","article-title":"Riffusion-stable diffusion for real-time music generation","author":"Forsgren Seth","year":"2022","unstructured":"Seth Forsgren and Hayk Martiros. 2022. Riffusion-stable diffusion for real-time music generation. URL https:\/\/riffusion.com (2022).","journal-title":"URL"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58621-8_44"},{"key":"e_1_3_3_35_2","first-page":"359","volume-title":"Proceedings of the 24th International Society for Music Information Retrieval Conference","author":"Garc\u00eda Hugo Flores","year":"2023","unstructured":"Hugo Flores Garc\u00eda, Prem Seetharaman, Rithesh Kumar, and Bryan Pardo. 2023. VampNet: Music generation via masked acoustic token modeling. In Proceedings of the 24th International Society for Music Information Retrieval Conference. 359\u2013366."},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952261"},{"key":"e_1_3_3_37_2","first-page":"2269","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Gillick Jon","year":"2019","unstructured":"Jon Gillick, Adam Roberts, Jesse Engel, Douglas Eck, and David Bamman. 2019. Learning to groove with inverse sequence transformations. In Proceedings of the International Conference on Machine Learning. PMLR, 2269\u20132279."},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01457"},{"key":"e_1_3_3_39_2","first-page":"2672","volume-title":"Advances in Neural Information Processing Systems","author":"Goodfellow Ian J","year":"2014","unstructured":"Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in Neural Information Processing Systems. 2672\u20132680."},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.findings-acl.647"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-024-0417-1"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/58.1.83"},{"key":"e_1_3_3_43_2","first-page":"50","volume-title":"Proceedings of the 19th International Society for Music Information Retrieval Conference","author":"Hawthorne Curtis","year":"2018","unstructured":"Curtis Hawthorne, Erich Elsen, Jialin Song, Adam Roberts, Ian Simon, Colin Raffel, Jesse H. Engel, Sageev Oore, and Douglas Eck. 2018. Onsets and frames: Dual-objective piano transcription. In Proceedings of the 19th International Society for Music Information Retrieval Conference. 50\u201357."},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3108242"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952132"},{"key":"e_1_3_3_47_2","first-page":"6626","volume-title":"Advances in Neural Information Processing Systems","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems. 6626\u20136637."},{"key":"e_1_3_3_48_2","first-page":"6840","volume-title":"Advances in Neural Information Processing Systems","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems. 6840\u20136851."},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3206025.3206046"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(81)90024-2"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i1.16091"},{"key":"e_1_3_3_52_2","volume-title":"Proceedings of the 7th International Conference on Learning Representations","author":"Huang Cheng-Zhi Anna","year":"2019","unstructured":"Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M. Dai, Matthew D. Hoffman, Monica Dinculescu, and Douglas Eck. 2019. Music transformer: Generating music with long-term structure. In Proceedings of the 7th International Conference on Learning Representations. OpenReview.net."},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.52202\/075280-0894"},{"key":"e_1_3_3_54_2","first-page":"559","volume-title":"Proceedings of the 23rd International Society for Music Information Retrieval Conference","author":"Huang Qingqing","year":"2022","unstructured":"Qingqing Huang, Aren Jansen, Joonseok Lee, Ravi Ganti, Judith Yue Li, and Daniel P. W. Ellis. 2022. MuLan: A joint embedding of music audio and natural language. In Proceedings of the 23rd International Society for Music Information Retrieval Conference. 559\u2013566."},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413671"},{"key":"e_1_3_3_56_2","unstructured":"Shulei Ji Jing Luo and Xinyu Yang. 2020. A comprehensive survey on deep music generation: Multi-level representations algorithms evaluations and future directions. arXiv:2011.06801. Retrieved from https:\/\/arxiv.org\/abs\/2011.06801"},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v40i26.39378"},{"key":"e_1_3_3_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597493"},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2024.123640"},{"key":"e_1_3_3_60_2","unstructured":"Will Kay Joao Carreira Karen Simonyan Brian Zhang Chloe Hillier Sudheendra Vijayanarasimhan Fabio Viola Tim Green Trevor Back Paul Natsev et\u00a0al. 2017. The kinetics human action video dataset. arXiv:1705.06950. Retrieved from https:\/\/arxiv.org\/abs\/1705.06950"},{"key":"e_1_3_3_61_2","first-page":"518","volume-title":"Proceedings of the 26th International Society for Music Information Retrieval Conference","author":"Kim Haven","year":"2025","unstructured":"Haven Kim, Zachary Novack, Weihan Xu, Julian McAuley, and Hao-Wen Dong. 2025. Video-guided text-to-music generation using public domain movie collections. In Proceedings of the 26th International Society for Music Information Retrieval Conference. 518\u2013527."},{"key":"e_1_3_3_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"e_1_3_3_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053115"},{"key":"e_1_3_3_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2020.3030497"},{"key":"e_1_3_3_65_2","doi-asserted-by":"publisher","DOI":"10.52202\/075280-1214"},{"key":"e_1_3_3_66_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33012588"},{"key":"e_1_3_3_67_2","first-page":"1","volume-title":"Proceedings of the 2013 IEEE International Conference on Multimedia and Expo","author":"Kuo Fang-Fei","year":"2013","unstructured":"Fang-Fei Kuo, Man-Kwan Shan, and Suh-Yin Lee. 2013. Background music recommendation for video based on multimodal latent semantic analysis. In Proceedings of the 2013 IEEE International Conference on Multimedia and Expo. IEEE, 1\u20136."},{"key":"e_1_3_3_68_2","doi-asserted-by":"publisher","DOI":"10.1145\/3714457"},{"key":"e_1_3_3_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2856090"},{"key":"e_1_3_3_70_2","first-page":"12888","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning. PMLR, 12888\u201312900."},{"key":"e_1_3_3_71_2","first-page":"198","volume-title":"Proceedings of the International Conference on AI-Generated Content","author":"Li Jiajun","year":"2025","unstructured":"Jiajun Li, Tianze Xu, Xuesong Chen, Xinrui Yao, Jingchou Han, and Shuchang Liu. 2025. Mozart\u2019s touch: A lightweight multimodal music generation framework based on pre-trained large models. In Proceedings of the International Conference on AI-Generated Content. SPIE, 198\u2013207."},{"key":"e_1_3_3_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01315"},{"key":"e_1_3_3_73_2","unstructured":"Ruilong Li Shan Yang David A Ross and Angjoo Kanazawa. 2021. Learn to dance with aist++: Music conditioned 3d dance generation. arXiv:2101.08779. Retrieved from https:\/\/arxiv.org\/abs\/2101.08779"},{"key":"e_1_3_3_74_2","unstructured":"Ruiqi Li Siqi Zheng Xize Cheng Ziang Zhang Shengpeng Ji and Zhou Zhao. 2024. Muvi: Video-to-music generation with semantic alignment and rhythmic synchronization. arXiv:2410.12957. Retrieved from https:\/\/arxiv.org\/abs\/2410.12957"},{"key":"e_1_3_3_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/3680528.3687562"},{"key":"e_1_3_3_76_2","doi-asserted-by":"publisher","DOI":"10.1145\/3800682"},{"key":"e_1_3_3_77_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02582"},{"key":"e_1_3_3_78_2","unstructured":"Sifei Li Binxin Yang Chunji Yin Chong Sun Yuxin Zhang Weiming Dong and Chen Li. 2024. Vidmusician: Video-to-music generation with semantic-rhythmic alignment via hierarchical visual features. arXiv:2412.06296. Retrieved from https:\/\/arxiv.org\/abs\/2412.06296"},{"key":"e_1_3_3_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3405734"},{"key":"e_1_3_3_80_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123399"},{"key":"e_1_3_3_81_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV61041.2025.00120"},{"key":"e_1_3_3_82_2","article-title":"Towards video to piano music generation with chain-of-perform support benchmarks","author":"Liu Chang","year":"2025","unstructured":"Chang Liu, Haomin Zhang, Shiyu Xia, Zihao Chen, Chaofan Ding, Xin Yue, Huizhe Chen, and Xinhan Di. 2025. Towards video to piano music generation with chain-of-perform support benchmarks. CVPR 2025 MMFM Workshop (2025).","journal-title":"CVPR 2025 MMFM Workshop"},{"key":"e_1_3_3_83_2","series-title":"Proceedings of Machine Learning Research","first-page":"21450","volume-title":"International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA","volume":"202","author":"Liu Haohe","year":"2023","unstructured":"Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo P. Mandic, Wenwu Wang, and Mark D. Plumbley. 2023. AudioLDM: Text-to-audio generation with latent diffusion models. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA(Proceedings of Machine Learning Research, Vol. 202). PMLR, 21450\u201321474."},{"key":"e_1_3_3_84_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2024.3399607"},{"key":"e_1_3_3_85_2","unstructured":"Shansong Liu Atin Sakkeer Hussain Qilong Wu Chenshuo Sun and Ying Shan. 2023. M \\(^2\\) UGen: Multi-modal music understanding and generation with the power of large language models. arXiv:2311.11255. Retrieved from https:\/\/arxiv.org\/abs\/2311.11255"},{"key":"e_1_3_3_86_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2025.130688"},{"key":"e_1_3_3_87_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00702"},{"key":"e_1_3_3_88_2","unstructured":"Xiaohao Liu Teng Tu Yunshan Ma and Tat-Seng Chua. 2025. Extending visual dynamics for video-to-music generation. arXiv:2504.07594. Retrieved from https:\/\/arxiv.org\/abs\/2504.07594"},{"key":"e_1_3_3_89_2","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818013"},{"key":"e_1_3_3_90_2","doi-asserted-by":"publisher","DOI":"10.52202\/075280-2121"},{"key":"e_1_3_3_91_2","doi-asserted-by":"publisher","DOI":"10.3389\/fpubh.2022.992200"},{"key":"e_1_3_3_92_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP48485.2024.10446029"},{"issue":"18","key":"e_1_3_3_93_2","first-page":"7","article-title":"librosa: Audio and music signal analysis in python.","volume":"2015","author":"McFee Brian","year":"2015","unstructured":"Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. 2015. librosa: Audio and music signal analysis in python. SciPy 2015, 18\u201324 (2015), 7.","journal-title":"SciPy"},{"key":"e_1_3_3_94_2","first-page":"6754","volume-title":"Advances in Neural Information Processing Systems","author":"Mei Hongyuan","year":"2017","unstructured":"Hongyuan Mei and Jason M. Eisner. 2017. The neural hawkes process: A neurally self-modulating multivariate point process. In Advances in Neural Information Processing Systems. 6754\u20136764."},{"key":"e_1_3_3_95_2","doi-asserted-by":"publisher","DOI":"10.1109\/MLSP58920.2024.10734721"},{"key":"e_1_3_3_96_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.apergo.2020.103301"},{"key":"e_1_3_3_97_2","first-page":"7176","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Naeem Muhammad Ferjad","year":"2020","unstructured":"Muhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi, and Jaejun Yoo. 2020. Reliable fidelity and diversity metrics for generative models. In Proceedings of the International Conference on Machine Learning. PMLR, 7176\u20137185."},{"key":"e_1_3_3_98_2","unstructured":"Aaron van den Oord Yazhe Li and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv:1807.03748. Retrieved from https:\/\/arxiv.org\/abs\/1807.03748"},{"key":"e_1_3_3_99_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.264"},{"key":"e_1_3_3_100_2","first-page":"400","volume-title":"ISMIR","author":"Pedersoli Fabrizio","year":"2020","unstructured":"Fabrizio Pedersoli and Masataka Goto. 2020. Dance beat tracking from visual information alone. In ISMIR. 400\u2013408."},{"key":"e_1_3_3_101_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"e_1_3_3_102_2","unstructured":"Jordi Pons and Xavier Serra. 2019. musicnn: Pre-trained convolutional neural networks for music audio tagging. arXiv:1909.06654. Retrieved from https:\/\/arxiv.org\/abs\/1909.06654"},{"key":"e_1_3_3_103_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01381"},{"key":"e_1_3_3_104_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02227"},{"key":"e_1_3_3_105_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PmLR, 8748\u20138763."},{"key":"e_1_3_3_106_2","unstructured":"ITUT Recommendation. 1994. Telephone Transmission Quality Subjective Opinion Tests. A Method for Subjective Performance Assessment of the Quality of Speech Voice Output Devices."},{"key":"e_1_3_3_107_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_3_108_2","doi-asserted-by":"publisher","DOI":"10.1145\/3746027.3755758"},{"key":"e_1_3_3_109_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00985"},{"key":"e_1_3_3_110_2","first-page":"29441","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Ryali Chaitanya","year":"2023","unstructured":"Chaitanya Ryali, Yuan-Ting Hu, Daniel Bolya, Chen Wei, Haoqi Fan, Po-Yao Huang, Vaibhav Aggarwal, Arkabandhu Chowdhury, Omid Poursaeed, Judy Hoffman, et\u00a0al. 2023. Hiera: A hierarchical vision transformer without the bells-and-whistles. In Proceedings of the International Conference on Machine Learning. PMLR, 29441\u201329454."},{"key":"e_1_3_3_111_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2016.7552868"},{"key":"e_1_3_3_112_2","unstructured":"Flavio Schneider Ojasv Kamal Zhijing Jin and Bernhard Sch\u00f6lkopf. 2023. Mo\u2303usai: Text-to-music generation with long-context latent diffusion. arXiv:2301.11757. Retrieved from https:\/\/arxiv.org\/abs\/2301.11757"},{"key":"e_1_3_3_113_2","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654919"},{"key":"e_1_3_3_114_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.naacl-demo.19"},{"key":"e_1_3_3_115_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.494"},{"key":"e_1_3_3_116_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01077"},{"key":"e_1_3_3_117_2","first-page":"2256","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Sohl-Dickstein Jascha","year":"2015","unstructured":"Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the International Conference on Machine Learning. pmlr, 2256\u20132265."},{"key":"e_1_3_3_118_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i5.28299"},{"key":"e_1_3_3_119_2","first-page":"3325","article-title":"Audeo: Audio generation for a silent performance video","volume":"33","author":"Su Kun","year":"2020","unstructured":"Kun Su, Xiulong Liu, and Eli Shlizerman. 2020. Audeo: Audio generation for a silent performance video. Advances in Neural Information Processing Systems 33 (2020), 3325\u20133337.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_120_2","unstructured":"Kun Su Xiulong Liu and Eli Shlizerman. 2020. Multi-instrumentalist net: Unsupervised generation of music from body movements. arXiv:1409.0473. Retrieved from https:\/\/arxiv.org\/abs\/1701.00133"},{"key":"e_1_3_3_121_2","first-page":"29258","article-title":"How does it sound?","volume":"34","author":"Su Kun","year":"2021","unstructured":"Kun Su, Xiulong Liu, and Eli Shlizerman. 2021. How does it sound? Advances in Neural Information Processing Systems 34 (2021), 29258\u201329273.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_122_2","doi-asserted-by":"crossref","unstructured":"Serkan Sulun Paula Viana and Matthew E. P. Davies. 2025. Video soundtrack generation by aligning emotions and temporal boundaries. arXiv:1409.0473. Retrieved from https:\/\/arxiv.org\/abs\/1701.00133","DOI":"10.1109\/TMM.2026.3701535"},{"key":"e_1_3_3_123_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00779"},{"key":"e_1_3_3_124_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3680889"},{"key":"e_1_3_3_125_2","first-page":"10564","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Sur\u00eds D\u00eddac","year":"2022","unstructured":"D\u00eddac Sur\u00eds, Carl Vondrick, Bryan Russell, and Justin Salamon. 2022. It\u2019s time for artistic correspondence in music and video. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10564\u201310574."},{"key":"e_1_3_3_126_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.308"},{"key":"e_1_3_3_127_2","doi-asserted-by":"publisher","DOI":"10.1145\/3610543.3626164"},{"key":"e_1_3_3_128_2","doi-asserted-by":"publisher","DOI":"10.52202\/075280-0707"},{"key":"e_1_3_3_129_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2025.3590912"},{"key":"e_1_3_3_130_2","unstructured":"Zeyue Tian Yizhu Jin Zhaoyang Liu Ruibin Yuan Xu Tan Qifeng Chen Wei Xue and Yike Guo. 2025. Audiox: Diffusion transformer for anything-to-audio generation. arXiv:2503.10522. Retrieved from https:\/\/arxiv.org\/abs\/2503.10522"},{"key":"e_1_3_3_131_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01750"},{"key":"e_1_3_3_132_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSS.2024.3451515"},{"key":"e_1_3_3_133_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v40i31.39799"},{"key":"e_1_3_3_134_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-0732"},{"key":"e_1_3_3_135_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et\u00a0al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_3_136_2","volume-title":"Proceedings of the ISMIR","author":"Tsuchida Shuhei","year":"2019","unstructured":"Shuhei Tsuchida, Satoru Fukayama, Masahiro Hamasaki, and Masataka Goto. 2019. AIST dance video database: Multi-genre, multi-dancer, and multi-camera database for dance information processing. In Proceedings of the ISMIR."},{"key":"e_1_3_3_137_2","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_3_138_2","unstructured":"Baisen Wang Le Zhuo Zhaokai Wang Chenxi Bao Wu Chengjing Xuecheng Nie Jiao Dai Jizhong Han Yue Liao and Si Liu. 2024. Multimodal music generation with explicit bridges and retrieval augmentation. arXiv:2412.09428. Retrieved from https:\/\/arxiv.org\/abs\/2412.09428"},{"key":"e_1_3_3_139_2","unstructured":"Jinting Wang Chenxing Li and Li Liu. 2025. GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment. arXiv:2510.26818. Retrieved from https:\/\/arxiv.org\/abs\/2510.26818"},{"key":"e_1_3_3_140_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49660.2025.10889094"},{"key":"e_1_3_3_141_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01398"},{"key":"e_1_3_3_142_2","volume-title":"Proceedings of the ISMIR 2025 Hybrid Conference","author":"Wang Zhaokai","year":"2025","unstructured":"Zhaokai Wang, Chenxi Bao, Le Zhuo, Jingrui Han, Yang Yue, Yihong Tang, Victor Shea-Jay Huang, and Yue Liao. 2025. A survey on vision-to-music generation: Methods, datasets, evaluation, and challenges. In Proceedings of the ISMIR 2025 Hybrid Conference."},{"key":"e_1_3_3_143_2","doi-asserted-by":"publisher","DOI":"10.1145\/3746027.3755656"},{"key":"e_1_3_3_144_2","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Wu Shengqiong","year":"2024","unstructured":"Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. 2024. Next-gpt: Any-to-any multimodal llm. In Proceedings of the 41st International Conference on Machine Learning."},{"key":"e_1_3_3_145_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICMEW63481.2024.10645427"},{"key":"e_1_3_3_146_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10095969"},{"key":"e_1_3_3_147_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01262"},{"key":"e_1_3_3_148_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00683"},{"key":"e_1_3_3_149_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.544"},{"key":"e_1_3_3_150_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00143"},{"key":"e_1_3_3_151_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"e_1_3_3_152_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-018-3849-7"},{"key":"e_1_3_3_153_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i7.28486"},{"key":"e_1_3_3_154_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Yizhi LI","year":"2024","unstructured":"LI Yizhi, Ruibin Yuan, Ge Zhang, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin, Anton Ragni, Emmanouil Benetos, et\u00a0al. 2024. MERT: Acoustic music understanding model with large-scale self-supervised training. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_3_155_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-4060"},{"key":"e_1_3_3_156_2","doi-asserted-by":"publisher","DOI":"10.1145\/3746027.3755523"},{"key":"e_1_3_3_157_2","volume-title":"Proceedings of the 10th International Conference on Learning Representations","author":"Yu Jiahui","year":"2022","unstructured":"Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. 2022. Vector-quantized Image Modeling with Improved VQGAN. In Proceedings of the 10th International Conference on Learning Representations. OpenReview.net."},{"key":"e_1_3_3_158_2","first-page":"40339","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Yu Jiashuo","year":"2023","unstructured":"Jiashuo Yu, Yaohui Wang, Xinyuan Chen, Xiao Sun, and Yu Qiao. 2023. Long-term rhythmic video soundtracker. In Proceedings of the International Conference on Machine Learning. PMLR, 40339\u201340353."},{"key":"e_1_3_3_159_2","doi-asserted-by":"publisher","DOI":"10.1631\/FITEE.2400299"},{"key":"e_1_3_3_160_2","doi-asserted-by":"publisher","DOI":"10.52202\/075280-1722"},{"key":"e_1_3_3_161_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3129994"},{"key":"e_1_3_3_162_2","volume-title":"Proceedings of the 9th International Conference on Learning Representations","author":"Zeghidour Neil","year":"2021","unstructured":"Neil Zeghidour, Olivier Teboul, F\u00e9lix de Chaumont Quitry, and Marco Tagliasacchi. 2021. LEAF: A learnable frontend for audio classification. In Proceedings of the 9th International Conference on Learning Representations. OpenReview.net."},{"key":"e_1_3_3_163_2","doi-asserted-by":"publisher","DOI":"10.1186\/s13636-024-00370-6"},{"key":"e_1_3_3_164_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-demo.49"},{"key":"e_1_3_3_165_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49660.2025.10889053"},{"key":"e_1_3_3_166_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_35"},{"key":"e_1_3_3_167_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00374"},{"key":"e_1_3_3_168_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00300"},{"key":"e_1_3_3_169_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Zhu Bin","year":"2024","unstructured":"Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, Hongfa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et\u00a0al. 2024. LanguageBind: Extending video-language pretraining to N-modality by language-based semantic alignment. In Proceedings of the 12th International Conference on Learning Representations. OpenReview.net."},{"key":"e_1_3_3_170_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19836-6_11"},{"key":"e_1_3_3_171_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023","author":"Zhu Ye","year":"2023","unstructured":"Ye Zhu, Yu Wu, Kyle Olszewski, Jian Ren, Sergey Tulyakov, and Yan Yan. 2023. Discrete contrastive diffusion for cross-modal music and image generation. In Proceedings of the 11th International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net."},{"key":"e_1_3_3_172_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01433"},{"key":"e_1_3_3_173_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i21.34474"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3816020","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T13:07:37Z","timestamp":1782306457000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3816020"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":172,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2026,10,31]]}},"alternative-id":["10.1145\/3816020"],"URL":"https:\/\/doi.org\/10.1145\/3816020","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-02-17","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-30","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}