{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T14:24:58Z","timestamp":1784816698927,"version":"3.55.0"},"reference-count":197,"publisher":"Association for Computing Machinery (ACM)","issue":"11","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62272409"],"award-info":[{"award-number":["62272409"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,8,30]]},"abstract":"<jats:p>With the rapid development of artificial intelligence, music generation has evolved from single-modal to cross-modal approaches and is gradually moving toward multi-modal fusion. This survey systematically reviews this developmental trajectory. The discussion begins with the representation methods for key modalities, including audio, symbolic, text, and visual data. Music generation techniques are then organized across single-modal, cross-modal, and multi-modal settings. In addition, key datasets and evaluation methodologies relevant to these tasks are compiled. Finally, the survey discusses core challenges in the field, including modal fusion, data scarcity, and evaluation frameworks, and outlines potential directions for future research.<\/jats:p>\n                  <jats:p\/>","DOI":"10.1145\/3800682","type":"journal-article","created":{"date-parts":[[2026,3,5]],"date-time":"2026-03-05T21:42:19Z","timestamp":1772746939000},"page":"1-38","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-3452-1641","authenticated-orcid":false,"given":"Shuyu","family":"Li","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology, Zhejiang University","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9908-136X","authenticated-orcid":false,"given":"Shulei","family":"Ji","sequence":"additional","affiliation":[{"name":"Zhejiang University","place":["Hangzhou, China"]},{"name":"Innovation Center of Yangtze River Delta, Zhejiang University","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5613-3262","authenticated-orcid":false,"given":"Zihao","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Zhejiang University","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0760-6289","authenticated-orcid":false,"given":"Songruoyao","family":"Wu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Zhejiang University","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-2162-5110","authenticated-orcid":false,"given":"Jiaxing","family":"Yu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Zhejiang University","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0778-2303","authenticated-orcid":false,"given":"Kejun","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Zhejiang University","place":["Hangzhou, China"]},{"name":"Innovation Center of Yangtze River Delta, Zhejiang University","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,4,17]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Sami Abu-El-Haija Nisarg Kothari Joonseok Lee Paul Natsev George Toderici Balakrishnan Varadarajan and Sudheendra Vijayanarasimhan. 2016. YouTube-8M: A Large-Scale Video Classification Benchmark. arxiv:1609.08675. Retrieved from https:\/\/arxiv.org\/abs\/1609.08675"},{"key":"e_1_3_2_3_2","first-page":"7941","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Afifi Mahmoud","year":"2021","unstructured":"Mahmoud Afifi, Marcus A. Brubaker, and Michael S. Brown. 2021. HistoGAN: Controlling colors of gan-generated and real images via color histograms. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 7941\u20137950."},{"key":"e_1_3_2_4_2","unstructured":"Andrea Agostinelli Timo I. Denk Zal\u00e1n Borsos Jesse Engel Mauro Verzetti Antoine Caillon Qingqing Huang Aren Jansen Adam Roberts Marco Tagliasacchi et\u00a0al. 2023. MusicLM: Generating Music From Text. arxiv:2301.11325. Retrieved from https:\/\/arxiv.org\/abs\/2301.11325"},{"key":"e_1_3_2_5_2","first-page":"6836","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Arnab Anurag","year":"2021","unstructured":"Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lu\u010di\u0107, and Cordelia Schmid. 2021. ViViT: A video vision transformer. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 6836\u20136846."},{"key":"e_1_3_2_6_2","first-page":"12449","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Baevski Alexei","year":"2020","unstructured":"Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. Wav2Vec 2.0: A framework for self-supervised learning of speech representations. In Proceedings of the Advances in Neural Information Processing Systems. 12449\u201312460."},{"key":"e_1_3_2_7_2","unstructured":"Ye Bai Haonan Chen Jitong Chen Zhuo Chen Yi Deng Xiaohong Dong Lamtharn Hantrakul Weituo Hao Qingqing Huang Zhongyi Huang et\u00a0al. 2024. Seed-Music: A unified framework for high quality and controlled music generation. arxiv:2409.09214. Retrieved from https:\/\/arxiv.org\/abs\/2409.09214"},{"key":"e_1_3_2_8_2","first-page":"359","volume-title":"Proceedings of the 3rd International Conference on Knowledge Discovery and Data Mining","author":"Berndt Donald J.","year":"1994","unstructured":"Donald J. Berndt and James Clifford. 1994. Using dynamic time warping to find patterns in time series. In Proceedings of the 3rd International Conference on Knowledge Discovery and Data Mining. 359\u2013370."},{"key":"e_1_3_2_9_2","first-page":"591","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Bertin-Mahieux Thierry","year":"2011","unstructured":"Thierry Bertin-Mahieux, Daniel P. W. Ellis, Brian Whitman, and Paul Lamere. 2011. The million song dataset. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR). 591\u2013596."},{"key":"e_1_3_2_10_2","volume-title":"Proceedings of the International Conference on Machine Learning (ICML)","author":"Bogdanov Dmitry","year":"2019","unstructured":"Dmitry Bogdanov, Minz Won, Philip Tovstogan, Alastair Porter, and Xavier Serra. 2019. The MTG-jamendo dataset for automatic music tagging. In Proceedings of the International Conference on Machine Learning (ICML). PMLR."},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"2523","DOI":"10.1109\/TASLP.2023.3288409","article-title":"AudioLM: A language modeling approach to audio generation","volume":"31","author":"Borsos Zal\u00e1n","year":"2023","unstructured":"Zal\u00e1n Borsos, Rapha\u00ebl Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et\u00a0al. 2023. AudioLM: A language modeling approach to audio generation. IEEE Transactions on Audio, Speech and Language Processing 31 (2023), 2523\u20132533.","journal-title":"IEEE Transactions on Audio, Speech and Language Processing"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1162\/leon_a_02135"},{"issue":"1","key":"e_1_3_2_13_2","doi-asserted-by":"crossref","first-page":"172","DOI":"10.1109\/TPAMI.2019.2929257","article-title":"OpenPose: Realtime multi-person 2D pose estimation using part affinity fields","volume":"43","author":"Cao Z.","year":"2020","unstructured":"Z. Cao, G. Hidalgo, T. Simon, S. E. Wei, and Y. Sheikh. 2020. OpenPose: Realtime multi-person 2D pose estimation using part affinity fields. IEEE Transactions on Pattern Analysis and Machine Intelligence 43, 1 (2020), 172\u2013186.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_14_2","first-page":"6","volume-title":"Proceedings of the 19th International Conference on the Foundations of Digital Games","author":"Cardoso Igor","year":"2024","unstructured":"Igor Cardoso, Rubens O. Moraes, and Lucas N. Ferreira. 2024. The NES video-music database: A dataset of symbolic video game music paired with gameplay videos. In Proceedings of the 19th International Conference on the Foundations of Digital Games. Association for Computing Machinery, 6 pages. DOI:10.1145\/3649921.3650011"},{"key":"e_1_3_2_15_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Carlini Nicholas","year":"2023","unstructured":"Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2023. Quantifying memorization across neural language models. In Proceedings of the 11th International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=TatRHT_1cK"},{"key":"e_1_3_2_16_2","first-page":"6299","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Carreira Joao","year":"2017","unstructured":"Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6299\u20136308."},{"key":"e_1_3_2_17_2","first-page":"11315","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chang Huiwen","year":"2022","unstructured":"Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman. 2022. MaskGIT: Masked generative image transformer. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 11315\u201311325."},{"key":"e_1_3_2_18_2","unstructured":"Rewon Child Scott Gray Alec Radford and Ilya Sutskever. 2019. Generating Long Sequences with Sparse Transformers. arxiv:1904.10509. Retrieved from https:\/\/arxiv.org\/abs\/1904.10509"},{"key":"e_1_3_2_19_2","first-page":"26816","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chowdhury Sanjoy","year":"2024","unstructured":"Sanjoy Chowdhury, Sayan Nag, K. J. Joseph, Balaji Vasan Srinivasan, and Dinesh Manocha. 2024. MelFusion: Synthesizing music from image and language cues using diffusion models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 26816\u201326825. DOI:10.1109\/CVPR52733.2024.02533"},{"issue":"70","key":"e_1_3_2_20_2","first-page":"1","article-title":"Scaling instruction-finetuned language models","volume":"25","author":"Chung Hyung Won","year":"2024","unstructured":"Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et\u00a0al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research 25, 70 (2024), 1\u201353.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","first-page":"244","DOI":"10.1109\/ASRU51503.2021.9688253","volume-title":"Proceedings of the 2021 IEEE Automatic Speech Recognition and Understanding Workshop","author":"Chung Yu-An","year":"2021","unstructured":"Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu. 2021. W2V-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training. In Proceedings of the 2021 IEEE Automatic Speech Recognition and Understanding Workshop. 244\u2013250. DOI:10.1109\/ASRU51503.2021.9688253"},{"key":"e_1_3_2_22_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Copet Jade","year":"2024","unstructured":"Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre D\u00e9fossez. 2024. Simple and controllable music generation. In Proceedings of the Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_23_2","first-page":"2978","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Dai Zihang","year":"2019","unstructured":"Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov. 2019. Transformer-XL: Attentive language models beyond a fixed-length context. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 2978\u20132988."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3672554"},{"issue":"4","key":"e_1_3_2_25_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3197517.3201371","article-title":"Visual rhythm and beat","volume":"37","author":"Davis Abe","year":"2018","unstructured":"Abe Davis and Maneesh Agrawala. 2018. Visual rhythm and beat. ACM Transactions on Graphics 37, 4 (2018), 1\u201311.","journal-title":"ACM Transactions on Graphics"},{"key":"e_1_3_2_26_2","first-page":"316","volume-title":"Proceedings of the International Society for Music Information Retrieval (ISMIR)","author":"Defferrard Micha\u00ebl","year":"2017","unstructured":"Micha\u00ebl Defferrard, Kirell Benzi, Pierre Vandergheynst, and Xavier Bresson. 2017. FMA: A dataset for music analysis. In Proceedings of the International Society for Music Information Retrieval (ISMIR). 316\u2013323."},{"key":"e_1_3_2_27_2","article-title":"High fidelity neural audio compression","author":"D\u00e9fossez Alexandre","year":"2023","unstructured":"Alexandre D\u00e9fossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2023. High fidelity neural audio compression. Transactions on Machine Learning Research (2023). Retrieved from https:\/\/openreview.net\/forum?id=ivCd8z8zR2","journal-title":"Transactions on Machine Learning Research"},{"key":"e_1_3_2_28_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Jill Burstein, Christy Doran, and Thamar Solorio (Eds.), Association for Computational Linguistics, Minneapolis, Minnesota, 4171\u20134186. DOI:10.18653\/v1\/N19-1423"},{"key":"e_1_3_2_29_2","unstructured":"Prafulla Dhariwal Heewoo Jun Christine Payne Jong Wook Kim Alec Radford and Ilya Sutskever. 2020. Jukebox: A Generative Model for Music. arxiv:2005.00341. Retrieved from https:\/\/arxiv.org\/abs\/2005.00341"},{"key":"e_1_3_2_30_2","first-page":"2037","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Di Shangzhe","year":"2021","unstructured":"Shangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang, Leyan Zhu, Zexin He, Hongming Liu, and Shuicheng Yan. 2021. Video background music generation with controllable music transformer. In Proceedings of the 29th ACM International Conference on Multimedia. 2037\u20132045."},{"key":"e_1_3_2_31_2","unstructured":"Sander Dieleman Heiga Zen Karen Simonyan Oriol Vinyals Alex Graves Nal Kalchbrenner Andrew Senior Koray Kavukcuoglu et\u00a0al. 2016. WaveNet: A Generative Model for Raw Audio. arxiv:1609.03499. Retrieved from https:\/\/arxiv.org\/abs\/1609.03499"},{"key":"e_1_3_2_32_2","first-page":"409","volume-title":"Proceedings of the International Society for Music Information Retrieval (ISMIR)","author":"Doh SeungHeon","year":"2023","unstructured":"SeungHeon Doh, Keunwoo Choi, Jongpil Lee, and Juhan Nam. 2023. LP-MusicCaps: LLM-based pseudo music captioning. In Proceedings of the International Society for Music Information Retrieval (ISMIR). 409\u2013416."},{"key":"e_1_3_2_33_2","volume-title":"Proceedings of the International Conference on Machine Learning (ICML)","author":"Donahue Chris","year":"2023","unstructured":"Chris Donahue, Antoine Caillon, Adam Roberts, Ethan Manilow, Philippe Esling, Andrea Agostinelli, Mauro Verzetti, Ian Simon, Olivier Pietquin, Neil Zeghidour, et\u00a0al. 2023. SingSong: Generating musical accompaniments from singing. In Proceedings of the International Conference on Machine Learning (ICML). PMLR."},{"key":"e_1_3_2_34_2","first-page":"475","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Donahue Chris","year":"2018","unstructured":"Chris Donahue, Huanru Henry Mao, and Julian McAuley. 2018. The NES music database: A multi-instrumental dataset with expressive performance attributes. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR). Paris, France, 475\u2013482."},{"key":"e_1_3_2_35_2","first-page":"951","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Dong Hao-Wen","year":"2022","unstructured":"Hao-Wen Dong, Cong Zhou, Taylor Berg-Kirkpatrick, and Julian McAuley. 2022. Deep performer: Score-to-audio music performance synthesis. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 951\u2013955."},{"key":"e_1_3_2_36_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et\u00a0al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=YicbFdNTTy"},{"key":"e_1_3_2_37_2","first-page":"320","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics","author":"Du Zhengxiao","year":"2022","unstructured":"Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. GLM: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. 320\u2013335."},{"key":"e_1_3_2_38_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Elizalde Benjamin","year":"2023","unstructured":"Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang. 2023. CLAP: Learning audio concepts from natural language supervision. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. 1\u20135. DOI:10.1109\/ICASSP49357.2023.10095889"},{"key":"e_1_3_2_39_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Engel Jesse","year":"2020","unstructured":"Jesse Engel, Chenjie Gu, Adam Roberts, et\u00a0al. 2020. DDSP: Differentiable digital signal processing. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_40_2","unstructured":"Han Fang Pengfei Xiong Luhui Xu and Yu Chen. 2021. CLIP2Video: Mastering Video-Text Retrieval via Image CLIP. arxiv:2106.11097. Retrieved from https:\/\/arxiv.org\/abs\/2106.11097"},{"key":"e_1_3_2_41_2","first-page":"486","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Fonseca Eduardo","year":"2017","unstructured":"Eduardo Fonseca, Jordi Pons, Xavier Favory, Frederic Font, Dmitry Bogdanov, Andres Ferraro, Sergio Oramas, Alastair Porter, and Xavier Serra. 2017. Freesound datasets: A platform for the creation of open audio datasets. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR). Suzhou, China, 486\u2013493."},{"key":"e_1_3_2_42_2","first-page":"534","volume-title":"Proceedings of the 21st International Society for Music Information Retrieval Conference","author":"Foscarin Francesco","year":"2020","unstructured":"Francesco Foscarin, Andrew McLeod, Philippe Rigaux, Florent Jacquemard, and Masahiko Sakai. 2020. ASAP: A dataset of aligned scores and performances for piano transcription. In Proceedings of the 21st International Society for Music Information Retrieval Conference. 534\u2013541."},{"key":"e_1_3_2_43_2","doi-asserted-by":"crossref","first-page":"2001","DOI":"10.18653\/v1\/2023.emnlp-main.123","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Fradet Nathan","year":"2023","unstructured":"Nathan Fradet, Nicolas Gutowski, Fabien Chhel, and Jean-Pierre Briot. 2023. Byte pair encoding for symbolic music. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2001\u20132020."},{"key":"e_1_3_2_44_2","first-page":"183","volume-title":"Proceedings of the Annales de l\u2019ISUP","author":"Fr\u00e9chet Maurice","year":"1957","unstructured":"Maurice Fr\u00e9chet. 1957. Sur la distance de deux lois de probabilit\u00e9. In Proceedings of the Annales de l\u2019ISUP. 183\u2013198."},{"key":"e_1_3_2_45_2","article-title":"A new algorithm for data compression","author":"Gage Philip","year":"1994","unstructured":"Philip Gage. 1994. A new algorithm for data compression. C Users J. 12, 2 (1994), 23\u201338.","journal-title":"C Users J."},{"key":"e_1_3_2_46_2","first-page":"758","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Gan Chuang","year":"2020","unstructured":"Chuang Gan, Deng Huang, Peihao Chen, Joshua B. Tenenbaum, and Antonio Torralba. 2020. Foley music: Learning to generate music from videos. In Proceedings of the European Conference on Computer Vision. Springer, 758\u2013775."},{"key":"e_1_3_2_47_2","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Garc\u00eda Hugo Flores","year":"2023","unstructured":"Hugo Flores Garc\u00eda, Prem Seetharaman, Rithesh Kumar, and Bryan Pardo. 2023. VampNet: Music generation via masked acoustic token modeling. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)."},{"key":"e_1_3_2_48_2","first-page":"776","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Gemmeke Jort F.","year":"2017","unstructured":"Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Audio set: An ontology and human-labeled dataset for audio events. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 776\u2013780."},{"key":"e_1_3_2_49_2","first-page":"15180","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Girdhar Rohit","year":"2023","unstructured":"Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. ImageBind: One embedding space to bind them all. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 15180\u201315190."},{"key":"e_1_3_2_50_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Proceedings of the Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_51_2","first-page":"5036","volume-title":"Proceedings of the Interspeech Conference","author":"Gulati Anmol","year":"2020","unstructured":"Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang. 2020. Conformer: Convolution-augmented transformer for speech recognition. In Proceedings of the Interspeech Conference. 5036\u20135040."},{"key":"e_1_3_2_52_2","first-page":"166","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Gurumurthy Swaminathan","year":"2017","unstructured":"Swaminathan Gurumurthy, Ravi Kiran Sarvadevabhatla, and R. Venkatesh Babu. 2017. DeLiGAN: Generative adversarial networks for diverse and limited data. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 166\u2013174."},{"key":"e_1_3_2_53_2","volume-title":"International Society for Music Information Retrieval Conference (ISMIR)","author":"Hawthorne Curtis","year":"2018","unstructured":"Curtis Hawthorne, Erich Elsen, Jialin Song, Adam Roberts, Ian Simon, Colin Raffel, Jesse Engel, Sageev Oore, and Douglas Eck. 2018. Onsets and frames: Dual-objective piano transcription. In International Society for Music Information Retrieval Conference (ISMIR)."},{"key":"e_1_3_2_54_2","first-page":"598","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Hawthorne Curtis","year":"2022","unstructured":"Curtis Hawthorne, Ian Simon, Adam Roberts, Neil Zeghidour, Joshua Gardner, Ethan Manilow, and Jesse Engel. 2022. Multi-instrument music synthesis with spectrogram diffusion. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR). 598\u2013607."},{"key":"e_1_3_2_55_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Hawthorne Curtis","year":"2019","unstructured":"Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck. 2019. Enabling factorized piano music modeling and generation with the MAESTRO dataset. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_56_2","first-page":"16000","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"He Kaiming","year":"2022","unstructured":"Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll\u00e1r, and Ross Girshick. 2022. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 16000\u201316009."},{"key":"e_1_3_2_57_2","first-page":"770","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 770\u2013778."},{"key":"e_1_3_2_58_2","first-page":"131","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Hershey Shawn","year":"2017","unstructured":"Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis, Jort F. Gemmeke, Aren Jansen, R. Channing Moore, Manoj Plakal, Devin Platt, Rif A. Saurous, Bryan Seybold, et\u00a0al. 2017. CNN architectures for large-scale audio classification. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 131\u2013135."},{"key":"e_1_3_2_59_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems.","volume":"30","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local nash equilibrium. In Proceedings of the Advances in Neural Information Processing Systems.I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30, Curran Associates, Inc. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2017\/file\/8a1d694707eb0fefe65871369074926d-Paper.pdf"},{"key":"e_1_3_2_60_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Higgins Irina","year":"2017","unstructured":"Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. 2017. Beta-VAE: Learning basic visual concepts with a constrained variational framework. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_61_2","first-page":"6840","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. In Proceedings of the Advances in Neural Information Processing Systems. 6840\u20136851."},{"issue":"8","key":"e_1_3_2_62_2","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter Sepp","year":"1997","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural Computation 9, 8 (1997), 1735\u20131780.","journal-title":"Neural Computation"},{"key":"e_1_3_2_63_2","doi-asserted-by":"crossref","first-page":"353","DOI":"10.1145\/3206025.3206046","volume-title":"Proceedings of the 2018 ACM on International Conference on Multimedia Retrieval","author":"Hong Sungeun","year":"2018","unstructured":"Sungeun Hong, Woobin Im, and Hyun S. Yang. 2018. CBVMR: Content-based video-music retrieval using soft intra-modal structure constraint. In Proceedings of the 2018 ACM on International Conference on Multimedia Retrieval. Association for Computing Machinery, New York, NY, USA, 353\u2013361. DOI:10.1145\/3206025.3206046"},{"key":"e_1_3_2_64_2","doi-asserted-by":"crossref","unstructured":"Wen-Yi Hsiao Jen-Yu Liu Yin-Cheng Yeh and Yi-Hsuan Yang. 2021. Compound word transformer: Learning to compose full-song music over dynamic directed hypergraphs. In Proceedings of the AAAI Conference on Artificial Intelligence. 178\u2013186.","DOI":"10.1609\/aaai.v35i1.16091"},{"key":"e_1_3_2_65_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Huang Cheng-Zhi Anna","year":"2019","unstructured":"Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M. Dai, Matthew D. Hoffman, Monica Dinculescu, and Douglas Eck. 2019. Music transformer: Generating music with long-term structure. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_66_2","first-page":"28708","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Huang Po-Yao","year":"2022","unstructured":"Po-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, and Christoph Feichtenhofer. 2022. Masked autoencoders that listen. In Proceedings of the Advances in Neural Information Processing Systems. 28708\u201328720."},{"key":"e_1_3_2_67_2","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Huang Qingqing","year":"2022","unstructured":"Qingqing Huang, Aren Jansen, Joonseok Lee, Ravi Ganti, Judith Yue Li, and Daniel P. W. Ellis. 2022. MuLan: A joint embedding of music audio and natural language. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)."},{"key":"e_1_3_2_68_2","unstructured":"Qingqing Huang Daniel S. Park Tao Wang Timo I. Denk Andy Ly Nanxin Chen Zhengdong Zhang Zhishuai Zhang Jiahui Yu Christian Frank et\u00a0al. 2023. Noise2Music: Text-Conditioned Music Generation with Diffusion Models. arxiv:2302.03917. Retrieved from https:\/\/arxiv.org\/abs\/2302.03917"},{"key":"e_1_3_2_69_2","unstructured":"Yichen Huang Zachary Novack Koichi Saito Jiatong Shi Shinji Watanabe Yuki Mitsufuji John Thickstun and Chris Donahue. 2025. Aligning Text-to-Music Evaluation with Human Preferences. arxiv:2503.16669. Retrieved from https:\/\/arxiv.org\/abs\/2503.16669"},{"key":"e_1_3_2_70_2","doi-asserted-by":"crossref","first-page":"1180","DOI":"10.1145\/3394171.3413671","volume-title":"Proceedings of the 28th ACM International Conference on Multimedia","author":"Huang Yu-Siang","year":"2020","unstructured":"Yu-Siang Huang and Yi-Hsuan Yang. 2020. Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions. In Proceedings of the 28th ACM International Conference on Multimedia. Association for Computing Machinery, New York, NY, USA, 1180\u20131188. DOI:10.1145\/3394171.3413671"},{"key":"e_1_3_2_71_2","first-page":"318","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Hung Hsiao-Tzu","year":"2021","unstructured":"Hsiao-Tzu Hung, Joann Ching, Seungheon Doh, Nabin Kim, Juhan Nam, and Yi-Hsuan Yang. 2021. EMOPIA: A multi-modal pop piano dataset for emotion recognition and emotion-based music generation. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR). 318\u2013325."},{"key":"e_1_3_2_72_2","first-page":"2207","volume-title":"Proceedings of the Interspeech Conference","author":"Jang Won","year":"2021","unstructured":"Won Jang, Dan Lim, Jaesam Yoon, Bongwan Kim, and Juntae Kim. 2021. UnivNet: A neural vocoder with multi-resolution spectrogram discriminators for high-fidelity waveform generation. In Proceedings of the Interspeech Conference. Brno, Czechia, 2207\u20132211."},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597493"},{"key":"e_1_3_2_74_2","first-page":"5426","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Ju Zeqian","year":"2022","unstructured":"Zeqian Ju, Peiling Lu, Xu Tan, Rui Wang, Chen Zhang, Songruoyao Wu, Kejun Zhang, Xiang-Yang Li, Tao Qin, and Tie-Yan Liu. 2022. TeleMelody: Lyric-to-melody generation with a template-based two-stage method. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 5426\u20135437."},{"key":"e_1_3_2_75_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Karras Tero","year":"2018","unstructured":"Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2018. Progressive growing of GANs for improved quality, stability, and variation. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_76_2","first-page":"26565","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Karras Tero","year":"2022","unstructured":"Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. 2022. Elucidating the design space of diffusion-based generative models. In Proceedings of the Advances in Neural Information Processing Systems. 26565\u201326577."},{"key":"e_1_3_2_77_2","first-page":"5156","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Katharopoulos Angelos","year":"2020","unstructured":"Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Fran\u00e7ois Fleuret. 2020. Transformers are RNNs: Fast autoregressive transformers with linear attention. In Proceedings of the International Conference on Machine Learning. PMLR, 5156\u20135165."},{"key":"e_1_3_2_78_2","first-page":"176","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Kim Jong Wook","year":"2019","unstructured":"Jong Wook Kim, Rachel Bittner, Aparna Kumar, and Juan Pablo Bello. 2019. Neural music synthesis for flexible timbre control. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 176\u2013180."},{"key":"e_1_3_2_79_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma and Max Welling. 2014. Auto-encoding variational bayes. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_80_2","first-page":"17022","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Kong Jungil","year":"2020","unstructured":"Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. HiFi-GAN: Generative adversarial networks for efficient and high-fidelity speech synthesis. In Proceedings of the Advances in Neural Information Processing Systems. 17022\u201317033."},{"key":"e_1_3_2_81_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Kong Zhifeng","year":"2021","unstructured":"Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. 2021. DiffWave: A versatile diffusion model for audio synthesis. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_82_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Kreuk Felix","year":"2023","unstructured":"Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre D\u00e9fossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. 2023. AudioGen: Textually guided audio generation. In Proceedings of the International Conference on Learning Representations (ICLR). Retrieved from https:\/\/openreview.net\/forum?id=CYK7RfcOzQ4"},{"key":"e_1_3_2_83_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Kumar Rithesh","year":"2024","unstructured":"Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar. 2024. High-fidelity audio compression with improved RVQGAN. In Proceedings of the Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_84_2","first-page":"3","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Lafferty John","year":"2001","unstructured":"John Lafferty, Andrew McCallum, Fernando Pereira, et\u00a0al. 2001. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In Proceedings of the International Conference on Machine Learning. 3."},{"key":"e_1_3_2_85_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Lam Max W. Y.","year":"2024","unstructured":"Max W. Y. Lam, Qiao Tian, Tang Li, Zongyu Yin, Siyuan Feng, Ming Tu, Yuliang Ji, Rui Xia, Mingbo Ma, Xuchen Song, et\u00a0al. 2024. Efficient neural music generation. In Proceedings of the Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_86_2","first-page":"84","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations)","author":"Lee Hsin-Pei","year":"2019","unstructured":"Hsin-Pei Lee, Jhih-Sheng Fang, and Wei-Yun Ma. 2019. iComposer: An automatic songwriting system for chinese popular music. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations). 84\u201388."},{"key":"e_1_3_2_87_2","article-title":"Dancing to music","volume":"32","author":"Lee Hsin-Ying","year":"2019","unstructured":"Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang, and Jan Kautz. 2019. Dancing to music. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) 32 (2019).","journal-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)"},{"issue":"2","key":"e_1_3_2_88_2","first-page":"522","article-title":"Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications","volume":"21","author":"Li Bochen","year":"2018","unstructured":"Bochen Li, Xinzhao Liu, Karthik Dinesh, Zhiyao Duan, and Gaurav Sharma. 2018. Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications. IEEE Transactions on Multimedia 21, 2 (2018), 522\u2013535.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_89_2","first-page":"12888","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning. PMLR, 12888\u201312900."},{"key":"e_1_3_2_90_2","first-page":"13401","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Li Ruilong","year":"2021","unstructured":"Ruilong Li, Shan Yang, David A. Ross, and Angjoo Kanazawa. 2021. AI choreographer: Music-conditioned 3D dance generation with AIST++. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 13401\u201313412."},{"key":"e_1_3_2_91_2","first-page":"1","volume-title":"Proceedings of the SIGGRAPH Asia 2024 Conference Papers","author":"Li Sifei","year":"2024","unstructured":"Sifei Li, Weiming Dong, Yuxin Zhang, Fan Tang, Chongyang Ma, Oliver Deussen, Tong-Yee Lee, and Changsheng Xu. 2024. Dance-to-music generation with encoder-based textual inversion. In Proceedings of the SIGGRAPH Asia 2024 Conference Papers. 1\u201311."},{"key":"e_1_3_2_92_2","first-page":"27348","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Sizhe","year":"2024","unstructured":"Sizhe Li, Yiming Qin, Minghang Zheng, Xin Jin, and Yang Liu. 2024. Diff-BGM: A diffusion model for video background music generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 27348\u201327357."},{"key":"e_1_3_2_93_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Li Yizhi","year":"2024","unstructured":"Yizhi Li, Ruibin Yuan, Ge Zhang, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin, Anton Ragni, Emmanouil Benetos, et\u00a0al. 2024. MERT: Acoustic music understanding model with large-scale self-supervised training. In Proceedings of the International Conference on Learning Representations (ICLR). Retrieved from https:\/\/openreview.net\/forum?id=w3YZ9MSlBu"},{"key":"e_1_3_2_94_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3360695"},{"key":"e_1_3_2_95_2","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI)","author":"Lin Liwei","year":"2024","unstructured":"Liwei Lin, Gus Xia, Yixiao Zhang, and Junyan Jiang. 2024. Arrange, inpaint, and refine: Steerable long-term music audio generation and editing via content-based controls. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI). Article 851, 9 pages. DOI:10.24963\/ijcai.2024\/851"},{"key":"e_1_3_2_96_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Lipman Yaron","year":"2023","unstructured":"Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. 2023. Flow matching for generative modeling. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_97_2","first-page":"21450","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Liu Haohe","year":"2023","unstructured":"Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D. Plumbley. 2023. AudioLDM: Text-to-audio generation with latent diffusion models. In Proceedings of the International Conference on Machine Learning. PMLR, 21450\u201321474."},{"key":"e_1_3_2_98_2","doi-asserted-by":"crossref","DOI":"10.1109\/TASLP.2024.3399607","article-title":"AudioLDM 2: Learning holistic audio generation with self-supervised pretraining","author":"Liu Haohe","year":"2024","unstructured":"Haohe Liu, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Qiao Tian, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D. Plumbley. 2024. AudioLDM 2: Learning holistic audio generation with self-supervised pretraining. IEEE\/ACM Transactions on Audio, Speech, and Language Processing 32 (2024), 2871\u20132883.","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"e_1_3_2_99_2","first-page":"286","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Liu Shansong","year":"2024","unstructured":"Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, and Ying Shan. 2024. Music understanding LLaMA: Advancing text-to-music generation with question answering and captioning. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 286\u2013290."},{"key":"e_1_3_2_100_2","unstructured":"Shansong Liu Atin Sakkeer Hussain Qilong Wu Chenshuo Sun and Ying Shan. 2024. MuMu-LLaMA: Multi-Modal Music Understanding and Generation via Large Language Models. arxiv:2412.06660. Retrieved from https:\/\/arxiv.org\/abs\/2412.06660"},{"key":"e_1_3_2_101_2","first-page":"1332","volume-title":"Proceedings of the International Conference on Data Engineering","author":"Long Cheng","year":"2013","unstructured":"Cheng Long, Raymond Chi-Wing Wong, and Raymond Ka Wai Sze. 2013. T-Music: A melody composer based on frequent pattern mining. In Proceedings of the International Conference on Data Engineering. IEEE, 1332\u20131335."},{"key":"e_1_3_2_102_2","article-title":"SMPL: A skinned multi-person linear model","author":"Loper Matthew","year":"2015","unstructured":"Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. 2015. SMPL: A skinned multi-person linear model. ACM Transactions on Graphics 34, 6 (2015), 851\u2013866.","journal-title":"ACM Transactions on Graphics"},{"key":"e_1_3_2_103_2","unstructured":"Peiling Lu Xin Xu Chenfei Kang Botao Yu Chengyi Xing Xu Tan and Jiang Bian. 2023. Musecoco: Generating Symbolic Music from Text. arxiv:2306.00110. Retrieved from https:\/\/arxiv.org\/abs\/2306.00110"},{"key":"e_1_3_2_104_2","first-page":"23","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Ma Nanye","year":"2024","unstructured":"Nanye Ma, Mark Goldstein, Michael S. Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden, and Saining Xie. 2024. SiT: Exploring flow and diffusion-based generative models with scalable interpolant transformers. In Proceedings of the European Conference on Computer Vision. Springer, 23\u201340."},{"key":"e_1_3_2_105_2","unstructured":"Yinghao Ma Anders \u00d8land Anton Ragni Bleiz MacSen Del Sette Charalampos Saitis Chris Donahue Chenghua Lin Christos Plachouras Emmanouil Benetos Elona Shatri et\u00a0al. 2024. Foundation Models for Music: A Survey. arxiv:2408.14340. Retrieved from https:\/\/arxiv.org\/abs\/2408.14340"},{"key":"e_1_3_2_106_2","first-page":"5045","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Maman Ben","year":"2024","unstructured":"Ben Maman, Johannes Zeitler, Meinard M\u00fcller, and Amit H. Bermano. 2024. Performance conditioning for diffusion-based multi-instrument music synthesis. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 5045\u20135049."},{"key":"e_1_3_2_107_2","first-page":"45","volume-title":"Proceedings of the 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics.","author":"Manilow Ethan","year":"2019","unstructured":"Ethan Manilow, Gordon Wichern, Prem Seetharaman, and Jonathan Le Roux. 2019. Cutting music source separation some slakh: A dataset to study the impact of training data quality and quantity. In Proceedings of the 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics.IEEE, 45\u201349."},{"key":"e_1_3_2_108_2","first-page":"182","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Manzelli Rachel","year":"2018","unstructured":"Rachel Manzelli, Vijay Thakkar, Ali Siahkamari, and Brian Kulis. 2018. Conditioning deep generative raw audio models for structured automatic music. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR). 182\u2013189."},{"key":"e_1_3_2_109_2","first-page":"8286","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics","author":"Melechovsky Jan","year":"2024","unstructured":"Jan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, and Soujanya Poria. 2024. Mustango: Toward controllable text-to-music generation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics. 8286\u20138309."},{"key":"e_1_3_2_110_2","unstructured":"Jan Melechovsky Abhinaba Roy and Dorien Herremans. 2024. MidiCaps: A large-scale MIDI dataset with text captions. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)."},{"key":"e_1_3_2_111_2","unstructured":"Mehdi Mirza and Simon Osindero. 2014. Conditional Generative Adversarial Nets. arxiv:1411.1784. Retrieved from https:\/\/arxiv.org\/abs\/1411.1784"},{"key":"e_1_3_2_112_2","first-page":"468","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Mittal Gautam","year":"2021","unstructured":"Gautam Mittal, Jesse Engel, Curtis Hawthorne, and Ian Simon. 2021. Symbolic music generation with diffusion models. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR). Online, 468\u2013475."},{"key":"e_1_3_2_113_2","first-page":"1","volume-title":"Proceedings of the 2020 IEEE 22nd International Workshop on Multimedia Signal Processing","author":"Montesinos Juan F.","year":"2020","unstructured":"Juan F. Montesinos, Olga Slizovskaia, and Gloria Haro. 2020. Solos: A dataset for audio-visual music analysis. In Proceedings of the 2020 IEEE 22nd International Workshop on Multimedia Signal Processing. IEEE, 1\u20136."},{"key":"e_1_3_2_114_2","first-page":"615","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"M\u00fcller Meinard","year":"2011","unstructured":"Meinard M\u00fcller, Peter Grosche, and Nanzhu Jiang. 2011. A segment-based fitness measure for capturing repetitive structures of music recordings. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR). 615\u2013620."},{"key":"e_1_3_2_115_2","first-page":"97","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"M\u00fcller Meinard","year":"2012","unstructured":"Meinard M\u00fcller and Nanzhu Jiang. 2012. A scape plot representation for visualizing repetitive structures of music recordings. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR). 97\u2013102."},{"key":"e_1_3_2_116_2","first-page":"7176","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Naeem Muhammad Ferjad","year":"2020","unstructured":"Muhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi, and Jaejun Yoo. 2020. Reliable fidelity and diversity metrics for generative models. In Proceedings of the International Conference on Machine Learning. PMLR, 7176\u20137185."},{"key":"e_1_3_2_117_2","first-page":"22","volume-title":"Proceedings of the International Conference on Machine Learning.","author":"Novack Zachary","year":"2024","unstructured":"Zachary Novack, Julian McAuley, Taylor Berg-Kirkpatrick, and Nicholas J. Bryan. 2024. DITTO: Diffusion inference-time T-optimization for music generation. In Proceedings of the International Conference on Machine Learning.22 pages."},{"key":"e_1_3_2_118_2","first-page":"1116","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Parker Julian D.","year":"2024","unstructured":"Julian D. Parker, Janne Spijkervet, Katerina Kosta, Furkan Yesiler, Boris Kuznetsov, Ju-Chiang Wang, Matt Avent, Jitong Chen, and Duc Le. 2024. StemGen: A music generation model that listens. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 1116\u20131120."},{"key":"e_1_3_2_119_2","volume-title":"Proceedings of the ICML Machine Learning for Music Discovery Workshop","author":"Pati Ashis","year":"2019","unstructured":"Ashis Pati and Alexander Lerch. 2019. Latent space regularization for explicit control of musical attributes. In Proceedings of the ICML Machine Learning for Music Discovery Workshop."},{"key":"e_1_3_2_120_2","first-page":"4195","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Peebles William","year":"2023","unstructured":"William Peebles and Saining Xie. 2023. Scalable diffusion models with transformers. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4195\u20134205."},{"key":"e_1_3_2_121_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Perez Ethan","year":"2018","unstructured":"Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018. FiLM: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_2_122_2","first-page":"14104","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Pidhorskyi Stanislav","year":"2020","unstructured":"Stanislav Pidhorskyi, Donald A. Adjeroh, and Gianfranco Doretto. 2020. Adversarial latent autoencoders. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 14104\u201314113."},{"key":"e_1_3_2_123_2","first-page":"10619","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Preechakul Konpat","year":"2022","unstructured":"Konpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, and Supasorn Suwajanakorn. 2022. Diffusion autoencoders: Toward a meaningful and decodable representation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10619\u201310629."},{"key":"e_1_3_2_124_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_2_125_2","article-title":"Language models are unsupervised multitask learners","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog (2019).","journal-title":"OpenAI Blog"},{"key":"e_1_3_2_126_2","volume-title":"Learning-Based Methods for Comparing Sequences with Applications to Audio-to-MIDI Alignment and Matching","author":"Raffel Colin","year":"2016","unstructured":"Colin Raffel. 2016. Learning-Based Methods for Comparing Sequences with Applications to Audio-to-MIDI Alignment and Matching. Columbia University."},{"issue":"140","key":"e_1_3_2_127_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21, 140 (2020), 1\u201367.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_128_2","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI)","author":"Raphael Christopher","year":"2009","unstructured":"Christopher Raphael. 2009. Representation and synthesis of melodic expression. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI)."},{"key":"e_1_3_2_129_2","first-page":"3982","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing","author":"Reimers Nils","year":"2019","unstructured":"Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. 3982\u20133992."},{"key":"e_1_3_2_130_2","doi-asserted-by":"crossref","first-page":"1198","DOI":"10.1145\/3394171.3413721","volume-title":"Proceedings of the 28th ACM International Conference on Multimedia","author":"Ren Yi","year":"2020","unstructured":"Yi Ren, Jinzheng He, Xu Tan, Tao Qin, Zhou Zhao, and Tie-Yan Liu. 2020. PopMAG: Pop music accompaniment generation. In Proceedings of the 28th ACM International Conference on Multimedia. 1198\u20131206."},{"key":"e_1_3_2_131_2","first-page":"1278","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Rezende Danilo Jimenez","year":"2014","unstructured":"Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. 2014. Stochastic backpropagation and approximate inference in deep generative models. In Proceedings of the International Conference on Machine Learning. PMLR, 1278\u20131286."},{"key":"e_1_3_2_132_2","first-page":"4364","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Roberts Adam","year":"2018","unstructured":"Adam Roberts, Jesse Engel, Colin Raffel, Curtis Hawthorne, and Douglas Eck. 2018. A hierarchical latent vector model for learning long-term structure in music. In Proceedings of the International Conference on Machine Learning. 4364\u20134373."},{"key":"e_1_3_2_133_2","first-page":"2350","volume-title":"Proceedings of the Interspeech Conference","author":"Roblek Dominik","year":"2019","unstructured":"Dominik Roblek, Kevin Kilgour, Matt Sharifi, and Mauricio Zuluaga. 2019. Fr\u00e9chet audio distance: A reference-free metric for evaluating music enhancement algorithms. In Proceedings of the Interspeech Conference. 2350\u20132354."},{"key":"e_1_3_2_134_2","first-page":"10684","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Rombach Robin","year":"2022","unstructured":"Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj\u00f6rn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10684\u201310695."},{"key":"e_1_3_2_135_2","first-page":"234","volume-title":"Medical image computing and computer-assisted intervention\u2013MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18","author":"Ronneberger Olaf","year":"2015","unstructured":"Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-Net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention\u2013MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18. Springer, 234\u2013241."},{"key":"e_1_3_2_136_2","article-title":"Improved techniques for training GANs","volume":"29","author":"Salimans Tim","year":"2016","unstructured":"Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016. Improved techniques for training GANs. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) 29 (2016).","journal-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_3_2_137_2","first-page":"8050","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics","author":"Schneider Flavio","year":"2024","unstructured":"Flavio Schneider, Ojasv Kamal, Zhijing Jin, and Bernhard Sch\u00f6lkopf. 2024. Mo\u00fbsai: Efficient text-to-music diffusion models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 8050\u20138068. August 11\u201316, 2024."},{"key":"e_1_3_2_138_2","first-page":"2616","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Shao Dian","year":"2020","unstructured":"Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin. 2020. FineGym: A hierarchical video dataset for fine-grained action understanding. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2616\u20132625."},{"key":"e_1_3_2_139_2","first-page":"464","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics (Short Papers)","author":"Shaw Peter","year":"2018","unstructured":"Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018. Self-attention with relative position representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics (Short Papers). 464\u2013468."},{"key":"e_1_3_2_140_2","first-page":"13798","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Sheng Zhonghao","year":"2021","unstructured":"Zhonghao Sheng, Kaitao Song, Xu Tan, Yi Ren, Wei Ye, Shikun Zhang, and Tao Qin. 2021. SongMASS: Automatic song writing with pre-training and alignment constraint. In Proceedings of the AAAI Conference on Artificial Intelligence. 13798\u201313805."},{"key":"e_1_3_2_141_2","doi-asserted-by":"crossref","first-page":"3495","DOI":"10.1109\/TMM.2022.3161851","article-title":"Theme transformer: Symbolic music generation with theme-conditioned transformer","volume":"25","author":"Shih Yi-Jen","year":"2022","unstructured":"Yi-Jen Shih, Shih-Lun Wu, Frank Zalkow, Meinard M\u00fcller, and Yi-Hsuan Yang. 2022. Theme transformer: Symbolic music generation with theme-conditioned transformer. IEEE Transactions on Multimedia 25 (2022), 3495\u20133508.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_142_2","first-page":"5926","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Song Kaitao","year":"2019","unstructured":"Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019. MASS: Masked sequence to sequence pre-training for language generation. In Proceedings of the International Conference on Machine Learning. PMLR, 5926\u20135936."},{"key":"e_1_3_2_143_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Song Yang","year":"2021","unstructured":"Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2021. Score-based generative modeling through stochastic differential equations. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_144_2","first-page":"4952","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Su Kun","year":"2024","unstructured":"Kun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin, Joonseok Lee, Chris Donahue, Fei Sha, Aren Jansen, Yu Wang, Mauro Verzetti, et\u00a0al. 2024. V2Meow: Meowing to the visual beat via video-to-music generation. In Proceedings of the AAAI Conference on Artificial Intelligence. 4952\u20134960."},{"key":"e_1_3_2_145_2","unstructured":"Kun Su Xiulong Liu and Eli Shlizerman. 2020. Multi-Instrumentalist Net: Unsupervised Generation of Music from Body Movements. arxiv:2012.03478. Retrieved from https:\/\/arxiv.org\/abs\/2012.03478"},{"key":"e_1_3_2_146_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Su Kun","year":"2021","unstructured":"Kun Su, Xiulong Liu, and Eli Shlizerman. 2021. How does it sound? Generation of rhythmic soundtracks for human movement videos. In Proceedings of the Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_147_2","first-page":"2818","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Szegedy Christian","year":"2016","unstructured":"Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2818\u20132826."},{"key":"e_1_3_2_148_2","unstructured":"Or Tal Alon Ziv Itai Gat Felix Kreuk and Yossi Adi. 2024. Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation. arxiv:2406.10970. Retrieved from https:\/\/arxiv.org\/abs\/2406.10970"},{"key":"e_1_3_2_149_2","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Tan Hao Hao","year":"2020","unstructured":"Hao Hao Tan and Dorien Herremans. 2020. Music fadernets: Controllable music generation based on high-level features via low-level feature modelling. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)."},{"key":"e_1_3_2_150_2","volume-title":"Introducing MPT-7B: A New Standard for Open-Source Commercially Usable LLMs","author":"Team MosaicML NLP","year":"2023","unstructured":"MosaicML NLP Team. 2023. Introducing MPT-7B: A New Standard for Open-Source Commercially Usable LLMs. Retrieved May 05, 2023 from https:\/\/www.mosaicml.com\/blog\/mpt-7b"},{"key":"e_1_3_2_151_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Thickstun John","year":"2017","unstructured":"John Thickstun, Zaid Harchaoui, and Sham Kakade. 2017. Learning features of music from scratch. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_152_2","article-title":"XMusic: Towards a generalized and controllable symbolic music generation framework","author":"Tian Sida","year":"2025","unstructured":"Sida Tian, Can Zhang, Wei Yuan, Wei Tan, and Wenjie Zhu. 2025. XMusic: Towards a generalized and controllable symbolic music generation framework. IEEE Transactions on Multimedia 27 (2025), 6857\u20136871.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_153_2","first-page":"10078","article-title":"VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training","volume":"35","author":"Tong Zhan","year":"2022","unstructured":"Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) 35 (2022), 10078\u201310093.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_3_2_154_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et\u00a0al. 2023. LLaMA: Open and Efficient Foundation Language Models. arxiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_2_155_2","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Tsai Fang Duo","year":"2024","unstructured":"Fang Duo Tsai, Shih-Lun Wu, Haven Kim, Bo-Yu Chen, Hao-Chung Cheng, and Yi-Hsuan Yang. 2024. Audio prompt adapter: Unleashing music editing abilities for text-to-music with lightweight finetuning. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)."},{"key":"e_1_3_2_156_2","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Tsuchida Shuhei","year":"2019","unstructured":"Shuhei Tsuchida, Satoru Fukayama, Masahiro Hamasaki, and Masataka Goto. 2019. AIST dance video database: Multi-genre, multi-dancer, and multi-camera database for dance information processing. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)."},{"key":"e_1_3_2_157_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Oord Aaron Van Den","year":"2017","unstructured":"Aaron Van Den Oord, Oriol Vinyals, et\u00a0al. 2017. Neural discrete representation learning. In Proceedings of the Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_158_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_159_2","first-page":"1174","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Wang Bryan","year":"2019","unstructured":"Bryan Wang and Yi-Hsuan Yang. 2019. PerformanceNet: Score-to-audio music generation with multi-band convolutional residual network. In Proceedings of the AAAI Conference on Artificial Intelligence. 1174\u20131181."},{"issue":"12","key":"e_1_3_2_160_2","doi-asserted-by":"crossref","first-page":"6381","DOI":"10.1007\/s00521-024-09418-2","article-title":"A review of intelligent music generation systems","volume":"36","author":"Wang Lei","year":"2024","unstructured":"Lei Wang, Ziyi Zhao, Hanwei Liu, Junwei Pang, Yi Qin, and Qidi Wu. 2024. A review of intelligent music generation systems. Neural Computing and Applications 36, 12 (2024), 6381\u20136401.","journal-title":"Neural Computing and Applications"},{"key":"e_1_3_2_161_2","doi-asserted-by":"crossref","first-page":"5670","DOI":"10.1109\/TMM.2023.3338089","article-title":"Continuous emotion-based image-to-music generation","volume":"26","author":"Wang Yajie","year":"2024","unstructured":"Yajie Wang, Mulin Chen, and Xuelong Li. 2024. Continuous emotion-based image-to-music generation. IEEE Transactions on Multimedia 26 (2024), 5670\u20135679.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_162_2","first-page":"38","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Wang Ziyu","year":"2020","unstructured":"Ziyu Wang, Ke Chen, Junyan Jiang, Yiyi Zhang, Maoran Xu, Shuqi Dai, Xianbin Gu, and Gus Xia. 2020. POP909: A pop-song dataset for music arrangement generation. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR). 38\u201345."},{"key":"e_1_3_2_163_2","first-page":"7771","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI)","author":"Wang Zihao","year":"2024","unstructured":"Zihao Wang, Shuyu Li, Tao Zhang, Qi Wang, Pengfei Yu, Jinyang Luo, Yan Liu, Ming Xi, and Kejun Zhang. 2024. MuChin: A Chinese colloquial description benchmark for evaluating language models in the field of music. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI). 7771\u20137779."},{"key":"e_1_3_2_164_2","unstructured":"Zihao Wang Haoxuan Liu Jiaxing Yu Tao Zhang Yan Liu and Kejun Zhang. 2024. MuDiT and MuSiT: Alignment with Colloquial Expression in Description-to-Song Generation. arxiv:2407.03188. Retrieved from https:\/\/arxiv.org\/abs\/2407.03188"},{"key":"e_1_3_2_165_2","article-title":"REMAST: Real-time emotion-based music arrangement with soft transition","author":"Wang Zihao","year":"2024","unstructured":"Zihao Wang, Le Ma, Chen Zhang, Bo Han, Yunfei Xu, Yikai Wang, Xinyi Chen, Haorong Hong, Wenbo Liu, Xinda Wu, and Kejun Zhang. 2024. REMAST: Real-time emotion-based music arrangement with soft transition. IEEE Transactions on Affective Computing 16, 2 (2024), 1016\u20131030.","journal-title":"IEEE Transactions on Affective Computing"},{"key":"e_1_3_2_166_2","doi-asserted-by":"crossref","first-page":"1057","DOI":"10.1145\/3503161.3548368","volume-title":"Proceedings of the 30th ACM International Conference on Multimedia","author":"Wang Zihao","year":"2022","unstructured":"Zihao Wang, Kejun Zhang, Yuxing Wang, Chen Zhang, Qihao Liang, Pengfei Yu, Yongsheng Feng, Wenbo Liu, Yikai Wang, Yuntao Bao, and Yiheng Yang. 2022. SongDriver: Real-time music accompaniment generation without logical latency nor exposure bias. In Proceedings of the 30th ACM International Conference on Multimedia. 1057\u20131067."},{"key":"e_1_3_2_167_2","volume-title":"Proceedings of the AAAI-23 Workshop on Creative AI Across Modalities","author":"Wu Shangda","year":"2023","unstructured":"Shangda Wu and Maosong Sun. 2023. Exploring the efficacy of pre-trained checkpoints in text-to-music generation task. In Proceedings of the AAAI-23 Workshop on Creative AI Across Modalities."},{"key":"e_1_3_2_168_2","doi-asserted-by":"crossref","first-page":"2692","DOI":"10.1109\/TASLP.2024.3399026","article-title":"Music controlnet: Multiple time-varying controls for music generation","volume":"32","author":"Wu Shih-Lun","year":"2024","unstructured":"Shih-Lun Wu, Chris Donahue, Shinji Watanabe, and Nicholas J. Bryan. 2024. Music controlnet: Multiple time-varying controls for music generation. IEEE\/ACM Transactions on Audio, Speech, and Language Processing 32 (2024), 2692\u20132703.","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"e_1_3_2_169_2","first-page":"142","volume-title":"Proceedings of the International Society for Music Information Retrieval Conference (ISMIR)","author":"Wu Shih-Lun","year":"2020","unstructured":"Shih-Lun Wu and Yi-Hsuan Yang. 2020. The jazz transformer on the front line: Exploring the shortcomings of AI-composed music through quantitative measures. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR). Montr\u00e9al, Canada, 142\u2013149."},{"key":"e_1_3_2_170_2","doi-asserted-by":"crossref","unstructured":"Xinda Wu Zhijie Huang Kejun Zhang Jiaxing Yu Xu Tan Tieyao Zhang Zihao Wang and Lingyun Sun. 2023. MelodyGLM: Multi-Task Pre-Training for Symbolic Melody Generation. IEEE Transactions on Audio Speech and Language Processing 34 (2026) 1469\u20131481.","DOI":"10.1109\/TASLPRO.2026.3667433"},{"key":"e_1_3_2_171_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo Workshops","author":"Wu Xinda","year":"2024","unstructured":"Xinda Wu, Jiaming Wang, Jiaxing Yu, Tieyao Zhang, and Kejun Zhang. 2024. Popular hooks: A multimodal dataset of musical hooks for music understanding and generation. In Proceedings of the IEEE International Conference on Multimedia and Expo Workshops. 1\u20136."},{"key":"e_1_3_2_172_2","first-page":"18","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wu Yusong","year":"2024","unstructured":"Yusong Wu, Tim Cooijmans, Kyle Kastner, Adam Roberts, Ian Simon, Alexander Scarlatos, Chris Donahue, Cassie Tarakajian, Shayegan Omidshafiei, Aaron Courville, et\u00a0al. 2024. Adaptive accompaniment with ReaLchords. In Proceedings of the International Conference on Machine Learning. 18 pages."},{"key":"e_1_3_2_173_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Wu Yusong","year":"2022","unstructured":"Yusong Wu, Ethan Manilow, Yi Deng, Rigel Swavely, Kyle Kastner, Tim Cooijmans, Aaron Courville, Cheng-Zhi Anna Huang, and Jesse Engel. 2022. MIDI-DDSP: Detailed control of musical performance via hierarchical modeling. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_174_2","first-page":"9","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence.","author":"Xia Jingfei","year":"2023","unstructured":"Jingfei Xia, Mingchen Zhuge, Tiantian Geng, Shun Fan, Yuantai Wei, Zhenyu He, and Feng Zheng. 2023. Skating-mixer: Long-term sport audio-visual modeling with MLPs. In Proceedings of the AAAI Conference on Artificial Intelligence.AAAI Press, 9 pages. DOI:10.1609\/aaai.v37i3.25392"},{"issue":"12","key":"e_1_3_2_175_2","first-page":"4578","article-title":"Learning to score figure skating sport videos","volume":"30","author":"Xu Chengming","year":"2019","unstructured":"Chengming Xu, Yanwei Fu, Bing Zhang, Zitian Chen, Yu-Gang Jiang, and Xiangyang Xue. 2019. Learning to score figure skating sport videos. IEEE Transactions on Circuits and Systems for Video Technology 30, 12 (2019), 4578\u20134590.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_176_2","doi-asserted-by":"crossref","first-page":"6787","DOI":"10.18653\/v1\/2021.emnlp-main.544","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Xu Hu","year":"2021","unstructured":"Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021. VideoCLIP: Contrastive pre-training for zero-shot video-text understanding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 6787\u20136800."},{"key":"e_1_3_2_177_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Yan Sijie","year":"2018","unstructured":"Sijie Yan, Yuanjun Xiong, and Dahua Lin. 2018. Spatial temporal graph convolutional networks for skeleton-based action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_2_178_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2023.3268730"},{"issue":"9","key":"e_1_3_2_179_2","doi-asserted-by":"crossref","first-page":"4773","DOI":"10.1007\/s00521-018-3849-7","article-title":"On the evaluation of generative models in music","volume":"32","author":"Yang Li-Chia","year":"2020","unstructured":"Li-Chia Yang and Alexander Lerch. 2020. On the evaluation of generative models in music. Neural Computing and Applications 32, 9 (2020), 4773\u20134784.","journal-title":"Neural Computing and Applications"},{"key":"e_1_3_2_180_2","first-page":"1376","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Yu Botao","year":"2022","unstructured":"Botao Yu, Peiling Lu, Rui Wang, Wei Hu, Xu Tan, Wei Ye, Shikun Zhang, Tao Qin, and Tie-Yan Liu. 2022. Museformer: Transformer with fine- and coarse-grained attention for music generation. In Proceedings of the Advances in Neural Information Processing Systems. 1376\u20131388."},{"key":"e_1_3_2_181_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Yu Jiahui","year":"2022","unstructured":"Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. 2022. Vector-quantized image modeling with improved VQGAN. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_182_2","first-page":"40339","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Yu Jiashuo","year":"2023","unstructured":"Jiashuo Yu, Yaohui Wang, Xinyuan Chen, Xiao Sun, and Yu Qiao. 2023. Long-term rhythmic video soundtracker. In Proceedings of the International Conference on Machine Learning. PMLR, 40339\u201340353."},{"key":"e_1_3_2_183_2","article-title":"Suno: Potential, prospects, and trends","author":"Yu Jiaxing","year":"2024","unstructured":"Jiaxing Yu, Songruoyao Wu, Guanting Lu, Zijin Li, Li Zhou, and Kejun Zhang. 2024. Suno: Potential, prospects, and trends. Frontiers of Information Technology and Electronic Engineering 25, 7 (2024), 1025\u20131030.","journal-title":"Frontiers of Information Technology and Electronic Engineering"},{"key":"e_1_3_2_184_2","first-page":"25742","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Yu Jiaxing","year":"2025","unstructured":"Jiaxing Yu, Xinda Wu, Yunfei Xu, Tieyao Zhang, Songruoyao Wu, Le Ma, and Kejun Zhang. 2025. SongGLM: Lyric-to-melody generation with 2D alignment encoding and multi-task pre-training. In Proceedings of the AAAI Conference on Artificial Intelligence. 25742\u201325750."},{"issue":"1","key":"e_1_3_2_185_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3424116","article-title":"Conditional LSTM-GAN for melody generation from lyrics","volume":"17","author":"Yu Yi","year":"2021","unstructured":"Yi Yu, Abhishek Srivastava, and Simon Canales. 2021. Conditional LSTM-GAN for melody generation from lyrics. ACM Transactions on Multimedia Computing, Communications, and Applications 17, 1 (2021), 1\u201320.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_2_186_2","doi-asserted-by":"crossref","first-page":"495","DOI":"10.1109\/TASLP.2021.3129994","article-title":"SoundStream: An end-to-end neural audio codec","volume":"30","author":"Zeghidour Neil","year":"2021","unstructured":"Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. 2021. SoundStream: An end-to-end neural audio codec. IEEE\/ACM Transactions on Audio, Speech, and Language Processing 30 (2021), 495\u2013507.","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"e_1_3_2_187_2","first-page":"791","volume-title":"Findings of the Association for Computational Linguistics","author":"Zeng Mingliang","year":"2021","unstructured":"Mingliang Zeng, Xu Tan, Rui Wang, Zeqian Ju, Tao Qin, and Tie-Yan Liu. 2021. MusicBERT: Symbolic music understanding with large-scale pre-training. In Findings of the Association for Computational Linguistics. 791\u2013800."},{"key":"e_1_3_2_188_2","doi-asserted-by":"crossref","first-page":"1047","DOI":"10.1145\/3503161.3548357","volume-title":"Proceedings of the 30th ACM International Conference on Multimedia","author":"Zhang Chen","year":"2022","unstructured":"Chen Zhang, Luchin Chang, Songruoyao Wu, Xu Tan, Tao Qin, Tie-Yan Liu, and Kejun Zhang. 2022. ReLyMe: Improving lyric-to-melody generation by incorporating lyric-melody relationships. In Proceedings of the 30th ACM International Conference on Multimedia. 1047\u20131056."},{"key":"e_1_3_2_189_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3284996"},{"key":"e_1_3_2_190_2","first-page":"3836","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhang Lvmin","year":"2023","unstructured":"Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 3836\u20133847."},{"key":"e_1_3_2_191_2","first-page":"54","volume-title":"Proceedings of the 1st Workshop on NLP for Music and Audio","author":"Zhang Yixiao","year":"2020","unstructured":"Yixiao Zhang, Ziyu Wang, Dingsu Wang, and Gus Xia. 2020. BUTTER: A representation learning framework for bi-directional music-sentence retrieval and generation. In Proceedings of the 1st Workshop on NLP for Music and Audio. 54\u201358."},{"key":"e_1_3_2_192_2","first-page":"570","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Zhao Hang","year":"2018","unstructured":"Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. 2018. The sound of pixels. In Proceedings of the European Conference on Computer Vision. 570\u2013586."},{"key":"e_1_3_2_193_2","first-page":"11127","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Zhao Shihao","year":"2023","unstructured":"Shihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao, Shaozhe Hao, Lu Yuan, and Kwan-Yee K. Wong. 2023. Uni-controlnet: All-in-one control to text-to-image diffusion models. In Proceedings of the Advances in Neural Information Processing Systems. 11127\u201311150."},{"key":"e_1_3_2_194_2","first-page":"2837","volume-title":"Proceedings of the 24th International Conference on Knowledge Discovery and Data Mining","author":"Zhu Hongyuan","year":"2018","unstructured":"Hongyuan Zhu, Qi Liu, Nicholas Jing Yuan, Chuan Qin, Jiawei Li, Kun Zhang, Guang Zhou, Furu Wei, Yuanchun Xu, and Enhong Chen. 2018. XiaoIce band: A melody and arrangement generation framework for pop music. In Proceedings of the 24th International Conference on Knowledge Discovery and Data Mining. 2837\u20132846."},{"issue":"5","key":"e_1_3_2_195_2","first-page":"1","article-title":"Pop music generation: From melody to multi-style arrangement","volume":"14","author":"Zhu Hongyuan","year":"2020","unstructured":"Hongyuan Zhu, Qi Liu, Nicholas Jing Yuan, Kun Zhang, Guang Zhou, and Enhong Chen. 2020. Pop music generation: From melody to multi-style arrangement. ACM Transactions on Knowledge Discovery from Data 14, 5 (2020), 1\u201331.","journal-title":"ACM Transactions on Knowledge Discovery from Data"},{"key":"e_1_3_2_196_2","first-page":"86","volume-title":"Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics: System Demonstrations","author":"Zhu Pengfei","year":"2023","unstructured":"Pengfei Zhu, Chao Pang, Yekun Chai, Lei Li, Shuohuan Wang, Yu Sun, Hao Tian, and Hua Wu. 2023. ERNIE-music: Text-to-waveform music generation with diffusion models. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics: System Demonstrations. 86\u201395."},{"key":"e_1_3_2_197_2","first-page":"182","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Zhu Ye","year":"2022","unstructured":"Ye Zhu, Kyle Olszewski, Yu Wu, Panos Achlioptas, Menglei Chai, Yan Yan, and Sergey Tulyakov. 2022. Quantized GAN for complex music generation from dance videos. In Proceedings of the European Conference on Computer Vision. Springer, 182\u2013199."},{"key":"e_1_3_2_198_2","first-page":"15637","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhuo Le","year":"2023","unstructured":"Le Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao, Chenxi Bao, Stanley Peng, Songhao Han, Aixi Zhang, Fei Fang, and Si Liu. 2023. Video background music generation: Dataset, method and evaluation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 15637\u201315647."}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3800682","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T16:18:49Z","timestamp":1776442729000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3800682"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,17]]},"references-count":197,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2026,8,30]]}},"alternative-id":["10.1145\/3800682"],"URL":"https:\/\/doi.org\/10.1145\/3800682","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,17]]},"assertion":[{"value":"2025-04-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-20","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-17","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}