{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,30]],"date-time":"2026-01-30T09:39:18Z","timestamp":1769765958303,"version":"3.49.0"},"reference-count":46,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2022,12,13]],"date-time":"2022-12-13T00:00:00Z","timestamp":1670889600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,12,13]],"date-time":"2022-12-13T00:00:00Z","timestamp":1670889600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R &D Program of China","doi-asserted-by":"crossref","award":["2018YFB1403903"],"award-info":[{"award-number":["2018YFB1403903"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["No. CUC220B018"],"award-info":[{"award-number":["No. CUC220B018"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012166","name":"National Key R &D Program of China","doi-asserted-by":"crossref","award":["No.2021YFF0900700"],"award-info":[{"award-number":["No.2021YFF0900700"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100014718","name":"Innovative Research Group Project of the National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No. 62207029 and 62271454"],"award-info":[{"award-number":["No. 62207029 and 62271454"]}],"id":[{"id":"10.13039\/100014718","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Beijing Natural Science Foundation","award":["No. L223033"],"award-info":[{"award-number":["No. L223033"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In recent years, there has been a national craze for metaverse concerts. However, existing meta-universe concert efforts often focus on immersive visual experiences and lack consideration of the musical and aural experience. But for concerts, it is the beautiful music and the immersive listening experience that deserve the most attention. Therefore, enhancing intelligent and immersive musical experiences is essential for the further development of the metaverse. With this in mind, we propose a metaverse concert generation framework \u2014 from intelligent music generation to stereo conversion and sound field design for virtual concert stages. First, combining the ideas of reinforcement learning and value functions, the Transformer-XL music generation network is improved and used in training all the music in the POP909 dataset. Experiments show that both improved algorithms have advantages over the original method in terms of objective evaluation and subjective evaluation metrics. In addition, this paper validates a neural rendering method that can be used to generate spatial audio based on a binaural-integrated neural network with a fully convolutional technique. And the purely data-driven end-to-end model performs to be more reliable compared with traditional spatial audio generation methods such as HRTF. Finally, we propose a metadata-based audio rendering algorithm to simulate real-world acoustic environments.<\/jats:p>","DOI":"10.1186\/s13636-022-00261-8","type":"journal-article","created":{"date-parts":[[2022,12,13]],"date-time":"2022-12-13T11:03:57Z","timestamp":1670929437000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":23,"title":["MetaMGC: a music generation framework for concerts in\u00a0metaverse"],"prefix":"10.1186","volume":"2022","author":[{"given":"Cong","family":"Jin","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fengjuan","family":"Wu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3653-9951","authenticated-orcid":false,"given":"Jing","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zixuan","family":"Guan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhe","family":"Han","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,12,13]]},"reference":[{"key":"261_CR1","unstructured":"L.H. Lee, Z. Lin, R. Hu, Z. Gong, A. Kumar, T. Li, S. Li, P. Hui, When creators meet the metaverse: A survey on computational arts. arXiv preprint arXiv:2111.13486 (2021)"},{"key":"261_CR2","unstructured":"Sony: Making Madison Beer\u2019s Immersive Reality Concert Experience. https:\/\/www.youtube.com\/watch?v=qm1PWZoWI (2021)"},{"key":"261_CR3","unstructured":"P. Danowski. Connexion. https:\/\/www.youtube.com\/watch?v=OUE8V8x28wQ (2019)"},{"key":"261_CR4","unstructured":"P. Danowski. Sound of the metaverse: A breif history of virtual reality music instruments and virtual music venues. https:\/\/panopticon.am\/a-brief-history-of-virtual-reality-music-instruments-and-virtual-music-venues (2021)"},{"key":"261_CR5","unstructured":"T. Scott. Travis scott and fortnite present: Astronomical. (2020)"},{"key":"261_CR6","doi-asserted-by":"crossref","unstructured":"C. Jin, T. Wang, X. Li, C.J.J. Tie, Y. Tie, S. Liu, M. Yan, Y. Li, J. Wang, S. Huang, A transformer generative adversarial network for multi-track music generation. CAAI Transactions on Intelligence Technology. 7(3):369\u2013380(2022)","DOI":"10.1049\/cit2.12065"},{"key":"261_CR7","unstructured":"Z. Wang, K. Chen, J. Jiang, Y. Zhang, M. Xu, S. Dai, X. Gu, G. Xia, Pop909: A pop-song dataset for music arrangement generation. arXiv preprint arXiv:2008.07142 (2020)"},{"key":"261_CR8","doi-asserted-by":"publisher","unstructured":"I. D. Gebru et al., \"Implicit HRTF Modeling Using Temporal Convolutional Networks,\" ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 3385\u20133389. https:\/\/doi.org\/10.1109\/ICASSP39728.2021.9414750","DOI":"10.1109\/ICASSP39728.2021.9414750"},{"key":"261_CR9","doi-asserted-by":"crossref","unstructured":"A. Kumar, in Immersive 3D Design Visualization: With Autodesk Maya and Unreal Engine 4. Interactive Visualization with UE4 (Apress, Berkeley, 2021), pp. 103\u2013115","DOI":"10.1007\/978-1-4842-6597-0_4"},{"key":"261_CR10","doi-asserted-by":"crossref","unstructured":"N. Boulanger-Lewandowski, Y. Bengio, P. Vincent, Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription. arXiv preprint arXiv:1206.6392 (2012)","DOI":"10.1109\/ICASSP.2013.6638244"},{"key":"261_CR11","unstructured":"E. Waite, et al., Generating long-term structure in songs and stories. Web Blog Post Magenta.4,231-234(2016)"},{"key":"261_CR12","unstructured":"G. Hadjeres, F. Nielsen, Interactive music generation with positional constraints using anticipation-rnns. arXiv preprint arXiv:1709.06404 (2017)"},{"key":"261_CR13","doi-asserted-by":"crossref","unstructured":"Johnson D D. Generating polyphonic music using tied parallel networks[C]\/\/International conference on evolutionary and biologically inspired music and art. Springer, Cham, 2017: 128-143.","DOI":"10.1007\/978-3-319-55750-2_9"},{"key":"261_CR14","doi-asserted-by":"crossref","unstructured":"S.R. Bowman, L. Vilnis, O. Vinyals, A.M. Dai, R. Jozefowicz, S. Bengio, Generating sentences from a continuous space. arXiv preprint arXiv:1511.06349 (2015)","DOI":"10.18653\/v1\/K16-1002"},{"key":"261_CR15","unstructured":"A. Roberts, J. Engel, C. Raffel, C. Hawthorne, D. Eck, in International conference on machine learning. A hierarchical latent vector model for learning long-term structure in music (PMLR, 2018), pp. 4364\u20134373. http:\/\/proceedings.mlr.press\/v80\/roberts18a\/roberts18a.pdf"},{"key":"261_CR16","doi-asserted-by":"crossref","unstructured":"B. Jia, J. Lv, Y. Pu, X. Yang, in 2019 International Joint Conference on Neural Networks (IJCNN). Impromptu accompaniment of pop music using coupled latent variable model with binary regularizer (IEEE, 2019), pp. 1\u20136","DOI":"10.1109\/IJCNN.2019.8852373"},{"key":"261_CR17","unstructured":"L.C. Yang, S.Y. Chou, Y.H. Yang, Midinet: A convolutional generative adversarial network for symbolic-domain music generation. arXiv preprint arXiv:1703.10847 (2017)"},{"key":"261_CR18","doi-asserted-by":"publisher","unstructured":"L. Yu, W. Zhang, J. Wang, Y. Yu, in Proceedings of the AAAI conference on artificial intelligence. Seqgan: Sequence generative adversarial nets with policy gradient, AAAI, vol. 31 (2017). https:\/\/doi.org\/10.1609\/aaai.v31i1.10804","DOI":"10.1609\/aaai.v31i1.10804"},{"key":"261_CR19","doi-asserted-by":"publisher","unstructured":"H.W. Dong, W.Y. Hsiao, L.C. Yang, Y.H. Yang, in Proceedings of the AAAI Conference on Artificial Intelligence. Musegan: Multi-track sequential generative adversarial networks for symbolic music generation and accompaniment, AAAI, vol. 32 (2018). https:\/\/doi.org\/10.1609\/aaai.v32i1.11312","DOI":"10.1609\/aaai.v32i1.11312"},{"key":"261_CR20","unstructured":"C.Z.A. Huang, A. Vaswani, J. Uszkoreit, N. Shazeer, I. Simon, C. Hawthorne, A.M. Dai, M.D. Hoffman, M. Dinculescu, D. Eck, Music transformer. arXiv preprint arXiv:1809.04281 (2018)"},{"key":"261_CR21","unstructured":"C. Donahue, H.H. Mao, Y.E. Li, G.W. Cottrell, J. McAuley, Lakhnes: Improving multi-instrumental music generation with cross-domain pre-training. arXiv preprint arXiv:1907.04868 (2019)"},{"key":"261_CR22","doi-asserted-by":"publisher","unstructured":"Y.S. Huang, Y.H. Yang, in Proceedings of the 28th ACM International Conference on Multimedia. Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions (2020),ACM Multimedia, pp. 1180\u20131188. https:\/\/doi.org\/10.1145\/3394171.3413671","DOI":"10.1145\/3394171.3413671"},{"key":"261_CR23","doi-asserted-by":"crossref","unstructured":"Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q.V. Le, R. Salakhutdinov, Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860 (2019)","DOI":"10.18653\/v1\/P19-1285"},{"issue":"14","key":"261_CR24","doi-asserted-by":"publisher","first-page":"5014","DOI":"10.3390\/app10145014","volume":"10","author":"S Li","year":"2020","unstructured":"S. Li, J. Peissig, Measurement of head-related transfer functions: A review. Appl Sci 10(14), 5014 (2020)","journal-title":"Appl Sci"},{"key":"261_CR25","unstructured":"C.  Armstrong,  T.   McKenzie,  D.   Murphy, and  G.   Kearney, \"A Perceptual Spectral Difference Model for Binaural Signals,\" Engineering Brief 457, (2018 October.). http:\/\/www.aes.org\/e-lib\/browse.cfm?elib=19722"},{"issue":"5","key":"261_CR26","doi-asserted-by":"publisher","first-page":"532","DOI":"10.3390\/app7050532","volume":"7","author":"W Zhang","year":"2017","unstructured":"W. Zhang, P.N. Samarasinghe, H. Chen, T.D. Abhayapala, Surround by sound: A review of spatial audio recording and reproduction. Appl. Sci. 7(5), 532 (2017)","journal-title":"Appl. Sci."},{"key":"261_CR27","doi-asserted-by":"crossref","unstructured":"Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R.J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, et\u00a0al., Tacotron: Towards end-to-end speech synthesis. arXiv preprint arXiv:1703.10135 (2017)","DOI":"10.21437\/Interspeech.2017-1452"},{"key":"261_CR28","doi-asserted-by":"crossref","unstructured":"J.Y. Lee, S.J. Cheon, B.J. Choi, N.S. Kim, E. Song, in INTERSPEECH. Acoustic Modeling Using Adversarially Trained Variational Recurrent Neural Network for Speech Synthesis (2018), INTERSPEECH ,pp. 917\u2013921","DOI":"10.21437\/Interspeech.2018-1598"},{"key":"261_CR29","unstructured":"S. Vasquez, M. Lewis, Melnet: A generative model for audio in the frequency domain. arXiv preprint arXiv:1906.01083 (2019)"},{"key":"261_CR30","first-page":"2","volume":"125","author":"A Van Den Oord","year":"2016","unstructured":"A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A.W. Senior, K. Kavukcuoglu, Wavenet: A generative model for raw audio. SSW 125, 2 (2016)","journal-title":"SSW"},{"key":"261_CR31","doi-asserted-by":"crossref","unstructured":"A. Defossez, G. Synnaeve, Y. Adi, Real time speech enhancement in the waveform domain. arXiv preprint arXiv:2006.12847 (2020)","DOI":"10.21437\/Interspeech.2020-2409"},{"key":"261_CR32","doi-asserted-by":"crossref","unstructured":"D. Rethage, J. Pons, X. Serra, in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). A wavenet for speech denoising (IEEE, 2018), pp. 5069\u20135073.Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8462417"},{"key":"261_CR33","unstructured":"N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. Oord, S. Dieleman, K. Kavukcuoglu, in International Conference on Machine Learning. Efficient neural audio synthesis. Efficient neural audio synthesis (PMLR, 2018), pp. 2410\u20132419.Available from https:\/\/proceedings.mlr.press\/v80\/kalchbrenner18a.html"},{"key":"261_CR34","unstructured":"N. Mor, L. Wolf, A. Polyak, Y. Taigman, A universal music translation network. arXiv preprint arXiv:1805.07848 (2018)"},{"key":"261_CR35","unstructured":"A. Richard, D. Markovic, I.D. Gebru, S. Krenn, G.A. Butler, F. Torre, Y. Sheikh, in International Conference on Learning Representations. Neural synthesis of binaural speech from mono audio ,In International Conference on Learning Representations. (2020)"},{"key":"261_CR36","unstructured":"P. Morgado, N. Nvasconcelos, T. Langlois, O. Wang, Self-supervised generation of spatial audio for 360 video. Adv. Neural Inf. Process. Syst.31,1-15 (2018)"},{"key":"261_CR37","doi-asserted-by":"crossref","unstructured":"R. Gao, K. Grauman, in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2.5 d visual sound (2019), IEEE Piscataway, pp. 324\u2013333.Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2019.00041"},{"key":"261_CR38","doi-asserted-by":"publisher","unstructured":"I.M. Sobol. A Primer for the Monte Carlo Method (1st ed.). CRC Press, Boca Raton. (1994)\u00a0https:\/\/doi.org\/10.1201\/9781315136448","DOI":"10.1201\/9781315136448"},{"issue":"247","key":"261_CR39","doi-asserted-by":"publisher","first-page":"335","DOI":"10.1080\/01621459.1949.10483310","volume":"44","author":"N Metropolis","year":"1949","unstructured":"N. Metropolis, S. Ulam, The monte carlo method. J. Am. Stat. Assoc. 44(247), 335\u2013341 (1949)","journal-title":"J. Am. Stat. Assoc."},{"issue":"2","key":"261_CR40","first-page":"229","volume":"17","author":"RS Sutton","year":"1999","unstructured":"R.S. Sutton, A.G. Barto, Reinforcement learning: An introduction. Robotica 17(2), 229\u2013235 (1999)","journal-title":"Robotica"},{"issue":"9","key":"261_CR41","doi-asserted-by":"publisher","first-page":"4773","DOI":"10.1007\/s00521-018-3849-7","volume":"32","author":"LC Yang","year":"2020","unstructured":"L.C. Yang, A. Lerch, On the evaluation of generative models in music. Neural Comput. Applic. 32(9), 4773\u20134784 (2020)","journal-title":"Neural Comput. Applic."},{"issue":"12","key":"261_CR42","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s42452-020-03715-w","volume":"2","author":"IP Yamshchikov","year":"2020","unstructured":"I.P. Yamshchikov, A. Tikhonov, Music generation with variational recurrent autoencoder supported by history. SN Appl. Sci. 2(12), 1\u20137 (2020)","journal-title":"SN Appl. Sci."},{"key":"261_CR43","unstructured":"H.W. Dong, K. Chen, J. McAuley, T. Berg-Kirkpatrick, Muspy: A toolkit for symbolic music generation. arXiv preprint arXiv:2008.01951 (2020)"},{"key":"261_CR44","unstructured":"W. Li, The intersection of audio music and computers-audio music technology [m] (Fudan University Press, Shanghai, 2019), pp.297\u2013316"},{"key":"261_CR45","doi-asserted-by":"publisher","unstructured":"L. Mou, J. Li, J. Li, F. Gao, R. Jain, B. Yin, in 2021 IEEE 4th International Conference on Multimedia Information Processing and Retrieval (MIPR). MemoMusic: A Personalized Music Recommendation Framework Based on Emotion and Memory (IEEE, 2021), pp. 341\u2013347. Tokyo, Japan. https:\/\/doi.org\/10.1109\/MIPR51284.2021.00064","DOI":"10.1109\/MIPR51284.2021.00064"},{"key":"261_CR46","doi-asserted-by":"publisher","unstructured":"L. Mou, Y. Zhao, Q. Hao, Y. Tian, J. Li, J. Li, Y. Sun, F. Gao, B. Yin, in 2022 IEEE International Conference on Multimedia and Expo Workshops (ICMEW). Memomusic Version 2.0: Extending Personalized Music Recommendation with Automatic Music Generation (IEEE, 2022), Taipei City, Taiwan. pp. 1\u20136. https:\/\/doi.org\/10.1109\/ICMEW56448.2022.9859356.","DOI":"10.1109\/ICMEW56448.2022.9859356"}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-022-00261-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-022-00261-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-022-00261-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,13]],"date-time":"2022-12-13T11:17:44Z","timestamp":1670930264000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-022-00261-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,13]]},"references-count":46,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,12]]}},"alternative-id":["261"],"URL":"https:\/\/doi.org\/10.1186\/s13636-022-00261-8","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,13]]},"assertion":[{"value":"25 July 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 November 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 December 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"31"}}