{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:56:59Z","timestamp":1760237819062,"version":"build-2065373602"},"reference-count":36,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2020,6,24]],"date-time":"2020-06-24T00:00:00Z","timestamp":1592956800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61761041","U1903213","61663044"],"award-info":[{"award-number":["61761041","U1903213","61663044"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Natural Science Foundation of the Xinjiang","award":["2016D01C061"],"award-info":[{"award-number":["2016D01C061"]}]},{"name":"University Scientific Research Project of Xinjiang","award":["XJEDU2017T002"],"award-info":[{"award-number":["XJEDU2017T002"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>This paper proposes a separation model adopting gated nested U-Net (GNU-Net) architecture, which is essentially a deeply supervised symmetric encoder\u2013decoder network that can generate full-resolution feature maps. Through a series of nested skip pathways, it can reduce the semantic gap between the feature maps of encoder and decoder subnetworks. In the GNU-Net architecture, only the backbone not including nested part is applied with gated linear units (GLUs) instead of conventional convolutional networks. The outputs of GNU-Net are further fed into a time-frequency (T-F) mask layer to generate two masks of singing voice and accompaniment. Then, those two estimated masks along with the magnitude and phase spectra of mixture can be transformed into time-domain signals. We explored two types of T-F mask layer, discriminative training network and difference mask layer. The experiment results show the latter to be better. We evaluated our proposed model by comparing with three models, and also with ideal T-F masks. The results demonstrate that our proposed model outperforms compared models, and it\u2019s performance comes near to ideal ratio mask (IRM). More importantly, our proposed model can output separated singing voice and accompaniment simultaneously, while the three compared models can only separate one source with trained model.<\/jats:p>","DOI":"10.3390\/sym12061051","type":"journal-article","created":{"date-parts":[[2020,6,24]],"date-time":"2020-06-24T10:54:59Z","timestamp":1592996099000},"page":"1051","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Monaural Singing Voice and Accompaniment Separation Based on Gated Nested U-Net Architecture"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0168-4830","authenticated-orcid":false,"given":"Haibo","family":"Geng","sequence":"first","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Key Laboratory of Signal Detection and Processing in Xinjiang Uygur Autonomous Region, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7505-1767","authenticated-orcid":false,"given":"Ying","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Key Laboratory of Signal Detection and Processing in Xinjiang Uygur Autonomous Region, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6604-0951","authenticated-orcid":false,"given":"Hao","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Key Laboratory of Multilingual Information Technology in Xinjiang Uygur Autonomous Region, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,6,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Sharma, B., Das, R.K., and Li, H. (2019, January 15\u201319). On the importance of audio-source separation for singer identification in polyphonic music. Proceedings of the Interspeech 2019, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-1925"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"643","DOI":"10.1109\/TASLP.2015.2396681","article-title":"Separation of singing voice using nonnegative matrix partial co-factorization for singer identification","volume":"23","author":"Hu","year":"2015","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_3","unstructured":"Mesaros, A., Virtanen, T., and Klapuri, A. (2007). Singer identification in polyphonic music using vocal separation and pattern recognition methods. ISMIR, 375\u2013378."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Kruspe, A.M., and Fraunhofer, I. (2016, January 8\u201312). Retrieval of textual song lyrics from sung inputs. Proceedings of the Interspeech 2016, San Francisco, CA, USA.","DOI":"10.21437\/Interspeech.2016-1272"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Mesaros, A., and Virtanen, T. (2010). Automatic recognition of lyrics in singing. EURASIP J. Audio Speech Music. Process., 546047.","DOI":"10.1186\/1687-4722-2010-546047"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Wang, Y., Kan, M.Y., Nwe, T.L., Shenoy, A., and Yin, J. (2004, January 10\u201316). Lyrically: Automatic synchronization of acoustic musical signals and textual lyrics. Proceedings of the 12th annual ACM international conference on Multimedia, New York, NY, USA.","DOI":"10.1145\/1027527.1027576"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2084","DOI":"10.1109\/TASLP.2016.2577879","article-title":"Singing voice separation and vocal f0 estimation based on mutual combination of robust principal component analysis and subharmonic summation","volume":"24","author":"Ikemiya","year":"2016","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lin, K.W.E., Anderson, H., Agus, N., So, C., and Lui, S. (2014, January 3\u20135). Visualising singing style under common musical events using pitch-dynamics trajectories and modified traclus clustering. Proceedings of the 2014 13th International Conference on Machine Learning and Applications, Detroit, MI, USA.","DOI":"10.1109\/ICMLA.2014.44"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Uhlich, S., Giron, F., and Mitsufuji, Y. (2015, January 19\u201324). Deep neural network based instrument extraction from music. Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, Australia.","DOI":"10.1109\/ICASSP.2015.7178348"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Yoo, J., Kim, M., Kang, K., and Choi, S. (2010, January 14\u201319). Nonnegative matrix partial co-factorization for drum source separation. Proceedings of the IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), Dallas, TX, USA.","DOI":"10.1109\/ICASSP.2010.5495305"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1307","DOI":"10.1109\/TASLP.2018.2825440","article-title":"An overview of lead and accompaniment separation in music","volume":"26","author":"Rafii","year":"2018","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Huang, P.S., Chen, S.D., Smaragdis, P., and Hasegawa-Johnson, M. (2012, January 25\u201330). Singing-voice separation from monaural recordings using robust principal component analysis. Proceedings of the 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Kyoto, Japan.","DOI":"10.1109\/ICASSP.2012.6287816"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1702","DOI":"10.1109\/TASLP.2018.2842159","article-title":"Supervised speech separation based on deep learning: An overview","volume":"26","author":"Wang","year":"2018","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_14","unstructured":"Huang, P.S., Kim, M., Hasegawa-Johnson, M., and Smaragdis, P. (2014, January 27\u201331). Singing-voice separation from monaural recordings using deep recurrent neural networks. Proceedings of the 15th International Society for Music Information Retrieval Conference (ISMIR 2014), Taipei, Taiwan."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"2136","DOI":"10.1109\/TASLP.2015.2468583","article-title":"Joint optimization of masks and deep recurrent neural networks for monaural source separation","volume":"23","author":"Huang","year":"2015","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Fan, Z.C., Lai, Y.L., and Jang, J.S.R. (2018, January 15\u201320). Svsgan: Singing voice separation via generative adversarial network. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8462091"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"He, B., Wang, S., Yuan, W., Wang, J., and Unoki, M. (2019, January 8\u201312). Data augmentation for monaural singing voice separation based on variational autoencoder-generative adversarial network. Proceedings of the 2019 IEEE International Conference on Multimedia and Expo (ICME), Shanghai, China.","DOI":"10.1109\/ICME.2019.00235"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Stoller, D., Ewert, S., and Dixon, S. (2018, January 15\u201320). Adversarial semi-supervised audio source separation applied to singing voice extraction. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461722"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Mimilakis, S.I., Drossos, K., Santos, J.F., Schuller, G., Virtanen, T., and Bengio, Y. (2018, January 15\u201320). Monaural singing voice separation with skip-filtering connections and recurrent inference of time-frequency mask. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461822"},{"key":"ref_20","unstructured":"Dauphin, Y.N., Fan, A., Auli, M., and Grangier, D. (2017, January 6\u201311). Language modeling with gated convolutional networks. Proceedings of the 34th International Conference on Machine Learning-Volume 70, Sydney, Australia."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"380","DOI":"10.1109\/TASLP.2019.2955276","article-title":"Learning complex spectral mapping with gated convolutional recurrent networks for monaural speech enhancement","volume":"28","author":"Tan","year":"2019","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1109\/TASLP.2018.2876171","article-title":"Gated Residual Networks with Dilated Convolutions for Monaural Speech Enhancement","volume":"27","author":"Tan","year":"2019","journal-title":"IEEE\/ACM IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Shi, Z., Lin, H., Liu, L., Liu, R., Han, J., and Shi, A. (2019). Deep attention gated dilated temporal convolutional networks with intra-parallel convolutional modules for end-to-end monaural speech separation. Proc. Interspeech, 3183\u20133187.","DOI":"10.21437\/Interspeech.2019-1373"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Xu, Y., Kong, Q., Wang, W., and Plumbley, M.D. (2017, January 5\u20139). Large-scale weakly supervised audio classification using gated convolutional neural network. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2018.8461975"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical image computing and computer-assisted intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_26","unstructured":"Jansson, A., Humphrey, E., Montecchio, N., Bittner, R., Kumar, A., and Weyde, T. (2017, January 23\u201327). Singing voice separation with deep u-net convolutional networks. Proceedings of the 18th International Society for Music Information Retrieval Conference, ISMIR, Suzhou, China."},{"key":"ref_27","unstructured":"Stoller, D., Ewert, S., and Dixon, S. (2018). Wave-u-net: A multi-scale neural network for end-to-end audio source separation. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Siddiquee, M.M.R., Tajbakhsh, N., and Liang, J. (2018). Unet++: A nested u-net architecture for medical image segmentation. Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, Springer.","DOI":"10.1007\/978-3-030-00889-5_1"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_30","unstructured":"Oord, A.V.D., Kalchbrenner, N., and Kavukcuoglu, K. (2016). Pixel Recurrent Neural Networks. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Chan, T.S., Yeh, T.C., Fan, Z.C., Chen, H.W., Su, L., Yang, Y.H., and Jang, R. (2015, January 19\u201324). Vocal activity informed singing voice separation with the ikala dataset. Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, Australia.","DOI":"10.1109\/ICASSP.2015.7178063"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1462","DOI":"10.1109\/TSA.2005.858005","article-title":"Performance measurement in blind audio source separation","volume":"14","author":"Vincent","year":"2006","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_33","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Isik, Y., Roux, J.L., Chen, Z., Watanabe, S., and Hershey, J.R. (2016). Single-channel multi-speaker separation using deep clustering. arXiv.","DOI":"10.21437\/Interspeech.2016-1176"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"787","DOI":"10.1109\/TASLP.2018.2795749","article-title":"Speaker-independent speech separation with deep attractor network","volume":"26","author":"Luo","year":"2018","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_36","unstructured":"Luo, Y., Chen, Z., and Ellis, D.P. (2020, May 10). Deep clustering for singing voice separation. Available online: https:\/\/www.music-ir.org\/mirex\/abstracts\/2016\/LCP1.pdf."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/12\/6\/1051\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:42:18Z","timestamp":1760175738000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/12\/6\/1051"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,6,24]]},"references-count":36,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2020,6]]}},"alternative-id":["sym12061051"],"URL":"https:\/\/doi.org\/10.3390\/sym12061051","relation":{},"ISSN":["2073-8994"],"issn-type":[{"type":"electronic","value":"2073-8994"}],"subject":[],"published":{"date-parts":[[2020,6,24]]}}}