{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T05:37:53Z","timestamp":1740116273043,"version":"3.37.3"},"reference-count":33,"publisher":"World Scientific Pub Co Pte Ltd","issue":"10","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61401227"],"award-info":[{"award-number":["61401227"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J CIRCUIT SYST COMP"],"published-print":{"date-parts":[[2021,8]]},"abstract":"<jats:p> This paper proposes a novel high-quality nonparallel many-to-many voice conversion method based on transitive star generative adversarial networks with adaptive instance normalization (Trans-StarGAN-VC with AdaIN). First, we improve the structure of generator with TransNets to make full use of hierarchical features associated with speech naturalness. In TransNets, many shortcut connections share hierarchical features between encoding and decoding part to capture sufficient linguistic and semantic information, which helps to provide natural sounding converted speech and accelerate the convergence of training process. Second, by incorporating AdaIN for style transfer, we enable the generator to learn sufficient speaker characteristic information directly from speech instead of using attribute labels, which also provides a promising framework for one-shot VC. Objective and subjective experiments with nonparallel training data show that our method significantly outperforms StarGAN-VC in both speech naturalness and speaker similarity. The mean values of mean opinion score (MOS) and ABX are increased by 24.5% and 10.7%, respectively. The comparison of spectrogram also shows that our method can provide more complete harmonic structures and details, and effectively bridge the gap between converted speech and target speech. <\/jats:p>","DOI":"10.1142\/s0218126621501887","type":"journal-article","created":{"date-parts":[[2020,12,24]],"date-time":"2020-12-24T03:17:04Z","timestamp":1608779824000},"page":"2150188","source":"Crossref","is-referenced-by-count":3,"title":["High-Quality Many-to-Many Voice Conversion Using Transitive Star Generative Adversarial Networks with Adaptive Instance Normalization"],"prefix":"10.1142","volume":"30","author":[{"given":"Yanping","family":"Li","sequence":"first","affiliation":[{"name":"College of Telecommunications and Information Engineering, Nanjing University of Posts and Telecommunications, Nanjing, Jiangsu, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhengtao","family":"He","sequence":"additional","affiliation":[{"name":"College of Telecommunications and Information Engineering, Nanjing University of Posts and Telecommunications, Nanjing, Jiangsu, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yan","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Software Engineering, Jinling Institute of Technology, Nanjing, Jiangsu, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhen","family":"Yang","sequence":"additional","affiliation":[{"name":"College of Telecommunications and Information Engineering, Nanjing University of Posts and Telecommunications, Nanjing, Jiangsu, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2021,2,19]]},"reference":[{"doi-asserted-by":"publisher","key":"S0218126621501887BIB003","DOI":"10.1109\/LSP.2019.2961213"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB004","DOI":"10.1109\/ICASSP.2014.6854066"},{"key":"S0218126621501887BIB005","first-page":"7749","volume-title":"IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP)","author":"Deng C.","year":"2020"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB006","DOI":"10.1109\/ASRU46091.2019.9003792"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB007","DOI":"10.1109\/ICASSP.2013.6639230"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB008","DOI":"10.1109\/TGRS.2018.2890513"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB009","DOI":"10.1109\/APSIPA.2015.7415320"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB010","DOI":"10.21437\/Interspeech.2019-2048"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB011","DOI":"10.1109\/ICASSP.2007.366962"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB012","DOI":"10.1109\/TASL.2010.2041688"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB013","DOI":"10.1109\/TASL.2009.2038669"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB014","DOI":"10.1109\/TASLP.2016.2593263"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB015","DOI":"10.1109\/ASRU.2003.1318521"},{"key":"S0218126621501887BIB016","first-page":"332","volume":"27","author":"Hashimoto T.","year":"2019","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB017","DOI":"10.21437\/Interspeech.2019-1774"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB018","DOI":"10.21437\/Interspeech.2019-2307"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB019","DOI":"10.1109\/TETCI.2020.2977678"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB020","DOI":"10.1109\/ISCSLP.2018.8706604"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB021","DOI":"10.1109\/ICASSP.2018.8462342"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB022","DOI":"10.1109\/SLT.2018.8639535"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB023","DOI":"10.23919\/EUSIPCO.2018.8553236"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB024","DOI":"10.21437\/Interspeech.2019-2236"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB025","DOI":"10.23919\/APSIPA.2018.8659543"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB026","DOI":"10.21437\/Interspeech.2016-1053"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB027","DOI":"10.1007\/978-3-030-33843-5_15"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB029","DOI":"10.1142\/S0218126620501212"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB030","DOI":"10.1109\/ICCV.2017.167"},{"key":"S0218126621501887BIB031","first-page":"7729","volume-title":"IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP)","author":"Wang R.","year":"2020"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB033","DOI":"10.21437\/Interspeech.2019-2067"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB034","DOI":"10.21437\/Interspeech.2019-2663"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB036","DOI":"10.1587\/transinf.2015EDP7457"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB038","DOI":"10.1109\/FSKD.2007.347"},{"doi-asserted-by":"publisher","key":"S0218126621501887BIB039","DOI":"10.21437\/Interspeech.2019-1798"}],"container-title":["Journal of Circuits, Systems and Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0218126621501887","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,7]],"date-time":"2021-09-07T11:34:52Z","timestamp":1631014492000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/abs\/10.1142\/S0218126621501887"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,2,19]]},"references-count":33,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2021,8]]}},"alternative-id":["10.1142\/S0218126621501887"],"URL":"https:\/\/doi.org\/10.1142\/s0218126621501887","relation":{},"ISSN":["0218-1266","1793-6454"],"issn-type":[{"type":"print","value":"0218-1266"},{"type":"electronic","value":"1793-6454"}],"subject":[],"published":{"date-parts":[[2021,2,19]]}}}