{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,5,4]],"date-time":"2025-05-04T04:01:57Z","timestamp":1746331317868,"version":"3.40.4"},"reference-count":43,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"5","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2025,5,1]]},"DOI":"10.1587\/transinf.2024edp7104","type":"journal-article","created":{"date-parts":[[2024,11,12]],"date-time":"2024-11-12T22:11:16Z","timestamp":1731449476000},"page":"420-430","source":"Crossref","is-referenced-by-count":0,"title":["A Text-to-Lyrics Generation Method Leveraging Image-based Semantics and Reducing Plagiarism Risk"],"prefix":"10.1587","volume":"E108.D","author":[{"given":"Kento","family":"WATANABE","sequence":"first","affiliation":[{"name":"National Institute of Advanced Industrial Science and Technology (AIST)"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Masataka","family":"GOTO","sequence":"additional","affiliation":[{"name":"National Institute of Advanced Industrial Science and Technology (AIST)"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","unstructured":"[1] K. Watanabe and M. Goto, \u201cText-to-lyrics generation with image-based semantics and reduced risk of plagiarism,\u201d Proceedings of the 24th International Society for Music Information Retrieval Conference (ISMIR), pp.398-406, 2023."},{"key":"2","unstructured":"[2] K. Watanabe and M. Goto, \u201cLyrics information processing: Analysis, generation, and applications,\u201d Proceedings of the 1st Workshop on NLP for Music and Audio (NLP4MusA), pp.6-12, 2020."},{"key":"3","doi-asserted-by":"crossref","unstructured":"[3] K. Watanabe, Y. Matsubayashi, K. Inui, T. Nakano, S. Fukayama, and M. Goto, \u201cLyriSys: An interactive support system for writing lyrics based on topic transition,\u201d Proceedings of the 22nd International Conference on Intelligent User Interfaces (ACM IUI), pp.559-563, 2017. 10.1145\/3025171.3025194","DOI":"10.1145\/3025171.3025194"},{"key":"4","doi-asserted-by":"crossref","unstructured":"[4] H.G. Oliveira, T. Mendes, and A. Boavida, \u201cCo-PoeTryMe: A co-creative interface for the composition of poetry,\u201d Proceedings of the 10th International Conference on Natural Language Generation (INLG), pp.70-71, 2017. 10.18653\/v1\/w17-3508","DOI":"10.18653\/v1\/W17-3508"},{"key":"5","doi-asserted-by":"publisher","unstructured":"[5] H.G. Oliveira, T. Mendes, A. Boavida, A. Nakamura, and M. Ackerman, \u201cCo-PoeTryMe: Interactive poetry generation,\u201d Cognitive Systems Research, vol.54, pp.199-216, 2019. 10.1016\/j.cogsys.2018.11.012","DOI":"10.1016\/j.cogsys.2018.11.012"},{"key":"6","doi-asserted-by":"crossref","unstructured":"[6] R. Zhang, X. Mao, L. Li, L. Jiang, L. Chen, Z. Hu, Y. Xi, C. Fan, and M. Huang, \u201cYouling: An AI-assisted lyrics creation system,\u201d Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (EMNLP), pp.85-91, 2020. 10.18653\/v1\/2020.emnlp-demos.12","DOI":"10.18653\/v1\/2020.emnlp-demos.12"},{"key":"7","doi-asserted-by":"crossref","unstructured":"[7] N. Ram, T. Gummadi, R. Bhethanabotla, R.J. Savery, and G. Weinberg, \u201cSay what? collaborative pop lyric generation using multitask transfer learning,\u201d Proceedings of the 9th International Conference on Human-Agent Interaction (HAI), pp.165-173, 2021. 10.1145\/3472307.3484175","DOI":"10.1145\/3472307.3484175"},{"key":"8","doi-asserted-by":"crossref","unstructured":"[8] L. Zhang, R. Zhang, X. Mao, and Y. Chang, \u201cQiuNiu: A Chinese lyrics generation system with passage-level input,\u201d Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics - System Demonstrations (ACL), pp.76-82, 2022. 10.18653\/v1\/2022.acl-demo.7","DOI":"10.18653\/v1\/2022.acl-demo.7"},{"key":"9","doi-asserted-by":"crossref","unstructured":"[9] N. Liu, W. Han, G. Liu, D. Peng, R. Zhang, X. Wang, and H. Ruan, \u201cChipSong: A controllable lyric generation system for Chinese popular song,\u201d Proceedings of the First Workshop on Intelligent and Interactive Writing Assistants (In2Writing), pp.85-95, 2022. 10.18653\/v1\/2022.in2writing-1.13","DOI":"10.18653\/v1\/2022.in2writing-1.13"},{"key":"10","unstructured":"[10] C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P.J. Liu, \u201cExploring the limits of transfer learning with a unified text-to-text transformer,\u201d Journal of Machine Learning Research, vol.21, no.140, pp.1-67, 2020."},{"key":"11","unstructured":"[11] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, \u0141. Kaiser, and I. Polosukhin, \u201cAttention is all you need,\u201d Advances in neural information processing systems, vol.30, pp.1-11, 2017."},{"key":"12","unstructured":"[12] D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. M\u00fcller, J. Penna, and R. Rombach, \u201cSDXL: improving latent diffusion models for high-resolution image synthesis,\u201d The Twelfth International Conference on Learning Representations (ICLR), 2024."},{"key":"13","doi-asserted-by":"publisher","unstructured":"[13] A. Papadopoulos, P. Roy, and F. Pachet, \u201cAvoiding plagiarism in markov sequence generation,\u201d Proceedings of the 28th AAAI Conference on Artificial Intelligence, vol.28, no.1, pp.2731-2737, 2014. 10.1609\/aaai.v28i1.9126","DOI":"10.1609\/aaai.v28i1.9126"},{"key":"14","doi-asserted-by":"crossref","unstructured":"[14] Q. Feng, C. Guo, F. Benitez-Quiroz, and A.M. Mart\u0131\u0301nez, \u201cWhen do GANs replicate? on the choice of dataset size,\u201d 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), pp.6681-6690, 2021. 10.1109\/iccv48922.2021.00663","DOI":"10.1109\/ICCV48922.2021.00663"},{"key":"15","doi-asserted-by":"publisher","unstructured":"[15] T. Nakano, K. Yoshii, and M. Goto, \u201cMusical similarity and commonness estimation based on probabilistic generative models of musical elements,\u201d International Journal of Semantic Computing (IJSC), vol.10, no.1, pp.27-52, 2016. 10.1142\/s1793351x1640002x","DOI":"10.1142\/S1793351X1640002X"},{"key":"16","doi-asserted-by":"crossref","unstructured":"[16] K. Watanabe, Y. Matsubayashi, S. Fukayama, M. Goto, K. Inui, and T. Nakano, \u201cA melody-conditioned lyrics language model,\u201d Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pp.163-172, 2018. 10.18653\/v1\/n18-1015","DOI":"10.18653\/v1\/N18-1015"},{"key":"17","doi-asserted-by":"crossref","unstructured":"[17] X. Lu, J. Wang, B. Zhuang, S. Wang, and J. Xiao, \u201cA syllable-structured, contextually-based conditionally generation of Chinese lyrics,\u201d Proceedings of the 16th Pacific Rim International Conference on Artificial Intelligence (PRICAI), pp.257-265, 2019. 10.1007\/978-3-030-29894-4_20","DOI":"10.1007\/978-3-030-29894-4_20"},{"key":"18","doi-asserted-by":"crossref","unstructured":"[18] Y. Chen and A. Lerch, \u201cMelody-conditioned lyrics generation with SeqGANs,\u201d Proceedings of the 2020 IEEE International Symposium on Multimedia (ISM), pp.189-196, 2020. 10.1109\/ism.2020.00040","DOI":"10.1109\/ISM.2020.00040"},{"key":"19","doi-asserted-by":"publisher","unstructured":"[19] Y.-F. Huang and K.-C. You, \u201cAutomated generation of Chinese lyrics based on melody emotions,\u201d IEEE Access, vol.9, pp.98060-98071, 2021. 10.1109\/access.2021.3095964","DOI":"10.1109\/ACCESS.2021.3095964"},{"key":"20","doi-asserted-by":"crossref","unstructured":"[20] X. Ma, Y. Wang, M.-Y. Kan, and W.S. Lee, \u201cAI-Lyricist: Generating music and vocabulary constrained lyrics,\u201d Proceedings of the 29th ACM International Conference on Multimedia (ACM-MM), pp.1002-1011, 2021. 10.1145\/3474085.3475502","DOI":"10.1145\/3474085.3475502"},{"key":"21","doi-asserted-by":"publisher","unstructured":"[21] Z. Sheng, K. Song, X. Tan, Y. Ren, W. Ye, S. Zhang, and T. Qin, \u201cSongMASS: Automatic song writing with pre-training and alignment constraint,\u201d Proceedings of the 35th AAAI Conference on Artificial Intelligence, vol.35, no.15, pp.13798-13805, 2021. 10.1609\/aaai.v35i15.17626","DOI":"10.1609\/aaai.v35i15.17626"},{"key":"22","unstructured":"[22] G. Barbieri, F. Pachet, P. Roy, and M.D. Esposti, \u201cMarkov constraints for generating lyrics with style,\u201d Proceedings of the 20th European Conference on Artificial Intelligence (ECAI), pp.115-120, 2012."},{"key":"23","doi-asserted-by":"crossref","unstructured":"[23] J. Hopkins and D. Kiela, \u201cAutomatically generating rhythmic verse with neural networks,\u201d Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL), pp.168-178, 2017. 10.18653\/v1\/p17-1016","DOI":"10.18653\/v1\/P17-1016"},{"key":"24","doi-asserted-by":"crossref","unstructured":"[24] E. Manjavacas, M. Kestemont, and F. Karsdorp, \u201cGeneration of hip-hop lyrics with hierarchical modeling and conditional templates,\u201d Proceedings of the 12th International Conference on Natural Language Generation (INLG), pp.301-310, 2019. 10.18653\/v1\/w19-8638","DOI":"10.18653\/v1\/W19-8638"},{"key":"25","doi-asserted-by":"crossref","unstructured":"[25] L. Xue, K. Song, D. Wu, X. Tan, N.L. Zhang, T. Qin, W.-Q. Zhang, and T.-Y. Liu, \u201cDeepRapper: Neural rap generation with rhyme and rhythm modeling,\u201d Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP), pp.69-81, 2021. 10.18653\/v1\/2021.acl-long.6","DOI":"10.18653\/v1\/2021.acl-long.6"},{"key":"26","doi-asserted-by":"publisher","unstructured":"[26] J.-W. Chang, J.C. Hung, and K.-C. Lin, \u201cSingability-enhanced lyric generator with music style transfer,\u201d Computer Communications, vol.168, pp.33-53, 2021. 10.1016\/j.comcom.2021.01.002","DOI":"10.1016\/j.comcom.2021.01.002"},{"key":"27","unstructured":"[27] O. Vechtomova, G. Sahu, and D. Kumar, \u201cGeneration of lyrics lines conditioned on music audio clips,\u201d Proceedings of the 1st Workshop on NLP for Music and Audio (NLP4MusA), pp.33-37, 2020."},{"key":"28","unstructured":"[28] O. Vechtomova, G. Sahu, and D. Kumar, \u201cLyricJam: A system for generating lyrics for live instrumental music,\u201d Proceedings of the 12th International Conference on Computational Creativity (ICCC), pp.122-130, 2021. 10.1007\/978-3-031-29956-8_19"},{"key":"29","doi-asserted-by":"crossref","unstructured":"[29] K. Watanabe and M. Goto, \u201cAtypical lyrics completion considering musical audio signals,\u201d Proceedings of the 27th International Conference on Multimedia Modeling (MMM), pp.174-186, 2021. 10.1007\/978-3-030-67832-6_15","DOI":"10.1007\/978-3-030-67832-6_15"},{"key":"30","unstructured":"[30] K. Watanabe, Y. Matsubayashi, K. Inui, and M. Goto, \u201cModeling structural topic transitions for automatic lyrics generation,\u201d Proceedings of the 28th Pacific Asia Conference on Language, Information and Computation (PACLIC), pp.422-431, 2014."},{"key":"31","doi-asserted-by":"crossref","unstructured":"[31] P. Potash, A. Romanov, and A. Rumshisky, \u201cGhostWriter: Using an LSTM for automatic rap lyric generation,\u201d Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp.1919-1924, 2015. 10.18653\/v1\/d15-1221","DOI":"10.18653\/v1\/D15-1221"},{"key":"32","doi-asserted-by":"crossref","unstructured":"[32] M. Ghazvininejad, X. Shi, Y. Choi, and K. Knight, \u201cGenerating topical poetry,\u201d Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp.1183-1191, 2016. 10.18653\/v1\/d16-1126","DOI":"10.18653\/v1\/D16-1126"},{"key":"33","doi-asserted-by":"crossref","unstructured":"[33] H. Fan, J. Wang, B. Zhuang, S. Wang, and J. Xiao, \u201cA hierarchical attention based seq2seq model for Chinese lyrics generation,\u201d Proceedings of the 16th Pacific Rim International Conference on Artificial Intelligence (PRICAI), pp.279-288, 2019. 10.1007\/978-3-030-29894-4_23","DOI":"10.1007\/978-3-030-29894-4_23"},{"key":"34","unstructured":"[34] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, \u201cAn image is worth 16x16 words: Transformers for image recognition at scale,\u201d Proceedings of the 9th International Conference on Learning Representations (ICLR), 2021."},{"key":"35","unstructured":"[35] I. Loshchilov and F. Hutter, \u201cDecoupled weight decay regularization,\u201d Proceedings of the 7th International Conference on Learning Representations (ICLR), 2019."},{"key":"36","unstructured":"[36] J. Devlin, M. Chang, K. Lee, and K. Toutanova, \u201cBERT: Pre-training of deep bidirectional transformers for language understanding,\u201d Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pp.4171-4186, 2019."},{"key":"37","unstructured":"[37] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, \u201cLanguage models are unsupervised multitask learners,\u201d OpenAI blog, vol.1, no.8, p.9, 2019."},{"key":"38","doi-asserted-by":"crossref","unstructured":"[38] T. Kudo and Y. Matsumoto, \u201cJapanese dependency analysis using cascaded chunking,\u201d Proceedings of the 6th Conference on Natural Language Learning (CoNLL), vol.20, pp.1-7, 2002. 10.3115\/1118853.1118869","DOI":"10.3115\/1118853.1118869"},{"key":"39","unstructured":"[39] A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, \u201cThe curious case of neural text degeneration,\u201d Proceedings of the 8th International Conference on Learning Representations (ICLR), 2020."},{"key":"40","doi-asserted-by":"publisher","unstructured":"[40] Y. Li and B. Liu, \u201cA normalized Levenshtein distance metric,\u201d IEEE Transactions on Pattern Analysis and Machine Intelligence, vol.29, no.6, pp.1091-1095, 2007. 10.1109\/tpami.2007.1078","DOI":"10.1109\/TPAMI.2007.1078"},{"key":"41","unstructured":"[41] C. Manning and H. Schutze, Foundations of statistical natural language processing, MIT press, 1999."},{"key":"42","unstructured":"[42] M. Marelli, S. Menini, M. Baroni, L. Bentivogli, R. Bernardi, and R. Zamparelli, \u201cA SICK cure for the evaluation of compositional distributional semantic models,\u201d Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC), pp.216-223, 2014."},{"key":"43","unstructured":"[43] M. Goto, H. Hashiguchi, T. Nishimura, and R. Oka, \u201cRWC Music Database: Popular, classical and jazz music databases,\u201d Proceedings of the 3rd International Conference on Music Information Retrieval (ISMIR), pp.287-288, 2002."}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E108.D\/5\/E108.D_2024EDP7104\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,3]],"date-time":"2025-05-03T03:46:51Z","timestamp":1746244011000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E108.D\/5\/E108.D_2024EDP7104\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,1]]},"references-count":43,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2025]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2024edp7104","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"type":"print","value":"0916-8532"},{"type":"electronic","value":"1745-1361"}],"subject":[],"published":{"date-parts":[[2025,5,1]]},"article-number":"2024EDP7104"}}