{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T12:25:19Z","timestamp":1783427119996,"version":"3.54.6"},"reference-count":30,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2024,4,16]],"date-time":"2024-04-16T00:00:00Z","timestamp":1713225600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,4,16]],"date-time":"2024-04-16T00:00:00Z","timestamp":1713225600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Process Lett"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Metaphor has significant implications for revealing cognitive and thinking mechanisms. Visual metaphor image generation not only presents metaphorical connotations intuitively but also reflects AI\u2019s understanding of metaphor through the generated images. This paper investigates the task of generating images based on text with visual metaphors. We explore metaphor image generation and create a dataset containing sentences with visual metaphors. Then, we propose a visual metaphor generation image framework based on metaphor understanding, which is more tailored to the essence of metaphor, better utilizes visual features, and has stronger interpretability. Specifically, the framework extracts the source domain, target domain, and metaphor interpretation from metaphorical sentences, separating the elements of the metaphor to deepen the understanding of its themes and intentions. Additionally, the framework introduces image data from the source domain to capture visual similarities and generate visual enhancement prompts specific to the domain. Finally, these prompts are combined with metaphorical interpretation sentences to form the final prompt text. Experimental results demonstrate that this approach effectively captures the essence of metaphor and generates metaphorical images consistent with the textual meaning.<\/jats:p>","DOI":"10.1007\/s11063-024-11609-w","type":"journal-article","created":{"date-parts":[[2024,4,16]],"date-time":"2024-04-16T19:01:27Z","timestamp":1713294087000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Efficient Visual Metaphor Image Generation Based on Metaphor Understanding"],"prefix":"10.1007","volume":"56","author":[{"given":"Chang","family":"Su","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xingyue","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shupin","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yijiang","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,4,16]]},"reference":[{"key":"11609_CR1","doi-asserted-by":"crossref","unstructured":"Hessel J, Marasovi\u0107 A, Hwang JD, Lee L, Da J, Zellers R, Mankoff R, Choi Y (2023) Do androids laugh at electric sheep? Humor \u201cunderstanding\u201d benchmarks from the new yorker caption contest","DOI":"10.18653\/v1\/2023.acl-long.41"},{"key":"11609_CR2","doi-asserted-by":"crossref","unstructured":"Yuri B, Simon D (2020) Sky + fire = sunset. exploring parallels between visually grounded metaphors and image classifiers. In: Beigman KB, Ekaterina S, Patricia L, Smaranda M, Chee W, Anna F, Debanjan G (eds) Proceedings of the second workshop on figurative language processing, pp 126\u2013135, Online. Association for Computational Linguistics","DOI":"10.18653\/v1\/2020.figlang-1.19"},{"key":"11609_CR3","unstructured":"Robin R, Andreas B, Dominik L, Patrick E, Bj\u00f6rn O (2022) High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 10684\u201310695"},{"key":"11609_CR4","unstructured":"Aditya R, Mikhail P, Gabriel G, Scott G, Chelsea V, Alec R, Mark C, Ilya S (2021) Zero-shot text-to-image generation. In: International conference on machine learning, pp 8821\u20138831. PMLR"},{"key":"11609_CR5","unstructured":"Alex N, Prafulla D, Aditya R, Pranav S, Pamela M, Bob M, Ilya S, Mark C (2022) Glide: Towards photorealistic image generation and editing with text-guided diffusion models arxiv:2205.13168v1"},{"key":"11609_CR6","unstructured":"Yu J, Xu Y, Koh JY, Luong T, Baid G, Wang Z, Vasudevan V, Ku A, Yang Y, Ayan BK, Hutchinson B (2022) Scaling autoregressive models for content-rich text-to-image generation"},{"key":"11609_CR7","doi-asserted-by":"crossref","unstructured":"Saharia C, Chan W, Saxena S, Li L, Whang J, Denton EL, Ghasemipour K, Gontijo Lopes R, Karagol Ayan B, Salimans T, Ho J (2022) Photorealistic text-to-image diffusion models with deep language understanding","DOI":"10.1145\/3528233.3530757"},{"key":"11609_CR8","unstructured":"Gal R, Alaluf Y, Atzmon Y, Patashnik O, Bermano AH, Chechik G, Cohen-Or D (2022) An image is worth one word: Personalizing text-to-image generation using textual inversion"},{"key":"11609_CR9","doi-asserted-by":"crossref","unstructured":"Chakrabarty T, Saakyan A, Winn O, Panagopoulou A, Yang Y, Apidianaki M, Muresan S(2023) I spy a metaphor: large language models and diffusion models co-create visual metaphors","DOI":"10.18653\/v1\/2023.findings-acl.465"},{"key":"11609_CR10","doi-asserted-by":"publisher","first-page":"52138","DOI":"10.1109\/ACCESS.2018.2870052","volume":"6","author":"A Adadi","year":"2018","unstructured":"Adadi A, Berrada M (2018) Peeking inside the black-box: a survey on explainable artificial intelligence (xai). IEEE Access 6:52138\u201352160","journal-title":"IEEE Access"},{"key":"11609_CR11","doi-asserted-by":"crossref","unstructured":"Gu Y, Han X, Liu Z, Huang M (2022) Ppt: pre-trained prompt tuning for few-shot learning","DOI":"10.18653\/v1\/2022.acl-long.576"},{"key":"11609_CR12","doi-asserted-by":"crossref","unstructured":"Qian J, Dong L, Shen Y, Wei F, Chen W (2022) Controllable natural language generation with contrastive prefixes","DOI":"10.18653\/v1\/2022.findings-acl.229"},{"key":"11609_CR13","unstructured":"An S, Li Y, Lin Z, Liu Q, Chen B, Fu Q, Chen W, Zheng N, Lou J-G (2022) Input-tuning: adapting unfamiliar inputs to frozen pretrained models"},{"key":"11609_CR14","doi-asserted-by":"crossref","unstructured":"Khashabi D, Lyu S, Min S, Qin L, Richardson K, Welleck S, Hajishirzi H, Khot T, Sabharwal A, Singh S, Choi Y (2022) Prompt waywardness: the curious case of discretized interpretation of continuous prompts","DOI":"10.18653\/v1\/2022.naacl-main.266"},{"key":"11609_CR15","doi-asserted-by":"crossref","unstructured":"Petroni F, Rockt\u00e4schel T, Lewis P, Bakhtin A, Wu Y, Miller AH, Riedel S (2019) Language models as knowledge bases?","DOI":"10.18653\/v1\/D19-1250"},{"key":"11609_CR16","doi-asserted-by":"crossref","unstructured":"Schick T, Sch\u00fctze H (2021) Exploiting cloze questions for few shot text classification and natural language inference","DOI":"10.18653\/v1\/2021.eacl-main.20"},{"key":"11609_CR17","doi-asserted-by":"crossref","unstructured":"Shin T, Razeghi Y, Logan IV RL, Wallace E, Singh S (2020) Autoprompt: eliciting knowledge from language models with automatically generated prompts","DOI":"10.18653\/v1\/2020.emnlp-main.346"},{"key":"11609_CR18","doi-asserted-by":"crossref","unstructured":"Deng M, Wang J, Hsieh C-P, Wang Y, Guo H, Shu T, Song M, Xing EP, Hu Z (2022) Rlprompt: optimizing discrete text prompts with reinforcement learning","DOI":"10.18653\/v1\/2022.emnlp-main.222"},{"key":"11609_CR19","doi-asserted-by":"crossref","unstructured":"Brysbaert M, Warriner AB, Kuperman V (2014) Concreteness ratings for 40 thousand generally known english word lemmas. Behav Res Methods 904\u2013911","DOI":"10.3758\/s13428-013-0403-5"},{"key":"11609_CR20","doi-asserted-by":"publisher","first-page":"32","DOI":"10.1007\/s11263-016-0981-7","volume":"123","author":"R Krishna","year":"2017","unstructured":"Krishna R, Zhu Y, Groth O, Johnson J, Hata K, Kravitz J, Chen S, Kalantidis Y, Li LJ, Shamma DA, Bernstein MS (2017) Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int J Comput Vis 123:32\u201373","journal-title":"Int J Comput Vis"},{"key":"11609_CR21","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li L-J, Li K, Fei-Fei L (2009) Imagenet: a large-scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"11609_CR22","unstructured":"Li Y, Lin C, Guerin F (2022) Cm-gen: a neural framework for Chinese metaphor generation with explicit context modelling. In: Proceedings of the 29th international conference on computational linguistics, pp 6468\u20136479"},{"key":"11609_CR23","doi-asserted-by":"publisher","first-page":"300","DOI":"10.1016\/j.neucom.2016.09.030","volume":"219","author":"S Chang","year":"2017","unstructured":"Chang S, Huang S, Chen Y (2017) Automatic detection and interpretation of nominal metaphor based on the theory of meaning. Neurocomputing 219:300\u2013311","journal-title":"Neurocomputing"},{"key":"11609_CR24","unstructured":"Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, et al (2021) Learning transferable visual models from natural language supervision. In: International conference on machine learning, pp 8748\u20138763. PMLR"},{"key":"11609_CR25","unstructured":"Schuhmann C, Beaumont R, Vencu R, Gordon C, Wightman R, Cherti M, Coombes T, Katta A, Mullis C, Wortsman M, et al (2022) Laion-5b: an open large-scale dataset for training next generation image-text models. arXiv preprint arXiv:2210.08402"},{"key":"11609_CR26","unstructured":"Dosovitskiy Alexey, Beyer Lucas, Kolesnikov Alexander, Weissenborn Dirk, Zhai Xiaohua, Unterthiner Thomas, Dehghani Mostafa, Minderer Matthias, Heigold Georg, Gelly Sylvain, et al (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929"},{"key":"11609_CR27","unstructured":"Loshchilov I, Hutter F (2017) Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101"},{"key":"11609_CR28","doi-asserted-by":"crossref","unstructured":"Hessel J, Holtzman A, Forbes M, Bras RL, Choi Y (2021) Clipscore: a reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718","DOI":"10.18653\/v1\/2021.emnlp-main.595"},{"key":"11609_CR29","unstructured":"Zhang T, Kishore V, Wu F, Weinberger KQ, Artzi Y (2019) Bertscore: evaluating text generation with bert. arXiv preprint arXiv:1904.09675"},{"key":"11609_CR30","unstructured":"Ridnik T, Ben-Baruch E, Noy A, Zelnik-Manor L (2021) Imagenet-21k pretraining for the masses"}],"container-title":["Neural Processing Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-024-11609-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11063-024-11609-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-024-11609-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,16]],"date-time":"2024-11-16T11:00:45Z","timestamp":1731754845000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11063-024-11609-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,16]]},"references-count":30,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2024,6]]}},"alternative-id":["11609"],"URL":"https:\/\/doi.org\/10.1007\/s11063-024-11609-w","relation":{},"ISSN":["1573-773X"],"issn-type":[{"value":"1573-773X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,16]]},"assertion":[{"value":"27 March 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 April 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"150"}}