{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T02:42:39Z","timestamp":1760150559935,"version":"build-2065373602"},"reference-count":50,"publisher":"MDPI AG","issue":"22","license":[{"start":{"date-parts":[[2023,11,20]],"date-time":"2023-11-20T00:00:00Z","timestamp":1700438400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001691","name":"Japan Society for the Promotion of Science (JSPS)","doi-asserted-by":"publisher","award":["JP21H03456","JP23K11141","JP23K11211"],"award-info":[{"award-number":["JP21H03456","JP23K11141","JP23K11211"]}],"id":[{"id":"10.13039\/501100001691","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>At present, text-guided image manipulation is a notable subject of study in the vision and language field. Given an image and text as inputs, these methods aim to manipulate the image according to the text, while preserving text-irrelevant regions. Although there has been extensive research to improve the versatility and performance of text-guided image manipulation, research on its performance evaluation is inadequate. This study proposes Manipulation Direction (MD), a logical and robust metric, which evaluates the performance of text-guided image manipulation by focusing on changes between image and text modalities. Specifically, we define MD as the consistency of changes between images and texts occurring before and after manipulation. By using MD to evaluate the performance of text-guided image manipulation, we can comprehensively evaluate how an image has changed before and after the image manipulation and whether this change agrees with the text. Extensive experiments on Multi-Modal-CelebA-HQ and Caltech-UCSD Birds confirmed that there was an impressive correlation between our calculated MD scores and subjective scores for the manipulated images compared to the existing metrics.<\/jats:p>","DOI":"10.3390\/s23229287","type":"journal-article","created":{"date-parts":[[2023,11,20]],"date-time":"2023-11-20T11:31:36Z","timestamp":1700479896000},"page":"9287","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Manipulation Direction: Evaluating Text-Guided Image Manipulation Based on Similarity between Changes in Image and Text Modalities"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9841-4089","authenticated-orcid":false,"given":"Yuto","family":"Watanabe","sequence":"first","affiliation":[{"name":"Graduate School of Information Science and Technology, Hokkaido University, N-14, W-9, Kita-ku, Sapporo 060-0814, Hokkaido, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4474-3995","authenticated-orcid":false,"given":"Ren","family":"Togo","sequence":"additional","affiliation":[{"name":"Faculty of Information Science and Technology, Hokkaido University, N-14, W-9, Kita-ku, Sapporo 060-0814, Hokkaido, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8039-3462","authenticated-orcid":false,"given":"Keisuke","family":"Maeda","sequence":"additional","affiliation":[{"name":"Faculty of Information Science and Technology, Hokkaido University, N-14, W-9, Kita-ku, Sapporo 060-0814, Hokkaido, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5332-8112","authenticated-orcid":false,"given":"Takahiro","family":"Ogawa","sequence":"additional","affiliation":[{"name":"Faculty of Information Science and Technology, Hokkaido University, N-14, W-9, Kita-ku, Sapporo 060-0814, Hokkaido, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1496-1761","authenticated-orcid":false,"given":"Miki","family":"Haseyama","sequence":"additional","affiliation":[{"name":"Faculty of Information Science and Technology, Hokkaido University, N-14, W-9, Kita-ku, Sapporo 060-0814, Hokkaido, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,11,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_2","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA."},{"key":"ref_3","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., and Parikh, D. (2015, January 7\u201313). VQA: Visual question answering. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.279"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Shao, Z., Yu, Z., Wang, M., and Yu, J. (2023, January 18\u201322). Prompting large language models with answer heuristics for knowledge-based visual question answering. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01438"},{"key":"ref_6","unstructured":"Frome, A., Corrado, G.S., Shlens, J., Bengio, S., Dean, J., Ranzato, M., and Mikolov, T. (2013, January 5\u20138). Devise: A deep visual-semantic embedding model. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Stateline, NV, USA."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Diao, H., Zhang, Y., Ma, L., and Lu, H. (2021, January 2\u20139). Similarity reasoning and filtration for image-text matching. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual.","DOI":"10.1609\/aaai.v35i2.16209"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Donahue, J., Anne Hendricks, L., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., and Darrell, T. (2015, January 7\u201312). Long-term recurrent convolutional networks for visual recognition and description. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298878"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Luo, J., Li, Y., Pan, Y., Yao, T., Feng, J., Chao, H., and Mei, T. (2023, January 18\u201322). Semantic-conditional diffusion networks for image captioning. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02237"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Hu, R., Rohrbach, M., and Darrell, T. (2016, January 11\u201314). Segmentation from natural language expressions. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_7"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Wang, Z., Lu, Y., Li, Q., Tao, X., Guo, Y., Gong, M., and Liu, T. (2022, January 19\u201324). CRIS: Clip-driven referring image segmentation. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01139"},{"key":"ref_12","unstructured":"Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., and Lee, H. (2016, January 19\u201324). Generative adversarial text to image synthesis. Proceedings of the International Conference on Machine Learning (ICML), New York City, NY, USA."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Kang, M., Zhu, J.Y., Zhang, R., Park, J., Shechtman, E., Paris, S., and Park, T. (2023, January 18\u201322). Scaling up GANs for text-to-image synthesis. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00976"},{"key":"ref_14","unstructured":"Kingma, D.P., and Welling, M. (2013). Auto-encoding variational bayes. arXiv."},{"key":"ref_15","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014, January 8\u201313). Generative adversarial nets. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montreal, QC, Canada."},{"key":"ref_16","unstructured":"Ho, J., Jain, A., and Abbeel, P. (2020, January 6\u201312). Denoising diffusion probabilistic models. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual."},{"key":"ref_17","unstructured":"Nam, S., Kim, Y., and Kim, S.J. (2018, January 3\u20138). Text-adaptive generative adversarial networks: Manipulating images with natural language. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montreal, QC, Canada."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Li, B., Qi, X., Lukasiewicz, T., and Torr, P.H. (2020, January 14\u201319). ManiGAN: Text-guided image manipulation. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Virtual.","DOI":"10.1109\/CVPR42600.2020.00790"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Li, B., Qi, X., Torr, P.H., and Lukasiewicz, T. (2020, January 6\u201312). Lightweight generative adversarial networks for text-guided image manipulation. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual.","DOI":"10.1109\/CVPR42600.2020.00790"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Haruyama, T., Togo, R., Maeda, K., Ogawa, T., and Haseyama, M. (2021, January 19\u201322). Segmentation-aware text-guided image manipulation. Proceedings of the IEEE International Conference on Image Processing (ICIP), Anchorage, AK, USA.","DOI":"10.1109\/ICIP42928.2021.9506601"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Xia, W., Yang, Y., Xue, J.H., and Wu, B. (2021, January 19\u201325). TediGAN: Text-guided diverse face image generation and manipulation. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.00229"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Patashnik, O., Wu, Z., Shechtman, E., Cohen-Or, D., and Lischinski, D. (2021, January 11\u201317). StyleCLIP: Text-driven manipulation of stylegan imagery. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Virtual.","DOI":"10.1109\/ICCV48922.2021.00209"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Kocasari, U., Dirik, A., Tiftikci, M., and Yanardag, P. (2022, January 4\u20138). StyleMC: Multi-channel based fast text-guided image generation and manipulation. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA.","DOI":"10.1109\/WACV51458.2022.00350"},{"key":"ref_24","unstructured":"Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., and Cohen-Or, D. (2022). Prompt-to-prompt image editing with cross attention control. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Brooks, T., Holynski, A., and Efros, A.A. (2023, January 18\u201322). InstructPix2Pix: Learning to follow image editing instructions. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01764"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Cheng, Z., Yang, Q., and Sheng, B. (2015, January 7\u201313). Deep colorization. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.55"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhang, R., Isola, P., and Efros, A.A. (2016, January 11\u201314). Colorful image colorization. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46487-9_40"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., and Efros, A.A. (2016, January 27\u201330). Context encoders: Feature learning by inpainting. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.278"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Lempitsky, V., Vedaldi, A., and Ulyanov, D. (2018, January 18\u201323). Deep image prior. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00984"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Gatys, L.A., Ecker, A.S., and Bethge, M. (2015). A neural algorithm of artistic style. arXiv.","DOI":"10.1167\/16.12.326"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Madaan, A., Setlur, A., Parekh, T., Poczos, B., Neubig, G., Yang, Y., Salakhutdinov, R., Black, A.W., and Prabhumoye, S. (2020). Politeness transfer: A tag and generate approach. arXiv.","DOI":"10.18653\/v1\/2020.acl-main.169"},{"key":"ref_32","unstructured":"Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S. (2011). The Caltech-UCSD Birds-200-2011 Dataset, California Institute of Technology."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","article-title":"Image quality assessment: From error visibility to structural similarity","volume":"13","author":"Wang","year":"2004","journal-title":"IEEE Trans. Image Process."},{"key":"ref_34","unstructured":"Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X. (2016, January 5\u201310). Improved techniques for training GANs. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Barcelona, Spain."},{"key":"ref_35","unstructured":"Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. (2017, January 4\u20139). GANs trained by a two time-scale update rule converge to a local nash equilibrium. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Watanabe, Y., Togo, R., Maeda, K., Ogawa, T., and Haseyama, M. (2022, January 16\u201319). Assessment of image manipulation using natural language description: Quantification of manipulation direction. Proceedings of the IEEE International Conference on Image Processing (ICIP), Bordeaux, France.","DOI":"10.1109\/ICIP46576.2022.9897900"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Dong, H., Yu, S., Wu, C., and Guo, Y. (2017, January 22\u201329). Semantic image synthesis via adversarial learning. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.608"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Nilsback, M.E., and Zisserman, A. (2008, January 16\u201319). Automated flower classification over a large number of classes. Proceedings of the 6th Indian Conference on Computer Vision, Graphics and Image Processing (ICVGIP), Bhubaneswar, India.","DOI":"10.1109\/ICVGIP.2008.47"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft COCO: Common objects in context. Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Karras, T., Laine, S., and Aila, T. (2019, January 16\u201320). A style-based generator architecture for generative adversarial networks. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00453"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 27\u201330). Rethinking the inception architecture for computer vision. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Xu, T., Zhang, P., Huang, Q., Zhang, H., Gan, Z., Huang, X., and He, X. (2018, January 18\u201323). AttnGAN: Fine-grained text to image generation with attentional generative adversarial networks. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00143"},{"key":"ref_43","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning (ICML), Virtual."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition Conference (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_45","unstructured":"Karras, T., Aila, T., Laine, S., and Lehtinen, J. (2017). Progressive growing of GANs for improved quality, stability, and variation. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"3998","DOI":"10.1109\/TIP.2018.2831899","article-title":"NIMA: Neural image assessment","volume":"27","author":"Talebi","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1023\/A:1026543900054","article-title":"The earth mover\u2019s distance as a metric for image retrieval","volume":"40","author":"Rubner","year":"2000","journal-title":"Int. J. Comput. Vis."},{"key":"ref_48","unstructured":"Kiros, R., Salakhutdinov, R., and Zemel, R.S. (2014). Unifying visual-semantic embeddings with multimodal neural language models. arXiv."},{"key":"ref_49","first-page":"67","article-title":"From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions","volume":"2","author":"Young","year":"2014","journal-title":"Trans. Assoc. Comput. Linguist. TACL"},{"key":"ref_50","unstructured":"Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C., and Dosovitskiy, A. (2021, January 6\u201314). Do vision transformers see like convolutional neural networks?. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/22\/9287\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:26:29Z","timestamp":1760131589000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/22\/9287"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,20]]},"references-count":50,"journal-issue":{"issue":"22","published-online":{"date-parts":[[2023,11]]}},"alternative-id":["s23229287"],"URL":"https:\/\/doi.org\/10.3390\/s23229287","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2023,11,20]]}}}