{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,24]],"date-time":"2025-11-24T16:46:56Z","timestamp":1764002816390,"version":"build-2065373602"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"11","funder":[{"name":"Strategic Priority Research Program of the Chinese Academy of Sciences","award":["XDB0660101, XDB0660000, and XDB0660100"],"award-info":[{"award-number":["XDB0660101, XDB0660000, and XDB0660100"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62172380"],"award-info":[{"award-number":["62172380"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100004608","name":"Jiangsu Provincial Natural Science Foundation","doi-asserted-by":"crossref","award":["BK20241818"],"award-info":[{"award-number":["BK20241818"]}],"id":[{"id":"10.13039\/501100004608","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100004739","name":"Youth Innovation Promotion Association CAS","doi-asserted-by":"crossref","award":["Y2021121"],"award-info":[{"award-number":["Y2021121"]}],"id":[{"id":"10.13039\/501100004739","id-type":"DOI","asserted-by":"crossref"}]},{"name":"USTC Research Funds of the Double First-Class Initiative","award":["YD2150002011"],"award-info":[{"award-number":["YD2150002011"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>\n                    Visual Concept Implantation (VCI) is essential in text-to-image fields. While VCI methods in diffusion models have matured, Visual Concept Disentanglement (VCD) remains an unexplored area. VCD involves analyzing Prompt Spaces trained in VCI to produce disentangled SubPrompts or SubCones for exploring interpretability. However, challenges arise due to Prompt Space design complexity and feature information extraction. We propose Picasso, a unified framework for VCD in diffusion models (DM). Our contributions include: for performance evaluation: Transforming VCD in DM into a regular clustering task by Visualization based on SubCones (VbSC); for unified framework design: Picasso processes diverse Prompt Spaces using spatially clusterable features; for Prompt Space exploration: Introducing a temporal-spatial SOTA Prompt Design subset based on temporal features. Our method provides a feasible mechanism for VCD in DM. Through functional and interpretability validation methods, we will comprehensively evaluate the effectiveness of our proposed method in visual concept implantation tasks and verify the correctness of the parameter space design principles. Picasso will be released at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/haoyu-cai.github.io\/our_picasso\">https:\/\/haoyu-cai.github.io\/our_picasso<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3724122","type":"journal-article","created":{"date-parts":[[2025,5,27]],"date-time":"2025-05-27T12:39:42Z","timestamp":1748349582000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Picasso: Analyzing Prompt Design for Text-to-Image Generative Diffusion Models from a Temporal-Spatial Perspective"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5010-3396","authenticated-orcid":false,"given":"Haoyu","family":"Cai","sequence":"first","affiliation":[{"name":"The School of Computer Science, University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2240-6672","authenticated-orcid":false,"given":"Wenqi","family":"Lou","sequence":"additional","affiliation":[{"name":"The School of Software Engineering, University of Science and Technology of China, Hefei, China and the Suzhou Institute for Advanced Research, University of Science and Technology of China, Suzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9403-5575","authenticated-orcid":false,"given":"Chao","family":"Wang","sequence":"additional","affiliation":[{"name":"The School of Computer Science, University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8360-3143","authenticated-orcid":false,"given":"Xuehai","family":"Zhou","sequence":"additional","affiliation":[{"name":"The School of Computer Science, University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,11,7]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","unstructured":"Omri Avrahami Kfir Aberman Ohad Fried Daniel Cohen-Or and Dani Lischinski. 2023. Break-A-Scene: Extracting multiple concepts from a single image. arXiv:2305.16311. Retrieved from https:\/\/arxiv.org\/abs\/2305.16311","DOI":"10.1145\/3610548.3618154"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01767"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.108102"},{"key":"e_1_3_2_5_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Baranchuk Dmitry","year":"2021","unstructured":"Dmitry Baranchuk, Andrey Voynov, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. 2021. Label-efficient semantic segmentation with diffusion models. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_6_2","volume-title":"Pattern Recognition with Fuzzy Objective Function Algorithms","author":"Bezdek James C.","year":"2013","unstructured":"James C. Bezdek. 2013. Pattern Recognition with Fuzzy Objective Function Algorithms. Springer Science & Business Media."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3630258"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01212"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00581"},{"key":"e_1_3_2_10_2","first-page":"8780","article-title":"Diffusion models beat GANs on image synthesis","volume":"34","author":"Dhariwal Prafulla","year":"2021","unstructured":"Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems, Vol. 34, 8780\u20138794.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3479238"},{"key":"e_1_3_2_12_2","unstructured":"Zhengcong Fei Mingyuan Fan and Junshi Huang. 2023. Gradient-free textual inversion. arXiv:2304.05818. Retrieved from https:\/\/arxiv.org\/abs\/2304.05818"},{"key":"e_1_3_2_13_2","volume-title":"Proceedings of the 111th International Conference on Learning Representations","author":"Gal Rinon","year":"2022","unstructured":"Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit Haim Bermano, Gal Chechik, and Daniel Cohen-or. 2022. An image is worth one word: Personalizing text-to-image generation using textual inversion. In Proceedings of the 111th International Conference on Learning Representations."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3422622"},{"key":"e_1_3_2_15_2","unstructured":"Ligong Han Yinxiao Li Han Zhang Peyman Milanfar Dimitris Metaxas and Feng Yang. 2023. Svdiff: Compact parameter space for diffusion fine-tuning. arXiv:2303.11305. Retrieved from https:\/\/arxiv.org\/abs\/2303.11305"},{"key":"e_1_3_2_16_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Hertz Amir","year":"2022","unstructured":"Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-or. 2022. Prompt-to-prompt image editing with cross-attention control. In Proceedings of the 11th International Conference on Learning Representations."},{"key":"e_1_3_2_17_2","first-page":"6840","article-title":"Denoising diffusion probabilistic models","volume":"33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, Vol. 33, 6840\u20136851.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_18_2","unstructured":"Nisha Huang Fan Tang Weiming Dong Tong-Yee Lee and Changsheng Xu .2023. Region-aware diffusion for zero-shot text-driven image editing. arXiv:2302.11797. Retrieved from https:\/\/arxiv.org\/abs\/2302.11797"},{"key":"e_1_3_2_19_2","unstructured":"Ziqi Huang Tianxing Wu Yuming Jiang Kelvin C. K. Chan and Ziwei Liu. 2023. ReVersion: Diffusion-based relation inversion from images. arXiv:2303.13495. Retrieved from https:\/\/arxiv.org\/abs\/2303.13495"},{"key":"e_1_3_2_20_2","unstructured":"Jaeseok Jeong Mingi Kwon and Youngjung Uh. 2023. Training-free style transfer emerges from h-space in diffusion models. arXiv:2303.15403. Retrieved from https:\/\/arxiv.org\/abs\/2303.15403"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00192"},{"key":"e_1_3_2_22_2","first-page":"1280","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"36","author":"Li Chao","year":"2022","unstructured":"Chao Li, Kelu Yao, Jin Wang, Boyu Diao, Yongjun Xu, and Quanshi Zhang. 2022. Interpretable generative adversarial networks. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, 1280\u20131288."},{"key":"e_1_3_2_23_2","unstructured":"Dongxu Li Junnan Li and Steven C. H. Hoi. 2023. Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing. arXiv:2305.14720. Retrieved from https:\/\/arxiv.org\/abs\/2305.14720"},{"key":"e_1_3_2_24_2","unstructured":"Senmao Li Joost van de Weijer Taihang Hu Fahad Shahbaz Khan Qibin Hou Yaxing Wang and Jian Yang. 2023. StyleDiffusion: Prompt-embedding inversion for text-based editing. arXiv:2303.15649. Retrieved from https:\/\/arxiv.org\/abs\/2303.15649"},{"key":"e_1_3_2_25_2","unstructured":"Zhiheng Liu Ruili Feng Kai Zhu Yifei Zhang Kecheng Zheng Yu Liu Deli Zhao Jingren Zhou and Yang Cao. 2023. Cones: Concept neurons in diffusion models for customized generation. arXiv:2303.05125. Retrieved from https:\/\/arxiv.org\/abs\/2303.05125"},{"key":"e_1_3_2_26_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Meng Chenlin","year":"2021","unstructured":"Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. 2021. SDEdit: Guided image synthesis and editing with stochastic differential equations. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00585"},{"key":"e_1_3_2_28_2","doi-asserted-by":"crossref","unstructured":"Chong Mou Xintao Wang Liangbin Xie Jian Zhang Zhongang Qi Ying Shan and Xiaohu Qie. 2023. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. arXiv:2302.08453. Retrieved from https:\/\/arxiv.org\/abs\/2302.08453","DOI":"10.1609\/aaai.v38i5.28226"},{"key":"e_1_3_2_29_2","doi-asserted-by":"crossref","unstructured":"Hadas Orgad Bahjat Kawar and Yonatan Belinkov. 2023. Editing implicit assumptions in text-to-image diffusion models. arXiv:2303.08084. Retrieved from https:\/\/arxiv.org\/abs\/2303.08084","DOI":"10.1109\/ICCV51070.2023.00649"},{"key":"e_1_3_2_30_2","unstructured":"Yong-Hyun Park Mingi Kwon Junghyo Jo and Youngjung Uh. 2023. Unsupervised discovery of semantic latent directions in diffusion models. arXiv:2302.12469. Retrieved from https:\/\/arxiv.org\/abs\/2302.12469"},{"key":"e_1_3_2_31_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_2_32_2","unstructured":"Aditya Ramesh Prafulla Dhariwal Alex Nichol Casey Chu and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv:2204.06125. Retrieved from https:\/\/arxiv.org\/abs\/2204.06125"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3472291"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"e_1_3_2_36_2","first-page":"36479","article-title":"Photorealistic text-to-image diffusion models with deep language understanding","volume":"35","author":"Saharia Chitwan","year":"2022","unstructured":"Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding. In Advances in Neural Information Processing Systems, Vol. 35, 36479\u201336494.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_37_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Song Jiaming","year":"2020","unstructured":"Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020. Denoising diffusion implicit models. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_38_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Song Yang","year":"2020","unstructured":"Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2020. Score-based generative modeling through stochastic differential equations. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_39_2","unstructured":"Andrey Voynov Qinghao Chu Daniel Cohen-Or and Kfir Aberman. 2023. \\(P+\\) : Extended textual conditioning in text-to-image generation. arXiv:2303.09522. Retrieved from https:\/\/arxiv.org\/abs\/2303.09522"},{"issue":"9","key":"e_1_3_2_40_2","doi-asserted-by":"crossref","first-page":"2082","DOI":"10.1109\/TPAMI.2019.2911937","article-title":"A novel dynamic model capturing spatial and temporal patterns for facial expression analysis","volume":"42","author":"Wang Shangfei","year":"2019","unstructured":"Shangfei Wang, Zhuangqiang Zheng, Shi Yin, Jiajia Yang, and Qiang Ji. 2019. A novel dynamic model capturing spatial and temporal patterns for facial expression analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 9 (2019), 2082\u20132095.","journal-title":". IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2022.3181070"},{"key":"e_1_3_2_42_2","unstructured":"Jianan Yang Haobo Wang Ruixuan Xiao Sai Wu Gang Chen and Junbo Zhao .2023. Controllable textual inversion for personalized text-to-image generation. arXiv:2304.05265. Retrieved from https:\/\/arxiv.org\/abs\/2304.05265"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626235"},{"key":"e_1_3_2_44_2","unstructured":"Yiyuan Yang Ming Jin Haomin Wen Chaoli Zhang Yuxuan Liang Lintao Ma Yi Wang Chenghao Liu Bin Yang Zenglin Xu et al. 2024. A survey on diffusion models for time series and spatio-temporal data. arXiv:2404.18886. Retrieved from https:\/\/arxiv.org\/abs\/2404.18886"},{"key":"e_1_3_2_45_2","unstructured":"Hu Ye Jun Zhang Sibo Liu Xiao Han and Wei Yang. 2023. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv:2308.06721. Retrieved from https:\/\/arxiv.org\/abs\/2308.06721"},{"key":"e_1_3_2_46_2","first-page":"528","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"Zeng Yanhong","year":"2020","unstructured":"Yanhong Zeng, Jianlong Fu, and Hongyang Chao. 2020. Learning joint spatial-temporal transformations for video inpainting. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920). Springer, 528\u2013543."},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00920"},{"key":"e_1_3_2_49_2","unstructured":"Yuxin Zhang Weiming Dong Fan Tang Nisha Huang Haibin Huang Chongyang Ma Tong-Yee Lee Oliver Deussen and Changsheng Xu. 2023. ProSpect: Expanded conditioning for the personalization of attribute-aware image generation. arXiv:2305.16225. Retrieved from https:\/\/arxiv.org\/abs\/2305.16225"},{"issue":"3","key":"e_1_3_2_50_2","first-page":"3311","article-title":"Label-guided generative adversarial network for realistic image synthesis","volume":"45","author":"Zhu Junchen","year":"2022","unstructured":"Junchen Zhu, Lianli Gao, Jingkuan Song, Yuan-Fang Li, Feng Zheng, Xuelong Li, and Heng Tao Shen. 2022. Label-guided generative adversarial network for realistic image synthesis. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 3 (2022), 3311\u20133328.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3724122","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T15:10:11Z","timestamp":1762528211000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3724122"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,7]]},"references-count":49,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3724122"],"URL":"https:\/\/doi.org\/10.1145\/3724122","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2025,11,7]]},"assertion":[{"value":"2024-05-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}