{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T06:10:44Z","timestamp":1784268644906,"version":"3.55.0"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2025,8,1]]},"abstract":"<jats:p>We present TokenVerse - a method for multi-concept personalization, leveraging a pre-trained text-to-image diffusion model. Our framework can disentangle complex visual elements and attributes from as little as a single image, while enabling seamless plug-and-play generation of combinations of concepts extracted from multiple images. As opposed to existing works, TokenVerse can handle multiple images with multiple concepts each, and supports a wide-range of concepts, including objects, accessories, materials, pose, and lighting. Our work exploits a DiT-based text-to-image model, in which the input text affects the generation through both attention and modulation (shift and scale). We observe that the modulation space is semantic and enables localized control over complex concepts. Building on this insight, we devise an optimization-based framework that takes as input an image and a text description, and finds for each word a distinct direction in the modulation space. These directions can then be used to generate new images that combine the learned concepts in a desired configuration. We demonstrate the effectiveness of TokenVerse in challenging personalization settings, and showcase its advantages over existing methods.<\/jats:p>","DOI":"10.1145\/3730843","type":"journal-article","created":{"date-parts":[[2025,7,27]],"date-time":"2025-07-27T04:02:41Z","timestamp":1753588961000},"page":"1-11","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["TokenVerse: Versatile Multi-concept Personalization in Token Modulation Space"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-4261-6153","authenticated-orcid":false,"given":"Daniel","family":"Garibi","sequence":"first","affiliation":[{"name":"Tel Aviv University, Tel Aviv, Israel"},{"name":"DeepMind, Tel Aviv, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-7758-671X","authenticated-orcid":false,"given":"Shahar","family":"Yadin","sequence":"additional","affiliation":[{"name":"Technion - Israel Institute of Technology, Haifa, Israel"},{"name":"DeepMind, Tel Aviv, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3792-1770","authenticated-orcid":false,"given":"Roni","family":"Paiss","sequence":"additional","affiliation":[{"name":"DeepMind, Tel Aviv, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5910-2659","authenticated-orcid":false,"given":"Omer","family":"Tov","sequence":"additional","affiliation":[{"name":"DeepMind, Tel Aviv, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7807-0606","authenticated-orcid":false,"given":"Shiran","family":"Zada","sequence":"additional","affiliation":[{"name":"DeepMind, Tel Aviv, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-7737-9880","authenticated-orcid":false,"given":"Ariel","family":"Ephrat","sequence":"additional","affiliation":[{"name":"DeepMind, Tel Aviv, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0525-8054","authenticated-orcid":false,"given":"Tomer","family":"Michaeli","sequence":"additional","affiliation":[{"name":"Technion - Israel Institute of Technology, Haifa, Israel"},{"name":"DeepMind, Tel Aviv, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8757-7790","authenticated-orcid":false,"given":"Inbar","family":"Mosseri","sequence":"additional","affiliation":[{"name":"DeepMind, Tel Aviv, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3703-0783","authenticated-orcid":false,"given":"Tali","family":"Dekel","sequence":"additional","affiliation":[{"name":"Weizmann Institute of Science, Tel Aviv, Israel"},{"name":"DeepMind, Tel Aviv, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,7,27]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"2024. OpenAI. Introducing gpt-4o and more tools to chatgpt free users. https:\/\/openai.com\/index\/gpt-4o-and-more-tools-to-chatgpt-free\/."},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00453"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00832"},{"key":"e_1_2_2_4_1","doi-asserted-by":"crossref","unstructured":"Rameen Abdal Peihao Zhu Niloy Mitra and Peter Wonka. 2020b. StyleFlow: Attribute-conditioned Exploration of StyleGAN-Generated Images using Conditional Continuous Normalizing Flows. arXiv:2008.02401 [cs.CV]","DOI":"10.1145\/3447648"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00664"},{"key":"e_1_2_2_6_1","unstructured":"Yuval Alaluf Elad Richardson Gal Metzer and Daniel Cohen-Or. 2023. A Neural Space-Time Representation for Text-to-Image Personalization. arXiv:2305.15391 [cs.CV] https:\/\/arxiv.org\/abs\/2305.15391"},{"key":"e_1_2_2_7_1","volume-title":"HyperStyle: StyleGAN Inversion with HyperNetworks for Real Image Editing. arXiv preprint arXiv:2111.15666","author":"Alaluf Yuval","year":"2021","unstructured":"Yuval Alaluf, Omer Tov, Ron Mokady, Rinon Gal, and Amit H Bermano. 2021b. HyperStyle: StyleGAN Inversion with HyperNetworks for Real Image Editing. arXiv preprint arXiv:2111.15666 (2021)."},{"key":"e_1_2_2_8_1","volume-title":"PALP: Prompt Aligned Personalization of Text-to-Image Models. arXiv preprint arXiv:2401.06105","author":"Arar Moab","year":"2024","unstructured":"Moab Arar, Andrey Voynov, Amir Hertz, Omri Avrahami, Shlomi Fruchter, Yael Pritch, Daniel Cohen-Or, and Ariel Shamir. 2024. PALP: Prompt Aligned Personalization of Text-to-Image Models. arXiv preprint arXiv:2401.06105 (2024)."},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3610548.3618154"},{"key":"e_1_2_2_10_1","unstructured":"Black Forest Labs. 2024. Flux https:\/\/github.com\/black-forest-labs\/flux. https:\/\/github.com\/black-forest-labs\/flux"},{"key":"e_1_2_2_11_1","unstructured":"Patrick Esser Sumith Kulal Andreas Blattmann Rahim Entezari Jonas M\u00fcller Harry Saini Yam Levi Dominik Lorenz Axel Sauer Frederic Boesel Dustin Podell Tim Dockhorn Zion English Kyle Lacey Alex Goodwin Yannik Marek and Robin Rombach. 2024. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis. arXiv:2403.03206 [cs.CV] https:\/\/arxiv.org\/abs\/2403.03206"},{"key":"e_1_2_2_12_1","doi-asserted-by":"crossref","unstructured":"Yarden Frenkel Yael Vinker Ariel Shamir and Daniel Cohen-Or. 2024. Implicit Style-Content Separation using B-LoRA. arXiv:2403.14572 [cs.CV] https:\/\/arxiv.org\/abs\/2403.14572","DOI":"10.1007\/978-3-031-72684-2_11"},{"key":"e_1_2_2_13_1","unstructured":"Rinon Gal Yuval Alaluf Yuval Atzmon Or Patashnik Amit H. Bermano Gal Chechik and Daniel Cohen-Or. 2022. An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion. arXiv:2208.01618 [cs.CV] https:\/\/arxiv.org\/abs\/2208.01618"},{"key":"e_1_2_2_14_1","doi-asserted-by":"crossref","unstructured":"Rinon Gal Moab Arar Yuval Atzmon Amit H. Bermano Gal Chechik and Daniel Cohen-Or. 2023. Encoder-based Domain Tuning for Fast Personalization of Text-to-Image Models. arXiv:2302.12228 [cs.CV] https:\/\/arxiv.org\/abs\/2302.12228","DOI":"10.1145\/3592133"},{"key":"e_1_2_2_15_1","volume-title":"Yujun Shi, Yunpeng Chen, Zihan Fan, Wuyou Xiao, Rui Zhao, Shuning Chang, Weijia Wu, Yixiao Ge, Ying Shan, and Mike Zheng Shou.","author":"Gu Yuchao","year":"2023","unstructured":"Yuchao Gu, Xintao Wang, Jay Zhangjie Wu, Yujun Shi, Yunpeng Chen, Zihan Fan, Wuyou Xiao, Rui Zhao, Shuning Chang, Weijia Wu, Yixiao Ge, Ying Shan, and Mike Zheng Shou. 2023. Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models. arXiv:2305.18292 [cs.CV] https:\/\/arxiv.org\/abs\/2305.18292"},{"key":"e_1_2_2_16_1","doi-asserted-by":"crossref","unstructured":"Shaozhe Hao Kai Han Zhengyao Lv Shihao Zhao and Kwan-Yee K. Wong. 2024. ConceptExpress: Harnessing Diffusion Models for Single-image Unsupervised Concept Extraction. arXiv:2407.07077 [cs.CV] https:\/\/arxiv.org\/abs\/2407.07077","DOI":"10.1007\/978-3-031-73202-7_13"},{"key":"e_1_2_2_17_1","volume-title":"GANSpace: Discovering Interpretable GAN Controls. arXiv preprint arXiv:2004.02546","author":"H\u00e4rk\u00f6nen Erik","year":"2020","unstructured":"Erik H\u00e4rk\u00f6nen, Aaron Hertzmann, Jaakko Lehtinen, and Sylvain Paris. 2020. GANSpace: Discovering Interpretable GAN Controls. arXiv preprint arXiv:2004.02546 (2020)."},{"key":"e_1_2_2_18_1","unstructured":"Amir Hertz Ron Mokady Jay Tenenbaum Kfir Aberman Yael Pritch and Daniel Cohen-Or. 2022. Prompt-to-Prompt Image Editing with Cross Attention Control. arXiv:2208.01626 [cs.CV]"},{"key":"e_1_2_2_19_1","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685 [cs.CL] https:\/\/arxiv.org\/abs\/2106.09685"},{"key":"e_1_2_2_20_1","doi-asserted-by":"crossref","unstructured":"Maxwell Jones Sheng-Yu Wang Nupur Kumari David Bau and Jun-Yan Zhu. 2024. Customizing Text-to-Image Models with a Single Image Pair. arXiv:2405.01536 [cs.CV] https:\/\/arxiv.org\/abs\/2405.01536","DOI":"10.1145\/3680528.3687642"},{"key":"e_1_2_2_21_1","volume-title":"Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196","author":"Karras Tero","year":"2017","unstructured":"Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2017. Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196 (2017)."},{"key":"e_1_2_2_22_1","volume-title":"Proc. NeurIPS.","author":"Karras Tero","year":"2020","unstructured":"Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila. 2020a. Training Generative Adversarial Networks with Limited Data. In Proc. NeurIPS."},{"key":"e_1_2_2_23_1","volume-title":"Alias-free generative adversarial networks. Advances in Neural Information Processing Systems 34","author":"Karras Tero","year":"2021","unstructured":"Tero Karras, Miika Aittala, Samuli Laine, Erik H\u00e4rk\u00f6nen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2021. Alias-free generative adversarial networks. Advances in Neural Information Processing Systems 34 (2021)."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00453"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00813"},{"key":"e_1_2_2_26_1","volume-title":"OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models. arXiv:2403.10983 [cs.CV] https:\/\/arxiv.org\/abs\/2403.10983","author":"Kong Zhe","year":"2024","unstructured":"Zhe Kong, Yong Zhang, Tianyu Yang, Tao Wang, Kaihao Zhang, Bizhu Wu, Guanying Chen, Wei Liu, and Wenhan Luo. 2024. OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models. arXiv:2403.10983 [cs.CV] https:\/\/arxiv.org\/abs\/2403.10983"},{"key":"e_1_2_2_27_1","doi-asserted-by":"crossref","unstructured":"Nupur Kumari Bingliang Zhang Richard Zhang Eli Shechtman and Jun-Yan Zhu. 2023. Multi-Concept Customization of Text-to-Image Diffusion. arXiv:2212.04488 [cs.CV] https:\/\/arxiv.org\/abs\/2212.04488","DOI":"10.1109\/CVPR52729.2023.00192"},{"key":"e_1_2_2_28_1","volume-title":"StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery. arXiv preprint arXiv:2103.17249","author":"Patashnik Or","year":"2021","unstructured":"Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. 2021. StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery. arXiv preprint arXiv:2103.17249 (2021)."},{"key":"e_1_2_2_29_1","unstructured":"Yuang Peng Yuxin Cui Haomiao Tang Zekun Qi Runpei Dong Jing Bai Chunrui Han Zheng Ge Xiangyu Zhang and Shu-Tao Xia. 2024. DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation. arXiv:2406.16855 [cs.CV] https:\/\/arxiv.org\/abs\/2406.16855"},{"key":"e_1_2_2_30_1","unstructured":"Ryan Po Guandao Yang Kfir Aberman and Gordon Wetzstein. 2023. Orthogonal Adaptation for Modular Customization of Diffusion Models. arXiv:2312.02432 [cs.CV]"},{"key":"e_1_2_2_31_1","volume-title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis. arXiv:2307.01952 [cs.CV]","author":"Podell Dustin","year":"2023","unstructured":"Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M\u00fcller, Joe Penna, and Robin Rombach. 2023. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis. arXiv:2307.01952 [cs.CV]"},{"key":"e_1_2_2_32_1","unstructured":"Aditya Ramesh Prafulla Dhariwal Alex Nichol Casey Chu and Mark Chen. 2022. Hierarchical Text-Conditional Image Generation with CLIP Latents. arXiv:2204.06125 [cs.CV]"},{"key":"e_1_2_2_33_1","volume-title":"Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation. arXiv preprint arXiv:2008.00951","author":"Richardson Elad","year":"2020","unstructured":"Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or. 2020. Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation. arXiv preprint arXiv:2008.00951 (2020)."},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3544777"},{"key":"e_1_2_2_35_1","doi-asserted-by":"crossref","unstructured":"Robin Rombach Andreas Blattmann Dominik Lorenz Patrick Esser and Bj\u00f6rn Ommer. 2022. High-Resolution Image Synthesis with Latent Diffusion Models. arXiv:2112.10752 [cs.CV]","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_2_2_36_1","doi-asserted-by":"crossref","unstructured":"Nataniel Ruiz Yuanzhen Li Varun Jampani Yael Pritch Michael Rubinstein and Kfir Aberman. 2023. DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation. arXiv:2208.12242 [cs.CV] https:\/\/arxiv.org\/abs\/2208.12242","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"e_1_2_2_37_1","volume-title":"Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi.","author":"Saharia Chitwan","year":"2022","unstructured":"Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. 2022. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. arXiv:2205.11487 [cs.CV]"},{"key":"e_1_2_2_38_1","doi-asserted-by":"crossref","unstructured":"Viraj Shah Nataniel Ruiz Forrester Cole Erika Lu Svetlana Lazebnik Yuanzhen Li and Varun Jampani. 2023. ZipLoRA: Any Subject in Any Style by Effectively Merging LoRAs. arXiv:2311.13600 [cs.CV] https:\/\/arxiv.org\/abs\/2311.13600","DOI":"10.1007\/978-3-031-73232-4_24"},{"key":"e_1_2_2_39_1","doi-asserted-by":"crossref","unstructured":"Yoad Tewel Rinon Gal Gal Chechik and Yuval Atzmon. 2024. Key-Locked Rank One Editing for Text-to-Image Personalization. arXiv:2305.01644 [cs.CV] https:\/\/arxiv.org\/abs\/2305.01644","DOI":"10.1145\/3588432.3591506"},{"key":"e_1_2_2_40_1","volume-title":"Designing an Encoder for StyleGAN Image Manipulation. arXiv preprint arXiv:2102.02766","author":"Tov Omer","year":"2021","unstructured":"Omer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen-Or. 2021. Designing an Encoder for StyleGAN Image Manipulation. arXiv preprint arXiv:2102.02766 (2021)."},{"key":"e_1_2_2_41_1","unstructured":"Yael Vinker Andrey Voynov Daniel Cohen-Or and Ariel Shamir. 2023. Concept Decomposition for Visual Exploration and Inspiration. arXiv:2305.18203 [cs.CV] https:\/\/arxiv.org\/abs\/2305.18203"},{"key":"e_1_2_2_42_1","unstructured":"Andrey Voynov Qinghao Chu Daniel Cohen-Or and Kfir Aberman. 2023. P+: Extended Textual Conditioning in Text-to-Image Generation. arXiv:2303.09522 [cs.CV] https:\/\/arxiv.org\/abs\/2303.09522"},{"key":"e_1_2_2_43_1","unstructured":"Yang Yang Wen Wang Liang Peng Chaotian Song Yao Chen Hengjia Li Xiaolong Yang Qinglin Lu Deng Cai Boxi Wu and Wei Liu. 2024. LoRA-Composer: Leveraging Low-Rank Adaptation for Multi-Concept Customization in Training-Free Diffusion Models. arXiv:2403.11627 [cs.CV] https:\/\/arxiv.org\/abs\/2403.11627"},{"key":"e_1_2_2_44_1","volume-title":"In-domain gan inversion for real image editing. arXiv preprint arXiv:2004.00049","author":"Zhu Jiapeng","year":"2020","unstructured":"Jiapeng Zhu, Yujun Shen, Deli Zhao, and Bolei Zhou. 2020. In-domain gan inversion for real image editing. arXiv preprint arXiv:2004.00049 (2020)."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3730843","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T17:58:48Z","timestamp":1774634328000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3730843"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,27]]},"references-count":44,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,8,1]]}},"alternative-id":["10.1145\/3730843"],"URL":"https:\/\/doi.org\/10.1145\/3730843","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,27]]},"assertion":[{"value":"2025-01-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-29","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}