{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T08:08:55Z","timestamp":1779350935458,"version":"3.51.4"},"reference-count":47,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2026,3,6]],"date-time":"2026-03-06T00:00:00Z","timestamp":1772755200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,3,6]],"date-time":"2026-03-06T00:00:00Z","timestamp":1772755200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Natonal Recovery and Resilience Plan Greece 2.0","award":["MIS 5154714"],"award-info":[{"award-number":["MIS 5154714"]}]},{"name":"The Cyprus Insitute"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2026,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Diffusion models have achieved state-of-the-art image synthesis, yet unlike GANs, they lack a well-structured latent space for intuitive image editing. Existing diffusion-based editing methods often rely on supervised fine-tuning or text-based guidance, while recent unsupervised techniques leveraging the model\u2019s bottleneck layer suffer from one or more key limitations: (i) they focus only on global attributes, (ii) fail to disentangle local and global semantics, or (iii) require extensive human intervention. To fill this gap, we first propose an unsupervised method for localized image editing in pre-trained unconditional diffusion models that disentangles local and global semantics in the model\u2019s latent space. Given an input image and a user-specified region of interest, our approach uses the denoising network\u2019s Jacobian to map that region to a corresponding latent subspace. We then separate this subspace into shared (global) and region-specific components to uncover latent directions that control local attributes. These directions generalize across images, enabling semantically consistent edits without retraining. We go one step further by extending our method to minimize manual supervision by automatically inferring edit directions from a single reference image and generating region masks without human input. Experiments on multiple datasets show that our method yields more localized, high-fidelity edits than state-of-the-art approaches.<\/jats:p>","DOI":"10.1007\/s11263-025-02694-y","type":"journal-article","created":{"date-parts":[[2026,3,6]],"date-time":"2026-03-06T10:18:49Z","timestamp":1772792329000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Disentangling Local and Global Semantics in Diffusion Models for Image Editing"],"prefix":"10.1007","volume":"134","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-8024-7033","authenticated-orcid":false,"given":"Manos","family":"Plitsis","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Theodoros","family":"Kouzelis","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Panagiotis","family":"Koromilas","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vassilis","family":"Katsouros","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mihalis","family":"A. Nicolaou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yannis","family":"Panagakis","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,3,6]]},"reference":[{"key":"2694_CR1","doi-asserted-by":"crossref","unstructured":"Avrahami, O., Lischinski, D., & Fried, O. (2021). Blended diffusion for text-driven editing of natural images. 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18187-18197 https:\/\/api.semanticscholar.org\/CorpusID:244714366","DOI":"10.1109\/CVPR52688.2022.01767"},{"key":"2694_CR2","doi-asserted-by":"crossref","unstructured":"Chen, S., Zhang, H., Guo, M., Lu, Y., Wang, P., & Qu, Q. (2024). Exploring low-dimensional subspace in diffusion models for controllable image editing. The thirty-eighth annual conference on neural information processing systems. https:\/\/openreview.net\/forum?id=50aOEfb2km","DOI":"10.52202\/079017-0859"},{"key":"2694_CR3","unstructured":"Choi, J., Hwang, G., Cho, H., & Kang, M. (2022). Finding the global semantic representation in gan through fr\u00e9chet mean. The eleventh international conference on learning representations."},{"key":"2694_CR4","unstructured":"Choi, J., Lee, J., Yoon, C., Park, J.H., Hwang, G., & Kang, M. (2021). Do not escape from the manifold: Discovering the local coordinates on the latent space of gans. International conference on learning representations."},{"key":"2694_CR5","unstructured":"Couairon, G., Verbeek, J., Schwenk, H., & Cord, M. (2022). Diffedit: Diffusion-based semantic image editing with mask guidance. The eleventh international conference on learning representations."},{"key":"2694_CR6","doi-asserted-by":"crossref","unstructured":"Deng, J., Guo, J., Xue, N., & Zafeiriou, S. (2019). Arcface: Additive angular margin loss for deep face recognition. Proceedings of the ieee\/cvf conference on computer vision and pattern recognition (cvpr).","DOI":"10.1109\/CVPR.2019.00482"},{"key":"2694_CR7","first-page":"8780","volume":"34","author":"P Dhariwal","year":"2021","unstructured":"Dhariwal, P., & Nichol, A. (2021). Diffusion models beat gans on image synthesis. Advances in neural information processing systems,34, 8780\u20138794.","journal-title":"Advances in neural information processing systems"},{"issue":"11","key":"2694_CR8","doi-asserted-by":"publisher","first-page":"139","DOI":"10.1145\/3422622","volume":"63","author":"I Goodfellow","year":"2020","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., & Bengio, Y. (2020). Generative adversarial networks. Communications of the ACM,63(11), 139\u2013144.","journal-title":"Communications of the ACM"},{"key":"2694_CR9","doi-asserted-by":"crossref","unstructured":"Haas, R., Huberman-Spiegelglas, I., Mulayoff, R., Gra\u00dfhof, S., Brandt, S. S., & Michaeli, T. (2024). Discovering interpretable directions in the semantic latent space of diffusion models. 2024 ieee 18th international conference on automatic face and gesture recognition (fg) (pp. 1\u20139).","DOI":"10.1109\/FG59268.2024.10581912"},{"key":"2694_CR10","doi-asserted-by":"crossref","unstructured":"Haas, R., Huberman-Spiegelglas, I., Mulayoff, R., & Michaeli, T. (2023). Discovering interpretable directions in the semantic latent space of diffusion models. arXiv preprint arXiv:2303.11073.","DOI":"10.1109\/FG59268.2024.10581912"},{"key":"2694_CR11","first-page":"9841","volume":"33","author":"E H\u00e4rk\u00f6nen","year":"2020","unstructured":"H\u00e4rk\u00f6nen, E., Hertzmann, A., Lehtinen, J., & Paris, S. (2020). Ganspace: Discovering interpretable gan controls. Advances in neural information processing systems,33, 9841\u20139850.","journal-title":"Advances in neural information processing systems"},{"key":"2694_CR12","unstructured":"Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., & Cohen-or, D. (2022). Prompt-to-prompt image editing with cross-attention control. The eleventh international conference on learning representations."},{"key":"2694_CR13","first-page":"6840","volume":"33","author":"J Ho","year":"2020","unstructured":"Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in neural information processing systems,33, 6840\u20136851.","journal-title":"Advances in neural information processing systems"},{"key":"2694_CR14","doi-asserted-by":"crossref","unstructured":"Huberman-Spiegelglas, I., Kulikov, V., & Michaeli, T. (2024). An edit friendly ddpm noise space: Inversion and manipulations. Proceedings of the ieee\/cvf conference on computer vision and pattern recognition (pp. 12469\u201312478).","DOI":"10.1109\/CVPR52733.2024.01185"},{"key":"2694_CR15","doi-asserted-by":"crossref","unstructured":"Jeong, J., Kwon, M., & Uh, Y. (2024). Training-free content injection using h-space in diffusion models. Proceedings of the ieee\/cvf winter conference on applications of computer vision (pp. 5151\u20135161).","DOI":"10.1109\/WACV57701.2024.00507"},{"key":"2694_CR16","first-page":"12104","volume":"33","author":"T Karras","year":"2020","unstructured":"Karras, T., Aittala, M., Hellsten, J., Laine, S., Lehtinen, J., & Aila, T. (2020). Training generative adversarial networks with limited data. Advances in neural information processing systems,33, 12104\u201312114.","journal-title":"Advances in neural information processing systems"},{"key":"2694_CR17","doi-asserted-by":"crossref","unstructured":"Kawar, B., Zada, S., Lang, O., Tov, O., Chang, H., Dekel, T., & Irani, M. (2023). Imagic: Text-based real image editing with diffusion models. Proceedings of the ieee\/cvf conference on computer vision and pattern recognition (cvpr) (p.6007-6017).","DOI":"10.1109\/CVPR52729.2023.00582"},{"key":"2694_CR18","unstructured":"Kouzelis, T., Plitsis, E., Nicolaou, M., & Panagakis, Y. (2024). Enabling local editing in diffusion models by joint and individual component analysis. 35th british machine vision conference 2024, BMVC 2024, glasgow, uk, november 25-28, 2024. BMVA. https:\/\/papers.bmvc2024.org\/0213.pdf"},{"key":"2694_CR19","unstructured":"Kwon, M., Jeong, J., & Uh, Y. (2022). Diffusion models already have a semantic latent space. The eleventh international conference on learning representations."},{"key":"2694_CR20","doi-asserted-by":"crossref","unstructured":"Liu, Z., Luo, P., Wang, X., & Tang, X. (2015). Deep learning face attributes in the wild. Proceedings of international conference on computer vision (iccv).","DOI":"10.1109\/ICCV.2015.425"},{"issue":"1","key":"2694_CR21","doi-asserted-by":"publisher","first-page":"523","DOI":"10.1214\/12-AOAS597","volume":"7","author":"EF Lock","year":"2013","unstructured":"Lock, E. F., Hoadley, K. A., Marron, J. S., & Nobel, A. B. (2013). Joint and individual variation explained (jive) for integrated analysis of multiple data types. The annals of applied statistics,7(1), 523.","journal-title":"The annals of applied statistics"},{"key":"2694_CR22","unstructured":"Manor, H., & Michaeli, T. (2023). On the posterior distribution in denoising: Application to uncertainty quantification. The twelfth international conference on learning representations."},{"key":"2694_CR23","unstructured":"Manor, H., & Michaeli, T. (2024). Zero-shot unsupervised and text-based audio editing using ddpm inversion. International conference on machine learning (pp. 34603\u201334629)."},{"key":"2694_CR24","unstructured":"Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J. Y.,& Ermon, S. (2021). Sdedit: Guided image synthesis and editing with stochastic differential equations. International conference on learning representations."},{"key":"2694_CR25","doi-asserted-by":"crossref","unstructured":"Mokady, R., Hertz, A., Aberman, K., Pritch, Y., & Cohen-Or, D. (2023). Null-text inversion for editing real images using guided diffusion models. Proceedings of the ieee\/cvf conference on computer vision and pattern recognition (pp. 6038\u20136047).","DOI":"10.1109\/CVPR52729.2023.00585"},{"key":"2694_CR26","unstructured":"Nichol, A. Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., Mcgrew, B., & Chen, M. (2022). Glide: Towards photorealistic image generation and editing with text-guided diffusion models. International conference on machine learning (pp. 16784\u201316804)."},{"key":"2694_CR27","unstructured":"Oldfield, J., Tzelepis, C., Panagakis, Y., Nicolaou, M., & Patras, I. (2022). Panda: Unsupervised learning of parts and appearances in the feature maps of gans. The eleventh international conference on learning representations."},{"key":"2694_CR28","first-page":"24129","volume":"36","author":"YH Park","year":"2023","unstructured":"Park, Y. H., Kwon, M., Choi, J., Jo, J., & Uh, Y. (2023). Understanding the latent space of diffusion models through the lens of riemannian geometry. Advances in Neural Information Processing Systems,36, 24129\u201324142.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2694_CR29","doi-asserted-by":"crossref","unstructured":"Preechakul, K., Chatthee, N., Wizadwongsa, S., & Suwajanakorn, S. (2022). Diffusion autoencoders: Toward a meaningful and decodable representation. Proceedings of the ieee\/cvf conference on computer vision and pattern recognition (pp. 10619\u201310629).","DOI":"10.1109\/CVPR52688.2022.01036"},{"key":"2694_CR30","unstructured":"Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., & Agarwal, S. others (2021). Learning transferable visual models from natural language supervision. International conference on machine learning (pp. 8748\u20138763)."},{"key":"2694_CR31","unstructured":"Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., & Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. arXiv preprint , arXiv:2204.06125, 1(2), 3."},{"key":"2694_CR32","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. Proceedings of the ieee\/cvf conference on computer vision and pattern recognition (pp. 10684\u201310695).","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"2694_CR33","first-page":"36479","volume":"35","author":"C Saharia","year":"2022","unstructured":"Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., et al. (2022). Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems,35, 36479\u201336494.","journal-title":"Advances in neural information processing systems"},{"key":"2694_CR34","doi-asserted-by":"crossref","unstructured":"Sehwag, V., Hazirbas, C., Gordo, A., Ozgenel, F., & Canton, C. (2022). Generating high fidelity data from low-density regions using diffusion models. Proceedings of the ieee\/cvf conference on computer vision and pattern recognition (pp. 11492\u201311501).","DOI":"10.1109\/CVPR52688.2022.01120"},{"key":"2694_CR35","doi-asserted-by":"crossref","unstructured":"Shen, Y., Gu, J., Tang, X., & Zhou, B. (2020). Interpreting the latent space of gans for semantic face editing. Proceedings of the ieee\/cvf conference on computer vision and pattern recognition (pp. 9243\u20139252).","DOI":"10.1109\/CVPR42600.2020.00926"},{"key":"2694_CR36","unstructured":"Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., & Ganguli, S. (2015). Deep unsupervised learning using nonequilibrium thermodynamics. International conference on machine learning (pp. 2256\u20132265)."},{"key":"2694_CR37","unstructured":"Song, J., Meng, C., & Ermon, S. (2020). Denoising diffusion implicit models. International conference on learning representations."},{"key":"2694_CR38","unstructured":"Song, Y., Dhariwal, P., Chen, M., & Sutskever, I. (2023). Consistency models. International conference on machine learning (pp. 32211\u201332252)."},{"key":"2694_CR39","unstructured":"Song, Y., & Ermon, S. (2019). G enerative modeling by estimating gradients of the data distribution. Advances in neural information processing systems, 32."},{"key":"2694_CR40","unstructured":"Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., & Poole, B. (2020). Score-based generative modeling through stochastic differential equations. International conference on learning representations."},{"key":"2694_CR41","unstructured":"Su, X., Song, J., Meng, C., & Ermon, S. (2022). Dual diffusion implicit bridges for image-to-image translation. The eleventh international conference on learning representations."},{"key":"2694_CR42","doi-asserted-by":"crossref","unstructured":"Tumanyan, N., Geyer, M., Bagon, S., & Dekel, T. (2023). Plug-and-play diffusion features for text-driven image-to-image translation. Proceedings of the ieee\/cvf conference on computer vision and pattern recognition (cvpr) (p.1921-1930).","DOI":"10.1109\/CVPR52729.2023.00191"},{"key":"2694_CR43","doi-asserted-by":"crossref","unstructured":"Wu, C.H., & De\u00a0la Torre, F. (2023). A latent space of stochastic diffusion models for zero-shot image editing and guidance. Proceedings of the ieee\/cvf international conference on computer vision (pp. 7378\u20137387).","DOI":"10.1109\/ICCV51070.2023.00678"},{"issue":"3","key":"2694_CR44","doi-asserted-by":"publisher","first-page":"1176","DOI":"10.1137\/15M1054201","volume":"37","author":"K Ye","year":"2016","unstructured":"Ye, K., & Lim, L. H. (2016). Schubert varieties and distances between subspaces of different dimensions. SIAM Journal on Matrix Analysis and Applications,37(3), 1176\u20131197.","journal-title":"SIAM Journal on Matrix Analysis and Applications"},{"key":"2694_CR45","unstructured":"Yu, F., Zhang, Y., Song, S., Seff, A., & Xiao, J. (2015). Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint, arXiv:1506.03365."},{"key":"2694_CR46","first-page":"16648","volume":"34","author":"J Zhu","year":"2021","unstructured":"Zhu, J., Feng, R., Shen, Y., Zhao, D., Zha, Z. J., Zhou, J., & Chen, Q. (2021). Low-rank subspaces in gans. Advances in Neural Information Processing Systems,34, 16648\u201316658.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2694_CR47","unstructured":"Zhu, J., Shen, Y., Xu, Y., Zhao, D., & Chen, Q. (2022). Region-based semantic factorization in gans. International conference on machine learning (pp. 27612\u201327632)."}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02694-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-025-02694-y","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02694-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T07:35:04Z","timestamp":1779348904000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-025-02694-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,6]]},"references-count":47,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4]]}},"alternative-id":["2694"],"URL":"https:\/\/doi.org\/10.1007\/s11263-025-02694-y","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,6]]},"assertion":[{"value":"4 March 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 September 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 March 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"All authors certify that they have no affiliations with or involvement in any organization or entity with any financial interest or non-financial interest in the subject matter or materials discussed in this manuscript.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing Interests"}}],"article-number":"160"}}