{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T06:02:42Z","timestamp":1784268162191,"version":"3.55.0"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"5","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2025,10,31]]},"abstract":"<jats:p>We propose a generative model that, given a coarsely edited image, synthesizes a photorealistic output that follows the prescribed layout. Our method transfers fine details from the original image and preserve the identity of its parts. Yet, it adapts it to the lighting and context defined by the new layout. Our key insight is that videos are a powerful source of supervision for this task: objects and camera motions provide many observations of how the world changes with viewpoint, lighting, and physical interactions. We construct an image dataset in which each sample is a pair of source and target frames extracted from the same video at randomly chosen time intervals. We warp the source frame toward the target using two motion models that mimic the expected test-time user edits. We supervise our model to translate the warped image into the ground truth, starting from a pretrained diffusion model. Our model design explicitly enables fine detail transfer from the source frame to the generated image, while closely following the user-specified layout. We show that by using simple segmentations and coarse 2D manipulations, we can synthesize a photorealistic edit faithful to the user\u2019s input while addressing second-order effects like harmonizing the lighting and physical interactions between edited objects. Project page and code can be found at https:\/\/magic-fixup.github.io.<\/jats:p>","DOI":"10.1145\/3750722","type":"journal-article","created":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T11:23:25Z","timestamp":1753356205000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Magic Fixup: Streamlining Photo Editing by Watching Dynamic Videos"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3215-5678","authenticated-orcid":false,"given":"Hadi","family":"Alzayer","sequence":"first","affiliation":[{"name":"University of Maryland","place":["College Park, United States"]},{"name":"Adobe Inc","place":["College Park, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1548-2113","authenticated-orcid":false,"given":"Zhihao","family":"Xia","sequence":"additional","affiliation":[{"name":"Adobe Inc","place":["San Jose, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7679-800X","authenticated-orcid":false,"given":"Xuaner (Cecilia)","family":"Zhang","sequence":"additional","affiliation":[{"name":"Adobe Inc","place":["San Jose, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6783-1795","authenticated-orcid":false,"given":"Eli","family":"Shechtman","sequence":"additional","affiliation":[{"name":"Adobe Inc","place":["San Jose, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0536-3658","authenticated-orcid":false,"given":"Jia-Bin","family":"Huang","sequence":"additional","affiliation":[{"name":"University of Maryland","place":["College Park, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4190-6955","authenticated-orcid":false,"given":"Michael","family":"Gharbi","sequence":"additional","affiliation":[{"name":"Adobe Inc","place":["San Jose, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,10]]},"reference":[{"key":"e_1_3_5_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00453"},{"key":"e_1_3_5_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00832"},{"key":"e_1_3_5_4_1","unstructured":"Alex Andonian Sabrina Osmany Audrey Cui YeonHwan Park Ali Jahanian Antonio Torralba and David Bau. 2021. Paint by word. arXiv:2103.10951. Retrieved from https:\/\/arxiv.org\/abs\/2103.10951 (2021)."},{"key":"e_1_3_5_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1275808.1276390"},{"key":"e_1_3_5_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01767"},{"key":"e_1_3_5_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00175"},{"key":"e_1_3_5_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW59228.2023.00091"},{"key":"e_1_3_5_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1531326.1531330"},{"key":"e_1_3_5_10_1","first-page":"I\u2013I","article-title":"Navier-stokes, fluid dynamics, and image and video inpainting","volume":"1","author":"Bertalm\u00edo Marcelo","year":"2001","unstructured":"Marcelo Bertalm\u00edo, A. Bertozzi, and Guillermo Sapiro. 2001. Navier-stokes, fluid dynamics, and image and video inpainting. In Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. 1 (2001), I\u2013I. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:695955","journal-title":"Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_5_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01764"},{"key":"e_1_3_5_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.02062"},{"key":"e_1_3_5_13_1","unstructured":"Lucy Chai Jonas Wulff and Phillip Isola. 2021. Using latent space regression to analyze and leverage compositionality in GANs. arXiv:2103.10426. Retrieved from https:\/\/arxiv.org\/abs\/2103.10426 (2021)."},{"key":"e_1_3_5_14_1","unstructured":"Xi Chen Lianghua Huang Yu Liu Yujun Shen Deli Zhao and Hengshuang Zhao. 2023. AnyDoor: Zero-shot Object-level Image Customization. arXiv:2307.09481. Retrieved from https:\/\/arxiv.org\/abs\/2307.09481 (2023)."},{"key":"e_1_3_5_15_1","first-page":"1","volume-title":"Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition","author":"Cho Taeg Sang","year":"2008","unstructured":"Taeg Sang Cho, Moshe Butman, Shai Avidan, and William T. Freeman. 2008. The patch transform and its applications to image editing. In Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1\u20138."},{"key":"e_1_3_5_16_1","unstructured":"Guillaume Couairon Jakob Verbeek Holger Schwenk and Matthieu Cord. 2022. Diffedit: Diffusion-based semantic image editing with mask guidance. arXiv:2210.11427. Retrieved from https:\/\/arxiv.org\/abs\/2210.11427 (2022)."},{"key":"e_1_3_5_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19836-6_6"},{"key":"e_1_3_5_18_1","first-page":"8780","article-title":"Diffusion models beat gans on image synthesis","volume":"34","author":"Dhariwal Prafulla","year":"2021","unstructured":"Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems 34 (2021), 8780\u20138794.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_5_19_1","article-title":"Diffusion self-guidance for controllable image generation","author":"Epstein Dave","year":"2023","unstructured":"Dave Epstein, Allan Jabri, Ben Poole, Alexei A. Efros, and Aleksander Holynski. 2023. Diffusion self-guidance for controllable image generation. Advances in Neural Information Processing Systems (2023).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_5_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01268"},{"key":"e_1_3_5_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528223.3530164"},{"key":"e_1_3_5_22_1","article-title":"Motion Guidance: Diffusion-based image editing with differentiable motion estimators","author":"Geng Daniel","year":"2024","unstructured":"Daniel Geng and Andrew Owens. 2024. Motion Guidance: Diffusion-based image editing with differentiable motion estimators. International Conference on Learning Representations (2024).","journal-title":"International Conference on Learning Representations"},{"key":"e_1_3_5_23_1","unstructured":"Nicholas Guttenberg. 2023. Diffusion with Offset Noise. (jan2023). Retrieved January 22 2024 from Retrieved from https:\/\/www.crosslabs.org\/blog\/diffusion-with-offset-noise"},{"key":"e_1_3_5_24_1","unstructured":"Amir Hertz Ron Mokady Jay Tenenbaum Kfir Aberman Yael Pritch and Daniel Cohen-Or. 2022. Prompt-to-prompt image editing with cross attention control. (2022)."},{"key":"e_1_3_5_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00457"},{"key":"e_1_3_5_26_1","first-page":"6840","article-title":"Denoising diffusion probabilistic models","volume":"33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840\u20136851.","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_5_27_1","unstructured":"Li Hu Xin Gao Peng Zhang Ke Sun Bang Zhang and Liefeng Bo. 2023. Animate anyone: Consistent and controllable image-to-video synthesis for character animation. arXiv:2311.17117. Retrieved from https:\/\/arxiv.org\/abs\/2311.17117 (2023)."},{"key":"e_1_3_5_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.632"},{"key":"e_1_3_5_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1141911.1141934"},{"key":"e_1_3_5_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00246"},{"key":"e_1_3_5_31_1","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv:1412.6980. Retrieved from https:\/\/arxiv.org\/abs\/1412.6980 (2014). https:\/\/api.semanticscholar.org\/CorpusID:6628106"},{"key":"e_1_3_5_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"e_1_3_5_33_1","volume-title":"ICML","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the ICML."},{"key":"e_1_3_5_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00532"},{"key":"e_1_3_5_35_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.14458"},{"key":"e_1_3_5_36_1","unstructured":"Grace Luo Trevor Darrell Oliver Wang Dan B. Goldman and Aleksander Holynski. 2023. Readout guidance: Learning control from diffusion features. arXiv:2312.02150. Retrieved from https:\/\/arxiv.org\/abs\/2312.02150 (2023)."},{"key":"e_1_3_5_37_1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Meng Chenlin","year":"2022","unstructured":"Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. 2022. SDEdit: Guided image synthesis and editing with stochastic differential equations. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_5_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00585"},{"key":"e_1_3_5_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2901464"},{"key":"e_1_3_5_40_1","unstructured":"Chong Mou Xintao Wang Jiechong Song Ying Shan and Jian Zhang. 2023. DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models. arXiv:2307.02421. Retrieved from https:\/\/arxiv.org\/abs\/2307.02421 (2023)."},{"key":"e_1_3_5_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00548"},{"key":"e_1_3_5_42_1","unstructured":"Maxime Oquab Timoth\u00e9e Darcet Theo Moutakanni Huy V. Vo Marc Szafraniec Vasil Khalidov Pierre Fernandez Daniel Haziza Francisco Massa Alaaeldin El-Nouby et\u00a0al. 2023. DINOv2: Learning Robust Visual Features without Supervision. (2023)."},{"key":"e_1_3_5_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3588432.3591500"},{"key":"e_1_3_5_44_1","article-title":"Diffusion handles: Enabling 3D edits for diffusion models by lifting activations to 3D","author":"Pandey Karran","year":"2024","unstructured":"Karran Pandey, Paul Guerrero, Metheus Gadelha, Yannick Hold-Geoffroy, Karan Singh, and Niloy J. Mitra. 2024. Diffusion handles: Enabling 3D edits for diffusion models by lifting activations to 3D. CVPR (2024).","journal-title":"CVPR"},{"key":"e_1_3_5_45_1","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_5_46_1","doi-asserted-by":"crossref","unstructured":"Ren\u00e9 Ranftl Alexey Bochkovskiy and Vladlen Koltun. 2021. Vision Transformers for Dense Prediction. arXiv:2103.13413. Retrieved from https:\/\/arxiv.org\/abs\/2103.13413 (2021).","DOI":"10.1109\/ICCV48922.2021.01196"},{"key":"e_1_3_5_47_1","article-title":"Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer","author":"Ranftl Ren\u00e9","year":"2020","unstructured":"Ren\u00e9 Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. 2020. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) (2020).","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)"},{"key":"e_1_3_5_48_1","unstructured":"Robin Rombach Andreas Blattmann Dominik Lorenz Patrick Esser and Bj orn Ommer. 2021. High-Resolution Image Synthesis with Latent Diffusion Models. arXiv:2112.10752. Retrieved from https:\/\/arxiv.org\/abs\/2112.10752 (2021)."},{"key":"e_1_3_5_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/1360612.1360615"},{"key":"e_1_3_5_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528233.3530757"},{"key":"e_1_3_5_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00416"},{"key":"e_1_3_5_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00844"},{"key":"e_1_3_5_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2008.4587842"},{"key":"e_1_3_5_54_1","unstructured":"Jiaming Song Chenlin Meng and Stefano Ermon. 2020. Denoising diffusion implicit models. arXiv:2010.02502. Retrieved from https:\/\/arxiv.org\/abs\/2010.02502 (2020)."},{"key":"e_1_3_5_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01756"},{"key":"e_1_3_5_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/1778765.1778862"},{"key":"e_1_3_5_57_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58536-5_24"},{"key":"e_1_3_5_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00063"},{"key":"e_1_3_5_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01761"},{"key":"e_1_3_5_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/1457515.1409071"},{"key":"e_1_3_5_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02148"},{"key":"e_1_3_5_62_1","unstructured":"Zhongcong Xu Jianfeng Zhang Jun Hao Liew Hanshu Yan Jia-Wei Liu Chenxu Zhang Jiashi Feng and Mike Zheng Shou. 2023. MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model. (2023)."},{"key":"e_1_3_5_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01763"},{"key":"e_1_3_5_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00406"},{"key":"e_1_3_5_65_1","doi-asserted-by":"crossref","unstructured":"Lvmin Zhang Anyi Rao and Maneesh Agrawala. 2023a. Adding Conditional Control to Text-to-Image Diffusion Models. (2023).","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"e_1_3_5_66_1","doi-asserted-by":"crossref","unstructured":"Lvmin Zhang Anyi Rao and Maneesh Agrawala. 2023b. Adding Conditional Control to Text-to-Image Diffusion Models. (2023).","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"e_1_3_5_67_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46454-1_36"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3750722","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,10]],"date-time":"2025-09-10T13:14:45Z","timestamp":1757510085000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3750722"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,10]]},"references-count":66,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2025,10,31]]}},"alternative-id":["10.1145\/3750722"],"URL":"https:\/\/doi.org\/10.1145\/3750722","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,10]]},"assertion":[{"value":"2024-08-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-20","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}