{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:10:15Z","timestamp":1750219815136,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":42,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,10,29]],"date-time":"2023-10-29T00:00:00Z","timestamp":1698537600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,11,2]]},"DOI":"10.1145\/3607827.3616840","type":"proceedings-article","created":{"date-parts":[[2023,10,26]],"date-time":"2023-10-26T22:09:13Z","timestamp":1698358153000},"page":"34-44","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["ImEW: A Framework for Editing Image in the Wild"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-0955-7200","authenticated-orcid":false,"given":"Tasnim","family":"Mohiuddin","sequence":"first","affiliation":[{"name":"Huawei Singapore Research Center, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-8140-2199","authenticated-orcid":false,"given":"Tianyi","family":"Zhang","sequence":"additional","affiliation":[{"name":"Huawei Singapore Research Center, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-5306-0297","authenticated-orcid":false,"given":"Maowen","family":"Nie","sequence":"additional","affiliation":[{"name":"Huawei Singapore Research Center, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-5651-9033","authenticated-orcid":false,"given":"Jing","family":"Huang","sequence":"additional","affiliation":[{"name":"Huawei Singapore Research Center, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3863-8882","authenticated-orcid":false,"given":"Qianqian","family":"Chen","sequence":"additional","affiliation":[{"name":"Huawei Singapore Research Center, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2717-4192","authenticated-orcid":false,"given":"Wei","family":"Shi","sequence":"additional","affiliation":[{"name":"Huawei Singapore Research Center, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,10,29]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01762"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW59228.2023.00091"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW59228.2023.00091"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"crossref","unstructured":"Fan Bao Shen Nie Kaiwen Xue Yue Cao Chongxuan Li Hang Su and Jun Zhu. 2023. All are Worth Words: A ViT Backbone for Diffusion Models. In CVPR.  Fan Bao Shen Nie Kaiwen Xue Yue Cao Chongxuan Li Hang Su and Jun Zhu. 2023. All are Worth Words: A ViT Backbone for Diffusion Models. In CVPR.","DOI":"10.1109\/CVPR52729.2023.02171"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1531326.1531330"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Hila Chefer Yuval Alaluf Yael Vinker Lior Wolf and Daniel Cohen-Or. 2023. Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models. arxiv: 2301.13826 [cs.CV]  Hila Chefer Yuval Alaluf Yael Vinker Lior Wolf and Daniel Cohen-Or. 2023. Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models. arxiv: 2301.13826 [cs.CV]","DOI":"10.1145\/3592116"},{"key":"e_1_3_2_1_7_1","volume-title":"Segment and Track Anything. ArXiv","author":"Cheng Yangming","year":"2023","unstructured":"Yangming Cheng , Liulei Li , Yuanyou Xu , Xiaodi Li , Zongxin Yang , Wenguan Wang , and Yi Yang . 2023. Segment and Track Anything. ArXiv , Vol. abs\/ 2305 .06558 ( 2023 ). Yangming Cheng, Liulei Li, Yuanyou Xu, Xiaodi Li, Zongxin Yang, Wenguan Wang, and Yi Yang. 2023. Segment and Track Anything. ArXiv , Vol. abs\/2305.06558 (2023)."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00367"},{"key":"e_1_3_2_1_9_1","volume-title":"Lin (Eds.)","volume":"33","author":"Chi Lu","year":"2020","unstructured":"Lu Chi , Borui Jiang , and Yadong Mu . 2020 . Fast Fourier Convolution. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H . Lin (Eds.) , Vol. 33 . Curran Associates, Inc., 4479--4488. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/ 2020\/file\/2fd5d41ec6cfab47e32164d5624269b1-Paper.pdf Lu Chi, Borui Jiang, and Yadong Mu. 2020. Fast Fourier Convolution. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 4479--4488. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2020\/file\/2fd5d41ec6cfab47e32164d5624269b1-Paper.pdf"},{"key":"e_1_3_2_1_10_1","unstructured":"Jaemin Cho Abhay Zala and Mohit Bansal. 2023. Visual Programming for Text-to-Image Generation and Evaluation. (2023).  Jaemin Cho Abhay Zala and Mohit Bansal. 2023. Visual Programming for Text-to-Image Generation and Evaluation. (2023)."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2003.1211538"},{"key":"e_1_3_2_1_12_1","unstructured":"Prafulla Dhariwal and Alex Nichol. 2021a. Diffusion Models Beat GANs on Image Synthesis. arxiv: 2105.05233 [cs.LG]  Prafulla Dhariwal and Alex Nichol. 2021a. Diffusion Models Beat GANs on Image Synthesis. arxiv: 2105.05233 [cs.LG]"},{"key":"e_1_3_2_1_13_1","volume-title":"Diffusion Models Beat GANs on Image Synthesis. ArXiv","author":"Dhariwal Prafulla","year":"2021","unstructured":"Prafulla Dhariwal and Alex Nichol . 2021b. Diffusion Models Beat GANs on Image Synthesis. ArXiv , Vol. abs\/ 2105 .05233 ( 2021 ). Prafulla Dhariwal and Alex Nichol. 2021b. Diffusion Models Beat GANs on Image Synthesis. ArXiv , Vol. abs\/2105.05233 (2021)."},{"key":"e_1_3_2_1_14_1","volume-title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. ArXiv","author":"Dosovitskiy Alexey","year":"1929","unstructured":"Alexey Dosovitskiy , Lucas Beyer , Alexander Kolesnikov , Dirk Weissenborn , Xiaohua Zhai , Thomas Unterthiner , Mostafa Dehghani , Matthias Minderer , Georg Heigold , Sylvain Gelly , Jakob Uszkoreit , and Neil Houlsby . 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. ArXiv , Vol. abs\/ 2010 .1 1929 (2020). Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. ArXiv , Vol. abs\/2010.11929 (2020)."},{"key":"e_1_3_2_1_15_1","volume-title":"Frido: Feature Pyramid Diffusion for Complex Scene Image Synthesis. In AAAI.","author":"Fan Wan-Cyuan","year":"2023","unstructured":"Wan-Cyuan Fan , Yen-Chun Chen , Dongdong Chen , Yu Cheng , Lu Yuan , and Yu-Chiang Frank Wang . 2023 . Frido: Feature Pyramid Diffusion for Complex Scene Image Synthesis. In AAAI. Wan-Cyuan Fan, Yen-Chun Chen, Dongdong Chen, Yu Cheng, Lu Yuan, and Yu-Chiang Frank Wang. 2023. Frido: Feature Pyramid Diffusion for Complex Scene Image Synthesis. In AAAI."},{"key":"e_1_3_2_1_16_1","volume-title":"Weinberger (Eds.)","volume":"27","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow , Jean Pouget-Abadie , Mehdi Mirza , Bing Xu , David Warde-Farley , Sherjil Ozair , Aaron Courville , and Yoshua Bengio . 2014 a. Generative Adversarial Nets. In Advances in Neural Information Processing Systems, Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K.Q . Weinberger (Eds.) , Vol. 27 . Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/ 2014\/file\/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014a. Generative Adversarial Nets. In Advances in Neural Information Processing Systems, Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K.Q. Weinberger (Eds.), Vol. 27. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2014\/file\/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf"},{"key":"e_1_3_2_1_17_1","volume-title":"Weinberger (Eds.)","volume":"27","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow , Jean Pouget-Abadie , Mehdi Mirza , Bing Xu , David Warde-Farley , Sherjil Ozair , Aaron Courville , and Yoshua Bengio . 2014 b. Generative Adversarial Nets. In Advances in Neural Information Processing Systems, Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K.Q . Weinberger (Eds.) , Vol. 27 . Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/ 2014\/file\/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014b. Generative Adversarial Nets. In Advances in Neural Information Processing Systems, Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K.Q. Weinberger (Eds.), Vol. 27. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2014\/file\/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf"},{"key":"e_1_3_2_1_18_1","volume-title":"Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"He Kaiming","year":"2015","unstructured":"Kaiming He , X. Zhang , Shaoqing Ren , and Jian Sun . 2015 . Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2015), 770--778. Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. 2015. Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2015), 770--778."},{"key":"e_1_3_2_1_19_1","volume-title":"Prompt-to-Prompt Image Editing with Cross Attention Control. arXiv preprint arXiv:2208.01626","author":"Hertz Amir","year":"2022","unstructured":"Amir Hertz , Ron Mokady , Jay Tenenbaum , Kfir Aberman , Yael Pritch , and Daniel Cohen-Or . 2022. Prompt-to-Prompt Image Editing with Cross Attention Control. arXiv preprint arXiv:2208.01626 ( 2022 ). Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. 2022. Prompt-to-Prompt Image Editing with Cross Attention Control. arXiv preprint arXiv:2208.01626 (2022)."},{"key":"e_1_3_2_1_20_1","volume-title":"Lin (Eds.)","volume":"33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho , Ajay Jain , and Pieter Abbeel . 2020 a. Denoising Diffusion Probabilistic Models. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H . Lin (Eds.) , Vol. 33 . Curran Associates, Inc., 6840--6851. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/ 2020\/file\/4c5bcfec8584af0d967f1ab10179ca4b-Paper.pdf Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020a. Denoising Diffusion Probabilistic Models. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 6840--6851. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2020\/file\/4c5bcfec8584af0d967f1ab10179ca4b-Paper.pdf"},{"key":"e_1_3_2_1_21_1","volume-title":"Lin (Eds.)","volume":"33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho , Ajay Jain , and Pieter Abbeel . 2020 b. Denoising Diffusion Probabilistic Models. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H . Lin (Eds.) , Vol. 33 . Curran Associates, Inc., 6840--6851. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/ 2020\/file\/4c5bcfec8584af0d967f1ab10179ca4b-Paper.pdf Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020b. Denoising Diffusion Probabilistic Models. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 6840--6851. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2020\/file\/4c5bcfec8584af0d967f1ab10179ca4b-Paper.pdf"},{"key":"e_1_3_2_1_22_1","volume-title":"Perceptual Losses for Real-Time Style Transfer and Super-Resolution. In European Conference on Computer Vision.","author":"Johnson Justin","year":"2016","unstructured":"Justin Johnson , Alexandre Alahi , and Li Fei-Fei . 2016 . Perceptual Losses for Real-Time Style Transfer and Super-Resolution. In European Conference on Computer Vision. Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016. Perceptual Losses for Real-Time Style Transfer and Super-Resolution. In European Conference on Computer Vision."},{"key":"e_1_3_2_1_23_1","volume-title":"Imagic: Text-Based Real Image Editing with Diffusion Models. In Conference on Computer Vision and Pattern Recognition","author":"Kawar Bahjat","year":"2023","unstructured":"Bahjat Kawar , Shiran Zada , Oran Lang , Omer Tov , Huiwen Chang , Tali Dekel , Inbar Mosseri , and Michal Irani . 2023 . Imagic: Text-Based Real Image Editing with Diffusion Models. In Conference on Computer Vision and Pattern Recognition 2023. Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. 2023. Imagic: Text-Based Real Image Editing with Diffusion Models. In Conference on Computer Vision and Pattern Recognition 2023."},{"key":"e_1_3_2_1_24_1","volume-title":"Kingma and Max Welling","author":"Diederik","year":"2013","unstructured":"Diederik P. Kingma and Max Welling . 2013 a. Auto-Encoding Variational Bayes. CoRR , Vol. abs\/ 1312 .6114 (2013). Diederik P. Kingma and Max Welling. 2013a. Auto-Encoding Variational Bayes. CoRR , Vol. abs\/1312.6114 (2013)."},{"key":"e_1_3_2_1_25_1","volume-title":"Kingma and Max Welling","author":"Diederik","year":"2013","unstructured":"Diederik P. Kingma and Max Welling . 2013 b. Auto-Encoding Variational Bayes. CoRR , Vol. abs\/ 1312 .6114 (2013). Diederik P. Kingma and Max Welling. 2013b. Auto-Encoding Variational Bayes. CoRR , Vol. abs\/1312.6114 (2013)."},{"key":"e_1_3_2_1_26_1","volume-title":"arXiv:2304.02643","author":"Kirillov Alexander","year":"2023","unstructured":"Alexander Kirillov , Eric Mintun , Nikhila Ravi , Hanzi Mao , Chloe Rolland , Laura Gustafson , Tete Xiao , Spencer Whitehead , Alexander C. Berg , Wan-Yen Lo , Piotr Doll\u00e1r , and Ross Girshick . 2023. Segment Anything . arXiv:2304.02643 ( 2023 ). Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Doll\u00e1r, and Ross Girshick. 2023. Segment Anything. arXiv:2304.02643 (2023)."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2366145.2366150"},{"key":"e_1_3_2_1_28_1","unstructured":"Donghoon Lee Jiseob Kim Jisu Choi Jongmin Kim Minwoo Byeon Woonhyuk Baek and Saehoon Kim. 2022. Karlo-v1.0.alpha on COYO-100M and CC15M. https:\/\/github.com\/kakaobrain\/karlo.  Donghoon Lee Jiseob Kim Jisu Choi Jongmin Kim Minwoo Byeon Woonhyuk Baek and Saehoon Kim. 2022. Karlo-v1.0.alpha on COYO-100M and CC15M. https:\/\/github.com\/kakaobrain\/karlo."},{"key":"e_1_3_2_1_29_1","volume-title":"Large Multimodal Models: Notes on CVPR 2023 Tutorial. ArXiv","volume":"2306","author":"Li Chunyuan","year":"2023","unstructured":"Chunyuan Li . 2023 . Large Multimodal Models: Notes on CVPR 2023 Tutorial. ArXiv , Vol. abs\/ 2306 .14895 (2023). Chunyuan Li. 2023. Large Multimodal Models: Notes on CVPR 2023 Tutorial. ArXiv , Vol. abs\/2306.14895 (2023)."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02156"},{"key":"e_1_3_2_1_31_1","volume-title":"LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models. arXiv preprint arXiv:2305.13655","author":"Lian Long","year":"2023","unstructured":"Long Lian , Boyi Li , Adam Yala , and Trevor Darrell . 2023. LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models. arXiv preprint arXiv:2305.13655 ( 2023 ). Long Lian, Boyi Li, Adam Yala, and Trevor Darrell. 2023. LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models. arXiv preprint arXiv:2305.13655 (2023)."},{"key":"e_1_3_2_1_32_1","volume-title":"Antonio Torralba, and Sanja Fidler.","author":"Ling Huan","year":"2021","unstructured":"Huan Ling , Karsten Kreis , Daiqing Li , Seung Wook Kim , Antonio Torralba, and Sanja Fidler. 2021 . EditGAN: High- Precision Semantic Image Editing. In Advances in Neural Information Processing Systems (NeurIPS) . Huan Ling, Karsten Kreis, Daiqing Li, Seung Wook Kim, Antonio Torralba, and Sanja Fidler. 2021. EditGAN: High-Precision Semantic Image Editing. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_2_1_33_1","unstructured":"Luping Liu Zijian Zhang Yi Ren Rongjie Huang Xiang Yin and Zhou Zhao. 2023 b. Detector Guidance for Multi-Object Text-to-Image Generation. arxiv: 2306.02236 [cs.CV]  Luping Liu Zijian Zhang Yi Ren Rongjie Huang Xiang Yin and Zhou Zhao. 2023 b. Detector Guidance for Multi-Object Text-to-Image Generation. arxiv: 2306.02236 [cs.CV]"},{"key":"e_1_3_2_1_34_1","volume-title":"2023 a. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499","author":"Liu Shilong","year":"2023","unstructured":"Shilong Liu , Zhaoyang Zeng , Tianhe Ren , Feng Li , Hao Zhang , Jie Yang , Chunyuan Li , Jianwei Yang , Hang Su , Jun Zhu , 2023 a. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499 ( 2023 ). Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. 2023 a. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499 (2023)."},{"key":"e_1_3_2_1_35_1","volume-title":"Can SAM Boost Video Super-Resolution? ArXiv","author":"Lu Zhihe","year":"2023","unstructured":"Zhihe Lu , Zeyu Xiao , Jiawang Bai , Zhiwei Xiong , and Xinchao Wang . 2023. Can SAM Boost Video Super-Resolution? ArXiv , Vol. abs\/ 2305 .06524 ( 2023 ). Zhihe Lu, Zeyu Xiao, Jiawang Bai, Zhiwei Xiong, and Xinchao Wang. 2023. Can SAM Boost Video Super-Resolution? ArXiv , Vol. abs\/2305.06524 (2023)."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01117"},{"key":"e_1_3_2_1_37_1","volume-title":"Segment Anything in Medical Images. ArXiv","author":"Ma Jun","year":"2023","unstructured":"Jun Ma and Bo Wang . 2023. Segment Anything in Medical Images. ArXiv , Vol. abs\/ 2304 .12306 ( 2023 ). Jun Ma and Bo Wang. 2023. Segment Anything in Medical Images. ArXiv , Vol. abs\/2304.12306 (2023)."},{"key":"e_1_3_2_1_38_1","unstructured":"Chong Mou Xintao Wang Liangbin Xie Yanze Wu Jian Zhang Zhongang Qi Ying Shan and Xiaohu Qie. 2023. T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models. arxiv: 2302.08453 [cs.CV]  Chong Mou Xintao Wang Liangbin Xie Yanze Wu Jian Zhang Zhongang Qi Ying Shan and Xiaohu Qie. 2023. T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models. arxiv: 2302.08453 [cs.CV]"},{"key":"e_1_3_2_1_39_1","unstructured":"Byong Mok Oh and Julie Dorsey. 2002. A system for image-based modeling and photo editing.  Byong Mok Oh and Julie Dorsey. 2002. A system for image-based modeling and photo editing."},{"key":"e_1_3_2_1_40_1","volume-title":"Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold. In ACM SIGGRAPH 2023 Conference Proceedings.","author":"Pan Xingang","year":"2023","unstructured":"Xingang Pan , Ayush Tewari , Thomas Leimk\u00fchler , Lingjie Liu , Abhimitra Meka , and Christian Theobalt . 2023 . Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold. In ACM SIGGRAPH 2023 Conference Proceedings. Xingang Pan, Ayush Tewari, Thomas Leimk\u00fchler, Lingjie Liu, Abhimitra Meka, and Christian Theobalt. 2023. Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold. In ACM SIGGRAPH 2023 Conference Proceedings."},{"key":"e_1_3_2_1_41_1","volume-title":"Context Encoders: Feature Learning by Inpainting.","author":"Pathak Deepak","year":"2016","unstructured":"Deepak Pathak , Philipp Kr\"ahenb \u00fchl , Jeff Donahue , Trevor Darrell , and Alexei Efros . 2016 . Context Encoders: Feature Learning by Inpainting. Deepak Pathak, Philipp Kr\"ahenb\u00fchl, Jeff Donahue, Trevor Darrell, and Alexei Efros. 2016. Context Encoders: Feature Learning by Inpainting."},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/1201775.882269"}],"event":{"name":"MM '23: The 31st ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Ottawa ON Canada","acronym":"MM '23"},"container-title":["Proceedings of the 1st Workshop on Large Generative Models Meet Multimodal Applications"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3607827.3616840","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3607827.3616840","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:05Z","timestamp":1750178765000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3607827.3616840"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,29]]},"references-count":42,"alternative-id":["10.1145\/3607827.3616840","10.1145\/3607827"],"URL":"https:\/\/doi.org\/10.1145\/3607827.3616840","relation":{},"subject":[],"published":{"date-parts":[[2023,10,29]]},"assertion":[{"value":"2023-10-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}