{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,24]],"date-time":"2026-08-24T20:36:59Z","timestamp":1787603819669,"version":"build-2736575974"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"11","license":[{"start":{"date-parts":[[2024,11,13]],"date-time":"2024-11-13T00:00:00Z","timestamp":1731456000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R & D Program of China","doi-asserted-by":"crossref","award":["2022ZD0160601"],"award-info":[{"award-number":["2022ZD0160601"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Municipal Science and Technology","award":["Z231100007423004"],"award-info":[{"award-number":["Z231100007423004"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62306315"],"award-info":[{"award-number":["62306315"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Natural Science Foundation","award":["4244099"],"award-info":[{"award-number":["4244099"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,11,30]]},"abstract":"<jats:p>Semantic image synthesis aims to generate images from given semantic layouts, which is a challenging task that requires training models to capture the relationship between layouts and images. Previous works are usually based on Generative Adversarial Networks (GAN) or autoregressive (AR) models. However, the GAN model's training process is unstable, and the AR model\u2019s performance is seriously affected by the independent image encoder and the unidirectional generation bias. Due to the above limitations, these methods tend to synthesize unrealistic, poorly aligned images and only consider single-style image generation. In this paper, we propose a Multi-model Style-aware Diffusion Learning (MSDL) framework for semantic image synthesis, including a training module and a sampling module. In the training module, a layout-to-image model is introduced to transfer the learned knowledge from a model pretrained with massive weak correlated text-image pairs data, making the training process more efficient. In the sampling module, we designed a map-guidance technique and creatively designed a multi-model style-guidance strategy for creating images in multiple styles, e.g., oil painting, Disney Cartoon, and pixel style. We evaluate our method on Cityscapes, ADE20K, and COCO-Stuff, making visual comparisons and computing with multiple metrics such as FID, LPIPS, etc. Experimental results demonstrate that our model is highly competitive, especially in terms of fidelity and diversity.<\/jats:p>","DOI":"10.1145\/3686155","type":"journal-article","created":{"date-parts":[[2024,8,2]],"date-time":"2024-08-02T16:04:57Z","timestamp":1722614697000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Multi-Model Style-Aware Diffusion Learning for Semantic Image Synthesis"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-1848-9962","authenticated-orcid":false,"given":"Yunfang","family":"Niu","sequence":"first","affiliation":[{"name":"Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences, Beijing, China and School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9346-3597","authenticated-orcid":false,"given":"Lingxiang","family":"Wu","sequence":"additional","affiliation":[{"name":"Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences, Beijing, China and School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-9465-7643","authenticated-orcid":false,"given":"Yufeng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences, Beijing, China and School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8544-410X","authenticated-orcid":false,"given":"Yousong","family":"Zhu","sequence":"additional","affiliation":[{"name":"Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences, Beijing, China and School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8293-3952","authenticated-orcid":false,"given":"Guibo","family":"Zhu","sequence":"additional","affiliation":[{"name":"Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences, Beijing, China, School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China, and Shanghai Artificial Intelligence Laboratory, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9118-2780","authenticated-orcid":false,"given":"Jinqiao","family":"Wang","sequence":"additional","affiliation":[{"name":"Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences, Beijing, China, School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China, and Peng Cheng Laboratory, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,11,13]]},"reference":[{"key":"e_1_3_2_2_2","article-title":"Semantic image synthesis with semantically coupled VQ-model","author":"Alaniz Stephan","year":"2022","unstructured":"Stephan Alaniz, Thomas Hummel, and Zeynep Akata. 2022. Semantic image synthesis with semantically coupled VQ-model. In Proceedings of the ICLR Workshop on Deep Generative Models for Highly Structured Data.","journal-title":"Proceedings of the ICLR Workshop on Deep Generative Models for Highly Structured Data"},{"key":"e_1_3_2_3_2","first-page":"17981","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"34","author":"Austin Jacob","year":"2021","unstructured":"Jacob Austin, Daniel D Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg. 2021. Structured denoising diffusion models in discrete state-spaces. Proceedings of the Advances in Neural Information Processing Systems 34 (2021), 17981\u201317993."},{"key":"e_1_3_2_4_2","article-title":"Analytic-DPM: An analytic estimate of the optimal reverse variance in diffusion probabilistic models","author":"Bao Fan","year":"2022","unstructured":"Fan Bao, Chongxuan Li, Jun Zhu, and Bo Zhang. 2022. Analytic-DPM: An analytic estimate of the optimal reverse variance in diffusion probabilistic models. In Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2201.06503.","journal-title":"Proceedings of the International Conference on Learning Representations (ICLR)"},{"key":"e_1_3_2_5_2","article-title":"Conditional image generation with score-based diffusion models","author":"Batzolis Georgios","year":"2021","unstructured":"Georgios Batzolis, Jan Stanczuk, Carola-Bibiane Sch\u00f6nlieb, and Christian Etmann. 2021. Conditional image generation with score-based diffusion models. arXiv:2111.13606. Retrieved from https:\/\/arxiv.org\/abs\/2111.13606","journal-title":"arXiv:2111.13606"},{"key":"e_1_3_2_6_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Caesar Holger","year":"2018","unstructured":"Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. 2018. COCO-stuff: Thing and stuff classes in context. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE."},{"key":"e_1_3_2_7_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NIPS)","volume":"36","author":"Chae JungWoo","year":"2024","unstructured":"JungWoo Chae, Hyunin Cho, Sooyeon Go, Kyungmook Choi, and Youngjung Uh. 2024. Semantic image synthesis with unconditional generator. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Vol. 36."},{"key":"e_1_3_2_8_2","first-page":"3558","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Changpinyo Soravit","year":"2021","unstructured":"Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. 2021. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3558\u20133568."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612524"},{"key":"e_1_3_2_10_2","first-page":"3697","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Pei","year":"2022","unstructured":"Pei Chen, Yangkang Zhang, Zejian Li, and Lingyun Sun. 2022. Few-shot incremental learning for label-to-image translation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3697\u20133707."},{"key":"e_1_3_2_11_2","first-page":"1511","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV)","author":"Chen Qifeng","year":"2017","unstructured":"Qifeng Chen and Vladlen Koltun. 2017. Photographic image synthesis with cascaded refinement networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 1511\u20131520."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.350"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3589002"},{"key":"e_1_3_2_14_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"34","author":"Dhariwal Prafulla","year":"2021","unstructured":"Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion models beat gans on image synthesis. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 34."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cag.2022.12.010"},{"key":"e_1_3_2_16_2","first-page":"12873\u201312883","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Esser Patrick","year":"2021","unstructured":"Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12873\u201312883."},{"key":"e_1_3_2_17_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo (ICME)","author":"Fang Fei","year":"2022","unstructured":"Fei Fang, Ziqing Li, Fei Luo, and Chunxia Xiao. 2022. Discriminator modification in GAN for text-to-image generation. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME), 1\u20136."},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3061286"},{"key":"e_1_3_2_19_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"30","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Proceedings of the Advances in Neural Information Processing Systems, Vol. 30."},{"key":"e_1_3_2_20_2","first-page":"6840","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Proceedings of the Advances in Neural Information Processing Systems, Vol. 33, 6840\u20136851."},{"key":"e_1_3_2_21_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems Workshops","author":"Ho Jonathan","year":"2021","unstructured":"Jonathan Ho and Tim Salimans. 2021. Classifier-free diffusion guidance. In Proceedings of the Advances in Neural Information Processing Systems Workshops."},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.632"},{"key":"e_1_3_2_23_2","volume-title":"Proceedings of the International Conference on Machine Learning Workshop","author":"Kong Zhifeng","year":"2021","unstructured":"Zhifeng Kong and Wei Ping. 2021. On fast sampling of diffusion probabilistic models. In Proceedings of the International Conference on Machine Learning Workshop. arXiv:2106.00132."},{"key":"e_1_3_2_24_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Lam Max WY","year":"2021","unstructured":"Max WY Lam, Jun Wang, Rongjie Huang, Dan Su, and Dong Yu. 2021. Bilateral denoising diffusion models. In Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2108.11514."},{"key":"e_1_3_2_25_2","first-page":"12","article-title":"Reference-guided landmark image inpainting with deep feature matching","volume":"32","author":"Li Jiacheng","year":"2022","unstructured":"Jiacheng Li, Zhiwei Xiong, and Dong Liu. 2022. Reference-guided landmark image inpainting with deep feature matching. IEEE Transactions on Circuits and Systems for Video Technology 32, 12 (2022), 8422\u20138435.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_26_2","article-title":"Implicit maximum likelihood estimation","author":"Li Ke","year":"2018","unstructured":"Ke Li and Jitendra Malik. 2018. Implicit maximum likelihood estimation. arXiv:1809.09087. Retrieved from http:\/\/arxiv.org\/abs\/1809.09087","journal-title":"arXiv:1809.09087"},{"key":"e_1_3_2_27_2","volume-title":"In Proceedings of the IEEE International Conference on Computer Vision (ICCV)","author":"Li Ke","year":"2019","unstructured":"Ke Li, Tianhao Zhang, and Jitendra Malik. 2019. Diverse image synthesis from semantic layouts via conditional IMLE. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 4219\u20134228."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV56688.2023.00037"},{"key":"e_1_3_2_29_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"32","author":"Liu Xihui","year":"2019","unstructured":"Xihui Liu, Guojun Yin, Jing Shao, Xiaogang Wang, and Hongsheng Li. 2019. Learning to predict layout-to-image conditional convolutions for semantic image synthesis. Proceedings of the Advances in Neural Information Processing Systems, Vol. 32."},{"key":"e_1_3_2_30_2","first-page":"11214","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lv Zhengyao","year":"2022","unstructured":"Zhengyao Lv, Xiaoming Li, Zhenxing Niu, Bing Cao, and Wangmeng Zuo. 2022. Semantic-shape adaptive feature modulation for semantic image synthesis. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 11214\u201311223."},{"key":"e_1_3_2_31_2","article-title":"Glide: Towards photorealistic image generation and editing with text-guided diffusion models","author":"Nichol Alex","year":"2021","unstructured":"Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In International Conference on Machine Learning (PMLR). arXiv:2112.10741.","journal-title":"International Conference on Machine Learning (PMLR)."},{"key":"e_1_3_2_32_2","first-page":"8162","volume-title":"Proceedings of theInternational Conference on Machine Learning (ICML)","author":"Nichol Alexander Quinn","year":"2021","unstructured":"Alexander Quinn Nichol and Prafulla Dhariwal. 2021. Improved denoising diffusion probabilistic models. In Proceedings of theInternational Conference on Machine Learning (ICML). PMLR, 8162\u20138171."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00244"},{"key":"e_1_3_2_34_2","first-page":"8808","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Qi Xiaojuan","year":"2018","unstructured":"Xiaojuan Qi, Qifeng Chen, Jiaya Jia, and Vladlen Koltun. 2018. Semi-parametric image synthesis. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 8808\u20138816."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_3_2_37_2","article-title":"Palette: Image-to-image diffusion models","author":"Saharia Chitwan","year":"2021","unstructured":"Chitwan Saharia, William Chan, Huiwen Chang, Chris A. Lee, Jonathan Ho, Tim Salimans, David J. Fleet, and Mohammad Norouzi. 2021. Palette: Image-to-image diffusion models. arXiv:2111.05826. Retrieved from https:\/\/arxiv.org\/abs\/2111.05826","journal-title":"arXiv:2111.05826"},{"key":"e_1_3_2_38_2","volume-title":"Proceedings of the Advances in neural information processing systems (NIPS)","author":"Saharia Chitwan","year":"2022","unstructured":"Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Jonathan Ho, David J. Fleet, and Mohammad Norouzi. 2022. Photorealistic text-to-image diffusion models with deep language understanding. In Proceedings of the Advances in neural information processing systems (NIPS). arXiv:2205.11487."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3311781"},{"key":"e_1_3_2_40_2","first-page":"2256","volume-title":"Proceedings of theInternational Conference on Machine Learning (ICML)","author":"Sohl-Dickstein Jascha","year":"2015","unstructured":"Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of theInternational Conference on Machine Learning (ICML). PMLR, 2256\u20132265."},{"key":"e_1_3_2_41_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Song Jiaming","year":"2020","unstructured":"Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020. Denoising diffusion implicit models. Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_42_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Sushko Vadim","year":"2020","unstructured":"Vadim Sushko, Edgar Sch\u00f6nfeld, Dan Zhang, Juergen Gall, Bernt Schiele, and Anna Khoreva. 2020. You only need adversarial supervision for semantic image synthesis. Proceedings of the International Conference on Learning Representations (ICLR)."},{"issue":"4","key":"e_1_3_2_43_2","first-page":"1526","article-title":"Incremental learning of multi-domain image-to-image translations","volume":"31","author":"Tan Daniel Stanley","year":"2020","unstructured":"Daniel Stanley Tan, Yong-Xiang Lin, and Kai-Lung Hua. 2020. Incremental learning of multi-domain image-to-image translations. IEEE Transactions on Circuits and Systems for Video Technology 31, 4 (2020), 1526\u20131539.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_44_2","first-page":"7962","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Tan Zhentao","year":"2021","unstructured":"Zhentao Tan, Menglei Chai, Dongdong Chen, Jing Liao, Qi Chu, Bin Liu, Gang Hua, and Nenghai Yu. 2021. Diverse semantic image synthesis via probability distribution modeling. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7962\u20137971."},{"key":"e_1_3_2_45_2","first-page":"1994","volume-title":"Proceedings of the ACM. International Conference on Multimedia","author":"Tang Hao","year":"2020","unstructured":"Hao Tang, Song Bai, and Nicu Sebe. 2020a. Dual attention gans for semantic image synthesis. In Proceedings of the ACM. International Conference on Multimedia, 1994\u20132002."},{"key":"e_1_3_2_46_2","first-page":"7870","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Tang Hao","year":"2020","unstructured":"Hao Tang, Dan Xu, Yan Yan, Philip H. S. Torr, and Nicu Sebe. 2020b. Local class-specific and global image-level generative adversarial networks for semantic-guided scene generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7870\u20137879."},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3506710"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00917"},{"key":"e_1_3_2_49_2","article-title":"Semantic image synthesis via diffusion models","author":"Wang Weilun","year":"2022","unstructured":"Weilun Wang, Jianmin Bao, Wengang Zhou, Dongdong Chen, Dong Chen, Lu Yuan, and Houqiang Li. 2022. Semantic image synthesis via diffusion models. arXiv:2207.00050. Retrieved from https:\/\/arxiv.org\/abs\/2207.00050","journal-title":"arXiv:2207.00050"},{"key":"e_1_3_2_50_2","first-page":"13749","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV)","author":"Wang Yi","year":"2021","unstructured":"Yi Wang, Lu Qi, Ying-Cong Chen, Xiangyu Zhang, and Jiaya Jia. 2021. Image synthesis via semantic composition. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 13749\u201313758."},{"key":"e_1_3_2_51_2","article-title":"N\u00fcwa: Visual synthesis pre-training for neural visual world creation","author":"Wu Chenfei","year":"2022","unstructured":"Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, and Nan Duan. 2022. N\u00fcwa: Visual synthesis pre-training for neural visual world creation. In Proceedings of the European conference on computer vision (ECCV). arXiv:2111.12417.","journal-title":"Proceedings of the European conference on computer vision (ECCV)"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3554739"},{"key":"e_1_3_2_53_2","first-page":"22873","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Yang Serin","year":"2023","unstructured":"Serin Yang, Hyunmin Hwang, and Jong Chul Ye. 2023. Zero-shot contrastive loss for text-guided diffusion image style transfer. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 22873\u201322882."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2891935"},{"key":"e_1_3_2_55_2","first-page":"472","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Yu Fisher","year":"2017","unstructured":"Fisher Yu, Vladlen Koltun, and Thomas Funkhouser. 2017. Dilated residual networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 472\u2013480."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458280"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"e_1_3_2_58_2","first-page":"586\u2013595","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang Richard","year":"2018","unstructured":"Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 586\u2013595."},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.544"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00515"},{"key":"e_1_3_2_61_2","first-page":"5467","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhu Zhen","year":"2020","unstructured":"Zhen Zhu, Zhiliang Xu, Ansheng You, and Xiang Bai. 2020. Semantically multi-modal image synthesis. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5467\u20135476."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3686155","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3686155","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:17:50Z","timestamp":1750295870000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3686155"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,13]]},"references-count":60,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2024,11,30]]}},"alternative-id":["10.1145\/3686155"],"URL":"https:\/\/doi.org\/10.1145\/3686155","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,13]]},"assertion":[{"value":"2023-04-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-07-16","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}