{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T19:02:26Z","timestamp":1784746946017,"version":"3.55.0"},"reference-count":74,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2023,7,28]],"date-time":"2023-07-28T00:00:00Z","timestamp":1690502400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2020AAA0106200"],"award-info":[{"award-number":["2020AAA0106200"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61832016, 62102162, and U20B2070"],"award-info":[{"award-number":["61832016, 62102162, and U20B2070"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Natural Science Foundation","award":["L221013"],"award-info":[{"award-number":["L221013"]}]},{"name":"National Science and Technology Council","award":["111-2221-E-006-112-MY3"],"award-info":[{"award-number":["111-2221-E-006-112-MY3"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2023,10,31]]},"abstract":"<jats:p>\n            This work presents Unified Contrastive Arbitrary Style Transfer (UCAST), a novel style representation learning and transfer framework, that can fit in most existing arbitrary image style transfer models, such as CNN-based, ViT-based, and flow-based methods. As the key component in image style transfer tasks, a suitable style representation is essential to achieve satisfactory results. Existing approaches based on deep neural networks typically use second-order statistics to generate the output. However, these hand-crafted features computed from a single image cannot leverage style information sufficiently, which leads to artifacts such as local distortions and style inconsistency. To address these issues, we learn style representation directly from a large number of images based on contrastive learning by considering the relationships between specific styles and the holistic style distribution. Specifically, we present an adaptive contrastive learning scheme for style transfer by introducing an input-dependent temperature. Our framework consists of three key components: a parallel contrastive learning scheme for style representation and transfer, a domain enhancement (DE) module for effective learning of style distribution, and a generative network for style transfer. Qualitative and quantitative evaluations show the results of our approach are superior to those obtained via state-of-the-art methods. The code is available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/zyxElsa\/CAST_pytorch\">https:\/\/github.com\/zyxElsa\/CAST_pytorch<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3605548","type":"journal-article","created":{"date-parts":[[2023,6,20]],"date-time":"2023-06-20T11:41:51Z","timestamp":1687261311000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":34,"title":["A Unified Arbitrary Style Transfer Framework via Adaptive Contrastive Learning"],"prefix":"10.1145","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6433-2678","authenticated-orcid":false,"given":"Yuxin","family":"Zhang","sequence":"first","affiliation":[{"name":"MAIS, Institute of Automation, CAS, China and School of Artificial Intelligence, UCAS, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3975-2483","authenticated-orcid":false,"given":"Fan","family":"Tang","sequence":"additional","affiliation":[{"name":"Institute Of Computing Technology, CAS, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6502-145X","authenticated-orcid":false,"given":"Weiming","family":"Dong","sequence":"additional","affiliation":[{"name":"MAIS, Institute of Automation, CAS, China and School of Artificial Intelligence, UCAS, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7787-6428","authenticated-orcid":false,"given":"Haibin","family":"Huang","sequence":"additional","affiliation":[{"name":"Kuaishou Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8243-9513","authenticated-orcid":false,"given":"Chongyang","family":"Ma","sequence":"additional","affiliation":[{"name":"Kuaishou Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6699-2944","authenticated-orcid":false,"given":"Tong-Yee","family":"Lee","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Cheng-Kung University, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8343-9665","authenticated-orcid":false,"given":"Changsheng","family":"Xu","sequence":"additional","affiliation":[{"name":"MAIS, Institute of Automation, CAS, China and School of Artificial Intelligence, UCAS, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,7,28]]},"reference":[{"key":"e_1_3_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00092"},{"key":"e_1_3_2_3_1","unstructured":"Art Institute of Chicago. 2023. (2023). Retrieved June 03 2023 from https:\/\/www.artic.edu\/."},{"key":"e_1_3_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01389"},{"key":"e_1_3_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"e_1_3_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.296"},{"key":"e_1_3_2_7_1","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)","author":"Chen Haibo","year":"2021","unstructured":"Haibo Chen, Lei Zhao, Zhizhong Wang, Zhang Hui Ming, Zhiwen Zuo, Ailin Li, Wei Xing, and Dongming Lu. 2021a. Artistic style transfer with internal-external learning and contrastive learning. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00093"},{"key":"e_1_3_2_9_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i2.16208"},{"key":"e_1_3_2_10_1","first-page":"11326","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Deng Yingying","year":"2022","unstructured":"Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Xingjia Pan, Lei Wang, and Changsheng Xu. 2022. StyTr \\(^2\\) : Image style transfer with transformers. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 11326\u201311336."},{"key":"e_1_3_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3414015"},{"key":"e_1_3_2_12_1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dumoulin Vincent","year":"2017","unstructured":"Vincent Dumoulin, Jonathon Shlens, and Manjunath Kudlur. 2017. A learned representation for artistic style. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925948"},{"key":"e_1_3_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093420"},{"key":"e_1_3_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.265"},{"key":"e_1_3_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.397"},{"key":"e_1_3_2_17_1","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NIPS)","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Proceedings of the Advances in Neural Information Processing Systems (NIPS)."},{"key":"e_1_3_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00084"},{"key":"e_1_3_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548282"},{"key":"e_1_3_2_21_1","unstructured":"Nisha Huang Yuxin Zhang Fan Tang Chongyang Ma Haibin Huang Yong Zhang Weiming Dong and Changsheng Xu. 2022b. DiffStyler: Controllable dual diffusion for text-driven image stylization. arXiv:2211.10682. Retrieved from https:\/\/arxiv.org\/abs\/2211.10682."},{"key":"e_1_3_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.167"},{"key":"e_1_3_2_23_1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Jeong Jongheon","year":"2021","unstructured":"Jongheon Jeong and Jinwoo Shin. 2021. Training GANs with stronger augmentations via contrastive discriminator. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_24_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5862"},{"key":"e_1_3_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2019.2921336"},{"key":"e_1_3_2_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"e_1_3_2_27_1","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)","author":"Kang Minguk","year":"2020","unstructured":"Minguk Kang and Jaesik Park. 2020. ContraGAN: Contrastive learning for conditional image generation. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_2_28_1","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations (ICLR) ."},{"key":"e_1_3_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01029"},{"key":"e_1_3_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00452"},{"key":"e_1_3_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01027"},{"key":"e_1_3_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01753"},{"key":"e_1_3_2_33_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_3"},{"key":"e_1_3_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00393"},{"key":"e_1_3_2_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-70139-4"},{"key":"e_1_3_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073683"},{"key":"e_1_3_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450525"},{"key":"e_1_3_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01611"},{"key":"e_1_3_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00658"},{"key":"e_1_3_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2899857"},{"key":"e_1_3_2_41_1","unstructured":"Thaneeya McArdle. 2022. Explore art styles. (2022). Retrieved from https:\/\/www.art-is-fun.com\/art-styles. Accessed 3 October 2022."},{"key":"e_1_3_2_42_1","unstructured":"National Gallery of Art. 2023. (2023). Retrieved June 03 2023 from https:\/\/www.nga.gov\/."},{"key":"e_1_3_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00603"},{"key":"e_1_3_2_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58545-7_19"},{"key":"e_1_3_2_45_1","unstructured":"Taesung Park Jun-Yan Zhu Oliver Wang Jingwan Lu Eli Shechtman Alexei Efros and Richard Zhang. 2020. Swapping autoencoder for deep image manipulation. In Proceedings of the Advances Neural Information Processing Systems (NeurIPS) . 7198\u20137211."},{"key":"e_1_3_2_46_1","unstructured":"Pexels. 2023. (2023). Retrieved June 03 2023 from https:\/\/www.pexels.com."},{"key":"e_1_3_2_47_1","doi-asserted-by":"publisher","DOI":"10.2308\/iace-50038"},{"key":"e_1_3_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00917"},{"key":"e_1_3_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_2_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01237-3_43"},{"key":"e_1_3_2_51_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01237-3_43"},{"key":"e_1_3_2_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2873701"},{"key":"e_1_3_2_53_1","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01383"},{"key":"e_1_3_2_55_1","first-page":"1349","volume-title":"Proceedings of the International Conference on Machine Learning (ICML)","author":"Ulyanov Dmitry","year":"2016","unstructured":"Dmitry Ulyanov, Vadim Lebedev, Andrea Vedaldi, and Victor S Lempitsky. 2016. Texture networks: Feed-forward synthesis of textures and stylized images. In Proceedings of the International Conference on Machine Learning (ICML). 1349\u20131357."},{"key":"e_1_3_2_56_1","unstructured":"Aaron van den Oord Yazhe Li and Oriol Vinyals. 2019. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748 (2019)."},{"key":"e_1_3_2_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2004.1272726"},{"key":"e_1_3_2_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00252"},{"key":"e_1_3_2_59_1","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-022-0287-3"},{"key":"e_1_3_2_60_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6905"},{"key":"e_1_3_2_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547939"},{"key":"e_1_3_2_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01435"},{"key":"e_1_3_2_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01041"},{"key":"e_1_3_2_64_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19787-1_11"},{"key":"e_1_3_2_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00632"},{"key":"e_1_3_2_66_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01404"},{"key":"e_1_3_2_67_1","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Zhang Chaoning","year":"2022","unstructured":"Chaoning Zhang, Kang Zhang, Chenshuang Zhang, Trung X. Pham, Chang D. Yoo, and In So Kweon. 2022b. How does SimSiam avoid collapse without negative samples? A unified understanding with self-supervised contrastive learning. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_68_1","first-page":"349","volume-title":"Proceedings of the European Conference on Computer Vision Workshops","author":"Zhang Hang","year":"2018","unstructured":"Hang Zhang and Kristin Dana. 2018. Multi-style generative network for real-time transfer. In Proceedings of the European Conference on Computer Vision Workshops. 349\u2013365."},{"key":"e_1_3_2_69_1","volume-title":"Proceedings of the NeurIPS Self-Supervised Learning\u2014Theory and Practice Workshop","author":"Zhang Oliver","year":"2021","unstructured":"Oliver Zhang, Mike Wu, Jasmine Bayrooti, and Noah Goodman. 2021. Temperature as uncertainty in contrastive learning. In Proceedings of the NeurIPS Self-Supervised Learning\u2014Theory and Practice Workshop."},{"key":"e_1_3_2_70_1","first-page":"10146","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang Yuxin","year":"2023","unstructured":"Yuxin Zhang, Nisha Huang, Fan Tang, Haibin Huang, Chongyang Ma, Weiming Dong, and Changsheng Xu. 2023a. Inversion-based style transfer with diffusion models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10146\u201310156."},{"key":"e_1_3_2_71_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00787"},{"key":"e_1_3_2_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528233.3530736"},{"key":"e_1_3_2_73_1","doi-asserted-by":"publisher","DOI":"10.1162\/leon_a_02323"},{"key":"e_1_3_2_74_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2723009"},{"key":"e_1_3_2_75_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.244"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3605548","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3605548","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:03Z","timestamp":1750182543000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3605548"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,28]]},"references-count":74,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2023,10,31]]}},"alternative-id":["10.1145\/3605548"],"URL":"https:\/\/doi.org\/10.1145\/3605548","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,28]]},"assertion":[{"value":"2022-10-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-24","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-28","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}