{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T11:53:57Z","timestamp":1782388437264,"version":"3.54.5"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2021,7,22]],"date-time":"2021-07-22T00:00:00Z","timestamp":1626912000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2020AAA0106200"],"award-info":[{"award-number":["2020AAA0106200"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61832016 and U20B2070"],"award-info":[{"award-number":["61832016 and U20B2070"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2021,8,31]]},"abstract":"<jats:p>Multimodal and multi-domain stylization are two important problems in the field of image style transfer. Currently, there are few methods that can perform multimodal and multi-domain stylization simultaneously. In this study, we propose a unified framework for multimodal and multi-domain style transfer with the support of both exemplar-based reference and randomly sampled guidance. The key component of our method is a novel style distribution alignment module that eliminates the explicit distribution gaps between various style domains and reduces the risk of mode collapse. The multimodal diversity is ensured by either guidance from multiple images or random style codes, while the multi-domain controllability is directly achieved by using a domain label. We validate our proposed framework on painting style transfer with various artistic styles and genres. Qualitative and quantitative comparisons with state-of-the-art methods demonstrate that our method can generate high-quality results of multi-domain styles and multimodal instances from reference style guidance or a random sampled style.<\/jats:p>","DOI":"10.1145\/3450525","type":"journal-article","created":{"date-parts":[[2021,7,22]],"date-time":"2021-07-22T14:44:29Z","timestamp":1626965069000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["Distribution Aligned Multimodal and Multi-domain Image Stylization"],"prefix":"10.1145","volume":"17","author":[{"given":"Minxuan","family":"Lin","sequence":"first","affiliation":[{"name":"NLPR, Institute of Automation, Chinese Academy of Sciences &amp; School of Artificial Intelligence, University of Chinese Academy of Sciences"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fan","family":"Tang","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Jilin University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weiming","family":"Dong","sequence":"additional","affiliation":[{"name":"NLPR, Institute of Automation, Chinese Academy of Sciences &amp; CASIA-LLvisionJoint Lab"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiao","family":"Li","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Changsheng","family":"Xu","sequence":"additional","affiliation":[{"name":"NLPR, Institute of Automation, Chinese Academy of Sciences &amp; CASIA-LLvision Joint Lab"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chongyang","family":"Ma","sequence":"additional","affiliation":[{"name":"Kuaishou Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,7,22]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2018.00122"},{"key":"e_1_2_2_2_1","volume-title":"Jamie Ryan Kiros, and Geoffrey E. Hinton","author":"Ba Jimmy Lei","year":"2016","unstructured":"Jimmy Lei Ba , Jamie Ryan Kiros, and Geoffrey E. Hinton . 2016 . Layer Normalization. Retrieved from https:\/\/arXiv:stat.ML\/1607.06450. Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer Normalization. Retrieved from https:\/\/arXiv:stat.ML\/1607.06450."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.299"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00490"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00916"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00821"},{"key":"e_1_2_2_7_1","volume-title":"Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI\u201921)","author":"Deng Yingying","year":"2021","unstructured":"Yingying Deng , Fan Tang , Weiming Dong , Haibin Huang , Ma Chongyang , and Changsheng Xu . 2021 . Arbitrary video style transfer via multi-channel correlation . In Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI\u201921) . Yingying Deng, Fan Tang, Weiming Dong, Haibin Huang, Ma Chongyang, and Changsheng Xu. 2021. Arbitrary video style transfer via multi-channel correlation. In Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI\u201921)."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3414015"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-019-0129-0"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/383259.383296"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.265"},{"key":"e_1_2_2_12_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917)","author":"Huang X.","year":"2017","unstructured":"X. Huang and S. Belongie . 2017. Arbitrary style transfer in real-time with adaptive instance normalization . In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917) . IEEE, Los Alamitos, CA, 1510\u20131519. DOI:https:\/\/doi.org\/10.1109\/ICCV. 2017 .167 10.1109\/ICCV.2017.167 X. Huang and S. Belongie. 2017. Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917). IEEE, Los Alamitos, CA, 1510\u20131519. DOI:https:\/\/doi.org\/10.1109\/ICCV.2017.167"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01219-9_11"},{"key":"e_1_2_2_14_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917)","author":"Isola Phillip","year":"2017","unstructured":"Phillip Isola , Jun-Yan Zhu , Tinghui Zhou , and Alexei A. Efros . 2017. Image-to-image translation with conditional adversarial networks . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917) . IEEE, Los Alamitos, CA, 5967\u20135976. DOI:https:\/\/doi.org\/10.1109\/CVPR. 2017 .632 10.1109\/CVPR.2017.632 Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. 2017. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917). IEEE, Los Alamitos, CA, 5967\u20135976. DOI:https:\/\/doi.org\/10.1109\/CVPR.2017.632"},{"key":"e_1_2_2_15_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2017","unstructured":"Diederik P. Kingma and Jimmy Ba . 2017 . Adam : A Method for Stochastic Optimization. Retrieved from https:\/\/arXiv:cs.LG\/1412.6980. Diederik P. Kingma and Jimmy Ba. 2017. Adam: A Method for Stochastic Optimization. Retrieved from https:\/\/arXiv:cs.LG\/1412.6980."},{"key":"e_1_2_2_16_1","volume-title":"Kingma and Max Welling","author":"Diederik","year":"2014","unstructured":"Diederik P. Kingma and Max Welling . 2014 . Auto-Encoding Variational Bayes. Retrieved from https:\/\/arXiv:stat.ML\/1312.6114. Diederik P. Kingma and Max Welling. 2014. Auto-Encoding Variational Bayes. Retrieved from https:\/\/arXiv:stat.ML\/1312.6114."},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00452"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01027"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_3"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-019-01284-z"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00393"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.5555\/3294771.3294808"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/3327144.3327184"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.5555\/3294771.3294838"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.01065"},{"key":"e_1_2_2_26_1","volume-title":"Jian Yao, Nicu Sebe, Bruno Lepri, and Xavier Alameda-Pineda.","author":"Liu Yahui","year":"2020","unstructured":"Yahui Liu , Marco De Nadai , Jian Yao, Nicu Sebe, Bruno Lepri, and Xavier Alameda-Pineda. 2020 . GMM-UNIT: Unsupervised Multi-Domain and Multi-Modal Image-to-Image Translation via Attribute Gaussian Mixture Modeling. Retrieved from https:\/\/arXiv:cs.CV\/2003.06788. Yahui Liu, Marco De Nadai, Jian Yao, Nicu Sebe, Bruno Lepri, and Xavier Alameda-Pineda. 2020. GMM-UNIT: Unsupervised Multi-Domain and Multi-Modal Image-to-Image Translation via Attribute Gaussian Mixture Modeling. Retrieved from https:\/\/arXiv:cs.CV\/2003.06788."},{"key":"e_1_2_2_27_1","volume-title":"Proceedings of the 7th International Conference on Learning Representations (ICLR\u201919)","author":"Ma Liqian","year":"2019","unstructured":"Liqian Ma , Xu Jia , Stamatios Georgoulis , Tinne Tuytelaars , and Luc Van Gool . 2019 . Exemplar guided unsupervised image-to-image translation with semantic consistency . In Proceedings of the 7th International Conference on Learning Representations (ICLR\u201919) . Retrieved from https:\/\/openreview.net\/forum?id=S1lTg3RqYQ. Liqian Ma, Xu Jia, Stamatios Georgoulis, Tinne Tuytelaars, and Luc Van Gool. 2019. Exemplar guided unsupervised image-to-image translation with semantic consistency. In Proceedings of the 7th International Conference on Learning Representations (ICLR\u201919). Retrieved from https:\/\/openreview.net\/forum?id=S1lTg3RqYQ."},{"key":"e_1_2_2_28_1","unstructured":"Alireza Makhzani Jonathon Shlens Navdeep Jaitly Ian Goodfellow and Brendan Frey. 2016. Adversarial Autoencoders. Retrieved from https:\/\/arXiv:cs.LG\/1511.05644.  Alireza Makhzani Jonathon Shlens Navdeep Jaitly Ian Goodfellow and Brendan Frey. 2016. Adversarial Autoencoders. Retrieved from https:\/\/arXiv:cs.LG\/1511.05644."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00152"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.304"},{"key":"e_1_2_2_31_1","unstructured":"Mehdi Mirza and Simon Osindero. 2014. Conditional Generative Adversarial Nets. Retrieved from https:\/\/arXiv:cs.LG\/1411.1784.  Mehdi Mirza and Simon Osindero. 2014. Conditional Generative Adversarial Nets. Retrieved from https:\/\/arXiv:cs.LG\/1411.1784."},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305890.3305954"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00603"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2019.00410"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01237-3_43"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00860"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969442.2969628"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.308"},{"key":"e_1_2_2_39_1","volume-title":"Instance Normalization: The Missing Ingredient for Fast Stylization.","author":"Ulyanov Dmitry","year":"2017","unstructured":"Dmitry Ulyanov , Andrea Vedaldi , and Victor Lempitsky . 2017 . Instance Normalization: The Missing Ingredient for Fast Stylization. Retrieved from https:\/\/arXiv:cs.CV\/1607.08022. Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. 2017. Instance Normalization: The Missing Ingredient for Fast Stylization. Retrieved from https:\/\/arXiv:cs.CV\/1607.08022."},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2952707"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00156"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454556"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-019-0136-1"},{"key":"e_1_2_2_45_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917)","author":"Zhu Jun-Yan","year":"2017","unstructured":"Jun-Yan Zhu , Taesung Park , Phillip Isola , and Alexei A. Efros . 2017. Unpaired image-to-image translation using cycle-consistent adversarial networks . In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917) . IEEE, Los Alamitos, CA, 2242\u20132251. DOI:https:\/\/doi.org\/10.1109\/ICCV. 2017 .244 10.1109\/ICCV.2017.244 Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. 2017. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917). IEEE, Los Alamitos, CA, 2242\u20132251. DOI:https:\/\/doi.org\/10.1109\/ICCV.2017.244"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3450525","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3450525","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:47:49Z","timestamp":1750193269000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3450525"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,22]]},"references-count":45,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2021,8,31]]}},"alternative-id":["10.1145\/3450525"],"URL":"https:\/\/doi.org\/10.1145\/3450525","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,7,22]]},"assertion":[{"value":"2020-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-07-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}