{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T15:59:01Z","timestamp":1778083141769,"version":"3.51.4"},"reference-count":79,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,7,26]],"date-time":"2023-07-26T00:00:00Z","timestamp":1690329600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Ministry of Culture, Sports and Tourism and Korea Creative Content Agency"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2023,8]]},"abstract":"<jats:p>Sketches reflect the drawing style of individual artists; therefore, it is important to consider their unique styles when extracting sketches from color images for various applications. Unfortunately, most existing sketch extraction methods are designed to extract sketches of a single style. Although there have been some attempts to generate various style sketches, the methods generally suffer from two limitations: low quality results and difficulty in training the model due to the requirement of a paired dataset. In this paper, we propose a novel multi-modal sketch extraction method that can imitate the style of a given reference sketch with unpaired data training in a semi-supervised manner. Our method outperforms state-of-the-art sketch extraction methods and unpaired image translation methods in both quantitative and qualitative evaluations.<\/jats:p>","DOI":"10.1145\/3592392","type":"journal-article","created":{"date-parts":[[2023,7,26]],"date-time":"2023-07-26T14:29:21Z","timestamp":1690381761000},"page":"1-12","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":16,"title":["Semi-supervised reference-based sketch extraction using a contrastive learning framework"],"prefix":"10.1145","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3809-9515","authenticated-orcid":false,"given":"Chang Wook","family":"Seo","sequence":"first","affiliation":[{"name":"Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5352-7306","authenticated-orcid":false,"given":"Amirsaman","family":"Ashtari","sequence":"additional","affiliation":[{"name":"Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1925-3326","authenticated-orcid":false,"given":"Junyong","family":"Noh","sequence":"additional","affiliation":[{"name":"Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,7,26]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3550454.3555504"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.1986.4767851"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00776"},{"key":"e_1_2_2_4_1","volume-title":"International conference on machine learning. PMLR, 1597--1607","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning. PMLR, 1597--1607."},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00981"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00821"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240661"},{"key":"e_1_2_2_8_1","unstructured":"DanbooruCommunity. 2021. Danbooru2020: A Large-Scale Crowdsourced and Tagged Anime Illustration Dataset. https:\/\/www.gwern.net\/Danbooru2020. https:\/\/www.gwern.net\/Danbooru2020 Accessed: 2022\/04\/03."},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00522"},{"key":"e_1_2_2_10_1","first-page":"1","article-title":"Deep exemplar-based colorization","volume":"37","author":"He Mingming","year":"2018","unstructured":"Mingming He, Dongdong Chen, Jing Liao, Pedro V Sander, and Lu Yuan. 2018. Deep exemplar-based colorization. ACM Transactions on Graphics (TOG) 37, 4 (2018), 1--16.","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295408"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.167"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01219-9_11"},{"key":"e_1_2_2_14_1","volume-title":"Image-to-Image Translation with Conditional Adversarial Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 5967--5976","author":"Isola P.","unstructured":"P. Isola, J. Zhu, T. Zhou, and A. A. Efros. 2017. Image-to-Image Translation with Conditional Adversarial Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 5967--5976."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"e_1_2_2_16_1","volume-title":"Diffusionclip: Text-guided image manipulation using diffusion models.","author":"Kim Gwanghyun","year":"2021","unstructured":"Gwanghyun Kim and Jong Chul Ye. 2021. Diffusionclip: Text-guided image manipulation using diffusion models. (2021)."},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00915"},{"key":"e_1_2_2_18_1","volume-title":"International conference on machine learning. PMLR","author":"Kim Taeksoo","year":"2017","unstructured":"Taeksoo Kim, Moonsu Cha, Hyunsoo Kim, Jung Kwon Lee, and Jiwon Kim. 2017. Learning to discover cross-domain relations with generative adversarial networks. In International conference on machine learning. PMLR, 1857--1865."},{"key":"e_1_2_2_19_1","volume-title":"Adam: A Method for Stochastic Optimization. International Conference on Learning Representations (12","author":"Kingma Diederik","year":"2014","unstructured":"Diederik Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Optimization. International Conference on Learning Representations (12 2014)."},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.521"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-019-01284-z"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00584"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00242"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00790"},{"key":"e_1_2_2_25_1","first-page":"22020","article-title":"Lightweight generative adversarial networks for text-guided image manipulation","volume":"33","author":"Li Bowen","year":"2020","unstructured":"Bowen Li, Xiaojuan Qi, Philip Torr, and Thomas Lukasiewicz. 2020b. Lightweight generative adversarial networks for text-guided image manipulation. Advances in Neural Information Processing Systems 33 (2020), 22020--22031.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_2_26_1","first-page":"1","article-title":"Deep extraction of manga structural lines","volume":"36","author":"Li Chengze","year":"2017","unstructured":"Chengze Li, Xueting Liu, and Tien-Tsin Wong. 2017a. Deep extraction of manga structural lines. ACM Transactions on Graphics (TOG) 36, 4 (2017), 1--12.","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2020.2987417"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350854"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0963-9"},{"key":"e_1_2_2_30_1","volume-title":"Proceedings of the Asian Conference on Computer Vision.","author":"Liu Bingchen","year":"2020","unstructured":"Bingchen Liu, Kunpeng Song, Yizhe Zhu, and Ahmed Elgammal. 2020b. Sketch-to-art: Synthesizing stylized art images from sketches. In Proceedings of the Asian Conference on Computer Vision."},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-021-0228-6"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i1.16111"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413505"},{"key":"e_1_2_2_34_1","unstructured":"lllyasviel. 2017. sketchKeras. https:\/\/github.com\/lllyasviel\/sketchKeras."},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.1982.1056489"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.3390\/app12010327"},{"key":"e_1_2_2_37_1","volume-title":"Exemplar guided unsupervised image-to-image translation with semantic consistency. arXiv preprint arXiv:1805.11145","author":"Ma Liqian","year":"2018","unstructured":"Liqian Ma, Xu Jia, Stamatios Georgoulis, Tinne Tuytelaars, and Luc Van Gool. 2018a. Exemplar guided unsupervised image-to-image translation with semantic consistency. arXiv preprint arXiv:1805.11145 (2018)."},{"key":"e_1_2_2_38_1","volume-title":"Exemplar guided unsupervised image-to-image translation with semantic consistency. arXiv preprint arXiv:1805.11145","author":"Ma Liqian","year":"2018","unstructured":"Liqian Ma, Xu Jia, Stamatios Georgoulis, Tinne Tuytelaars, and Luc Van Gool. 2018b. Exemplar guided unsupervised image-to-image translation with semantic consistency. arXiv preprint arXiv:1805.11145 (2018)."},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459833"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00788"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58542-6_24"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58545-7_19"},{"key":"e_1_2_2_43_1","first-page":"7198","article-title":"Swapping autoencoder for deep image manipulation","volume":"33","author":"Park Taesung","year":"2020","unstructured":"Taesung Park, Jun-Yan Zhu, Oliver Wang, Jingwan Lu, Eli Shechtman, Alexei Efros, and Richard Zhang. 2020b. Swapping autoencoder for deep image manipulation. Advances in Neural Information Processing Systems 33 (2020), 7198--7211.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00099"},{"key":"e_1_2_2_45_1","unstructured":"ref2sketch. 2022. Ref2sketch official page. https:\/\/github.com\/ref2sketch\/ref2sketch."},{"key":"e_1_2_2_46_1","doi-asserted-by":"crossref","unstructured":"Tamar Rott Shaham Michael Gharbi Richard Zhang Eli Shechtman and Tomer Michaeli. 2021. Spatially-Adaptive Pixelwise Networks for Fast Image Translation. In Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR46437.2021.01464"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1049\/el.2018.6167"},{"key":"e_1_2_2_48_1","volume-title":"HSEmotion: High-speed emotion recognition library. Software Impacts","author":"Savchenko Andrey V","year":"2022","unstructured":"Andrey V Savchenko. 2022. HSEmotion: High-speed emotion recognition library. Software Impacts (2022), 100433."},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00387"},{"key":"e_1_2_2_50_1","volume-title":"Jingwan Lu, Joon-Young Lee, Seonghyeon Kim, and Junyong Noh.","author":"Seo Kwanggyoon","year":"2022","unstructured":"Kwanggyoon Seo, Seoung Wug Oh, Jingwan Lu, Joon-Young Lee, Seonghyeon Kim, and Junyong Noh. 2022. StylePortraitVideo: Editing Portrait Videos with Expression Optimization. (2022)."},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00020"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00359"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132703"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925972"},{"key":"e_1_2_2_55_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_2_2_56_1","volume-title":"You only need adversarial supervision for semantic image synthesis. arXiv preprint arXiv:2012.04781","author":"Sushko Vadim","year":"2020","unstructured":"Vadim Sushko, Edgar Sch\u00f6nfeld, Dan Zhang, Juergen Gall, Bernt Schiele, and Anna Khoreva. 2020. You only need adversarial supervision for semantic image synthesis. arXiv preprint arXiv:2012.04781 (2020)."},{"key":"e_1_2_2_57_1","volume-title":"Philip HS Torr, and Nicu Sebe","author":"Tang Hao","year":"2022","unstructured":"Hao Tang, Philip HS Torr, and Nicu Sebe. 2022. Multi-Channel Attention Selection GANs for Guided Image-to-Image Translation. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) (2022)."},{"key":"e_1_2_2_58_1","doi-asserted-by":"crossref","unstructured":"Hao Tang Dan Xu Nicu Sebe Yanzhi Wang Jason J. Corso and Yan Yan. 2019. MultiChannel Attention Selection GAN with Cascaded Semantic Guidance for Cross-View Image Translation. In CVPR.","DOI":"10.1109\/CVPR.2019.00252"},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2019.00388"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528223.3530068"},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_47"},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00917"},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/2024676.2024700"},{"key":"e_1_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"e_1_2_2_66_1","unstructured":"Xiao Yang Yiheng Zhu Xiaohui Shen Xiaoyu Xiang Ding Liu. 2021. Anime2Sketch: A Sketch Extractor for Anime Arts with Deep Networks. https:\/\/github.com\/Mukosame\/Anime2Sketch."},{"key":"e_1_2_2_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01391"},{"key":"e_1_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.164"},{"key":"e_1_2_2_69_1","volume-title":"Pano2cad: Room layout from a single panorama image. In 2017 IEEE winter conference on applications of computer vision (WACV)","author":"Xu Jiu","unstructured":"Jiu Xu, Bj\u00f6rn Stenger, Tommi Kerola, and Tony Tung. 2017. Pano2cad: Room layout from a single panorama image. In 2017 IEEE winter conference on applications of computer vision (WACV). IEEE, 354--362."},{"key":"e_1_2_2_70_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2019.2930512"},{"key":"e_1_2_2_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417784"},{"key":"e_1_2_2_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01100"},{"key":"e_1_2_2_73_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00824"},{"key":"e_1_2_2_74_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00442"},{"key":"e_1_2_2_75_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3272127.3275090","article-title":"Two-stage sketch colorization","volume":"37","author":"Zhang Lvmin","year":"2018","unstructured":"Lvmin Zhang, Chengze Li, Tien-Tsin Wong, Yi Ji, and Chunping Liu. 2018b. Two-stage sketch colorization. ACM Transactions on Graphics (TOG) 37, 6 (2018), 1--14.","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_2_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_2_2_77_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58523-5_43"},{"key":"e_1_2_2_78_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.244"},{"key":"e_1_2_2_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/3355089.3356561"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3592392","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3592392","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:48:59Z","timestamp":1750182539000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3592392"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,26]]},"references-count":79,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,8]]}},"alternative-id":["10.1145\/3592392"],"URL":"https:\/\/doi.org\/10.1145\/3592392","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,26]]},"assertion":[{"value":"2023-07-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}