{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,25]],"date-time":"2025-09-25T18:08:47Z","timestamp":1758823727188,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":39,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T00:00:00Z","timestamp":1602460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,12]]},"DOI":"10.1145\/3394171.3413785","type":"proceedings-article","created":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T12:27:38Z","timestamp":1602505658000},"page":"3164-3172","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Retrieval Guided Unsupervised Multi-domain Image to Image Translation"],"prefix":"10.1145","author":[{"given":"Raul","family":"Gomez","sequence":"first","affiliation":[{"name":"Centre Tecnologic de Catalunya - Computer Vision Center, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yahui","family":"Liu","sequence":"additional","affiliation":[{"name":"University of Trento &amp;Fondazione Bruno Kessler, Trento, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marco","family":"De Nadai","sequence":"additional","affiliation":[{"name":"Fondazione Bruno Kessler, Trento, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dimosthenis","family":"Karatzas","sequence":"additional","affiliation":[{"name":"Universitat Autonoma de Barcelona - Computer Vision Centre, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bruno","family":"Lepri","sequence":"additional","affiliation":[{"name":"Fondazione Bruno Kessler, Trento, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nicu","family":"Sebe","sequence":"additional","affiliation":[{"name":"University of Trento &amp; Huawei Ireland, Trento, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,10,12]]},"reference":[{"key":"e_1_3_2_2_1_1","unstructured":"Amjad Almahairi Sai Rajeswar Alessandro Sordoni Philip Bachman and Aaron Courville. 2018. Augmented cyclegan: Learning many-to-many mappings from unpaired data. In ICML.  Amjad Almahairi Sai Rajeswar Alessandro Sordoni Philip Bachman and Aaron Courville. 2018. Augmented cyclegan: Learning many-to-many mappings from unpaired data. In ICML."},{"key":"e_1_3_2_2_2_1","volume-title":"High-Resolution Daytime Translation Without Domain Labels. arXiv preprint arXiv:2003.08791","author":"Anokhin Ivan","year":"2020","unstructured":"Ivan Anokhin , Pavel Solovev , Denis Korzhenkov , Alexey Kharlamov , Taras Khakhulin , Gleb Sterkin , Alexey Silvestrov , Sergey Nikolenko , and Victor Lempitsky . 2020. High-Resolution Daytime Translation Without Domain Labels. arXiv preprint arXiv:2003.08791 ( 2020 ). Ivan Anokhin, Pavel Solovev, Denis Korzhenkov, Alexey Kharlamov, Taras Khakhulin, Gleb Sterkin, Alexey Silvestrov, Sergey Nikolenko, and Victor Lempitsky. 2020. High-Resolution Daytime Translation Without Domain Labels. arXiv preprint arXiv:2003.08791 (2020)."},{"volume-title":"Night-to-day image translation for retrieval-based localization. In 2019 ICRA","author":"Anoosheh Asha","key":"e_1_3_2_2_3_1","unstructured":"Asha Anoosheh , Torsten Sattler , Radu Timofte , Marc Pollefeys , and Luc Van Gool . 2019. Night-to-day image translation for retrieval-based localization. In 2019 ICRA . IEEE , 5958--5964. Asha Anoosheh, Torsten Sattler, Radu Timofte, Marc Pollefeys, and Luc Van Gool. 2019. Night-to-day image translation for retrieval-based localization. In 2019 ICRA. IEEE, 5958--5964."},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1195"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/1756006.1756042"},{"key":"e_1_3_2_2_6_1","volume-title":"Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In CVPR. 8789--8797.","author":"Choi Yunjey","year":"2018","unstructured":"Yunjey Choi , Minje Choi , Munyoung Kim , Jung-Woo Ha , Sunghun Kim , and Jaegul Choo . 2018 . Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In CVPR. 8789--8797. Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. 2018. Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In CVPR. 8789--8797."},{"key":"e_1_3_2_2_7_1","volume-title":"Self-Supervised Learning from Web Data for Multimodal Retrieval. Multi-Modal Scene Understanding","author":"Gomez Raul","year":"2019","unstructured":"Raul Gomez , Lluis Gomez , Jaume Gibert , and Dimosthenis Karatzas . 2019. Self-Supervised Learning from Web Data for Multimodal Retrieval. Multi-Modal Scene Understanding ( 2019 ). Raul Gomez, Lluis Gomez, Jaume Gibert, and Dimosthenis Karatzas. 2019. Self-Supervised Learning from Web Data for Multimodal Retrieval. Multi-Modal Scene Understanding (2019)."},{"key":"e_1_3_2_2_8_1","volume-title":"Joost Van De Weijer, and Yoshua Bengio","author":"Gonzalez-Garcia Abel","year":"2018","unstructured":"Abel Gonzalez-Garcia , Joost Van De Weijer, and Yoshua Bengio . 2018 . Image-to-image translation for cross-domain disentanglement. In Advances in neural information processing systems. 1287--1298. Abel Gonzalez-Garcia, Joost Van De Weijer, and Yoshua Bengio. 2018. Image-to-image translation for cross-domain disentanglement. In Advances in neural information processing systems. 1287--1298."},{"key":"e_1_3_2_2_9_1","unstructured":"Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in neural information processing systems. 2672--2680.  Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in neural information processing systems. 2672--2680."},{"key":"e_1_3_2_2_10_1","volume-title":"Beyond Instance-Level Image Retrieval: Leveraging Captions to Learn a Global Visual Representation for Semantic Retrieval. CVPR","author":"Gordo Albert","year":"2017","unstructured":"Albert Gordo and Diane Larlus . 2017. Beyond Instance-Level Image Retrieval: Leveraging Captions to Learn a Global Visual Representation for Semantic Retrieval. CVPR ( 2017 ). Albert Gordo and Diane Larlus. 2017. Beyond Instance-Level Image Retrieval: Leveraging Captions to Learn a Global Visual Representation for Semantic Retrieval. CVPR (2017)."},{"key":"e_1_3_2_2_11_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR. 770--778.  Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR. 770--778."},{"key":"e_1_3_2_2_12_1","unstructured":"Martin Heusel Hubert Ramsauer Thomas Unterthiner Bernhard Nessler and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NIPS.  Martin Heusel Hubert Ramsauer Thomas Unterthiner Bernhard Nessler and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NIPS."},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"crossref","unstructured":"Xun Huang Ming-Yu Liu Serge Belongie and Jan Kautz. 2018. Multimodal unsupervised image-to-image translation. In ECCV. 172--189.  Xun Huang Ming-Yu Liu Serge Belongie and Jan Kautz. 2018. Multimodal unsupervised image-to-image translation. In ECCV. 172--189.","DOI":"10.1007\/978-3-030-01219-9_11"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"crossref","unstructured":"Phillip Isola Jun-Yan Zhu Tinghui Zhou and Alexei A Efros. 2017. Image-to-image translation with conditional adversarial networks. In CVPR. 1125--1134.  Phillip Isola Jun-Yan Zhu Tinghui Zhou and Alexei A Efros. 2017. Image-to-image translation with conditional adversarial networks. In CVPR. 1125--1134.","DOI":"10.1109\/CVPR.2017.632"},{"key":"e_1_3_2_2_15_1","unstructured":"Andrej Karpathy Armand Joulin and Li Fei-Fei. 2014. Deep Fragment Embeddings for Bidirectional Image Sentence Mapping. In NIPS.  Andrej Karpathy Armand Joulin and Li Fei-Fei. 2014. Deep Fragment Embeddings for Bidirectional Image Sentence Mapping. In NIPS."},{"key":"e_1_3_2_2_16_1","volume-title":"Zemel","author":"Kiros Ryan","year":"2014","unstructured":"Ryan Kiros , Ruslan Salakhutdinov , and Richard S . Zemel . 2014 . Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models . arXiv (2014). Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel. 2014. Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models. arXiv (2014)."},{"key":"e_1_3_2_2_17_1","unstructured":"Hsin-Ying Lee Hung-Yu Tseng Jia-Bin Huang Maneesh Singh and Ming-Hsuan Yang. 2018. Diverse image-to-image translation via disentangled representations. In ECCV. 35--51.  Hsin-Ying Lee Hung-Yu Tseng Jia-Bin Huang Maneesh Singh and Ming-Hsuan Yang. 2018. Diverse image-to-image translation via disentangled representations. In ECCV. 35--51."},{"key":"e_1_3_2_2_18_1","unstructured":"Ming-Yu Liu Thomas Breuel and Jan Kautz. 2017. Unsupervised image-to-image translation networks. In NIPS. 700--708.  Ming-Yu Liu Thomas Breuel and Jan Kautz. 2017. Unsupervised image-to-image translation networks. In NIPS. 700--708."},{"key":"e_1_3_2_2_19_1","unstructured":"Ming-Yu Liu Xun Huang Arun Mallya Tero Karras Timo Aila Jaakko Lehtinen and Jan Kautz. 2019. Few-shot unsupervised image-to-image translation. In ICCV. 10551--10560.  Ming-Yu Liu Xun Huang Arun Mallya Tero Karras Timo Aila Jaakko Lehtinen and Jan Kautz. 2019. Few-shot unsupervised image-to-image translation. In ICCV. 10551--10560."},{"key":"e_1_3_2_2_20_1","volume-title":"Jian Yao, Nicu Sebe, Bruno Lepri, and Xavier Alameda-Pineda.","author":"Liu Yahui","year":"2020","unstructured":"Yahui Liu , Marco De Nadai , Jian Yao, Nicu Sebe, Bruno Lepri, and Xavier Alameda-Pineda. 2020 . GMM-UNIT: Unsupervised Multi-Domain and Multi-Modal Image-to-Image Translation via Attribute Gaussian Mixture Modeling . arXiv preprint arXiv:2003.06788 (2020). Yahui Liu, Marco De Nadai, Jian Yao, Nicu Sebe, Bruno Lepri, and Xavier Alameda-Pineda. 2020. GMM-UNIT: Unsupervised Multi-Domain and Multi-Modal Image-to-Image Translation via Attribute Gaussian Mixture Modeling. arXiv preprint arXiv:2003.06788 (2020)."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"crossref","unstructured":"Ziwei Liu Ping Luo Xiaogang Wang and Xiaoou Tang. 2015. Deep learning face attributes in the wild. In ICCV. 3730--3738.  Ziwei Liu Ping Luo Xiaogang Wang and Xiaoou Tang. 2015. Deep learning face attributes in the wild. In ICCV. 3730--3738.","DOI":"10.1109\/ICCV.2015.425"},{"key":"e_1_3_2_2_22_1","volume-title":"Zhen Wang, and Stephen Paul Smolley.","author":"Mao Xudong","year":"2017","unstructured":"Xudong Mao , Qing Li , Haoran Xie , Raymond YK Lau , Zhen Wang, and Stephen Paul Smolley. 2017 . Least squares generative adversarial networks. In ICCV. 2794--2802. Xudong Mao, Qing Li, Haoran Xie, Raymond YK Lau, Zhen Wang, and Stephen Paul Smolley. 2017. Least squares generative adversarial networks. In ICCV. 2794--2802."},{"key":"e_1_3_2_2_23_1","unstructured":"Youssef Alami Mejjati Christian Richardt James Tompkin Darren Cosker and Kwang In Kim. 2018. Unsupervised attention-guided image-to-image translation. In NeurIPS. 3693--3703.  Youssef Alami Mejjati Christian Richardt James Tompkin Darren Cosker and Kwang In Kim. 2018. Unsupervised attention-guided image-to-image translation. In NeurIPS. 3693--3703."},{"key":"e_1_3_2_2_24_1","volume-title":"Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784","author":"Mirza Mehdi","year":"2014","unstructured":"Mehdi Mirza and Simon Osindero . 2014. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784 ( 2014 ). Mehdi Mirza and Simon Osindero. 2014. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784 (2014)."},{"key":"e_1_3_2_2_25_1","unstructured":"Sangwoo Mo Minsu Cho and Jinwoo Shin. 2019. Instance-aware Image-to-Image Translation. In ICLR.  Sangwoo Mo Minsu Cho and Jinwoo Shin. 2019. Instance-aware Image-to-Image Translation. In ICLR."},{"key":"e_1_3_2_2_26_1","volume-title":"Self-Supervised Visual Representations for Cross-Modal Retrieval. In International Conference on Multimedia Retrieval .","author":"Patel Yash","year":"2019","unstructured":"Yash Patel , Lluis Gomez , Marcc al Rusi n ol, Dimosthenis Karatzas , and C V Jawahar . 2019 . Self-Supervised Visual Representations for Cross-Modal Retrieval. In International Conference on Multimedia Retrieval . Yash Patel, Lluis Gomez, Marcc al Rusi n ol, Dimosthenis Karatzas, and C V Jawahar. 2019. Self-Supervised Visual Representations for Cross-Modal Retrieval. In International Conference on Multimedia Retrieval ."},{"key":"e_1_3_2_2_27_1","volume-title":"Ganimation: Anatomically-aware facial animation from a single image. In ECCV . 818--833.","author":"Pumarola Albert","year":"2018","unstructured":"Albert Pumarola , Antonio Agudo , Aleix M Martinez , Alberto Sanfeliu , and Francesc Moreno-Noguer . 2018 . Ganimation: Anatomically-aware facial animation from a single image. In ECCV . 818--833. Albert Pumarola, Antonio Agudo, Aleix M Martinez, Alberto Sanfeliu, and Francesc Moreno-Noguer. 2018. Ganimation: Anatomically-aware facial animation from a single image. In ECCV . 818--833."},{"key":"e_1_3_2_2_28_1","volume-title":"International Journal of Computer Science and Information Technologies","author":"Satpute Bharti S","year":"2013","unstructured":"Bharti S Satpute . 2013 . Content-Based Face Image Retrieval Using Attribute-Enhanced Sparse Codewords . International Journal of Computer Science and Information Technologies , (2013). Bharti S Satpute. 2013. Content-Based Face Image Retrieval Using Attribute-Enhanced Sparse Codewords. International Journal of Computer Science and Information Technologies, (2013)."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"crossref","unstructured":"Walter J Scheirer Neeraj Kumar Peter N Belhumeur and Terrance E Boult. 2012. Multi-Attribute Spaces: Calibration for Attribute Fusion and Similarity Search. In CVPR.  Walter J Scheirer Neeraj Kumar Peter N Belhumeur and Terrance E Boult. 2012. Multi-Attribute Spaces: Calibration for Attribute Fusion and Similarity Search. In CVPR.","DOI":"10.1109\/CVPR.2012.6248021"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"crossref","unstructured":"Florian Schroff Dmitry Kalenichenko and James Philbin. 2015. FaceNet: A Unified Embedding for Face Recognition and Clustering. In CVPR.  Florian Schroff Dmitry Kalenichenko and James Philbin. 2015. FaceNet: A Unified Embedding for Face Recognition and Clustering. In CVPR.","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"e_1_3_2_2_31_1","volume-title":"Davis","author":"Siddiquie Behjat","year":"2011","unstructured":"Behjat Siddiquie , Rogerio S. Feris , and Larry S . Davis . 2011 . Image ranking and retrieval based on multi-attribute queries. In CVPR. Behjat Siddiquie, Rogerio S. Feris, and Larry S. Davis. 2011. Image ranking and retrieval based on multi-attribute queries. In CVPR."},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"crossref","unstructured":"Brandon M Smith Shengqi Zhu and Li Zhang. 2011. Face Image Retrieval by Shape Manipulation. In CVPR.  Brandon M Smith Shengqi Zhu and Li Zhang. 2011. Face Image Retrieval by Shape Manipulation. In CVPR.","DOI":"10.1109\/CVPR.2011.5995471"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"crossref","unstructured":"Jiang Wang Yang Song Thomas Leung Chuck Rosenberg Jingbin Wang James Philbin Bo Chen and Ying Wu. 2014. Learning Fine-Grained Image Similarity with Deep Ranking. In CVPR .  Jiang Wang Yang Song Thomas Leung Chuck Rosenberg Jingbin Wang James Philbin Bo Chen and Ying Wu. 2014. Learning Fine-Grained Image Similarity with Deep Ranking. In CVPR .","DOI":"10.1109\/CVPR.2014.180"},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"crossref","unstructured":"Liwei Wang Yin Li Svetlana Lazebnik and Slazebni@illinois Edu. 2016. Learning Deep Structure-Preserving Image-Text Embeddings. CVPR.  Liwei Wang Yin Li Svetlana Lazebnik and Slazebni@illinois Edu. 2016. Learning Deep Structure-Preserving Image-Text Embeddings. CVPR.","DOI":"10.1109\/CVPR.2016.541"},{"key":"e_1_3_2_2_35_1","unstructured":"Ting-Chun Wang Ming-Yu Liu Jun-Yan Zhu Andrew Tao Jan Kautz and Bryan Catanzaro. 2018. High-resolution image synthesis and semantic manipulation with conditional gans. In CVPR. 8798--8807.  Ting-Chun Wang Ming-Yu Liu Jun-Yan Zhu Andrew Tao Jan Kautz and Bryan Catanzaro. 2018. High-resolution image synthesis and semantic manipulation with conditional gans. In CVPR. 8798--8807."},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"crossref","unstructured":"Zhong Wu Qifa Ke Jian Sun and Heung-Yeung Shum. 2010. Scalable Face Image Retrieval with Identity-Based Quantization and Multi-Reference Re-ranking. In CVPR.  Zhong Wu Qifa Ke Jian Sun and Heung-Yeung Shum. 2010. Scalable Face Image Retrieval with Identity-Based Quantization and Multi-Reference Re-ranking. In CVPR.","DOI":"10.1109\/TPAMI.2011.111"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"crossref","unstructured":"Richard Zhang Phillip Isola Alexei A Efros Eli Shechtman and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR. 586--595.  Richard Zhang Phillip Isola Alexei A Efros Eli Shechtman and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR. 586--595.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_3_2_2_38_1","unstructured":"Jun-Yan Zhu Taesung Park Phillip Isola and Alexei A Efros. 2017a. Unpaired image-to-image translation using cycle-consistent adversarial networks. In ICCV. 2223--2232.  Jun-Yan Zhu Taesung Park Phillip Isola and Alexei A Efros. 2017a. Unpaired image-to-image translation using cycle-consistent adversarial networks. In ICCV. 2223--2232."},{"key":"e_1_3_2_2_39_1","unstructured":"Jun-Yan Zhu Richard Zhang Deepak Pathak Trevor Darrell Alexei A Efros Oliver Wang and Eli Shechtman. 2017b. Toward multimodal image-to-image translation. In NIPS. 465--476.  Jun-Yan Zhu Richard Zhang Deepak Pathak Trevor Darrell Alexei A Efros Oliver Wang and Eli Shechtman. 2017b. Toward multimodal image-to-image translation. In NIPS. 465--476."}],"event":{"name":"MM '20: The 28th ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Seattle WA USA","acronym":"MM '20"},"container-title":["Proceedings of the 28th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413785","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3394171.3413785","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:01:17Z","timestamp":1750197677000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413785"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,12]]},"references-count":39,"alternative-id":["10.1145\/3394171.3413785","10.1145\/3394171"],"URL":"https:\/\/doi.org\/10.1145\/3394171.3413785","relation":{},"subject":[],"published":{"date-parts":[[2020,10,12]]},"assertion":[{"value":"2020-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}