{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,19]],"date-time":"2026-02-19T15:43:14Z","timestamp":1771515794284,"version":"3.50.1"},"reference-count":50,"publisher":"MDPI AG","issue":"19","license":[{"start":{"date-parts":[[2020,10,2]],"date-time":"2020-10-02T00:00:00Z","timestamp":1601596800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Virtual Try-on is the ability to realistically superimpose clothing onto a target person. Due to its importance to the multi-billion dollar e-commerce industry, the problem has received significant attention in recent years. To date, most virtual try-on methods have been supervised approaches, namely using annotated data, such as clothes parsing semantic segmentation masks and paired images. These approaches incur a very high cost in annotation. Even existing weakly-supervised virtual try-on methods still use annotated data or pre-trained networks as auxiliary information and the costs of the annotation are still significantly high. Plus, the strategy using pre-trained networks is not appropriate in the practical scenarios due to latency. In this paper we propose Unsupervised VIRtual Try-on using disentangled representation (UVIRT). After UVIRT extracts a clothes and a person feature from a person image and a clothes image respectively, it exchanges a clothes and a person feature. Finally, UVIRT achieve virtual try-on. This is all achieved in an unsupervised manner so UVIRT has the advantage that it does not require any annotated data, pre-trained networks nor even category labels. In the experiments, we qualitatively and quantitatively compare between supervised methods and our UVIRT method on the MPV dataset (which has paired images) and on a Consumer-to-Consumer (C2C) marketplace dataset (which has unpaired images). As a result, UVIRT outperform the supervised method on the C2C marketplace dataset, and achieve comparable results on the MPV dataset, which has paired images in comparison with the conventional supervised method.<\/jats:p>","DOI":"10.3390\/s20195647","type":"journal-article","created":{"date-parts":[[2020,10,2]],"date-time":"2020-10-02T09:39:25Z","timestamp":1601631565000},"page":"5647","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["UVIRT\u2014Unsupervised Virtual Try-on Using Disentangled Clothing and Person Features"],"prefix":"10.3390","volume":"20","author":[{"given":"Hideki","family":"Tsunashima","sequence":"first","affiliation":[{"name":"Computer Vision Research Team, Artificial Intelligence Research Center, National Institute of Advanced Industrial Science and Technology (AIST), Tsukuba 305-8560, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kosuke","family":"Arase","sequence":"additional","affiliation":[{"name":"Mercari, Inc., Tokyo 106-6188, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Antony","family":"Lam","sequence":"additional","affiliation":[{"name":"Mercari, Inc., Tokyo 106-6188, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8844-165X","authenticated-orcid":false,"given":"Hirokatsu","family":"Kataoka","sequence":"additional","affiliation":[{"name":"Computer Vision Research Team, Artificial Intelligence Research Center, National Institute of Advanced Industrial Science and Technology (AIST), Tsukuba 305-8560, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,10,2]]},"reference":[{"key":"ref_1","unstructured":"Statista (2020, August 04). eBay: Annual Net Revenue 2013\u20132019. Available online: https:\/\/www.statista.com\/statistics\/507881\/ebays-annual-net-revenue\/."},{"key":"ref_2","unstructured":"Amazon (2020, August 04). Prime Wardrobe. Available online: https:\/\/www.amazon.co.jp\/b\/?ie=UTF8&bbn=5429200051&node=5425661051&tag=googhydr-22&ref=pd_sl_pn6acna5m_b&adgrpid=60466354043&hvpone=&hvptwo=&hvadid=338936031340&hvpos=&hvnetw=g&hvrand=9233286119764777328&hvqmt=b&hvdev=c&hvdvcmdl=&hvlocint=&hvlocphy=1009310&hvtargid=kwd-327794409245&hydadcr=3627_11172809&gclid=Cj0KCQjw6_vzBRCIARIsAOs54z7lvkoeDTmckFM2Vakdras4uWdqWSe2ESCzS6dRoX_fTNLNE8Y5p-8aAq1REALw_wcB."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Han, X., Wu, Z., Wu, Z., Yu, R., and Davis, L.S. (2018, January 19\u201321). VITON: An Image-based Virtual Try-on Network. Proceedings of the CVPR 2021: IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00787"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wang, B., Zheng, H., Liang, X., Chen, Y., and Lin, L. (2018, January 8\u201314). Toward Characteristic-Preserving Image-based Virtual Try-On Network. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01261-8_36"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Hsieh, C., Chen, C., Chou, C., Shuai, H., and Cheng, W. (2019, January 22\u201325). Fit-me: Image-Based Virtual Try-on With Arbitrary Poses. Proceedings of the 2019 IEEE International Conference on Image Processing (ICIP), Taipei, Taiwan.","DOI":"10.1109\/ICIP.2019.8803681"},{"key":"ref_6","unstructured":"Yildirim, G., Jetchev, N., Vollgraf, R., and Bergmann, U. (November, January 27). Generating High-Resolution Fashion Model Images Wearing Custom Outfits. Proceedings of the IEEE International Conference on Computer Vision (ICCV) Workshops, Seoul, Korea."},{"key":"ref_7","unstructured":"Yu, R., Wang, X., and Xie, X. (November, January 27). VTNFP: An Image-Based Virtual Try-On Network With Body and Clothing Feature Preservation. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Seoul, Korea."},{"key":"ref_8","unstructured":"Jae Lee, H., Lee, R., Kang, M., Cho, M., and Park, G. (November, January 27). LA-VITON: A Network for Looking-Attractive Virtual Try-On. Proceedings of the IEEE International Conference on Computer Vision (ICCV) Workshops, Seoul, Korea."},{"key":"ref_9","unstructured":"Issenhuth, T., Mary, J., and Calauz\u00e8nes, C. (2019). End-to-End Learning of Geometric Deformations of Feature Maps for Virtual Try-On. arXiv."},{"key":"ref_10","unstructured":"Dong, H., Liang, X., Shen, X., Wang, B., Lai, H., Zhu, J., Hu, Z., and Yin, J. (November, January 27). Towards Multi-Pose Guided Virtual Try-On Network. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Seoul, Korea."},{"key":"ref_11","unstructured":"Pumarola, A., Goswami, V., Vicente, F., De la Torre, F., and Moreno-Noguer, F. (November, January 27). Unsupervised Image-to-Video Clothing Transfer. Proceedings of the IEEE International Conference on Computer Vision (ICCV) Workshops, Seoul, Korea."},{"key":"ref_12","unstructured":"Wang, K., Ma, L., Oramas, J., Gool, L.V., and Tuytelaars, T. (2018). Unsupervised shape transformer for image translation and cross-domain retrieval. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Jetchev, N., and Bergmann, U. (2017, January 22\u201329). The Conditional Analogy GAN: Swapping Fashion Articles on People Images. Proceedings of the IEEE International Conference on Computer Vision (ICCV) Workshops, Venice, Italy.","DOI":"10.1109\/ICCVW.2017.269"},{"key":"ref_14","unstructured":"Mercari (2020, August 04). The Total Merchandise Number of C2C Market Application \u201cMercari\u201d Goes Beyond 10 Billion (Japanese Website). Available online: https:\/\/about.mercari.com\/press\/news\/article\/20180719_billionitems\/."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1798","DOI":"10.1109\/TPAMI.2013.50","article-title":"Representation Learning: A Review and New Perspectives","volume":"35","author":"Bengio","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_16","unstructured":"Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A. (2017, January 24\u201326). beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework. Proceedings of the International Conference on Learning Representations, Toulon, France."},{"key":"ref_17","unstructured":"Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., and Abbeel, P. (2016, January 5\u201310). InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets. Proceedings of the Advances in Neural Information Processing Systems 29, Centre Convencions Internacional Barcelona, Barcelona, Spain."},{"key":"ref_18","unstructured":"Kim, H., and Mnih, A. (2018, January 10\u201315). Disentangling by Factorising. Proceedings of the 35th International Conference on Machine Learning, Stockholmsm\u00e4ssan, Stockholm, Sweden."},{"key":"ref_19","unstructured":"Mathieu, E., Rainforth, T., Siddharth, N., and Teh, Y.W. (2019, January 9\u201315). Disentangling Disentanglement in Variational Autoencoders. Proceedings of the 36th International Conference on Machine Learning, Long Beach, CA, USA."},{"key":"ref_20","unstructured":"Kumar, A., Sattigeri, P., and Balakrishnan, A. (May, January 30). Variational Inference of Disentangled Latent Concepts from Unlabeled Observations. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_21","unstructured":"Esmaeili, B., Wu, H., Jain, S., Bozkurt, A., Siddharth, N., Paige, B., Brooks, D.H., Dy, J., and van de Meent, J.W. (2019, January 9\u201315). Structured Disentangled Representations. Proceedings of the Machine Learning Research, Long Beach, CA, USA."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Karras, T., Laine, S., and Aila, T. (2019, January 16\u201320). A Style-Based Generator Architecture for Generative Adversarial Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00453"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Isola, P., Zhu, J.Y., Zhou, T., and Efros, A.A. (2017, January 21\u201326). Image-To-Image Translation With Conditional Adversarial Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.632"},{"key":"ref_24","unstructured":"Zhu, J.Y., Zhang, R., Pathak, D., Darrell, T., Efros, A.A., Wang, O., and Shechtman, E. (2017, January 4\u20139). Toward Multimodal Image-to-Image Translation. Proceedings of the Advances in Neural Information Processing Systems 30, Long Beach, CA, USA."},{"key":"ref_25","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (July, January 26). The Cityscapes Dataset for Semantic Urban Scene Understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_26","unstructured":"Liu, M.Y., Breuel, T., and Kautz, J. (2017, January 4\u20139). Unsupervised Image-to-Image Translation Networks. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Huang, X., Liu, M.Y., Belongie, S., and Kautz, J. (2018, January 8\u201314). Multimodal Unsupervised Image-to-image Translation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01219-9_11"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Choi, Y., Choi, M., Kim, M., Ha, J., Kim, S., and Choo, J. (2018, January 18\u201322). StarGAN: Unified Generative Adversarial Networks for Multi-domain Image-to-Image Translation. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00916"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhu, J.Y., Park, T., Isola, P., and Efros, A.A. (2017, January 22\u201329). Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.244"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Cho, W., Choi, S., Park, D.K., Shin, I., and Choo, J. (2019, January 16\u201320). Image-To-Image Translation via Group-Wise Deep Whitening-And-Coloring Transformation. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01089"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"G\u00fcler, R.A., Neverova, N., and Kokkinos, I. (2018, January 18\u201322). Densepose: Dense human pose estimation in the wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00762"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid Scene Parsing Network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Huang, X., and Belongie, S. (2017, January 22\u201329). Arbitrary Style Transfer in Real-Time with Adaptive Instance Normalization. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.167"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Mao, X., Li, Q., Xie, H., Lau, R.Y., Wang, Z., and Paul Smolley, S. (2017, January 22\u201329). Least Squares Generative Adversarial Networks. Proeedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.304"},{"key":"ref_35","unstructured":"Mercari, I. (2020, August 04). Mercari (Japanese Website). Available online: https:\/\/www.mercari.com\/jp\/."},{"key":"ref_36","unstructured":"Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., Chen, X., and Chen, X. (2016, January 5\u201310). Improved Techniques for Training GANs. Proceedings of the Advances in Neural Information Processing Systems 29, Barcelona, Spain."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zhang, R., Isola, P., Efros, A.A., Shechtman, E., and Wang, O. (2018, January 18\u201322). The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","article-title":"Image quality assessment: From error visibility to structural similarity","volume":"13","author":"Wang","year":"2004","journal-title":"IEEE Trans. Image Process."},{"key":"ref_39","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20138). ImageNet Classification with Deep Convolutional Neural Networks. Proceedings of the Advances in Neural Information Processing Systems 25, Lake Tahoe, CA, USA."},{"key":"ref_40","unstructured":"Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. (2017, January 4\u20139). GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. Proceedings of the Advances in Neural Information Processing Systems 30, Long Beach, CA, USA."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"450","DOI":"10.1016\/0047-259X(82)90077-X","article-title":"The frechet distance between multivariate normal distributions","volume":"12","author":"Dowson","year":"1982","journal-title":"J. Multivar. Anal."},{"key":"ref_42","unstructured":"Kingma, D., and Ba, J. (2015, January 7\u20139). Adam: A Method for Stochastic Optimization. Proceedings of the International Conference on Learning Representation (ICLR), San Diego, CA, USA."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Wang, T., Liu, M., Zhu, J., Tao, A., Kautz, J., and Catanzaro, B. (2018, January 18\u201322). High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs. Proceedings of the 2018 the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00917"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Zhu, X., Xu, C., and Tao, D. (2020, January 23\u201328). Learning Disentangled Representations with Latent Variation Predictability. Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK.","DOI":"10.1007\/978-3-030-58607-2_40"},{"key":"ref_45","unstructured":"Cao, Z., Hidalgo Martinez, G., Simon, T., Wei, S., and Sheikh, Y.A. OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields. IEEE Trans. Pattern Anal. Mach. Intell., 2019."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Cao, Z., Simon, T., Wei, S.E., and Sheikh, Y. (2017, January 21\u201326). Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.143"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Simon, T., Joo, H., Matthews, I., and Sheikh, Y. (2017, January 21\u201326). Hand Keypoint Detection in Single Images using Multiview Bootstrapping. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.494"},{"key":"ref_48","unstructured":"Wei, S.E., Ramakrishna, V., Kanade, T., and Sheikh, Y. (July, January 26). Convolutional pose machines. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Gong, K., Liang, X., Zhang, D., Shen, X., and Lin, L. (2017, January 21\u201326). Look Into Person: Self-Supervised Structure-Sensitive Learning and a New Benchmark for Human Parsing. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.715"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"871","DOI":"10.1109\/TPAMI.2018.2820063","article-title":"Look into Person: Joint Body Parsing & Pose Estimation Network and a New Benchmark","volume":"41","author":"Liang","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/19\/5647\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:16:02Z","timestamp":1760177762000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/19\/5647"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,2]]},"references-count":50,"journal-issue":{"issue":"19","published-online":{"date-parts":[[2020,10]]}},"alternative-id":["s20195647"],"URL":"https:\/\/doi.org\/10.3390\/s20195647","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,10,2]]}}}