{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T16:31:05Z","timestamp":1778689865747,"version":"3.51.4"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2024,1,11]],"date-time":"2024-01-11T00:00:00Z","timestamp":1704931200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Graduate Student Innovation Fund of Donghua University","award":["GSIF-DH-M-2021006"],"award-info":[{"award-number":["GSIF-DH-M-2021006"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61971121, 62072151"],"award-info":[{"award-number":["61971121, 62072151"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Anhui Provincial Natural Science Fund for the Distinguished Young Scholars","award":["2008085J30"],"award-info":[{"award-number":["2008085J30"]}]},{"name":"Open Foundation of Yunnan Key Laboratory of Software Engineering","award":["2023SE103"],"award-info":[{"award-number":["2023SE103"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,5,31]]},"abstract":"<jats:p>Deep learning based virtual try-on system has achieved some encouraging progress recently, but there still remain several big challenges that need to be solved, such as trying on arbitrary clothes of all types, trying on the clothes from one category to another and generating image-realistic results with few artifacts. To handle this issue, we in this article first collect a new dataset with all types of clothes, i.e., tops, bottoms, and whole clothes, each one has multiple categories with rich information of clothing characteristics such as patterns, logos, and other details. Based on this dataset, we then propose the Arbitrary Virtual Try-On Network (AVTON) that is utilized for all-type clothes, which can synthesize realistic try-on images by preserving and trading off characteristics of the target clothes and the reference person. Our approach includes three modules: (1) Limbs Prediction Module, which is utilized for predicting the human body parts by preserving the characteristics of the reference person. This is especially good for handling cross-category try-on task (e.g., long sleeves \u2194 short sleeves or long pants \u2194 skirts), where the exposed arms or legs with the skin colors and details can be reasonably predicted; (2) Improved Geometric Matching Module, which is designed to warp clothes according to the geometry of the target person. We improve the TPS based warping method with a compactly supported radial function (Wendland\u2019s \u03a8-function); (3) Trade-Off Fusion Module, which is to tradeoff the characteristics of the warped clothes and the reference person. This module is to make the generated try-on images look more natural and realistic based on a fine-tune symmetry of the network structure. Extensive simulations are conducted and our approach can achieve better performance compared with the state-of-the-art virtual try-on methods.<\/jats:p>","DOI":"10.1145\/3636426","type":"journal-article","created":{"date-parts":[[2023,12,9]],"date-time":"2023-12-09T12:10:46Z","timestamp":1702123846000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Arbitrary Virtual Try-on Network: Characteristics Preservation and Tradeoff between Body and Clothing"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1304-0872","authenticated-orcid":false,"given":"Yu","family":"Liu","sequence":"first","affiliation":[{"name":"Donghua University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0381-4360","authenticated-orcid":false,"given":"Mingbo","family":"Zhao","sequence":"additional","affiliation":[{"name":"Donghua University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5703-7969","authenticated-orcid":false,"given":"Zhao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-3203-7360","authenticated-orcid":false,"given":"Yuping","family":"Liu","sequence":"additional","affiliation":[{"name":"Fudan University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8906-3777","authenticated-orcid":false,"given":"Shuicheng","family":"Yan","sequence":"additional","affiliation":[{"name":"The National University of Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,1,11]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3519029"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3425636"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3326332"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3478642"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1287\/mksc.19.1.4.15178"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2013.67"},{"key":"e_1_3_2_8_2","volume-title":"Proceedings of the International Conference on Computer Vision.","author":"Zhu Shizhan","year":"2017","unstructured":"Shizhan Zhu, Sanja Fidler, Raquel Urtasun, Dahua Lin, and Chen Change Loy. 2017. Be your own prada: Fashion synthesis with structural coherence. In Proceedings of the International Conference on Computer Vision."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.124"},{"key":"e_1_3_2_10_2","first-page":"5337","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Ge Yuying","year":"2019","unstructured":"Yuying Ge, Ruimao Zhang, Xiaogang Wang, Xiaoou Tang, and Ping Luo. 2019. Deepfashion2: A versatile benchmark for detection, pose estimation, segmentation and re-identification of clothing images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5337\u20135345."},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Peike Li Yunqiu Xu Yunchao Wei and Yi Yang. 2020. Self-correction for human parsing. IEEE Transactions on Pattern Analysis and Machine Intelligence 44 6 (2020) 3260\u20133271.","DOI":"10.1109\/TPAMI.2020.3048039"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2015.2408360"},{"key":"e_1_3_2_13_2","unstructured":"Jianshu Li Jian Zhao Yunchao Wei Congyan Lang Yidong Li Terence Sim Shuicheng Yan and Jiashi Feng. 2017. Multiple-human parsing in the wild. arXiv:1705.07206. Retrieved from https:\/\/arxiv.org\/abs\/1705.07206"},{"key":"e_1_3_2_14_2","first-page":"932","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Gong Ke","year":"2017","unstructured":"Ke Gong, Xiaodan Liang, Dongyu Zhang, Xiaohui Shen, and Liang Lin. 2017. Look into person: Self-supervised structure-sensitive learning and a new benchmark for human parsing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 932\u2013940."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01261-8_36"},{"key":"e_1_3_2_16_2","first-page":"2262","volume-title":"Proceedings of the 22nd International Joint Conference on Artificial Intelligence","author":"Iwata Tomoharu","year":"2011","unstructured":"Tomoharu Iwata, Shinji Wanatabe, and Hiroshi Sawada. 2011. Fashion coordinates recommender system using photographs from fashion magazines. In Proceedings of the 22nd International Joint Conference on Artificial Intelligence. Citeseer, 2262."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/MMUL.2020.3024221"},{"key":"e_1_3_2_18_2","first-page":"4642","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Veit Andreas","year":"2015","unstructured":"Andreas Veit, Balazs Kovacs, Sean Bell, Julian McAuley, Kavita Bala, and Serge Belongie. 2015. Learning visual clothing style with heterogeneous dyadic co-occurrences. In Proceedings of the IEEE International Conference on Computer Vision. 4642\u20134650."},{"key":"e_1_3_2_19_2","first-page":"937","volume-title":"Proceedings of the 2016 IEEE 16th International Conference on Data Mining","author":"He Ruining","year":"2016","unstructured":"Ruining He, Charles Packer, and Julian McAuley. 2016. Learning compatibility across categories for heterogeneous item recommendation. In Proceedings of the 2016 IEEE 16th International Conference on Data Mining. IEEE, 937\u2013942."},{"key":"e_1_3_2_20_2","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence","author":"Shih Yong-Siang","year":"2018","unstructured":"Yong-Siang Shih, Kai-Yueh Chang, Hsuan-Tien Lin, and Min Sun. 2018. Compatibility family learning for item recommendation and generation. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2690144"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123394"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313444"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"Wei Zeng Mingbo Zhao Yuan Gao and Zhao Zhang. 2020. TileGAN: category-oriented attention-based high-quality tiled clothes generation from dressed person. Neural Computing and Applications 32 (2020) 17587\u201317600.","DOI":"10.1007\/s00521-020-04928-1"},{"key":"e_1_3_2_25_2","first-page":"0","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Liu Jingyuan","year":"2018","unstructured":"Jingyuan Liu and Hong Lu. 2018. Deep fashion analysis with feature map upsampling and landmark-driven attention. In Proceedings of the European Conference on Computer Vision. 0\u20130."},{"key":"e_1_3_2_26_2","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Liu Ziwei","year":"2016","unstructured":"Ziwei Liu, Sijie Yan, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2016. Fashion landmark detection in the wild. In Proceedings of the European Conference on Computer Vision."},{"key":"e_1_3_2_27_2","doi-asserted-by":"crossref","unstructured":"Hui Wu Yupeng Gao Xiaoxiao Guo Ziad Al-Halah Steven Rennie Kristen Grauman and Rogerio Feris. 2021. Fashion iq: A new dataset towards retrieving images by natural language feedback. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 11307\u201311317.","DOI":"10.1109\/CVPR46437.2021.01115"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00337"},{"key":"e_1_3_2_29_2","first-page":"3596","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Hosseinzadeh Mehrdad","year":"2020","unstructured":"Mehrdad Hosseinzadeh and Yang Wang. 2020. Composed query image retrieval using locally bounded features. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3596\u20133605."},{"key":"e_1_3_2_30_2","unstructured":"Surgan Jandial Ayush Chopra Pinkesh Badjatiya Pranit Chawla Mausoom Sarkar and Balaji Krishnamurthy. 2020. TRACE: Transform aggregate and compose visiolinguistic representations for image search with text feedback. arXiv:2009.01485. Retrieved from https:\/\/arxiv.org\/abs\/2009.01485"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00515"},{"key":"e_1_3_2_32_2","unstructured":"Haoye Dong Xiaodan Liang Yixuan Zhang Xujie Zhang Zhenyu Xie Bowen Wu Ziqi Zhang Xiaohui Shen and Jian Yin. 2019. Fashion editing with multi-scale attention normalization. arXiv:1906.00884. Retrieved from https:\/\/arxiv.org\/abs\/1906.00884"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3422622"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00787"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.01061"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00787"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.07.092"},{"key":"e_1_3_2_38_2","first-page":"5184","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Neuberger Assaf","year":"2020","unstructured":"Assaf Neuberger, Eran Borenstein, Bar Hilleli, Eduard Oks, and Sharon Alpert. 2020. Image based virtual try-on network from unpaired data. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 5184\u20135193."},{"key":"e_1_3_2_39_2","first-page":"679","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Raj Amit","year":"2018","unstructured":"Amit Raj, Patsorn Sangkloy, Huiwen Chang, James Hays, Duygu Ceylan, and Jingwan Lu. 2018. Swapnet: Image based garment transfer. In Proceedings of the European Conference on Computer Vision. Springer, 679\u2013695."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073711"},{"key":"e_1_3_2_41_2","unstructured":"Sangwoo Mo Minsu Cho and Jinwoo Shin. 2018. Instagan: Instance-aware image-to-image translation. arXiv:1812.10889. Retrieved from https:\/\/arxiv.org\/abs\/1812.10889"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185531"},{"key":"e_1_3_2_43_2","unstructured":"Xintong Han Zuxuan Wu Weilin Huang Matthew R. Scott and Larry S. Davis. 2019. Finet: Compatible and diverse fashion image inpainting. In Proceedings of the IEEE\/CVF International Conference on Computer Vision . 4481\u20134491."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF02123482"},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1007\/BFb0086566","volume-title":"Proceedings of the Constructive Theory of Functions of Several Variables","author":"Duchon Jean","year":"1977","unstructured":"Jean Duchon. 1977. Splines minimizing rotation-invariant semi-norms in sobolev spaces. In Proceedings of the Constructive Theory of Functions of Several Variables. Springer, 85\u2013100."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.632"},{"key":"e_1_3_2_48_2","unstructured":"Tero Karras Timo Aila Samuli Laine and Jaakko Lehtinen. 2017. Progressive growing of gans for improved quality stability and variation. arXiv:1710.10196. Retrieved from https:\/\/arxiv.org\/abs\/1710.10196"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00453"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00244"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.244"},{"key":"e_1_3_2_52_2","first-page":"8929","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Wenguan","year":"2020","unstructured":"Wenguan Wang, Hailong Zhu, Jifeng Dai, Yanwei Pang, Jianbing Shen, and Ling Shao. 2020. Hierarchical human parsing with typed part-relation reasoning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 8929\u20138939."},{"key":"e_1_3_2_53_2","first-page":"5703","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Wang Wenguan","year":"2019","unstructured":"Wenguan Wang, Zhijie Zhang, Siyuan Qi, Jianbing Shen, Yanwei Pang, and Ling Shao. 2019. Learning compositional neural information fusion for human parsing. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 5703\u20135713."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00449"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2020.12.001"},{"key":"e_1_3_2_56_2","doi-asserted-by":"crossref","unstructured":"Hui Wu Yupeng Gao Xiaoxiao Guo Ziad Al-Halah Steven Rennie Kristen Grauman and Rogerio Feris. 2021. Fashion iq: A new dataset towards retrieving images by natural language feedback. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 11307\u201311317.","DOI":"10.1109\/CVPR46437.2021.01115"},{"key":"e_1_3_2_57_2","first-page":"8387","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Alldieck Thiemo","year":"2018","unstructured":"Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. 2018. Video based reconstruction of 3d people models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 8387\u20138397."},{"key":"e_1_3_2_58_2","first-page":"7023","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Mir Aymen","year":"2020","unstructured":"Aymen Mir, Thiemo Alldieck, and Gerard Pons-Moll. 2020. Learning to transfer texture from clothing images to 3d humans. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 7023\u20137034."},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3143712"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00762"},{"key":"e_1_3_2_61_2","doi-asserted-by":"crossref","first-page":"402","DOI":"10.1109\/CVPR.1999.786970","volume-title":"Proceedings. 1999 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (Cat. No PR00149)","author":"Fornefett Mike","year":"1999","unstructured":"Mike Fornefett, Karl Rohr, and H. Siegfried Stiehl. 1999. Elastic registration of medical images using radial basis functions with compact support. In Proceedings. 1999 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (Cat. No PR00149). IEEE, 402\u2013407."},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0262-8856(00)00057-3"},{"key":"e_1_3_2_63_2","first-page":"793","article-title":"Bulletin de l\u2019Acad\u00e9mie des Sciences de l\u2019URSS","volume":"6","author":"Delaunay B","year":"1934","unstructured":"B Delaunay, S Vide, A Lam\u00e9moire, and V De Georges. 1934. Bulletin de l\u2019Acad\u00e9mie des Sciences de l\u2019URSS. Classe Des Sciences Math\u00e9matiques et Naturelles 6 (1934), 793\u2013800.","journal-title":"Classe Des Sciences Math\u00e9matiques et Naturelles"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00917"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_3_2_67_2","first-page":"2234","article-title":"Improved techniques for training gans","volume":"29","author":"Salimans Tim","year":"2016","unstructured":"Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016. Improved techniques for training gans. Advances in Neural Information Processing Systems 29 (2016), 2234\u20132242.","journal-title":"Advances in Neural Information Processing Systems"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3636426","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3636426","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:35:41Z","timestamp":1750178141000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3636426"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,11]]},"references-count":66,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,5,31]]}},"alternative-id":["10.1145\/3636426"],"URL":"https:\/\/doi.org\/10.1145\/3636426","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,11]]},"assertion":[{"value":"2022-03-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-24","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-11","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}