{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,11]],"date-time":"2025-09-11T22:32:30Z","timestamp":1757629950494,"version":"3.44.0"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"9","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>\n            Virtual try-on, a significant application in computer vision, aims to seamlessly simulate the appearance of clothing on a person from a single image. We propose a diffusion-based tryon approach, solving virtual tryon as a problem of conditional image inpainting. Our method introduces GarNet and OutlineNet as two learnable Stable Diffusion ControlNet encoders conditioned on the garment and person outline images, enhancing the controllability and realism of the generated try-on. We propose a two-stage garment diffusion recycling training strategy, utilizing\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(x_{0}\\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            -parameterization. We estimate the initial clean image that is conditioned on the maximum noised input and feed the same to the same diffusion model again to estimate total noise. This reduces over-fitting and makes our model more generalized. We also introduce a zero garment-outline conditioning (ZGOC) block along with a Garment-Outline Cross Attention layer to optimize garment draping and ensure global consistency in the try-on results. The ZGOC block provides control and adaptability by prioritizing garment details that are most affected by body shape, ensuring precise garment alignment with the person\u2019s outline. Our comprehensive experiments on the VITON-HD and Dresscode dataset demonstrate that our proposed approach achieves state-of-the-art realism and controllability in VITON, marking a significant advancement in virtual fashion experiences and online shopping applications.\n          <\/jats:p>","DOI":"10.1145\/3758098","type":"journal-article","created":{"date-parts":[[2025,8,6]],"date-time":"2025-08-06T15:16:30Z","timestamp":1754493390000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Garment Recycle Training and Conditional Garment-Person Outline Attention-Guided Virtual Tryon"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-9489-8386","authenticated-orcid":false,"given":"Sanhita","family":"Pathak","sequence":"first","affiliation":[{"name":"BSTTM, IIT Delhi, New Delhi, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6729-4857","authenticated-orcid":false,"given":"Vinay","family":"Kaushik","sequence":"additional","affiliation":[{"name":"Indian Institute of Information Technology Sonepat, Sonepat, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2677-3071","authenticated-orcid":false,"given":"Brejesh","family":"Lall","sequence":"additional","affiliation":[{"name":"IIT Delhi, New Delhi, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,10]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"154","volume-title":"14th European Conference on Computer Vision (ECCV \u201916)","author":"Bai Min","year":"2016","unstructured":"Min Bai, Wenjie Luo, Kaustav Kundu, and Raquel Urtasun. 2016. Exploiting semantic information and deep matching for optical flow. In 14th European Conference on Computer Vision (ECCV \u201916). Springer, 154\u2013170."},{"key":"e_1_3_1_3_2","volume-title":"European Conference on Computer Vision","author":"Bai Shuai","year":"2022","unstructured":"Shuai Bai, Huiling Zhou, Zhikang Li, Chang Zhou, and Hongxia Yang. 2022. Single stage virtual try-on via deformable attention flows. In European Conference on Computer Vision. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:250644446"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00090"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01391"},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","unstructured":"Yisol Choi Sangkyung Kwak Kyungmin Lee Hyungwon Choi and Jinwoo Shin. 2024. Improving diffusion models for authentic virtual try-on in the wild. arXiv:2403.05139. Retrieved from https:\/\/arxiv.org\/abs\/2403.05139","DOI":"10.1007\/978-3-031-73016-0_13"},{"key":"e_1_3_1_7_2","first-page":"14638","volume-title":"IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Cui Aiyu","year":"2021","unstructured":"Aiyu Cui, Daniel McKee, and Svetlana Lazebnik. 2021. Dressing in order: Recurrent person image generation for pose transfer, virtual try-on and outfit editing. In IEEE\/CVF International Conference on Computer Vision (ICCV), 14638\u201314647."},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","unstructured":"Phuong Dam Jihoon Jeong Anh Tran and Daeyoung Kim. 2024. Time-efficient and identity-consistent virtual try-on using a variant of altered diffusion models. arXiv:2403.07371. Retrieved from https:\/\/arxiv.org\/abs\/2403.07371","DOI":"10.1007\/978-3-031-73220-1_3"},{"key":"e_1_3_1_9_2","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"De Luigi Luca","year":"2023","unstructured":"Luca De Luigi, Ren Li, Benoit Guillard, Mathieu Salzmann, and Pascal Fua. 2023. DrapeNet: Garment generation and self-supervised draping. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_1_10_2","unstructured":"Prafulla Dhariwal and Alex Nichol. 2021. Diffusion models beat GANs on image synthesis. arXiv:2105.05233. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:234357997"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","unstructured":"Haoye Dong Xiaodan Liang Bochao Wang Hanjiang Lai Jia Zhu and Jian Yin. 2019. Towards multi-pose guided virtual try-on network. arXiv:1902.11026. Retrieved from https:\/\/arxiv.org\/abs\/1902.11026","DOI":"10.1109\/ICCV.2019.00912"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.59275\/j.melba.2023-fbe4"},{"key":"e_1_3_1_13_2","unstructured":"Rinon Gal Yuval Alaluf Yuval Atzmon Or Patashnik Amit H. Bermano Gal Chechik and Daniel Cohen-Or. 2022. An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv:2208.01618. Retrieved from https:\/\/arxiv.org\/abs\/2208.01618"},{"key":"e_1_3_1_14_2","first-page":"16928","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ge Chongjian","year":"2021","unstructured":"Chongjian Ge, Yibing Song, Yuying Ge, Han Yang, Wei Liu, and Ping Luo. 2021. Disentangled cycle consistency for highly-realistic virtual try-on. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 16928\u201316937."},{"key":"e_1_3_1_15_2","unstructured":"Yuying Ge Yibing Song Ruimao Zhang Chongjian Ge Wei Liu and Ping Luo. 2021. Parser-free virtual try-on via distilling appearance flows. arXiv:2103.04559. Retrieved from https:\/\/arxiv.org\/abs\/2103.04559"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612255"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185531"},{"key":"e_1_3_1_18_2","first-page":"7297","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"G\u00fcler R\u0131za Alp","year":"2018","unstructured":"R\u0131za Alp G\u00fcler, Natalia Neverova, and Iasonas Kokkinos. 2018. Densepose: Dense human pose estimation in the wild. In IEEE Conference on Computer Vision and Pattern Recognition, 7297\u20137306."},{"key":"e_1_3_1_19_2","first-page":"10470","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Han Xintong","year":"2019","unstructured":"Xintong Han, Weilin Huang, Xiaojun Hu, and Matthew R. Scott. 2019. ClothFlow: A flow-based model for clothed person generation. In 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), 10470\u201310479."},{"key":"e_1_3_1_20_2","first-page":"7543","volume-title":"2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Han Xintong","year":"2017","unstructured":"Xintong Han, Zuxuan Wu, Zhe Wu, Ruichi Yu, and Larry S. Davis. 2017. VITON: An image-based virtual try-on network. In 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 7543\u20137552."},{"key":"e_1_3_1_21_2","first-page":"3460","volume-title":"2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"He Sen","year":"2022","unstructured":"Sen He, Yi-Zhe Song, and Tao Xiang. 2022. Style-based global appearance flow for virtual try-on. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3460\u20133469."},{"key":"e_1_3_1_22_2","unstructured":"Martin Heusel Hubert Ramsauer Thomas Unterthiner Bernhard Nessler and Sepp Hochreiter. 2018. GANs trained by a two time-scale update rule converge to a local nash equilibrium. arXiv:1706.08500. Retrieved from https:\/\/arxiv.org\/abs\/1706.08500"},{"key":"e_1_3_1_23_2","unstructured":"Edward J. Hu Yelong Shen Phil Wallis Zeyuan Allen-Zhu Yuanzhi Li Lu Wang and Weizhu Chen. 2021. LoRA: Low-rank adaptation of large language models. arXiv:2106.09685. Retrieved from https:\/\/arxiv.org\/abs\/2106.09685"},{"key":"e_1_3_1_24_2","unstructured":"Thibaut Issenhuth J\u00e9r\u00e9mie Mary and Cl\u00e9ment Calauz\u00e8nes. 2019. End-to-end learning of geometric deformations of feature maps for virtual try-on. arXiv:1906.01347. Retrieved from https:\/\/arxiv.org\/abs\/1906.01347"},{"key":"e_1_3_1_25_2","unstructured":"Jeongho Kim Gyojung Gu Minho Park Sunghyun Park and Jaegul Choo. 2023. StableVITON: Learning semantic correspondence with latent diffusion model for virtual try-on. arxiv:2312.01725. Retrieved from https:\/\/arxiv.org\/abs\/2312.01725"},{"key":"e_1_3_1_26_2","first-page":"21696","article-title":"Variational diffusion models","volume":"34","author":"Kingma Diederik","year":"2021","unstructured":"Diederik Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. 2021. Variational diffusion models. In Advances in Neural Information Processing Systems, Vol. 34, 21696\u201321707.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_27_2","unstructured":"Sangyun Lee Gyojung Gu Sunghyun Park Seunghwan Choi and Jaegul Choo. 2022. High-resolution virtual try-on with misalignment and occlusion-handled conditions. arXiv:2206.14180. Retrieved from https:\/\/arxiv.org\/abs\/2206.14180"},{"key":"e_1_3_1_28_2","first-page":"6571","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu Xiaoning","year":"2024","unstructured":"Xiaoning Liu, Zongwei Wu, Ao Li, Florin-Alexandru Vasluianu, Yulun Zhang, Shuhang Gu, Le Zhang, Ce Zhu, Radu Timofte, Zhi Jin, et al. 2024. NTIRE 2024 challenge on low light image enhancement: Methods and results. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6571\u20136594."},{"key":"e_1_3_1_29_2","volume-title":"International Conference on Learning Representations","author":"Liu Xingchao","year":"2024","unstructured":"Xingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng, and Qiang Liu. 2024. Instaflow: One step is enough for high-quality diffusion-based text-to-image generation. In International Conference on Learning Representations."},{"key":"e_1_3_1_30_2","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Lu Yanzuo","year":"2024","unstructured":"Yanzuo Lu, Manlin Zhang, Andy J. Ma, Xiaohua Xie, and Jian-Huang Lai. 2024. Coarse-to-fine latent diffusion for pose-guided person image synthesis. In Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_1_31_2","unstructured":"Tsiry Mayet Pourya Shamsolmoali Simon Bernard Eric Granger Romain H\u00e9rault and Clement Chatelain. 2024. TD-Paint: Faster diffusion inpainting through time aware pixel conditioning. arXiv:2410.09306. Retrieved from https:\/\/arxiv.org\/abs\/2410.09306"},{"key":"e_1_3_1_32_2","doi-asserted-by":"crossref","unstructured":"Davide Morelli Alberto Baldrati Giuseppe Cartella Marcella Cornia Marco Bertini and Rita Cucchiara. 2023. LaDI-VTON: Latent diffusion textual-inversion enhanced virtual try-on. arXiv:2305.13501. Retrieved from https:\/\/arxiv.org\/abs\/2305.13501","DOI":"10.1145\/3581783.3612137"},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","first-page":"2230","DOI":"10.1109\/CVPRW56347.2022.00243","volume-title":"2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)","author":"Morelli Davide","year":"2022","unstructured":"Davide Morelli, Matteo Fincato, Marcella Cornia, Federico Landi, Fabio Cesari, and Rita Cucchiara. 2022. Dress code: High-resolution multi-category virtual try-on. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2230\u20132234. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:248240016"},{"key":"e_1_3_1_34_2","unstructured":"Maxime Oquab Timoth\u00e9e Darcet Th\u00e9o Moutakanni Huy Vo Marc Szafraniec Vasil Khalidov Pierre Fernandez Daniel Haziza Francisco Massa Alaaeldin El-Nouby et al. 2024. DINOv2: Learning robust visual features without supervision. arXiv:2304.07193. Retrieved from https:\/\/arxiv.org\/abs\/2304.07193"},{"key":"e_1_3_1_35_2","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Patel Chaitanya","year":"2020","unstructured":"Chaitanya Patel, Zhouyingcheng Liao, and Gerard Pons-Moll. 2020. TailorNet: Predicting clothing in 3D as a function of human pose, shape and garment style. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE."},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","unstructured":"Sanhita Pathak Vinay Kaushik and Brejesh Lall. 2023. Single stage warped cloth learning and semantic-contextual attention feature fusion for virtual tryon. arXiv:2310.05024. Retrieved from https:\/\/arxiv.org\/abs\/2310.05024","DOI":"10.1109\/ICME57554.2024.10687502"},{"key":"e_1_3_1_37_2","doi-asserted-by":"crossref","unstructured":"Sanhita Pathak Vinay Kaushik and Brejesh Lall. 2024. GraVITON: Graph based garment warping with attention guided inversion for Virtual-tryon. arXiv:2406.02184. Retrieved from https:\/\/arxiv.org\/abs\/2406.02184","DOI":"10.1007\/978-981-96-2644-1_16"},{"key":"e_1_3_1_38_2","first-page":"1","volume-title":"2024 39th International Conference on Image and Vision Computing New Zealand (IVCNZ)","author":"Pathak Sanhita","year":"2024","unstructured":"Sanhita Pathak, Vinay Kaushik, and Brejesh Lall. 2024. MAC-VTON: Multi-modal attention conditioning for virtual try-on with diffusion-based inpainting. In 2024 39th International Conference on Image and Vision Computing New Zealand (IVCNZ). IEEE, 1\u20136."},{"key":"e_1_3_1_39_2","unstructured":"Dustin Podell Zion English Kyle Lacey Andreas Blattmann Tim Dockhorn Jonas M\u00fcller Joe Penna and Robin Rombach. 2023. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv:2307.01952. Retrieved from https:\/\/arxiv.org\/abs\/2307.01952"},{"key":"e_1_3_1_40_2","unstructured":"Alec Radford Jong Wook Kim Chris Hallacy Aditya Ramesh Gabriel Goh Sandhini Agarwal Girish Sastry Amanda Askell Pamela Mishkin Jack Clark et al. 2021. Learning transferable visual models from natural language supervision. arXiv:2103.00020. Retrieved from https:\/\/arxiv.org\/abs\/2103.00020"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_1_42_2","doi-asserted-by":"crossref","unstructured":"Olaf Ronneberger Philipp Fischer and Thomas Brox. 2015. U-Net: Convolutional networks for biomedical image segmentation. arXiv:1505.04597. Retrieved from https:\/\/arxiv.org\/abs\/1505.04597","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","unstructured":"Nataniel Ruiz Yuanzhen Li Varun Jampani Yael Pritch Michael Rubinstein and Kfir Aberman. 2023. DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation. arXiv:2208.12242. Retrieved from https:\/\/arxiv.org\/abs\/2208.12242","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.13643"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i5.28288"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3300513"},{"key":"e_1_3_1_47_2","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2023. Attention is all you need. arXiv:1706.03762. Retrieved from https:\/\/arxiv.org\/abs\/1706.03762"},{"key":"e_1_3_1_48_2","doi-asserted-by":"crossref","unstructured":"Siqi Wan Yehao Li Jingwen Chen Yingwei Pan Ting Yao Yang Cao and Tao Mei. 2024. Improving virtual try-on with garment-focused diffusion models. arXiv:2409.08258. Retrieved from https:\/\/arxiv.org\/abs\/2409.08258","DOI":"10.1007\/978-3-031-72967-6_11"},{"key":"e_1_3_1_49_2","doi-asserted-by":"crossref","unstructured":"Bochao Wang Huabing Zhang Xiaodan Liang Yimin Chen Liang Lin and Meng Yang. 2018. Toward characteristic-preserving image-based virtual try-on network. arXiv:1807.07688. Retrieved from https:\/\/arxiv.org\/1807.07688","DOI":"10.1007\/978-3-030-01261-8_36"},{"key":"e_1_3_1_50_2","unstructured":"Haoyu Wang Zhilu Zhang Donglin Di Shiliang Zhang and Wangmeng Zuo. 2024. Mv-vton: Multi-view virtual try-on with diffusion models. arXiv:2404.17364. Retrieved from https:\/\/arxiv.org\/abs\/2404.17364"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_1_52_2","article-title":"Towards scalable unpaired virtual try-on via patch-routed spatially-adaptive GAN","author":"Xie Zhenyu","year":"2021","unstructured":"Zhenyu Xie, Zaiyu Huang, Fuwei Zhao, Haoye Dong, Michael C. Kampffmeyer, and Xiaodan Liang. 2021. Towards scalable unpaired virtual try-on via patch-routed spatially-adaptive GAN. In Neural Information Processing Systems.","journal-title":"Neural Information Processing Systems"},{"key":"e_1_3_1_53_2","unstructured":"Yuhao Xu Tao Gu Weifeng Chen and Chengcai Chen. 2024. OOTDiffusion: Outfitting fusion based latent diffusion for controllable virtual try-on. arXiv:2403.01779. Retrieved from https:\/\/arxiv.org\/abs\/2403.01779"},{"key":"e_1_3_1_54_2","unstructured":"Hanshu Yan Xingchao Liu Jiachun Pan Jun Hao Liew Qiang Liu and Jiashi Feng. 2024. PeRFlow: Piecewise rectified flow as universal plug-and-play accelerator. arXiv:2405.07510. Retrieved from https:\/\/arxiv.org\/abs\/2405.07510"},{"key":"e_1_3_1_55_2","unstructured":"Binxin Yang Shuyang Gu Bo Zhang Ting Zhang Xuejin Chen Xiaoyan Sun Dong Chen and Fang Wen. 2022. Paint by example: Exemplar-based image editing with diffusion models. arXiv:2211.13227. Retrieved from https:\/\/arxiv.org\/abs\/2211.13227"},{"key":"e_1_3_1_56_2","volume-title":"IEEE CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Yang Han","year":"2020","unstructured":"Han Yang, Ruimao Zhang, Xiaobao Guo, Wei Liu, Wangmeng Zuo, and Ping Luo. 2020. Towards photo-realistic virtual try-on by adaptively generating-preserving image content. In IEEE CVF Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_1_57_2","first-page":"36","volume-title":"European Conference on Computer Vision","author":"Yang Zhaotong","year":"2025","unstructured":"Zhaotong Yang, Zicheng Jiang, Xinzhe Li, Huiyu Zhou, Junyu Dong, Huaidong Zhang, and Yong Du. 2025. \\(D^{4}\\) -VTON: Dynamic semantics disentangling for differential diffusion based virtual Try-On. In European Conference on Computer Vision. Springer, 36\u201352."},{"key":"e_1_3_1_58_2","unstructured":"Zilong Yang Yujiu Yang Xiaotang Lin and Xiaoguang Wang. 2023. DiffKD: Knowledge distillation for diffusion models. arXiv:2305.15712. Retrieved from https:\/\/arxiv.org\/abs\/2305.15712"},{"key":"e_1_3_1_59_2","first-page":"10510","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Yu Ruiyun","year":"2019","unstructured":"Ruiyun Yu, Xiaoqi Wang, and Xiaohui Xie. 2019. VTNFP: An image-based virtual try-on network with body and clothing feature preservation. In 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), 10510\u201310519."},{"key":"e_1_3_1_60_2","unstructured":"Jianhao Zeng Dan Song Weizhi Nie Hongshuo Tian Tongtong Wang and Anan Liu. 2023. CAT-DM: Controllable accelerated virtual try-on with diffusion model. arXiv:2311.18405. Retrieved from https:\/\/arxiv.org\/abs\/2311.18405"},{"key":"e_1_3_1_61_2","doi-asserted-by":"crossref","unstructured":"Lintao Zhang Xiangcheng Du LeoWu TomyEnrique Yiqun Wang Yingbin Zheng and Cheng Jin. 2024. Minutes to Seconds: Speeded-up DDPM-based image inpainting with coarse-to-fine sampling. arXiv:2407.05875. Retrieved from https:\/\/arxiv.org\/abs\/2407.05875","DOI":"10.1109\/ICME57554.2024.10687818"},{"key":"e_1_3_1_62_2","doi-asserted-by":"crossref","unstructured":"Richard Zhang Phillip Isola Alexei A. Efros Eli Shechtman and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. arXiv:1801.03924. Retrieved from https:\/\/arxiv.org\/abs\/1801.03924","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_3_1_63_2","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhenyu Xie","year":"2023","unstructured":"Xie Zhenyu, Huang Zaiyu, Dong Xin, Zhao Fuwei, Dong Haoye, Zhang Xijin, Zhu Feida, and Liang Xiaodan. 2023. GP-VTON: Towards general purpose virtual try-on via collaborative local-flow global-parsing learning. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_1_64_2","doi-asserted-by":"crossref","unstructured":"Luyang Zhu Dawei Yang Tyler Zhu Fitsum Reda William Chan Chitwan Saharia Mohammad Norouzi and Ira Kemelmacher-Shlizerman. 2023. TryOnDiffusion: A tale of two UNets. arXiv:2306.08276. Retrieved from https:\/\/arxiv.org\/abs\/2306.08276","DOI":"10.1109\/CVPR52729.2023.00447"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3758098","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,10]],"date-time":"2025-09-10T16:01:33Z","timestamp":1757520093000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3758098"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,10]]},"references-count":63,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3758098"],"URL":"https:\/\/doi.org\/10.1145\/3758098","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2025,9,10]]},"assertion":[{"value":"2024-11-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-25","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}