{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,9]],"date-time":"2026-05-09T00:15:10Z","timestamp":1778285710474,"version":"3.51.4"},"reference-count":46,"publisher":"Institution of Engineering and Technology (IET)","issue":"1","license":[{"start":{"date-parts":[[2025,6,26]],"date-time":"2025-06-26T00:00:00Z","timestamp":1750896000000},"content-version":"vor","delay-in-days":176,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"},{"start":{"date-parts":[[2025,1,1]],"date-time":"2025-01-01T00:00:00Z","timestamp":1735689600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100004479","name":"Natural Science Foundation of Jiangxi Province","doi-asserted-by":"publisher","award":["20224BAB202018"],"award-info":[{"award-number":["20224BAB202018"]}],"id":[{"id":"10.13039\/501100004479","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62462032"],"award-info":[{"award-number":["62462032"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100013064","name":"Key Research and Development Program of Jiangxi Province","doi-asserted-by":"publisher","award":["20223BBE51039"],"award-info":[{"award-number":["20223BBE51039"]}],"id":[{"id":"10.13039\/501100013064","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100013064","name":"Key Research and Development Program of Jiangxi Province","doi-asserted-by":"publisher","award":["20232BBE50020"],"award-info":[{"award-number":["20232BBE50020"]}],"id":[{"id":"10.13039\/501100013064","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100021183","name":"Science Fund for Distinguished Young Scholars of Jiangxi Province","doi-asserted-by":"publisher","award":["20232ACB212007"],"award-info":[{"award-number":["20232ACB212007"]}],"id":[{"id":"10.13039\/501100021183","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["ietresearch.onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["IET Image Processing"],"published-print":{"date-parts":[[2025,1]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>Transforming fashion design sketches into realistic garments remains a challenging task due to the reliance on labor\u2010intensive manual workflows that limit efficiency and scalability in traditional fashion pipelines. While recent advances in image generation and virtual try\u2010on technologies have introduced partial automation, existing methods still lack controllability and struggle to maintain semantic consistency in garment pose and structure, restricting their applicability in real\u2010world design scenarios. In this work, we present CG\u2010VTON, a controllable virtual try\u2010on framework designed to generate high\u2010quality try\u2010on images directly from clothing design sketches. The model integrates multi\u2010modal conditional inputs, including dense human pose maps and textual garment descriptions, to guide the generation process. A novel pose constraint module is introduced to enhance garment\u2010body alignment, while a structured diffusion\u2010based pipeline performs progressive generation through latent denoising and global\u2010context refinement. Extensive experiments conducted on benchmark datasets demonstrate that CG\u2010VTON significantly outperforms existing state\u2010of\u2010the\u2010art methods in terms of visual quality, pose consistency, and computational efficiency. By enabling high\u2010fidelity and controllable try\u2010on results from abstract sketches, CG\u2010VTON offers a practical and robust solution for bridging the gap between conceptual design and realistic garment\u00a0visualization.<\/jats:p>","DOI":"10.1049\/ipr2.70144","type":"journal-article","created":{"date-parts":[[2025,6,27]],"date-time":"2025-06-27T01:40:52Z","timestamp":1750988452000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["CG\u2010VTON: Controllable Generation of Virtual Try\u2010On Images Based on Multimodal Conditions"],"prefix":"10.1049","volume":"19","author":[{"given":"Haopeng","family":"Lei","sequence":"first","affiliation":[{"name":"School of Artificial Intelligence Jiangxi Normal University  NanChang Jiangxi China"},{"name":"Jiangxi Provincial Key Laboratory of Intelligent Information Processing and Affective Computing Jiangxi Normal University  NanChang Jiangxi China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6098-9757","authenticated-orcid":false,"given":"Xuan","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence Jiangxi Normal University  NanChang Jiangxi China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yaqin","family":"Liang","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence Jiangxi Normal University  NanChang Jiangxi China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuanlong","family":"Cao","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence Jiangxi Normal University  NanChang Jiangxi China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"265","published-online":{"date-parts":[[2025,6,26]]},"reference":[{"key":"e_1_2_10_2_1","doi-asserted-by":"crossref","unstructured":"S.Choi S.Park M.Lee et\u00a0al. \u201cViton\u2010Hd: High\u2010Resolution Virtual Try\u2010on Via Misalignment\u2010Aware Normalization \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2021) 14131\u201314140.","DOI":"10.1109\/CVPR46437.2021.01391"},{"key":"e_1_2_10_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3422622"},{"key":"e_1_2_10_4_1","doi-asserted-by":"crossref","unstructured":"X.Han Z.Wu Z.Wu et\u00a0al. \u201cViton: An Image\u2010Based Virtual Try\u2010on Network \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 7543\u20137552.","DOI":"10.1109\/CVPR.2018.00787"},{"key":"e_1_2_10_5_1","unstructured":"P.DhariwalandA.Nichol \u201cDiffusion Models Beat Gans on Image Synthesis \u201d inAdvances in Neural Information Processing Systems34(Curran Associates Inc. 2021) 8780\u20108794."},{"key":"e_1_2_10_6_1","doi-asserted-by":"crossref","unstructured":"R.Rombach A.Blattmann D.Lorenz et\u00a0al. \u201cHigh\u2010Resolution Image Synthesis With Latent Diffusion Models \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2022) 10684\u201310695.","DOI":"10.1109\/CVPR52688.2022.01042"},{"issue":"5","key":"e_1_2_10_7_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1618452.1618470","article-title":"Sketch2Photo: Internet Image Montage","volume":"28","author":"Chen T.","year":"2009","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_10_8_1","doi-asserted-by":"crossref","unstructured":"W.ChenandJ.Hays \u201cSketchygan: Towards Diverse and Realistic Sketch to Image Synthesis \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 9416\u20139425.","DOI":"10.1109\/CVPR.2018.00981"},{"key":"e_1_2_10_9_1","doi-asserted-by":"crossref","unstructured":"C.Gao Q.Liu Q.Xu et\u00a0al. \u201cSketchycoco: Image Generation From Freehand Scenesketches \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2020) 5174\u20135183.","DOI":"10.1109\/CVPR42600.2020.00522"},{"key":"e_1_2_10_10_1","doi-asserted-by":"crossref","unstructured":"J.Choi S.Kim Y.Jeong et\u00a0al. \u201cILVR: Conditioning Method for Denoising Diffusion Probabilistic Models \u201d in2021 IEEE\/CVF International Conference on Computer Vision (ICCV)(IEEE 2021).","DOI":"10.1109\/ICCV48922.2021.01410"},{"key":"e_1_2_10_11_1","unstructured":"C.Meng Y.Song J.Song et\u00a0al. \u201cSDEdit: Image Synthesis and Editing With Stochastic Differential Equations \u201d inInternational Conference on Learning Representations(IEEE Information Theory Society 2021)."},{"key":"e_1_2_10_12_1","doi-asserted-by":"crossref","unstructured":"K.Pnvr B.Singh P.Ghosh et\u00a0al. \u201cLd\u2010Znet: A Latent Diffusion Approach for Text\u2010Based Image Segmentation \u201d inProceedings of the IEEE\/CVF International Conference on Computer Vision(IEEE 2023) 4157\u20134168.","DOI":"10.1109\/ICCV51070.2023.00384"},{"key":"e_1_2_10_13_1","doi-asserted-by":"crossref","unstructured":"G.Parmar R.Zhang andJ. Y.Zhu \u201cOn Aliased Resizing and Surprising Subtleties in Gan Evaluation \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2022) 11410\u201311420.","DOI":"10.1109\/CVPR52688.2022.01112"},{"key":"e_1_2_10_14_1","doi-asserted-by":"crossref","unstructured":"X.Han Z.Wu Z.Wu et\u00a0al. \u201cViton: An Image\u2010Based Virtual Try\u2010on Network \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 7543\u20137552.","DOI":"10.1109\/CVPR.2018.00787"},{"key":"e_1_2_10_15_1","doi-asserted-by":"crossref","unstructured":"L.Zhu D.Yang T.Zhu et\u00a0al. \u201cTryondiffusion: A Tale of Two Unets \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2023) 4606\u20134615.","DOI":"10.1109\/CVPR52729.2023.00447"},{"key":"e_1_2_10_16_1","doi-asserted-by":"crossref","unstructured":"P.Isola J. Y.Zhu T.Zhou et\u00a0al. \u201cImage\u2010to\u2010Image Translation With Conditional Adversarial Networks \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2017) 1125\u20131134.","DOI":"10.1109\/CVPR.2017.632"},{"key":"e_1_2_10_17_1","unstructured":"J.Ho A.Jain andP.Abbeel \u201cDenoising Diffusion Probabilistic Models \u201d inAdvances in Neural Information Processing Systems Vol.33(Curran Associates Inc. 2020) 6840\u20136851."},{"key":"e_1_2_10_18_1","unstructured":"J.Song C.Meng andS.Ermon \u201cDenoising Diffusion Implicit Models \u201darXiv preprint arXiv:2010.02502(2020)."},{"key":"e_1_2_10_19_1","doi-asserted-by":"crossref","unstructured":"R. A.G\u00fcler N.Neverova andI.Kokkinos \u201cDensepose: Dense Human Pose Estimation in the Wild \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 7297\u20137306.","DOI":"10.1109\/CVPR.2018.00762"},{"key":"e_1_2_10_20_1","doi-asserted-by":"crossref","unstructured":"R.Suvorov E.Logacheva A.Mashikhin et\u00a0al. \u201cResolution\u2010Robust Large Mask Inpainting With Fourier Convolutions \u201d inProceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision(IEEE 2022) 2149\u20132159.","DOI":"10.1109\/WACV51458.2022.00323"},{"key":"e_1_2_10_21_1","doi-asserted-by":"crossref","unstructured":"Z.Cao T.Simon S. E.Wei et\u00a0al. \u201cRealtime Multi\u2010Person 2D Pose Estimation Using Part Affinity Fields \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2017) 7291\u20137299.","DOI":"10.1109\/CVPR.2017.143"},{"key":"e_1_2_10_22_1","doi-asserted-by":"crossref","unstructured":"Z.Su W.Liu Z.Yu et\u00a0al. \u201cPixel Difference Networks for Efficient Edge Detection \u201d inProceedings of the IEEE\/CVF International Conference on Computer Vision(IEEE 2021) 5117\u20135127.","DOI":"10.1109\/ICCV48922.2021.00507"},{"key":"e_1_2_10_23_1","unstructured":"J.Yu Z.Wang V.Vasudevan et\u00a0al. \u201cCoca: Contrastive Captioners are Image\u2010Text Foundation Models \u201darXiv preprint arXiv:2205.01917(2022)."},{"key":"e_1_2_10_24_1","doi-asserted-by":"crossref","unstructured":"R.Zhang P.Isola A. A.Efros et\u00a0al. \u201cThe Unreasonable Effectiveness of Deep Features as a Perceptual Metric \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 586\u2013595.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_2_10_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_2_10_26_1","unstructured":"M.Bikowski D. J.Sutherland M.Arbel et\u00a0al. \u201cDemystifying Mmd Gans \u201darXiv preprint arXiv:1801.01401(2018)."},{"key":"e_1_2_10_27_1","unstructured":"M.Heusel H.Ramsauer T.Unterthiner et\u00a0al. \u201cGans Trained by a Two Time\u2010Scale Update Rule Converge to a Local Nash Equilibrium \u201d inAdvances in Neural Information Processing Systems30(Curran Associates Inc. 2017)."},{"key":"e_1_2_10_28_1","unstructured":"D. P.KingmaandM.Welling \u201cAuto\u2010Encoding Variational Bayes \u201darXiv.1312.6114(2022)."},{"key":"e_1_2_10_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.993558"},{"key":"e_1_2_10_30_1","doi-asserted-by":"crossref","unstructured":"N.JetchevandU.Bergmann \u201cThe Conditional Analogy Gan: Swapping Fashion Articles on People Images \u201d inProceedings of the IEEE International Conference on Computer Vision Workshops(IEEE 2017) 2287\u20132292.","DOI":"10.1109\/ICCVW.2017.269"},{"key":"e_1_2_10_31_1","unstructured":"L.Ma X.Jia Q.Sun et\u00a0al. \u201cPose Guided Person Image Generation \u201d inAdvances in Neural Information Processing Systems Vol.30(Curran Associates Inc. 2017)."},{"key":"e_1_2_10_32_1","unstructured":"A.Radford J. W.Kim C.Hallacy et\u00a0al. \u201cLearning Transferable Visual Models From Natural Language Supervision \u201d inInternational Conference on Machine Learning(PMLR 2021) 8748\u20138763."},{"key":"e_1_2_10_33_1","unstructured":"A.Dosovitskiy L.Beyer A.Kolesnikov et\u00a0al. \u201cAn Image is Worth 16x16 Words: Transformers for Image Recognition at Scale \u201darXiv preprint arXiv:2010.11929(2020)."},{"key":"e_1_2_10_34_1","doi-asserted-by":"crossref","unstructured":"K.He X.Zhang S.Ren et\u00a0al. \u201cDeep Residual Learning for Image Recognition \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2016) 770\u2013778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_10_35_1","doi-asserted-by":"crossref","unstructured":"R.Zhang P.Isola A. A.Efros et\u00a0al. \u201cThe Unreasonable Effectiveness of Deep Features as a Perceptual Metric \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 586\u2013595.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_2_10_36_1","unstructured":"Z.KongandW.Ping \u201cOn Fast Sampling of Diffusion Probabilistic Models \u201darXiv preprint arXiv:2106.00132(2021)."},{"key":"e_1_2_10_37_1","doi-asserted-by":"crossref","unstructured":"L.Zhang A.Rao andM.Agrawala \u201cAdding Conditional Control to Text\u2010to\u2010Image Diffusion Models \u201d inProceedings of the IEEE\/CVF International Conference on Computer Vision(IEEE 2023) 3836\u20133847.","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"e_1_2_10_38_1","doi-asserted-by":"crossref","unstructured":"A.Baldrati D.Morelli G.Cartella et\u00a0al. \u201cMultimodal Garment Designer: Human\u2010Centric Latent Diffusion Models for Fashion Image Editing \u201d inProceedings of the IEEE\/CVF International Conference on Computer Vision(IEEE 2023) 23393\u201323402.","DOI":"10.1109\/ICCV51070.2023.02138"},{"key":"e_1_2_10_39_1","doi-asserted-by":"crossref","unstructured":"Y.Xu T.Gu W.Chen et\u00a0al. \u201cOotdiffusion: Outfitting Fusion Based Latent Diffusion for Controllable Virtual Try\u2010On \u201d inProceedings of the AAAI Conference on Artificial Intelligence(AAAI Press 2025) 8996\u20139004.","DOI":"10.1609\/aaai.v39i9.32973"},{"key":"e_1_2_10_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459884"},{"key":"e_1_2_10_41_1","doi-asserted-by":"crossref","unstructured":"N.Zheng X.Song Z.Chen et\u00a0al. \u201cVirtually Trying on New Clothing With Arbitrary Poses \u201d inMM '19: Proceedings of the 27th ACM International Conference on Multimedia(ACM 2019) https:\/\/doi.org\/10.1145\/3343031.3350946.","DOI":"10.1145\/3343031.3350946"},{"key":"e_1_2_10_42_1","doi-asserted-by":"crossref","unstructured":"L.Zhu Y.Li N.Liu et\u00a0al. \u201cM&M VTO: Multi\u2010Garment Virtual Try\u2010On and Editing \u201d inIEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(IEEE 2024) https:\/\/doi.org\/10.1109\/CVPR52733.2024.00134.","DOI":"10.1109\/CVPR52733.2024.00134"},{"key":"e_1_2_10_43_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.07.092"},{"key":"e_1_2_10_44_1","doi-asserted-by":"crossref","unstructured":"M.Gadelha S.Maji andR.Wang \u201c3D Shape Induction From 2D Views of Multiple Objects \u201d in2017 International Conference on 3D Vision (3DV)(IEEE 2017) 402\u2013411.","DOI":"10.1109\/3DV.2017.00053"},{"key":"e_1_2_10_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447648"},{"key":"e_1_2_10_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528223.3530104"},{"key":"e_1_2_10_47_1","unstructured":"A.Radford J. W.Kim C.Hallacy et\u00a0al. \u201cLearning Transferable Visual Models From Natural Language Supervision \u201d inInternational Conference on Machine Learning(PMLR 2021) 8748\u20138763."}],"container-title":["IET Image Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/ipr2.70144","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/full-xml\/10.1049\/ipr2.70144","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/ipr2.70144","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T23:54:51Z","timestamp":1778284491000},"score":1,"resource":{"primary":{"URL":"https:\/\/ietresearch.onlinelibrary.wiley.com\/doi\/10.1049\/ipr2.70144"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1]]},"references-count":46,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1]]}},"alternative-id":["10.1049\/ipr2.70144"],"URL":"https:\/\/doi.org\/10.1049\/ipr2.70144","archive":["Portico"],"relation":{},"ISSN":["1751-9659","1751-9667"],"issn-type":[{"value":"1751-9659","type":"print"},{"value":"1751-9667","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1]]},"assertion":[{"value":"2025-02-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-26","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70144"}}