{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T14:55:52Z","timestamp":1761404152931,"version":"build-2065373602"},"publisher-location":"New York, NY, USA","reference-count":31,"publisher":"ACM","funder":[{"name":"National Natural Science Foundation of China","award":["No. 62206283"],"award-info":[{"award-number":["No. 62206283"]}]},{"name":"National Natural Science Foundation of China","award":["No. 62306315"],"award-info":[{"award-number":["No. 62306315"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,10,27]]},"DOI":"10.1145\/3746262.3761977","type":"proceedings-article","created":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T14:52:23Z","timestamp":1761403943000},"page":"42-50","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["BlingDiff: High-Fidelity Virtual Jewelry Try-On with Detail-Optimized Diffusion"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-1848-9962","authenticated-orcid":false,"given":"Yunfang","family":"Niu","sequence":"first","affiliation":[{"name":"Foundation Model Research Center, Institute of Automation, CAS, Beijing, China and School of Artificial Intelligence, UCAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9346-3597","authenticated-orcid":false,"given":"Lingxiang","family":"Wu","sequence":"additional","affiliation":[{"name":"Foundation Model Research Center, Institute of Automation, CAS, Beijing, China and Wuhan AI Research, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-1282-7310","authenticated-orcid":false,"given":"Dong","family":"Yi","sequence":"additional","affiliation":[{"name":"Foundation Model Research Center, Institute of Automation, CAS, Beijing, China and Wuhan AI Research, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8531-9138","authenticated-orcid":false,"given":"Lu","family":"Zhou","sequence":"additional","affiliation":[{"name":"Foundation Model Research Center, Institute of Automation, CAS, Beijing, China and Wuhan AI Research, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9118-2780","authenticated-orcid":false,"given":"Jinqiao","family":"Wang","sequence":"additional","affiliation":[{"name":"Foundation Model Research Center, Institute of Automation, CAS, Beijing, China and Wuhan AI Research, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,10,26]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Proceedings of European Conference on Computer Vision (ECCV). 1-17","author":"Ahn Donghoon","year":"2024","unstructured":"Donghoon Ahn, Hyoungwon Cho, Jaewon Min, Wooseok Jang, Jungwoo Kim, SeonHwa Kim, Hyun Hee Park, Kyong Hwan Jin, and Seungryong Kim. 2024. Self-rectifying diffusion sampling with perturbed-attention guidance. In Proceedings of European Conference on Computer Vision (ECCV). 1-17."},{"key":"e_1_3_2_1_2_1","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 22669-22679","author":"Bao Fan","year":"2023","unstructured":"Fan Bao, Shen Nie, Kaiwen Xue, Yue Cao, Chongxuan Li, Hang Su, and Jun Zhu. 2023. All are worth words: A vit backbone for diffusion models. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 22669-22679."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2929257"},{"key":"e_1_3_2_1_4_1","volume-title":"Proceedings of International Conference on Learning Representations (ICLR).","author":"Chen Junsong","year":"2024","unstructured":"Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al., 2024. PixArt-\u03b1: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis. In Proceedings of International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_1_5_1","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 14131-14140","author":"Choi Seunghwan","year":"2021","unstructured":"Seunghwan Choi, Sunghyun Park, Minsoo Lee, and Jaegul Choo. 2021. Viton-hd: High-resolution virtual try-on via misalignment-aware normalization. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 14131-14140."},{"key":"e_1_3_2_1_6_1","first-page":"2567","article-title":"Image quality assessment: Unifying structure and texture similarity","volume":"44","author":"Ding Keyan","year":"2020","unstructured":"Keyan Ding, Kede Ma, Shiqi Wang, and Eero P Simoncelli. 2020. Image quality assessment: Unifying structure and texture similarity. IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), Vol. 44, 5 (2020), 2567-2581.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"crossref","first-page":"721","DOI":"10.1016\/S0031-3203(00)00023-6","article-title":"On the Canny edge detector","volume":"34","author":"Ding Lijun","year":"2001","unstructured":"Lijun Ding and Ardeshir Goshtasby. 2001. On the Canny edge detector. Pattern Recognition (PR), Vol. 34, 3 (2001), 721-725.","journal-title":"Pattern Recognition (PR)"},{"key":"e_1_3_2_1_8_1","volume-title":"Proceedings of International Conference on Machine Learning (ICML).","author":"Esser Patrick","year":"2024","unstructured":"Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M\u00fcller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al., 2024. Scaling rectified flow transformers for high-resolution image synthesis. In Proceedings of International Conference on Machine Learning (ICML)."},{"key":"e_1_3_2_1_9_1","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 12873-12883","author":"Esser Patrick","year":"2021","unstructured":"Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for high-resolution image synthesis. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 12873-12883."},{"key":"e_1_3_2_1_10_1","volume-title":"Proceedings of ACM International Conference on Multimedia (ACMMM). 7599-7607","author":"Gou Junhong","year":"2023","unstructured":"Junhong Gou, Siyu Sun, Jianfu Zhang, Jianlou Si, Chen Qian, and Liqing Zhang. 2023. Taming the power of diffusion models for high-quality virtual try-on with appearance flow. In Proceedings of ACM International Conference on Multimedia (ACMMM). 7599-7607."},{"key":"e_1_3_2_1_11_1","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 7297-7306","author":"G\u00fcler Riza Alp","year":"2018","unstructured":"Riza Alp G\u00fcler, Natalia Neverova, and Iasonas Kokkinos. 2018. Densepose: Dense human pose estimation in the wild. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 7297-7306."},{"key":"e_1_3_2_1_12_1","volume-title":"Proceedings of International Conference on Computer Vision (ICCV). 10471-10480","author":"Han Xintong","year":"2019","unstructured":"Xintong Han, Xiaojun Hu, Weilin Huang, and Matthew R Scott. 2019. Clothflow: A flow-based model for clothed person generation. In Proceedings of International Conference on Computer Vision (ICCV). 10471-10480."},{"key":"e_1_3_2_1_13_1","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 7543-7552","author":"Han Xintong","year":"2018","unstructured":"Xintong Han, Zuxuan Wu, Zhe Wu, Ruichi Yu, and Larry S Davis. 2018. Viton: An image-based virtual try-on network. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 7543-7552."},{"key":"e_1_3_2_1_14_1","volume-title":"PM-Jewelry: Personalized Multimodal Adaptation for Virtual Jewelry Try-On with Latent Diffusion. In NeurIPS workshops.","author":"He Yangfan","year":"2024","unstructured":"Yangfan He, Yinghui Xia, Jinfeng Wei, Tianyu Shi, and Yang Jingsong. 2024. PM-Jewelry: Personalized Multimodal Adaptation for Virtual Jewelry Try-On with Latent Diffusion. In NeurIPS workshops."},{"key":"e_1_3_2_1_15_1","volume-title":"FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on. arXiv preprint arXiv:2411.10499","author":"Jiang Boyuan","year":"2024","unstructured":"Boyuan Jiang, Xiaobin Hu, Donghao Luo, Qingdong He, Chengming Xu, Jinlong Peng, Jiangning Zhang, Chengjie Wang, Yunsheng Wu, and Yanwei Fu. 2024. FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on. arXiv preprint arXiv:2411.10499 (2024)."},{"key":"e_1_3_2_1_16_1","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 8176-8185","author":"Kim Jeongho","year":"2024","unstructured":"Jeongho Kim, Guojung Gu, Minho Park, Sunghyun Park, and Jaegul Choo. 2024. Stableviton: Learning semantic correspondence with latent diffusion model for virtual try-on. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 8176-8185."},{"key":"e_1_3_2_1_17_1","first-page":"3260","article-title":"Self-correction for human parsing","volume":"44","author":"Li Peike","year":"2020","unstructured":"Peike Li, Yunqiu Xu, Yunchao Wei, and Yi Yang. 2020. Self-correction for human parsing. IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), Vol. 44, 6 (2020), 3260-3271.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"crossref","first-page":"e839","DOI":"10.55908\/sdgs.v11i5.839","article-title":"A review of current cultural jewellery trend","volume":"11","author":"Mei Li","year":"2023","unstructured":"Li Mei and Nooraziah Binti Ahmad. 2023. A review of current cultural jewellery trend. Journal of Law and Sustainable Development, Vol. 11, 5 (2023), e839-e839.","journal-title":"Journal of Law and Sustainable Development"},{"key":"e_1_3_2_1_19_1","volume-title":"Proceedings of ACM International Conference on Multimedia (ACMMM). 8580-8589","author":"Morelli Davide","year":"2023","unstructured":"Davide Morelli, Alberto Baldrati, Giuseppe Cartella, Marcella Cornia, Marco Bertini, and Rita Cucchiara. 2023. Ladi-vton: Latent diffusion textual-inversion enhanced virtual try-on. In Proceedings of ACM International Conference on Multimedia (ACMMM). 8580-8589."},{"key":"e_1_3_2_1_20_1","volume-title":"Proceedings of International Conference on Acoustics, Speech, and Signal Processing (ICASSP). 3780-3784","author":"Niu Yunfang","year":"2024","unstructured":"Yunfang Niu, Dong Yi, Lingxiang Wu, Zhiwei Liu, Pengxiang Cai, and Jinqiao Wang. 2024. PFDM: Parser-Free Virtual Try-On via Diffusion Model. In Proceedings of International Conference on Acoustics, Speech, and Signal Processing (ICASSP). 3780-3784."},{"key":"e_1_3_2_1_21_1","volume-title":"Enhancing the Virtual Jewelry Try-On Experience with Computer Vision. In International IEEE Applied Sensing Conference (APSCON). 1-4.","author":"Patel Hrutika","year":"2024","unstructured":"Hrutika Patel, Jap Purohit, and Sanket Patel. 2024. Enhancing the Virtual Jewelry Try-On Experience with Computer Vision. In International IEEE Applied Sensing Conference (APSCON). 1-4."},{"key":"e_1_3_2_1_22_1","volume-title":"Proceedings of International Conference on Computer Vision (ICCV). 4195-4205","author":"Peebles William","year":"2023","unstructured":"William Peebles and Saining Xie. 2023. Scalable diffusion models with transformers. In Proceedings of International Conference on Computer Vision (ICCV). 4195-4205."},{"key":"e_1_3_2_1_23_1","volume-title":"Proceedings of International Conference on Learning Representations (ICLR).","author":"Podell Dustin","year":"2024","unstructured":"Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M\u00fcller, Joe Penna, and Robin Rombach. 2024. Sdxl: Improving latent diffusion models for high-resolution image synthesis. In Proceedings of International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_1_24_1","unstructured":"Nikhila Ravi Valentin Gabeur Yuan-Ting Hu Ronghang Hu Chaitanya Ryali Tengyu Ma Haitham Khedr Roman R\u00e4dle Chloe Rolland Laura Gustafson et al. 2024. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714 (2024)."},{"key":"e_1_3_2_1_25_1","volume-title":"Proceedings of International Conference on Learning Representations (ICLR)","author":"Salimans Tim","year":"2022","unstructured":"Tim Salimans and Jonathan Ho. 2022. Progressive distillation for fast sampling of diffusion models. Proceedings of International Conference on Learning Representations (ICLR) (2022)."},{"key":"e_1_3_2_1_26_1","volume-title":"The market value of brands on an international scale: considering the example of jewelry industry. Sciences of Europe, 39-3 (39)","author":"Serdiuk AV","year":"2019","unstructured":"AV Serdiuk, ES Votchenko, and IV Bogdashev. 2019. The market value of brands on an international scale: considering the example of jewelry industry. Sciences of Europe, 39-3 (39) (2019), 8-12."},{"key":"e_1_3_2_1_27_1","volume-title":"Proceedings of International Conference on Computer Vision (ICCV). 2149-2159","author":"Suvorov Roman","year":"2022","unstructured":"Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. 2022. Resolution-robust large mask inpainting with fourier convolutions. In Proceedings of International Conference on Computer Vision (ICCV). 2149-2159."},{"key":"e_1_3_2_1_28_1","volume-title":"Proceedings of European Conference on Computer Vision (ECCV). 184-199","author":"Wan Siqi","year":"2024","unstructured":"Siqi Wan, Yehao Li, Jingwen Chen, Yingwei Pan, Ting Yao, Yang Cao, and Tao Mei. 2024. Improving Virtual Try-On with Garment-Focused Diffusion Models. In Proceedings of European Conference on Computer Vision (ECCV). 184-199."},{"key":"e_1_3_2_1_29_1","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 4818-4829","author":"Xiao Bin","year":"2024","unstructured":"Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, and Lu Yuan. 2024. Florence-2: Advancing a unified representation for a variety of vision tasks. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 4818-4829."},{"key":"e_1_3_2_1_30_1","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 23550-23559","author":"Xie Zhenyu","year":"2023","unstructured":"Zhenyu Xie, Zaiyu Huang, Xin Dong, Fuwei Zhao, Haoye Dong, Xijin Zhang, Feida Zhu, and Xiaodan Liang. 2023. Gp-vton: Towards general purpose virtual try-on via collaborative local-flow global-parsing learning. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 23550-23559."},{"key":"e_1_3_2_1_31_1","volume-title":"Ootdiffusion: Outfitting fusion based latent diffusion for controllable virtual try-on. arXiv preprint arXiv:2403.01779","author":"Xu Yuhao","year":"2024","unstructured":"Yuhao Xu, Tao Gu, Weifeng Chen, and Chengcai Chen. 2024. Ootdiffusion: Outfitting fusion based latent diffusion for controllable virtual try-on. arXiv preprint arXiv:2403.01779 (2024)."}],"event":{"name":"MM '25:The 33rd ACM International Conference on Multimedia","location":"Dublin Ireland","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 3rd International Workshop on Rich Media With Generative AI"],"original-title":[],"deposited":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T14:52:55Z","timestamp":1761403975000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746262.3761977"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,26]]},"references-count":31,"alternative-id":["10.1145\/3746262.3761977","10.1145\/3746262"],"URL":"https:\/\/doi.org\/10.1145\/3746262.3761977","relation":{},"subject":[],"published":{"date-parts":[[2025,10,26]]},"assertion":[{"value":"2025-10-26","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}