{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,28]],"date-time":"2025-11-28T12:36:39Z","timestamp":1764333399152,"version":"build-2065373602"},"reference-count":61,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2024,3,18]],"date-time":"2024-03-18T00:00:00Z","timestamp":1710720000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["62006131","62071260","LQ21F020009","LZ22F020001"],"award-info":[{"award-number":["62006131","62071260","LQ21F020009","LZ22F020001"]}]},{"name":"National Natural Science Foundation of Zhejiang Province","award":["62006131","62071260","LQ21F020009","LZ22F020001"],"award-info":[{"award-number":["62006131","62071260","LQ21F020009","LZ22F020001"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>The diffusion model has made progress in the field of image synthesis, especially in the area of conditional image synthesis. However, this improvement is highly dependent on large annotated datasets. To tackle this challenge, we present the Guided Diffusion model for Unlabeled Images (GDUI) framework in this article. It utilizes the inherent feature similarity and semantic differences in the data, as well as the downstream transferability of Contrastive Language-Image Pretraining (CLIP), to guide the diffusion model in generating high-quality images. We design two semantic-aware algorithms, namely, the pseudo-label-matching algorithm and label-matching refinement algorithm, to match the clustering results with the true semantic information and provide more accurate guidance for the diffusion model. First, GDUI encodes the image into a semantically meaningful latent vector through clustering. Then, pseudo-label matching is used to complete the matching of the true semantic information of the image. Finally, the label-matching refinement algorithm is used to adjust the irrelevant semantic information in the data, thereby improving the quality of the guided diffusion model image generation. Our experiments on labeled datasets show that GDUI outperforms diffusion models without any guidance and significantly reduces the gap between it and models guided by ground-truth labels.<\/jats:p>","DOI":"10.3390\/a17030125","type":"journal-article","created":{"date-parts":[[2024,3,18]],"date-time":"2024-03-18T04:25:15Z","timestamp":1710735915000},"page":"125","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["GDUI: Guided Diffusion Model for Unlabeled Images"],"prefix":"10.3390","volume":"17","author":[{"given":"Xuanyuan","family":"Xie","sequence":"first","affiliation":[{"name":"Mobile Network Application Technology Laboratory, School of Information Science and Engineering, Ningbo University, 818 Fenghua Road, Ningbo 315211, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jieyu","family":"Zhao","sequence":"additional","affiliation":[{"name":"Mobile Network Application Technology Laboratory, School of Information Science and Engineering, Ningbo University, 818 Fenghua Road, Ningbo 315211, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,3,18]]},"reference":[{"key":"ref_1","unstructured":"Po, R., Yifan, W., Golyanik, V., Aberman, K., Barron, J.T., Bermano, A.H., Chan, E.R., Dekel, T., Holynski, A., and Kanazawa, A. (2023). State of the art on diffusion models for visual computing. arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Gu, S., Chen, D., Bao, J., Wen, F., Zhang, B., Chen, D., Yuan, L., and Guo, B. (2022, January 18\u201324). Vector quantized diffusion model for text-to-image synthesis. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01043"},{"key":"ref_3","unstructured":"Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. arXiv."},{"key":"ref_4","unstructured":"Nichol, A.Q., and Dhariwal, P. (2021, January 18\u201324). Improved denoising diffusion probabilistic models. Proceedings of the International Conference on Machine Learning, Virtual Event."},{"key":"ref_5","first-page":"8780","article-title":"Diffusion models beat gans on image synthesis","volume":"34","author":"Dhariwal","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_6","unstructured":"Ho, J., and Salimans, T. (2021, January 14). Classifier-Free Diffusion Guidance. Proceedings of the NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, Virtual Event."},{"key":"ref_7","first-page":"27484","article-title":"Change event dataset for discovery from spatio-temporal remote sensing imagery","volume":"35","author":"Mall","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Shin, H., Kim, H., Kim, S., Jun, Y., Eo, T., and Hwang, D. (2023, January 17\u201324). SDC-UDA: Volumetric Unsupervised Domain Adaptation Framework for Slice-Direction Continuous Cross-Modality Medical Image Segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00716"},{"key":"ref_9","unstructured":"Bordes, F., Balestriero, R., and Vincent, P. (2022). Transactions on Machine Learning Research, OpenReview.net."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Hu, V.T., Zhang, D.W., Asano, Y.M., Burghouts, G.J., and Snoek, C.G. (2023, January 17\u201324). Self-guided diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01766"},{"key":"ref_11","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning, PMLR, Virtual Event."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022, January 18\u201324). High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"ref_13","first-page":"36479","article-title":"Photorealistic text-to-image diffusion models with deep language understanding","volume":"35","author":"Saharia","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wang, Z., Zhang, Z., Zhang, X., Zheng, H., Zhou, M., Zhang, Y., and Wang, Y. (2023, January 17\u201324). DR2: Diffusion-based Robust Degradation Remover for Blind Face Restoration. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00170"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Li, Y., Fan, Y., Xiang, X., Demandolx, D., Ranjan, R., Timofte, R., and Van Gool, L. (2023, January 17\u201324). Efficient and explicit modelling of image hierarchies for image restoration. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01753"},{"key":"ref_16","unstructured":"Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. (2021, January 18\u201324). Zero-shot text-to-image generation. Proceedings of the International Conference on Machine Learning, PMLR, New Orleans, LA, USA."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K. (2023, January 17\u201324). Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"ref_18","unstructured":"Wang, W., Bao, J., Zhou, W., Chen, D., Chen, D., Yuan, L., and Li, H. (2022). Semantic image synthesis via diffusion models. arXiv."},{"key":"ref_19","first-page":"18100","article-title":"Card: Classification and regression diffusion models","volume":"35","author":"Han","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Kim, G., Kwon, T., and Ye, J.C. (2022, January 19\u201324). Diffusionclip: Text-guided diffusion models for robust image manipulation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00246"},{"key":"ref_21","unstructured":"Nichol, A.Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., Mcgrew, B., Sutskever, I., and Chen, M. (2022, January 17\u201323). GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. Proceedings of the International Conference on Machine Learning, PMLR, Baltimore, MD, USA."},{"key":"ref_22","unstructured":"Sheynin, S., Ashual, O., Polyak, A., Singer, U., Gafni, O., Nachmani, E., and Taigman, Y. (2023, January 1\u20135). kNN-Diffusion: Image Generation via Large-Scale Retrieval. Proceedings of the International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_23","first-page":"15309","article-title":"Retrieval-augmented diffusion models","volume":"35","author":"Blattmann","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Zhang, R., Chen, C., Li, C., Tensmeyer, C., Yu, T., Gu, J., Xu, J., and Sun, T. (2022, January 19\u201324). Towards language-free training for text-to-image generation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01738"},{"key":"ref_25","first-page":"3371","article-title":"Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion","volume":"11","author":"Vincent","year":"2010","journal-title":"J. Mach. Learn. Res."},{"key":"ref_26","unstructured":"Ji, P., Zhang, T., Li, H., Salzmann, M., and Reid, I. (2017). Deep subspace clustering networks. Adv. Neural Inf. Process. Syst., 30."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Tian, K., Zhou, S., and Guan, J. (2017, January 18\u201322). Deepcluster: A general clustering framework based on deep learning. Proceedings of the Machine Learning and Knowledge Discovery in Databases, Skopje, Macedonia.","DOI":"10.1007\/978-3-319-71246-8_49"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Jiang, Z., Zheng, Y., Tan, H., Tang, B., and Zhou, H. (2017, January 19\u201325). Variational Deep Embedding: An Unsupervised and Generative Approach to Clustering. Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial Intelligence Organization, Melbourne, Australia.","DOI":"10.24963\/ijcai.2017\/273"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhou, P., Hou, Y., and Feng, J. (2018, January 18\u201323). Deep adversarial subspace clustering. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00172"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhang, J., Li, C.G., You, C., Qi, X., Zhang, H., Guo, J., and Lin, Z. (2019, January 15\u201320). Self-supervised convolutional subspace clustering network. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00562"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"7264","DOI":"10.1109\/TIP.2022.3221290","article-title":"Spice: Semantic pseudo-labeling for image clustering","volume":"31","author":"Niu","year":"2022","journal-title":"IEEE Trans. Image Process."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Han, S., Park, S., Park, S., Kim, S., and Cha, M. (2020, January 23\u201328). Mitigating embedding and class assignment mismatch in unsupervised image classification. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58586-0_45"},{"key":"ref_33","first-page":"21271","article-title":"Bootstrap your own latent-a new approach to self-supervised learning","volume":"33","author":"Grill","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Banterle, F., Marnerides, D., Bashford-Rogers, T., and Debattista, K. (2024). Self-Supervised High Dynamic Range Imaging: What Can Be Learned from a Single 8-bit Video?. ACM Trans. Graph., just accepted.","DOI":"10.1145\/3648570"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., and Gao, J. (2021, January 20\u201325). Vinvl: Revisiting visual representations in vision-language models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00553"},{"key":"ref_36","first-page":"9694","article-title":"Align before fuse: Vision and language representation learning with momentum distillation","volume":"34","author":"Li","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_37","first-page":"7892","article-title":"MACK: Multimodal aligned conceptual knowledge for unpaired image-text matching","volume":"35","author":"Huang","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_38","unstructured":"Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q., Sung, Y.H., Li, Z., and Duerig, T. (2021, January 18\u201324). Scaling up visual and vision-language representation learning with noisy text supervision. Proceedings of the International Conference on Machine Learning, PMLR, Virtual Event."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Ao, T., Zhang, Z., and Liu, L. (2023). GestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents. ACM Trans. Graph., 42.","DOI":"10.1145\/3592097"},{"key":"ref_40","first-page":"6840","article-title":"Denoising diffusion probabilistic models","volume":"33","author":"Ho","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_41","unstructured":"Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. (2015, January 6\u201311). Deep unsupervised learning using nonequilibrium thermodynamics. Proceedings of the International Conference on Machine Learning, PMLR, Lille, France."},{"key":"ref_42","unstructured":"Song, Y., Sohl-Dickstein, J., Kingma, D.P., Kumar, A., Ermon, S., and Poole, B. (2021, January 3\u20137). Score-Based Generative Modeling through Stochastic Differential Equations. Proceedings of the International Conference on Learning Representations, Virtual Event."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Caron, M., Bojanowski, P., Joulin, A., and Douze, M. (2018, January 8\u201314). Deep clustering for unsupervised learning of visual features. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_9"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"7455","DOI":"10.1007\/s10489-022-03916-3","article-title":"A multiscale convolutional gragh network using only structural information for entity alignment","volume":"53","author":"Qi","year":"2023","journal-title":"Appl. Intell."},{"key":"ref_45","unstructured":"Coates, A., Ng, A., and Lee, H. (2011, January 11\u201313). An analysis of single-layer networks in unsupervised feature learning. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics. JMLR Workshop and Conference Proceedings, Fort Lauderdale, FL, USA."},{"key":"ref_46","first-page":"6626","article-title":"Gans trained by a two time-scale update rule converge to a local nash equilibrium","volume":"30","author":"Heusel","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_47","unstructured":"Nash, C., Menick, J., Dieleman, S., and Battaglia, P. (2021, January 18\u201324). Generating images with sparse representations. Proceedings of the International Conference on Machine Learning, PMLR, Virtual Event."},{"key":"ref_48","unstructured":"Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X. (2016). Improved techniques for training gans. Adv. Neural Inf. Process. Syst., 29."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 27\u201330). Rethinking the inception architecture for computer vision. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_51","unstructured":"Kynk\u00e4\u00e4nniemi, T., Karras, T., Laine, S., Lehtinen, J., and Aila, T. (2019). Improved precision and recall metric for assessing generative models. Adv. Neural Inf. Process. Syst., 32."},{"key":"ref_52","unstructured":"Chen, X., Fan, H., Girshick, R., and He, K. (2020). Improved baselines with momentum contrastive learning. arXiv."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Hua, Y., Wang, H., and McLoone, S. (2024, January 4\u20138). Improving the Fairness of the Min-Max Game in GANs Training. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV57701.2024.00289"},{"key":"ref_55","unstructured":"Xiao, Z., Kreis, K., and Vahdat, A. (2022, January 25\u201329). Tackling the Generative Learning Trilemma with Denoising Diffusion GANs. Proceedings of the International Conference on Learning Representations, Virtual Event."},{"key":"ref_56","first-page":"14745","article-title":"Transgan: Two pure transformers can make one strong gan, and that can scale up","volume":"34","author":"Jiang","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_57","first-page":"21039","article-title":"Data-Centric Learning from Unlabeled Graphs with Diffusion Model","volume":"Volume 36","author":"Oh","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_58","first-page":"26699","article-title":"Improving Adversarial Robustness Through the Contrastive-Guided Diffusion Process","volume":"Volume 202","author":"Krause","year":"2023","journal-title":"Proceedings of the 40th International Conference on Machine Learning, PMLR"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Dong, W., Tang, F., Huang, N., Huang, H., Ma, C., Lee, T.Y., Deussen, O., and Xu, C. (2023). ProSpect: Prompt Spectrum for Attribute-Aware Personalization of Diffusion Models. ACM Trans. Graph., 42.","DOI":"10.1145\/3618342"},{"key":"ref_60","doi-asserted-by":"crossref","first-page":"980","DOI":"10.1109\/TMI.2023.3325703","article-title":"Zero-Shot Medical Image Translation via Frequency-Guided Diffusion Models","volume":"43","author":"Li","year":"2024","journal-title":"IEEE Trans. Med. Imaging"},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"2588","DOI":"10.1109\/TNNLS.2022.3190452","article-title":"Learning Better Registration to Learn Better Few-Shot Medical Image Segmentation: Authenticity, Diversity, and Robustness","volume":"35","author":"He","year":"2024","journal-title":"IEEE Trans. Neural Networks Learn. Syst."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/17\/3\/125\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:15:07Z","timestamp":1760105707000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/17\/3\/125"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,18]]},"references-count":61,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2024,3]]}},"alternative-id":["a17030125"],"URL":"https:\/\/doi.org\/10.3390\/a17030125","relation":{},"ISSN":["1999-4893"],"issn-type":[{"type":"electronic","value":"1999-4893"}],"subject":[],"published":{"date-parts":[[2024,3,18]]}}}