{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T03:32:50Z","timestamp":1779334370006,"version":"3.51.4"},"reference-count":36,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2023,10,28]],"date-time":"2023-10-28T00:00:00Z","timestamp":1698451200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>With the recent advancements in the field of diffusion generative models, it has been shown that defining the generative process in the latent space of a powerful pretrained autoencoder can offer substantial advantages. This approach, by abstracting away imperceptible image details and introducing substantial spatial compression, renders the learning of the generative process more manageable while significantly reducing computational and memory demands. In this work, we propose to replace autoencoder coding with a model-based coding scheme based on traditional lossy image compression techniques; this choice not only further diminishes computational expenses but also allows us to probe the boundaries of latent-space image generation. Our objectives culminate in the proposal of a valuable approximation for training continuous diffusion models within a discrete space, accompanied by enhancements to the generative model for categorical values. Beyond the good results obtained for the problem at hand, we believe that the proposed work holds promise for enhancing the adaptability of generative diffusion models across diverse data types beyond the realm of imagery.<\/jats:p>","DOI":"10.3390\/a16110501","type":"journal-article","created":{"date-parts":[[2023,10,30]],"date-time":"2023-10-30T09:27:28Z","timestamp":1698658048000},"page":"501","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Denoising Diffusion Models on Model-Based Latent Space"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1006-7826","authenticated-orcid":false,"given":"Carmelo","family":"Scribano","sequence":"first","affiliation":[{"name":"Department of Physics, Informatics and Mathematics, University of Modena and Reggio Emilia, 41125 Modena, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Danilo","family":"Pezzi","sequence":"additional","affiliation":[{"name":"Department of Physics, Informatics and Mathematics, University of Modena and Reggio Emilia, 41125 Modena, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9082-8087","authenticated-orcid":false,"given":"Giorgia","family":"Franchini","sequence":"additional","affiliation":[{"name":"Department of Physics, Informatics and Mathematics, University of Modena and Reggio Emilia, 41125 Modena, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7327-3347","authenticated-orcid":false,"given":"Marco","family":"Prato","sequence":"additional","affiliation":[{"name":"Department of Physics, Informatics and Mathematics, University of Modena and Reggio Emilia, 41125 Modena, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,10,28]]},"reference":[{"key":"ref_1","first-page":"6840","article-title":"Denoising diffusion probabilistic models","volume":"33","author":"Ho","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_2","first-page":"8780","article-title":"Diffusion models beat gans on image synthesis","volume":"34","author":"Dhariwal","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Esser, P., Rombach, R., and Ommer, B. (2021, January 20\u201325). Taming Transformers for High-Resolution Image Synthesis. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01268"},{"key":"ref_4","unstructured":"Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. (2021, January 18\u201324). Zero-shot text-to-image generation. Proceedings of the International Conference on Machine Learning. PMLR, Online."},{"key":"ref_5","first-page":"19822","article-title":"CogView: Mastering Text-to-Image Generation via Transformers","volume":"Volume 34","author":"Ranzato","year":"2021","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022, January 18\u201324). High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"ref_7","first-page":"2234","article-title":"Improved techniques for training gans","volume":"29","author":"Salimans","year":"2016","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_9","first-page":"6626","article-title":"Gans trained by a two time-scale update rule converge to a local nash equilibrium","volume":"30","author":"Heusel","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Parmar, G., Zhang, R., and Zhu, J.Y. (2022, January 18\u201324). On Aliased Resizing and Surprising Subtleties in GAN Evaluation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01112"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1956","DOI":"10.1007\/s11263-020-01316-z","article-title":"The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale","volume":"128","author":"Kuznetsova","year":"2020","journal-title":"Int. J. Comput. Vis."},{"key":"ref_12","first-page":"6306","article-title":"Neural discrete representation learning","volume":"30","author":"Vinyals","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"180","DOI":"10.1109\/TCOM.1986.1096503","article-title":"Image Compression Using Adaptive Vector Quantization","volume":"34","author":"Goldberg","year":"1986","journal-title":"IEEE Trans. Commun."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"957","DOI":"10.1109\/26.3776","article-title":"Image coding using vector quantization: A review","volume":"36","author":"Nasrabadi","year":"1988","journal-title":"IEEE Trans. Commun."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1007\/s10915-023-02125-5","article-title":"DCT-Former: Efficient Self-Attention with Discrete Cosine Transform","volume":"94","author":"Scribano","year":"2023","journal-title":"J. Sci. Comput."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Garg, I., Chowdhury, S.S., and Roy, K. (2021, January 11\u201317). DCT-SNN: Using DCT To Distribute Spatial Information Over Time for Low-Latency Spiking Neural Networks. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00463"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1109\/MASSP.1984.1162229","article-title":"Vector quantization","volume":"1","author":"Gray","year":"1984","journal-title":"IEEE Assp Mag."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"2325","DOI":"10.1109\/18.720541","article-title":"Quantization","volume":"44","author":"Gray","year":"1998","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1109\/TPAMI.2010.57","article-title":"Product quantization for nearest neighbor search","volume":"33","author":"Jegou","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"744","DOI":"10.1109\/TPAMI.2013.240","article-title":"Optimized Product Quantization","volume":"36","author":"Ge","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/BF02289451","article-title":"A generalized solution of the orthogonal procrustes problem","volume":"31","year":"1966","journal-title":"Psychometrika"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Babenko, A., and Lempitsky, V. (2014, January 23\u201328). Additive quantization for extreme vector compression. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.124"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"11259","DOI":"10.3390\/s101211259","article-title":"Approximate Nearest Neighbor Search by Residual Vector Quantization","volume":"10","author":"Chen","year":"2010","journal-title":"Sensors"},{"key":"ref_24","first-page":"17981","article-title":"Structured denoising diffusion models in discrete state-spaces","volume":"34","author":"Austin","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_25","first-page":"12454","article-title":"Argmax flows and multinomial diffusion: Learning categorical distributions","volume":"34","author":"Hoogeboom","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_26","first-page":"28266","article-title":"A continuous time framework for discrete denoising models","volume":"35","author":"Campbell","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_27","unstructured":"Vahdat, A., Kreis, K., and Kautz, J. (2021). Score-based generative modeling in latent space. arXiv."},{"key":"ref_28","unstructured":"Tang, Z., Gu, S., Bao, J., Chen, D., and Wen, F. (2022). Improved vector quantized diffusion models. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"20150202","DOI":"10.1098\/rsta.2015.0202","article-title":"Principal component analysis: A review and recent developments","volume":"374","author":"Jolliffe","year":"2016","journal-title":"Philos. Trans. R. Soc. A Math. Phys. Eng. Sci."},{"key":"ref_30","unstructured":"Jang, E., Gu, S., and Poole, B. (2016). Categorical reparameterization with gumbel-softmax. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"023016","DOI":"10.1117\/1.3600632","article-title":"Color demosaicking by local directional interpolation and nonlocal adaptive thresholding","volume":"20","author":"Zhang","year":"2011","journal-title":"J. Electron. Imaging"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"535","DOI":"10.1109\/TBDATA.2019.2921572","article-title":"Billion-scale similarity search with GPUs","volume":"7","author":"Johnson","year":"2019","journal-title":"IEEE Trans. Big Data"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1109\/TIT.1962.1057702","article-title":"Picture coding using pseudo-random noise","volume":"8","author":"Roberts","year":"1962","journal-title":"IRE Trans. Inf. Theory"},{"key":"ref_34","unstructured":"Yu, F., Seff, A., Zhang, Y., Song, S., Funkhouser, T., and Xiao, J. (2015). Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Hang, T., Gu, S., Li, C., Bao, J., Chen, D., Hu, H., Geng, X., and Guo, B. (2023). Efficient diffusion training via min-snr weighting strategy. arXiv.","DOI":"10.1109\/ICCV51070.2023.00684"},{"key":"ref_36","unstructured":"Chen, T., Zhang, R., and Hinton, G. (2022). Analog bits: Generating discrete data using diffusion models with self-conditioning. arXiv."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/16\/11\/501\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:13:27Z","timestamp":1760130807000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/16\/11\/501"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,28]]},"references-count":36,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2023,11]]}},"alternative-id":["a16110501"],"URL":"https:\/\/doi.org\/10.3390\/a16110501","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,28]]}}}