{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T14:03:30Z","timestamp":1771682610605,"version":"3.50.1"},"reference-count":0,"publisher":"Slovenian Association Informatika","issue":"7","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IJCAI"],"abstract":"<jats:p>Text-to-image generation has quickly evolved with diffusion-based generative models that combine semantic conditioning and latent-space denoising, allowing machines to generate high-quality visuals from natural language prompts. Despite these developments, existing diffusion systems still face challenges in prompt clarification accuracy, model adaptability, and computational efficacy, which limit their performance in real-time and resource-limited settings. The research aims to design and optimize an image generation framework based on Stable Diffusion (SD) that improves prompt processing, improves image quality, and enables lightweight fine-tuning. The system utilizes the LAION-Aesthetics v2 4.5 dataset, which contains high-quality text\u2013image pairs suitable for visual generation tasks. Preprocessing involves text cleaning, tokenization, and semantic structuring, utilizing a transformer-based tokenizer to ensure accurate language-to-visual mapping. The architecture integrates Stable Diffusion, Variational Autoencoder (VAE) for latent-space decoding, and Low-Rank Adaptation (LoRA) for efficient fine-tuning with minimal computational cost. Results show that SD-VAE-LoRA achieved a PSNR of 33.7 dB, SSIM of 93 %, FID of 17.8, Inception Score of 36.02, and R-Precision of 90 %, superior to baseline SD and advanced diffusion models such as Latent Diffusion Method (LDM) [24], Menstrual Cycle-Inspired Latent Diffusion Method (MCI-LDM) [24], and Conditional Generative Adversarial Networks, Attention mechanisms, and Contrastive Learning (C-GAN+ATT+CL). The optimized system advances semantic alignment, decreases training time, and preserves image realism, confirming its strength for scalable, adaptive, and high-fidelity image generation applications.<\/jats:p>","DOI":"10.31449\/inf.v50i7.10368","type":"journal-article","created":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T13:16:05Z","timestamp":1771679765000},"source":"Crossref","is-referenced-by-count":0,"title":["Stable Diffusion Image Generation System Optimized with Variational Autoencoders and Low-Rank Adaptation"],"prefix":"10.31449","volume":"50","author":[{"given":"ChunLing","family":"Zhang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"16141","published-online":{"date-parts":[[2026,2,21]]},"container-title":["Informatica"],"original-title":[],"link":[{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/download\/10368\/6506","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/download\/10368\/6506","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T13:16:05Z","timestamp":1771679765000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/view\/10368"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,21]]},"references-count":0,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2026,2,21]]}},"URL":"https:\/\/doi.org\/10.31449\/inf.v50i7.10368","relation":{},"ISSN":["1854-3871","0350-5596"],"issn-type":[{"value":"1854-3871","type":"electronic"},{"value":"0350-5596","type":"print"}],"subject":[],"published":{"date-parts":[[2026,2,21]]}}}