{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T11:46:33Z","timestamp":1784547993280,"version":"3.55.0"},"reference-count":48,"publisher":"Springer Science and Business Media LLC","issue":"36","license":[{"start":{"date-parts":[[2024,10,5]],"date-time":"2024-10-05T00:00:00Z","timestamp":1728086400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,10,5]],"date-time":"2024-10-05T00:00:00Z","timestamp":1728086400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100018777","name":"Nile University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100018777","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2024,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The generation and enhancement of satellite imagery are critical in remote sensing, requiring high-quality, detailed images for accurate analysis. This research introduces a two-stage diffusion model methodology for synthesizing high-resolution satellite images from textual prompts. The pipeline comprises a low-resolution diffusion model (LRDM) that generates initial images based on text inputs and a super-resolution diffusion model (SRDM) that refines these images into high-resolution outputs. The LRDM merges text and image embeddings within a shared latent space, capturing essential scene content and structure. The SRDM then enhances these images, focusing on spatial features and visual clarity. Experiments conducted using the Remote Sensing Image Captioning Dataset demonstrate that our method outperforms existing models, producing satellite images with accurate geographical details and improved spatial resolution.<\/jats:p>","DOI":"10.1007\/s00521-024-10363-3","type":"journal-article","created":{"date-parts":[[2024,10,5]],"date-time":"2024-10-05T02:01:50Z","timestamp":1728093710000},"page":"23103-23111","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":50,"title":["RSDiff: remote sensing image generation from text using diffusion model"],"prefix":"10.1007","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3700-7324","authenticated-orcid":false,"given":"Ahmad","family":"Sebaq","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohamed","family":"ElHelw","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,10,5]]},"reference":[{"issue":"1","key":"10363_CR1","doi-asserted-by":"publisher","first-page":"8","DOI":"10.1109\/MGRS.2016.2616418","volume":"5","author":"P Ghamisi","year":"2017","unstructured":"Ghamisi P, Plaza J, Chen Y, Li J, Plaza AJ (2017) Advanced spectral classifiers for hyperspectral images: a review. IEEE Geosci Remote Sens Mag 5(1):8\u201332","journal-title":"IEEE Geosci Remote Sens Mag"},{"key":"10363_CR2","first-page":"1","volume":"60","author":"Y Xu","year":"2022","unstructured":"Xu Y, Ghamisi P (2022) Universal adversarial examples in remote sensing: methodology and benchmark. IEEE Trans Geosci Remote Sens 60:1\u201315","journal-title":"IEEE Trans Geosci Remote Sens"},{"issue":"2","key":"10363_CR3","doi-asserted-by":"publisher","first-page":"270","DOI":"10.1109\/MGRS.2022.3145854","volume":"10","author":"L Zhang","year":"2022","unstructured":"Zhang L, Zhang L (2022) Artificial intelligence for remote sensing data analysis: a review of challenges and opportunities. IEEE Geosci Remote Sens Mag 10(2):270\u2013294","journal-title":"IEEE Geosci Remote Sens Mag"},{"key":"10363_CR4","unstructured":"Sermanet P, Chintala S, LeCun Y (2012) Convolutional neural networks applied to house numbers digit classification. In: Proceedings of the 21st international conference on pattern recognition (ICPR2012). IEEE, pp 3288\u20133291"},{"key":"10363_CR5","unstructured":"Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y(2014) Generative adversarial nets. Adv Neural Inf Process Syst 27:2672\u20132680"},{"issue":"10","key":"10363_CR6","doi-asserted-by":"publisher","first-page":"1894","DOI":"10.3390\/rs13101894","volume":"13","author":"C Chen","year":"2021","unstructured":"Chen C, Ma H, Yao G, Lv N, Yang H, Li C, Wan S (2021) Remote sensing image augmentation based on text description for waterside change detection. Remote Sens 13(10):1894. https:\/\/doi.org\/10.3390\/rs13101894","journal-title":"Remote Sens"},{"issue":"3","key":"10363_CR7","doi-asserted-by":"publisher","first-page":"950","DOI":"10.1109\/JSTARS.2019.2895693","volume":"12","author":"MB Bejiga","year":"2019","unstructured":"Bejiga MB, Melgani F, Vascotto A (2019) Retro-remote sensing: generating images from ancient texts. IEEE J Sel Top Appl Earth Obse Remote Sens 12(3):950\u2013960","journal-title":"IEEE J Sel Top Appl Earth Obse Remote Sens"},{"key":"10363_CR8","first-page":"1","volume":"19","author":"R Zhao","year":"2021","unstructured":"Zhao R, Shi Z (2021) Text-to-remote-sensing-image generation with structured generative adversarial networks. IEEE Geosci Remote Sens Lett 19:1\u20135","journal-title":"IEEE Geosci Remote Sens Lett"},{"key":"10363_CR9","first-page":"6840","volume":"33","author":"J Ho","year":"2020","unstructured":"Ho J, Jain A, Abbeel P (2020) Denoising diffusion probabilistic models. Adv Neural Inf Process Syst 33:6840\u20136851","journal-title":"Adv Neural Inf Process Syst"},{"issue":"1","key":"10363_CR10","first-page":"2249","volume":"23","author":"J Ho","year":"2022","unstructured":"Ho J, Saharia C, Chan W, Fleet DJ, Norouzi M, Salimans T (2022) Cascaded diffusion models for high fidelity image generation. J Mach Learn Res 23(1):2249\u20132281","journal-title":"J Mach Learn Res"},{"key":"10363_CR11","unstructured":"Reed SE, Akata Z, Mohan S, Tenka S, Schiele B, Lee H (2016) Learning what and where to draw. Adv Neural Inf Process Syst 29:217\u2013225"},{"key":"10363_CR12","doi-asserted-by":"crossref","unstructured":"Zhang H, Xu T, Li H, Zhang S, Wang X, Huang X, Metaxas DN (2017) Stackgan: text to photo-realistic image synthesis with stacked generative adversarial networks. In: Proceedings of the IEEE international conference on computer vision, pp 5907\u20135915","DOI":"10.1109\/ICCV.2017.629"},{"key":"10363_CR13","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A (2020) Language models are few-shot learners. Adv Neural Inf Process Syst 33:1877\u20131901","journal-title":"Adv Neural Inf Process Syst"},{"key":"10363_CR14","unstructured":"Ramesh A, Pavlov M, Goh G, Gray S, Voss C, Radford A, Chen M, Sutskever I (2021) Zero-shot text-to-image generation. In: International conference on machine learning. PMLR, pp 8821\u20138831"},{"key":"10363_CR15","unstructured":"Ramesh A, Dhariwal P, Nichol A, Chu C, Chen M(2022) Hierarchical text-conditional image generation with clip latents. arXiv preprint 1(2):3. arXiv:2204.06125"},{"key":"10363_CR16","unstructured":"Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J (2021) Learning transferable visual models from natural language supervision. In: International conference on machine learning. PMLR, pp 8748\u20138763"},{"issue":"1","key":"10363_CR17","doi-asserted-by":"publisher","first-page":"72","DOI":"10.1038\/s41597-024-02918-9","volume":"11","author":"Z Chen","year":"2024","unstructured":"Chen Z, Yang J, Feng Z, Zhu H (2024) Railfod23: a dataset for foreign object detection on railroad transmission lines. Sci Data 11(1):72","journal-title":"Sci Data"},{"issue":"11","key":"10363_CR18","doi-asserted-by":"publisher","first-page":"10751","DOI":"10.1109\/TII.2023.3241682","volume":"19","author":"L Yang","year":"2023","unstructured":"Yang L, Li X, Sun M, Sun C (2023) Hybrid policy-based reinforcement learning of adaptive energy management for the energy transmission-constrained island group. IEEE Trans Industr Inf 19(11):10751\u201310762. https:\/\/doi.org\/10.1109\/TII.2023.3241682","journal-title":"IEEE Trans Industr Inf"},{"issue":"12","key":"10363_CR19","doi-asserted-by":"publisher","first-page":"3065","DOI":"10.1109\/TFUZZ.2020.2967282","volume":"28","author":"Y Cui","year":"2020","unstructured":"Cui Y, Wu D, Huang J (2020) Optimize tsk fuzzy systems for classification problems: minibatch gradient descent with uniform regularization and batch normalization. IEEE Trans Fuzzy Syst 28(12):3065\u20133075. https:\/\/doi.org\/10.1109\/TFUZZ.2020.2967282","journal-title":"IEEE Trans Fuzzy Syst"},{"key":"10363_CR20","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2024.3390595","author":"N Zhang","year":"2024","unstructured":"Zhang N, Yan J, Hu C, Sun Q, Yang L, Gao DW, Guerrero JM, Li Y (2024) Price-matching-based regional energy market with hierarchical reinforcement learning algorithm. IEEE Trans Ind Inform. https:\/\/doi.org\/10.1109\/TII.2024.3390595","journal-title":"IEEE Trans Ind Inform"},{"issue":"4","key":"10363_CR21","doi-asserted-by":"publisher","first-page":"2008","DOI":"10.1109\/TII.2018.2862436","volume":"15","author":"Y Li","year":"2019","unstructured":"Li Y, Zhang H, Liang X, Huang B (2019) Event-triggered-based distributed cooperative energy management for multienergy systems. IEEE Trans Ind Inf 15(4):2008\u20132022. https:\/\/doi.org\/10.1109\/TII.2018.2862436","journal-title":"IEEE Trans Ind Inf"},{"key":"10363_CR22","first-page":"36479","volume":"35","author":"C Saharia","year":"2022","unstructured":"Saharia C, Chan W, Saxena S, Li L, Whang J, Denton EL, Ghasemipour K, Gontijo Lopes R, Karagol Ayan B, Salimans T (2022) Photorealistic text-to-image diffusion models with deep language understanding. Adv Neural Inf Process Syst 35:36479\u201336494","journal-title":"Adv Neural Inf Process Syst"},{"key":"10363_CR23","unstructured":"Devlin J, Chang M-W, Lee K, Toutanova K (2018) Bert: pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805"},{"issue":"1","key":"10363_CR24","first-page":"5485","volume":"21","author":"C Raffel","year":"2020","unstructured":"Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, Zhou Y, Li W, Liu PJ (2020) Exploring the limits of transfer learning with a unified text-to-text transformer. J Mach Learn Res 21(1):5485\u20135551","journal-title":"J Mach Learn Res"},{"key":"10363_CR25","unstructured":"Raffel C, Luong M-T, Liu PJ, Weiss RJ, Eck D (2017) Online and linear-time attention by enforcing monotonic alignments. In: International conference on machine learning. PMLR, pp 2837\u20132846"},{"key":"10363_CR26","unstructured":"Sohl-Dickstein J, Weiss E, Maheswaranathan N, Ganguli S (2015) Deep unsupervised learning using nonequilibrium thermodynamics. In: International conference on machine learning. PMLR, pp 2256\u20132265"},{"key":"10363_CR27","unstructured":"Song Y, Ermon S (2019) Generative modeling by estimating gradients of the data distribution. Adv Neural Inf Process Syst 32:11918\u201311930"},{"key":"10363_CR28","first-page":"8780","volume":"34","author":"P Dhariwal","year":"2021","unstructured":"Dhariwal P, Nichol A (2021) Diffusion models beat gans on image synthesis. Adv Neural Inf Process Syst 34:8780\u20138794","journal-title":"Adv Neural Inf Process Syst"},{"key":"10363_CR29","unstructured":"Nichol A, Dhariwal P, Ramesh A, Shyam P, Mishkin P, McGrew B, Sutskever I, Chen M(2021) Glide: towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741"},{"key":"10363_CR30","doi-asserted-by":"crossref","unstructured":"Saharia C, Chan W, Chang H, Lee C, Ho J, Salimans T, Fleet D, Norouzi M(2022) Palette: image-to-image diffusion models. In: ACM SIGGRAPH 2022 conference proceedings, pp 1\u201310","DOI":"10.1145\/3528233.3530757"},{"issue":"4","key":"10363_CR31","first-page":"4713","volume":"45","author":"C Saharia","year":"2022","unstructured":"Saharia C, Ho J, Chan W, Salimans T, Fleet DJ, Norouzi M (2022) Image super-resolution via iterative refinement. IEEE Trans Pattern Anal Mach Intell 45(4):4713\u20134726","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"10363_CR32","doi-asserted-by":"crossref","unstructured":"Whang J, Delbracio M, Talebi H, Saharia C, Dimakis AG, Milanfar P(2022) Deblurring via stochastic refinement. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 16293\u201316303","DOI":"10.1109\/CVPR52688.2022.01581"},{"key":"10363_CR33","unstructured":"Ho J, Salimans T (2022) Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598"},{"key":"10363_CR34","unstructured":"Nichol AQ, Dhariwal P (2021) Improved denoising diffusion probabilistic models. In: International conference on machine learning. PMLR, pp 8162\u20138171"},{"key":"10363_CR35","unstructured":"Song Y, Sohl-Dickstein J, Kingma DP, Kumar A, Ermon S, Poole B (2020) Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456"},{"issue":"4","key":"10363_CR36","doi-asserted-by":"publisher","first-page":"2183","DOI":"10.1109\/TGRS.2017.2776321","volume":"56","author":"X Lu","year":"2017","unstructured":"Lu X, Wang B, Zheng X, Li X (2017) Exploring models and data for remote sensing image caption generation. IEEE Trans Geosci Remote Sens 56(4):2183\u20132195","journal-title":"IEEE Trans Geosci Remote Sens"},{"key":"10363_CR37","doi-asserted-by":"crossref","unstructured":"Xu Y, Yu W, Ghamisi P, Kopp M, Hochreiter S (2022) Txt2img-mhn: remote sensing image generation from text using modern hopfield networks. arXiv preprint arXiv:2208.04441","DOI":"10.1109\/TIP.2023.3323799"},{"key":"10363_CR38","unstructured":"Salimans T, Goodfellow I, Zaremba W, Cheung V, Radford A, Chen X (2016) Improved techniques for training gans. Adv Neural Inf Process Syst 29:2234\u20132242"},{"key":"10363_CR39","unstructured":"Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S (2017) Gans trained by a two time-scale update rule converge to a local nash equilibrium. Adv Neural Inf Process Syst 30"},{"key":"10363_CR40","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li L-J, Li K, Fei-Fei L (2009) Imagenet: A large-scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition. IEEE, pp 248\u2013255","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"10363_CR41","unstructured":"Barratt S, Sharma R (2018) A note on the inception score. arXiv preprint arXiv:1801.01973"},{"key":"10363_CR42","doi-asserted-by":"crossref","unstructured":"Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z (2016) Rethinking the inception architecture for computer vision. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 2818\u20132826","DOI":"10.1109\/CVPR.2016.308"},{"key":"10363_CR43","doi-asserted-by":"crossref","unstructured":"Zhou Y, Zhang R, Chen C, Li C, Tensmeyer C, Yu T, Gu J, Xu J, Sun T (2021) Lafite: towards language-free training for text-to-image generation. arXiv preprint arXiv:2111.13792","DOI":"10.1109\/CVPR52688.2022.01738"},{"key":"10363_CR44","unstructured":"Shazeer N, Stern M(2018) Adafactor: adaptive learning rates with sublinear memory cost. In: International conference on machine learning. PMLR, pp 4596\u20134604"},{"key":"10363_CR45","unstructured":"Kingma DP, Ba J (2014) Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980"},{"key":"10363_CR46","doi-asserted-by":"crossref","unstructured":"Xu T, Zhang P, Huang Q, Zhang H, Gan Z, Huang X, He X (2018) Attngan: fine-grained text to image generation with attentional generative adversarial networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1316\u20131324","DOI":"10.1109\/CVPR.2018.00143"},{"key":"10363_CR47","doi-asserted-by":"crossref","unstructured":"Ruan S, Zhang Y, Zhang K, Fan Y, Tang F, Liu Q, Chen E (2021) Dae-gan: Dynamic aspect-aware gan for text-to-image synthesis supplementary document","DOI":"10.1109\/ICCV48922.2021.01370"},{"key":"10363_CR48","doi-asserted-by":"crossref","unstructured":"Tao M, Tang H, Wu F, Jing X-Y, Bao B-K, Xu C (2022) Df-gan: a simple and effective baseline for text-to-image synthesis. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 16515\u201316525","DOI":"10.1109\/CVPR52688.2022.01602"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-024-10363-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-024-10363-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-024-10363-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,26]],"date-time":"2024-11-26T20:06:28Z","timestamp":1732651588000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-024-10363-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,5]]},"references-count":48,"journal-issue":{"issue":"36","published-print":{"date-parts":[[2024,12]]}},"alternative-id":["10363"],"URL":"https:\/\/doi.org\/10.1007\/s00521-024-10363-3","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,5]]},"assertion":[{"value":"18 May 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 August 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 October 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that there is no conflict of interest regarding the publication of this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}