{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T05:04:58Z","timestamp":1782450298024,"version":"3.54.5"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"18","license":[{"start":{"date-parts":[[2021,5,20]],"date-time":"2021-05-20T00:00:00Z","timestamp":1621468800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,5,20]],"date-time":"2021-05-20T00:00:00Z","timestamp":1621468800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100008982","name":"Qatar National Research Fund","doi-asserted-by":"publisher","award":["NPRP10-0205-170346"],"award-info":[{"award-number":["NPRP10-0205-170346"]}],"id":[{"id":"10.13039\/100008982","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"published-print":{"date-parts":[[2021,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The automatic generation of realistic images directly from a story text is a very challenging problem, as it cannot be addressed using a single image generation approach due mainly to the semantic complexity of the story text constituents. In this work, we propose a new approach that decomposes the task of story visualization into three phases: semantic text understanding, object layout prediction, and image generation and refinement. We start by simplifying the text using a scene graph triple notation that encodes semantic relationships between the story objects. We then introduce an object layout module to capture the features of these objects from the corresponding scene graph. Specifically, the object layout module aggregates individual object features from the scene graph as well as averaged or likelihood object features generated by a graph convolutional neural network. All these features are concatenated to form semantic triples that are then provided to the image generation framework. For the image generation phase, we adopt a scene graph image generation framework as stage-I, which is refined using a StackGAN as stage-II conditioned on the object layout module and the generated output image from stage-I. Our approach renders object details in high-resolution images while keeping the image structure consistent with the input text. To evaluate the performance of our approach, we use the COCO dataset and compare it with three baseline approaches, namely, sg2im, StackGAN and AttnGAN, in terms of image quality and user evaluation. According to the obtained assessment results, our object layout guidance-based approach significantly outperforms the abovementioned baseline approaches in terms of the accuracy of semantic matching and realism of the generated images representing the story text sentences.<\/jats:p>","DOI":"10.1007\/s11042-021-11038-0","type":"journal-article","created":{"date-parts":[[2021,5,20]],"date-time":"2021-05-20T20:02:21Z","timestamp":1621540941000},"page":"27423-27443","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":22,"title":["Improving text-to-image generation with object layout guidance"],"prefix":"10.1007","volume":"80","author":[{"given":"Jezia","family":"Zakraoui","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6434-1790","authenticated-orcid":false,"given":"Moutaz","family":"Saleh","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Somaya","family":"Al-Maadeed","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jihad Mohammed","family":"Jaam","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,5,20]]},"reference":[{"key":"11038_CR1","doi-asserted-by":"crossref","unstructured":"Caesar H, Uijlings J, Ferrari V (2018) Coco-stuff: Thing and Stuff Classes in Context, ArXiv:1612.03716","DOI":"10.1109\/CVPR.2018.00132"},{"key":"11038_CR2","unstructured":"Chigozie N, Ijomah W, Gachagan A, Marshall S (2018) Activation Functions: Comparison of trends in Practice and Research for Deep Learning, ArXiv abs\/1811.03378"},{"key":"11038_CR3","unstructured":"Denton E, Chintala S, Szlam A, Fergus R (2015) Deep generative image models using a Laplacian pyramid of adversarial networks, ArXiv e-prints"},{"key":"11038_CR4","unstructured":"Gangyan Z, Zhaohui L, Yuan Z (2019) PororoGAN: an improved story visualization model on Pororo-SV dataset, in 3rd International Conference on Computer Science and Artificial Intelligence, Normal IL USA"},{"key":"11038_CR5","unstructured":"Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2014) Generative adversarial nets, In Advances in Neural Information Processing Systems, p. pages 2672\u20132680"},{"key":"11038_CR6","doi-asserted-by":"crossref","unstructured":"Gu S, Bao J, Chen D, Wen F (2020) GIQA: Generated Image Quality Assessment, ArXiv:2003.08932","DOI":"10.1007\/978-3-030-58621-8_22"},{"key":"11038_CR7","unstructured":"Gulrajani I, Ahmed F, Arjovsky M, Dumoulin V, Courville A (2017) Improved training of wasserstein GANs, in Proceedings of the 31st International Conference on Neural Information Processing Systems, pages 5769\u20135779"},{"key":"11038_CR8","unstructured":"Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S (2017) GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium, in Advances in neural information processing systems. pages 6626\u20136637"},{"key":"11038_CR9","unstructured":"Hinz T, Heinrich S, Wermter S (2019) Generating multiple objects at spatially distinct locations, in International Conference on Learning Representations, ICLR"},{"key":"11038_CR10","doi-asserted-by":"crossref","unstructured":"Hong S, Yang D, Choi J, Lee H (2018) Inferring semantic layout for hierarchical text-to-image synthesis, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 7986\u20137994","DOI":"10.1109\/CVPR.2018.00833"},{"key":"11038_CR11","doi-asserted-by":"crossref","unstructured":"Huang C-J, Li C-T, Shan M-K (2013) VizStory: visualization of digital narrative for fairy Tales, in Conference on Technologies and Applications of Artificial Intelligence, pages 67\u201372, USA","DOI":"10.1109\/TAAI.2013.26"},{"key":"11038_CR12","doi-asserted-by":"crossref","unstructured":"Huang X, Li Y, Poursaeed O, Hopcroft J, Belongie S (2016) Stacked generative adversarial networks, CVPR, vol. arXiv:1612.04357","DOI":"10.1109\/CVPR.2017.202"},{"key":"11038_CR13","doi-asserted-by":"crossref","unstructured":"Isola P, Zhu J-Y, Zhou T, Efros AA (2017) Image-to-image translation with conditional adversarial network, ArXiv:1611.07004","DOI":"10.1109\/CVPR.2017.632"},{"key":"11038_CR14","doi-asserted-by":"crossref","unstructured":"Johnson J, Gupta A, Fei-Fei L (2018) Image Generation from Scene Graphs, in IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 1219\u20131228, doi: 10.1109\/CVPR.2018.00133, Salt Lake City, UT, USA","DOI":"10.1109\/CVPR.2018.00133"},{"key":"11038_CR15","unstructured":"Kingma DP, Welling M (2013) Auto-encoding variational bayes, ArXiv, p. arXiv:1312.6114"},{"key":"11038_CR16","doi-asserted-by":"crossref","unstructured":"Li Y, Gan Z, Shen Y, Liu J, Cheng Y, Wu Y, Carin L, Carlson D, Gao J (2019) StoryGAN: A Sequential Conditional GAN for Story Visualization, Conference on Computer Vision and Pattern Recognition (CVPR), vol. abs\/1812.02784, pp. 6322\u20136331","DOI":"10.1109\/CVPR.2019.00649"},{"key":"11038_CR17","first-page":"3948","volume":"32","author":"Y Li","year":"2019","unstructured":"Li Y, Ma T, Bai Y, Duan N, Wei S, Wang X (2019) PasteGAN: a semi-parametric method to generate image from scene graph. Adv Neural Inf Proces Syst 32:3948\u20133958","journal-title":"Adv Neural Inf Proces Syst"},{"key":"11038_CR18","doi-asserted-by":"crossref","unstructured":"Li W, Zhang P, Zhang L, Huang Q, He X, Lyu S, Gao J (2019) Object-driven Text-to-Image Synthesis via Adversarial, CVPR, p. arXiv:1902.10740","DOI":"10.1109\/CVPR.2019.01245"},{"key":"11038_CR19","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll\u2019ar P, Zitnick CL (2014) Microsoft COCO: Common Objects in Context, in European conference on computer vision, pages 740\u2013755","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"11038_CR20","doi-asserted-by":"crossref","unstructured":"Lu Y, Wu S, Tai Y-W, Tang C-K (2018) Image generation from sketch constraint using contextual gan, Computer Vision ECCV 2018. Lecture notes in computer science, vol 11220, pp https:\/\/doiorg\/101007\/978-3","DOI":"10.1007\/978-3-030-01270-0_13"},{"key":"11038_CR21","unstructured":"Mansimov E, Parisotto E, Lei Ba J, Salakhutdinov R (2016) Generating Images from Captions with Attention, ArXiv:1511.02793"},{"key":"11038_CR22","unstructured":"Mirza M, Osindero S (2014) Conditional Generative Adversarial Nets, ArXiv:1411.1784"},{"key":"11038_CR23","unstructured":"Odena A, Olah C, Shlens J (2017) Conditional Image Synthesis With Auxiliary Classifier GANs, ArXiv:1610.09585"},{"key":"11038_CR24","doi-asserted-by":"crossref","unstructured":"Ouyang X, Zhang X, Ma D, Agam G (2018) Generating image sequence from description with lstm conditional Gan, ArXiv: 1806.03027","DOI":"10.1109\/ICPR.2018.8545419"},{"issue":"7","key":"11038_CR25","first-page":"579","volume":"8","author":"M-C Popescu","year":"2009","unstructured":"Popescu M-C, Balas VE, Perescu-Popescu L, Mastorakis N (2009) Multilayer perceptron and neural networks. World Scientific and Engineering Academy and Society (WSEAS) 8(7):579\u2013588","journal-title":"World Scientific and Engineering Academy and Society (WSEAS)"},{"key":"11038_CR26","doi-asserted-by":"crossref","unstructured":"Qiao T, Zhang J, Xu D, Tao D (2019) MirrorGAN: learning text-to-image generation by Redescription, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1505\u20131514","DOI":"10.1109\/CVPR.2019.00160"},{"key":"11038_CR27","doi-asserted-by":"crossref","unstructured":"Qiao X, Zheng Q, Cao Y, Lau RWH (2019) Tell me where I am: object-level scene context prediction, in Conference on Computer Vision and Pattern Recognition (CVPR), pages 2628\u20132636, Long Beach, CA, USA","DOI":"10.1109\/CVPR.2019.00274"},{"key":"11038_CR28","unstructured":"Reed S, Akata Z, Yan X, Logeswaran L, Schiele B, Lee H (2016) Generative adversarial text to image synthesis, in Proceedings of the 33rd International Conference on International Conference on Machine Learning, pages 1060\u20131069"},{"key":"11038_CR29","unstructured":"Salimans T, Goodfellow I, Zaremba W, Cheung V, Radford A, Chen X (2016) Improved techniques for training gans, Adv neural inf process syst, p. 2234\u20132242"},{"key":"11038_CR30","unstructured":"Sharma S, Suhubdy D, Michalski V, Kahou SE, Bengio Y (2018) ChatPainter: Improving Text to Image Generation using Dialogue, ArXiv:1802.08216"},{"key":"11038_CR31","unstructured":"Sutskever I, Vinyals O, Le QV (2014) Sequence to sequence learning with neural networks, In Advances in Neural Information Processing Systems(NeurIPS)"},{"key":"11038_CR32","doi-asserted-by":"crossref","unstructured":"Tan F, Feng S, Ordonez V (2019) Text2Scene: generating compositional scenes from textual descriptions, in IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR.2019.00687"},{"key":"11038_CR33","unstructured":"\"The Stanford Natural Language Processing Group,\" [Online]. Available: https:\/\/nlp.stanford.edu\/software\/scenegraph-parser.shtml. [Accessed 1 October 2019]."},{"key":"11038_CR34","doi-asserted-by":"crossref","unstructured":"Vo DM, Sugimoto A (2020) Visual-relation conscious image generation from Structured-Text, in arXiv preprint arXiv:1908.01741","DOI":"10.1007\/978-3-030-58604-1_18"},{"key":"11038_CR35","doi-asserted-by":"publisher","unstructured":"Weisberg DS, Hopkins EJ (2020) Preschoolers' extension and export of information from realistic and fantastical stories, Infant and Child Development, p. doi:https:\/\/doi.org\/10.1002\/icd.2182","DOI":"10.1002\/icd.2182"},{"key":"11038_CR36","doi-asserted-by":"publisher","unstructured":"Xu T, Zhang P, Huang Q, Zhang H, Gan Z, Huang X, He X (2018) Attngan: Fine-grained text to image generation with attentional generative adversarial networks, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1316\u20131324, doi: https:\/\/doi.org\/10.1109\/CVPR.2018.00143., Salt Lake City, UT","DOI":"10.1109\/CVPR.2018.00143"},{"key":"11038_CR37","doi-asserted-by":"crossref","unstructured":"Zhang H, Xu T, Li H, Zhang S, Wang X, Huang X, Metaxas D (2017) StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial Networks, ArXiv:1710.10916","DOI":"10.1109\/ICCV.2017.629"},{"key":"11038_CR38","doi-asserted-by":"crossref","unstructured":"Zhang H, Xu T, Li H, Zhang S, Wang X, Huang X, Metaxas D (2017) StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks, in IEEE International Conference on Computer Vision (ICCV), pp. 5908\u20135916, doi: 10.1109\/ICCV.2017.629, Venice, Italy","DOI":"10.1109\/ICCV.2017.629"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-021-11038-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-021-11038-0\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-021-11038-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,7,24]],"date-time":"2021-07-24T15:15:42Z","timestamp":1627139742000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-021-11038-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,20]]},"references-count":38,"journal-issue":{"issue":"18","published-print":{"date-parts":[[2021,7]]}},"alternative-id":["11038"],"URL":"https:\/\/doi.org\/10.1007\/s11042-021-11038-0","relation":{},"ISSN":["1380-7501","1573-7721"],"issn-type":[{"value":"1380-7501","type":"print"},{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,5,20]]},"assertion":[{"value":"22 July 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 February 2021","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 May 2021","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 May 2021","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}