{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,4,2]],"date-time":"2022-04-02T18:34:02Z","timestamp":1648924442119},"reference-count":38,"publisher":"American Institute of Mathematical Sciences (AIMS)","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["MFC"],"published-print":{"date-parts":[[2021]]},"abstract":"<jats:p xml:lang=\"fr\">&lt;p style='text-indent:20px;'&gt;The recent progress in learning image feature representations has opened the way for tasks such as label-to-image or text-to-image synthesis. However, one particular challenge widely observed in existing methods is the difficulty of synthesizing fine-grained textures and small-scale instances. In this paper, we propose a novel Global-Affine and Local-Specific Generative Adversarial Network (GALS-GAN) to explicitly construct global semantic layouts and learn distinct instance-level features. To achieve this, we adopt the graph convolutional network to calculate the instance locations and spatial relationships from scene graphs, which allows our model to obtain the high-fidelity semantic layouts. Also, a local-specific generator, where we introduce the feature filtering mechanism to separately learn semantic maps for different categories, is utilized to disentangle and generate specific visual features. Moreover, we especially apply a weight map predictor to better combine the global and local pathways considering the highly complementary between these two generation sub-networks. Extensive experiments on the COCO-Stuff and Visual Genome datasets demonstrate the superior generation performance of our model against previous methods, our approach is more capable of capturing photo-realistic local characteristics and rendering small-sized entities with more details.&lt;\/p&gt;<\/jats:p>","DOI":"10.3934\/mfc.2021009","type":"journal-article","created":{"date-parts":[[2021,6,9]],"date-time":"2021-06-09T10:36:57Z","timestamp":1623235017000},"page":"145","source":"Crossref","is-referenced-by-count":0,"title":["Global-Affine and Local-Specific Generative Adversarial Network for semantic-guided image generation"],"prefix":"10.3934","volume":"4","author":[{"given":"Susu","family":"Zhang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiancheng","family":"Ni","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lijun","family":"Hou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zili","family":"Zhou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jie","family":"Hou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Feng","family":"Gao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"2321","reference":[{"key":"key-10.3934\/mfc.2021009-1","doi-asserted-by":"publisher","unstructured":"H. Caesar, J. Uijlings and V. Ferrari, COCO-Stuff: Thing and stuff classes in context, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2018), 1209\u20131218.","DOI":"10.1109\/CVPR.2018.00132"},{"key":"key-10.3934\/mfc.2021009-2","doi-asserted-by":"publisher","unstructured":"W. L. Chen and J. Hays, Sketchygan: Towards diverse and realistic sketch to image synthesis, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2018), 9416\u20139425.","DOI":"10.1109\/CVPR.2018.00981"},{"key":"key-10.3934\/mfc.2021009-3","doi-asserted-by":"publisher","unstructured":"B. Chen, T. Liu, K. Liu, H. Liu and S. Pei, Image Super-Resolution Using Complex Dense Block on Generative Adversarial Networks, <i>IEEE International Conference on Image Processing<\/i>, (2019), 2866\u20132870.","DOI":"10.1109\/ICIP.2019.8803711"},{"key":"key-10.3934\/mfc.2021009-4","doi-asserted-by":"publisher","unstructured":"Y. Choi, M. Choi, M. Kim, J. M. Ha, S. H. Kim and J. Choo, Stargan: Unified generative adversarial networks for multi-domain image-to-image translation, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2018), 8789\u20138797.","DOI":"10.1109\/CVPR.2018.00916"},{"key":"key-10.3934\/mfc.2021009-5","doi-asserted-by":"publisher","unstructured":"Y. Choi, Y. Uh, J. Yoo and J. W. Ha, StarGAN v2: Diverse image synthesis for multiple domains, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2020), 8185\u20138194.","DOI":"10.1109\/CVPR42600.2020.00821"},{"key":"key-10.3934\/mfc.2021009-6","doi-asserted-by":"publisher","unstructured":"H. Dhamo, A. Farshad, I. Laina, N. Navab, G. D. Hager, F. Tombari and C. Rupprecht, Semantic image manipulation using scene graphs, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2020), 5212\u20135221.","DOI":"10.1109\/CVPR42600.2020.00526"},{"key":"key-10.3934\/mfc.2021009-7","doi-asserted-by":"publisher","unstructured":"C. Gao, Q. Liu, Q. Xu, L. Wang, J. Liu and C. Zou, SketchyCOCO: Image generation from freehand scene sketches, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2020), 5173\u20135182.","DOI":"10.1109\/CVPR42600.2020.00522"},{"key":"key-10.3934\/mfc.2021009-8","unstructured":"I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville and Y. Bengio, Generative adversarial nets, <i>Advances in Neural Information Processing Systems<\/i>, (2014), 2672\u20132680."},{"key":"key-10.3934\/mfc.2021009-9","doi-asserted-by":"publisher","unstructured":"S. Hong, D. Yang, J. Choi and H. Lee, Inferring semantic layout for hierarchical text-to-image synthesis, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2018), 7986\u20137994.","DOI":"10.1109\/CVPR.2018.00833"},{"key":"key-10.3934\/mfc.2021009-10","doi-asserted-by":"publisher","unstructured":"J. Johnson, A. Gupta and F. F. Li, Image generation from scene graphs, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2018), 1219\u20131228.","DOI":"10.1109\/CVPR.2018.00133"},{"key":"key-10.3934\/mfc.2021009-11","doi-asserted-by":"publisher","unstructured":"T. Kaneko, Y. Ushiku and T. Harada, Label-noise robust generative adversarial networks, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2019), 2462\u20132471.","DOI":"10.1109\/CVPR.2019.00257"},{"key":"key-10.3934\/mfc.2021009-12","doi-asserted-by":"publisher","unstructured":"S. W. Kim, Y. Zhou, J. Philion, A. Torralba and S. Fidler, Learning to Simulate Dynamic Environments With GameGAN, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2020), 1228\u20131237.","DOI":"10.1109\/CVPR42600.2020.00131"},{"key":"key-10.3934\/mfc.2021009-13","unstructured":"D. Kingma and J. Ba, Adam: A method for stochastic optimization, <i>International Conference on Learning Representations<\/i>, 2019."},{"key":"key-10.3934\/mfc.2021009-14","unstructured":"T. N. Kipf and M. Welling, Semi-supervised classification with graph convolutional networks, preprint, arXiv: 1609.02907."},{"key":"key-10.3934\/mfc.2021009-15","doi-asserted-by":"publisher","unstructured":"R. Krishna.Visual genome: Connecting language and vision using crowdsourced dense image annotations, <i>International Journal of Computer Vision<\/i>, <b>123<\/b> (2017), 32-73.","DOI":"10.1007\/s11263-016-0981-7"},{"key":"key-10.3934\/mfc.2021009-16","doi-asserted-by":"publisher","unstructured":"T. Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, C. L. Zitnick.Microsoft coco: Common objects in context, <i>European Conference on Computer Vision<\/i>, <b>8693<\/b> (2014), 740-755.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"key-10.3934\/mfc.2021009-17","doi-asserted-by":"publisher","unstructured":"M. Li, H. Huang, L. Ma, W. Liu, T. Zhang and Y. Jiang, Unsupervised image-to-image translation with stacked cycle-consistent adversarial networks, <i>European Conference on Computer Vision<\/i>, (2018), 186\u2013201.","DOI":"10.1007\/978-3-030-01240-3_12"},{"key":"key-10.3934\/mfc.2021009-18","doi-asserted-by":"publisher","unstructured":"W. Li, P. Zhang, L. Zhang, Q. Huang, X. He, S. Lyu and J. Gao, Object-driven text-to-image synthesis via adversarial training, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2019), 12166\u201312174.","DOI":"10.1109\/CVPR.2019.01245"},{"key":"key-10.3934\/mfc.2021009-19","unstructured":"Y. Li, T. Ma, Y. Bai, N. Duan, S. Wei, and X. Wang, Pastegan: A semi-parametric method to generate image from scene graph, <i>Advances in Neural Information Processing Systems<\/i>, 2019."},{"key":"key-10.3934\/mfc.2021009-20","doi-asserted-by":"publisher","unstructured":"B. Li, B. Zhuang, M. Li and J. Gu, Seq-SG2SL: Inferring semantic layout from scene graph through sequence to sequence learning, <i>IEEE International Conference on Computer Vision<\/i>, (2019), 7434\u20137442.","DOI":"10.1109\/ICCV.2019.00753"},{"key":"key-10.3934\/mfc.2021009-21","doi-asserted-by":"publisher","unstructured":"S. Liu, T. Wang, D. Bau, J. Y. Zhu and A. Torralba, Diverse Image Generation via Self-Conditioned GANs, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2020), 14274\u201314283.","DOI":"10.1109\/CVPR42600.2020.01429"},{"key":"key-10.3934\/mfc.2021009-22","unstructured":"S. Nam, Y. Kim and S. J. Kim, Text-adaptive generative adversarial networks: Manipulating images with natural language, <i>Advances in Neural Information Processing Systems<\/i>, (2018), 42\u201351."},{"key":"key-10.3934\/mfc.2021009-23","doi-asserted-by":"publisher","unstructured":"J. C. Ni, S. S. Zhang, Z. L. Zhou, J. Hou, F. Gao.Instance Mask Embedding and Attribute-Adaptive Generative Adversarial Network for Text-to-Image Synthesis, <i>IEEE Access<\/i>, <b>8<\/b> (2020), 37697-37711.","DOI":"10.1109\/ACCESS.2020.2975841"},{"key":"key-10.3934\/mfc.2021009-24","doi-asserted-by":"publisher","unstructured":"T. Park, M. Y. Liu, T. C. Wang and J. Y. Zhu, Semantic image synthesis with spatially-adaptive normalization, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2019), 2332\u20132341.","DOI":"10.1109\/CVPR.2019.00244"},{"key":"key-10.3934\/mfc.2021009-25","doi-asserted-by":"crossref","unstructured":"T. Qiao, J. Zhang, D. Xu, and D. Tao, Mirrorgan: Learning text-to-image generation by redescription, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2019), 1505\u20131514.","DOI":"10.1109\/CVPR.2019.00160"},{"key":"key-10.3934\/mfc.2021009-26","unstructured":"S. Ravuri and O. Vinyals, Classification accuracy score for conditional generative models, preprint, arXiv: 1905.10887."},{"key":"key-10.3934\/mfc.2021009-27","doi-asserted-by":"publisher","unstructured":"S. Ren, K. He, R. Girshick, J. Sun.Faster R-CNN: Towards real-time object detection with region proposal networks, <i>IEEE Transactions on Pattern Analysis and Machine Intelligence<\/i>, <b>39<\/b> (2016), 1137-1149.","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"key-10.3934\/mfc.2021009-28","doi-asserted-by":"publisher","unstructured":"S. Sah, D. Peri, A. Shringi, C. Zhang, M. Dominguez, A. Savakis and R. Ptucha, Semantically invariant text-to-image generation, <i>IEEE International Conference on Image Processing<\/i>, (2018), 3783\u20133787.","DOI":"10.1109\/ICIP.2018.8451656"},{"key":"key-10.3934\/mfc.2021009-29","doi-asserted-by":"publisher","unstructured":"Y. Shen, J. Gu, X. Tang and B. Zhou, Interpreting the Latent space of GANs for semantic face editing, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2020), 9240\u20139249.","DOI":"10.1109\/CVPR42600.2020.00926"},{"key":"key-10.3934\/mfc.2021009-30","doi-asserted-by":"publisher","unstructured":"T. R. Shaham, T. Dekel and T. Michaeli, SinGAN: Learning a generative model from a single natural image, <i>IEEE International Conference on Computer Vision<\/i>, (2019), 4569\u20134579.","DOI":"10.1109\/ICCV.2019.00467"},{"key":"key-10.3934\/mfc.2021009-31","unstructured":"W. Sun and T. F. Wu, Learning Layout and Style Reconfigurable GANs for Controllable Image Synthesis, preprint, arXiv: 2003.11571."},{"key":"key-10.3934\/mfc.2021009-32","unstructured":"T. Sylvain, P. C. Zhang, Y. Bengio, R. D. Hjelm and S. Sharma, Object-centric image generation from layouts, preprint, arXiv: 2003.07449."},{"key":"key-10.3934\/mfc.2021009-33","doi-asserted-by":"publisher","unstructured":"C. Szegedy, et al., Going deeper with convolutions, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2015), 1\u20139.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"key-10.3934\/mfc.2021009-34","doi-asserted-by":"publisher","unstructured":"H. Tang, H. Liu, N. Sebe.Unified generative adversarial networks for controllable image-to-image translation, <i>IEEE Transactions on Image Processing<\/i>, <b>29<\/b> (2020), 8916-8929.","DOI":"10.1109\/TIP.2020.3021789"},{"key":"key-10.3934\/mfc.2021009-35","doi-asserted-by":"publisher","unstructured":"N. N. Vo and J. Hays, Localizing and orienting street views using overhead imagery, <i>European Conference on Computer Vision<\/i>, (2016), 494\u2013509.","DOI":"10.1007\/978-3-319-46448-0_30"},{"key":"key-10.3934\/mfc.2021009-36","unstructured":"D. M. Vo and A. Sugimoto, Visual-relation conscious image generation from structured-text, preprint, arXiv: 1908.01741."},{"key":"key-10.3934\/mfc.2021009-37","doi-asserted-by":"publisher","unstructured":"H. Yu, Y. Huang, L. Pi and L. Wang, Recurrent deconvolutional generative adversarial networks with application to video generation, <i>Pattern Recognition and Computer Vision<\/i>, (2019), 18\u201328.","DOI":"10.1007\/978-3-030-31723-2_2"},{"key":"key-10.3934\/mfc.2021009-38","doi-asserted-by":"publisher","unstructured":"L. Z. Zhang, J. C. Wang, Y. S. Xu, J. Min, T. Wen, J. C. Gee and J. B. Shi, Nested Scale-Editing for Conditional Image Synthesis, <i>IEEE Conference on Computer Vision and Pattern Recognition<\/i>, (2020), 5476\u20135486.","DOI":"10.1109\/CVPR42600.2020.00552"}],"container-title":["Mathematical Foundations of Computing"],"original-title":[],"deposited":{"date-parts":[[2021,8,25]],"date-time":"2021-08-25T09:12:58Z","timestamp":1629882778000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.aimsciences.org\/article\/doi\/10.3934\/mfc.2021009"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021]]},"references-count":38,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2021]]}},"alternative-id":["2577-8838_2021_3_145"],"URL":"https:\/\/doi.org\/10.3934\/mfc.2021009","relation":{},"ISSN":["2577-8838"],"issn-type":[{"value":"2577-8838","type":"print"}],"subject":[],"published":{"date-parts":[[2021]]}}}