{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T18:24:32Z","timestamp":1784139872179,"version":"3.55.0"},"reference-count":69,"publisher":"Springer Science and Business Media LLC","issue":"12","license":[{"start":{"date-parts":[[2022,9,17]],"date-time":"2022-09-17T00:00:00Z","timestamp":1663372800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,9,17]],"date-time":"2022-09-17T00:00:00Z","timestamp":1663372800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Despite their recent successes, generative adversarial networks (GANs) for semantic image synthesis still suffer from poor image quality when trained with only adversarial supervision. Previously, additionally employing the VGG-based perceptual loss has helped to overcome this issue, significantly improving the synthesis quality, but at the same time limited the progress of GAN models for semantic image synthesis. In this work, we propose a novel, simplified GAN model, which needs only adversarial supervision to achieve high quality results. We re-design the discriminator as a semantic segmentation network, directly using the given semantic label maps as the ground truth for training. By providing stronger supervision to the discriminator as well as to the generator through spatially- and semantically-aware discriminator feedback, we are able to synthesize images of higher fidelity and with a better alignment to their input label maps, making the use of the perceptual loss superfluous. Furthermore, we enable high-quality multi-modal image synthesis through global and local sampling of a 3D noise tensor injected into the generator, which allows complete or partial image editing. We show that images synthesized by our model are more diverse and follow the color and texture distributions of real images more closely. We achieve a strong improvement in image synthesis quality over prior state-of-the-art models across the commonly used ADE20K, Cityscapes, and COCO-Stuff datasets using only adversarial supervision. In addition, we investigate semantic image synthesis under severe class imbalance and sparse annotations, which are common aspects in practical applications but were overlooked in prior works. To this end, we evaluate our model on LVIS, a dataset originally introduced for long-tailed object recognition. We thereby demonstrate high performance of our model in the sparse and unbalanced data regimes, achieved by means of the proposed 3D noise and the ability of our discriminator to balance class contributions directly in the loss function. Our code and pretrained models are available at <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/boschresearch\/OASIS\">https:\/\/github.com\/boschresearch\/OASIS<\/jats:ext-link>.\n<\/jats:p>","DOI":"10.1007\/s11263-022-01673-x","type":"journal-article","created":{"date-parts":[[2022,9,17]],"date-time":"2022-09-17T02:02:25Z","timestamp":1663380145000},"page":"2903-2923","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":49,"title":["OASIS: Only Adversarial Supervision for Semantic Image Synthesis"],"prefix":"10.1007","volume":"130","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2803-335X","authenticated-orcid":false,"given":"Vadim","family":"Sushko","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Edgar","family":"Sch\u00f6nfeld","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dan","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Juergen","family":"Gall","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bernt","family":"Schiele","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anna","family":"Khoreva","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,9,17]]},"reference":[{"key":"1673_CR1","doi-asserted-by":"crossref","unstructured":"Alharbi, Y., & Wonka, P. (2020). Disentangled image generation through structured noise injection. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR42600.2020.00518"},{"key":"1673_CR2","unstructured":"Arjovsky, M., & Bottou, L. (2017). Towards principled methods for training generative adversarial networks. In International Conference on Learning Representations (ICLR)."},{"key":"1673_CR3","doi-asserted-by":"publisher","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","volume":"39","author":"V Badrinarayanan","year":"2016","unstructured":"Badrinarayanan, V., Kendall, A., & Cipolla, R. (2016). Segnet: A deep convolutional encoder-decoder architecture for image segmentation. Transactions on Pattern Analysis and Machine Intelligence, 39, 2481\u20132495.","journal-title":"Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1673_CR4","unstructured":"Brock, A., Donahue, J., & Simonyan, K. (2019). Large scale GAN training for high fidelity natural image synthesis. In International conference on learning representations (ICLR)."},{"key":"1673_CR5","unstructured":"Bruna, J., Sprechmann, P., & LeCun, Y. (2016). Super-resolution with deep convolutional sufficient statistics. In International conference on learning representations (ICLR)."},{"key":"1673_CR6","doi-asserted-by":"crossref","unstructured":"Caesar, H., Uijlings, J., & Ferrari, V. (2018). Coco-stuff: Thing and stuff classes in context. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2018.00132"},{"key":"1673_CR7","unstructured":"Casanova, A., Careil, M., Verbeek, J., Drozdzal, M., & Romero Soriano, A. (2021). Instance-conditioned gan. In Advances in neural information processing systems (NeurIPS)."},{"key":"1673_CR8","unstructured":"Chen, L. C., Papandreou, G., Kokkinos, I., Murphy, K., & Yuille, A. L. (2015). Semantic image segmentation with deep convolutional nets and fully connected crfs. In International conference on learning representations (ICLR)."},{"key":"1673_CR9","doi-asserted-by":"crossref","unstructured":"Chen, L. C., Zhu, Y., Papandreou, G., Schroff, F., & Adam, H. (2018). Encoder-decoder with atrous separable convolution for semantic image segmentation. In European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"1673_CR10","doi-asserted-by":"crossref","unstructured":"Chen, Q., & Koltun, V. (2017). Photographic image synthesis with cascaded refinement networks. In International conference on computer vision (ICCV).","DOI":"10.1109\/ICCV.2017.168"},{"key":"1673_CR11","doi-asserted-by":"crossref","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., & Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2016.350"},{"key":"1673_CR12","doi-asserted-by":"crossref","unstructured":"Cubuk, E. D., Zoph, B., Shlens, J., & Le, Q. (2020). Randaugment: Practical automated data augmentation with a reduced search space. In Advances in neural information processing systems (NeurIPS).","DOI":"10.1109\/CVPRW50498.2020.00359"},{"key":"1673_CR13","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, LJ., Li, K., & Fei-Fei, L. (2009). ImageNet: A large-scale hierarchical image database. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"1673_CR14","doi-asserted-by":"crossref","unstructured":"Gatys, L., Ecker, A. S., Bethge, M. (2015). Texture synthesis using convolutional neural networks. In Advances in neural information processing systems (NeurIPs).","DOI":"10.1109\/CVPR.2016.265"},{"key":"1673_CR15","doi-asserted-by":"crossref","unstructured":"Gatys, L. A., Ecker, A. S., & Bethge, M. (2016). Image style transfer using convolutional neural networks. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2016.265"},{"key":"1673_CR16","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. In Advances in neural information processing systems (NeurIPS)."},{"key":"1673_CR17","doi-asserted-by":"crossref","unstructured":"Gupta, A., Dollar, P., & Girshick, R. (2019). LVIS: A dataset for large vocabulary instance segmentation. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2019.00550"},{"key":"1673_CR18","unstructured":"Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., & Hochreiter, S. (2017). GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Advances in neural information processing systems (NeurIPS)."},{"key":"1673_CR19","doi-asserted-by":"crossref","unstructured":"Huang, X., Liu, M. Y., Belongie, S., & Kautz, J. (2018). Multimodal unsupervised image-to-image translation. In European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-01219-9_11"},{"key":"1673_CR20","doi-asserted-by":"crossref","unstructured":"Isola, P., Zhu, J. Y., Zhou, T., & Efros, A. A. (2017). Image-to-image translation with conditional adversarial networks. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2017.632"},{"key":"1673_CR21","doi-asserted-by":"crossref","unstructured":"Johnson, J., Alahi, A., & Fei-Fei, L. (2016). Perceptual losses for real-time style transfer and super-resolution. In European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"1673_CR22","doi-asserted-by":"crossref","unstructured":"Karras, T., Laine, S., & Aila, T. (2019). A style-based generator architecture for generative adversarial networks. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2019.00453"},{"key":"1673_CR23","unstructured":"Karras, T., Aittala, M., Hellsten, J., Laine, S., Lehtinen, J., & Aila, T. (2020a). Training generative adversarial networks with limited data. In Advances in neural information processing systems (NeurIPS)."},{"key":"1673_CR24","doi-asserted-by":"crossref","unstructured":"Karras, T., Laine, S., Aittala, M., Hellsten, J., Lehtinen, J., & Aila, T. (2020b). Analyzing and improving the image quality of stylegan. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR42600.2020.00813"},{"key":"1673_CR25","unstructured":"Karras, T., Aittala, M., Laine, S., H\u00e4rk\u00f6nen, E., Hellsten, J., Lehtinen, J., & Aila, T. (2021). Alias-free generative adversarial networks. In Advances in neural information processing systems (NeurIPS)."},{"key":"1673_CR26","unstructured":"Kingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. In International conference on learning representations (ICLR)."},{"key":"1673_CR27","unstructured":"Li, K., & Malik, J. (2018). Implicit maximum likelihood estimation. arXiv:1809.09087."},{"key":"1673_CR28","doi-asserted-by":"crossref","unstructured":"Li, K., Zhang, T., & Malik, J. (2019). Diverse image synthesis from semantic layouts via conditional imle. In International conference on computer vision (ICCV).","DOI":"10.1109\/ICCV.2019.00432"},{"key":"1673_CR29","doi-asserted-by":"crossref","unstructured":"Li, Y., Li, Y., Lu, J., Shechtman, E., Lee, Y. J., & Singh, K. K. (2021). Collaging class-specific gans for semantic image synthesis. In International conference on computer vision (ICCV).","DOI":"10.1109\/ICCV48922.2021.01415"},{"key":"1673_CR30","unstructured":"Liu, B., Zhu, Y., Song, K., & Elgammal, A. (2021). Towards faster and stabilized gan training for high-fidelity few-shot image synthesis. In International conference on learning representations (ICLR)."},{"key":"1673_CR31","unstructured":"Liu, X., Yin, G., Shao, J., Wang, X., & Li, H. (2019). Learning to predict layout-to-image conditional convolutions for semantic image synthesis. In Advances in neural information processing systems (NeurIPS)."},{"key":"1673_CR32","unstructured":"Mirza, M., & Osindero, S. (2014). Conditional generative adversarial nets. arXiv:1411.1784"},{"key":"1673_CR33","unstructured":"Miyato, T., & Koyama, M. (2018). cGANs with projection discriminator. In International conference on learning representations (ICLR)."},{"key":"1673_CR34","unstructured":"Miyato, T., Kataoka, T., Koyama, M., & Yoshida, Y. (2018). Spectral normalization for generative adversarial networks. In International conference on learning representations (ICLR)."},{"key":"1673_CR35","doi-asserted-by":"crossref","unstructured":"Ntavelis, E., Romero, A., Kastanis, I., Van\u00a0Gool, L., & Timofte, R. (2020). Sesame: Semantic editing of scenes by adding, manipulating or erasing objects. In European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-58542-6_24"},{"key":"1673_CR36","doi-asserted-by":"publisher","first-page":"51","DOI":"10.1016\/0031-3203(95)00067-4","volume":"29","author":"T Ojala","year":"1996","unstructured":"Ojala, T., Pietik\u00e4inen, M., & Harwood, D. (1996). A comparative study of texture measures with classification based on featured distributions. Pattern Recognition, 29, 51\u201359.","journal-title":"Pattern Recognition"},{"key":"1673_CR37","doi-asserted-by":"crossref","unstructured":"Park, T., Liu, M. Y., Wang, T. C., & Zhu, J. Y. (2019a). Gaugan: Semantic image synthesis with spatially adaptive normalization. In ACM SIGGRAPH.","DOI":"10.1145\/3306305.3332370"},{"key":"1673_CR38","doi-asserted-by":"crossref","unstructured":"Park, T., Liu, M. Y., Wang, T. C., & Zhu, J. Y. (2019b). Semantic image synthesis with spatially-adaptive normalization. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2019.00244"},{"key":"1673_CR39","doi-asserted-by":"crossref","unstructured":"Park, T., Efros, A. A., Zhang, R., & Zhu, J. Y. (2020). Contrastive learning for unpaired image-to-image translation. In European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-58545-7_19"},{"key":"1673_CR40","doi-asserted-by":"crossref","unstructured":"Qi, X., Chen, Q., Jia, J., & Koltun, V. (2018). Semi-parametric image synthesis. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2018.00918"},{"key":"1673_CR41","unstructured":"Reed, S. E., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., & Lee, H. (2016). Generative adversarial text to image synthesis. In International conference on machine learning (ICML)."},{"key":"1673_CR42","doi-asserted-by":"crossref","unstructured":"Richardson, E., Alaluf, Y., Patashnik, O., Nitzan, Y., Azar, Y., Shapiro, S., & Cohen-Or, D. (2021). Encoding in style: A stylegan encoder for image-to-image translation. In Conference on Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR46437.2021.00232"},{"key":"1673_CR43","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., & Brox, T. (2015). U-net: Convolutional networks for biomedical image segmentation. In MICCAI.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"1673_CR44","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1023\/A:1026543900054","volume":"40","author":"Y Rubner","year":"2000","unstructured":"Rubner, Y., Tomasi, C., & Guibas, L. J. (2000). The earth mover\u2019s distance as a metric for image retrieval. International Journal of Computer Vision (IJCV), 40, 99\u2013121.","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"1673_CR45","unstructured":"Sauer, A., Chitta, K., M\u00fcller, J., & Geiger, A. (2021). Projected gans converge faster. In Advances in neural information processing systems (NeurIPS)."},{"key":"1673_CR46","doi-asserted-by":"crossref","unstructured":"Sch\u00f6nfeld, E., Schiele, B., & Khoreva, A. (2020). A u-net based discriminator for generative adversarial networks. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR42600.2020.00823"},{"key":"1673_CR47","unstructured":"Sch\u00f6nfeld, E., Sushko, V., Zhang, D., Gall, J., Schiele, B., & Khoreva, A. (2021). You only need adversarial supervision for semantic image synthesis. In International conference on learning representations (ICLR)."},{"key":"1673_CR48","unstructured":"Simonyan, K., & Zisserman, A. (2015). Very deep convolutional networks for large-scale image recognition. In International conference on learning representations (ICLR)."},{"key":"1673_CR49","doi-asserted-by":"crossref","unstructured":"Souly, N., Spampinato, C., & Shah, M. (2017). Semi supervised semantic segmentation using generative adversarial network. In International conference on computer vision (ICCV).","DOI":"10.1109\/ICCV.2017.606"},{"key":"1673_CR50","doi-asserted-by":"crossref","unstructured":"Sudre, C. H., Li, W., Vercauteren, T., Ourselin, S., & Cardoso, M. J. (2017). Generalised dice overlap as a deep learning loss function for highly unbalanced segmentations. In Deep learning in medical image analysis and multimodal learning for clinical decision support.","DOI":"10.1007\/978-3-319-67558-9_28"},{"key":"1673_CR51","unstructured":"Tan, Z., Chen, D., Chu, Q., Chai, M., Liao, J., He, M., Yuan, L., & Yu, N. (2020). Rethinking spatially-adaptive normalization. arXiv:2004.02867."},{"key":"1673_CR52","doi-asserted-by":"crossref","unstructured":"Tang, H., Bai, S., & Sebe, N. (2020a). Dual attention gans for semantic image synthesis. In ACM international conference on multimedia.","DOI":"10.1145\/3394171.3416270"},{"key":"1673_CR53","doi-asserted-by":"crossref","unstructured":"Tang, H., Qi, X., Xu, D., Torr, P. H., & Sebe, N. (2020b). Edge guided gans with semantic preserving for semantic image synthesis. arXiv:2003.13898.","DOI":"10.1145\/3394171.3416270"},{"key":"1673_CR54","doi-asserted-by":"crossref","unstructured":"Tang, H., Xu, D., Yan, Y., Torr, PH. ., & Sebe, N. (2020c). Local class-specific and global image-level generative adversarial networks for semantic-guided scene generation. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR42600.2020.00789"},{"key":"1673_CR55","doi-asserted-by":"crossref","unstructured":"Wang, J., Zhang, W., Zang, Y., Cao, Y., Pang, J., Gong, T., Chen, K., Liu, Z., Loy, C. C., & Lin, D. (2021a). Seesaw loss for long-tailed instance segmentation. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR46437.2021.00957"},{"key":"1673_CR56","doi-asserted-by":"crossref","unstructured":"Wang, T. C., Liu, M. Y., Zhu, J. Y., Tao, A., Kautz, J., & Catanzaro, B. (2018). High-resolution image synthesis and semantic manipulation with conditional GANs. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2018.00917"},{"key":"1673_CR57","doi-asserted-by":"crossref","unstructured":"Wang, Y., Qi, L., Chen, Y. C., Zhang, X., & Jia, J. (2021b). Image synthesis via semantic composition. In ICCV.","DOI":"10.1109\/ICCV48922.2021.01349"},{"key":"1673_CR58","doi-asserted-by":"crossref","unstructured":"Wang, Z., Simoncelli, E. P., & Bovik, A. C. (2003). Multiscale structural similarity for image quality assessment. In Asilomar conference on signals, systems & computers.","DOI":"10.1109\/ACSSC.2003.1292216"},{"key":"1673_CR59","doi-asserted-by":"crossref","unstructured":"Xiao, T., Liu, Y., Zhou, B., Jiang, Y., & Sun, J. (2018). Unified perceptual parsing for scene understanding. In European Conference on Computer Vision (ECCV).","DOI":"10.1007\/978-3-030-01228-1_26"},{"key":"1673_CR60","unstructured":"Yaz, Y., Foo, C. S., Winkler, S., Yap, K. H., Piliouras, G., & Chandrasekhar, V. (2018). The unusual effectiveness of averaging in gan training. In International conference on learning representations (ICLR)."},{"key":"1673_CR61","doi-asserted-by":"crossref","unstructured":"Yu, F., Koltun, V., & Funkhouser, T. (2017) Dilated residual networks. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2017.75"},{"key":"1673_CR62","doi-asserted-by":"crossref","unstructured":"Yun, S,, Han, D., Oh, S. J., Chun, S., Choe, J., & Yoo, Y. (2019). Cutmix: Regularization strategy to train strong classifiers with localizable features. In International conference on computer vision (ICCV).","DOI":"10.1109\/ICCV.2019.00612"},{"key":"1673_CR63","unstructured":"Zhang, D., & Khoreva, A. (2019). PA-GAN: Improving GAN training by progressive augmentation. In Advances in neural information processing systems (NeurIPS)."},{"key":"1673_CR64","doi-asserted-by":"crossref","unstructured":"Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., & Metaxas, D. N. (2018a). StackGAN++: Realistic image synthesis with stacked generative adversarial networks. Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 41, 1947\u20131962.","DOI":"10.1109\/TPAMI.2018.2856256"},{"key":"1673_CR65","unstructured":"Zhang, H., Wu, C., Zhang, Z., Zhu, Y., Zhang, Z., Lin, H., Sun, Y., He, T., Mueller, J., Manmatha, R., Li, M., & Smola, A. (2020). Resnest: Split-attention networks. arXiv:2004.08955."},{"key":"1673_CR66","doi-asserted-by":"crossref","unstructured":"Zhang, H., Koh, J. Y., Baldridge, J., Lee, H., & Yang, Y. (2021). Cross-modal contrastive learning for text-to-image generation. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR46437.2021.00089"},{"key":"1673_CR67","doi-asserted-by":"crossref","unstructured":"Zhang, R., Isola, P., Efros, A. A., Shechtman, E., & Wang, O. (2018b). The unreasonable effectiveness of deep features as a perceptual metric. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2018.00068"},{"key":"1673_CR68","doi-asserted-by":"crossref","unstructured":"Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., & Torralba, A. (2017). Scene parsing through ade20k dataset. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2017.544"},{"key":"1673_CR69","doi-asserted-by":"crossref","unstructured":"Zhu, Z., Xu, Z., You, A., & Bai, X. (2020). Semantically multi-modal image synthesis. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR42600.2020.00551"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-022-01673-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-022-01673-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-022-01673-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,10,26]],"date-time":"2022-10-26T08:12:54Z","timestamp":1666771974000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-022-01673-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,17]]},"references-count":69,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["1673"],"URL":"https:\/\/doi.org\/10.1007\/s11263-022-01673-x","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,17]]},"assertion":[{"value":"6 May 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 August 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 September 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}